Display device and speech interaction method for display device

By integrating long text understanding and multi-intent recognition technologies with large models, the display device automatically recommends relevant objects and services when it recognizes keywords related to holidays or anniversaries. This solves the problem of requiring multiple interactions in traditional voice interaction methods and achieves more efficient intelligent voice interaction.

WO2026066605A1PCT designated stage Publication Date: 2026-04-02HISENSE VISUAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Traditional smart TVs' voice interaction methods require users to interact multiple times when handling complex tasks, which is inefficient and cannot proactively uncover hidden commands, especially in providing a seamless service for holiday greetings.

Method used

By integrating long text understanding and multi-intent recognition technologies from large models, display devices can automatically recommend relevant objects and provide services when they recognize keywords such as holidays, birthdays, or anniversaries, simplifying the user interaction process. For example, when recognizing "tomorrow is my birthday," it can automatically recommend birthday cakes and generate ambient wallpapers.

Benefits of technology

It improves the intelligence level of voice interaction, simplifies user operation, and can automatically provide relevant recommendations and reminders for holidays, birthdays or anniversaries based on simple voice content, thereby improving interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025110950_02042026_PF_FP_ABST
    Figure CN2025110950_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of display devices, and particularly relates to a display device and a speech interaction method for a display device. The display device may comprise a display, one or more external device interfaces, and at least one processor, wherein the at least one processor may be configured to execute instructions, such that the display device executes: in response to a control instruction for entering a home page, the display displaying a home page interface; in response to a speech interaction instruction, the display displaying text information of input speech while displaying the home page interface; when it is recognized that the text information includes a keyword including one of holiday, birthday and anniversary, displaying at least one control in a first display area in the home page interface; the control displaying integrated recommendation data of at least one recommended object associated with the keyword; and in response to an instruction for triggering a detail page link, controlling the display to display a detail page of the recommended object to which the detail page link belongs. Therefore, the intelligence level of speech interaction of display devices is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Display device and voice interaction method for display device

[0001] Cross-reference to Related Applications

[0002] This application claims priority to Chinese Patent Application No. 202411391925.3, filed on September 30, 2024, and Chinese Patent Application No. 202411389528.2, filed on September 30, 2024, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0003] The present application relates to the technical field of display devices, and in particular to a display device and a voice interaction method for the display device. BACKGROUND

[0004] Currently, display devices commonly used by people include televisions, mobile phones, computers, etc. Among them, televisions are often the best choice for people to watch videos and play electronic games due to their ultra-large screen characteristics. Traditional televisions are integrated with voice interaction functions, and through natural language understanding processing, smart televisions and other display devices can perform voice interaction with users. For example, a smart television can recognize the user's voice for natural language understanding, obtain the user's interaction instruction, and execute corresponding operations according to the user's interaction instruction.

[0005] In related technologies, a smart television recognizes different interaction instructions through voice interaction to perform different operations. For complex interaction tasks, the user usually needs to perform voice interaction multiple times to issue single interaction instructions, and the interaction process is cumbersome and the service is not coherent. The traditional voice interaction method lacks the ability to understand the implied requirements in the user's voice, especially in the aspect of holiday greetings, and cannot actively further dig out associated implied instructions, resulting in the technical problem of low voice interaction efficiency. SUMMARY

[0006] According to some embodiments of the present application, a display device can include a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, the communication protocol including at least a Bluetooth protocol; and at least one processor connected with the display, the memory, and the one or more external device interfaces, and configured to execute the computer programs or instructions to cause the display device to: in response to a control instruction to enter a home page, display a home page interface on the display, the home page interface including at least one media information; in response to a voice interaction instruction, display text information of input voice on the display while displaying the home page interface on the display; in a case where it is identified that the text information contains a keyword of one of a holiday, a birthday, and an anniversary, display at least one control in a first display area of the home page interface; the control can display integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data including at least detail integrated information of the recommended object and a detail page link of the recommended object; and in response to an instruction to trigger the detail page link, control the display to display a detail page of the recommended object to which the detail page link belongs.

[0007] According to some embodiments of the present application, a display device can include a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, the communication protocol including at least a Bluetooth protocol; and at least one processor connected with the display, the memory, and the one or more external device interfaces, and configured to execute the computer programs or instructions to cause the display device to: in response to a control instruction to enter a home page, display a home page interface on the display, the home page interface including at least one media information; in response to a voice interaction instruction, display text information of input voice on the display while displaying the home page interface on the display; in a case where it is identified that the text information contains a keyword of one of a holiday, a birthday, and an anniversary, display at least one control in a first display area of the home page interface; the control can display integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data including at least detail integrated information of the recommended object and a detail page link of the recommended object; and in response to an instruction to trigger the detail page link, control the display to display a detail page of the recommended object to which the detail page link belongs.

[0008] According to some embodiments of the present application, a display device is provided, which can include a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, the communication protocol including at least a Bluetooth protocol; and at least one processor connected with the display, the memory, and the one or more external device interfaces, and configured to execute the computer programs or instructions to cause the display device to: in response to a voice interaction instruction, display an interactive dialogue interface while displaying a home page interface or a video playing interface, the display displaying text information of an input voice; in response to identifying that the text information contains a keyword of a holiday, a birthday, or an anniversary, display at least one control in a third display area of the interactive dialogue interface; display, by the control, integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data including at least detail integrated information of the recommended object and a detail page link of the recommended object; and in response to an instruction triggering the detail page link, control the display to display a detail page of the recommended object to which the detail page link belongs.

[0009] According to some embodiments of the present application, a voice interaction method for a display device is provided, the display device can include a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, the communication protocol including at least a Bluetooth protocol; and at least one processor connected with the display, the memory, and the one or more external device interfaces, and configured to execute the computer programs or instructions to control the display device; the method can include: in response to a control instruction for entering a home page, the display displays a home page interface, the home page interface including at least one media information; in response to a voice interaction instruction, the display displays text information of an input voice while displaying the home page interface; in response to identifying that the text information contains a keyword of a holiday, a birthday, or an anniversary, at least one control is displayed in a first display area of the home page interface; the control can display integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data including at least detail integrated information of the recommended object and a detail page link of the recommended object; and in response to an instruction triggering the detail page link, the display displays a detail page of the recommended object to which the detail page link belongs.

[0010] According to some embodiments of the present application, a voice interaction method for a display device is also provided. The display device can include a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, which can include at least a Bluetooth protocol; and at least one processor connected with the display, the memory, and the one or more external device interfaces, and configured to execute the computer programs or instructions to control the display device. The method can include: in response to a control instruction for a video play, the display displays a video play interface; in response to a voice interaction instruction, the display pauses a video play or reduces a play volume of the video play interface, and displays text information of an input voice; in a case where it is identified that the text information contains a keyword of one of a holiday, a birthday, and an anniversary, at least one control is displayed in a second display area in the video play interface; the control can display integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data including at least detail integrated information of the recommended object and a detail page link of the recommended object; and in response to an instruction for triggering the detail page link, the display displays a detail page of the recommended object to which the detail page link belongs.

[0011] According to some embodiments of the present application, a voice interaction method for a display device is also provided. The display device can include a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, which can include at least a Bluetooth protocol; and at least one processor connected with the display, the memory, and the one or more external device interfaces, and configured to execute the computer programs or instructions to control the display device. The method can include: in response to a control instruction for a video play, the display displays a video play interface; in response to a voice interaction instruction, the display pauses a video play or reduces a play volume of the video play interface, and displays text information of an input voice; in a case where it is identified that the text information contains a keyword of one of a holiday, a birthday, and an anniversary, at least one control is displayed in a second display area in the video play interface; the control can display integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data including at least detail integrated information of the recommended object and a detail page link of the recommended object; and in response to an instruction for triggering the detail page link, the display displays a detail page of the recommended object to which the detail page link belongs. Attached Figure Description

[0012] Figure 1 is a schematic diagram of an operation scenario between a display device and a control device according to some embodiments of this application;

[0013] Figure 2 is a schematic diagram of the hardware configuration of a display device provided according to some embodiments of this application;

[0014] Figure 3 is a schematic diagram of the hardware configuration of a control device provided according to some embodiments of this application;

[0015] Figure 4 is a schematic diagram of the software configuration of a display device according to some embodiments of this application;

[0016] Figure 5 is a flowchart illustrating a voice interaction method for a display device according to some embodiments of this application;

[0017] Figure 6a is a schematic diagram of a homepage interface provided according to some embodiments of this application;

[0018] Figure 6b is a schematic diagram of a details page provided according to some embodiments of this application;

[0019] Figure 7a is a schematic diagram of a module architecture provided according to some embodiments of this application;

[0020] Figure 7b is a schematic diagram of a text information processing flow provided according to some embodiments of this application;

[0021] Figure 7c is a schematic diagram of an intelligent agent processing flow provided according to some embodiments of this application;

[0022] Figure 8a is a flowchart illustrating another voice interaction method for a display device according to some other embodiments of this application;

[0023] Figure 8b is a schematic diagram of a video playback interface provided according to some embodiments of this application;

[0024] Figure 8c is a schematic diagram of a process for reducing playback volume according to some embodiments of this application;

[0025] Figure 9a is a flowchart illustrating another voice interaction method for a display device according to some embodiments of this application;

[0026] Figure 9b is a schematic diagram of a homepage interface or video playback interface including an interactive dialog interface provided according to some embodiments of this application;

[0027] Figure 10 is a signaling interaction diagram of voice interaction according to some embodiments of this application;

[0028] FIG. 11 is a flow diagram of another method of voice interaction for a display device according to some embodiments of the present disclosure;

[0029] FIG. 12 is a diagram of a play interface of a target video according to some embodiments of the present disclosure;

[0030] FIG. 13 is a diagram of text corresponding to an interactive voice according to some embodiments of the present disclosure;

[0031] FIG. 14 is a diagram of an interactive result item according to some embodiments of the present disclosure;

[0032] FIG. 15 is a diagram of text information corresponding to a target audio according to some embodiments of the present disclosure;

[0033] FIG. 16 is a diagram of text information corresponding to a target audio according to some embodiments of the present disclosure;

[0034] FIG. 17 is a diagram of a ticket reservation operation according to some embodiments of the present disclosure;

[0035] FIG. 18 is a diagram of a weather information detail viewing operation according to some embodiments of the present disclosure;

[0036] FIG. 19 is a diagram of a lodging reservation operation according to some embodiments of the present disclosure;

[0037] FIG. 20 is a diagram of a sliding view of various interactive result items according to some embodiments of the present disclosure;

[0038] FIG. 21 is a diagram of travel tool recommendation information according to some embodiments of the present disclosure;

[0039] FIG. 22 is a diagram of weather related reminder information according to some embodiments of the present disclosure;

[0040] FIG. 23 is a flow diagram of another method of voice interaction for a display device according to some embodiments of the present disclosure;

[0041] FIG. 24 is a module interaction diagram of another method of voice interaction for a display device according to some embodiments of the present disclosure;

[0042] FIG. 25 is a module interaction diagram of another method of voice interaction for a display device according to some embodiments of the present disclosure;

[0043] FIG. 26 is a module interaction diagram of another method of voice interaction for a display device according to some embodiments of the present disclosure;

[0044] FIG. 27 is a diagram of an interaction between a voice interaction application and an audio management component according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0045] The embodiments will be described in detail below with reference to examples thereof as illustrated in the accompanying drawings. In the following description, the same drawing reference numerals are used to denote elements having the same or similar functions in different drawings. The embodiments described in the following examples do not represent all of the embodiments consistent with the present application. Rather, they are merely examples of systems and methods consistent with some aspects of the present application as detailed in the appended claims. The following description in the examples is merely illustrative of some embodiments consistent with the present application and as such other variations and modifications of the systems and methods disclosed herein, can be made by those of ordinary skill in the art without departing from the scope of the present application. The following examples are illustrative only and are not intended to limit the scope of the present application.

[0046] It should be noted that the brief description of terms in the present application is only for the convenience of understanding the following described embodiments, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.

[0047] The terms "first", "second", "third", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar or like objects or entities, and do not necessarily mean a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances. The terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not necessarily limit to all components clearly listed, but can include other components not clearly listed or inherent to these products or devices. The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code capable of performing a function related to the element.

[0048] The display device in the embodiments of the present application refers to a device having the ability of picture display and data processing. For example, the display device can include but is not limited to a television, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.

[0049] FIG. 1 is a schematic diagram of an operating scenario between a display device and a control device according to some embodiments of the present application. As shown in FIG. 1, a user can operate the display device 200 through a touch operation, a mobile terminal 300 and a control device 100. For example, the control device 100 can be a remote controller, a stylus, a handle, etc.

[0050] The mobile terminal 300 can be used as a kind of control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a kind of communication device for establishing a communication connection with the display device 200 and performing data interaction.

[0051] In some embodiments, the mobile terminal 300 can install a software application with the display device 200, implement connection communication through a network communication protocol, and achieve the purpose of one-to-one control operation and data communication. The mobile terminal 300 can also transmit the display of audio and video content to the display device 200, and achieve the function of synchronous display.

[0052] As further shown in FIG. 1, the display device 200 also communicates data with the server 400 through various communication modes. The display device 200 can be allowed to communicate through a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0053] The display device 200 can provide a broadcast receiving television function, and can also provide an intelligent network television function with computer support, including but not limited to a network television, a smart television, an interactive personality TV (IPTV), and the like.

[0054] FIG. 2 is a hardware configuration block diagram of a display device according to some embodiments of the present application.

[0055] In some embodiments, the display device 200 can include at least one of a tuning demodulator 210, a communication device 220, a detector 230, a device interface 240, at least one processor 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0056] In some embodiments, the detector 230 can be used to collect signals of an external environment or interaction with the outside. For example, the detector 230 can include a light receiver, and can be used to collect a sensor of ambient light intensity; or the detector 230 can include an image collector such as a camera, and can be used to collect an external environment scene, a user attribute, or a user interaction gesture; or the detector 230 can include a sound collector such as a microphone, and can be used to receive external sound.

[0057] In some embodiments, the display 260 can include a display function component for presenting a picture, and a driving component for driving image display. The display 260 can be used to receive an image signal output from the at least one processor 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface, and a user control UI interface, etc.

[0058] In some embodiments, the communication device 220 is a component that can be used to communicate with the external device or the server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication manners. For example, when the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 containing WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 containing Bluetooth function.

[0059] The communication device 220 can make the display device 200 communicate with the external device or the server 400 through wireless or wired connection. The wired connection can connect the display device 200 with the external device through data line, interface, etc. The wireless connection can connect the display device 200 with the external device through wireless signal or wireless network. The display device 200 can directly establish connection relationship with the external device, or indirectly establish connection relationship through gateway, route, connection device, etc.

[0060] In some embodiments, the processor can include at least one of a central processor, a video processor, an audio processor, a graphics processor, a power supply processor, a first interface to an n-th interface for input / output.

[0061] In some embodiments, the at least one processor 250 can control the operation of the display device and respond to the user's operation through various software control programs stored on the memory. The at least one processor 250 can control the overall operation of the display device 200.

[0062] In some embodiments, the at least one processor 250 and the tuner demodulator 210 can be located in different split devices, i.e. the tuner demodulator 210 can also be in an external device of the main device where the at least one processor 250 is located, such as an external set-top box, etc.

[0063] In some embodiments, the user can input user commands through the graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the graphical user interface (GUI).

[0064] In some embodiments, the audio output device 270 can be a native loudspeaker of the display device 200, or an audio output device connected to the display device 200. For the audio output device connected to the display device 200, the display device 200 can also be provided with an external audio output terminal, and the audio output device can be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.

[0065] In some embodiments, the user input interface 280 can be configured to receive instructions from a user input.

[0066] FIG. 3 is a hardware configuration block diagram of a control device according to some embodiments of the present application. As shown in FIG. 3, the control device 100 can include a control apparatus 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0067] In some embodiments, the control device 100 can be configured to control the display device 200 and receive user input operation instructions and convert the operation instructions into instructions recognizable and responsive by the display device 200, thereby serving as an intermediary between the user and the display device 200.

[0068] In some embodiments, the control device 100 can be a smart device. For example, the control device 100 can install various applications for controlling the display device 200 according to user needs.

[0069] In some embodiments, as shown in FIG. 1, the mobile terminal 300 or other smart electronic device can perform similar functions of the control device 100 after installing the application for controlling the display device 200.

[0070] In some embodiments, the control apparatus 110 can include a processor 112 and a RAM 113 and a ROM 114, a communication interface 130, and a communication bus. The control apparatus 110 can be configured to control the operation of the control device 100 and the communication and cooperation between the internal components, as well as the data processing functions of the external and internal components.

[0071] In some embodiments, the communication interface 130 can be controlled by the control apparatus 110 to communicate control signals and data signals with the display device 200. The communication interface 130 can include at least one of a WiFi chip 131, a Bluetooth module 132, a Near Field Communication (NFC) module 133, and other near field communication modules.

[0072] In some embodiments, the user input / output interface 140 can include at least one of a microphone 141, a touchpad 142, a sensor 143, a key 144, and other input interfaces.

[0073] In some embodiments, the control device 100 can include at least one of the communication interface 130 and the user input / output interface 140. The control device 100 can be configured with a communication interface 130, such as a WiFi, Bluetooth, NFC, etc. module, which can encode user input instructions through a WiFi protocol, or a Bluetooth protocol, or an NFC protocol and send them to the display device 200.

[0074] In some embodiments, the memory 190 can be configured to store various operation programs, data and applications for driving and controlling the control device 100 under the control of the control device. The memory 190 can store various control signal instructions input by the user.

[0075] The power supply 180 can be configured to provide operating power support for the elements of the control device 100 under the control of the control device.

[0076] In order to perform user interaction, in some embodiments, the display device 200 can run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system can provide a user interface to allow the user to interact with the display device 200 and support the running of various application programs.

[0077] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0078] The operating system can be divided into different modules or levels according to the functions implemented, for example, as shown in FIG. 4, in some embodiments, the system is divided into four layers, from top to bottom, which can be an application (Applications) layer (referred to as “application layer”), an application framework (Application Framework) layer (referred to as “framework layer”), a system library layer and a kernel layer.

[0079] In some embodiments, the application layer can be configured to provide services and interfaces for application programs, so that the display device 200 can run application programs and interact with users based on the application programs. At least one application program can be run in the application layer, which can be a window (Window) program, a system setting program or a clock program provided with the operating system; or an application program developed by a third-party developer. In specific implementation, the application programs in the application layer are not limited to the above examples.

[0080] The framework layer can provide application programming interfaces (Application Programming Interface, API) and programming frameworks for application programs. The application framework layer can include some pre-defined functions. The application framework layer is equivalent to a processing center that decides which application program in the application layer to act. The application program can access the resources in the system and obtain the services of the system through the API interface in the execution.

[0081] As shown in FIG. 4, the application framework layer in some embodiments of the present application can include a view system, managers, content providers, etc., wherein the view system can be designed and implemented to interface and interact with the application, and the view system can include lists, grids, text boxes, buttons, etc. The managers can include at least one of the following modules: an activity manager can be used to interact with all activities running in the system; a location manager can be used to provide access to system location services for system services or applications; a package manager can be used to retrieve various information related to application packages currently installed on the device; a notification manager can be used to control the display and clearing of notification messages; and a window manager can be used to manage icons, windows, toolbars, wallpapers and desktop components on the user interface.

[0082] In some embodiments, the activity manager can be used to manage the life cycle of various applications and general navigation back function, such as controlling the exit, opening, back, etc. of the application. The window manager can be used to manage all window programs, such as obtaining the size of the display screen, determining whether there is a status bar, locking the screen, intercepting the screen, controlling the display window change, such as reducing the display window, shaking the display, twisting the display, etc.

[0083] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as C / C++ instruction library, to realize the functions of the framework layer.

[0084] In some embodiments, the kernel layer can be a functional layer between the hardware and software of the display device 200. The kernel layer can realize functions such as hardware abstraction, multitasking, memory management, etc. For example, as shown in FIG. 4, the kernel layer can be configured with hardware drivers, and the kernel layer contains at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, Universal Serial Bus (USB) driver, High Definition Multimedia Interface (HDMI) driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power supply driver, etc.

[0085] It should be noted that the above examples are only simple divisions of operating system functions, and do not constitute limitations on the specific operating system forms of the display device in the embodiments of the present application. Depending on the functions of the display device, the types of the operating system, and other factors, the number and specific types of the levels included in the operating system can take other forms.

[0086] According to some embodiments of the present application, a display device can include a display configured to display content from a broadcast system or a network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, the communication protocol including at least a Bluetooth protocol; and at least one processor connected with the display, the memory, and the one or more external device interfaces, and configured to execute the computer programs or instructions to cause the display device to perform steps as shown in FIG. 5, to implement a voice interaction scheme for the display device.

[0087] At step S501, in response to a control instruction to enter a home page, the display displays a home page interface, the home page interface including at least one media asset information.

[0088] At step S502, in response to a voice interaction instruction, the display displays text information of the input voice while displaying the home page interface.

[0089] At step S503, in the case where the text information is identified to contain one of a holiday, a birthday, and an anniversary keyword, at least one control is displayed in a first display area of the home page interface.

[0090] At step S504, the control displays integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data including at least detail integration information of the recommended object and a detail page link of the recommended object.

[0091] At step S505, in response to an instruction to trigger the detail page link, the display displays a detail page of a recommended object to which the detail page link belongs.

[0092] In some embodiments, the display device can, in response to a control instruction to enter a home page, control the display to display a home page interface, the home page interface including at least one media asset information, the media asset information including a media asset name of media asset data, the media asset data including video data and audio data, the media asset name including a film name and a song name, and the media asset information including character information of the media asset data, such as a performer, a singer, and an author.

[0093] The display device can respond to the voice interaction instruction, as shown in FIG. 6a, taking the birthday scene as an example. While displaying the home page interface on the display, the text information (such as "Tomorrow is my birthday") input by voice can be displayed. Then, in the case where it is recognized that the text information contains a birthday keyword, at least one control can be displayed in the first display area in the home page interface. The control can display integrated recommendation data of at least one recommended object (such as recommended object A, recommended object B, and recommended object C) associated with the keyword. The integrated recommendation data can at least include detail integration information of the recommended object and a detail page link of the recommended object. The recommended object can be a merchant providing a birthday cake. The detail integration information can be information extracted from the content in the detail page of the merchant. The detail page link can be used to jump to the detail page of the merchant. In turn, by responding to the instruction triggering the detail page link, the display device can control the display to display the detail page of the recommended object to which the detail page link belongs, as shown in FIG. 6b.

[0094] In some embodiments, the display device can analyze the text information of the input voice based on the natural language processing module. In the case where the text information does not contain a keyword of a holiday, a birthday, or an anniversary, but the semantics of the text information has an intent of a holiday, a birthday, or an anniversary, such as the input voice "I was born tomorrow in xx years ago", by performing natural language understanding processing on the text information, in the case where it is recognized that the text information contains an intent of a holiday, a birthday, or an anniversary, the subsequent processing flow can be triggered to display integrated recommendation data of at least one recommended object associated with the intent.

[0095] In the application scenario of holiday greetings, the traditional voice interaction method requires the user to interact multiple times to give a single instruction, resulting in incoherent interaction. For example, when the user says "Tomorrow is my birthday", since the traditional voice interaction method follows the user's command to query instructions and does not have the ability to further mine implicit instructions, it gives a chat greeting reply of "Happy Birthday", and cannot actively push and remind other associated information. Compared with the traditional voice interaction method, in some embodiments of the present application, the long text understanding and multi-intent recognition technology of the integrated large model are used for interactive life service processing. For a user's simple input voice interaction content, the complex task instructions implied in the user's voice can be accurately parsed, and the related service request can be automatically recognized and executed, such as the user only needs to say "Tomorrow is my birthday". By performing long text understanding on the input voice, the key information (such as holiday information) in the input voice can be extracted, and the corresponding gift for the holiday can be automatically recommended, and the services of booking a birthday cake, generating ambient wallpaper and music can be provided.

[0096] In one example, for a festival scenario, taking the Spring Festival as an example, while the display displays the home page interface, the text information (such as "the day after tomorrow is the Spring Festival") of the input voice can be displayed, and then at least one control can be displayed in the first display area in the home page interface in a case where it is recognized that the text information contains a festival keyword, the control can display integrated recommendation data of at least one recommended object associated with the keyword, and the integrated recommendation data can at least include detail integration information of the recommended object and a detail page link of the recommended object. For example, the recommended object can be a Spring Festival activity site, the detail integration information can be information extracted from the content in the detail page of the activity site, and the detail page link can be used to jump to the detail page of the activity site, so that the display device can control the display to display the detail page of the recommended object to which the detail page link belongs by responding to an instruction triggering the detail page link. The festival scenario can also include Chinese traditional festivals such as Dragon Boat Festival and Mid-Autumn Festival, or Western festivals such as Christmas and April Fool's Day, and generally recognized festivals such as New Year's Eve, which are not specifically limited in this embodiment. Different recommended objects and their recommendation data can be associated for different festivals.

[0097] In another example, for an anniversary scenario, taking the wedding anniversary as an example, while the display displays the home page interface, the text information (such as "the day after tomorrow is the 10th wedding anniversary") of the input voice can be displayed, and then at least one control can be displayed in the first display area in the home page interface in a case where it is recognized that the text information contains an anniversary keyword, the control can display integrated recommendation data of at least one recommended object associated with the keyword, and the integrated recommendation data can at least include detail integration information of the recommended object and a detail page link of the recommended object. For example, the recommended object can be a gift, the detail integration information can be information extracted from the content in the detail page of the gift, and the detail page link can be used to jump to the detail page of the gift, so that the display device can control the display to display the detail page of the recommended object to which the detail page link belongs by responding to an instruction triggering the detail page link. The anniversary scenario can also include various personalized anniversaries set by the user or specific statutory anniversaries, which are not specifically limited in this embodiment. Different recommended objects and their recommendation data can be associated for different anniversaries.

[0098] The scheme of the embodiment shows that, after the display device responds to the voice interaction instruction in the case of displaying the home page interface, at least one control can be displayed in the display area in the home page interface in the case that the text information of the input voice is identified to contain one of the keywords of the festival, birthday and memorial day, and the integrated recommendation data of at least one recommended object associated with the keyword can be displayed through the control. The integrated recommendation data can at least include the detail integration information of the recommended object and the detail page link of the recommended object. Then, the display device can respond to the instruction of triggering the detail page link to control the display to display the detail page of the recommended object to which the detail page link belongs, thereby improving the voice interaction intelligent level of the display device, enabling it to process complex voice interaction requirements containing one of the intentions of the festival, birthday and memorial day, simplifying the voice interaction processing flow, and providing a recommended plan for one of the intentions of the festival, birthday and memorial day based on the simple voice content of the user without the user repeatedly interacting to output an instruction, thereby improving the voice interaction efficiency of the display device.

[0099] In some embodiments, the at least one processor can be further configured to execute instructions to cause the display device to perform: in the case that the text information is identified to contain one of the keywords of the festival, birthday and memorial day, the reminder information of the keyword can be displayed in the home page interface; and the reminder information can be used to perform a reminder operation at a time corresponding to the festival or birthday or memorial day.

[0100] In the embodiment, as shown in FIG. 6a, taking the birthday scenario as an example, in the case that the text information is identified to contain the birthday keyword, the reminder information of the keyword can be displayed in the home page interface, and the reminder information can be used to perform a reminder operation at a time corresponding to the birthday. The task of the reminder can also be pushed from the television to the user's mobile phone end through the system platform.

[0101] For example, taking the birthday scenario as an example, for the input voice "Tomorrow is my birthday", in the case that the system time reaches tomorrow, the display device can be triggered to perform the operation of playing a birthday song according to the reminder time (such as celebrating the birthday at 8 pm) set by the reminder information, and the display device displays a birthday celebration interface and plays a birthday song.

[0102] For the festival scenario, taking the Spring Festival as an example, for the input voice "The day after tomorrow is the Spring Festival", in the case that the system time reaches the day after tomorrow, the display device can be triggered to perform the real-time information broadcast of the Spring Festival activity place according to the reminder time (such as reminding to celebrate the Spring Festival at 12 pm) set by the reminder information, and the display device displays a Spring Festival celebration interface and displays real-time information of the place.

[0103] For the anniversary scene, taking the wedding anniversary as an example, for the input voice "the day after tomorrow is the 10th wedding anniversary", in the case that the system time reaches the day after tomorrow, the reminder information corresponding to the reminder time (such as reminding the wedding anniversary at 10 o'clock in the morning) can be set, triggering the display device to perform the specific wallpaper replacement operation of the wedding anniversary, and the display device displays the wedding anniversary celebration interface and displays the corresponding specific wallpaper.

[0104] The scheme of the embodiment can display the reminder information of the keyword in the home interface in the case that the text information of the input voice contains one of the keywords of the festival, the birthday, and the anniversary, so as to perform the reminding operation at the corresponding time of the festival or the birthday or the anniversary, thereby the display device can provide the reminding processing of one of the festival, the birthday, and the anniversary based on the simple voice content of the user, and simplify the voice interaction processing flow.

[0105] In some embodiments, the at least one processor can be further configured to execute instructions to cause the display device to perform: the text information, the current date of the system in the display device, and the positioning information of the display device can be input into the multi-intent analysis model to obtain a plurality of intent-related task information; each intent-related task information can include a predicted intent based on an intent, and task arrangement information of an interactive task related to the predicted intent; the predicted intents can be combined by performing intent understanding on each predicted intent, and the predicted intents having an intent connection relationship are obtained at least one target intent; the service processing result of different service functions can be obtained based on the task arrangement information of the interactive task related to each target intent; the service processing result of different service functions can at least include integrated recommendation data of at least one recommended object associated with the keyword.

[0106] In the embodiment, as shown in FIG. 7a, the model architecture can include a large model multi-intent analysis module, a multi-intent decision processing module, and a multi-intent service module. By receiving the voice input of the user, the input voice can be converted into text information through voice recognition processing, then the multi-intent analysis model can be used to identify and analyze the implicit intent (such as task requirements) in the text information, and then the related services can be automatically triggered according to the analyzed implicit intent, the service content obtained by calling the services can be summarized, and the summarized service content can be displayed to the user through the display interface, and further operation options can be provided.

[0107] The scheme of the embodiment can input the text information, the current date of the system in the display device, and the positioning information of the display device into the multi-intention analysis model for data processing. Based on the interactive task associated with the keyword of one of the festivals, birthdays, and anniversaries, the service function corresponding to each interactive task can be called, and the service processing result of different service functions can be obtained, including the integrated recommendation data of at least one recommended object associated with the keyword. Therefore, by using the multi-intention analysis model to identify and analyze the implicit intention associated with the keyword of one of the festivals, birthdays, and anniversaries in the text information, the related service can be automatically triggered for calling according to the analyzed implicit intention, so as to obtain the service content provided to the user.

[0108] In some embodiments, the at least one processor performs inputting the text information, the current date of the system in the display device, and the positioning information of the display device into the multi-intention analysis model to obtain the multi-intention associated task information, which can be configured to execute instructions to cause the display device to perform: the service providing information of the display device can be obtained; the service providing information can be used to indicate the service functions possessed by the display device; based on the keyword contained in the text information and the service functions possessed by the display device, a plurality of predicted intentions extended based on the keyword can be determined; the multi-intention associated task information can be determined by combining the specified date of the keyword contained in the text information, the current date of the system in the display device, the positioning information of the display device, and the configuration rules of the interactive tasks associated with each predicted intention.

[0109] In the embodiment, the multi-intention analysis module based on a large model can include a plurality of multi-intention sentence rewriting modules (such as the Mx multi-intention sentence rewriting module in FIG. 7a), which can combine the text information of the input voice with the positioning information of the display device associated with the user, the current date of the system in the display device, and other information to further mine the implicit intention, so as to obtain the hidden expression meaning based on the text information of the input voice, and expand and rewrite the instruction of the user input voice.

[0110] For example, taking the user input voice "Tomorrow is my birthday" as an example, the multi-intention analysis module based on a large model can process as follows:

[0111] 1. The text information obtained by converting the user input voice (such as S701: user input in FIG. 7b), the positioning information of the display device, and the current date of the system in the display device (such as S702: extracting required information in FIG. 7b) are input into the multi-intention analysis model. The multi-intention analysis module based on a large model in the multi-intention analysis model can obtain the text information, the information associated with the user (such as regional information), and the system date information.

[0112] 2. Combining the specified date of the keywords contained in the text information, the current date of the system on the display device, and the location information of the display device, key information extraction and information reasoning are performed. For example, if the text information is "Tomorrow is my birthday", the user's location is A, and the current date of the system is September 2, 2024, and it is determined that the display device has birthday-related service functions (different holidays can be set to different types, such as Valentine's Day and Qixi Festival being one type, Spring Festival being another type, and birthday being another type), the implicit intent can be mined based on the specified date "tomorrow" contained in the text information, the keyword "birthday" contained in the text information, the location information of the display device "A", and the current date of the system on the display device "September 2, 2024" (as shown in S703 in Figure 7b: assembly prompt words). The model can reason out the extended predicted intent, which may include, but is not limited to: I want to see birthday gifts, I want to find a cake shop near A, set a birthday reminder for September 3, use a cheerful wallpaper, or play a birthday song (as shown in S704 in Figure 7b: large model reasoning).

[0113] 3. Each predicted intent can trigger a corresponding interactive task. Based on the preset trigger conditions (i.e., task orchestration information) of different interactive tasks, tasks can be selected and orchestrated according to the system's current time. For example, a model prompt word can be designed, and task selection can be set to: if the holiday is not on the current day, do not play music; if the holiday is on the current day, do not show options to buy gifts and cakes. Taking "Tomorrow is my birthday" as an example, since it is tomorrow's birthday, task orchestration determines that playing a birthday song is not currently necessary.

[0114] 4. Based on the large model multi-intent parsing module, it can output rewritten intent-related task information (as shown in S705 in Figure 7b: output structured data), such as finding a cake shop near location A, setting a birthday reminder for September 3, or using a cheerful wallpaper.

[0115] As shown in Figure 7a, the multi-intent decision processing module can further process the structured data output by the multi-intent parsing module based on the large model. For example, for each intent-related task information in the structured data, a single intent processing module can be used to process it separately.

[0116] As shown in the task arrangement and multi-intent understanding processing flow of FIG. 7c, for example, the structured data output above is input into an intelligent agent (S801), and the intelligent agent can decompose the structured data through task decomposition processing (such as the large model task decomposition in FIG. 7c) to obtain short texts that can support semantic understanding, such as short text 1: find a cake shop near A, short text 2: set a birthday reminder on September 3, and short text 3: use a cheerful atmosphere wallpaper, which are task information of different types of tasks (such as search task, reminder task, and atmosphere construction task in FIG. 7c) (S802). Then, based on the single-intent processing module, the intent of each task information can be understood, such as using a natural language understanding system to understand each short text to determine the predicted intent represented by the short text (S803).

[0117] The processing results of each single-intent processing module can be subjected to multi-intent DST decision-making, which can complete the information of the decomposed short texts based on text understanding and reasoning capabilities. For the case where a text part of a task is decomposed into multiple short texts, such as decomposing the complete text of a task into the first half and the second half, the two tasks decomposed can be combined into a complete task through multi-intent DST decision-making after natural language understanding processing. Or for the case where a short text has ambiguity, such as a word that can be a song title or other meanings, the correct intent of the ambiguous short text can be determined through multi-intent DST decision-making combined with other decomposed short texts. As shown in FIG. 7a, the multi-intent decision-making module can combine the predicted intents having an intent connection relationship through intent understanding of each predicted intent, obtain at least one target intent, and output the task information of the target intent after further task arrangement.

[0118] According to the key word of one of the festival, birthday, and memorial day contained in the text information and the service function of the display device, the scheme of the embodiment can expand multiple predicted intents, and then the task arrangement can be performed for each predicted intent to obtain the task information associated with each predicted intent, thereby realizing the recognition and analysis of the implicit intent in the text information and obtaining the task requirement corresponding to the implicit intent to further provide the task-related service calling processing.

[0119] In some embodiments, the at least one processor can execute the task arrangement information associated with the interaction task of each target intention, call the service function corresponding to the interaction task, and obtain the service processing result of each service function, which can be configured to execute instructions to enable the display device to execute: the task arrangement information associated with the interaction task of each target intention can be used to call the service function corresponding to the interaction task, and the service processing result of each service function can be obtained; the service providing information of each service function can be adjusted according to the service type and the service display requirement, and the service processing result of each service function can be obtained.

[0120] In this embodiment, as shown in FIG. 7a, the multi-intention service module can perform single-service processing according to the task arrangement information of the interaction task associated with each target intention. After the implicit user intention is determined through the multi-intention decision processing module, the related business microservice can be called according to the task information corresponding to the implicit intention, such as the search tool, the reminder tool, and the atmosphere building tool in FIG. 7c. Then, the corresponding service providing information can be obtained based on different business microservices (S804). Through multi-service merging processing, the obtained service providing information can be summarized and generalized, and the generated multi-intention reply content can be optimized (S805). Then, the summarized and generalized information can be provided to the user, i.e., pushed to the terminal of the user (S806).

[0121] For example, the service providing information is: confirm, 8:00 am on September 3, 2021, to remind you of your birthday, and 50 related recommendations including "Sweet Cake House" have been searched. The service processing result can be obtained by summarizing and generalizing the service type (such as the reminder type) and the service display requirement (such as the reminder information display requirement), which is: set a birthday reminder on September 3, 2021, and wish you a happy celebration; at the same time, I have found 50 selected cake shop information including "Sweet Cake House" for your selection.

[0122] The scheme of this embodiment can obtain the service providing information of each service function by obtaining the task arrangement information to call the service function corresponding to the interaction task, and then adjust the service providing information according to the service type and the service display requirement. The service content obtained by calling the service can be summarized and generalized, so that the summarized and generalized service content can be displayed to the user, effectively improving the display effect of voice interaction.

[0123] In some embodiments, a display device is provided, which can include a display configured to display content from a broadcast system or network and / or a user interface, a memory configured to store computer programs or instructions, one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, which at least includes a Bluetooth protocol, and at least one processor connected with the display, the memory and the one or more external device interfaces, and configured to execute the computer programs or instructions to cause the display device to perform steps as shown in FIG. 8a.

[0124] At step S801, in response to a control instruction of video playing, the display displays a video playing interface.

[0125] At step S802, in response to a voice interaction instruction, the display pauses playing a video picture or reduces a playing volume of the video playing interface, and displays text information of the input voice.

[0126] At step S803, in a case where it is identified that the text information contains one of a festival, a birthday and an anniversary keyword, at least one control is displayed in a second display area in the video playing interface.

[0127] At step S804, the control displays integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data at least including detail integrated information of the recommended object and a detail page link of the recommended object.

[0128] At step S805, in response to an instruction of triggering the detail page link, the display displays a detail page of the recommended object to which the detail page link belongs.

[0129] As shown in FIG. 8b, taking a Spring Festival scenario as an example, in response to a voice interaction instruction, the display device can pause playing a video picture or reduce a playing volume of the video playing interface, display text information of the input voice (such as “the day after tomorrow is the Spring Festival”), and then in a case where it is identified that the text information contains a festival keyword, at least one control can be displayed in a second display area in the video playing interface, the control can display integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data can at least include detail integrated information of the recommended object and a detail page link of the recommended object, the recommended object can be a Spring Festival activity site, the detail integrated information can be information extracted according to content in a detail page of the activity site, and the detail page link can be used to jump to the detail page of the activity site, and then by responding to an instruction of triggering the detail page link, the display device can control the display to display a detail page of the recommended object to which the detail page link belongs.

[0130] As shown in FIG. 8c, the display device can detect the current scene of the user, and if it is a scene of watching a movie or listening to music, etc., the player can be paused without affecting the user's viewing, or when the voice is broadcast, the volume of the player can be reduced.

[0131] As a possible implementation, referring to FIG. 8c, the voice assistant detects the start of TTS, sends an audio focus use request to the audio manager (audio management component) to obtain the audio focus (S801'), the audio manager detects the application currently holding the audio focus as the foreground running application, i.e., the foreground application, the foreground application detects the focus change event, listens to the focus change event of the view (S802'), determines the loss of audio focus (S803'), pauses or reduces the volume (S804'); the voice assistant detects the start of TTS, sends an audio focus abandonment notification to the audio manager to release the audio focus (S805'), the audio manager sends an audio focus recovery notification to the foreground application to inform the foreground application of the focus state (S806'), the foreground application determines that the focus change event occurs, determines the gain of the audio focus (S807'), and restores or increases the volume (S808').

[0132] The scheme of the embodiment can display at least one control in the display area in the video playing interface when the display device responds to the voice interaction instruction in the case of displaying the video playing interface and identifies that the text information of the input voice contains one of the keywords of the festival, birthday, and anniversary, and the control can display integrated recommendation data of at least one recommended object associated with the keyword. The integrated recommendation data can at least include detail integrated information of the recommended object and a detail page link of the recommended object. In turn, the display device can respond to an instruction of triggering the detail page link to control the display to display a detail page of the recommended object to which the detail page link belongs, thereby improving the voice interaction intelligent level of the display device, enabling it to process complex voice interaction requirements containing one of the intentions of the festival, birthday, and anniversary, simplifying the voice interaction processing procedure, and providing a recommended plan of one of the intentions of the festival, birthday, and anniversary based on the simple voice content of the user without the user repeatedly interacting to output an instruction, improving the voice interaction efficiency of the display device, and reducing the impact on the user's video viewing by pausing the video playing picture or reducing the playing volume of the video playing interface.

[0133] In some embodiments, a display device is provided, which can include a display configured to display content from a broadcast system or network and / or a user interface, a memory configured to store computer programs or instructions, one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, which at least includes a Bluetooth protocol, and at least one processor connected with the display, the memory and the one or more external device interfaces, and configured to execute the computer programs or instructions to cause the display device to perform steps as shown in FIG. 9a.

[0134] At step S901, in a case where a home page interface is displayed or a video playing interface is displayed, the display displays an interactive dialogue interface in response to a voice interaction instruction; the interactive dialogue interface displays at least text information of input voice.

[0135] At step S902, in a case where it is identified that the text information contains one of keywords of a holiday, a birthday and an anniversary, at least one control is displayed in a third display area in the interactive dialogue interface.

[0136] At step S903, the control displays integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data at least including detail integrated information of the recommended object and a detail page link of the recommended object.

[0137] At step S904, in response to an instruction of triggering the detail page link, the display displays a detail page of the recommended object to which the detail page link belongs.

[0138] As shown in FIG. 9b, taking an anniversary scenario as an example, in a case where a home page interface is displayed or a video playing interface is displayed, the display device can display an interactive dialogue interface in response to a voice interaction instruction, the interactive dialogue interface can display at least text information of input voice (such as “the next day is the 10th anniversary of marriage, celebrate in the evening, plan for me”), and then in a case where it is identified that the text information contains an anniversary keyword, at least one control can be displayed in a third display area in the home page interface, the control can display integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data can at least include detail integrated information of the recommended object and a detail page link of the recommended object, the recommended object can be a gift, the detail integrated information can be information extracted from a content in a detail page of the gift, the detail page link can be used to jump to the detail page of the gift, and then in response to an instruction of triggering the detail page link, the display device can control the display to display a detail page of the recommended object to which the detail page link belongs.

[0139] The scheme of the embodiment shows that, after the display device responds to the voice interaction instruction in the case of displaying the home page interface or the video playing interface, at least one control can be displayed in the display area in the interactive dialogue interface jumped to in the case that the text information of the input voice is identified to contain one of the keywords of the festival, birthday and memorial day, and the integrated recommendation data of at least one recommended object associated with the keyword can be displayed through the control. The integrated recommendation data can at least include the detail integration information of the recommended object and the detail page link of the recommended object. Then, the display device can display the detail page of the recommended object to which the detail page link belongs in response to the instruction of triggering the detail page link, thereby improving the voice interaction intelligent level of the display device, enabling it to process complex voice interaction requirements containing one of the intentions of the festival, birthday and memorial day, simplifying the voice interaction processing procedure, providing the recommended plan of one of the intentions of the festival, birthday and memorial day based on the simple voice content of the user without repeated multiple interactions of the user to output the instruction, improving the voice interaction efficiency of the display device, and the display device can provide a specific interactive dialogue interface for voice interaction, which helps to improve the processing effect of voice interactive life service.

[0140] In some embodiments, as shown in FIG. 5, a voice interaction method for a display device is provided, which can include a display, which can be configured to display content from a broadcast system or a network and / or a user interface; one or more external device interfaces, which can be configured to communicate with one or more external devices according to a communication protocol, the communication protocol at least including a Bluetooth protocol; and at least one processor connected with the display and the one or more external device interfaces, and which can be configured to execute instructions to control the display device. The method can include but is not limited to the foregoing steps S501-S505. In this way, after responding to the voice interaction instruction in the case of displaying the home page interface, at least one control can be displayed in the display area in the home page interface in the case that the text information of the input voice is identified to contain one of the keywords of the festival, birthday and memorial day, and the integrated recommendation data of at least one recommended object associated with the keyword can be displayed through the control. The integrated recommendation data can at least include the detail integration information of the recommended object and the detail page link of the recommended object. Then, the detail page of the recommended object to which the detail page link belongs can be displayed in response to the instruction of triggering the detail page link, thereby improving the voice interaction intelligent level of the display device, enabling it to process complex voice interaction requirements containing one of the intentions of the festival, birthday and memorial day, simplifying the voice interaction processing procedure, providing the recommended plan of one of the intentions of the festival, birthday and memorial day based on the simple voice content of the user without repeated multiple interactions of the user to output the instruction, and improving the voice interaction efficiency of the display device.

[0141] In some embodiments, the method can further include: in a case where it is identified that the text information contains one of keywords of festivals, birthdays, and anniversaries, displaying reminder information of the keywords in the home interface; the reminder information can be used to perform a reminder operation at a time corresponding to the festival or the birthday or the anniversary.

[0142] In some embodiments, the method can further include: inputting the text information, the current date of the system in the display device, and the positioning information of the display device into a multi-intent analysis model to obtain a plurality of intent-related task information; each intent-related task information can include a predicted intent expanded based on the keywords and task arrangement information of an interactive task related to the predicted intent; at least one target intent can be obtained by performing intent understanding on each predicted intent and merging predicted intents having an intent connection relationship; a service function corresponding to each interactive task related to each target intent can be invoked based on the task arrangement information of the interactive task to obtain service processing results of different service functions; the service processing results of different service functions can at least include integrated recommendation data of at least one recommended object related to the keywords.

[0143] In some embodiments, the method can further include: obtaining service providing information of the display device; the service providing information can be used to indicate service functions possessed by the display device; a plurality of predicted intents expanded based on the keywords can be determined according to the keywords contained in the text information and the service functions possessed by the display device; task arrangement can be performed in combination with a specified date of the keywords contained in the text information, the current date of the system in the display device, the positioning information of the display device, and configuration rules of interactive tasks related to each predicted intent to determine a plurality of intent-related task information.

[0144] In some embodiments, the method can further include: based on the task arrangement information of the interactive task related to each target intent, a service function corresponding to the interactive task can be invoked to obtain business providing information of each service function; the business providing information of each service function can be adjusted according to a service type and a service display requirement to obtain service processing results of each service function.

[0145] In some embodiments, as shown in FIG. 8a, another voice interaction method for a display device is provided, which can include a display, which can be configured to display content from a broadcast system or a network and / or a user interface; one or more external device interfaces, which can be configured to communicate with one or more external devices according to a communication protocol, the communication protocol at least including a Bluetooth protocol; and at least one processor connected with the display and the one or more external device interfaces, and configured to execute instructions to control the display device. The method can include, but is not limited to, the aforementioned steps S801-S805; thus, after responding to the voice interaction instruction in the case of displaying a home page interface or a video playing interface, at least one control can be displayed in a display area in the jumped-to interactive dialogue interface in the case of identifying that the text information of the input voice contains a keyword of one of a holiday, a birthday, and an anniversary, and through the control, integrated recommendation data of at least one recommended object associated with the keyword can be displayed, the integrated recommendation data can at least include detail integrated information of the recommended object and a detail page link of the recommended object, and further, a detail page of the recommended object to which the detail page link belongs can be displayed in response to an instruction triggering the detail page link, thereby improving the voice interaction intelligent level of the display device, enabling it to handle complex voice interaction requirements containing an intent of one of a holiday, a birthday, and an anniversary, simplifying the voice interaction processing flow, providing a recommended plan of an intent of one of a holiday, a birthday, and an anniversary based on simple voice content of the user without the user repeatedly interacting multiple times to output an instruction, improving the voice interaction efficiency of the display device, and the display device can provide a specific interactive dialogue interface for voice interaction, which helps to improve the processing effect of voice interactive life services.

[0146] In some embodiments, the method can further include: in the case of identifying that the text information contains a keyword of one of a holiday, a birthday, and an anniversary, displaying reminder information of the keyword in the home page interface; the reminder information is used to perform a reminder operation at a time corresponding to the holiday or the birthday or the anniversary.

[0147] In some embodiments, the method can further include: the text information, the current date of the system in the display device, and the positioning information of the display device can be input into a multi-intent analysis model to obtain a plurality of intent-associated task information; each intent-associated task information can include a predicted intent expanded based on an intent, and task scheduling information of an interactive task associated with the predicted intent; at least one target intent can be obtained by performing intent understanding on each predicted intent and merging predicted intents having an intent connection relationship; service processing results of different service functions can be obtained by calling service functions corresponding to interactive tasks based on the task scheduling information of the interactive tasks associated with each target intent; the service processing results of the different service functions can at least include integrated recommendation data of at least one recommended object associated with the keyword.

[0148] In some embodiments, the method can further include: the service providing information of the display device can be acquired; the service providing information can be used to indicate the service functions possessed by the display device; the multiple predicted intents extended based on the keyword can be determined according to the keyword contained in the text information and the service functions possessed by the display device; the multiple intent-related task information can be determined by task scheduling in combination with the specified date containing the keyword in the text information, the current date of the system in the display device, the positioning information of the display device, and the configuration rules of the interactive tasks associated with each predicted intent.

[0149] In some embodiments, the method can further include: the service providing information of each service function can be obtained by calling the service function corresponding to the interactive task based on the task scheduling information of the interactive task associated with each target intent; and the service processing result of each service function can be obtained by adjusting the service providing information of each service function according to the service type and the service display requirement.

[0150] In some embodiments, as shown in FIG. 9a, another voice interaction method for a display device is provided, which can include a display, which can be configured to display content from a broadcast system or a network and / or a user interface; one or more external device interfaces, which can be configured to communicate with one or more external devices according to a communication protocol, the communication protocol at least including a Bluetooth protocol; and at least one processor connected with the display and the one or more external device interfaces and configured to execute instructions to control the display device. The method can include but is not limited to steps S901-S904; thus, after responding to the voice interaction instruction in the case of displaying a home page interface or a video playing interface, at least one control can be displayed in the display area in the interactive dialogue interface jumped to in the case of recognizing that the text information of the input voice contains a keyword of one of a festival, a birthday, and an anniversary, and the integrated recommendation data of at least one recommended object associated with the keyword is displayed through the control, the integrated recommendation data at least including detail integration information of the recommended object and a detail page link of the recommended object, and then the detail page of the recommended object to which the detail page link belongs can be displayed in response to an instruction triggering the detail page link, thereby improving the voice interaction intelligent level of the display device, enabling it to process complex voice interaction requirements containing one of a festival, a birthday, and an anniversary intent, simplifying the voice interaction processing flow, providing a festival, a birthday, and an anniversary one of the intent recommendation plan based on the simple voice content of the user without the user repeatedly interacting to output an instruction, improving the voice interaction efficiency of the display device, and the display device can provide a specific interactive dialogue interface for voice interaction, which helps to improve the voice interaction life service processing effect.

[0151] In some embodiments, the method can further include: in response to identifying that the text information contains one of the keywords of the festival, birthday, and anniversary, displaying reminder information of the keyword in the home page interface; the reminder information can be used to perform a reminder operation at a time corresponding to the festival or birthday or anniversary.

[0152] In some embodiments, the method can further include: inputting the text information, the current date of the system in the display device, and the positioning information of the display device into the multi-intent analysis model to obtain a plurality of intent-related task information; each intent-related task information can include a predicted intent expanded based on the keyword and task arrangement information of an interactive task related to the predicted intent; at least one target intent can be obtained by performing intent understanding on each predicted intent and merging predicted intents having an intent connection relationship; service processing results of different service functions can be obtained by invoking service functions corresponding to the interactive tasks based on the task arrangement information of the interactive tasks related to each target intent; the service processing results of different service functions can at least include integrated recommendation data of at least one recommended object related to the keyword.

[0153] In some embodiments, the method can further include: obtaining service providing information of the display device; the service providing information can be used to indicate service functions possessed by the display device; a plurality of predicted intents expanded based on the keyword can be determined according to the keyword contained in the text information and the service functions possessed by the display device; a plurality of intent-related task information can be determined by task arrangement based on the specified date of the keyword contained in the text information, the current date of the system in the display device, the positioning information of the display device, and configuration rules of the interactive tasks related to each predicted intent.

[0154] In some embodiments, the method can further include: obtaining service providing information of each service function by invoking service functions corresponding to the interactive tasks based on the task arrangement information of the interactive tasks related to each target intent; the service processing results of each service function can be obtained by adjusting the service providing information of each service function according to service types and service display requirements.

[0155] The following specific embodiments are used to elaborate an application example of the display device in some embodiments of the present application in detail. FIG. 10 is a signaling interaction diagram of voice interaction provided according to some embodiments of the present application. As shown in FIG. 10, the specific process includes but is not limited to the following:

[0156] Step S1001, the processor responds to a control instruction for entering a home page and sends a display instruction for displaying a home page interface to the display;

[0157] Step S1002, the display displays the home page interface;

[0158] Step S1003, the sound collector sends the input voice to the processor;

[0159] Step S1004, the processor sends a display text information instruction to the display in response to the voice interaction instruction;

[0160] Step S1005, the display displays the text information of the input voice while displaying the home page interface on the display;

[0161] Step S1006, the processor sends a display control instruction to the display in the case where it is identified that the text information contains one of the keywords of festival, birthday, and anniversary;

[0162] Step S1007, the display displays at least one control in the first display area in the home page interface, and the control displays integrated recommendation data of at least one recommended object associated with the keyword;

[0163] Step S1008, the processor sends a display detail page instruction to the display in response to the instruction triggering the detail page link;

[0164] Step S1009, the display displays the detail page of the recommended object to which the detail page link belongs.

[0165] In some embodiments, as shown in FIG. 10, the integrated recommendation data of the at least one recommended object is obtained by performing the following steps:

[0166] Step S1001', it is identified that the text information of the input voice contains one of the keywords of festival, birthday, and anniversary;

[0167] In the step S1001', the processor performs the identification after receiving and identifying the input voice.

[0168] Step S1002', the processor inputs the input text information, the current date of the system in the display device, and the positioning information of the display device into the multi-intent analysis model;

[0169] Step S1003', the multi-intent analysis model based on the large model multi-intent analysis module expands a plurality of predicted intents according to the keywords and the service functions possessed by the display device, and arranges tasks based on the predicted intents;

[0170] Step S1004', the large model multi-intent analysis module outputs a plurality of intent-related task information and sends it to the multi-intent decision processing module;

[0171] Step S1005', the multi-intent decision processing module performs intent understanding on each predicted intent and merges the intents;

[0172] Step S1006', the multi-intent decision module outputs at least one target intent and sends to the multi-intent service module;

[0173] Step S1007', the multi-intent service module obtains service processing results (including integrated recommendation data of the recommended object) of different service functions by calling service functions corresponding to the interactive task;

[0174] Step S1008', the multi-intent service module outputs the integrated recommendation data of the at least one recommended object and sends to the processor.

[0175] In some embodiments, according to the display device provided by some embodiments of the present application, referring to FIG. 11, the at least one processor can be further configured to execute instructions to enable the display device to implement another voice interaction method for a display device:

[0176] Step S1101, in response to a play control instruction for a target video, the display device controls the display to display a play interface of the target video.

[0177] In some embodiments, the play control instruction for the target video carries information such as an identifier of the target video and a network request address, and the display device can respond to the play control instruction for the target video. On the one hand, a data acquisition request can be generated based on the identifier of the target video and sent to the network request address. The server or database corresponding to the network request address can parse the data acquisition request to obtain the identifier of the target video. Based on the identifier of the target video, video data can be searched and returned to the display device. On the other hand, a video player can be called to play the received video data. At this time, the display displays the play interface of the target video. For example, referring to FIG. 12, which is a schematic diagram of a play interface of a target video according to some embodiments of the present application.

[0178] In some embodiments, the user can view a video of interest on the home page, search page, history page, etc. of the display device, take the video as a target video, and select the target video through a remote controller. After the display device detects the selection operation of the remote controller, a play control instruction for the target video can be generated based on information such as the identifier of the target video and the network request address.

[0179] In step S1102, in response to the voice interaction instruction, the display device controls the display to display the interaction result item in a first display area of a playing interface of the target video, and the first display area displays at least one result item display control, and each result item display control is used to display one interaction result item; the voice interaction instruction carries at least location information; and the at least one interaction result item includes at least one of a ticket query result, a weather query result, a lodging query result, and a memo making result for the location information.

[0180] In some embodiments, the voice interaction instruction can be generated based on an interaction voice input by the user to the display device.

[0181] As a possible implementation, the user can speak a voice of a specific word to wake up the voice interaction function of the display device. For example, the user can speak a voice of a specific word, and after the voice interaction function of the display device receives the voice of the specific word spoken by the user, the voice interaction function can be switched to a wake-up state and play audio of the specific word. When the user hears the display device play the audio of the specific word, it can be determined that the voice interaction function of the display device is woken up.

[0182] For example, the user can speak a voice of "hello", and when the user hears the display device play audio of "hello", it can be determined that the voice interaction function of the display device is woken up. It should be noted that the specific word spoken by the user and the specific word played by the display device can be pre-defined, and "hello" is only an example and does not constitute a limitation on the embodiments of the present application.

[0183] In some embodiments, the voice interaction function of the display device can be an application program.

[0184] In some embodiments, after the voice interaction function of the display device receives the voice of the specific word spoken by the user, the voice interaction function is switched to a wake-up state, and the voice spoken by the user subsequently is taken as the interaction voice input by the user.

[0185] In some embodiments, after the display device receives the interaction voice input by the user, the display device can convert the interaction voice into text, and can display the text near an icon corresponding to the voice interaction function on the display. For the convenience of description, the icon corresponding to the voice interaction function can be referred to as a voice icon in subsequent embodiments of the present application. For example, assuming that the interaction voice input by the user is "I will go to Shanghai for business tomorrow, stay there for three days, near Lujiazui, arrange for me", and the interface of the display is as shown in FIG. 13.

[0186] In some embodiments, the display device can convert the interactive voice into text. It is judged whether the text contains a place word. In the case of containing a place word, it is determined that the user needs a travel service. Then, based on the text, the large model can be used to disassemble a plurality of task sentences corresponding to the travel service, and voice interaction instructions can be generated based on the plurality of task sentences.

[0187] The plurality of task sentences can include a ticket query sentence for the place word, a weather query sentence for the place word, a lodging query sentence for the place word, and a note making sentence for the place word. For example, the ticket query sentence is divided into a one-way ticket query sentence and a return ticket query sentence.

[0188] The voice interaction instructions generated based on the plurality of task sentences can include at least one of a one-way ticket query instruction for the place word, a weather query instruction for the place word, a lodging query instruction for the place word, a return ticket query instruction for the place word, and a note making instruction for the place word. The place word in these instructions is the place information.

[0189] In some embodiments, the place word can be any word that can represent a place, such as a city, a scenic spot, etc. For example, the converted text is “I will go to Shanghai for business tomorrow, stay there for three days, near Lujiazui, arrange it for me”.

[0190] In some embodiments, after generating the voice interaction instructions, the at least one processor can be configured to execute instructions to cause the display device to perform: obtaining a task execution result corresponding to each task sentence as an interaction result item.

[0191] In some embodiments, the at least one processor can be configured to execute instructions to cause the display device to perform: calling a business microservice corresponding to each task sentence to obtain a task execution result corresponding to each task sentence.

[0192] In the case of determining that the user needs a travel service, a text disassembly requirement corresponding to the travel service can be obtained. Based on the converted text and the text disassembly requirement, the large model can be used to disassemble a plurality of task sentences.

[0193] The text splitting requirements corresponding to the travel service can include at least one of ticketing, weather, accommodation, and notes. Therefore, as mentioned above, the plurality of task sub-sentences obtained by splitting the large model can include at least one of a ticket query sub-sentence for a location word, a weather query sub-sentence for a location word, an accommodation query sub-sentence for a location word, and a note making sub-sentence for a location word. The at least one processor can be configured to execute instructions to cause the display device to perform: calling the interaction result item obtained by the business microservice corresponding to each task sub-sentence, which can include at least one of a ticket query result for a location information, a weather query result, an accommodation query result, and a note making result.

[0194] In some embodiments, the at least one processor can be configured to execute instructions to cause the display device to perform: controlling the display to display the interaction result item in a first display area on the playing interface of the target video. The first display area can be a pre-designated area, such as the right half of the playing interface of the target video.

[0195] In some embodiments, the interaction result item can be displayed in the first display area in the form of a floating layer.

[0196] In a specific implementation, a content layer floating above the original content of the first display area can be added to the first display area, and the interaction result item can be displayed on the content layer. In this way, the interaction result item displayed on the content layer does not block or interfere with the original content of the first display area on the playing interface of the target video.

[0197] In some embodiments, the first display area can display at least one result item display control, and each result item display control can be used to display an interaction result item. In the case where the interaction result item is displayed on the content layer, the content layer can display at least one result item display control, and each result item display control can be used to display an interaction result item.

[0198] In some embodiments, the at least one result item display control can be arranged on the content layer according to a preset rule, for example, the at least one result item display control can be arranged in a column.

[0199] The at least one result item display control can include at least one of a ticket display control, a weather display control, an accommodation display control, and a note display control. The ticket display control can be used to display a ticket query result, the weather display control can be used to display a weather query result, the accommodation display control can be used to display an accommodation query result, and the note display control can be used to display a note making result.

[0200] In some embodiments, when the at least one result item display control includes a ticket display control, a weather display control, an accommodation display control, and a reminder display control, and the at least one result item display control is arranged in a column, the ticket display control can be arranged first, the weather display control can be arranged second, the accommodation display control can be arranged third, and the reminder display control can be arranged fourth. This is only an example and does not limit the embodiments of the present application.

[0201] In some embodiments, the text splitting requirements corresponding to the travel service involve splitting dimensions, and the ticket dimension further includes a departure ticket dimension and a return ticket dimension. In this case, the plurality of task sub-sentences obtained by splitting the large model can include a departure ticket query sub-sentence for the location word and a return ticket query sub-sentence for the location word. The interaction result item obtained by the at least one processor calling the business microservice corresponding to each task sub-sentence can include a departure ticket query result and a return ticket query result. The content layer displays a departure ticket display control and a return ticket display control. The departure ticket display control can be used to display the departure ticket query result, and the return ticket display control can be used to display the return ticket query result.

[0202] In some embodiments, since the departure ticket query result, the weather query result, and the return ticket query result corresponding to different dates are different, when the text corresponding to the interaction voice only includes a location word, the departure ticket query result for the location information can be a ticket query result for the latest date of the location information, such as a ticket query result for the next day of the location information. The weather query result for the location information can be a weather query result for at least one day of the location information, such as a weather query result for the last 3 days of the location information. The return ticket query result for the location information can be a ticket query result for the day after the latest date of the location information, such as a ticket query result for the day after tomorrow of the location information. The embodiments of the present application do not limit this.

[0203] In some embodiments, when the text corresponding to the interaction voice only includes a location word, the reminder result for the location information can include a departure reminder for the date corresponding to the departure ticket query result and a return reminder for the date corresponding to the return ticket query result.

[0204] For example, when the interaction voice input by the user is "I will go to Shanghai for business tomorrow, stay there for three days, near Lujiazui, arrange it for me", the interaction result item can include a departure ticket query result, a weather query result, an accommodation query result, a return ticket query result, and a reminder result. The interaction result item can be displayed in the area shown in FIG. 14 by a floating layer display. The user can view each interaction result item up and down through the remote controller.

[0205] Step S1103: playing target audio corresponding to the at least one interaction result item.

[0206] In some embodiments, after obtaining the at least one interaction result item, the at least one processor can generate summary reply content based on the text corresponding to the interaction voice and the at least one interaction result item, and can synthesize the summary reply content into audio to obtain the target audio.

[0207] In a specific implementation, a summary induction requirement previously set for the interaction result item can be obtained, and the summary induction requirement can include a requirement for guiding the large model to summarize and induce the text corresponding to the interaction voice and the at least one interaction result item. The summary induction requirement, the text corresponding to the interaction voice, and the at least one interaction result item can be assembled to obtain a prompt, and the prompt can be input to the large model. The large model can output the summary reply content based on the received prompt.

[0208] For example, the text corresponding to the interaction voice is: I will go on a business trip to Shanghai tomorrow and stay there for three days. Please arrange it for me near Lujiazui. After the processing process of the foregoing embodiments, all interaction result items are obtained, and the summary reply content obtained by the large model is: “You will go from Qingdao to Shanghai tomorrow and stay in Lujiazui for 3 days. Train tickets and refreshments have been booked for you. It will be rainy in Shanghai on the 20th. Remember to bring an umbrella. The temperature will be 28 to 36 degrees. The itinerary has been added to the memo. You can operate it by sliding up and down through the remote control, etc.

[0209] After obtaining the summary reply content, the summary reply content is synthesized into audio to obtain target audio corresponding to the at least one interaction result item, and the target audio is played.

[0210] According to the display device provided in the foregoing embodiments, the display can display a play interface of a target video in response to a play control instruction for the target video. The display can display an interaction result item in a first display area of the play interface of the target video in response to a voice interaction instruction. The first display area can display at least one result item display control. Each result item display control can be used to display an interaction result item. The voice interaction instruction can carry at least location information. The at least one interaction result item can include at least one of a ticket query result, a weather query result, a lodging query result, and a memo making result for the location information. Target audio corresponding to the at least one interaction result item can be played. Thus, a user only needs to input a simple voice containing a location word, and can see multiple result items such as a ticket query result, a weather query result, a lodging query result, and a memo making result on the display device. The intelligence of the display device is improved, and the interaction between the user and the display device is simpler, and the user interaction experience is improved.

[0211] In some embodiments, the user can input a page opening instruction for the target page to the display device, and the display device can control the display to display the page of the target page in response to the page opening instruction. The user can input an interactive voice while the display displays the page of the target page, and the display device can convert the interactive voice input by the user into text. If the text meets the target service scenario, the text splitting requirement corresponding to the target service scenario can be obtained. The text splitting requirement can be information set in advance for the target service scenario to guide the large model to split the text. Based on the text and the text splitting requirement, the large model can be used to split to obtain a plurality of task sentences. Based on the plurality of task sentences, a voice interaction instruction can be generated. In response to the voice interaction instruction, the display device can control the display to display an interactive result item in a first display area on the playing interface of the target video. The first display area can display at least one result item display control. Each result item display control can be used to display an interactive result item. The voice interaction instruction can carry at least location information. The at least one interactive result item can include at least one of a ticket query result, a weather query result, an accommodation query result, and a memo making result for the location information. The target audio corresponding to the at least one interactive result item can be played.

[0212] The target page can be any page, for example, the target page can be a home page. That is, the user can perform voice interaction with the display device while the display displays any page, so as to make the display display at least one interactive result item. For details, refer to the description of the foregoing embodiments, which will not be described here.

[0213] In some embodiments, the at least one processor can be further configured to execute instructions to cause the display device to perform: controlling the target video in a playing state on the playing interface to pause playing, and controlling the target video to continue playing after the target audio is played. Alternatively, the at least one processor can be further configured to execute instructions to cause the display device to perform: controlling the target video in a playing state on the playing interface to reduce the playing volume, and controlling the target video to restore the playing volume after the target audio is played.

[0214] As described above, the voice interaction function of the display device can be an application program, which can be referred to as a voice interaction application for convenience of description. After obtaining the summary reply content, the display device can detect the application currently holding the audio focus as a foreground running application, give the audio focus to the voice interaction application, and control the target video in a playing state on the playing interface to pause playing after the foreground running application loses the audio focus. After the voice interaction application obtains the audio focus, the voice interaction application can play the audio corresponding to the summary reply content, i.e., the target audio. After the target audio is played, the audio focus can be returned to the foreground running application, and the foreground running application controls the target video to continue playing.

[0215] Alternatively, after the foreground running application loses the audio focus, the target video in the playing state on the playing interface can be controlled to reduce the playing volume, and after the voice interaction application finishes playing the target audio, the foreground running application can be given the audio focus, and the foreground running application controls the target video to restore the playing volume.

[0216] In the above embodiments, during the process of playing the target audio corresponding to the at least one interaction result item, the target video in the playing state on the playing interface can be controlled to pause playing; or the target video in the playing state on the playing interface can be controlled to reduce the playing volume, which can prevent the playing of the target audio from affecting the viewing of the target video, and further improve the user experience.

[0217] In some embodiments, the at least one processor can be further configured to execute instructions to cause the display device to perform: the text information corresponding to the target audio can be displayed on a first display area; or the text information corresponding to the target audio can be displayed on a second display area on the playing interface of the target video; the second display area and the first display area do not overlap.

[0218] In some embodiments, the text information corresponding to the target audio is the summary reply content mentioned above.

[0219] In some embodiments, the first display area can further include a text broadcast display control, which can be used to display the text information corresponding to the target audio, and the text information corresponding to the target audio can be displayed on the text broadcast display control.

[0220] In a specific implementation, in a case where the interaction result item is displayed on the first display area in a floating layer display manner, the corresponding content layer can include a text broadcast display control, and the text information corresponding to the target audio can be displayed on the text broadcast display control.

[0221] For example, the text corresponding to the interactive voice is: I will go on a business trip to Shanghai tomorrow, stay there for three days, near Lujiazui, arrange for me, through the processing process of the foregoing embodiments, after obtaining all the interaction result items, the summary reply content obtained by using the large model is: "You will go from Qingdao to Shanghai tomorrow, stay in Lujiazui for 3 days, have booked train tickets and refreshments for you, it will rain in Shanghai on the 20th, remember to bring an umbrella, the temperature is 28 to 36 degrees. The itinerary has been added to the memo, you can use the remote control and other operations such as up and down sliding". The summary reply content can be displayed in the text broadcast display control on the content layer where the interaction result item is located, as shown in FIG. 15.

[0222] In some embodiments, the text information corresponding to the target audio can also be displayed on a second display area on the playing interface of the target video; the second display area and the first display area do not overlap.

[0223] In a specific implementation, a content layer can be added in the second display area, which is suspended above the original content of the second display area, and the text information corresponding to the target audio can be displayed on the content layer. In this way, the text information displayed on the content layer will not block or interfere with the original content of the second display area on the playing interface of the target video.

[0224] For example, the text corresponding to the interactive voice is: I will go to Shanghai for business tomorrow, stay there for three days, near Lujiazui, arrange it for me. After the processing process of the foregoing embodiment, all interactive result items are obtained, and the summary reply content obtained by using the large model is: "You will go from Qingdao to Shanghai tomorrow, stay in Lujiazui for 3 days, have booked train tickets and refreshments for you, it will rain in Shanghai on the 20th, remember to bring an umbrella, the temperature is 28 to 36 degrees. The itinerary has been added to the memo, you can slide up and down through the remote control and other operations". The summary reply content can be displayed on the content layer suspended above the original content of the second display area, as shown in FIG. 16.

[0225] In the foregoing embodiment, the text information corresponding to the target audio can be displayed on the first display area; or the text information corresponding to the target audio can be displayed on the second display area on the playing interface of the target video; and the second display area and the first display area do not overlap. This combination of voice broadcast and text display expands the dimension of information transmission and improves the efficiency of users capturing information.

[0226] In some embodiments, the ticket query result can include at least one selectable transportation ticket, the result item display control for displaying the ticket query result can include a reservation entry of each selectable transportation ticket; and the at least one processor can be further configured to execute instructions to cause the display device to perform: in response to a triggering operation on the reservation entry of a target transportation ticket in the at least one selectable transportation ticket, controlling the display to display a reservation page of the target transportation ticket; and in response to a ticket booking operation triggered on the reservation page, completing reservation of the target transportation ticket.

[0227] In some embodiments, the ticket query result can include at least one selectable transportation ticket, the result item display control for displaying the ticket query result can include a reservation entry of each selectable transportation ticket; and the at least one processor can be further configured to execute instructions to cause the display device to perform: in response to a triggering operation on the reservation entry of a target transportation ticket in the at least one selectable transportation ticket, controlling the display to display a reservation page of the target transportation ticket; and in response to a ticket booking operation triggered on the reservation page, completing reservation of the target transportation ticket.

[0228] The trigger operation on the pre-order entry of the target traffic ticket can be implemented by using a remote controller, or can be directly triggered by hand if the display device supports touch control, which is not limited in the embodiments of the present application. For the convenience of description, the pre-order entry of each selectable traffic ticket is referred to as a first entry.

[0229] Referring to FIG. 17, the ticket query result can include four selectable traffic tickets, and the result item display control for displaying the ticket query result can include a pre-order entry of each selectable traffic ticket. Assuming that the user wants to select the second traffic ticket, the second traffic ticket is taken as a target traffic ticket, and a trigger operation is performed on the pre-order entry of the second traffic ticket. The display device can respond to the trigger operation to control the display to display a pre-order page of the second traffic ticket. The pre-order of the second traffic ticket can be completed in response to a ticket booking operation triggered on the pre-order page.

[0230] In the above embodiments, when the user wants to pre-order a certain traffic ticket in the ticket query result, a trigger operation is performed on the pre-order entry of the traffic ticket, and the display device jumps to the corresponding pre-order page, thereby improving the convenience of ticket booking.

[0231] In some embodiments, the ticket query result can be divided into a travel ticket query result and a return ticket query result. The travel ticket query result can include at least one selectable travel ticket, and the result item display control for displaying the travel ticket query result can include a pre-order entry of each selectable travel ticket. The return ticket query result can include at least one selectable return ticket, and the result item display control for displaying the return ticket query result can include a pre-order entry of each selectable return ticket. The at least one processor can be further configured to execute instructions to cause the display device to perform the following operations. The display of a pre-order page of a target travel ticket can be controlled in response to a trigger operation on a pre-order entry of the target travel ticket in the at least one selectable travel ticket. The pre-order of the target travel ticket can be completed in response to a ticket booking operation triggered on the pre-order page of the target travel ticket. The display of a pre-order page of a target return ticket can be controlled in response to a trigger operation on a pre-order entry of the target return ticket in the at least one selectable return ticket. The pre-order of the target return ticket can be completed in response to a ticket booking operation triggered on the pre-order page of the target return ticket.

[0232] As described above, in the splitting dimensions involved in the text splitting requirement corresponding to the above-mentioned travel service, the ticket dimension can further include a travel ticket dimension and a return ticket dimension. In this case, the interaction result item can include a travel ticket query result and a return ticket query result. The travel ticket display control can be used to display the travel ticket query result, and the return ticket display control can be used to display the return ticket query result.

[0233] In the above embodiments, when the user wants to book a ticket in the ticket query result, the user performs a trigger operation on the booking entry of the ticket, and the display device jumps to the corresponding booking page, thereby improving the convenience of ticket booking.

[0234] In the above embodiments, when the user wants to book a ticket in the ticket query result, the user performs a trigger operation on the booking entry of the ticket, and the display device jumps to the corresponding booking page, thereby improving the convenience of ticket booking.

[0235] The trigger operation on the booking entry of the target ticket can be implemented by using a remote controller, or can be directly triggered by hand when the display device supports touch control. As described above, the booking entry of each selectable ticket is referred to as a first entry. The selectable ticket includes a selectable outbound ticket and a selectable return ticket. The booking entry of the selectable outbound ticket and the selectable return ticket is also referred to as a first entry.

[0236] In the above embodiments, when the user wants to book a ticket in the ticket query result, the user performs a trigger operation on the booking entry of the ticket, and the display device jumps to the corresponding booking page, thereby improving the convenience of ticket booking.

[0237] In some embodiments, the weather query result can include weather information of at least one day in the past. The result item display control for displaying the weather query result can include a detail viewing entry of each piece of weather information. The at least one processor can be further configured to execute instructions to cause the display device to perform the following operation: the display device can display a detail display page corresponding to target weather information in response to a trigger operation on the detail viewing entry of the target weather information.

[0238] The user can view the weather query result in the weather display control, the weather query result can include weather information of at least one day recently, when the user wants to view the details of a certain weather information, the weather information can be taken as target weather information, a trigger operation is performed on the details viewing entry of the target weather information, and the display device can control the display to display a details display page corresponding to the target weather information in response to the trigger operation.

[0239] The trigger operation on the details viewing entry of the target weather information can be implemented by using a remote controller, or can be directly triggered by hand when the display device supports touch, and the embodiments of the present application do not limit this.

[0240] Referring to FIG. 18, the weather query result can include weather information of the last four days; the result item display control for displaying the weather query result can include a details viewing entry of each weather information; when the user wants to view the details of the weather information of tomorrow, the weather information of tomorrow can be taken as target weather information, and a trigger operation is performed on the details viewing entry of the target weather information, and the display device can control the display to display a details display page of the weather information of tomorrow in response to the trigger operation.

[0241] In the above embodiments, when the user wants to view the details of the weather information of a certain day, a trigger operation is performed on the details viewing entry of the weather information, and the display device jumps to the corresponding details display page, thereby improving the convenience of weather query.

[0242] In some embodiments, the voice interaction instruction can also carry travel date information, and the weather query result includes weather information of a date determined according to the travel date information.

[0243] In the text corresponding to the interactive voice, in addition to containing the place word, the time word is also contained. In this case, the multiple task sentences obtained by the large model disassembling can include at least one of the travel ticket query sentence for the place word and the time word, the weather query sentence for the place word and the time word, the accommodation query sentence for the place word, the return ticket query sentence for the place word and the time word, and the memo making sentence for the place word and the time word. The voice interaction instruction generated based on the multiple task sentences can include at least one of the travel ticket query instruction for the place word and the time word, the weather query instruction for the place word and the time word, the accommodation query instruction for the place word, the return ticket query instruction for the place word and the time word, and the memo making instruction for the place word and the time word. The place word in these instructions can be the place information, and the time word can be the date information. The display device can obtain, in response to the voice interaction instruction, the travel ticket query result for the place information and the date information, the weather query result for the place information and the date information, the accommodation query result for the place information, the return ticket query result for the place information and the date information, and the memo making result for the place information and the date information. The weather query result contains the weather information of the date determined according to the travel date information.

[0244] In the above embodiment, in addition to containing the place word, the text corresponding to the interactive voice can also contain the time word. In this case, the voice interaction instruction also carries the date information. In response to the voice interaction instruction, the display device obtains the weather query result containing the weather information of the date determined according to the travel date information, so that the user can view the weather information of the date corresponding to the time word.

[0245] In some embodiments, the accommodation query result can include at least one optional accommodation place, and the result item display control for displaying the accommodation query result can include a reservation entry of each optional accommodation place. The at least one processor can be further configured to execute instructions to cause the display device to perform: in response to a triggering operation on the reservation entry of a target accommodation place in the at least one optional accommodation place, controlling the display to display a reservation page of the target accommodation place; and in response to a reservation operation triggered on the reservation page, completing the reservation of the target accommodation place.

[0246] In the above embodiment, the user can view the accommodation query result in the accommodation display control. The accommodation query result can include at least one optional accommodation place. If the user wants to reserve a certain accommodation place, the user can take the accommodation place as a target accommodation place, perform a triggering operation on a reservation entry of the target accommodation place, and in response to the triggering operation, the display device can control the display to display a reservation page of the target accommodation place. The user can trigger a reservation operation on the reservation page, and in response to the reservation operation, the display device can complete the reservation of the target accommodation place.

[0247] The triggering operation on the reservation entry of the target accommodation site can be implemented by a remote controller, or directly by hand if the display device supports touch control, which is not limited in the embodiments of the present application. For convenience of illustration, the reservation entry of each optional accommodation site is referred to as a second entry.

[0248] Referring to FIG. 19, the accommodation query result can include four optional accommodation sites, and the result item display control for displaying the accommodation query result can include a reservation entry of each optional accommodation site. Assuming that the user wants to select the second accommodation site, the second accommodation site is taken as a target accommodation site, and a triggering operation is performed on the reservation entry of the second accommodation site. The display device can control the display to display a reservation page of the second accommodation site in response to the triggering operation. The reservation of the second accommodation site can be completed in response to a ticket booking operation triggered on the reservation page.

[0249] In the above embodiments, when the user wants to reserve a certain accommodation site in the accommodation query result, a triggering operation is performed on the reservation entry of the accommodation site, and the display device jumps to the corresponding reservation page, thereby improving the convenience of reservation of the accommodation site.

[0250] In some embodiments, the at least one result item display control can be displayed in a column. The at least one processor can be further configured to execute instructions to cause the display device to perform: in response to an information down scrolling instruction, display the lower controls of the currently displayed result item display control in the first display area; and in response to an information up scrolling instruction, display the upper controls of the currently displayed result item display control in the first display area.

[0251] The at least one result item display control can be displayed in a column. Due to the limited height of the first display area, all the result item display controls can not be displayed at one time. In this case, the user can input an information down scrolling instruction to the display device by using a remote controller. In response to the information down scrolling instruction, the lower controls of the currently displayed result item display control are displayed in the first display area.

[0252] In some embodiments, when the currently displayed result item display control is a travel ticket display control, and the user continuously inputs an information down scrolling instruction, the display interface changes as shown in FIG. 20.

[0253] The user can also input an information up scrolling instruction to the display device by using a remote controller. The display device can display the upper controls of the currently displayed result item display control in the first display area in response to the information up scrolling instruction.

[0254] In the above embodiments, the at least one result item display control is displayed in a column, and the user can input information up and information down instructions to the display device through the remote controller to flip up and down each result item display control to view the interactive result item displayed on the result item display control, and the operation is more convenient.

[0255] In some embodiments, the result item display control for displaying the ticket query result can further include a travel tool recommendation area, and the travel tool recommendation area can be used to display travel tool recommendation information summarized based on the ticket query result. The result item display control for displaying the weather query result can further include a weather reminder area, and the weather reminder area can be used to display weather-related reminder information summarized based on the weather query result.

[0256] The ticket query result can include at least one selectable transportation ticket, and the large model can be used to summarize and conclude the at least one selectable transportation ticket to obtain the travel tool recommendation information.

[0257] In a specific implementation, a recommendation requirement set in advance for the travel tool recommendation can be obtained, and the recommendation requirement can include a requirement for guiding the large model to recommend the at least one selectable transportation ticket in the ticket query result, for example, the recommendation requirement can be "recommend a reasonable travel tool based on the distance to the destination, the ticket price of the train ticket and the air ticket, and the weather information of the destination on the arrival day". The recommendation requirement and the ticket query result can be assembled to obtain a prompt, and the prompt is input to the large model. The large model can output the travel tool recommendation information based on the received prompt.

[0258] In combination with the foregoing embodiments, the ticket query result can be divided into an outbound ticket query result and a return ticket query result, the outbound ticket display control can be used to display the outbound ticket query result, and the return ticket display control can be used to display the return ticket query result. The outbound ticket display control and the return ticket display control can both include a travel tool recommendation area. The travel tool recommendation area on the outbound ticket display control can display travel tool recommendation information summarized based on the outbound ticket query result. The travel tool recommendation area on the return ticket display control can display travel tool recommendation information summarized based on the return ticket query result.

[0259] As a possible implementation, the travel ticket query sentence for the location word and the time word is: "Tickets from Qingdao to Shanghai on August 23, 2024", and the corresponding travel ticket query result includes air tickets, train tickets, etc. The recommendation requirement is: "distance to the destination, ticket price of train tickets and air tickets, and weather information on the day of arrival at the destination, etc. Recommend a reasonable travel tool", the recommendation requirement and the ticket query result can be assembled to obtain a prompt, and the prompt is input into the large model. The travel tool recommendation information output by the large model is: "According to the information you provided, the price of the train ticket is 308 yuan, and the time consumption is about 4 hours and 17 minutes; while the price of the air ticket is 920 yuan, and the time consumption is about 1 hour and 40 minutes. Considering comprehensively, it is recommended that you choose train travel, although the train takes longer than the plane, but the price is more affordable, and the train travel is more stable than the plane, not easily affected by weather factors", which can be displayed in the travel tool recommendation area of the travel ticket display control, as shown in FIG. 21.

[0260] Similarly, the weather query result can include weather information of at least one day, and the large model can be used to summarize the weather information of these days to obtain weather-related reminder information.

[0261] In a specific implementation, a recommendation requirement previously set for weather reminders can be obtained. The recommendation requirement can include requirements for guiding the large model to recommend the weather information of at least one day in the weather query result, for example, the recommendation requirement can be "abnormal weather reminder, clothing reminder, and umbrella reminder". The recommendation requirement and the weather query result can be assembled to obtain a prompt, and the prompt is input into the large model. The large model can output weather-related reminder information based on the received prompt.

[0262] As a possible implementation, the weather query sentence for the location word and the time word is: "Weather in Shanghai from August 23, 2024 to August 26, 2024", and the recommendation requirement is: "abnormal weather reminder, clothing reminder, and umbrella reminder". The recommendation requirement and the weather query result can be assembled to obtain a prompt, and the prompt is input into the large model. The weather-related reminder information output by the large model is: "Temperature in Shanghai from August 23, 2024 to August 26, 2024 is 27-34 degrees, overcast, and air quality is general. It is recommended to wear light and breathable clothes, and carry an umbrella to prevent rain. If you are sensitive to air quality, it is recommended to wear a mask. Avoid going out during high-temperature periods, and always supplement water to prevent dehydration. Although it is overcast, sun protection is still necessary. Wish you a successful trip", which can be displayed in the weather reminder area, as shown in FIG. 22.

[0263] In the above embodiments, the travel tool recommendation area displays travel tool recommendation information summarized based on the ticket query result, and the weather reminder area displays weather-related reminder information summarized based on the weather query result, which can help the user make a more reasonable choice and make comprehensive travel preparations, and further improve the user interaction experience.

[0264] In some embodiments, the at least one processor can be further configured to execute instructions to cause the display device to perform: the reminder result can be sent to the mobile terminal through the Internet of Things to instruct the mobile terminal to make a reminder based on the reminder result.

[0265] Among them, the mobile terminal can be installed with an application program for communication with the display device. The display device can send the reminder result to the application program through the Internet of Things, and the application program can make a reminder based on the reminder result.

[0266] Among them, in the case that the text corresponding to the interactive voice contains time words in addition to place words, the reminder result for the place information and the date information can include: an outbound travel reminder on the day before the outbound travel date determined based on the time words, and a return travel reminder on the day before the return travel date determined based on the time words, the outbound travel reminder can include reminder information such as the outbound travel destination, and the return travel reminder can include reminder information such as the return travel destination.

[0267] For example, the reminder result can include: an outbound travel reminder on 2024.08.22 and a return travel reminder on 2024.08.25, and the display device can send the reminder result to the mobile terminal. After receiving the reminder result, the mobile terminal calls the alarm on the mobile terminal to remind the user to check the outbound travel reminder on 2024.08.22, and calls the alarm on the mobile terminal to remind the user to check the return travel reminder on 2024.08.25.

[0268] In the above embodiments, the reminder result is sent to the mobile terminal through the Internet of Things, and the mobile terminal can remind the user to check the outbound travel reminder or the return travel reminder through the alarm and other ways at the corresponding time, further improving the service intelligence.

[0269] In some embodiments, the at least one processor can be further configured to execute instructions to cause the display device to perform: the user input interactive voice can be converted into text; in the case that the text meets the target service scenario, the text splitting requirement corresponding to the target service scenario can be obtained; the text splitting requirement can be information for guiding the large model to split the text which is set in advance for the target service scenario; based on the text and the text splitting requirement, a large model can be used to split to obtain a plurality of task sentences; based on the plurality of task sentences, a voice interaction instruction can be generated; in response to the voice interaction instruction, a task execution result corresponding to each task sentence can be obtained as an interaction result item.

[0270] In some embodiments, after receiving the interactive voice of the user input, the interactive voice can be converted into text, and it is determined whether the text meets the target service scenario. Specifically, it can be determined whether the text contains a place word. In the case where the text contains a place word, it is determined that the user needs a travel service, that is, the text meets the target service scenario.

[0271] In some embodiments, in the case where the text corresponding to the interactive voice meets the target service scenario, the text splitting requirement corresponding to the target service scenario can be obtained. The text splitting requirement can be information set in advance for the target service scenario to guide the large model to split the text.

[0272] For example, the text splitting requirement corresponding to the target service scenario can be: 1. According to the departure place, destination, time information, etc., make a travel plan. The travel plan needs to include: booking a ticket or train ticket for travel, querying a hotel near the destination, querying the weather of the destination, querying a ticket or train ticket for return, making a one-day schedule reminder before departure and before return; 2. Split the input into sub-sentences according to the above requirements; 3. You can refer to this example: Example 1.

[0273] It should be noted that the above example 1 can be an example of model input and output in a pre-set target service scenario, and the embodiments of the present application are not limited.

[0274] In some embodiments, the above text and text splitting requirement can be directly input into the large model. The large model understands and splits the above text based on the text splitting requirement, and then outputs multiple task sub-sentences. Alternatively, the processing reference information corresponding to the target service scenario can be obtained, such as the user's region, the current date, etc. The above text, text splitting requirement and the processing reference information are input into the large model. The large model understands and splits the above text and the processing reference information based on the text splitting requirement, and then outputs multiple task sub-sentences. After obtaining the multiple task sub-sentences, the voice interaction instruction is generated based on the multiple task sub-sentences.

[0275] As a possible implementation, in response to the voice interaction instruction, the display device can identify the task intent corresponding to each task sub-sentence, and based on the task intent, call the corresponding business microservice to obtain the task execution result as the interactive result item.

[0276] For example, given multiple task clauses including "tickets from Qingdao to Shanghai on August 23, 2024", "hotels near Lujiazui on August 23, 2024", and "weather in Shanghai from August 23 to August 26, 2024", the task intent corresponding to the clause "tickets from Qingdao to Shanghai on August 23, 2024" can be identified as ticketing. Therefore, the ticketing query microservice is invoked to query tickets from Qingdao to Shanghai on August 23, 2024, thus obtaining the task execution result for this clause. Similarly, the task intent corresponding to "2024" can be identified. The task intent corresponding to the clause "Hotels near Lujiazui on 08.23" is to query hotels. Therefore, the hotel query microservice is invoked to query hotels near Lujiazui on 08.23, 2024, thus obtaining the task execution result for this clause. Similarly, the task intent corresponding to the clause "Weather in Shanghai from 08.23 to 08.26, 2024" is to query the weather. Therefore, the weather query microservice is invoked to query the weather in Shanghai from 08.23 to 08.26, 2024, thus obtaining the task execution result for this clause.

[0277] In the above embodiments, users only need to input a simple voice sentence containing a location word to see multiple results on the display device, such as ticket search results, weather search results, accommodation search results, and memo creation results, which improves the intelligence of the display device, simplifies the interaction between the user and the display device, and enhances the user interaction experience.

[0278] In some embodiments, at least one processor may be further configured to execute instructions to cause the display device to perform: determining that the text matches a target service scenario if the text contains location words.

[0279] The display device can determine whether the text corresponding to the interactive voice contains location words. If it does, it assumes that the user needs travel services and determines that the text matches the target service scenario.

[0280] In the above embodiments, it can be determined whether the text corresponding to the interactive voice contains location words to determine whether the text matches the target service scenario. If it matches, the text segmentation requirements corresponding to the target service scenario are obtained. Based on the text and text segmentation requirements, multiple task sentences are obtained using a large model. Voice interaction commands are generated based on the multiple task sentences. In response to the voice interaction commands, the task execution results corresponding to each task sentence are obtained as interaction result items. This allows users to see multiple result items on the display device by simply inputting a voice sentence containing location words, improving the intelligence of the display device, simplifying the interaction between the user and the display device, and enhancing the user interaction experience.

[0281] In some embodiments, the at least one processor performs text-based and text splitting requirements, and utilizes a large model to split into multiple task sentences, which can be further configured to execute instructions to cause the display device to perform: the processing reference information corresponding to the target service scenario can be obtained; the processing keywords can be extracted based on the text and the processing reference information; the model input information can be constructed based on the processing keywords; the assembly information can be obtained by assembling the model input information and the text splitting requirements according to the prompt word template of the large model; and the multiple task sentences can be obtained by inputting the assembly information into the large model.

[0282] In some embodiments, the processing reference information corresponding to the target service scenario can be obtained, such as the region where the user is located, the current date, etc.

[0283] In some embodiments, the processing keywords can be extracted based on the text converted from the interactive voice and the obtained processing reference information. For example, the text converted from the interactive voice is: I will go to Shanghai for business tomorrow, stay there for three days, near Lujiazui, arrange it for me. The region where the user is located is Qingdao, Shandong, and the current date is 2024.08.22. The extracted processing keywords are: departure place: Qingdao; destination: Shanghai; accommodation: Lujiazui; travel time: 08.23; return time: 08.26.

[0284] In some embodiments, after the processing keywords are extracted, the text converted from the interactive voice can be supplemented based on the processing keywords, thereby obtaining the model input information. For example, the model input information constructed based on the processing keywords of the above example is: 2024.08.23 from Qingdao to Shanghai for business, the destination is near Lujiazui, and the return time is 2024.08.26.

[0285] In some embodiments, the prompt word prompt template of the large model can be obtained, and the model input information and the text splitting requirements can be assembled according to the prompt word template to obtain the assembly information.

[0286] For example, the assembly information can be:

[0287] "You are a travel planning intelligent agent, and you will make travel arrangements according to the requirements.

[0288] Specific requirements:

[0289] 1. According to the departure place, destination, time information, etc., make travel arrangements; the travel arrangements need to include: booking tickets for travel by plane or train, querying hotels near the destination, querying the weather of the destination, querying tickets for return by plane or train, making a one-day schedule reminder before departure and before return;

[0290] 2. According to the above requirements, split the input into sentences;

[0291] 3. You can refer to this example: Example 1.

[0292] Next, my input is: 2024.08.23 from Qingdao to Shanghai for business, the destination is near Lujiazui, and the return time is 2024.08.26.

[0293] In some embodiments, the assembly information can be input to the large model, and the large model understands and splits the model input information, and then outputs multiple task sentences supported by the Natural Language Understanding (NLU) system.

[0294] As a possible implementation, after inputting the above assembly information into the large model, the multiple task sentences output by the large model are: 2024.08.23 ticket from Qingdao to Shanghai, 2024.08.23 hotel near Lujiazui, 2024.08.26 ticket from Shanghai to Qingdao, 2024.08.23 to 024.08.26 weather in Shanghai, and 2024.08.22 travel reminder and 2024.08.25 return reminder.

[0295] In the above embodiment, on the basis of the text after interactive speech conversion, the processing reference information corresponding to the target service scene is also obtained, such as the region where the user is located and the current date. The combination of these auxiliary information enables the display device to complete the user's demand even if the user only inputs simple voice, further improving the intelligence of the display device. In addition, after obtaining the model input information, the model input information and the text splitting requirement are packaged according to the prompt word template of the model and then input to the model, so that the model can better understand the received information, and the final disassembly result is also more accurate.

[0296] In some embodiments, the at least one processor performing obtaining a task execution result corresponding to each task sentence can be further configured to execute instructions to cause the display device to perform: for each task sentence, performing intent understanding on the current task sentence to obtain a task intent corresponding to the current task sentence, and calling a corresponding business microservice to obtain a task execution result of the current task sentence based on the task intent corresponding to the current task sentence.

[0297] In some embodiments, after the multiple task sentences are obtained by disassembly, the multiple task sentences can be input to the NLU system respectively, and the NLU system performs intent understanding on each task sentence to obtain a task intent corresponding to each task sentence. Then, a corresponding business microservice can be called to obtain a task execution result based on the task intent corresponding to each task sentence.

[0298] With the above examples, there are six task sentences in total. Through the NLU system, it can be determined that the task intent corresponding to the sentence "2024.08.23 Qingdao to Shanghai ticket" is to query the travel ticket, and then the ticket query microservice is called to query the 2024.08.23 Qingdao to Shanghai ticket, so as to obtain the task execution result of this sentence; through the NLU system, it can be determined that the task intent corresponding to the sentence "2024.08.23 near Lujiazui hotel" is to query the accommodation, and then the hotel query microservice is called to query the 2024.08.23 near Lujiazui hotel, so as to obtain the task execution result of this sentence; through the NLU system, it can be determined that the task intent corresponding to the sentence "2024.08.26 Shanghai to Qingdao ticket" is to query the return ticket, and then the ticket query microservice is called to query the 2024.08.26 Shanghai to Qingdao ticket, so as to obtain the task execution result of this sentence; through the NLU system, it can be determined that the task intent corresponding to the sentence "2024.08.23 to 024.08.26 Shanghai weather" is to query the weather, and then the weather query microservice is called to query the 2024.08.23 to 024.08.26 Shanghai weather, so as to obtain the task execution result of this sentence; through the NLU system, it can be determined that the task intent corresponding to the sentence "make 2024.08.22 travel reminder; make 2024.08.25 return reminder" is to make a note, and then the note making microservice is called to make 2024.08.22 travel reminder and 2024.08.25 return reminder, so as to obtain the task execution result of this sentence.

[0299] In the above embodiment, after the multiple task sentences are disassembled, the NLU system can understand the intent of each task sentence. Based on the task intent corresponding to each task sentence, the corresponding business microservice can be called to obtain the task execution result of the current task sentence, which is displayed for the user to view. Therefore, the user only needs to input a simple voice containing a place word, and multiple aspects of result items can be seen on the display device, which improves the intelligence of the display device.

[0300] In some embodiments, a display device can include a display configured to display content from a broadcast system or network and / or a user interface, a memory configured to store computer programs or instructions, and at least one processor connected with the display and the memory and configured to execute the computer programs or instructions to cause the display device to: in response to a play control instruction for a target video, control the display to display a play interface of the target video; in response to a voice interaction instruction, control the display to display, in the play interface of the target video, an interactive result item in a first display area, the first display area can display at least one result item display control, each result item display control can be used to display one interactive result item; the voice interaction instruction can carry at least location information; the at least one interactive result item can include at least one of a ticket query result, a weather query result, an accommodation query result, a scenic spot query result, a guide query result, and a memo making result for the location information; and a target audio corresponding to the at least one interactive result item can be played.

[0301] In some embodiments, the display device can convert the interactive voice into text. It is determined whether the text contains a location word. In the case of containing a location word, it is determined that the user needs a travel service. Then, based on the text, a large model can be used to disassemble a plurality of task sub-sentences corresponding to the travel service, and based on the plurality of task sub-sentences, a voice interaction instruction can be generated.

[0302] In some embodiments, the plurality of task sub-sentences can include a ticket query sub-sentence for the location word, a weather query sub-sentence for the location word, an accommodation query sub-sentence for the location word, a scenic spot query sub-sentence for the location word, a guide query sub-sentence for the location word, and a memo making sub-sentence for the location word. In some embodiments, the ticket query sub-sentence can be divided into a travel ticket query sub-sentence and a return ticket query sub-sentence.

[0303] In some embodiments, the voice interaction instruction generated based on the plurality of task sub-sentences can include at least one of a travel ticket query instruction for the location word, a weather query instruction for the location word, an accommodation query instruction for the location word, a return ticket query instruction for the location word, a scenic spot query instruction for the location word, a guide query instruction for the location word, and a memo making instruction for the location word. The location word in these instructions is the location information.

[0304] In some embodiments, after generating the voice interaction instruction, in response to the voice interaction instruction, a task execution result corresponding to each task sub-sentence can be obtained as an interactive result item. The interactive result item can include at least one of a ticket query result, a weather query result, an accommodation query result, a scenic spot query result, a guide query result, and a memo making result for the location information.

[0305] That is, compared with the travel service, the travel service corresponds to an increased number of task sentences for the place word, including a scenic spot query sentence for the place word and a guide query sentence for the place word, and the interaction result items are increased, including a scenic spot query result and a guide query result. The specific implementation process is similar, and reference can be made to the foregoing description of the embodiments, which will not be repeated here.

[0306] In some embodiments, a voice interaction method for a display device is provided. Referring to FIG. 23 and FIG. 24, FIG. 23 is a flowchart of another voice interaction method for a display device according to some embodiments of the present application. FIG. 24 is a module interaction diagram of another voice interaction method for a display device according to some embodiments of the present application. The display device can run a voice interaction application, a large model, a service execution module, and a user interface module; the method can include but is not limited to the following steps:

[0307] In step S1701, the voice interaction application receives the user input interaction voice, converts the interaction voice into text; in the case that the text contains a place word, it is determined that the text meets the target service scenario, and the text splitting requirement corresponding to the target service scenario is obtained; the text splitting requirement is information for guiding the large model to split the text, which is set in advance for the target service scenario;

[0308] The implementation of the steps of receiving the interaction voice by the voice interaction application, converting the interaction voice into text, determining whether the text meets the target service scenario, and obtaining the text splitting requirement is similar to the foregoing embodiments, and will not be repeated here.

[0309] In step S1702, the large model splits the text based on the text and the text splitting requirement to obtain a plurality of task sentences, and the plurality of task sentences include at least one of a ticket query sentence for a place word, a weather query sentence, a lodging query sentence, and a memo making sentence;

[0310] Specifically, the processing reference information corresponding to the target service scenario can be obtained; the processing keyword can be extracted based on the text and the processing reference information; the model input information can be constructed based on the processing keyword; the model input information and the text splitting requirement can be assembled according to the prompt word template of the language processing model corresponding to the target service scenario to obtain assembly information; and the assembly information can be input into the large model to obtain the plurality of task sentences. For details, reference can be made to the foregoing embodiments, which will not be repeated here.

[0311] In step S1703, the service execution module obtains a task execution result corresponding to each task sentence as an interaction result item;

[0312] In some embodiments, the service execution module can include various business microservices, such as a ticket query microservice, a hotel query microservice, a weather query microservice, a reminder making microservice, etc. The multiple task sentences can be input to the NLU system respectively, the NLU system performs intent understanding on each task sentence to obtain the task intent corresponding to each task sentence. Then, based on the task intent corresponding to each task sentence, the corresponding business microservice is called to obtain the task execution result as the interaction result item. For details, refer to the foregoing embodiments, which will not be described here.

[0313] In step S1704, the user interface module displays the interaction result item.

[0314] In some embodiments, the user interface module can display the interaction result item in the first display area in the form of a floating layer. For details, refer to the foregoing embodiments, which will not be described here.

[0315] In the above embodiments, the voice interaction application can receive an interaction voice input by a user (S1801), convert the interaction voice into text (S1802); in a case where the text meets a target service scenario, obtain a text splitting requirement corresponding to the target service scenario; the text splitting requirement can be information set in advance for the target service scenario to guide the large model to split the text; the large model can split the text based on the text and the text splitting requirement to obtain multiple task sentences (S1803); the service execution module can obtain a task execution result corresponding to each task sentence as an interaction result item (S1804); and the user interface module can display the interaction result item (S1805). In this way, the user only needs to input a simple voice containing a place word, and multiple result items in multiple aspects can be seen on the display device, improving the intelligence of the display device. The interaction between the user and the display device is also simpler, improving the user interaction experience.

[0316] In some embodiments, referring to FIG. 19, the voice interaction application can include a voice input module and a speech recognition module, and the voice interaction application receiving an interaction voice input by a user and converting the interaction voice into text can include: the voice input module receiving an interaction voice input by a user; and the speech recognition module converting the interaction voice into text. Referring to FIG. 25, the voice input module receives an interaction voice input by a user (S1901) and sends it to the speech recognition module (S1902); and the speech recognition module converts the interaction voice into text (S1903). This way of processing in modules can decouple the processing process, facilitating modular management.

[0317] In some embodiments, the display device can further run a summary module, and the method can further include: the summary module can generate summary reply content based on the text and the task execution result; the summary reply content can be sent to the voice interaction application and the user interface module respectively; the voice interaction application can perform a broadcast process based on the summary reply content; and the user interface module displays the summary reply content.

[0318] Specifically, referring to FIG. 26, the summary module can receive the task execution result sent by the service execution module (S2001) and the converted text sent by the voice interaction application (S2002). The summary module can generate summary reply content using a large model. The detailed implementation is similar to the foregoing embodiments, and please refer to the foregoing embodiments, which will not be described here. After generating the summary reply content, the summary module can send the summary reply content to the voice interaction application (S2003) and the user interface module (S2004) respectively. The user interface module displays the summary reply content in the first display area or a second display area in the playing interface of the target video. The second display area and the first display area do not overlap. Please refer to the foregoing embodiments for detailed implementation. The voice interaction application plays the audio corresponding to the summary reply content, i.e., the target audio, and controls the target video in the playing state in the playing interface to pause playing. After the target audio is played, the voice interaction application controls the target video to continue playing. Alternatively, the voice interaction application controls the target video in the playing state in the playing interface to reduce the playing volume. After the target audio is played, the voice interaction application controls the target video to restore the playing volume. Please refer to the foregoing embodiments for detailed implementation.

[0319] In the foregoing embodiments, after generating the summary reply content, the summary module sends the summary reply content to the voice interaction application and the user interface module respectively. The user interface module can display the summary reply content, and the voice interaction application can also broadcast the corresponding audio. This combination of voice broadcasting and text display expands the dimension of information transmission and improves the efficiency of users capturing information.

[0320] In some embodiments, referring to FIG. 27, the display device can further run an audio management component. The voice interaction application can perform a broadcast process based on the summary reply content, which can include but is not limited to the following steps:

[0321] In step S2101, the voice interaction application sends an audio focus use application to the audio management component.

[0322] The audio focus (AudioFocus) is a mechanism that can be used to manage the competition for audio resources between applications. When multiple applications simultaneously request to use the audio device, the audio focus mechanism can ensure that the user's experience is not affected.

[0323] In some embodiments, after receiving the summary reply content sent by the summary induction module, the voice interaction application can first send an audio focus use application to an audio management component (AudioManager), which can carry application information such as the identifier of the voice interaction application in the audio focus use application.

[0324] In step S2102, the audio management component detects an application currently holding the audio focus as a foreground running application.

[0325] In step S2103, an audio focus use notification is returned to the voice interaction application.

[0326] In step S2104, an audio focus occupied notification is sent to the foreground running application.

[0327] In some embodiments, after receiving the audio focus use application, the audio management component (AudioManager) can detect an application currently holding the audio focus as a foreground running application. On the one hand, an audio focus use notification can be returned to the voice interaction application. On the other hand, an audio focus occupied notification can be sent to the foreground running application.

[0328] In the case where the detection result of the audio management component (AudioManager) is that the audio focus is not held, the audio management component (AudioManager) can directly return an audio focus use notification to the voice interaction application. After receiving the audio focus use notification, the voice interaction application can directly synthesize target audio based on the summary reply content and play the target audio.

[0329] In step S2105, the foreground running application pauses playing the audio and video or reduces the playing volume of the audio and video.

[0330] In some embodiments, after receiving the audio focus occupied notification, the foreground running application can pause playing the audio and video or reduce the playing volume of the audio and video. Whether to pause or reduce the playing volume can be flexibly set based on actual conditions.

[0331] In step S2106, the voice interaction application synthesizes target audio based on the summary reply content and plays the target audio.

[0332] In some embodiments, after receiving the audio focus use notification, the voice interaction application can synthesize target audio based on the summary reply content and play the target audio.

[0333] In step S2107, an audio focus release notification is sent to the audio management component after the target audio is played.

[0334] In some embodiments, after the voice interaction application finishes playing the target audio, the audio focus release notification can be sent to the audio manager (AudioManager).

[0335] In step S2108, the audio manager sends an audio focus recovery notification to the foreground running application.

[0336] In step S2109, the foreground running application continues to play the audio and video, or recovers the playing volume of the audio and video.

[0337] In some embodiments, after the audio manager (AudioManager) receives the audio focus release notification, the audio focus recovery notification can be sent to the foreground running application. After the foreground running application receives the audio focus recovery notification, the foreground running application can continue to play the audio and video, or can recover the playing volume of the audio and video. Thus, the influence of the playing of the target audio on the watching of the target video is prevented, and the user experience is further improved.

[0338] In the above embodiments, the audio focus mechanism is used to control the foreground running application to suspend playing the audio and video or to reduce the playing volume of the audio and video, so that the influence of the playing of the target audio on the watching of the target video is prevented, and the user experience is further improved.

[0339] It should be understood that, although each step in the flowchart involved in each of the above embodiments is shown in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or stages in other steps.

[0340] Each module in the display device described above can be implemented wholly or partially by software, hardware, and a combination thereof. Each module described above can be embedded in or independent of the processor in the display device in hardware form, or can be stored in the display device in software form so as to be called and executed by the processor to perform the operations corresponding to each module.

[0341] In some embodiments, a computer-readable nonvolatile storage medium is provided, which stores a computer program that is executed by a processor to implement the steps in each of the above method embodiments.

[0342] In some embodiments, a computer program product can be provided, which can include a computer program that, when executed by a processor, implements the steps of the above-mentioned method embodiments.

[0343] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0344] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, database or other medium used in the embodiments provided by the present application can include at least one of non-volatile memory and volatile memory.

[0345] The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0346] The technical features of the above embodiments can be combined in any manner. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not contradict each other, they should be considered to be within the scope of the present application.

[0347] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be pointed out that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A display device, comprising: a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, the communication protocol comprising at least a Bluetooth protocol; and at least one processor connected with the display, the memory and the one or more external device interfaces, and configured to execute the computer programs or instructions to cause the display device to: in response to a control instruction to enter a home page, the display displays a home page interface, the home page interface comprising at least one media information; in response to a voice interaction instruction, while the display displays the home page interface, the display displays text information of input voice; in a case where it is identified that the text information contains a keyword of one of a holiday, a birthday and an anniversary, at least one control is displayed in a first display area in the home page interface; the control displays integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data comprising at least detail integrated information of the recommended object and a detail page link of the recommended object; in response to an instruction to trigger the detail page link, the display is controlled to display a detail page of the recommended object to which the detail page link belongs. 2.The display device of claim 1, wherein the at least one processor is further configured to execute the computer program instructions to cause the display device to: in a case where it is identified that the text information contains a keyword of one of a holiday, a birthday and an anniversary, display reminder information of the keyword in the home page interface; the reminder information is used to perform a reminder operation at a time corresponding to the holiday or the birthday or the anniversary. 3.The display device of claim 1, wherein the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: input the text information, a current date of a system in the display device and positioning information of the display device to a multi-intent analysis model to obtain a plurality of intent-associated task information; each of the intent-associated task information comprises a predicted intent expanded based on the keyword and task arrangement information of an interactive task associated with the predicted intent; merge the predicted intents having an intent connection relationship by performing intent understanding on each of the predicted intents to obtain at least one target intent; invoke a service function corresponding to each of the interactive tasks associated with each of the target intents based on the task arrangement information of the interactive tasks to obtain service processing results of different service functions; the service processing results of the different service functions comprise at least integrated recommendation data of at least one recommended object associated with the keyword. 4.The display device of claim 3, wherein the at least one processor is configured to execute the computer programs or instructions to cause the display device to: obtain service providing information of the display device, the service providing information indicating service functions supported by the display device; determine a plurality of predicted intents expanded based on the keyword included in the text information and the service functions supported by the display device; and determine the plurality of intent-associated task information by task scheduling based on a specified date of the keyword included in the text information, a current date of a system in the display device, location information of the display device, and configuration rules of the interactive tasks associated with the predicted intents. 5.The display device of claim 3, wherein the at least one processor is configured to execute the computer programs or instructions to cause the display device to: obtain service processing results of different service functions by invoking the service functions corresponding to the interactive tasks associated with each of the target intents based on task scheduling information of the interactive tasks; and obtain service providing information of each of the service functions based on the task scheduling information of the interactive tasks associated with each of the target intents. 6.The display device of claim 1, wherein the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: control the display to display a play interface of a target video in response to a play control instruction for the target video; control the display to display an interactive result item in a first display area of the play interface of the target video in response to a voice interactive instruction, the first display area displaying at least one result item display control, each result item display control being configured to display one interactive result item, the voice interactive instruction carrying at least location information; and play a target audio corresponding to the at least one interactive result item. 7.The display device of claim 6, wherein the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: control a target video in a playing state on the play interface to pause playing, and control the target video to continue playing after the target audio is played; or control a target video in a playing state on the play interface to reduce a playing volume, and control the target video to restore the playing volume after the target audio is played. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 8.The display device of claim 6, wherein the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: display the text information corresponding to the target audio in the first display area; or display the text information corresponding to the target audio in a second display area on a playing interface of the target video; and the second display area and the first display area do not overlap. 9.The display device of claim 6, wherein the ticket query result comprises at least one selectable transportation ticket, and a result item display control for displaying the ticket query result comprises a reservation entry of each selectable transportation ticket; and the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: in response to a triggering operation on the reservation entry of a target transportation ticket in the at least one selectable transportation ticket, control the display to display a reservation page of the target transportation ticket; and in response to a ticket reservation operation triggered on the reservation page, complete reservation of the target transportation ticket. 10.The display device of claim 9, wherein the ticket query result is divided into a trip ticket query result and a return ticket query result; the trip ticket query result comprises at least one selectable trip ticket, and a result item display control for displaying the trip ticket query result comprises a reservation entry of each selectable trip ticket; the return ticket query result comprises at least one selectable return ticket, and a result item display control for displaying the return ticket query result comprises a reservation entry of each selectable return ticket; and the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: in response to a triggering operation on the reservation entry of a target trip ticket in the at least one selectable trip ticket, control the display to display a reservation page of the target trip ticket; in response to a ticket reservation operation triggered on the reservation page of the target trip ticket, complete reservation of the target trip ticket; in response to a triggering operation on the reservation entry of a target return ticket in the at least one selectable return ticket, control the display to display a reservation page of the target return ticket; and in response to a ticket reservation operation triggered on the reservation page of the target return ticket, complete reservation of the target return ticket. 11.The display device of claim 6, wherein the weather query result comprises weather information of at least one day in the past, and a result item display control for displaying the weather query result comprises a detail viewing entry of each piece of weather information; and the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: in response to a triggering operation on the detail viewing entry of a target piece of weather information, control the display to display a detail display page corresponding to the target piece of weather information. 12.The display device of claim 11, wherein the voice interaction instruction further carries trip date information, and the weather query result comprises weather information of a date determined according to the trip date information. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 13. The display device of claim 6, wherein the accommodation query result comprises at least one selectable accommodation site, and the result item display control for displaying the accommodation query result comprises a reservation entry for each selectable accommodation site; and the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: in response to a triggering operation on the reservation entry for a target accommodation site among the at least one selectable accommodation site, control the display to display a reservation page of the target accommodation site; and in response to a reservation operation triggered on the reservation page, complete the reservation of the target accommodation site.

14. The display device of claim 6, wherein the at least one result item display control is displayed in a list, and the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: in response to an information-down scrolling instruction, display a lower control of the currently displayed result item display control in the first display area; and in response to an information-up scrolling instruction, display an upper control of the currently displayed result item display control in the first display area.

15. The display device of claim 6, wherein the result item display control for displaying the ticket query result further comprises a travel tool recommendation area for displaying travel tool recommendation information summarized based on the ticket query result, and the result item display control for displaying the weather query result further comprises a weather reminder area for displaying weather-related reminder information summarized based on the weather query result.

16. The display device of claim 6, wherein the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: send the reminder formulation result to a mobile terminal through the Internet of Things to instruct the mobile terminal to provide a reminder based on the reminder formulation result.

17. The display device of claim 6, wherein the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: convert an interactive voice input by a user into text; in a case where the text conforms to a target service scenario, obtain a text splitting requirement corresponding to the target service scenario; the text splitting requirement being information set in advance for the target service scenario to guide a large model to split text; based on the text and the text splitting requirement, split the text into a plurality of task sub-sentences using the large model; generate a voice interaction instruction based on the plurality of task sub-sentences; and in response to the voice interaction instruction, obtain a task execution result corresponding to each task sub-sentence as an interaction result item.

18. The display device of claim 17, wherein the at least one processor is further configured to execute the computer programs or instructions to cause the display device to: in a case where the text contains a location word, determine that the text conforms to a target service scenario. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 19. The display device of claim 17, wherein the at least one processor, based on the text and the text splitting requirement, utilizes the large model to split into a plurality of task sentences, is further configured to execute the computer programs or instructions to cause the display device to: obtain processing reference information corresponding to the target service scenario; extract processing keywords based on the text and the processing reference information; construct model input information based on the processing keywords; assemble the model input information and the text splitting requirement according to the prompt word template of the large model to obtain assembly information; input the assembly information into the large model to obtain a plurality of task sentences.

20. The display device of claim 17, wherein the at least one processor obtains a task execution result corresponding to each task sentence, is further configured to execute the computer programs or instructions to cause the display device to: for each task sentence, perform intent understanding on the current task sentence to obtain a task intent corresponding to the current task sentence, and based on the task intent corresponding to the current task sentence, call a corresponding business microservice to obtain a task execution result of the current task sentence.

21. A display device comprising: a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, the communication protocol including at least a Bluetooth protocol; and at least one processor connected with the display, the memory and the one or more external device interfaces, and configured to execute the computer programs or instructions to cause the display device to: in response to a control instruction of a video play, the display displays a video play interface; in response to a voice interaction instruction, the display pauses playing a video picture or reduces a playing volume of the video play interface, and the display displays text information of an input voice; in a case where it is identified that the text information contains one of a keyword of a festival, a birthday and an anniversary, at least one control is displayed in a second display area in the video play interface; the control displays integrated recommendation data of at least one recommended object associated with the keyword, the integrated recommendation data including at least detail integrated information of the recommended object and a detail page link of the recommended object; in response to an instruction of triggering the detail page link, the display is controlled to display a detail page of the recommended object to which the detail page link belongs.

22. A display device comprising: a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces configured to communicate with one or more external devices according to a communication protocol, the communication protocol including at least a Bluetooth protocol; ​ and at least one processor interfacing with the display, the memory, and the one or more external devices and configured to execute the computer programs or instructions to cause the display device: In the case of displaying a home page interface or a video playing interface, the display displays an interactive dialogue interface in response to a voice interaction instruction; the interactive dialogue interface displays at least text information of input voice; In the case of identifying that the text information contains a keyword of one of a festival, a birthday, and an anniversary, the display displays at least one control in a third display area in the interactive dialogue interface; The control displays integrated recommendation data of at least one recommended object associated with the keyword, and the integrated recommendation data at least includes detail integrated information of the recommended object and a detail page link of the recommended object; In response to an instruction of triggering the detail page link, the display displays a detail page of a recommended object to which the detail page link belongs.

23. A voice interaction method for a display device, the display device comprising a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces that can be configured to communicate with one or more external devices according to a communication protocol, the communication protocol at least including a Bluetooth protocol; and at least one processor interfacing with the display, the memory, and the one or more external device interfaces and configured to execute the computer programs or instructions to control the display device; the method comprises: In response to a control instruction of entering a home page, the display displays a home page interface, and the home page interface at least includes media information; In response to a voice interaction instruction, the display displays text information of input voice while displaying the home page interface; In the case of identifying that the text information contains a keyword of one of a festival, a birthday, and an anniversary, the display displays at least one control in a first display area in the home page interface; The control displays integrated recommendation data of at least one recommended object associated with the keyword, and the integrated recommendation data at least includes detail integrated information of the recommended object and a detail page link of the recommended object; In response to an instruction of triggering the detail page link, the display displays a detail page of a recommended object to which the detail page link belongs.

24. A voice interaction method for a display device, the display device comprising a display configured to display content from a broadcast system or network and / or a user interface; a memory configured to store computer programs or instructions; one or more external device interfaces that can be configured to communicate with one or more external devices according to a communication protocol, the communication protocol at least including a Bluetooth protocol; and at least one processor interfacing with the display, the memory, and the one or more external device interfaces and configured to execute the computer programs or instructions to control the display device; the method comprises: In response to a control instruction of entering a home page, the display displays a home page interface, and the home page interface at least includes media information; In response to a voice interaction instruction, the display displays text information of input voice while displaying the home page interface; In the case of identifying that the text information contains a keyword of one of a festival, a birthday, and an anniversary, the display displays at least one control in a first display area in the home page interface; The control displays integrated recommendation data of at least one recommended object associated with the keyword, and the integrated recommendation data at least includes detail integrated information of the recommended object and a detail page link of the recommended object; In response to an instruction of triggering the detail page link, the display displays a detail page of a recommended object to which the detail page link belongs. In a case where it is identified that the text information contains one of keywords of a festival, a birthday, and an anniversary, at least one control is displayed in a second display area in the video playing interface; The control displays integrated recommendation data of at least one recommended object associated with the keyword, and the integrated recommendation data at least includes detail integrated information of the recommended object and a detail page link of the recommended object; In response to an instruction of triggering the detail page link, the display is controlled to display a detail page of the recommended object to which the detail page link belongs.

25. A voice interaction method for a display device, the display device comprising a display configured to display content from a broadcast system or network and / or a user interface; The memory is configured to store computer programs or instructions; the one or more external device interfaces can be configured to communicate with one or more external devices according to a communication protocol, and the communication protocol at least includes a Bluetooth protocol; and at least one processor is connected with the display, the memory, and the one or more external device interfaces, and is configured to execute the computer programs or instructions to control the display device; the method includes: In a case where a home page interface is displayed or a video playing interface is displayed, in response to a voice interaction instruction, the display displays an interaction dialogue interface; the interaction dialogue interface at least displays text information of input voice; In a case where it is identified that the text information contains one of keywords of a festival, a birthday, and an anniversary, at least one control is displayed in a third display area in the interaction dialogue interface; The control displays integrated recommendation data of at least one recommended object associated with the keyword, and the integrated recommendation data at least includes detail integrated information of the recommended object and a detail page link of the recommended object; In response to an instruction of triggering the detail page link, the display is controlled to display a detail page of the recommended object to which the detail page link belongs.

Citation Information

Patent Citations

  • Festival notification method, system and device based on calendar and storage medium

    CN110008400A

  • Route display device for national culture tourism planning

    CN114822330A

  • Route planning method and device based on large language model, equipment and storage medium

    CN117556155A

  • Travel management method, travel management model training method, travel management model training device, travel management equipment and storage medium

    CN117634699A

  • Method and system to curate media collections

    US20140143247A1