A display device and a display device control method

By combining large models with text and image recognition technologies to generate intelligent path planning operation instructions, the problem of insufficient intelligent operability of traditional display devices is solved, enabling accurate control of third-party applications and content search, thus improving the user experience.

CN119815090BActive Publication Date: 2026-04-24HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HISENSE VISUAL TECH CO LTD
Filing Date
2024-11-11
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional display devices lack intelligent operability, making it difficult to meet users' complex and non-preset command needs, especially in the case of voice control of third-party applications, where intelligent operation of the internal interface cannot be achieved.

Method used

By receiving user-input interaction commands, leveraging the decision-making and planning capabilities of large models, and combining text recognition and image recognition technologies, intelligent path planning operation commands are generated, and the sequence of operation actions is executed with each output until the user interaction command is completed.

Benefits of technology

It improves the accuracy and rationality of operation commands, achieves matching with user interaction commands, and meets the intelligent control needs of display devices, especially improving recognition speed and accuracy when opening third-party applications and searching for content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119815090B_ABST
    Figure CN119815090B_ABST
Patent Text Reader

Abstract

The application discloses a display device and a display device control method. The method receives an interactive instruction input by a user, and obtains interactive text corresponding to the interactive instruction. For a process of outputting an operation instruction corresponding to the interactive text for the first time, at least the interactive text and an operation specification prompt for guiding a large model to output the operation instruction are taken as inputs of the large model, and at least a first operation instruction for operating a demonstration function is output. For a process of outputting an operation instruction corresponding to the interactive text for the first time, at least the interactive text, content output by the large model last time, and the operation specification prompt are input to the large model, and at least a current operation instruction is output. In each case of outputting the operation instruction, the output operation instruction is executed to complete an execution action sequence of operating the demonstration function, so that a corresponding operation function is realized, and the execution of the operation instruction output for the last time is completed. The method can improve intelligent control of the display device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and in particular to a display device and a display device control method. Background Technology

[0002] With the rapid development of display device functionality, the features they offer users are becoming increasingly diverse. Currently, display devices include smart TVs, smart set-top boxes, smart set-top boxes, and products with smart display screens. Taking smart TVs as an example, they can be intelligently controlled by receiving user voice commands.

[0003] However, the voice control currently supported on smart TVs can only perform some very simple operations, such as turning the TV on and off and adjusting the volume. Furthermore, the functions supported by voice control are all pre-set or pre-planned, and cannot complete complex user commands or non-preset commands, resulting in insufficient intelligent operability and difficulty in meeting user needs. Summary of the Invention

[0004] This application provides a display device and a display device control method to solve the problem that traditional display devices lack intelligent operability and are difficult to meet user needs.

[0005] In a first aspect, some embodiments provide a display device, including: a display and at least one processor. The display is configured to display content from a broadcast system or network and / or a user interface; and the at least one processor is configured to:

[0006] Receive user input of interactive commands and obtain the corresponding interactive text;

[0007] For the process of outputting the operation instructions corresponding to the interactive text for the first time, at least the interactive text and the operation specification prompts used to guide the large model to output operation instructions should be used as inputs to the large model, and at least the first operation instruction to operate the demonstration function should be output.

[0008] For the process of outputting operation instructions corresponding to interactive text for the first time, at least the interactive text, the content of the previous output of the large model, and the operation specification prompts should be input into the large model, and at least the current operation instructions should be output.

[0009] Each time an operation command is output, the output operation command is executed to complete the sequence of actions to operate the demonstration function, so as to realize the corresponding operation function, until the last output operation command is executed.

[0010] Technical Effects: By receiving user-input interactive commands, the corresponding interactive text and operational guidelines for guiding the large model to output operation commands serve as input to the large model. The large model's decision-making and planning capabilities enable intelligent path planning, which includes the operation commands for each demonstration function output by the large model. For subsequent outputs of interactive text-based operation commands, the large model's input also includes the content of its previous output, ensuring that the large model can decide on the current output operation command based on completed content, thus improving the accuracy and rationality of each output operation command. In each output operation command scenario, the output operation command is executed to complete the sequence of actions for operating the demonstration function, achieving the corresponding operation function, until the last output operation command is executed. This achieves intelligent control that matches user interactive commands, effectively meeting users' needs for intelligent control of display devices.

[0011] In some embodiments, the processor executes the output operation instructions to complete the sequence of actions for operating the demonstration function, configured as follows:

[0012] When the operation command is to open a third-party application, if the currently displayed page is a summary page of application icons, text recognition is performed on the corresponding page image of the currently displayed page to obtain the location of the application name of each application.

[0013] Match the application name of the third-party application to be opened with the application name on the currently displayed page to obtain the target position of the application name of the third-party application to be opened on the currently displayed page.

[0014] Open a third-party application based on the target location.

[0015] Technical effect: When the operation command is to open a third-party application, the method of combining text recognition and image recognition is used to identify the location of the application name of each application. Compared with the traditional method of using only image recognition, the recognition speed is improved, power consumption is reduced, and the accuracy of target location is improved.

[0016] In some embodiments, the process of the processor obtaining the application name of the third-party application to be opened is configured as follows:

[0017] Using a large model, extract the names of the third-party applications that need to be opened from the interactive text;

[0018] The app name will be transformed into a generic app name for the third-party app that needs to be opened.

[0019] Based on the correspondence between the general application name and the application name of the third-party application to be opened after installation on the display device, the application name of the third-party application to be opened is obtained.

[0020] Technical effect: By extracting the reference application name of the third-party application to be opened from the interactive text through a large model, and then determining the general application name corresponding to the reference application name, the application name that is displayed after installation on the display device and corresponds to the general application name is used as the application name of the third-party application to be opened. Through multi-level transformation, the application name in the user's interactive command is converted into an application name that the display device can recognize and process, thereby improving the intelligent controllability and control accuracy of the display device.

[0021] In some embodiments, the processor executes the output operation instructions to complete the sequence of actions for operating the demonstration function, configured as follows:

[0022] When the operation command is to input text, if a virtual keyboard is displayed on the current display page, then text recognition is performed on the corresponding page image of the current display page to obtain the position of each key on the virtual keyboard.

[0023] Obtain the character sequence corresponding to the content to be searched, match each character in the character sequence with the characters displayed on the keys, and obtain the target position of each character in the character sequence on the virtual keyboard;

[0024] Based on the target location, the corresponding buttons are triggered sequentially for content search.

[0025] Technical effect: When the operation command is input text, the position of each key on the virtual keyboard is obtained by performing text recognition on the corresponding page image of the currently displayed page to ensure the accuracy of key triggering. Each character in the character sequence corresponding to the content to be searched is matched with the character displayed on the key to obtain the target position of each character in the character sequence on the virtual keyboard. In this way, it can ensure that the content matching the search content is found, which helps to improve the accuracy of content search.

[0026] In some embodiments, the processor is further configured to:

[0027] After the current output operation command is executed, compare the page image before execution with the page image after execution;

[0028] If the two are inconsistent, the currently output operation instruction is executed successfully and execution continues;

[0029] If the two are inconsistent, the current output operation command will fail to execute. At the very least, the interactive text, the content of the last output of the large model, and the operation specification prompts should be input into the large model, and the current operation command should be re-output and executed.

[0030] Technical effect: After the current output operation command is executed, the page images before and after the execution are compared. If they are consistent, it indicates that the execution was successful and subsequent steps can be continued. If they are inconsistent, it indicates that the execution failed and the output steps of the current operation command will be re-executed. This abnormal state handling method ensures that exception handling can be performed automatically when execution fails, ensuring the smooth operation of the control method and improving the user experience.

[0031] In some embodiments, the input content of the current processing of the large model also includes the page image corresponding to the currently displayed page; the processor is further configured to: complete the sequence of actions to operate the demonstration function before executing the output current operation instructions.

[0032] If the large model identifies that the currently displayed page contains recommended content, skip the recommended content or wait for the recommended content to finish being displayed.

[0033] Technical effect: Identifying recommended content on the currently displayed page, skipping recommended content, or waiting for the recommended content to finish displaying can avoid interference from recommended content in the execution of actions, which helps to improve the success rate of action execution.

[0034] In some embodiments, the large model also outputs at least one of the action description text of the output operation instruction or the description text of the target of the action to be performed.

[0035] Technical Effects: By using at least one of the action description text or the description text of the target of the action to be performed as the output of the large model, the action description text can describe the action of the operation instruction, and each action description text also includes the relevant description content of the previous operation instruction; the description text can describe the target of the operation instruction, and each description text of the executed action also includes the relevant target description content of the previous executed action. This helps the large model to accurately grasp at least one of the actions performed by each operation instruction or the target of the executed action, thereby improving the decision-making accuracy of the large model for the next output operation instruction.

[0036] In some embodiments, for the process of outputting operation instructions corresponding to interactive text for the first time, the input content of the current processing of the large model also includes summary text of the completed actions.

[0037] Technical effect: The summary text includes a summary description of all actions completed before the current processing. In this way, the large model can output the current operation instruction based on the summary text of the completed actions, which helps to improve the decision-making accuracy of the large model.

[0038] Secondly, some embodiments also provide a display device control method applied to the display device provided in the first aspect, the display device including: a display and at least one processor, the method including:

[0039] Receive user input of interactive commands and obtain the corresponding interactive text;

[0040] For the process of outputting the operation instructions corresponding to the interactive text for the first time, at least the interactive text and the operation specification prompts used to guide the large model to output operation instructions should be used as inputs to the large model, and at least the first operation instruction to operate the demonstration function should be output.

[0041] For the process of outputting operation instructions corresponding to interactive text for the first time, at least the interactive text, the content of the previous output of the large model, and the operation specification prompts should be input into the large model, and at least the current operation instructions should be output.

[0042] Each time an operation command is output, the output operation command is executed to complete the sequence of actions to operate the demonstration function, so as to realize the corresponding operation function of the operation command, until the last output operation command is executed.

[0043] Technical Effects: By receiving user-input interactive commands, the corresponding interactive text and operational guidelines for guiding the large model to output operation commands serve as input to the large model. The large model's decision-making and planning capabilities enable intelligent path planning, which includes the operation commands for each demonstration function output by the large model. For subsequent outputs of interactive text-based operation commands, the large model's input also includes the content of its previous output, ensuring that the large model can decide on the current output operation command based on completed content, thus improving the accuracy and rationality of each output operation command. In each output operation command scenario, the output operation command is executed to complete the sequence of actions for operating the demonstration function, achieving the corresponding operation function, until the last output operation command is executed. This achieves intelligent control that matches user interactive commands, effectively meeting users' needs for intelligent control of display devices.

[0044] In some embodiments, executing the output operation instructions to complete a sequence of actions to operate the demonstration function includes:

[0045] When the operation command is to open a third-party application, if the currently displayed page is a summary page of application icons, text recognition is performed on the corresponding page image of the currently displayed page to obtain the location of the application name of each application.

[0046] Match the application name of the third-party application to be opened with the application name on the currently displayed page to obtain the target position of the application name of the third-party application to be opened on the currently displayed page.

[0047] Open a third-party application based on the target location.

[0048] Technical effect: When the operation command is to open a third-party application, the method of combining text recognition and image recognition is used to identify the location of the application name of each application. Compared with the traditional method of using only image recognition, the recognition speed is improved, power consumption is reduced, and the accuracy of target location is improved. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application;

[0051] Figure 2 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application;

[0052] Figure 3 This is a schematic diagram of the hardware configuration of the control device provided in some embodiments of this application;

[0053] Figure 4 This is a schematic diagram of the software configuration of a display device provided in some embodiments of this application;

[0054] Figure 5 A schematic flowchart illustrating a display device control method provided in some embodiments of this application;

[0055] Figure 6 A flowchart illustrating the decision-making process of a large model provided for some embodiments of this application;

[0056] Figure 7 A schematic diagram of an application icon summary page provided for some embodiments of this application;

[0057] Figure 8This is a schematic diagram of the processing flow of the OCR service provided in some embodiments of this application;

[0058] Figure 9 Illustration of display pages provided for some embodiments of this application Figure 1 ;

[0059] Figure 10 Illustration of display pages provided for some embodiments of this application Figure 2 ;

[0060] Figure 11 A schematic diagram of a page image before execution completion provided in some embodiments of this application;

[0061] Figure 12 A schematic diagram of a page image after execution is completed, provided for some embodiments of this application;

[0062] Figure 13 This is a schematic diagram of the processing flow for the result feedback service provided in some embodiments of this application;

[0063] Figure 14 A schematic diagram illustrating the processing of recommended content provided in some embodiments of this application;

[0064] Figure 15 This is a schematic diagram of the system framework of a display device provided in some embodiments of this application;

[0065] Figure 16 A schematic flowchart illustrating the overall process of a display control method provided in some embodiments of this application;

[0066] Figure 17 This is an internal structural diagram of a computer device provided in some embodiments of this application. Detailed Implementation

[0067] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0068] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0069] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0070] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0071] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0072] In this embodiment, the display device 200 generally refers to a device with screen display and data processing capabilities. For example, the display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0073] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, users can operate the display device 200 via touch operation, mobile terminal 300, and control device 100. For example, control device 100 can be a remote control, stylus, gamepad, etc.

[0074] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the display device 200. It can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the display device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the display device 200 for synchronized display.

[0075] like Figure 1 The diagram also shows that the display device 200 communicates with the server 400 via various communication methods. This allows the display device 200 to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0076] Display device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.

[0077] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of display device 200.

[0078] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0079] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0080] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0081] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The display device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the display device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the display device 200 supports Bluetooth connectivity, it needs to have a communication device 220 with Bluetooth functionality.

[0082] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0083] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0084] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0085] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0086] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0087] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0088] Figure 3 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of the central control device. (Example) Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0089] The control device 100 is configured to control the display device 200, and to receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.

[0090] In some embodiments, the control device 100 may be an intelligent device. For example, the control device 100 may be equipped with various applications for controlling the display device 200 according to user needs.

[0091] In some embodiments, such as Figure 1As shown, the mobile terminal 300 or other smart electronic devices can perform similar functions to the control device 100 after installing the application of the control display device 200.

[0092] The controller 110 includes a processor 112, RAM 113, ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation of the control device 100, as well as the communication and cooperation between internal components and the external and internal data processing functions.

[0093] Under the control of the controller 110, the communication interface 130 enables communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of other near-field communication modules such as WiFi chip 131, Bluetooth module 132, and NFC module 133.

[0094] User input / output interface 140, wherein the input interface includes at least one of other input interfaces such as microphone 141, touchpad 142, sensor 143, and button 144.

[0095] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, which can encode user input commands via WiFi, Bluetooth, or NFC protocols and send them to the display device 200.

[0096] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can also store various control signal instructions input by the user.

[0097] The power supply 180 is used to provide operating power support for the various components of the control device 100 under the control of the controller.

[0098] In some embodiments, to enable user interaction, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can (control the display device) provide a user interface, allowing users to interact with the display device 200 and supporting the running of various applications.

[0099] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for display devices.

[0100] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Library layer, and the Kernel layer.

[0101] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0102] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0103] like Figure 4 As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0104] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0105] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0106] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 4 As shown, hardware drivers can be configured in the kernel layer. The kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0107] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.

[0108] With the rapid development of display device functionality, the features they offer users are becoming increasingly diverse. Currently, display devices include smart TVs, smart set-top boxes, smart set-top boxes, and products with smart display screens. Taking smart TVs as an example, they can be intelligently controlled by receiving user voice commands.

[0109] However, the voice control currently supported on smart TVs can only perform some very simple operations, such as powering on / off and adjusting volume. Furthermore, the functions supported by voice control are all pre-set or pre-planned, and cannot fulfill complex user commands or non-preset commands. For example, voice control of third-party applications is limited to opening or closing, and cannot control the internal interface of third-party applications. Therefore, traditional display devices suffer from insufficient intelligent operability and cannot meet user needs.

[0110] To address the issue of insufficient intelligent operability in traditional display devices, which makes it difficult to meet user needs, this application provides a display device control method. This method receives interactive commands input by the user. The corresponding interactive text and operational guidelines for guiding a large model to output operation commands serve as input to the large model. The method utilizes the large model's decision-making and planning capabilities to achieve intelligent path planning. The intelligent path planning includes operation commands for each demonstration function output by the large model. For operation commands output after the initial interaction text, the large model's input also includes the content of its previous output, ensuring that the large model can decide on the current output operation command based on completed content, thus improving the accuracy and rationality of each output operation command. In each output operation command, the output operation command is executed to complete the sequence of actions for operating the demonstration function, achieving the corresponding operation function, until the last output operation command is executed. This achieves intelligent control that matches the user's interactive commands, effectively meeting the user's intelligent control needs for the display device.

[0111] like Figure 5 The diagram illustrates a flowchart of a display device control method provided in some embodiments. This method is applied to the aforementioned display device, which includes a display and at least one processor. The method includes:

[0112] S100: Receive the user's input interaction command and obtain the corresponding interaction text.

[0113] Interaction commands refer to user-inputted instructions that control the display device. Interaction commands can be in the form of voice, text, etc. For example, the processor can acquire user-inputted interaction commands through a control device or through a built-in voice receiving module. Based on the controlled object, interaction commands can include commands controlling local applications, commands controlling system settings, and commands controlling third-party applications. For example, a command controlling a local application might be "Please search for TV series A," a command controlling system settings might be "Please turn up the volume," and a command controlling a third-party application might be "Please open application B to search for TV series C," etc.

[0114] Interactive text refers to the text corresponding to an interactive instruction. If the interactive instruction is a voice instruction, the processor can use a speech recognition algorithm to recognize the interactive instruction and obtain the interactive text; if the interactive instruction is a text instruction, the processor can use the text in the text instruction or the text keywords in the text instruction as the interactive text.

[0115] S200. For the process of outputting the operation instructions corresponding to the interactive text for the first time, at least the interactive text and the operation specification prompts used to guide the large model to output operation instructions shall be used as inputs to the large model, and at least the first operation instruction to operate the demonstration function shall be output.

[0116] Large models refer to machine learning models with a large number of parameters and complex computational structures. Given a prompt word as a condition, a large model can automatically generate corresponding output results for the interactive text; therefore, large models possess a certain degree of decision-making and planning capabilities.

[0117] Demonstration functions refer to features displayed on a display device, including audio and video playback control, application switching, system settings, and content search. User-inputted interactive commands are used to operate at least one demonstration function.

[0118] Operation instructions refer to the commands used to operate the demonstration functions. Each operation instruction performs at least one action to operate the demonstration function. For example, operation instructions include application home, my applications, open application, click, swipe, input, home, and stop. When the operation instruction is to open an application, the corresponding at least one action may include exiting the current application, obtaining the name of the application to be opened, and opening the application corresponding to the application name.

[0119] In scenarios where intelligent path planning is based on interactive commands, the large model can output operation commands for the demonstration function based on the input interactive text and operation specification prompts. These operation specification prompts refer to pre-set prompts in the processor to guide the large model in outputting operation commands. For example, an operation specification prompt might be: "In the user command, 'homepage' specifically refers to the homepage of the local application; if the user command requires operation on system settings, use the action 'system settings'; if you need to return to the homepage of the local application, use the action 'homepage'; if you need to return to the homepage of a non-local application, use the action 'application homepage'." It's important to understand that the operation specification prompts are not limited to the above. Operation specification prompts can be flexibly set according to the actual operation specifications of the display device. Due to the complexity of functions in the display device and the diversity of applications, there can be multiple operation commands for the demonstration function. Therefore, using operation specification prompts as input for the large model helps guide the large model to output operation commands that conform to the operation specifications of the display device.

[0120] In some embodiments, the interaction instructions are relatively complex, and the large model needs to output at least one operation instruction to complete the interaction. For example, if the interaction instruction is "Please open application B to search for TV series C", the operation instructions output by the large model should at least include opening the application, clicking, and inputting, and these operation instructions should have a sequential execution order to smoothly complete the user's needs indicated by the interaction instruction. In the display device control method proposed in this application embodiment, the large model can output a single operation instruction at a time, and multiple operation instructions can be obtained through multiple outputs by the large model.

[0121] In some embodiments, the processor takes interactive text and operation instructions as input to the large model, and the operation instructions output by the large model are used as the first operation instructions. For example, if the interactive instruction is "Please open application B to search for TV series C", the first operation instruction output by the large model could be "Open application".

[0122] S300. For the process of outputting the corresponding operation instructions for the interactive text that is not the first time, at least the interactive text, the content of the previous output of the large model, and the operation specification prompts shall be input into the large model, and at least the current operation instructions shall be output.

[0123] The previous output of the large model includes the operation instructions from the previous output. For example, the previous output of the large model could be: Step-1: Operation instruction: Open application (B).

[0124] For processes involving outputting interactive text and corresponding operation instructions that are not the first time, the input to the large model includes the interactive text, the content of the previous output by the large model, and operation specification prompts. Introducing the operation instructions from the previous output helps the large model clearly understand the historical operation process, thus ensuring that the order of the operation instructions output by the large model conforms to the actual operation sequence requirements corresponding to the interactive instructions. For example, if the interactive instruction is "Please open application B to search for TV series C", the first output operation instruction could be "open application" (B), the second output operation instruction could be "click" (search box), and the third output operation instruction could be "input" (TV series C). Figure 6 The diagram shown illustrates the large model decision-making process provided in some embodiments of this application. The input to the large model includes at least interactive text, the content of the previous output of the large model, and operation specification prompts. The content of the large model's decision output includes at least operation instructions.

[0125] S400. Each time an operation instruction is output, the output operation instruction is executed to complete the sequence of actions to operate the demonstration function, so as to realize the corresponding operation function, until the last output operation instruction is completed.

[0126] In this context, an action sequence refers to a series of sequential actions performed to operate a demonstration function. An action sequence can also consist of a single action. Since executing an operation command to operate a demonstration function often requires completing at least one action, and these actions have a specific execution order, at least one action can constitute an action sequence. Executing each operation command requires completing the action sequence for the corresponding demonstration function. For example, if the operation command is "Open application (B)," the required action sequence could include exiting the current application, searching for application B, and clicking the application icon of application B.

[0127] Each time the large model outputs an operation instruction, the processor executes the sequence of actions to perform the demonstration function according to the output instructions, thus realizing the corresponding operation function, until the last operation instruction output by the large model is completed. In this way, intelligent control of interactive commands is achieved through the execution of multiple operation instructions.

[0128] The aforementioned display device and control method receive user-input interactive commands. The corresponding interactive text and operational guidelines for guiding the large model to output operation commands serve as input to the large model. The large model's decision-making and planning capabilities enable intelligent path planning. This intelligent path planning includes operation commands for each demonstration function output by the large model. For subsequent outputs of interactive text-based operation commands, the large model's input also includes the content of its previous output, ensuring that the large model can determine the current output operation command based on completed content, thus improving the accuracy and rationality of each output operation command. In each output operation command scenario, the output operation command is executed to complete the sequence of actions for operating the demonstration function, achieving the corresponding operation function, until the last output operation command is executed. This achieves intelligent control that matches user interactive commands, effectively meeting users' intelligent control needs for the display device.

[0129] In some embodiments, executing the output operation instructions to complete the sequence of actions to operate the demonstration function includes: when the operation instruction is to open a third-party application, if the currently displayed page is an application icon summary page, performing text recognition on the corresponding page image of the currently displayed page to obtain the location of the application name of each application; matching the application name of the third-party application to be opened with the application names in the currently displayed page to obtain the target location of the application name of the third-party application to be opened in the current displayed page; and opening the third-party application based on the target location.

[0130] The applications installed on the display device include native applications and third-party applications. Native applications refer to the applications that come pre-installed on the display device, while third-party applications refer to applications other than the native applications installed on the display device.

[0131] In some embodiments, the display device is currently displaying application A, and the user's interaction command instructs the user to open application B. Since application B is a third-party application, the operation command generated by the large model can be "Open application (application B)," which means opening the third-party application. Application A can be a native application or a third-party application different from application B. For example, determining that the operation command is for a third-party application can be done by first determining whether the application in the interaction text matches the currently displayed application. If they don't match, then further determining whether the application in the interaction text is a native application or a third-party application. If it is a third-party application, then determining that the operation command is to open the third-party application, and the application to be opened is the application in the interaction text.

[0132] The app icon summary page is a page that compiles the app icons and names of all third-party applications. The app icon summary page can consist of one or more pages. For example... Figure 7 The diagram shown is a schematic of an application icon summary page provided in some embodiments of this application. Some applications have application icons on the application icon summary page, while others have both an application icon and an application name.

[0133] The corresponding page image for the currently displayed page can be a screenshot of the currently displayed page. The third-party application to be opened refers to the third-party application that the operation command indicates to be opened. Since the operation command is based on the interaction command, the third-party application to be opened is the third-party application that the interaction command indicates to be opened.

[0134] The sequence of execution actions required to operate the demonstration function can include at least one execution action. For example, if the currently displayed page is an application icon summary page, at least one execution action includes: performing text recognition on the corresponding page image of the currently displayed page to obtain the location of the application name of each application; matching the application name of the third-party application to be opened with the application names on the currently displayed page to obtain the target location of the application name of the third-party application to be opened on the currently displayed page; and opening the third-party application based on the target location. Since the processor needs to click on the application icon of the corresponding application on the application icon summary page to open the third-party application, the processor needs to determine the accurate location of each application icon. In some embodiments, the processor can use an image recognition method to perform image recognition on the corresponding page image of the currently displayed page to obtain the location of the application icon of each application. However, since the application icons may contain content that changes with the recommendation information, the location of the application icons obtained by the image recognition method may be inaccurate. Since there are applications with both an application icon and an application name on the application icon summary page, this application embodiment proposes to perform text recognition on the corresponding page image of the currently displayed page to obtain the location of the application name of each application. For example, using an OCR text recognition algorithm for text recognition is beneficial to improving the accuracy and recognition speed of the obtained application name location.

[0135] For example, the location identified using the OCR text recognition algorithm is shown below:

[0136] {word=Homepage,PosBeanLists=[PosBean{x=143,y=51},PosBean{x=183,y=51},

[0137] PosBean{x=183,y=70},PosBean{x=143,y=70}},

[0138] {word=TV series,PosBeanLists=[PosBean{x=202,y=52},PosBean{x=248,y=52},

[0139] PosBean{x=248,y=69},PosBean{x=202,y=69},]},

[0140] {word=movie,PosBeanLists=[PosBean{x=270,y=53},

[0141] PosBean{x=305,y=53},PosBean{x=305,y=69},PosBean{x=270,y=69},]},

[0142] {word=variety show,PosBeanLists=[PosBean{x=324,y=52},

[0143] PosBean{x=359,y=52},PosBean{x=359,y=69},PosBean{x=324,y=69},]}.

[0144] like Figure 8 The diagram illustrates the processing flow of the OCR service provided in some embodiments of this application. The processor can use the OCR service to identify the location of the application name of each application on the currently displayed page, match the application name of the third-party application to be opened with the application names on the currently displayed page, and use the location of the successfully matched application name as the target location. In some embodiments, image recognition is performed on the corresponding page image of the currently displayed page to obtain the location of the application icon for each application. This ensures that for applications with application names, text recognition is used, guaranteeing the accuracy and speed of the target location; for applications without application names, image recognition is used, further improving the accuracy of the target location.

[0145] In some embodiments, after obtaining the location of the application icon of each application on the currently displayed page, the processor matches the application name of the third-party application to be opened with the application name corresponding to the application icon on the currently displayed page, and takes the location of the successfully matched application icon as the target location.

[0146] In some embodiments, the cloud server pre-stores the positions of multiple application icons in the application icon summary page. If there is an application icon position that matches the third-party application to be opened among the multiple applications pre-stored in the cloud server, the position of the successfully matched application icon is taken as the target position. If there is no application icon position that matches the third-party application to be opened among the multiple applications pre-stored in the cloud server, the target position is determined by OCR text recognition and image recognition methods.

[0147] The processor opens a third-party application based on the target location. For example, the action performed is clicking (the target location), and after the action is completed, the third-party application is launched. Alternatively, if the target location is the location of the application name, the processor determines the location of the corresponding application icon based on the target location, performs the action of clicking (the location of the application icon corresponding to the target location), and launches the third-party application after the action is completed.

[0148] In this embodiment, when the operation instruction is to open a third-party application, a method combining text recognition and image recognition is used to identify the location of the application name of each application. Compared with the traditional method that only uses image recognition, this method improves the recognition speed, reduces power consumption, and improves the accuracy of the target location.

[0149] In some embodiments, the process of obtaining the application name of the third-party application to be opened includes: extracting the referential application name of the third-party application to be opened from the interactive text through a large model; converting the referential application name into a generic application name of the third-party application to be opened; and obtaining the application name of the third-party application to be opened based on the correspondence between the generic application name and the application name presented after the third-party application to be opened is installed on the display device.

[0150] Because the application name on the display device may differ from the application name in the user's interaction command, it is necessary to convert the application name in the interaction command so that the processor can open the accurate application. For example, Tencent Video's application name on the display device is Cloud Video Aurora, etc.

[0151] The referential application name refers to the application name of the third-party application that needs to be opened in the interactive text obtained using the large model. For example, the referential application name obtained using the large model is Youku Video.

[0152] The generic application name refers to the name that an application typically uses on a display device. The reference application name may differ from the generic application name; for example, the reference application name might be Youku Video, while the generic application name is Coolcat TV. Exemplarily, there is a preset correspondence between the reference application name and the generic application name; by looking up this correspondence, the corresponding generic application name can be determined.

[0153] refer to Figure 7 The name of the third-party application to be opened, after installation on the display device, can be the application name in the application icon summary page, or the application name displayed in the application icon. The processor pre-stores a mapping between generic application names and the application names of the third-party applications to be opened after installation on the display device. For example, if the generic application name is "Cool Cat TV," the application name displayed on the display device is "CIBN Cool Cat." The application name of the third-party application to be opened is the application name displayed on the display device after installation, corresponding to the generic application name.

[0154] In this embodiment, the reference application name of the third-party application to be opened in the interactive text is extracted by a large model, and then the general application name corresponding to the reference application name is determined. The application name that is displayed after installation on the display device and corresponds to the general application name is used as the application name of the third-party application to be opened. Through multi-level conversion, the application name in the interactive command input by the user is converted into an application name that the display device can recognize and process, thereby improving the intelligent controllability and control accuracy of the display device.

[0155] In some embodiments, executing the output operation instructions to complete the execution action sequence for operating the demonstration function includes: when the operation instruction is input text, if a virtual keyboard is displayed on the current display page, performing text recognition on the corresponding page image of the current display page to obtain the position of each key on the virtual keyboard; obtaining the character sequence corresponding to the content to be searched, matching each character in the character sequence with the character displayed on the key to obtain the target position of each character in the character sequence on the virtual keyboard; and triggering the corresponding keys sequentially based on the target position for content search.

[0156] The display device supports a search function, which typically requires text input to retrieve matching search results. A virtual keyboard refers to an area with multiple keys displayed on the screen. In some embodiments, the processor can determine whether a virtual keyboard exists on the current display page by performing image recognition on the page image corresponding to the currently displayed page.

[0157] Each key on the virtual keyboard is triggerable, and each key displays a corresponding character. The characters displayed on the keys can be letters, pinyin, Chinese characters, words, numbers, or combinations thereof. Figure 9 The image shown is a schematic diagram of a display page provided in some embodiments of this application. Figure 1 The image shows virtual buttons on the display page, with each button displaying characters including pinyin and numbers.

[0158] The content to be searched refers to the content specified in the user's interactive command. For example, if the interactive command is "Search for TV series C", then TV series C is the content to be searched.

[0159] To ensure the desired search content is obtained, corresponding keys on the virtual keyboard need to be triggered. Therefore, the precise position of each key on the virtual keyboard needs to be determined. This application proposes an action sequence for executing output operation instructions to complete the demonstration function operation when the operation instruction is input text. This sequence includes, if virtual keys are displayed on the current display page, performing text recognition on the corresponding page image to obtain the position of each key on the virtual keyboard, obtaining the character sequence corresponding to the content to be searched, matching each character in the character sequence with the character displayed on the key to obtain the target position of each character in the character sequence on the virtual keyboard, and triggering the corresponding keys sequentially based on the target position for content search. The text recognition can employ an OCR text recognition algorithm. (Reference) Figure 8 The processor can obtain the position of each key in the virtual keyboard through the OCR service.

[0160] A character sequence refers to the sequence of characters corresponding to the content to be searched. The characters in the character sequence should consist of the characters displayed on the keys to ensure that the processor can trigger the corresponding keys according to the characters in the character sequence. For example, the character sequence corresponding to the content to be searched is the pinyin of the title of TV series C.

[0161] In some embodiments, the large model supports Chinese to Pinyin conversion service. Using the large model, the content to be searched can be converted into multiple corresponding Pinyin, and the Pinyin can be used to form a character sequence.

[0162] The processor matches each character in the character sequence with the character displayed on the keypad to obtain the target position of each character in the character sequence on the virtual keyboard. Following the order of the characters in the character sequence, it sequentially triggers the keys at the corresponding target positions. After the keys are triggered, the corresponding content is retrieved. For example... Figure 10 The image shown is a schematic diagram of a display page provided in some embodiments of this application. Figure 2 In the image, the search box displays the character sequence "LANGLABANG", and the search results show the content corresponding to the character sequence.

[0163] In this embodiment, when the operation instruction is input text, the position of each key on the virtual keyboard is obtained by performing text recognition on the corresponding page image of the currently displayed page to ensure the accuracy of key triggering. Each character in the character sequence corresponding to the content to be searched is matched with the character displayed on the key to obtain the target position of each character in the character sequence on the virtual keyboard. In this way, it can be guaranteed that the content matching the content to be searched is found, which is beneficial to improving the accuracy of content search.

[0164] In some embodiments, the display device control method further includes: after the current output operation instruction is executed, comparing the page image before execution with the page image after execution; if the two are inconsistent, the current output operation instruction is executed successfully and continues to be executed; if the two are inconsistent, the current output operation instruction fails to be executed, and at least the interactive text, the content previously output by the large model, and the operation specification prompts are input into the large model, and the current operation instruction is re-output and continues to be executed.

[0165] After the operation command is executed, the page image will usually change. For example... Figure 11 The image shown is a schematic diagram of the page image before execution completion provided in some embodiments of this application. The image shows the homepage of application A. Figure 12 The diagram shows a page image after execution completion according to some embodiments of this application. The image shows the homepage of application B. For example, if the operation instruction is to open the application (application B), the page image before execution completion is the homepage of application A (see reference). Figure 11 The page image after execution is a summary of application icons. For example, if the operation instruction is to click (the target location corresponding to application B), the page image before execution is a summary of application icons (see reference). Figure 7 The page image after execution is the homepage of application B (see reference). Figure 12 ).

[0166] The processor compares the page images before and after execution. The comparison method can be an image similarity algorithm. If the similarity result is less than a preset value, the two are determined to be inconsistent. If the similarity result is not less than the preset value, the two are determined to be consistent.

[0167] If the page images before and after execution are inconsistent, it indicates that a page switch has occurred. The processor determines that the current output operation instruction has been successfully executed and continues to execute the step of outputting the next operation instruction. That is, it executes at least the steps of inputting the interactive text, the content output by the large model this time, and the operation specification prompts into the large model, and at least outputting the current operation instruction.

[0168] If the page images before and after execution are consistent, it indicates that no page switching has occurred. The processor determines that the current output operation instruction has failed. To ensure the smooth execution of the interaction instruction, the processor will re-execute the steps of outputting the current operation instruction and continue execution. That is, it will re-execute at least the steps of inputting the interaction text, the content previously output by the large model, and the operation specification prompts into the large model, and re-outputting the current operation instruction.

[0169] like Figure 13The diagram illustrates the processing flow of a result feedback service provided in some embodiments of this application. The result feedback service acquires a comparison of the page images before and after execution, and feeds back the execution result corresponding to the comparison result to the processor. Additionally, the processor can perform exception handling based on the execution result.

[0170] In this embodiment, after the current output operation instruction is executed, the page images before and after the execution are compared. If they are consistent, it indicates that the execution was successful and subsequent steps can continue. If they are inconsistent, it indicates that the execution failed and the output steps of the current operation instruction will be re-executed. This method of handling abnormal states ensures that abnormalities can be automatically handled when execution fails, ensuring the smooth operation of the control method and improving the user experience.

[0171] In some embodiments, the input content of the current processing of the large model also includes the page image corresponding to the currently displayed page; before executing the output current operation instruction to complete the sequence of execution actions to operate the demonstration function, the method further includes: if the large model identifies that the currently displayed page displays recommended content, skipping the recommended content or waiting for the demonstration of the recommended content to be completed.

[0172] Before executing the current operation command to complete the sequence of actions for the demonstration function, the currently displayed page may show recommended content for a preset duration. The display of recommended content may cause the execution sequence of actions to fail. For example, if the previous output operation command was to open the application (Application B), and the current output operation command is to click (the search box), since the current page does not enter the homepage of Application B after opening Application B, but instead displays recommended content, executing the action sequence corresponding to clicking (the search box) will cause the action to fail.

[0173] Based on this, embodiments of this application propose that when the processor identifies recommended content displayed on the current page through a large model, it skips the recommended content or waits for the recommended content to finish being displayed. This avoids interference from the recommended content with the execution of actions and helps improve the success rate of action execution. Figure 14 The diagram illustrates the recommended content processing provided in some embodiments of this application. The image and text understanding module identifies the page image of the currently displayed page, determines whether recommended content is displayed, and if so, skips the recommended content or waits for the recommended content to finish displaying.

[0174] When faced with the choice to skip recommended content or wait for the recommended content demonstration to complete, the processor can determine the user's identity type. If the user's identity type supports skipping recommended content, then skipping the recommended content is selected. If the user's identity type does not support skipping recommended content, then waiting for the recommended content demonstration to complete is selected.

[0175] In this embodiment, identifying recommended content on the currently displayed page, skipping recommended content, or waiting for the recommended content to finish being displayed can avoid interference from recommended content on action execution, which helps to improve the success rate of action execution.

[0176] In some embodiments, reference Figure 6 The large model also outputs at least one of the following: action description text of the operation instruction or description text of the target of the action being performed.

[0177] The output of the large model also includes at least one of the following: the action description text of the output operation instruction or the description text of the target of the action being performed.

[0178] Action description text refers to the text that describes the operation instruction. The action description text of the current operation instruction includes the action description content of the previous operation instruction. For example, if the operation instruction is to open an application (application B), the action description text could be: I will open application B. Another example is if the operation instruction is to click (the search box), the action description text could be: I will click the search box in application B. Yet another example is if the operation instruction is to enter (the title of TV series C), the action description text could be: I will enter the title of TV series C in the search box of application B.

[0179] The description text refers to the text that describes the objective of the action being performed. The description text of the currently performed action includes the description of the objective of the previous performed action. For example, if the action instruction is to open an application (application B), the description text could be: "To enable subsequent TV series search operations." Another example is if the action instruction is to click (the search box), the description text could be: "To activate the keyboard, then enter the title of TV series C and perform a search." Yet another example is if the action instruction is to enter (the title of TV series C), the description text could be: "To perform a TV series search."

[0180] In this embodiment, by using at least one of the action description text or the description text of the target of the executed action as the output of the large model, the action description text can describe the action of the operation instruction, and each action description text also includes the relevant description content of the previous operation instruction; the description text can describe the target of the operation instruction, and each description text of the executed action also includes the relevant target description content of the previous executed work, which helps the large model to accurately grasp at least one of the actions performed by each operation instruction or the target of the executed action, thereby improving the decision accuracy of the large model for the next output operation instruction.

[0181] In some embodiments, reference Figure 6 For the process of outputting interactive text and corresponding operation instructions for non-first-time outputs, the input content of the current processing of the large model also includes summary text of the completed actions.

[0182] The summary text refers to a summary of the completed actions. For example, after the first operation instruction is executed, the summary text could be: "As per the user's instruction, the operation of opening application B has been performed, preparing to search for TV series C." Another example is after the second operation instruction is executed: "The operation of opening application B has been performed, and the virtual keyboard has been activated, ready to input the title of TV series C." Yet another example is after the third operation instruction is executed: "The operation of opening application B has been performed, and the title of TV series C has been entered into the search box to perform the search."

[0183] In some embodiments, the large model supports a historical summary proxy service, which generates a summary description text of the completed actions after each operation instruction is executed.

[0184] For the process of outputting interactive text and corresponding operation instructions for the first time, the input content of the current processing of the large model also includes summary text of the completed actions.

[0185] In this embodiment, the summary description text includes a summary description of all actions completed before the current processing. In this way, the large model can output the current operation instruction based on the summary description text of the completed actions, which helps to improve the decision-making accuracy of the large model.

[0186] To illustrate the display device, display device control method, and effects in this solution in detail, a specific embodiment is described below:

[0187] This application provides a display device, including a display and at least one processor. The display is configured to display content from a broadcast system or network and / or a user interface; and the at least one processor is configured to execute a display device control method.

[0188] like Figure 15 The diagram shows a system framework of a display device provided in some embodiments of this application. The system includes an application framework, model services, and basic services. The application framework includes a basic capability execution unit and multiple processing modules. The basic capability execution unit includes, but is not limited to, application homepage, my applications, opening an application, clicking, swiping (up, down, left, right), input, homepage, and stopping. The multiple processing modules include, but are not limited to, a decision-making agent module, a history summary agent module, a result feedback agent module, a history storage module, a name conversion module, a keyboard judgment module, a recommended content filtering module, a keyboard key acquisition module, and a Chinese-to-Pinyin conversion module. Model services include, but are not limited to, large models, image and text understanding models, OCR services, and cloud query services.

[0189] Among them, the decision agent module can generate the current operation instructions from the user's interaction instructions by using the understanding ability of the large model through operation specification prompts. Then, it parses out the corresponding execution action sequence of the operation instructions and executes the corresponding actions according to the execution action sequence to realize the corresponding operation function.

[0190] Based on the operational guidelines, the large model outputs operational instructions, including but not limited to: application homepage, my applications, open application, click, swipe, input, homepage, and stop. Executing each operational instruction requires completing a sequence of actions to operate the demo function. The completion of each action within the sequence ensures that the corresponding operation function is performed correctly.

[0191] In addition, when a decision needs to be made next, the previous output of the large model and the summary text of the completed actions can be used as input for the decision, ensuring that the large model can output the current action based on the completed records.

[0192] The historical summary agent module can summarize and compile completed actions, generating summary description text for completed actions, which serves as an important input prompt for the decision agent module.

[0193] The result feedback agent module can provide the execution result based on the comparison of page images before and after execution, including execution success and execution failure, providing a decision-making premise for the operation instructions planned by the decision agent module in the next step.

[0194] The large-scale model training process includes the following steps: First, a large amount of open-source data is collected. This data can come from various sources such as books, articles, and websites, as well as business data, self-collected and self-annotated data, and self-constructed data. The collected data is used for pre-training the model, giving it basic language processing capabilities. Then, by adding a suitable amount of fine-tuning data containing input and output instructions, the pre-trained model is fine-tuned to ensure its content generation and output style are professional and consistent. Finally, a professional evaluation set is constructed to verify the instruction compliance capability. For any issues raised in the feedback, the model is further optimized, ultimately resulting in a large-scale model that meets the requirements.

[0195] The name conversion module can use a large model to convert the reference application name of the third-party application to be opened in the interactive text into a general application name. Then, based on the correspondence between the general application name and the application name of the third-party application to be opened after installation on the display device, the application name of the third-party application to be opened can be obtained.

[0196] The keyboard detection module can use the image and text understanding model to identify the page image of the currently displayed page and determine whether there is a virtual keyboard in the page image.

[0197] The Chinese-to-Pinyin module can convert the content to be searched into the corresponding Pinyin.

[0198] The recommended content filtering module can skip recommended content or wait for the recommended content to finish being displayed when the current page is already showing recommended content.

[0199] The keyboard key acquisition module can perform text recognition on the corresponding page image of the currently displayed page when a virtual keyboard is displayed on the current page, and obtain the position of each key on the virtual keyboard.

[0200] Specifically, when the operation command is to open a third-party application, if the currently displayed page is a summary page of application icons, text recognition is performed on the corresponding page image of the currently displayed page to obtain the location of the application name of each application; and image recognition is performed on the page image to obtain the location of the application icon of each application. The application name of the third-party application to be opened is matched with the application names in the currently displayed page. For applications that contain the application name, the location of the application name is used as the target location; for applications that do not contain the application name, the location of the application icon is used as the target location.

[0201] The following example illustrates display device control using a specific interactive command. The user inputs the command "Open application B, search for TV series C," and the currently active application is application A. Figure 16The diagram shown is a general flowchart of a display control method provided in some embodiments of this application.

[0202] For the process of initially outputting the operation instruction corresponding to the interactive text, the interactive text corresponding to the interactive instruction and the operation specification prompts used to guide the large model to output the operation instruction are used as the input of the large model. The content output by the large model for the first time includes the first operation instruction to operate the demonstration function, the action description text, and the description text of the target of the action to be performed. For example, the content output by the large model for the first time can be: Step-1: [Operation: I will open application B in order to perform subsequent TV series search operations; Action: Open application (application B)].

[0203] For the process of outputting the corresponding operation instructions for the interactive text for the first time, the interactive text, the content output by the large model last time, the operation specification prompts, and the summary description text of the completed actions are used as input to the large model. The content output by the large model for the first time includes the operation instructions, action description text, and description text of the target of the action being performed.

[0204] For example, the second output of the large model can be: Step-1: [Action: I will open application B in order to perform the subsequent TV series search operation; Action: Open application (application B)]; Step-2: [Action: I will click the search box in application B; Action: Click (search box)].

[0205] The third output of the large model can be: Step-1: [Operation: I will open application B to perform the subsequent TV series search operation; Action: Open application (application B)]; Step-2: [Operation: I will click the search box in application B; Action: Click (search box)]; Step-3: [Operation: I will enter the title of TV series C in the search box of application B; Action: Enter (title of TV series C)].

[0206] The fourth output of the large model can be: Step-1: [Action: I will open application B to perform the subsequent TV series search operation; Action: Open application (application B)]; Step-2: [Action: I will click the search box in application B; Action: Click (search box)]; Step-3: [Action: I will enter the name of TV series C in the search box of application B; Action: Enter (the name of TV series C)]; Step-4: [Action: I will click the option corresponding to TV series C in the search results to complete the search operation in the user command; Action: Click (the option corresponding to TV series C)].

[0207] In the above embodiments, a display device and a display device control method receive user-input interactive commands. The corresponding interactive text and operational guidelines for guiding a large model to output operation commands are used as input to the large model. The large model's decision-making and planning capabilities enable intelligent path planning. This intelligent path planning includes operation commands for each demonstration function output by the large model. For operation commands output after the initial interaction text, the large model's input also includes the content of its previous output, ensuring that the large model can decide on the current output operation command based on completed content, thus improving the accuracy and rationality of each output operation command. Each time an operation command is output, the output operation command is executed to complete the sequence of actions for operating the demonstration function, achieving the corresponding operation function, until the last output operation command is executed. This achieves intelligent control matching user interactive commands, which is beneficial for meeting users' intelligent control needs of the display device. Simultaneously, by combining image recognition and text recognition, clickable icon positions can be quickly and accurately identified, ultimately enabling operation of user interactive commands.

[0208] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0209] Based on the same inventive concept, this application also provides a display device control apparatus for implementing the display device control method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more display device control apparatus embodiments provided below can be found in the limitations of the display device control method described above, and will not be repeated here.

[0210] In one exemplary embodiment, a display device control apparatus is provided, comprising: an acquisition module, a first output module, a second output module, and an execution module, wherein:

[0211] The acquisition module is used to receive user-input interaction commands and obtain the corresponding interaction text.

[0212] The first output module is used for the process of outputting operation instructions corresponding to the first interactive text. It takes at least the interactive text and the operation specification prompts used to guide the large model to output operation instructions as input to the large model, and outputs at least the first operation instruction to operate the demonstration function.

[0213] The second output module is used for the process of outputting operation instructions corresponding to interactive text for the first time. It inputs at least the interactive text, the content of the previous output of the large model, and the operation specification prompts into the large model, and outputs at least the current operation instructions.

[0214] The execution module is used to execute the output operation instructions each time an operation instruction is output to complete the sequence of execution actions to operate the demonstration function, so as to realize the corresponding operation function, until the last output operation instruction is completed.

[0215] The aforementioned display device control device receives user-input interactive commands. The corresponding interactive text and operational guidelines for guiding the large model to output operation commands serve as input to the large model. Utilizing the large model's decision-making and planning capabilities, it achieves intelligent path planning. This intelligent path planning includes the operation commands output by the large model for each demonstration function. For operation commands output after the initial interaction text, the large model's input also includes the content of its previous output, ensuring that the large model can decide on the current output operation command based on completed content, thus improving the accuracy and rationality of each output operation command. In each output operation command, the output operation command is executed to complete the sequence of actions for operating the demonstration function, achieving the corresponding operation function, until the last output operation command is executed. This achieves intelligent control that matches user interactive commands, effectively meeting users' needs for intelligent operation of the display device.

[0216] In some embodiments, the execution module is further configured to: when the operation instruction is to open a third-party application, if the currently displayed page is an application icon summary page, perform text recognition on the corresponding page image of the currently displayed page to obtain the location of the application name of each application; match the application name of the third-party application to be opened with the application names in the currently displayed page to obtain the target location of the application name of the third-party application to be opened in the current displayed page; and open the third-party application based on the target location.

[0217] In some embodiments, the process of obtaining the application name of the third-party application to be opened is further configured by the execution module to: extract the reference application name of the third-party application to be opened from the interactive text through a large model; convert the reference application name into a general application name of the third-party application to be opened; and obtain the application name of the third-party application to be opened based on the correspondence between the general application name and the application name presented after the third-party application to be opened is installed on the display device.

[0218] In some embodiments, the execution module is further configured to: when the operation instruction is input text, if a virtual keyboard is displayed on the current display page, perform text recognition on the corresponding page image of the current display page to obtain the position of each key on the virtual keyboard; obtain the character sequence corresponding to the content to be searched, match each character in the character sequence with the character displayed on the key to obtain the target position of each character in the character sequence on the virtual keyboard; and, based on the target position, sequentially trigger the corresponding keys for content search.

[0219] In some embodiments, the display device control device further includes an exception handling module, which is configured to: after the currently output operation instruction completes the execution action sequence for operating the demonstration function, compare the page image before execution completion with the page image after execution completion; if the two are inconsistent, the currently output operation instruction is executed successfully and continues to be executed; if the two are inconsistent, the currently output operation instruction fails to be executed, and at least the interactive text, the content previously output by the large model, and the operation specification prompts are input into the large model, the current operation instruction is re-output, and execution continues.

[0220] In some embodiments, the input content of the current processing of the large model also includes the page image corresponding to the currently displayed page; before executing the output current operation instruction to complete the sequence of execution actions to operate the demonstration function, the display device control device further includes a recommended content processing module, which is used to: skip the recommended content or wait for the recommended content demonstration to be completed when the large model identifies that the currently displayed page displays recommended content.

[0221] In some embodiments, the large model also outputs at least one of the action description text of the output operation instruction or the description text of the target of the action to be performed.

[0222] In some embodiments, for the process of outputting operation instructions corresponding to interactive text for the first time, the input content of the current processing of the large model also includes summary text of the completed actions.

[0223] Each module in the aforementioned display device control unit can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0224] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 17 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a display device control method. The display unit of the computer device forms a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0225] Those skilled in the art will understand that Figure 17 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0226] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0227] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0228] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0229] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0230] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0231] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0232] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A display device, characterized in that, include: The display is configured to show content from a broadcast system or network and / or a user interface; and at least one processor, configured as follows: Receive user input of interactive instructions and obtain the corresponding interactive text of the interactive instructions; For the process of outputting the corresponding operation command for the first time, at least the interactive text and the operation specification prompts used to guide the large model to output operation commands are used as inputs to the large model, and at least the first operation command to operate the demonstration function is output; the operation specification prompts are preset prompt words used to guide the large model to output operation commands that conform to the operation specifications of the display device; For processes where the operation instructions corresponding to the interactive text are not output for the first time, at least the interactive text, the content previously output by the large model, and the operation specification prompts are input into the large model, and at least the current operation instructions are output. Multiple operation instructions are obtained through multiple outputs of the large model, and each operation instruction has a sequential execution order. Each time an operation command is output, the output operation command is executed to complete the sequence of execution actions to operate the demonstration function, so as to realize the corresponding operation function, until the last output operation command is completed; the sequence of execution actions includes at least one execution action with a sequential order; The sequence of actions by which the processor executes output operation instructions to complete the operation of the demonstration function is configured as follows: When the operation instruction is to open a third-party application, if the currently displayed page is an application icon summary page, text recognition is performed on the corresponding page image of the currently displayed page to obtain the location of the application name of each application. The application name of the third-party application to be opened is matched with the application name on the currently displayed page to obtain the target position of the application name of the third-party application to be opened on the currently displayed page; the third-party application to be opened is the third-party application to be opened as indicated by the interaction command. Based on the target location, open the third-party application.

2. The display device according to claim 1, characterized in that, The process by which the processor retrieves the application name of the third-party application to be opened is configured as follows: Using the large model, extract the name of the third-party application that needs to be opened from the interactive text; The referred application name is converted into a generic application name of the third-party application that needs to be opened; Based on the correspondence between the general application name and the application name of the third-party application to be opened after installation on the display device, the application name of the third-party application to be opened is obtained.

3. The display device according to claim 1, characterized in that, The sequence of actions by which the processor executes output operation instructions to complete the operation of the demonstration function is configured as follows: When the operation instruction is input text, if a virtual keyboard is displayed on the current display page, then text recognition is performed on the corresponding page image of the current display page to obtain the position of each key on the virtual keyboard; Obtain the character sequence corresponding to the content to be searched, match each character in the character sequence with the character displayed on the key, and obtain the target position of each character in the character sequence on the virtual keyboard; Based on the target location, the corresponding buttons are triggered sequentially for content search.

4. The display device according to claim 1, characterized in that, The processor is also configured to: After the current output operation command is executed, compare the page image before execution with the page image after execution; If the two are inconsistent, the currently output operation instruction is executed successfully and execution continues; If the two are inconsistent, the current output operation instruction fails to execute. At least the interactive text, the content of the previous output of the large model, and the operation specification prompt should be input into the large model, and the current operation instruction should be re-output and continue to be executed.

5. The display device according to claim 1, characterized in that, The input content of the current processing of the large model also includes the page image corresponding to the currently displayed page; before executing the output current operation instructions to complete the sequence of actions to operate the demonstration function, the processor is also configured to: If the large model identifies that the currently displayed page contains recommended content, skip the recommended content or wait for the recommended content to be displayed.

6. The display device according to claim 1, characterized in that, The large model also outputs at least one of the following: action description text of the output operation instruction or description text of the target of the action being performed.

7. The display device according to claim 1, characterized in that, For processes where the operation instructions corresponding to the interactive text are not output for the first time, the input content of the current processing of the large model also includes a summary description text of the completed actions.

8. A method for controlling a display device, characterized in that, Applied to a display device as described in any one of claims 1-7; the method includes: Receive user input of interactive instructions and obtain the corresponding interactive text of the interactive instructions; For the process of outputting the corresponding operation command for the first time, at least the interactive text and the operation specification prompts used to guide the large model to output operation commands are used as inputs to the large model, and at least the first operation command to operate the demonstration function is output; the operation specification prompts are preset prompt words used to guide the large model to output operation commands that conform to the operation specifications of the display device; For the process of outputting the corresponding operation instructions for the interactive text for the first time, at least the interactive text, the content of the previous output of the large model, and the operation specification prompts are input into the large model, and at least the current operation instruction is output; through multiple outputs of the large model, multiple operation instructions are obtained, and each operation instruction has a sequential execution order; Each time an operation command is output, the output operation command is executed to complete the sequence of execution actions to operate the demonstration function, so as to realize the corresponding operation function, until the last output operation command is completed; the sequence of execution actions includes at least one execution action with a sequential order; The execution output operation instructions complete the sequence of execution actions to operate the demonstration function, including: When the operation instruction is to open a third-party application, if the currently displayed page is an application icon summary page, text recognition is performed on the corresponding page image of the currently displayed page to obtain the location of the application name of each application. The application name of the third-party application to be opened is matched with the application name on the currently displayed page to obtain the target position of the application name of the third-party application to be opened on the currently displayed page; the third-party application to be opened is the third-party application to be opened as indicated by the interaction command. Based on the target location, open the third-party application.

Citation Information

Patent Citations

  • Large model interaction processing method and system, terminal, equipment and medium

    CN117520497A

  • Display device and display device control method based on multi-model group interaction

    CN118646924A