Display device and server

By detecting screenshot events in the display device, acquiring images, and generating appropriate prompts, the problem of poor user interaction caused by monotonous prompts is solved, and the reliability and flexibility of user interaction are improved.

CN121842444APending Publication Date: 2026-04-10JUHAOKAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-04
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The current technology uses a relatively simple form of guidance text, which reduces the user's interactive experience.

Method used

By detecting screenshot events, the target image is obtained and an image recognition request is sent to the server to generate a target prompt. Based on the target content consisting of image recognition results, business information, and interface focus information, as well as the trigger source identifier, a prompt that is compatible with the current display interface and the screenshot initiation method is generated.

Benefits of technology

It enhances the user interaction experience, ensures the reliability and adaptability of the guidance messages, provides multiple screenshot triggering methods, and improves the flexibility and accuracy of user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121842444A_ABST
    Figure CN121842444A_ABST
Patent Text Reader

Abstract

The invention relates to a display device and a server. The display apparatus includes a display and a controller. And the controller is configured to perform screenshot on a current display interface of the display to obtain a target image under the condition that a screenshot event is detected, send an image recognition request to the server, receive an image recognition result and a target guide language fed back by the server based on the image recognition request, and control the display to display the image recognition result and the target guide language. Wherein the image recognition request carries a target image and a trigger source identifier of a screenshot event, the image recognition result is obtained by recognizing the target image, and the target guide language is generated based on the trigger source identifier and target content; the target content comprises at least one of the image recognition result, service information associated with the current display interface and interface focus information of the current display interface. By adopting the server, the guide language can adapt to the image recognition scene, so that the interaction experience feeling of the user is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, and in particular to a display device and a server. BACKGROUND

[0002] With the continuous development of display devices (such as smart televisions), in order to facilitate user interaction with the display device, when it is detected that the user has an interaction intention, the user interaction intention can be processed, and when the interaction result is fed back to the user, the guidance language is also fed back to prompt the user to interact. For example, when the user has a screenshot intention, a screenshot of the display page can be obtained, and the screenshot can be recognized, and then when the image recognition result is displayed to the user, the guidance language is also displayed.

[0003] However, the form of the guidance language in the prior art is relatively single, which reduces the user interaction experience. SUMMARY

[0004] The present application provides a display device and a server, which can adapt the guidance language to the image recognition scene, thereby ensuring the user interaction experience.

[0005] In a first aspect, some embodiments provide a display device, comprising:

[0006] a display and a controller;

[0007] the controller is configured to:

[0008] in a case where a screenshot event is detected, take a screenshot of a current display interface of the display to obtain a target image;

[0009] send an image recognition request to a server; wherein the image recognition request carries the target image and a trigger source identifier of the screenshot event; the trigger source identifier is used to identify whether the screenshot event is triggered by voice or a key;

[0010] receive an image recognition result and a target guidance language fed back by the server based on the image recognition request; wherein the image recognition result is obtained by recognizing the target image, and the target guidance language is generated based on the trigger source identifier and target content; the target content includes at least one of the image recognition result, business information associated with the current display interface, and interface focus information of the current display interface; in a case where the target content is the same and the trigger source identifier is different, the corresponding guidance language is different; in a case where the trigger source identifier is the same and the target content is different, the corresponding guidance language is different;

[0011] control the display to display the image recognition result and the target guidance language.

[0012] In the guide sentence generation method, the display device obtains a target image by taking a screenshot of the current display interface of the display when a screenshot event is detected; then, the server performs image recognition on the target image, generates a target guide sentence according to target content composed of at least one of an image recognition result, service information, and interface focus information, and a trigger source identifier, and the target guide sentence; and the display displays the image recognition result and the target guide sentence for the user to initiate a next query request in combination with the target guide sentence. Compared with the related art in which only fixed guide sentences are displayed to the user, the above method can make the target guide sentence adapt to the current display interface, the image recognition result, and the screenshot initiation mode by instructing the server to generate the target guide sentence according to the target content and the trigger source identifier, thereby ensuring the reliability of the target guide sentence, laying a reliable foundation for subsequent guiding the user to interact with the display device for the next step, and improving the user interaction experience.

[0013] In some embodiments, when the controller detects the screenshot event, it is configured to: receive a pressing instruction; wherein the pressing instruction represents that a preset button in the remote controller is pressed; or, detect that a screenshot control in the current display interface is triggered; or, receive a first user voice representing a screenshot intention. In the above embodiments, multiple optional screenshot triggering modes are provided for the user, which can ensure the flexibility of initiating the screenshot event and improve the user interaction experience.

[0014] In some embodiments, when the screenshot event is triggered based on the first user voice representing the screenshot intention, the controller is configured to: receive a second user voice; perform voice recognition on the second user voice; send an image recognition request to the server when the voice recognition result includes a target image recognition type for the current display interface; wherein the image recognition request further includes service information associated with the current display interface, and the service information includes the target image recognition type for the current display interface; and perform deletion processing on the target image when the voice recognition result does not include the target image recognition type. In the above embodiments, when the screenshot event is triggered based on the first user voice representing the screenshot intention, the second user voice of the user is continuously recognized, and the image recognition request is sent to the server only when the target image recognition type exists in the voice recognition result, which can ensure that the user recognition processing is adapted to the user intention and avoid waste of computing resources.

[0015] In some embodiments, the image recognition request further comprises business information associated with the current display interface and / or interface focus information of the current display interface, the business information comprising a target interaction scenario represented by the current display interface; the controller is further configured to: acquire, by the home application, the target interaction scenario represented by the current display interface and / or the interface focus information of the current display interface; wherein the home application is an application in the display device that is currently in visual interaction with the user. In the above embodiments, on the one hand, the target interaction scenario represented by the current display interface and / or the interface focus information are acquired by the home application, which can ensure the reliability of information acquisition; on the other hand, the target interaction scenario represented by the current display interface and / or the interface focus information are carried in the image recognition request, which provides rich content for subsequent generation of diversified target guide language.

[0016] In a second aspect, some embodiments further provide a server comprising:

[0017] a communication device configured to be communicatively connected with the display device;

[0018] and at least one processor connected with the communication device and configured to:

[0019] receive an image recognition request sent by the display device; wherein the image recognition request carries a target image and a trigger source identifier of a screenshot event; the target image is obtained by taking a screenshot of a current display interface when the display device detects the screenshot event; the trigger source identifier is used to identify whether the screenshot event is triggered by voice or by a key;

[0020] identify the target image to obtain an image recognition result; and

[0021] generate a target guide language according to the trigger source identifier and target content; wherein the target content comprises at least one of the image recognition result, business information associated with the current display interface, and interface focus information of the current display interface; the corresponding guide language is different in the case that the target content is the same but the trigger source identifier is different; the corresponding guide language is different in the case that the trigger source identifier is the same but the target content is different;

[0022] feed back the image recognition result and the target guide language to the display device; wherein the image recognition result and the target guide language are used for the display device to update and display the current display interface.

[0023] In the guide sentence generation method, after receiving the image recognition request carrying the target image and the trigger source identifier sent by the display device, the server identifies the target image to obtain an image recognition result, and generates a target guide sentence according to the trigger source identifier, the image recognition result, service information, and interface focus information. Then, the server feeds back the image recognition result and the target guide sentence to the display device, so that the display device updates and displays the current display interface. Compared with the related art in which only fixed guide sentences are displayed to the user, the above method can adapt the target guide sentence to the current display interface, the image recognition result, and the screenshot initiation mode, thereby ensuring the reliability of the target guide sentence and laying a reliable foundation for subsequent guiding the user to perform the next interaction, and improving the user experience of the interaction.

[0024] In some embodiments, when the processor generates the target guide sentence according to the trigger source identifier and the target content, the processor is configured to: find an auxiliary guide sentence corresponding to the trigger source identifier from a first mapping relationship, wherein the first mapping relationship includes a correspondence between different trigger source identifiers and different guide sentences; generate a guide sentence body of the target guide sentence according to the target content; and splice the guide sentence body and the found auxiliary guide sentence to obtain the target guide sentence. In the above embodiment, the auxiliary guide sentence is determined according to the trigger source identifier, the guide sentence body is determined according to the target content, and the target guide sentence is constructed according to the guide sentence body and the auxiliary guide sentence, which can ensure the adaptability between the target guide sentence and the triggering mode and the guiding scene, and thus improve the accuracy of the target guide sentence determination.

[0025] In some embodiments, the service information includes a target image recognition type for the current display interface, and when the processor generates the guide sentence body of the target guide sentence according to the target content, the processor is configured to: find a guide sentence corresponding to the target image recognition type from a second mapping relationship, wherein the second mapping relationship includes a correspondence between different image recognition types and different guide sentences; and determine the guide sentence body of the target guide sentence based on the found guide sentence. In the above embodiment, the second mapping relationship including the correspondence between different image recognition types and different guide sentences is introduced, and then the guide sentence body associated with the target image recognition type can be obtained by querying the second mapping relationship. On the one hand, this provides a convenient way for quickly generating guide sentences; on the other hand, different guide sentences are configured for different image recognition types, which enriches the image recognition scene and the interaction scene and further improves the user experience.

[0026] In some embodiments, the service information includes a target interaction scene represented by the current display interface; when the processor generates the guide speech subject of the target guide speech according to the target content, the processor is configured to: find a guide speech corresponding to the target interaction scene from a third mapping relationship; the third mapping relationship includes a corresponding relationship between different interaction scenes and different guide speeches; and determine the guide speech subject of the target guide speech based on the found guide speech. In the above embodiment, the third mapping relationship including the corresponding relationship between different interaction scenes and different guide speeches is introduced, and then the guide speech subject associated with the target interaction scene can be obtained by querying in the third mapping relationship. On the one hand, a convenient way is provided for quickly generating a guide speech; on the other hand, different guide speeches are configured for different interaction scenes, which enriches the interaction scene of the display device and the user and further improves the user experience.

[0027] In some embodiments, when the processor generates the guide speech subject of the target guide speech according to the target content, the processor is configured to: obtain media asset information of a media asset associated with the current display interface; analyze the image recognition result to obtain a target object; determine role information of the target object in the media asset according to the media asset information; and generate the guide speech subject of the target guide speech according to the role information and the media asset information. In the above embodiment, the guide speech subject of the target guide speech is generated according to the role information and the media asset information in the current display interface, which can make the generated guide speech adapt to the current display interface, thereby increasing the interest of the user in further watching or browsing the current display interface, which not only improves the user experience but also achieves the effect of expanding the product.

[0028] In some embodiments, when the processor generates the guide speech subject of the target guide speech according to the target content, the processor is configured to: determine an interface word corresponding to an interface focus in the current display interface according to the interface focus information; determine a focus content of the current display interface according to the determined interface word; and generate the guide speech subject of the target guide speech based on the focus content. In the above embodiment, the guide speech subject of the target guide speech is generated according to the interface word corresponding to the interface focus in the current display interface. Since the interface focus can usually represent the content currently focused by the user, the generated target guide speech can be adapted to the actual needs of the user based on the interface focus, thereby improving the rationality of the determination of the target guide speech.

[0029] In a third aspect, some embodiments further provide a guide speech generation method applied to a display device, including:

[0030] In the case of detecting the screenshot event, taking a screenshot of the current display interface of the display to obtain a target image;

[0031] sending an image recognition request to a server; wherein the image recognition request carries a target image and a trigger source identifier of a screenshot event; the trigger source identifier is used to identify that the screenshot event is triggered by voice or a key;

[0032] receiving an image recognition result and a target guide sentence fed back by the server based on the image recognition request; wherein the image recognition result is obtained by recognizing the target image, and the target guide sentence is generated based on the trigger source identifier and target content; the target content comprises at least one of the image recognition result, business information associated with a current display interface, and interface focus information of the current display interface; the corresponding guide sentences are different in the case that the target content is the same and the trigger source identifiers are different; the corresponding guide sentences are different in the case that the trigger source identifiers are the same and the target content is different;

[0033] controlling a display to display the image recognition result and the target guide sentence.

[0034] In a fourth aspect, some embodiments further provide a guide sentence generation method applied to a server, comprising:

[0035] receiving an image recognition request sent by a display device; wherein the image recognition request carries a target image and a trigger source identifier of a screenshot event; the target image is obtained by taking a screenshot of a current display interface in the case that the display device detects the screenshot event; the trigger source identifier is used to identify that the screenshot event is triggered by voice or a key;

[0036] recognizing the target image to obtain an image recognition result; and

[0037] generating a target guide sentence according to the trigger source identifier and target content; wherein the target content comprises at least one of the image recognition result, business information associated with a current display interface, and interface focus information of the current display interface; the corresponding guide sentences are different in the case that the target content is the same and the trigger source identifiers are different; the corresponding guide sentences are different in the case that the trigger source identifiers are the same and the target content is different;

[0038] feeding back the image recognition result and the target guide sentence to the display device; wherein the image recognition result and the target guide sentence are used for the display device to update and display the current display interface.

[0039] In a fifth aspect, a computer readable storage medium is provided, which stores a computer program, and when the computer program is run by a guide sentence generation device, the guide sentence generation device executes the guide sentence generation method provided in the third aspect or the fourth aspect.

[0040] In a sixth aspect, a computer program product is provided, which includes a computer program, when the computer program is executed by a guide sentence generation apparatus, causes the guide sentence generation apparatus to execute the guide sentence generation method provided in the third aspect or the fourth aspect. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0042] Figure 1 The schematic diagram of the operation scenario between the display device and the control device provided by some embodiments of the present application is shown.

[0043] Figure 2 The schematic diagram of the hardware configuration of the display device provided by some embodiments of the present application is shown.

[0044] Figure 3 The schematic diagram of the hardware configuration of the control device provided by some embodiments of the present application is shown.

[0045] Figure 4 The schematic diagram of the software configuration of the display device provided by some embodiments of the present application is shown.

[0046] Figure 5 The schematic diagram of the flow of the guide sentence generation method provided by some embodiments of the present application is shown.

[0047] Figure 6 The schematic diagram of the display page provided by some embodiments of the present application is shown.

[0048] Figure 7 The schematic diagram of the flow of the voice intent recognition provided by some embodiments of the present application is shown.

[0049] Figure 8 The schematic diagram of the flow of the information acquisition provided by some embodiments of the present application is shown.

[0050] Figure 9 The schematic diagram of the flow of the guide sentence generation method provided by some other embodiments of the present application is shown.

[0051] Figure 10 The schematic diagram of the flow of the determination of the target guide sentence provided by some embodiments of the present application is shown.

[0052] Figure 11 The schematic diagram of the flow of the determination of the guide sentence subject provided by some embodiments of the present application is shown.

[0053] Figure 12 Flowchart of a process for determining a guide speech subject according to some embodiments of the present application;

[0054] Figure 13 Flowchart of a process for determining a guide speech subject according to some embodiments of the present application;

[0055] Figure 14 Flowchart of a process for determining a guide speech subject according to some embodiments of the present application;

[0056] Figure 15 Flowchart of a process for determining a guide speech subject according to some embodiments of the present application; DETAILED DESCRIPTION

[0057] The embodiments will be described in detail with reference to the drawings, wherein like reference numerals refer to like parts throughout the several views. The following description is made with reference to the accompanying drawings in which the same or like reference numerals in different drawings represent the same or similar elements. The embodiments described in the following description are not meant to be exhaustive or to be limited to the precise form of the embodiments disclosed. These embodiments are chosen and described so that others skilled in the art can appreciate and understand the principles and practices of the present application.

[0058] It should be noted that the brief description of terms in the present application is only for the convenience of understanding the following described embodiments, and is not intended to limit the embodiments of the present application. Unless otherwise stated, these terms should be understood according to their ordinary and customary meanings.

[0059] The terms "first", "second", "third", and the like in the description and in the claims of the present application and above-described drawings are used for distinguishing between similar or identical objects or entities, and do not necessarily mean a specified order or sequence, unless otherwise noted. It will be understood that the terms so used are interchangeable under appropriate circumstances.

[0060] The terms "comprise" and "have" and any variations thereof are intended to cover a non-exclusive inclusion, for example, a product or device that comprises a list of components does not necessarily include all the components in the list, but can include additional components not expressly listed or inherent to such product or device.

[0061] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that can perform the functions described with respect to that element.

[0062] In some embodiments, the server can be a server 400 with picture recognition function, which can be a standalone server or a server cluster. The display device 200 refers to a device with picture display and data processing capability. For example, the display device 200 includes but is not limited to a smart television, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.

[0063] Figure 1 The schematic diagram of the operation scenario between the display device and the control device is provided for some embodiments of the present application. As shown in Figure 1 The user can operate the display device 200 through touch operation, mobile terminal 300 and control device 100. For example, the control device 100 can be a remote controller, a touch pen, a handle, etc.

[0064] The mobile terminal 300 can be used as a kind of control device to perform human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a kind of communication device to establish a communication connection with the display device 200 and perform data interaction. In some embodiments, the mobile terminal 300 can install a software application on the display device 200, realize connection communication through a network communication protocol, and achieve the purpose of one-to-one control operation and data communication. The mobile terminal 300 can also transmit the audio and video content displayed on the mobile terminal 300 to the display device 200 to realize the function of synchronous display.

[0065] As shown in Figure 1 It is also shown in that the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 can be allowed to communicate through a local area network (Local Area Network, LAN), a wireless local area network (Wireless Local Area Network, WLAN) and other networks.

[0066] The display device 200 can provide a broadcast receiving television function, and can additionally provide a smart network television function with computer support function, including but not limited to a network television, a smart television, an Internet Protocol Television (Internet Protocol Television, IPTV), etc.

[0067] Figure 2 The schematic diagram of the operation scenario between the display device and the control device is provided for some embodiments of the present application. As shown in Figure 1 The hardware configuration block diagram of the display device 200 is shown in some embodiments of the present application. In some embodiments, the display device 200 can include at least one of a tuning demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, a user input interface.

[0068] In some embodiments, the detector 230 is configured to collect signals of the external environment or the external interaction. For example, the detector 230 includes a light receiver configured to collect ambient light intensity; or the detector 230 includes an image collector, such as a camera, configured to collect an external environment scene, a user attribute, or a user interaction gesture; or the detector 230 includes a sound collector, such as a microphone, configured to receive external sound.

[0069] In some embodiments, the display 260 includes a display component configured to present a picture, and a driving component configured to drive the display. The display 260 is configured to receive image signals output from the controller 250 and display the image signals. For example, the display 260 can be configured to display video content, image content, and components of a menu control interface, and a user control UI.

[0070] In some embodiments, the communication device 220 is configured to communicate with the external device or the server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication manners. For example, when the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including a Wireless Fidelity (WiFi) function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.

[0071] The communication device 220 can be configured to connect the display device 200 to the external device or the server 400 in a wireless or wired manner. The wired connection can be achieved by connecting the display device 200 to the external device through a data line, an interface, or the like. The wireless connection can be achieved by connecting the display device 200 to the external device through a wireless signal or a wireless network. The display device 200 can be directly connected to the external device, or can be indirectly connected to the external device through a gateway, a router, a connection device, or the like.

[0072] In some embodiments, the controller 250 can include at least one of a central processor, a video processor, an audio processor, a graphics processor, a power supply processor, a first interface to an n-th interface for input / output, and the controller 250 controls the operation of the display device and responds to the user's operation by controlling various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.

[0073] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different split devices, i.e., the tuner demodulator 210 can also be located in an external device of the main device where the controller 250 is located, such as an external set-top box, or the like.

[0074] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0075] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0076] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0077] Figure 3 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of the central control device. (Example) Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0078] The control device 100 is configured to control the display device 200, and to receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.

[0079] In some embodiments, the control device 100 may be an intelligent device. For example, the control device 100 may be equipped with various applications for controlling the display device 200 according to user needs.

[0080] In some embodiments, such as Figure 1 As shown, the mobile terminal 300 or other intelligent electronic device, after installing the application that controls the display device 200, can perform functions similar to those of the control device 100. The controller 110 includes a processor 112, RAM 113 and ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation of the control device 100, as well as communication and cooperation between its internal components and external and internal data processing functions. Under the control of the controller 110, the communication interface 130 realizes communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of other near-field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.

[0081] The user input / output interface 140. Among them, the input interface includes at least one of the microphone 141, the touchpad 142, the sensor 143, the key 144 and other input interfaces. In some embodiments, the control device 100 includes at least one of the communication interface 130 and the input / output interface 140. The communication interface 130 configured in the control device 100, such as WiFi, Bluetooth, NFC and the like, can encode the user input instruction through the WiFi protocol, or the Bluetooth protocol, or the NFC protocol, and send it to the display device 200. The memory 190 is used to store various running programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user. The power supply 180 is used to provide operating power support for each element of the control device 100 under the control of the controller.

[0082] In order to perform user interaction, in some embodiments, the display device 200 can run an operating system. The operating system is a computer program used to manage and control hardware resources and software resources in the display device 200. The operating system can provide a user interface to allow the user to interact with the display device 200 and support running various application programs.

[0083] It should be noted that the operating system can be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device. The operating system can be divided into different modules or levels according to the functions implemented, for example, as shown in Figure 4 In some embodiments, the system is divided into four layers from top to bottom, namely the application layer (referred to as "application layer"), the application framework layer (referred to as "framework layer"), the system library layer and the kernel layer.

[0084] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run the applications and interact with the user based on the applications. At least one application can be run in the application layer, which can be a Window program, a system setting program, a clock program, etc. provided by the operating system, or an application developed by a third party developer. In a specific implementation, the applications in the application layer are not limited to the above examples. The framework layer provides an application programming interface (API) and a programming framework for the applications. The application framework layer includes some pre-defined functions. The application framework layer is equivalent to a processing center that decides which application in the application layer to act. The application can access the resources in the system and obtain the services of the system through the API interface during execution.

[0085] As shown in Figure 4 The application framework layer in the embodiments of the present application includes a view system, managers, a content provider, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The managers include at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing the system services or applications with access to the system location services; a package manager for retrieving various information related to the application package currently installed on the device; a notification manager for controlling the display and clearing of the notification messages; and a window manager for managing the icons, windows, toolbars, wallpapers and desktop components on the user interface.

[0086] In some embodiments, the activity manager is used to manage the life cycle of each application and the general navigation back function, such as controlling the exit, opening, back, etc. of the application. The window manager is used to manage all window programs, such as obtaining the size of the display screen, judging whether there is a status bar, locking the screen, intercepting the screen, controlling the display window change, such as reducing the display window, shaking the display, twisting the display, etc.

[0087] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions implemented by the framework layer.

[0088] In some embodiments, the kernel layer is a functional layer between the hardware and the software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, memory management, etc. For example, as shown in Figure 4 The kernel layer can be configured with hardware drivers. The drivers contained in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power supply driver, etc.

[0089] It should be noted that the above examples are only a simple division of the functions of the operating system, and do not constitute a limitation on the specific operating system form of the display device 200 in the embodiments of the present application. According to the function of the display device, the type of the operating system, and other factors, the number and specific type of the layers contained in the operating system can be in other forms.

[0090] With the continuous development of display devices (such as smart TVs), in order to facilitate user interaction with the display device, when a user has an interaction intention, the user interaction intention can be processed, and when the user is fed back the interaction result, the guidance language is also fed back to prompt the user to interact. For example, when the user has a screenshot intention, the display page can be obtained, and the screenshot can be identified, and then the guidance language is displayed together with the image recognition result when the image recognition result is displayed to the user.

[0091] However, the form of the guidance language in the prior art is relatively single, which reduces the experience of user interaction.

[0092] Based on this, in some embodiments, a guidance language generation method is provided. Wherein the guidance language generation method can be realized by cooperation of devices such as remote control devices, display devices and servers.

[0093] In an optional implementation, the guidance language generation method is applied to the controller in the display device. Wherein the display device includes a display and a controller. As shown in Figure 5 The specific steps include the following steps:

[0094] S510, in the case of detecting a screenshot event, taking a screenshot of the current display interface of the display to obtain a target image.

[0095] The screenshot event refers to the event that takes a screenshot of the currently displayed interface. The currently displayed interface is the interface shown on the monitor at the current moment. The target image is the image data obtained after taking a screenshot of the currently displayed interface.

[0096] In one alternative implementation, the user can initiate a screenshot event based on the language recognition function of the display device. Alternatively, the user can initiate a screenshot event to the controller by pressing a button related to the screenshot function on the remote control device.

[0097] In some embodiments, a screenshot application and a system running on the display device allow the controller to capture a screenshot of the current display interface via the screenshot application when a screenshot event is detected, thereby obtaining the target image. For example, the screenshot application can send a screenshot command to the system running on the display device, allowing the system to capture a screenshot of the current display interface to obtain the target image.

[0098] S520 sends an image recognition request to the server.

[0099] The image recognition request carries the target image and a trigger source identifier for the screenshot event. The trigger source identifier identifies how the screenshot event is initiated; more specifically, it identifies whether the screenshot event is triggered by voice or by a key press. In other words, the trigger source identifier can include either a voice trigger identifier or a key press trigger identifier.

[0100] It is understandable that different triggering methods will affect the subsequent guidance text. Therefore, when a screenshot event is detected, it is necessary to generate a corresponding trigger source identifier based on the triggering method of the screenshot event.

[0101] For example, the trigger source identifier can be determined based on how the screenshot event is initiated. For instance, if the screenshot event is triggered by voice, the trigger source identifier can be determined to be a voice trigger identifier. In this case, an image recognition request carrying the voice trigger identifier and the target image can be sent to the server. If the screenshot event is triggered by a button, the trigger source identifier can be determined to be a button trigger identifier. In this case, an image recognition request carrying the button trigger identifier and the target image can be sent to the server.

[0102] Alternatively, the trigger source identifier can be determined by the request parameters in the screenshot request received by the screenshot application when it launches. For example, if the request parameters include parameters indicating language initiation, the trigger source identifier is determined to be a voice trigger identifier; if the request parameters include parameters indicating key press initiation, the trigger source identifier is determined to be a key press trigger identifier.

[0103] In some embodiments, the image recognition request carrying the target image and the trigger source identifier can be directly sent to the server. Alternatively, the target image can be compressed first, and the image recognition request carrying the trigger source identifier and the compressed target image can be sent to the server.

[0104] It is worth noting that, in order to ensure the reliability of image recognition, a global identifier (such as a session ID) can be configured for each image recognition request. The server can perform corresponding processing and feedback in combination with the session ID after receiving the image recognition request sent by each display device.

[0105] S530, receiving the image recognition result and the target guide language fed back by the server based on the image recognition request.

[0106] The image recognition result is obtained by identifying the target image, and further, the image recognition result can include but is not limited to object recognition results of objects in the target image, object relationships between the objects, etc. The target guide language is the relevant vocabulary information displayed in the current display interface and used to guide the user to ask questions. Further, the target guide language is generated based on the trigger source identifier and the target content.

[0107] The target content includes at least one of the image recognition result, the business information associated with the current display interface, and the interface focus information of the current display interface. The business information is information related to the image recognition operation, for example, the image area to be recognized by the user, the image recognition type, etc., or it can also be the target interaction scene represented by the current display interface. The page focus information is the relevant information of the page focus in the current display interface, which can include but is not limited to text information, icon information or location information in the page focus, etc. Optionally, the target content can be obtained through the home application in the display device, or it can also be obtained based on the current interaction information with the user, etc.

[0108] It is worth noting that, in addition to the trigger source identifier, the target content also affects the content of the guide language, so that the corresponding guide language is different in the case of the same target content and different trigger source identifiers, and the corresponding guide language is different in the case of the same trigger source identifier and different target content.

[0109] In an optional implementation, after obtaining the image recognition request, the server can identify the target image to obtain the image recognition result. Further, after determining the target content and the trigger source identifier, the server can perform semantic analysis on the target content to obtain the guide language corresponding to the target content; and then the guide language corresponding to the target content and the standard guide language associated with the trigger source identifier can be mapped to the preset guide template, so as to generate the target guide language.

[0110] After the target guide language is determined, the server can feed back the image recognition result and the target guide language to the display device.

[0111] In S540, the display is controlled to display the image recognition result and the target guide language.

[0112] In an optional implementation, after the image recognition result and the target guide language are received, the display can be controlled to display the image recognition result and the target guide language on the current display interface. For example, the image recognition result and the target guide language can be displayed in the form of a floating layer in a fixed area on the current display interface. For example, refer to the display page shown in FIG. 6A, the image recognition result 1 and the target guide language 2 can be displayed on the left floating layer of the current display interface. Figure 6

[0113] In another optional implementation, the image recognition result and the target guide language can be displayed in the form of a pop-up window in a non-important area in the current display interface. For example, the non-important area can be determined according to the image distribution in the current display page, and then the image recognition result and the target guide language can be displayed in the form of a pop-up window in the non-important area.

[0114] It is worth noting that, in addition to displaying the target guide language in the floating layer / popup window, the target guide language can also be broadcasted. For example, in the case where the trigger source identification represents a language trigger, the target guide language can be sent to a broadcast system in the display device as a broadcast script, and the target guide language can be broadcasted to the user by the broadcast system.

[0115] In some embodiments, after the display is controlled to display the image recognition result and the target guide language, the user can combine the image recognition result and the target guide language to interact with the display device for the next step.

[0116] In the guide language generation method described above, the display device detects a screenshot event, and then takes a target image of the current display interface of the display. Then, the server performs image recognition on the target image, and generates a target guide language according to target content composed of at least one of the image recognition result, the business information and the interface focus information, and the trigger source identification. Finally, the display displays the image recognition result and the target guide language, so that the user can initiate a next query request in combination with the target guide language. Compared with the related art in which only fixed guide language is displayed to the user, the above method can make the target guide language adapt to the current display interface, the image recognition result and the screenshot initiation mode by instructing the server to generate the target guide language according to the target content and the trigger source identification, thereby ensuring the reliability of the target guide language, laying a reliable foundation for subsequent guiding the user to interact with the display device for the next step, and improving the user experience of the user interaction. ​

[0117] On the basis of the above-mentioned embodiments, in some embodiments, step S510 is further refined. Specifically, it includes the following contents: receiving a pressing instruction; or, detecting that a screenshot control in the current display interface is triggered; or, receiving a first user voice representing a screenshot intention.

[0118] The pressing instruction represents that a preset key in the remote controller is pressed. For example, it can be a screenshot control. The first user voice is voice information that can represent a user screenshot demand.

[0119] In an optional implementation, in a case where the remote controller sends a pressing instruction of a preset key, it is determined that a screenshot event is detected. For example, after the whole system of the display device receives a key signal, the key signal can be identified. In a case where the key signal is identified as a pressing event (down event) of a screenshot key, the screenshot event is initiated.

[0120] In another optional implementation, in a case where a screenshot control in the current display interface is detected to be triggered, it is determined that a screenshot event is detected. For example, in a case where the page focus of the current display interface is the screenshot control, if a pressing instruction of a confirmation key in the remote controller is detected, the screenshot event is initiated. Or, in a case where the display supports touch screen operation, in a case where a click operation of the user on the screenshot control is detected, the screenshot event is initiated.

[0121] In yet another optional implementation, in a case where a first user voice representing a screenshot intention is received, it is determined that a screenshot event is detected. For example, a language application in the display device can perform real-time identification on the user voice. In a case where a voice representing a user screenshot intention is identified, the screenshot event can be initiated. For example, in a case where the first user voice contains words such as “capture” and “query” that contain a screenshot intention, it can be determined that the first user voice represents a screenshot intention. Or, semantic identification is performed on the first user voice. In a case where the semantic identification result represents a screenshot intention, it is determined that the first user voice represents a screenshot intention.

[0122] In some embodiments, the user is provided with multiple selectable screenshot triggering modes, which can ensure the flexibility of the initiation of the screenshot event and improve the user interaction experience.

[0123] On the basis of the above-mentioned embodiments, in some embodiments, in a case where the screenshot event is triggered based on a first user voice representing a screenshot intention, step S520 is further refined. As shown in Figure 7 , it specifically includes the following contents:

[0124] S710, receiving a second user voice.

[0125] The second user voice is the user voice information collected after triggering the screenshot event.

[0126] In an optional embodiment, after triggering the screenshot event based on the first user voice, the target image obtained can be pre-stored, and the second user voice subsequently issued by the user can be continuously collected by the voice collection device, so as to determine whether the user has an image recognition intention.

[0127] Notably, in order to ensure the reliability of the second user voice collection, when the first user voice is collected, the corresponding conversation identifier (conversation ID) can be assigned to the first user voice according to the voiceprint information in the user voice. Then, when the subsequent user voice is collected, if the conversation ID does not change (the voiceprint information does not change), the subsequent user voice is taken as the second user voice; if the conversation ID changes, the subsequent user voice is taken as a new first user voice.

[0128] S720, voice recognition is performed on the second user voice.

[0129] In an optional embodiment, the second user voice can be subjected to voice-to-text processing first to obtain voice text information corresponding to the second user voice. Then, a large language model is used to perform semantic recognition on the voice text information to obtain a voice recognition result.

[0130] In another optional embodiment, the second user voice can be input into a trained voice recognition model, and the voice recognition model outputs a voice recognition result according to the second user voice and model parameters.

[0131] In some embodiments, the second user voice can be sent to a server, the server can recognize the second user voice, and the server can feed back a voice recognition result.

[0132] S730, in a case where the voice recognition result includes a target image recognition type for the current display interface, an image recognition request is sent to the server.

[0133] The target image recognition type is an image recognition type included in the second user voice, specifically an intention type for image recognition of the target image, for example, which can include but is not limited to image recognition types such as a person, a scene, and an article. Notably, in a case where the second user voice includes a target image recognition type for the current display interface, in order to ensure the accuracy of image recognition, the image recognition request initiated by the controller further includes business information associated with the current display interface, wherein the business information includes the target image recognition type for the current display interface.

[0134] In an optional implementation, in a case where it is determined that the speech recognition result contains the target image recognition type, it is proved that the user has the image recognition demand, at this time, the server can be sent an image recognition request carrying the business information, the target image and the trigger source identification. For example, in a case where it is determined that the speech recognition result contains the target image recognition type of identifying the person in the target image, the target image recognition type of identifying the person in the target image can be taken as the business information, and the target image and the trigger source identification form the image recognition request, and are sent to the server.

[0135] S740, in a case where the speech recognition result does not contain the target image recognition type, the target image is deleted.

[0136] In an optional implementation, in a case where it is determined that the speech recognition result does not contain the target image recognition type, the pre-stored target image can be deleted.

[0137] In some embodiments, in a case where the screenshot event is triggered based on the first user speech representing the screenshot intention, the second user speech of the user is continuously recognized, and in a case where the target image recognition type exists in the speech recognition result, the image recognition request is sent to the server, which can ensure that the user recognition processing is adapted to the user intention, and the waste of computing resources is avoided.

[0138] On the basis of the above-mentioned embodiments, in some embodiments, the image recognition request further includes the business information associated with the current display interface and / or the interface focus information of the current display interface, and the business information includes the target interaction scene represented by the current display interface. Based on this, an information acquisition manner is provided, as shown in Figure 8 The method specifically includes the following steps:

[0139] S810, by the home application, the target interaction scene represented by the current display interface and / or the interface focus information of the current display interface are acquired.

[0140] The so-called home application is an application in the display device that is currently in visual interaction with the user. For example, it can be a video playing application, etc. The so-called target interaction scene is the interaction scene in the current display interface. For example, it can be different channels in the video playing application, i.e. variety channel, video channel, etc.

[0141] In the case of detecting the screenshot event, in an optional embodiment, the target interactive scene represented by the current display interface can be acquired by the home application. For example, the interactive identifier associated with the current display interface can be acquired by the home application, and then the standard interactive scene associated with the interactive identifier can be taken as the target interactive scene. For example, in the case that the current display interface is a video display interface under a "video channel" in a video application A, the target interactive scene of the video display interface can be acquired by the video application A.

[0142] In another optional embodiment, the interface focus information of the current display interface can be acquired by the home application. For example, the page focus contained in the current display interface can be acquired by the home application, and then the page focus closest to the remote cursor can be taken as the target page focus according to the distance between the remote cursor and each page focus, and the focus information of the target page focus can be taken as the interface focus information.

[0143] In yet another optional embodiment, the target interactive scene represented by the current display interface and the interface focus information of the current display interface can be acquired by the home application simultaneously according to the above-mentioned manner.

[0144] In addition, when the current display interface is playing the media asset, the media asset information of the media asset in the current display interface can be acquired by the home application, and the media asset information can be sent to the server. At this time, the server can generate the target guide language according to the media asset information.

[0145] In some embodiments, on the one hand, the target interactive scene and / or the interface focus information represented by the current display interface can be acquired by the home application, which can ensure the reliability of information acquisition; on the other hand, the target interactive scene and / or the interface focus information represented by the current display interface can be carried into the image recognition request, which can provide rich content for subsequent generation of diversified target guide language.

[0146] In some embodiments, another guide language generation method is provided. Taking the case that the guide language generation method is applied to a processor in a server as an example, the server includes a communication device and at least one processor. The communication device is configured to be in communication connection with a display device, which can be wired communication connection or wireless communication connection, and the processor is connected with the communication device. As shown in Figure 9 the specific steps include:

[0147] S910, receiving the image recognition request sent by the display device.

[0148] The image recognition request carries the target image and a trigger source identifier of the screenshot event. The target image is obtained by taking a screenshot of the current display interface when the display device detects the screenshot event. The trigger source identifier is used to identify whether the screenshot event is triggered by voice or a key.

[0149] In an optional implementation, when the display device detects the screenshot event, the display device can take a screenshot of the current display interface to obtain a target image. Then, the display device can send an image recognition request carrying the target image and a trigger source identifier of the screenshot event to the server. At this time, the server can receive the image recognition request sent by the display device.

[0150] It is worth noting that the display device can also obtain the target interactive scenario represented by the current display interface and / or the interface focus information of the current display interface through the home application. Then, the display device can add the target interactive scenario and / or the interface focus information to the image recognition request and send the image recognition request to the server.

[0151] S920, identifying the target image to obtain an image recognition result.

[0152] In an optional implementation, the target image can be matched with each standard image in a preset image library, and the image information associated with the standard image with the highest matching degree can be taken as the image recognition result.

[0153] In another optional implementation, each image feature in the target image can be extracted first, and each image feature can be identified respectively. Then, the image recognition result can be determined in combination with the identification results of the image features.

[0154] S930, generating a target guide sentence according to the trigger source identifier and the target content.

[0155] The target content includes at least one of the image recognition result, the business information associated with the current display interface, and the interface focus information of the current display interface. It is worth noting that the corresponding guide sentences are different in the case of the same target content and different trigger source identifiers, and the corresponding guide sentences are different in the case of the same trigger source identifier and different target content.

[0156] In an optional implementation, in the case where the target content only includes the image recognition result, the target guide sentence can be generated according to the image recognition result and the trigger source identifier. For example, the object information of an important object in the image recognition result can be extracted. Then, the object information can be processed by using the guide mode associated with the trigger source identifier to obtain the target guide sentence.

[0157] In another optional implementation, in the case that the target content only contains business information, the target guide language can be generated according to the business information and the trigger source identifier. For example, the initial guide language under the business information can be determined according to the historical query record of the user under the business information; then, the initial guide language under the business information is processed by using the guide mode associated with the trigger source identifier, to obtain the target guide language.

[0158] In another optional implementation, in the case that the target content only contains interface focus information, the target guide language can be generated according to the interface focus information and the trigger source identifier. For example, the interface focus information can be semantically analyzed to obtain the initial guide language associated with the interface focus information; then, the initial guide language associated with the interface focus information is processed by using the guide mode associated with the trigger source identifier, to obtain the target guide language.

[0159] In another optional implementation, in the case that the target content contains image recognition result and business information, the target guide language can be generated according to the trigger source identifier, the image recognition result and the business information. In the case that the target content contains image recognition result and interface focus information, the target guide language can be generated according to the trigger source identifier, the image recognition result and the interface focus information. In the case that the target content contains business information and interface focus information, the target guide language can be generated according to the trigger source identifier, the business information and the interface focus information. In the case that the target content contains image recognition result, business information and interface focus information, the target guide language can be generated according to the trigger source identifier, the image recognition result, the business information and the interface focus information, and so on.

[0160] S940, feeding back the image recognition result and the target guide language to the display device.

[0161] The image recognition result and the target guide language are used for the display device to update and display the current display interface.

[0162] In an optional implementation, after the target guide language is determined, the image recognition result and the target guide language can be fed back to the display device.

[0163] After receiving the image recognition result and the target guide language, the display device can control the display to display the image recognition result and the target guide language in the form of a floating layer in a fixed region on the current display interface. Alternatively, the display device can control the display to display the image recognition result and the target guide language in the form of a pop-up window in a non-important region in the current display interface. In addition, in the case that the trigger source identifier is a standard voice trigger, the target guide language can also be broadcasted by a broadcast system.

[0164] In the guiding sentence generation method, after receiving the image recognition request sent by the display device and carrying the target image and the trigger source identifier, the server identifies the target image to obtain an image recognition result, and generates a target guiding sentence according to the trigger source identifier, the image recognition result, service information, and interface focus information. Then, the server feeds back the image recognition result and the target guiding sentence to the display device, so that the display device updates and displays the current display interface. Compared with the related art in which only fixed guiding sentences are displayed to the user, the above method can adapt the target guiding sentence to the current display interface, the image recognition result, and the screenshot initiation mode, thereby ensuring the reliability of the target guiding sentence, laying a reliable foundation for guiding the user to perform the next interaction, and improving the experience of user interaction.

[0165] On the basis of the above embodiment, in some embodiments, step S930 is further refined. As shown in the following steps: Figure 10

[0166] S1010, find the auxiliary guiding sentence corresponding to the trigger source identifier from the first mapping relationship.

[0167] The first mapping relationship refers to the corresponding relationship between different trigger source identifiers and different guiding sentences. The auxiliary guiding sentence is a guiding sentence with an auxiliary function.

[0168] In an optional implementation, the guiding sentence corresponding to each trigger source identifier can be configured first, thereby constructing the first mapping relationship. For example, when the trigger source identifier represents voice triggering, the corresponding guiding sentence can be "please try to say XXX"; when the trigger source identifier represents key triggering, the corresponding guiding sentence can be "please initiate voice interaction and say XXX", so as to guide the user to perform voice interaction.

[0169] After receiving the image recognition request, the trigger source identifier can be used as an index to query the first mapping relationship, thereby obtaining the guiding word associated with the trigger source identifier, and taking the guiding word as the auxiliary guiding sentence.

[0170] S1020, generate the guiding sentence body of the target guiding sentence according to the target content.

[0171] The guiding sentence body is the main guiding content in the target guiding sentence.

[0172] ​In an optional implementation, the user guidance direction can be determined according to at least one of the image recognition result, the service information, and the interface focus information, so as to obtain the guidance speech subject of the target guidance speech. For example, the image recognition result, the service information, and the interface focus information can be input into a trained semantic analysis model, and the semantic analysis model can output a semantic analysis result according to the image recognition result, the service information, and the interface focus information, and the semantic analysis result can be used as the guidance speech subject.

[0173] In another optional implementation, at least one of the image recognition result, the service information, and the interface focus information can be used as index information to query a pre-constructed guidance subject correspondence relationship, so as to obtain the guidance speech subject of the target guidance speech.

[0174] S1030, the guidance speech subject and the found auxiliary guidance speech are spliced to obtain the target guidance speech.

[0175] In an optional implementation, the guidance speech subject and the found auxiliary guidance speech can be mapped to corresponding positions in the guidance speech template, so as to obtain the target guidance speech. Alternatively, the guidance speech subject and the found auxiliary guidance speech can be directly spliced, so as to obtain the target guidance speech.

[0176] In some embodiments, by determining the auxiliary guidance speech according to the trigger source identifier, determining the guidance speech subject according to the target content, and constructing the target guidance speech according to the guidance speech subject and the auxiliary guidance speech, the adaptability of the target guidance speech to the triggering mode and the guidance scene can be ensured, and the accuracy of the target guidance speech determination is improved.

[0177] On the basis of the above embodiments, in some embodiments, the service information includes a target image recognition type for the current display interface. Step S1020 is further refined. Specifically, the guidance speech corresponding to the target image recognition type is found from the second mapping relationship; and the guidance speech subject of the target guidance speech is determined based on the found guidance speech.

[0178] The second mapping relationship is a guidance speech correspondence relationship in the image recognition type dimension, and further, the second mapping relationship includes the correspondence between different image recognition types and different guidance speeches.

[0179] It can be understood that, in order to ensure the efficiency of the subject of the guide language, for each image recognition type, the guide language corresponding to the image recognition type can be determined in advance. For example, in the case of the image recognition type being a person, the guide language can be "personnel introduction", "personnel recent dynamics", etc.; in the case of the image recognition type being a scene, the guide language can be "geographical position", "travel strategy", etc. Then, the second mapping relationship can be constructed according to the guide language corresponding to each image recognition type.

[0180] In an optional implementation, the target image recognition type can be taken as an index to query in the second mapping relationship to obtain a first query result. Then, in the case that the first query result contains only one guide language, the guide language contained in the first query result can be directly taken as the subject of the target guide language.

[0181] In another optional implementation, in the case that the first query result contains multiple guide languages, the guide language with the highest usage frequency can be selected from the guide languages according to the recent usage frequencies of the guide languages as the subject of the target guide language.

[0182] In some embodiments, the second mapping relationship including the corresponding relationship between different image recognition types and different guide languages is introduced, and then the subject of the guide language associated with the target image recognition type can be obtained by querying in the second mapping relationship. On the one hand, a convenient way is provided for quickly generating a guide language; on the other hand, different guide languages are configured for different image recognition types, which enriches the image recognition scene and the interaction scene and further improves the user experience.

[0183] On the basis of the above embodiments, in some embodiments, the business information includes a target interaction scene represented by a current display interface. Step S1020 is further refined. Specifically, the guide language corresponding to the target interaction scene is found from the third mapping relationship; and the subject of the target guide language is determined based on the found guide language.

[0184] The third mapping relationship is a guide language corresponding relationship in the dimension of an interaction scene. Further, the third mapping relationship includes the corresponding relationship between different interaction scenes and different guide languages.

[0185] It can be understood that, in order to ensure the efficiency of the subject of the guide language, for each interaction scene, the guide language corresponding to the interaction scene can be determined in advance. For example, in the case of the interaction scene being a video channel, the guide language can be "Is this drama good?" "Do you like this drama?"; in the case of the interaction scene being a music channel, the guide language can be "Is this music good to listen to?".

[0186] In an optional implementation, the target interactive scenario can be taken as an index to query in the third mapping relationship to obtain a second query result. Then, in a case where the second query result contains only one guide language, the guide language contained in the second query result can be directly taken as the guide language subject of the target guide language.

[0187] In another optional implementation, in a case where the second query result contains multiple guide languages, the guide language with the highest usage frequency can be selected from the guide languages as the guide language subject of the target guide language according to the recent usage frequencies of the guide languages.

[0188] In some embodiments, the third mapping relationship including the correspondence between different interactive scenarios and different guide languages is introduced, and then the guide language subject associated with the target interactive scenario can be obtained by querying in the third mapping relationship. On the one hand, this provides a convenient way for quickly generating guide languages. On the other hand, different guide languages are configured for different interactive scenarios, which enriches the interactive scenarios of the display device and the user and further improves the user experience.

[0189] On the basis of the above embodiments, in some embodiments, step S1020 is further refined. As shown in Figure 11 the specific steps include the following steps:

[0190] S1110, obtaining media asset information of a media asset associated with a current display interface.

[0191] The media asset information is information related to a media asset presented in the current display interface.

[0192] In an optional implementation, the display device can obtain a media asset identifier of the media asset in the current display interface through the home application and send the media asset identifier to the server. After receiving the media asset identifier, the server can query in the media asset library using the media asset identifier to obtain the media asset information of the media asset associated with the current display interface.

[0193] In another optional implementation, the display device can directly obtain the media asset information of the media asset in the current display interface through the home application and send the media asset information to the server.

[0194] S1120, analyzing the image recognition result to obtain a target object.

[0195] The target object is a person object with the highest importance in the target image. For example, it can be a person object with the highest image proportion, or a person object in a central position.

[0196] In an optional implementation, the image recognition result can be analyzed to obtain image proportions of each character object in the target image; then, the character object with the highest image proportion can be taken as the target object.

[0197] In another optional implementation, the image recognition result can be analyzed to obtain position information of each character object in the target image; then, the character object located at the center position can be taken as the target object.

[0198] S1130, determining role information of the target object in the media content according to the media content information.

[0199] The role information refers to information about a role played by the target object in the media content. For example, the role information can include, but is not limited to, role name, role introduction, and the like.

[0200] In an optional implementation, real character information of the target object can be determined according to the image recognition result; then, the real character information is used as an index to query the media content information, so as to obtain the role information.

[0201] S1140, generating a subject of the target guide language according to the role information and the media content information.

[0202] In an optional implementation, the role information and the media content information can be fused to obtain the subject of the target guide language. For example, when the role information is role B and the media content information is episode C, the subject of the target guide language can be determined as "the ending of role B in episode C".

[0203] In some embodiments, by generating the subject of the target guide language according to the role information and the media content information in the current display interface, the generated guide language can be adapted to the current display interface, thereby increasing the interest of the user in further watching or browsing the current display interface, improving the user experience, and widening the product.

[0204] On the basis of the above embodiments, in some embodiments, step S1020 is further refined. As shown in the following table, the step S1020 specifically includes the following steps: Figure 12

[0205] S1210, determining an interface word corresponding to the interface focus in the current display interface according to the interface focus information.

[0206] The interface word refers to word information associated with the interface focus in the current display interface.

[0207] ​In one alternative implementation, the focus position of the current display focus can be determined based on the interface focus information; then, the interface word corresponding to the interface focus can be located from the interface description information associated with the current display interface based on the focus position.

[0208] S1220: Based on the determined interface keywords, determine the content of interest for the currently displayed interface.

[0209] The so-called "content of interest" refers to the main content that is currently displayed on the screen.

[0210] In one optional implementation, semantic parsing can be performed on the determined interface terms to obtain the content of interest for the currently displayed interface. Alternatively, the interface terms can be used as indexes to query the interface term correspondence relationship, thereby obtaining the content of interest for the currently displayed interface. The interface term correspondence relationship includes the correspondence between different interface terms and different content of interest.

[0211] S1230, Based on the content of interest, generate the main body of the target guidance text.

[0212] In one alternative implementation, the content of interest can be directly used as the main body of the target prompt. Alternatively, if the number of fields corresponding to the content of interest is large, the content of interest can be simplified to obtain the main body of the target prompt.

[0213] For example, if the current display is a football match, and the interface focus is determined to be "statistics," then the focus of the current display is the performance of the football players in the match. In this case, the guiding text can be set as "How did this player score?"

[0214] In some embodiments, the main body of the target guidance message is generated based on the interface word corresponding to the current focus of the displayed interface. Since the interface focus usually represents the content that the user is currently paying attention to, generating guidance messages based on the interface focus can make the generated target guidance messages adapt to the user's actual needs, thereby improving the rationality of the determination of the target guidance messages.

[0215] Furthermore, based on the above embodiments, this application provides several optional methods for determining the subject of the guiding text. In one optional implementation, when the target content includes image recognition results and the target image recognition type of the currently displayed interface, the guiding text associated with the target image recognition type can be first searched from the second mapping relationship. Then, the guiding text associated with the target image recognition type can be optimized using the image recognition results to obtain the target guiding text.

[0216] In another optional implementation, if the target content package contains image recognition results and the target interaction scenario represented by the current display interface, the guide text corresponding to the target interaction scenario can be found first from the third mapping relationship. Then, the guide text corresponding to the target interaction scenario can be optimized using the image recognition results to obtain the target guide text.

[0217] In another alternative implementation, if the target content package contains image recognition results and interface focus information, the interface words corresponding to the current display focus can be determined first based on the interface focus information, and the content of interest for the current display can be determined based on the determined interface words; then, the image recognition results are processed based on the content of interest to obtain the target guidance text.

[0218] In another alternative implementation, when the target content includes business information (target image recognition type and / or target interaction scenario) and interface focus information, the guiding words corresponding to the target image recognition type and / or target interaction scenario can be determined first, and the content of interest on the currently displayed interface can be determined based on the interface focus information; then, the guiding words corresponding to the target image recognition type and / or target interaction scenario can be processed using the content of interest to obtain the target guiding text.

[0219] In another alternative implementation, when the target content includes image recognition results, business information (target image recognition type and / or target interaction scenario) and interface focus information, the focus content can be used to process the image recognition results and the guiding words corresponding to the target image recognition type and / or target interaction scenario to obtain the target guiding text.

[0220] Furthermore, if the current display interface contains media assets, the media asset information can be used to optimize the target prompt obtained by the above method, thereby obtaining the final target prompt.

[0221] Based on the technical solutions of the above embodiments, the display device can be further refined. The display device runs voice applications, screenshot applications, and the overall system. Based on this, the process of generating the introductory text is further refined. See also... Figure 13 The diagram shown illustrates the interaction between devices during the introductory message generation process, including:

[0222] S1301: When the system receives a press command or detects that the screenshot control in the currently displayed interface has been triggered, a screenshot request is sent to the screenshot application. Alternatively, when the voice application receives a first user voice indicating an intention to take a screenshot, a screenshot request is sent to the screenshot application.

[0223] S1302, the screenshot application takes a screenshot of the current display interface of the monitor, obtains the target image, and sends an image recognition request carrying the target image and the trigger source identifier of the screenshot event to the server.

[0224] S1303, the screenshot application obtains the target interaction scene and the interface focus information of the currently displayed interface through the home application, and sends the target interaction scene and interface focus information to the server.

[0225] The screenshot application interacts with the system to determine the home application associated with the current display interface; then, through the home application, it obtains the target interaction scene represented by the current display interface and the interface focus information of the current display interface.

[0226] S1304, the server identifies the target image and obtains the image recognition result.

[0227] S1305, the server finds the auxiliary guidance text corresponding to the trigger source identifier from the first mapping relationship, and generates the guidance text body of the target guidance text based on the image recognition result, the target interaction scene and the interface focus information.

[0228] S1306, the server concatenates the main body of the guiding text and the found auxiliary guiding text to obtain the target guiding text.

[0229] S1307, the server sends the image recognition results and target guidance text to the screenshot application, and also sends the target guidance text to the voice application.

[0230] S1308, the screenshot application displays the image recognition results and target guidance text.

[0231] S1309, the voice application broadcasts the target guidance.

[0232] The specific processes of S1301-S1309 described above can be found in the description of the above method embodiments. Their implementation principles and technical effects are similar, and will not be repeated here.

[0233] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0234] Based on the same inventive concept, this application also provides a prompt generation apparatus for implementing the prompt generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more prompt generation apparatus embodiments provided below can be found in the limitations of the prompt generation method described above, and will not be repeated here.

[0235] In one exemplary embodiment, such as Figure 14 As shown, a prompt generation device 14 is provided, configured in a display device, including: a screenshot module 1410, a request sending module 1420, a prompt receiving module 1430, and a display module 1440, wherein:

[0236] The screenshot module 1410 is used to capture the current display interface of the monitor and obtain the target image when a screenshot event is detected.

[0237] The request sending module 1420 is used to send an image recognition request to the server; wherein, the image recognition request carries a target image and a trigger source identifier for the screenshot event; the trigger source identifier is used to identify whether the screenshot event is triggered by voice or by key press;

[0238] The guidance message receiving module 1430 is used to receive the image recognition result and target guidance message fed back by the server based on the image recognition request. The image recognition result is obtained by recognizing the target image, and the target guidance message is generated based on the trigger source identifier and target content. The target content includes at least one of the following: the image recognition result, business information associated with the currently displayed interface, and interface focus information of the currently displayed interface. Different guidance messages are provided when the target content is the same but the trigger source identifier is different; different guidance messages are provided when the trigger source identifier is the same but the target content is different.

[0239] Display module 1440 is used to control the display to show image recognition results and target guidance text.

[0240] In one exemplary embodiment, the screenshot module 1410 is specifically used for:

[0241] A press command is received; wherein the press command indicates that a preset button in the remote control is pressed; or, a screenshot control in the current display interface is detected to be triggered; or, a first user voice indicating a screenshot intention is received.

[0242] In an exemplary embodiment, when the screenshot event is triggered based on a first user voice indicating the intent to take a screenshot, the request sending module 1420 is specifically configured to:

[0243] Receive a second user's voice; perform speech recognition on the second user's voice; if the speech recognition result includes the target image recognition type for the current display interface, send an image recognition request to the server; wherein, the image recognition request also includes business information associated with the current display interface, and the business information includes the target image recognition type for the current display interface; if the speech recognition result does not include the target image recognition type, delete the target image.

[0244] In an exemplary embodiment, the image recognition request further includes business information associated with the current display interface and / or interface focus information of the current display interface, wherein the business information includes the target interaction scenario represented by the current display interface; the guidance generation device 14 further includes an acquisition module, wherein the acquisition module is specifically used for:

[0245] The home application is used to obtain the target interaction scenario represented by the currently displayed interface and / or the interface focus information of the currently displayed interface; where the home application is the application on the display device that is currently interacting visually with the user.

[0246] In one exemplary embodiment, such as Figure 15 As shown, a prompt generation device 15 is provided, configured on a server, including: a request receiving module 1510, an image recognition module 1520, a prompt generation module 1530, and a feedback module 1540, wherein:

[0247] The request receiving module 1510 is used to receive an image recognition request sent by the display device; wherein, the image recognition request carries a target image and a trigger source identifier for a screenshot event; the target image is obtained by the display device taking a screenshot of the current display interface when it detects a screenshot event; the trigger source identifier is used to identify whether the screenshot event is triggered by voice or by a key.

[0248] The image recognition module 1520 is used to recognize the target image and obtain the image recognition result;

[0249] The prompt generation module 1530 is used to generate a target prompt based on the trigger source identifier and the target content; wherein, the target content includes at least one of the following: image recognition result, business information associated with the current display interface, and interface focus information of the current display interface; the corresponding prompts are different when the target content is the same but the trigger source identifier is different; the corresponding prompts are different when the trigger source identifier is the same but the target content is different.

[0250] The feedback module 1540 is used to provide the display device with image recognition results and target guidance text; wherein, the image recognition results and target guidance text are used by the display device to update the current display interface.

[0251] In an exemplary embodiment, the prompt generation module 1530 is specifically used for:

[0252] From the first mapping relationship, find the auxiliary guidance text corresponding to the trigger source identifier; wherein, the first mapping relationship includes the correspondence between different trigger source identifiers and different guidance texts; generate the guidance text body of the target guidance text according to the target content; and concatenate the guidance text body and the found auxiliary guidance text to obtain the target guidance text.

[0253] In one exemplary embodiment, the business information includes the target image recognition type for the current display interface; the prompt generation module 1530 is further configured to:

[0254] From the second mapping relationship, find the guiding text corresponding to the target image recognition type; wherein, the second mapping relationship includes the correspondence between different image recognition types and different guiding texts; based on the found guiding texts, determine the guiding text subject of the target guiding text.

[0255] In an exemplary embodiment, the business information includes the target interaction scenario represented by the current display interface; the guidance generation module 1530 is specifically used for:

[0256] From the third mapping relationship, find the corresponding guidance text for the target interaction scenario; where the third mapping relationship includes the correspondence between different interaction scenarios and different guidance texts; based on the found guidance texts, determine the guidance text subject of the target guidance text.

[0257] In an exemplary embodiment, the prompt generation module 1530 is specifically used for:

[0258] Obtain media asset information associated with the currently displayed interface; analyze the image recognition results to obtain the target object; determine the role information of the target object in the media asset based on the media asset information; generate the main body of the target guidance text based on the role information and the media asset information.

[0259] In an exemplary embodiment, the prompt generation module 1530 is specifically used for:

[0260] Based on the interface focus information, determine the interface word corresponding to the current display interface focus; based on the determined interface word, determine the content to be focused on in the current display interface; based on the content to be focused on, generate the main body of the target guidance message.

[0261] Each module in the aforementioned prompt generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer as software.

[0262] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0263] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0264] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0265] It should be noted that the data involved in this application (including but not limited to introductory text data) is all data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0266] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0267] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0268] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A display device, characterized in that, include: Displays and controllers; The controller is configured as follows: Upon detecting a screenshot event, a screenshot is taken of the current display interface of the monitor to obtain the target image; Send an image recognition request to the server; wherein the image recognition request carries the target image and the trigger source identifier of the screenshot event; the trigger source identifier is used to identify whether the screenshot event is triggered by voice or by key press; The system receives image recognition results and target guidance text from the server based on the image recognition request; wherein the image recognition results are obtained by recognizing the target image, and the target guidance text is generated based on the trigger source identifier and target content; the target content includes at least one of the image recognition results, business information associated with the current display interface, and interface focus information of the current display interface; the corresponding guidance text is different when the target content is the same but the trigger source identifier is different; the corresponding guidance text is different when the trigger source identifier is the same but the target content is different. Control the display to show the image recognition results and target guidance text.

2. The display device according to claim 1, characterized in that, When the controller detects a screenshot event, it is configured as follows: A press command is received; wherein the press command indicates that a preset button on the remote control is pressed; or, The screenshot control in the currently displayed interface has been triggered; or, Receive the first user's voice representing the intent of the screenshot.

3. The display device according to claim 2, characterized in that, When the screenshot event is triggered by a first user voice representing the intent to take a screenshot, the controller is configured to send an image recognition request to the server as follows: Receive voice from a second user; Perform speech recognition on the second user's voice; If the speech recognition result includes the target image recognition type for the current display interface, an image recognition request is sent to the server; wherein, the image recognition request also includes business information associated with the current display interface, and the business information includes the target image recognition type for the current display interface; If the speech recognition result does not include the target image recognition type, the target image is deleted.

4. The display device according to any one of claims 1-3, characterized in that, The image recognition request also includes business information associated with the current display interface and / or interface focus information of the current display interface, wherein the business information includes the target interaction scenario represented by the current display interface; the controller is further configured to: The home application is used to obtain the target interaction scenario represented by the currently displayed interface and / or the interface focus information of the currently displayed interface; wherein, the home application is the application on the display device that is currently interacting visually with the user.

5. A server, characterized in that, include: A communication device configured to communicate with a display device; and at least one processor, connected to the communication device, and configured to: The system receives an image recognition request sent by the display device; wherein the image recognition request carries a target image and a trigger source identifier for a screenshot event; the target image is obtained by taking a screenshot of the current display interface when the display device detects a screenshot event; the trigger source identifier is used to identify whether the screenshot event is triggered by voice or by a key. The target image is identified to obtain an image recognition result; and, Based on the trigger source identifier and target content, a target prompt is generated; wherein, the target content includes at least one of the image recognition result, the business information associated with the current display interface, and the interface focus information of the current display interface; the corresponding prompts are different when the target content is the same but the trigger source identifier is different; the corresponding prompts are different when the trigger source identifier is the same but the target content is different. The image recognition result and the target guidance message are fed back to the display device; wherein, the image recognition result and the target guidance message are used by the display device to update the current display interface.

6. The server according to claim 5, characterized in that, When the processor generates the target prompt based on the trigger source identifier and the target content, it is configured as follows: From the first mapping relationship, find the auxiliary guidance text corresponding to the trigger source identifier; wherein, the first mapping relationship includes the correspondence between different trigger source identifiers and different guidance texts; Based on the target content, generate the main body of the target guidance text; The target guidance message is obtained by combining the main guidance message and the found auxiliary guidance messages.

7. The server according to claim 6, characterized in that, The business information includes the target image recognition type for the currently displayed interface; When the processor executes the process of generating the main body of the target introductory text based on the target content, it is configured as follows: From the second mapping relationship, find the guiding text corresponding to the target image recognition type; wherein, the second mapping relationship includes the correspondence between different image recognition types and different guiding texts; Based on the found guiding text, determine the guiding text subject of the target guiding text.

8. The server according to claim 6, characterized in that, The business information includes the target interaction scenario represented by the current display interface; When the processor executes the process of generating the main body of the target introductory text based on the target content, it is configured as follows: From the third mapping relationship, find the corresponding prompt for the target interaction scenario; wherein, the third mapping relationship includes the correspondence between different interaction scenarios and different prompts; Based on the found guiding text, determine the guiding text subject of the target guiding text.

9. The server according to claim 6, characterized in that, When the processor executes the process of generating the main body of the target introductory text based on the target content, it is configured as follows: Obtain the media asset information associated with the media assets of the currently displayed interface; The image recognition results are analyzed to obtain the target object; Based on the media asset information, determine the role information of the target object in the media asset; Based on the role information and the media asset information, generate the main body of the target introductory text.

10. The server according to claim 6, characterized in that, When the processor executes the process of generating the main body of the target introductory text based on the target content, it is configured as follows: Based on the interface focus information, determine the interface word corresponding to the interface focus in the currently displayed interface; Based on the determined interface words, determine the content of interest for the currently displayed interface; Based on the content of interest, the main body of the target guidance message is generated.