Methods for labeling display devices and interface content

CN119166866BActive Publication Date: 2026-08-14HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本申请提供一种显示设备及界面内容标注方法,以解决标注内容的准确率低的问题

Benefits of technology

[0016]由以上技术方案可知,本申请一些实施例提供一种显示设备及界面内容标注方法,所述方法通过响应于第一请求标注指令,获取当前显示的界面图像以及用户界面的元信息,再将第一请求标注指令、界面图像以及元信息输入至多模态检索系统检索关联信息,再将关联信息、第一请求标注指令以及界面图像输入至多模态理解模型生成标注反馈信息,控制显示器在用户界面上显示标注反馈信息。所述方法通过多模态检索系统对多模态信息执行信息检索,其中,通过使用第一请求标注指令、界面图像以及元信息等作为多模态信息,可提升信息检索和理解的准确性,进而提高标注内容的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119166866B_ABST
    Figure CN119166866B_ABST
Patent Text Reader

Abstract

This application provides a display device and a method for annotating interface content. The method, in response to a first annotation request instruction, acquires the currently displayed interface image and metadata of the user interface. It then inputs the first annotation request instruction, the interface image, and the metadata into a multimodal retrieval system to retrieve related information. Finally, it inputs the related information, the first annotation request instruction, and the interface image into a multimodal understanding model to generate annotation feedback information, and controls the display to show the annotation feedback information on the user interface. This method performs information retrieval on multimodal information through a multimodal retrieval system. By using the first annotation request instruction, the interface image, and metadata as multimodal information, the accuracy of information retrieval and understanding can be improved, thereby increasing the accuracy of the annotated content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and in particular to a display device and a method for annotating interface content. Background Technology

[0002] A display device is a terminal device capable of outputting images. The output images can be either graphic or video. Taking video as an example, a user's needs when watching a video are not limited to simply watching; the video also contains different characters, scenes, and other content, and the user's needs also include understanding information about the different content within the video.

[0003] To identify the content of the output screen, text content within the output screen can be extracted and recognition can be performed using retrieval enhancement technology. This retrieval enhancement technology obtains labeled information through text retrieval, but can only perform single-modal retrieval, i.e., text information, and cannot recognize labeled content such as people or scenes within the screen.

[0004] For person recognition, computer vision and machine learning technologies can be used to identify, classify, and label content in videos. For example, if a user needs to understand the information of people in a video, images can be captured, and then the video content can be analyzed and labeled to identify the people in the video, and the information can be fed back to the user. However, this method can only process single-modal information and cannot analyze and label other content such as scenes, resulting in low accuracy of the labeled content. Summary of the Invention

[0005] This application provides a method for annotating display devices and interface content to solve the problem of low accuracy of annotated content.

[0006] In a first aspect, some embodiments of this application provide a display device, including a display and a controller. The display is configured to display a user interface; the controller is configured to:

[0007] In response to the first request annotation instruction, the currently displayed interface image and user interface metadata are obtained, wherein the first request annotation instruction is an instruction generated by the display device according to a preset period, and / or an instruction input by the user;

[0008] The first request annotation instruction, the interface image, and the metadata are input into a multimodal retrieval system to retrieve related information. The multimodal retrieval system is used to perform information retrieval based on the multimodal information, which includes keywords extracted from the first request annotation instruction, the interface image, and the metadata.

[0009] The associated information, the first request annotation instruction, and the interface image are input into the multimodal understanding model to generate annotation feedback information through the multimodal understanding model. The multimodal understanding model is a model trained and generated based on a deep learning algorithm using the multimodal information.

[0010] Control the display to show annotation feedback information on the user interface.

[0011] Secondly, some embodiments of this application also provide an interface content annotation method, applied to the display device described in the first aspect, the interface content annotation method comprising:

[0012] In response to the first request annotation instruction, the currently displayed interface image and user interface metadata are obtained, wherein the first request annotation instruction is an instruction generated by the display device according to a preset period, and / or an instruction input by the user;

[0013] The first request annotation instruction, the interface image, and the metadata are input into a multimodal retrieval system to retrieve related information. The multimodal retrieval system is used to perform information retrieval based on the multimodal information, which includes keywords extracted from the first request annotation instruction, the interface image, and the metadata.

[0014] The associated information, the first request annotation instruction, and the interface image are input into the multimodal understanding model to generate annotation feedback information through the multimodal understanding model. The multimodal understanding model is a model trained and generated based on a deep learning algorithm using the multimodal information.

[0015] Control the display to show labeled feedback information on the user interface.

[0016] As can be seen from the above technical solutions, some embodiments of this application provide a display device and a method for annotating interface content. The method, in response to a first annotation request instruction, obtains the currently displayed interface image and metadata of the user interface. It then inputs the first annotation request instruction, the interface image, and the metadata into a multimodal retrieval system to retrieve related information. Finally, it inputs the related information, the first annotation request instruction, and the interface image into a multimodal understanding model to generate annotation feedback information, and controls the display to show the annotation feedback information on the user interface. This method performs information retrieval on multimodal information through a multimodal retrieval system. By using the first annotation request instruction, the interface image, and metadata as multimodal information, the accuracy of information retrieval and understanding can be improved, thereby increasing the accuracy of the annotated content. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application;

[0019] Figure 2 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application;

[0020] Figure 3 This is a schematic diagram of the software configuration of a display device provided in some embodiments of this application;

[0021] Figure 4 This application provides schematic diagrams illustrating the recognition process performed on the output screen in some embodiments.

[0022] Figure 5 This application provides schematic flowcharts of interface content annotation methods for some embodiments;

[0023] Figure 6 This is a schematic diagram illustrating the annotation setting parameters provided in some embodiments of this application;

[0024] Figure 7 This is a first schematic diagram of the current interface image provided for some embodiments of this application;

[0025] Figure 8 This is a second schematic diagram of the current interface image provided for some embodiments of this application;

[0026] Figure 9 A flowchart illustrating the process of correcting related information provided in some embodiments of this application;

[0027] Figure 10 A schematic diagram illustrating annotation information provided for some embodiments of this application;

[0028] Figure 11 This is a schematic diagram illustrating feedback information provided for some embodiments of this application;

[0029] Figure 12 This is a schematic diagram illustrating the process of generating supplementary annotation feedback information for some embodiments of this application. Detailed Implementation

[0030] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0031] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0032] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0033] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0034] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0035] In this embodiment, the display device 200 generally refers to a device with screen display and data processing capabilities. For example, the display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0036] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, a user can operate the display device 200 via touch operation, a mobile terminal 300, and a control device 100. The control device 100 receives user input commands and converts them into control commands that the display device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.

[0037] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the display device 200. It can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the display device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the display device 200 for synchronized display.

[0038] In some embodiments, the mobile terminal 300 or other electronic devices may also simulate the functions of the control device 100 by running an application that controls the display device 200.

[0039] like Figure 1 The diagram also shows that the display device 200 communicates with the server 400 via various communication methods. This allows the display device 200 to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0040] Display device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.

[0041] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of display device 200.

[0042] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0043] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0044] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0045] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The display device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the display device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the display device 200 supports Bluetooth connectivity, it needs to have a communication device 220 with Bluetooth functionality.

[0046] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0047] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0048] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0049] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0050] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0051] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0052] To enable user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can control the display device to provide a user interface; for example, the operating system can directly control the display device 200 to provide a user interface, or it can provide a user interface by running an application. The operating system also allows users to interact with the display device 200.

[0053] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for the display device 200.

[0054] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 3 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Library layer, and the Kernel layer.

[0055] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0056] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0057] like Figure 3 As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0058] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0059] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0060] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 3As shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0061] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the functions of the display device 200, the type of the operating system, and other factors, the number of layers and the specific type of the operating system may take other forms.

[0062] Based on the aforementioned display device 200, the display device 200 can output a screen and display the screen via display 260. For improved user experience, see [link to relevant documentation]. Figure 4 The display device 200 can perform recognition and annotation on the output screen, for example, to recognize and annotate person B in the output interface.

[0063] The output screen can be either an image or a video, and the recognition methods for image and video screens can be different.

[0064] For static images, in some embodiments, image recognition algorithms can be used for image recognition, such as Convolutional Neural Networks (CNNs) and object detection algorithms. CNNs can extract image features through structures such as convolutional layers and pooling layers, and then perform classification or recognition. In the display device 200, CNNs can be used to recognize the currently displayed image content, such as faces, objects, and scenes. Through a pre-trained CNN model, feature representations of the image can be obtained and used to recognize images on the screen in real time, and then the recognized content can be labeled.

[0065] Among them, object detection algorithms can locate and identify multiple objects in an image. Combining classification and localization techniques, object detection algorithms can simultaneously identify multiple objects and their locations in an image. In the display device 200, object detection algorithms can be used to identify multiple targets in the currently displayed screen, such as participants in a meeting, moving objects in a video, etc., and then label the identified content.

[0066] In other words, for image processing, after acquiring a static image, image preprocessing is performed, such as noise reduction and contrast enhancement, followed by feature extraction. The resulting image is then classified and labeled. The labels for static images include people, objects, and scenes.

[0067] For video footage, i.e., dynamic footage, in some embodiments, computer vision and machine learning techniques can be used to perform recognition, classification, and annotation on the video footage. For example, the video file can be decoded into a series of video frames, keyframes can be extracted from the video, or frames can be sampled at fixed intervals to reduce computational load. Features can then be extracted from the video frames. These features can be low-level features such as color, texture, and shape, or high-level features such as motion and audio. Feature extraction methods include CNNs and optical flow methods. Among them, optical flow methods estimate the direction and speed of motion by analyzing the changes in pixels or feature points between adjacent frames. In the display device 200, optical flow methods are used to analyze motion information on the screen, such as detecting motion trajectories and calculating motion speed.

[0068] By training the model using a labeled video dataset, a pre-trained model can be obtained. The content to be recognized is then input into the pre-trained model to perform recognition on the video footage and output the recognition result.

[0069] In other words, for video footage, video frames can be captured first, and then static image recognition can be performed on each frame. By fusing inter-frame information, such as motion estimation and tracking, the recognition results can be output and labeled. The labeling methods for video footage can be used for behavior analysis, video surveillance, motion detection, and other applications in dynamic scenes.

[0070] Large model technology refers to deep learning models with a large number of parameters, layers, and units. These models are built from deep neural networks and possess expressive power and predictive performance. Large model technology can handle complex tasks and data, employing a pre-training and fine-tuning training model to adapt to downstream tasks. Large model technology has applications in multiple fields, such as Natural Language Processing (NLP), Computer Vision (CV), and multimodal processing. For multimodal processing, large models can simultaneously process information from multiple modalities, including text, images, and audio.

[0071] For example, multimodal models enable multimodal interactive question answering. Large model technology can be used to build intelligent question-answering systems that understand user questions and provide corresponding answers. However, intelligent question-answering systems based on large models primarily rely on the model's own knowledge reserves, limiting the accuracy, comprehensiveness, and timeliness of the answers. Retrieval Augmentation (RAG) technology can also be used. RAG searches for information related to the user's question in external knowledge bases through text retrieval to supplement the model's own knowledge reserves. However, RAG is mainly limited to single-modal retrieval, focusing primarily on text data, and is not suitable for video content analysis and annotation applications.

[0072] Video content contains images, audio, and text, which interact to form the overall meaning of the video. However, for video content analysis and annotation applications, accurate annotation information cannot be obtained solely based on text retrieval. Therefore, relying solely on text retrieval to obtain supplementary information cannot provide a comprehensive and accurate understanding of the video content, limiting the application of intelligent question-answering systems in video content analysis and annotation, and ultimately preventing accurate video content annotation.

[0073] In some embodiments, the video content analysis method using multimodal technology can obtain corresponding multimodal feature vectors by sampling and encoding multimodal information within the same time period at a specific time point in the video; using a fusion model, at least two of the image information feature vectors, audio information feature vectors, or text information feature vectors from the multimodal feature vectors are fused to obtain a fused feature vector; and using a bullet screen generation model, bullet screens at specific time points in the video are generated from the fused feature vectors. Combining multimodality, bullet screens at the current video time point are automatically generated, enhancing interactivity and enriching video content. However, this approach is limited to single-modal information from the video frame, resulting in low accuracy in video content annotation.

[0074] To address the issue of low accuracy in labeled content, this application provides a method for labeling interface content. The display device 200 uses this method to perform a retrieval based on multimodal information from a multimodal retrieval system. This multimodal information includes user-inputted instructions and / or instructions, interface images, and metadata generated by the display device 200 according to a preset cycle. The retrieved related information is then understood using a multimodal understanding model to generate accurate and comprehensive labeled content.

[0075] To facilitate the implementation of the interface content annotation method, the display device 200 includes at least a display 260 and a controller 250. The display 260 is configured to display the user interface. Figure 5 As shown, the controller 250 is configured to execute the program steps corresponding to the interface content annotation method, including the following:

[0076] S100: In response to the first request annotation instruction, obtain the currently displayed interface image and the metadata of the user interface.

[0077] The first request annotation instruction is used to trigger the display device 200 to display annotation information. In some embodiments, the first request annotation instruction is an instruction generated by the display device 200 according to a preset period. In some embodiments, the first request annotation instruction is an instruction input by the user. The first request annotation instruction generated by the display device 200 according to the preset period can be automatically generated by the system or program of the display device 200 according to a preset time period. For example, the user sets the display device 200 to annotate information on the content in the video when playing sports content. When the video being played is a football match, the display device 200 automatically initiates a video content annotation request, i.e., the first request annotation instruction.

[0078] In some embodiments, the display device 200 can also generate a first request annotation instruction based on historical interaction data. For example, if the video being played is a contemporary TV series, the display device 200 can infer from the user's historical interaction data that the user might be interested in specific information about the filming scene in the current context, and initiate a first request annotation instruction independently.

[0079] Historical interaction data refers to data records generated when a user interacts with the display device 200 before this viewing session. These data records can include user behavior, preferences, choices, and feedback. Historical interaction data can include interaction behaviors, feedback data, and settings preferences. Interaction behaviors are actions performed between the user and the device, such as clicking, swiping, and voice commands. Feedback data includes user evaluations, ratings, comments, and likes of the content. Settings preferences are personalized settings made by the user on the device, such as selecting annotation content and annotation methods.

[0080] The display device 200 can generate a first request annotation command at fixed time intervals, such as after 2 minutes of video playback; or when a TV series episode changes; the display device 200 can also generate a first request annotation command based on periodic changes triggered by a certain event, such as when the video screen is switched.

[0081] The instructions generated by the display device 200 according to a preset cycle can be set as a function switch. The function switch can also be activated by the user by setting the time interval, for example, during a specific playback application or a specific screen.

[0082] In addition to setting automatic annotations via time intervals, to improve annotation accuracy, in some embodiments, when the user obtains the first request annotation command generated by the display device 200 at fixed time intervals, annotation setting parameters can also be input. The annotation setting parameters are annotation preference setting parameters for the interface image, and the annotation setting parameters include the annotation content of the interface image; then, annotation preference information is extracted from the annotation setting parameters, and the annotation preference information includes at least one annotation type item; based on the annotation preference information, the first request annotation command is generated according to a preset period.

[0083] like Figure 6 As shown, the annotation settings parameters are defined by the user according to their preferences and needs, and are used to guide how to annotate interface images. Annotation settings parameters include, but are not limited to: annotation content, annotation type, annotation position, and annotation style.

[0084] The annotation content refers to the specified interface content, such as video title, current playback time, remaining time, actor information, and plot synopsis. The annotation type defines the specific form of the annotation, such as text annotation, icon annotation, and voice annotation. Text annotations can be displayed directly on the interface image or in the interface interacting with the user; icon annotations can be represented directly by icons or symbols; the annotation content can be read aloud via a voice assistant, whether the user is in a conversation or not. The annotation position specifies the display location of the annotation on the interface, ensuring a non-intrusive and easily viewable presentation. Annotation styles include font size, color, and background color to suit different visual needs and interface styles.

[0085] Understandably, users can further configure annotation settings by selecting parameters such as video title, text annotation, non-intrusiveness, and font size.

[0086] Users can determine the annotation type through annotation settings parameters. Annotation type can include at least one parameter. For example, if the annotation settings parameters only include annotation content, then after responding to the first request annotation instruction, only the annotation content will be annotated unless the user issues other request annotation instructions. Other annotation settings parameters are set to their default values, such as annotation style and annotation position.

[0087] In some embodiments, the video content analysis method of multimodal technology can integrate content through the application of web crawling technology, language big models, knowledge graphs and generative artificial intelligence (AIGC), which can automatically generate the latest character-themed pages, including character introductions, programs the character has participated in, related programs and background of the page, all of which are processed and synthesized by algorithms, improving the speed and frequency of character-themed page updates.

[0088] In some embodiments, a user can send instructions to the display device 200 via the control device 100, a smartphone application, or other interactive methods. For example, by controlling a remote control to move the focus to a target marker, a first request for marking instruction is generated. Another example is by clicking on a target marker via a smartphone application, generating the first request for marking instruction. Yet another example is that the user pre-sets voice or interactive button activation as the wake-up method, waking up the interactive assistant during video playback. If the video being played is a contemporary TV series, the user can wake up the interactive assistant using the pre-set wake-up method and ask, "Where is the current scene? Is it possible to visit it recently?", thereby generating the first request for marking instruction.

[0089] Unlike the instructions generated by the display device 200 according to a preset cycle, the instructions input by the user are instantaneous. By responding to the user's needs, interactivity and flexibility can be increased, and the video screen can be annotated and analyzed according to the user's request.

[0090] It is understandable that during user interface recognition, the user interface is either a displaying image or a video interface. If the user interface displays an image, the current interface image can be a screenshot of the image interface; upon receiving the first request annotation instruction, the screenshot operation is performed. If the user interface displays a video interface, upon receiving the first request annotation instruction, a specific video frame can be obtained. This video frame can be the current video frame or other video frames that include the current video frame.

[0091] In other words, in some embodiments, the currently displayed interface image can be determined based on the first request annotation instruction. If the first request annotation instruction is an instruction generated by the display device 200 according to a preset period, the currently displayed interface image can be the current video frame; if the first request annotation instruction is an instruction input by the user, the currently displayed interface image can be the current video frame, and the currently displayed interface image may also include other video frames of the current video frame.

[0092] For example, a user is watching a video and notices a certain scene. The display device 200, following instructions generated at a preset cycle, annotates or analyzes the current video frame. It can capture the currently displayed video frame and process it as the target frame, where the target frame is the current interface image.

[0093] like Figure 7 As shown, display device 200 is playing a live broadcast of a sports event and can display athlete information in the live broadcast. However, during a sports event, athletes will perform different postures and movements, and these movements occur within a very short period of time. For example, the current video frame may not be able to capture the full video. Figure 7The system retrieves information for athlete number 5, but if the athlete information cannot be obtained after receiving the first request annotation instruction, it can capture a certain number of video frames forward or backward from the current video frame, such as... Figure 8 As shown, five frames are captured forward and five frames are captured backward. One of these video frames can identify the athlete's number. Therefore, the display device 200 uses the video frame from which the athlete's number can be identified as the current video frame.

[0094] Meta-information of the user interface is a series of descriptive information related to the content being played. This meta-information includes, but is not limited to, video title, author, theme, task, release date, tags, language, and comments. Taking a video frame as an example, when the user interface displays a video frame, the video frame is the video file itself, which contains some built-in meta-information, such as author, title, creation date, encoding format, resolution, frame rate, and duration. When acquiring video files, certain third-party databases or websites also collect and organize video files and their related information. These databases can be used to search for specific video files and obtain their meta-information.

[0095] The interface type of the interface image includes video interface and image interface. The acquisition methods of the current interface image and metadata are different. In some embodiments, the interface type of the interface image is acquired, which is either a video interface or an image interface. If the interface type is a video interface, the content type of the current interface image is acquired, so as to acquire the image of the current video interface and the metadata of the video interface according to the content type. If the interface type is an image interface, the image of the image interface and the metadata of the image interface are acquired.

[0096] If the currently displayed interface image is detected to contain the identifier of a video stream or video file, it can be marked as a video interface; if a static image file or image metadata is detected, it can be marked as an image interface.

[0097] The content type is categorized into attributes that describe the video content, such as drama, sports, entertainment, and education. Identification can be achieved by analyzing the video file's metadata, title, tags, or the video content itself. After obtaining the content type, the processing can be optimized accordingly. For example, a suitable video summarization algorithm can be selected based on the content type to generate an image of the current video interface, while simultaneously acquiring the video interface's metadata. Another example is obtaining metadata about the video interface from different perspectives based on different content types.

[0098] Taking video footage as an example, the currently displayed interface image is a frame of the video, the user interface is the video file, and the metadata of the user interface can be information related to the video file, rather than specific information about a frame, such as the scene or characters.

[0099] The order in which the currently displayed interface image and user interface metadata are acquired is not limited. In some embodiments, they can be acquired in parallel, which can shorten response time and improve processing efficiency.

[0100] In some embodiments, the currently displayed interface image can be acquired first, followed by the metadata of the user interface. If an object within the interface image is the target object requested for annotation by the user, and the target object can be directly obtained, then the interface image is acquired first, followed by the metadata of the user interface, so that the user can directly view the annotation and then perform further analysis or processing. It is understood that for character annotations in directly obtainable interface images, the metadata of the user interface may not need to be acquired.

[0101] In some embodiments, the metadata of the user interface can be obtained first, followed by the currently displayed interface image. If the metadata is more important to the labeled content, it can be obtained first. This allows a decision to be made based on the content of the metadata as to whether further acquisition of the interface image is necessary.

[0102] After obtaining the currently displayed interface image and user interface metadata, the relevant content to be annotated is retrieved according to the first request annotation instruction.

[0103] S200: Input the first request annotation instruction, the interface image, and the metadata into the multimodal retrieval system to retrieve related information through the multimodal retrieval system.

[0104] When analyzing and annotating video footage, one can base it not only on video frames but also on audio information, time information, and more. Audio information provides the sound content of the video, such as dialogue, music, and ambient sounds. Audio information helps in a more comprehensive understanding of the video content. Time information can characterize the sequence of events and the duration of actions in the video. During annotation, the timestamp of each tag or event can be recorded for subsequent analysis and processing.

[0105] In some embodiments, multimodal fusion technology is employed to combine information from different modalities for analysis. For example, image and audio information from a video can be combined, and facial recognition and speech recognition technologies can be used to simultaneously identify people and dialogue content in the video. In the visual dimension, deep learning models can be used to extract image features from video frames; in the audio dimension, speech recognition technology can be used to convert audio into text and further extract text features; in the text dimension, features can be directly extracted from text information such as video titles, descriptions, or subtitles.

[0106] It can also combine time information and user needs. For example, it can continuously track objects in a video using object tracking algorithms, maintaining their consistency across different frames. Simultaneously, it can identify different scenes in a video using scene recognition algorithms and understand the transitions between them. If a user needs to search for and retrieve specific people or objects in a video, it can accurately identify and label the people and objects appearing in the video. If a user needs to understand emotional expressions or mood changes in a video, it can describe and label the emotions and feelings within the video.

[0107] However, the use of multimodal information may focus on one or two modalities, while the utilization rate of other modalities is relatively low. For example, video annotation primarily performs annotation tasks using image information. Key objects, scenes, or behaviors in the video are identified through video frame extraction and image processing techniques. While textual information is present, it may only be used as supplementary information to verify or supplement the results of image recognition.

[0108] Video annotation primarily utilizes text information to perform annotation tasks. When video content mainly involves verbal communication, image information is relatively secondary or located in scenarios where it is difficult to directly identify. For example, in the annotation of meeting minutes or court proceedings, the annotation mainly relies on transcribing the speech in the video, using natural language processing technology to identify keywords, themes, or sentiment tendencies, and generating annotation information accordingly.

[0109] In other words, the emphasis on multimodal use varies depending on the scenario, requiring the selection of different information for different scenarios. However, scenarios may change at any time, and if the required information cannot be accurately selected when the scenario changes, the video annotation content will be inaccurate.

[0110] Furthermore, when using a single modality to annotate a video, or using image information for multimodal annotation, it is necessary to extract the current video frame. In other words, by capturing the image and then extracting the features of the video frame, the annotation is performed after recognition. However, video images are in a state of real-time change, and using a single modality for annotation leads to inaccurate annotation of the video frame.

[0111] A multimodal retrieval system is used to perform information retrieval based on multimodal information. In order to obtain accurate retrieval, the construction of a multimodal retrieval system includes at least data collection and preprocessing, feature extraction and representation, retrieval model construction, system deployment and testing.

[0112] During data collection and preprocessing, various types of video data can be collected, including but not limited to movies, TV series, educational videos, and advertisements. Meta-information, such as title, author, description, tags, and time, is extracted from the video data. Image recognition and video content analysis techniques are then applied to the video frames to extract keyframes and image features.

[0113] Natural language processing techniques are used to segment and vectorize textual information such as video titles and descriptions, converting them into computable numerical features. Computer vision techniques are used to extract image features from keyframes of the video, such as color, texture, shape, and object recognition, and these image features are converted into vector or semantic space representations. Textual and image features are then fused to form a unified representation for cross-modal information retrieval.

[0114] Select a retrieval model to be trained, such as an algorithm based on a vector space model, probabilistic model, or deep learning model, and train and optimize the model. Build a multimodal retrieval engine that supports queries and retrieval based on video footage, video metadata, and user questions. Deploy the trained model and retrieval engine to a server and provide an API interface for external calls. Conduct comprehensive testing, including functional testing, performance testing, and stability testing, to ensure that the system can stably and accurately provide multimodal retrieval services.

[0115] In some embodiments, the multimodal retrieval system can be constructed based on a multimodal information knowledge base. The multimodal information knowledge base includes multimodal information, which includes all information from the first request annotation instruction, the interface image, and the metadata, or keywords extracted from the first request annotation instruction, the interface image, and the metadata. The construction process, including but not limited to data collection and preprocessing, feature extraction and representation, retrieval model construction, system deployment, and testing, can be built based on existing construction methods and will not be elaborated here. In this embodiment, the multimodal retrieval system is a retrieval system capable of retrieving related information based on the first request annotation instruction, interface image, and metadata. Furthermore, the multimodal retrieval system can directly perform retrieval based on the user's question, without needing to extract keywords from the question.

[0116] For the first request annotation instruction initiated by the display device 200, for example, the multimodal retrieval system can retrieve information such as the teams and player numbers of the opposing teams from the multimodal information knowledge base based on the images and metadata displayed in the current video frame, including information such as the opposing teams and the event and time indicated by the video title.

[0117] For user-input commands, such as those from a multimodal retrieval system, the system can retrieve information such as the filming location of a TV series from a multimodal information knowledge base based on the images displayed in the current video frame and information such as the TV series title, episode number, and time.

[0118] Once relevant information is retrieved, it can be displayed on the user interface. In some embodiments, users can determine whether to continue searching or use a multimodal understanding model for analysis based on the relevant information. In other words, when relevant information is displayed on the user interface, users can use the user interface to control the multimodal retrieval system to search again as needed, or if the relevant information cannot meet their needs, they can first use a multimodal understanding model to understand the relevant information.

[0119] like Figure 9 As shown, errors may occur in the displayed related information. To avoid errors during the understanding process, the user can correct the related information. In some embodiments, the display 260 is controlled to display the related information on the user interface; in response to the user's input focus movement command, the display 260 is controlled to display a second interactive interface on the user interface; error correction information is generated according to the error correction command input by the user based on the second interactive interface; the error correction information, the related information, the first request annotation command, and the interface image are input into the multimodal understanding model to generate annotation feedback information through the multimodal understanding model.

[0120] When a user needs to perform further operations or view more information, they can issue a focus movement command in various ways, such as by pressing buttons, moving a mouse, operating a touchscreen, or using a remote control. The display device 200 recognizes the focus movement command and controls the monitor 260 to display a second interactive interface on the user interface.

[0121] The second interactive interface can be a drop-down menu, a pop-up interface, or a dialog interface. If the user finds that the associated information is incorrect, or needs to modify certain annotation settings, they can input correction commands through the second interactive interface. The display device 200 receives and parses the correction commands, generating correction information including the proposed corrected text, instructions to re-search, or cancellation of the secondary search.

[0122] For example, after retrieving information such as the filming location of a TV series through a multimodal retrieval system, the multimodal understanding model determines that the scene of the current interface image is a scenic spot in a certain city in China. Based on the user's travel intention, the model requests information such as the recent weather conditions and travel suitability index of that city from the multimodal retrieval system.

[0123] S300: Input the association information, the first request annotation instruction, and the interface image into the multimodal understanding model to generate annotation feedback information through the multimodal understanding model.

[0124] The associated information includes associated information that has been displayed and then corrected by the user, or associated information that has not been displayed, i.e., associated information obtained directly through a multimodal retrieval system.

[0125] Multimodal understanding models are used to encode, understand, and analyze related information, initial request annotation instructions, and interface images, and generate annotation feedback information based on the analysis results. The construction of a multimodal understanding model can include data collection and preprocessing, feature extraction, and model training.

[0126] During data collection and preprocessing, data from multiple modalities, including video, text, images, and audio, are collected. The collected data is then cleaned, labeled, and formatted for subsequent model training. For example, video data requires segmentation and keyframe extraction; text data requires word segmentation and stop word removal.

[0127] Algorithms or models are used to extract features from data of each modality. For example, CNNs are used to extract visual features from images, and recurrent neural networks (RNNs) or transformers are used to extract linguistic features from text. Features from different modalities are then transformed into the same feature space to enable cross-modal comparison and analysis.

[0128] The multimodal understanding model employs large-scale models such as Transformer, BERT (Bidirectional Encoder Representations from Transformers), and GPT (Generative Pre-Trained Model) as its basic architecture, and is customized and optimized according to specific tasks. Appropriate loss functions are designed to evaluate the difference between the model's predictions and actual annotations, and model parameters are optimized using the backpropagation algorithm. The model is trained on large-scale datasets, and its parameters are continuously adjusted and optimized to improve its accuracy and generalization ability.

[0129] Model performance can also be evaluated using metrics such as accuracy, recall, and F1 score. Based on the evaluation results, the model can be tuned, including adjusting the model architecture and optimizing training strategies.

[0130] When the first request annotation instruction includes instructions generated by the display device 200 according to a preset cycle and instructions input by the user, the instructions can be sorted according to their priority or importance, with more important instructions being processed first. For example, a multimodal understanding model can first understand the instructions input by the user so that the generated annotation feedback information better meets the user's needs.

[0131] The initial annotation instructions can be further processed, such as through semantic understanding and normalization. For natural language instructions, semantic analysis can be performed to understand their meaning. Understanding the content includes identifying entities, relationships, and actions within the instructions, and understanding the logical relationships between them. If the instructions contain ambiguous or vague expressions, disambiguation can be achieved using contextual information or domain knowledge bases.

[0132] By using standardized language, such as words and terms, in the instructions as a standard vocabulary or terminology set for model training, the multimodal understanding model can accurately identify and process the words involved in these instructions. For periodically generated instructions, time-related expressions can be converted into timestamps or time interval formats that the model can understand.

[0133] To facilitate multimodal understanding models' comprehension of user interface metadata and improve the accuracy of labeled feedback, data cleaning and normalization can be performed on the metadata. Cleaning removes useless or erroneous data, such as duplicates and inconsistently formatted dates. Furthermore, the metadata is normalized into a model-understandable format; for example, dates are converted to a standardized date format, and labels and topics are converted into a standard vocabulary or terminology set used during model training.

[0134] Next, useful features are extracted from the metadata, such as keywords, sentiment, and time trends. These extracted features are then encoded into a numerical form that the model can process; for example, text information is converted into a vector representation through word embedding. The metadata is then correlated with other modal data in the user interface, such as the initial request annotation and interface images, so that the model can comprehensively consider multiple modalities for understanding and analysis.

[0135] For interface images, preprocessing can be performed, such as image cleaning, image enhancement, and resizing. Low-level features such as edges, textures, and shapes can also be extracted from the interface image. Further processing by the model can then extract higher-level semantic features from these low-level features, such as object category, location, and relationships.

[0136] Fusing image features with features from other modalities such as text and audio to form a multimodal feature representation helps the model to more comprehensively understand the content of the user interface. Based on these multimodal features, further understanding is then performed to improve the accuracy and depth of the model's understanding of the user interface.

[0137] After processing the first request annotation instructions, interface images, and metadata, annotation feedback information can be generated through multimodal understanding model.

[0138] To ensure greater accuracy in the annotation feedback information, the multimodal understanding model can be configured to generate different annotation feedback information for different content types. In some embodiments, emphasis information is set according to the content type, which characterizes the emphasis of the annotation feedback for the video footage of that content type. The associated information, the first request annotation instruction, and the interface image are input into the multimodal understanding model to generate annotation feedback information based on the emphasis information.

[0139] Different types of content have different annotation requirements. For example, educational videos may focus more on annotating knowledge points, while entertainment videos may focus more on annotating emotions or scenes.

[0140] In some embodiments, for videos on specific topics, such as technology, history, and art, the emphasis information can be set to highlight key information related to the topic. For videos with rich emotional expression, such as movie clips and music videos, the emphasis information can be set to focus on emotional tendencies, such as joy, sadness, and anger. For videos containing human actions or behaviors, such as sports competitions and dance performances, the emphasis information can be set to identify and label key actions or behaviors. For videos containing specific objects, such as product demonstrations and animal documentaries, the emphasis information can be set to detect and label key objects in the video.

[0141] Once the multimodal understanding model has completed its understanding of the first request annotation instruction, interface image, and metadata, it generates different annotation feedback information based on the biased information.

[0142] For the first request annotation instruction initiated by the display device 200, for example, the multimodal retrieval system can retrieve information such as the teams and player numbers of the opposing teams from the multimodal information knowledge base based on the images and metadata displayed in the current video frame, including information such as the opposing teams and the event and time indicated by the video title.

[0143] For user-input commands, such as those from a multimodal retrieval system, the system can retrieve information such as the filming location of a TV series from a multimodal information knowledge base based on the images displayed in the current video frame and information such as the TV series title, episode number, and time.

[0144] For example, by retrieving information such as the teams and player numbers of the two opposing sides through a multimodal retrieval system, the multimodal understanding model can annotate information such as the name of the player holding the ball in the interface image based on the information provided above and the current interface image, and can also predict in-depth analysis information such as the comparison of the winning probabilities of the two sides.

[0145] For example, by retrieving information such as the recent weather conditions and travel suitability index of a city through a multimodal retrieval system, the multimodal understanding model interprets the information and provides answers to the user.

[0146] S400: Control the display 260 to display annotation feedback information on the user interface.

[0147] Users can view annotation feedback information through the user interface. This feedback information can include annotation details and feedback information, such as... Figure 10 As shown, annotation information is defined as content marked on the user interface, such as... Figure 11 As shown, feedback information is defined as the response made during a conversation with a user.

[0148] The labeled information is generated after the multimodal retrieval system associates information and / or the multimodal understanding model understands the information. For example, it includes in-depth analysis information such as the current score on the field and the model's predicted winning probability comparison between the two sides, which are displayed non-intrusively at the bottom of the screen.

[0149] Feedback information is the answer provided to the user based on the information fed back by the multimodal retrieval system and / or the information understood by the multimodal understanding model. For example, the answer may be that the scene shown in the current video is a scenic spot in a certain city, and the weather is suitable for travel recently.

[0150] In some embodiments, the video content analysis method using multimodal technology can use a multimodal understanding model to mark target regions on the target visual content based on the question text and the target visual content. Using this multimodal understanding model, a response text is generated based on the target visual content with the marked target regions and the question text. Based on the target response text and the target visual content, target tags for the target visual content are recalled; these target tags are used to trigger functions related to the target visual content. The response text is displayed to answer the question text, and the target tags are also displayed so that the user can quickly execute the corresponding function, enriching the ways to interact with the user, thereby meeting user needs and improving the user experience.

[0151] After receiving the annotation feedback, the user may believe that the annotation feedback is inaccurate or that further annotation is needed, such as... Figure 12 As shown, in some embodiments, a second request annotation instruction is obtained, which is used to perform supplementary annotation; in response to the second request annotation instruction, preference information is extracted; supplementary information is generated by re-retrieval through the multimodal retrieval system using the preference information and the current interface image; and supplementary annotation feedback information is generated by understanding the supplementary information, the second request annotation instruction, and the current interface image through a multimodal understanding model.

[0152] The second request annotation instruction is similar to the first request annotation instruction. It can be a supplementary instruction generated by the display device 200 according to a preset cycle, or it can be a supplementary instruction input by the user, used to perform supplementary annotation.

[0153] Preference information can be obtained through supplementary instructions input by the user or through historical interaction information. In some embodiments, in response to the second request annotation instruction, the display 260 is controlled to display a preference settings interface, which includes preference settings items; the preference settings items selected by the user based on the preference settings interface are obtained; and the preference information is generated based on the selected preference items.

[0154] If no annotation setting parameters are entered when setting up automatic annotation on display device 200, that is, the generated annotation feedback information may not be the annotation feedback information needed by the user, the preference setting interface will be displayed when the user enters the second request annotation command. The preference setting interface includes multiple preference setting items, which can be the same as the annotation setting parameters, and can also include the correction direction, supplementary direction, etc. of the annotation feedback information.

[0155] In some embodiments, historical interaction information is retrieved according to the content type, the historical interaction information including the associated information and / or the annotation feedback information generated in response to the first request annotation instruction; preference information is extracted from the historical interaction information.

[0156] When generating association information and annotation feedback information, the first request annotation instruction is input, which is generated through user input.

[0157] User input commands contain user preference information. For different content types, such as sports events and dramas, user needs differ. Therefore, by extracting user preference information for different content types, and generating supplementary annotation feedback information using different types of preference information, the accuracy of the supplementary annotation feedback information can be improved, making it more in line with user needs and enhancing the user experience.

[0158] It is understandable that for each generated association information, annotation feedback information, and supplementary annotation feedback information, preference information can be extracted from the user's input content, and a mapping relationship can be formed between the preference information and the content type to improve the annotation accuracy.

[0159] After supplementary annotation feedback information is generated, the user will continue to interact with the display device 200 through the supplementary annotation feedback information. However, since the supplementary annotation feedback information can be provided to the user in the form of annotation information and feedback information, when the supplementary annotation feedback information is displayed in the form of annotation information, the user cannot have a real-time interaction with the display device 200, or the user has other needs, resulting in a poor user experience. In some embodiments, a prompt interface is generated based on the supplementary annotation feedback information and the interface image; in response to the interactive command input by the user based on the prompt interface, the display 260 is controlled to display a first interactive interface on the user interface; an inquiry message is generated based on the inquiry command input by the user based on the first interactive interface; the inquiry message and the interface image are input to a multimodal retrieval system to retrieve response information through the multimodal retrieval system; and the display 260 is controlled to display the response message on the first interactive interface.

[0160] For example, after generating supplementary annotation feedback information, the display device 200 can display a prompt interface, prompting the user to activate the viewing companion, i.e., the first interactive interface, through a predetermined button. The viewing companion includes, but is not limited to, smart speakers, TV companion boxes, smart TV applications, etc. The user can ask further questions about the supplementary annotation feedback information through the viewing companion, or fulfill other needs through the viewing companion, such as the needs of the viewing environment.

[0161] Based on the aforementioned display device 200, some embodiments of this application also provide an interface content annotation method, including the following steps:

[0162] In response to the first request annotation instruction, the currently displayed interface image and user interface metadata are obtained, wherein the first request annotation instruction is an instruction generated by the display device according to a preset period, and / or an instruction input by the user;

[0163] The first request annotation instruction, the interface image, and the metadata are input into a multimodal retrieval system to retrieve related information. The multimodal retrieval system is used to perform information retrieval based on the multimodal information, which includes keywords extracted from the first request annotation instruction, the interface image, and the metadata.

[0164] The associated information, the first request annotation instruction, and the interface image are input into the multimodal understanding model to generate annotation feedback information through the multimodal understanding model. The multimodal understanding model is a model trained and generated based on a deep learning algorithm using the multimodal information.

[0165] Control the display to show annotation feedback information on the user interface.

[0166] As can be seen from the above technical solutions, some embodiments of this application provide a display device and a method for annotating interface content. The method, in response to a first annotation request instruction, obtains the currently displayed interface image and metadata of the user interface. It then inputs the first annotation request instruction, the interface image, and the metadata into a multimodal retrieval system to retrieve related information. Finally, it inputs the related information, the first annotation request instruction, and the interface image into a multimodal understanding model to generate annotation feedback information, and controls the display to show the annotation feedback information on the user interface. This method performs information retrieval on multimodal information through a multimodal retrieval system. By using the first annotation request instruction, the interface image, and metadata as multimodal information, the accuracy of information retrieval and understanding can be improved, thereby increasing the accuracy of the annotated content.

[0167] The same or similar parts among the various embodiments in this specification can be referred to mutually, and will not be repeated here.

[0168] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or certain parts of the embodiments of the present invention.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0170] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. A display device, characterized in that, include: The monitor is configured to display the user interface; The controller is configured as follows: In response to the first request annotation instruction, the currently displayed interface image and user interface metadata are obtained, wherein the first request annotation instruction is an instruction generated by the display device according to a preset period, and / or an instruction input by the user; The first request annotation instruction, the interface image, and the metadata are input into a multimodal retrieval system to retrieve related information. The multimodal retrieval system is used to perform information retrieval based on the multimodal information, which includes keywords extracted from the first request annotation instruction, the interface image, and the metadata. The associated information, the first request annotation instruction, and the interface image are input into the multimodal understanding model to generate annotation feedback information through the multimodal understanding model. The multimodal understanding model is a model trained and generated based on a deep learning algorithm using the multimodal information. Control the display to show annotation feedback information on the user interface.

2. The display device according to claim 1, characterized in that, The controller executes a response to the first request annotation instruction, specifically configured as follows: Obtain annotation setting parameters input by the user, wherein the annotation setting parameters are annotation preference setting parameters for the interface image, and the annotation setting parameters include annotation content for the interface image; Extract annotation preference information from the annotation setting parameters, wherein the annotation preference information includes at least one annotation type item; Based on the annotation preference information, a first request annotation instruction is generated according to a preset period.

3. The display device according to claim 1, characterized in that, The controller is configured to acquire the currently displayed interface image and the metadata of the user interface, specifically as follows: Obtain the interface type of the interface image, wherein the interface type is a video interface or an image interface; If the interface type is a video interface, obtain the content type of the current interface image, and obtain the image of the current video interface and the metadata of the video interface according to the content type; If the interface type is an image interface, obtain the image of the image interface and the metadata of the image interface.

4. The display device according to claim 3, characterized in that, The controller executes the input of the association information, the first request annotation instruction, and the interface image into the multimodal understanding model, so as to generate annotation feedback information through the multimodal understanding model, specifically configured as follows: Based on the content type, emphasis information is set, which is used to characterize the emphasis of the annotation feedback of the video frame of the content type; The associated information, the first request annotation instruction, and the interface image are input into the multimodal understanding model to generate annotation feedback information based on the bias information.

5. The display device according to claim 3, characterized in that, The controller executes the input of the association information, the first request annotation instruction, and the interface image into the multimodal understanding model, and after generating annotation feedback information through the multimodal understanding model, it is further configured to: Obtain a second request annotation instruction, which is used to perform supplementary annotation; In response to the second request annotation instruction, preference information is extracted; The preference information and the interface image are input into a multimodal retrieval system to retrieve supplementary information. The supplementary information, the second request annotation instruction, and the interface image are input into the multimodal understanding model to generate supplementary annotation feedback information through the multimodal understanding model. Control the display to show supplementary annotation feedback information on the user interface.

6. The display device according to claim 5, characterized in that, The controller is configured to extract preference information as follows: In response to the second request annotation command, the display is controlled to show a preference settings interface, which includes preference settings items; Obtain the preference settings selected by the user based on the preference settings interface; The preference information is generated based on the selected preference settings.

7. The display device according to claim 5, characterized in that, The controller is configured to extract preference information as follows: Based on the content type, retrieve historical interaction information, which includes the associated information and / or the annotation feedback information generated in response to the first request annotation instruction; Preference information is extracted from the historical interaction information.

8. The display device according to claim 6, characterized in that, After the controller executes the command to control the display to show supplementary annotation feedback information on the user interface, it is also configured to: Based on the supplementary annotation feedback information and the interface image, a prompt interface is generated; In response to an interactive command input by the user based on the prompt interface, the display is controlled to show a first interactive interface on the user interface; Generate inquiry information based on the inquiry command input by the user on the first interactive interface; The query information and the interface image are input into a multimodal retrieval system to retrieve response information through the multimodal retrieval system; Control the display to show the response information on the first interactive interface.

9. The display device according to claim 1, characterized in that, The controller executes the input of the first request annotation command, the interface image, and the metadata to the multimodal retrieval system, and after retrieving related information through the multimodal retrieval system, it is further configured to: Control the display to show associated information on the user interface; In response to a user-inputted focus movement command, the display is controlled to show a second interactive interface on the user interface; Based on the error correction command input by the user through the second interactive interface, error correction information is generated; The error correction information, the association information, the first request annotation instruction, and the interface image are input into the multimodal understanding model to generate annotation feedback information through the multimodal understanding model.

10. A method for annotating interface content, characterized in that, Applied to the display device according to any one of claims 1-9, the interface content annotation method includes: In response to the first request annotation instruction, the currently displayed interface image and user interface metadata are obtained, wherein the first request annotation instruction is an instruction generated by the display device according to a preset period, and / or an instruction input by the user; The first request annotation instruction, the interface image, and the metadata are input into a multimodal retrieval system to retrieve related information. The multimodal retrieval system is used to perform information retrieval based on the multimodal information, which includes keywords extracted from the first request annotation instruction, the interface image, and the metadata. The associated information, the first request annotation instruction, and the interface image are input into the multimodal understanding model to generate annotation feedback information through the multimodal understanding model. The multimodal understanding model is a model trained and generated based on a deep learning algorithm using the multimodal information. Control the display to show labeled feedback information on the user interface.

Citation Information

Patent Citations

  • Video classification model training method, video classification method and device

    CN116977701A

  • Auxiliary labeling method and device, object recognition method and device and electronic equipment

    CN117454149A