Electronic device and control method therefor
The electronic device uses AI models and servers to enhance screen description by identifying screen types and obtaining detailed information, addressing the limitation of conventional methods by providing specific details on screen subjects.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-02
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional methods for providing description information on a current screen are limited to general information and do not include specific details about the subject, such as character names or locations.
An electronic device identifies an artificial intelligence model corresponding to the current screen type, obtains prompts for detailed description information, and transmits these to a server to receive and provide specific information, including metadata and image captioning, using processors and communication interfaces.
Enhances user experience by providing detailed and specific information about the current screen, improving understanding and engagement.
Smart Images

Figure KR2025015785_09042026_PF_FP_ABST
Abstract
Description
Electronic device and control method thereof
[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more specifically, to an electronic device and a method for controlling the same for providing description information for a currently displayed screen.
[0002] An artificial intelligence system is a computer system that implements human-level intelligence, in which machines learn and make judgments on their own, and whose recognition rate improves with use.
[0003] Artificial intelligence technology consists of machine learning (deep learning) technology, which utilizes algorithms to self-classify and learn the characteristics of input data, and component technologies that utilize machine learning algorithms to mimic functions such as cognition and judgment of the human brain.
[0004] The elemental technologies may include, for example, linguistic understanding technology that recognizes human language / characters, visual understanding technology that recognizes objects like human vision, reasoning / prediction technology that judges information to logically reason and predict, knowledge representation technology that processes human experience information into knowledge data, and motion control technology that controls autonomous driving of vehicles and the movement of robots.
[0005] Meanwhile, with the recent development of various image recognition technologies, services providing descriptions of the current screen are being offered. In particular, conventionally, there were methods for providing description information of content provided by content providers, as well as methods for providing description information through image captioning.
[0006] However, conventional methods provide descriptions of the current screen that are limited to general information. In other words, there is a problem in that specific information about the subject of the current screen (e.g., character names, specific locations, etc.) is not included.
[0007] Meanwhile, the information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0008] According to one embodiment of the present disclosure, an electronic device comprises: a communication interface; a memory for storing instructions; and at least one processor. When the instructions are executed collectively or individually by the at least one processor, the electronic device identifies an artificial intelligence model corresponding to the current screen among a plurality of artificial intelligence models based on a type of current screen identified using information related to content, obtains a prompt for obtaining description information corresponding to the current screen using information related to content, and transmits the prompt to a server corresponding to the identified artificial intelligence model to provide a first description information corresponding to the obtained prompt.
[0009] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may acquire information related to a person included in the current screen, image captioning information for the current screen, and information about text included in the current screen using the current screen captured while the content is provided, acquire text information corresponding to speech output from the current screen through Automatic Speech Recognition (ASR), and acquire metadata related to the content.
[0010] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may obtain information related to a person included in the current screen, image captioning information for the current screen, information about text included in the current screen, text information corresponding to the voice, and second description information for the current screen based on the metadata.
[0011] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may obtain first type information for the current screen using information about the content type included in the metadata, obtain second type information for the current screen using content description information and a knowledge graph included in the metadata, obtain third type information for the current screen using the second description information and the knowledge graph, and obtain type information for the current screen based on the first to third type information.
[0012] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may acquire the third type information through the plurality of screens, and if the number of the plurality of screens that have acquired the third type information is greater than or equal to a threshold value, identify the type of the current screen based on the third type information, and if the number of the plurality of screens that have acquired the third type information is less than a threshold value, identify the type of the current screen based on the first type information and the second type information.
[0013] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may obtain the prompt using the captured screen, the voice output from the current screen, the metadata, and the second description information.
[0014] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may transmit the prompt and the second description information to a server corresponding to the identified artificial intelligence model to obtain the first description information from the server.
[0015] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may update the weights of the information for obtaining the second description based on the received first description information.
[0016] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may first provide the second description information obtained by the electronic device, and when the first description information is received, remove the second description and provide the first description information.
[0017] Meanwhile, a control method for an electronic device according to one embodiment of the present disclosure comprises: a step of identifying an artificial intelligence model corresponding to a current screen among a plurality of artificial intelligence models based on a type of current screen identified using information related to content; a step of obtaining a prompt for obtaining description information corresponding to the current screen using information related to content; and a step of transmitting the prompt to a server corresponding to the identified artificial intelligence model to provide a first description information corresponding to the obtained prompt.
[0018] The step of providing information related to the above content may include: a step of obtaining information related to a person included in the current screen, image captioning information for the current screen, and information regarding text included in the current screen using the current screen captured while the above content is provided; a step of obtaining text information corresponding to a voice output from the current screen through Automatic Speech Recognition (ASR); and a step of obtaining metadata related to the above content.
[0019] The above control method may further include the step of obtaining second description information for the current screen based on information related to a person included in the current screen, image captioning information for the current screen, information about text included in the current screen, text information corresponding to the voice, and metadata.
[0020] The above identifying step may include: a step of obtaining first type information for the current screen using information about the content type included in the metadata; a step of obtaining second type information for the current screen using content description information and a knowledge graph included in the metadata; a step of obtaining third type information for the current screen using the second description information and the knowledge graph; and a step of obtaining type information for the current screen based on the first to third type information.
[0021] The third type information is obtained through the plurality of screens, and the step of obtaining type information for the current screen is to identify the type for the current screen based on the third type information if the number of plurality of screens that obtained the third type information is greater than or equal to a threshold value, and to identify the type for the current screen based on the first type information and the second type information if the number of plurality of screens that obtained the third type information is less than a threshold value.
[0022] The step of obtaining the above prompt can be performed by using the captured screen, the voice output from the current screen, the metadata, and the second description information to obtain the prompt.
[0023] The step of obtaining the first description information can be performed by transmitting the prompt and the second description information to a server corresponding to the identified artificial intelligence model, thereby obtaining the first description information from the server.
[0024] The above control method may further include the step of updating the weights of the information for obtaining the second description based on the received first description information.
[0025] The above control method further includes the step of first providing the second description information obtained by the electronic device; wherein, when the first description information is received, the step of providing may remove the second description and provide the first description information.
[0026] FIG. 1 is a drawing illustrating a system for providing description information according to one embodiment of the present disclosure.
[0027] FIG. 2 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0028] FIG. 3 is a drawing illustrating a plurality of modules for providing description information according to one embodiment of the present disclosure.
[0029] FIG. 4 is a sequence diagram illustrating a method for an electronic device and a server to provide description information according to one embodiment of the present disclosure.
[0030] FIG. 5 is a drawing for explaining information related to content according to one embodiment of the present disclosure.
[0031] FIG. 6a is a drawing for explaining second description information according to one embodiment of the present disclosure,
[0032] FIG. 6b is a drawing for explaining first description information according to one embodiment of the present disclosure, and,
[0033] FIG. 7 is a flowchart illustrating a control method for an electronic device for providing description information according to one embodiment of the present disclosure.
[0034] The embodiments described herein are subject to various modifications and may have various forms; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals may be used for similar components.
[0035] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is omitted.
[0036] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concept of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to make the present disclosure more faithful and complete and to fully convey the technical concept of the present disclosure to those skilled in the art.
[0037] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of the rights. The singular expression includes the plural expression unless the context clearly indicates otherwise.
[0038] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.
[0039] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.
[0040] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.
[0041] Where it is stated that a component (e.g., Component 1) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., Component 2), it should be understood that the component may be directly connected to the other component or connected through the other component (e.g., Component 3).
[0042] On the other hand, when it is stated that a certain component (e.g., a first component) is "directly connected" or "directly coupled" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between the certain component and the other component.
[0043] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware.
[0044] Instead, in some situations, the expression “device configured to do something” may mean that the device is “capable of doing something” together with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing those operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) capable of performing those operations by executing one or more software programs stored in a memory device.
[0045] In the embodiments, a 'module' or 'part' performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented by at least one processor, except for the 'module' or 'part' that needs to be implemented in specific hardware.
[0046] Meanwhile, the various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0047] Meanwhile, according to one embodiment of the present disclosure, a "prompt" may mean an input for initiating interaction with an artificial intelligence model (e.g., a generative AI model). The prompt may be a text input or voice input comprising one or more texts and / or one or more sentences. In one embodiment, the prompt may include natural language text. The natural language text may include various information that the generative AI model can use to generate a response to a user inquiry or to control the electronic device (100), such as context, intent, task, and constraints. Meanwhile, the prompt may be referred to by being replaced with various expressions representing the same or similar concept. The prompt may be replaced with expressions such as, for example, "input," "user input," "input phrase," "user command," "directive," "starting sentence," "task query," "trigger sentence," "message," etc., but is not limited to the examples mentioned above.
[0048] Meanwhile, according to one embodiment of the present disclosure, the artificial intelligence model may be a Large Language Model (LM). Here, the LLM is a language model composed of an artificial neural network containing numerous parameters. The LLM may be trained using self-supervised learning or semi-self-supervised learning with a substantial amount of unlabeled corpus text. In this case, the LLM may not only possess the ability to generate answers to user inquiries but may also include reasoning capabilities and the ability to formulate and execute plans independently. Meanwhile, the LLM may be referred to by various terms such as a large language model, an AI chatbot model, etc. In particular, according to one embodiment of the present disclosure, the LLM may be a model trained to obtain description information corresponding to the current screen by inputting a prompt.
[0049] Meanwhile, according to one embodiment of the present disclosure, "description information" may be information describing the currently displayed screen. In particular, the description information may include information about content related to the current screen, information about objects included in the current screen (e.g., characters, etc.), information describing the current screen, web information related to the current screen, advertising information, etc.
[0050] Hereinafter, embodiments according to the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them.
[0051] FIG. 1 is a diagram illustrating a system for providing description information according to an embodiment of the present disclosure. As illustrated in FIG. 1, the system for providing description information may include an electronic device (100) and a plurality of servers (200-1, 200-2, 200-3,...). The electronic device (100) is a device for providing a user with description information corresponding to content and the current screen, and as illustrated in FIG. 1, it may be implemented as a TV, but this is merely an embodiment, and it may be implemented as various devices such as a set-top box, a desktop PC, a laptop PC, a projector, a refrigerator, etc. The plurality of servers (200-1, 200-2, 200-3,...) are servers for providing description information using an LLM, and each of the plurality of servers (200-1, 200-2, 200-3,...) may store an LLM corresponding to the type of the current screen.
[0052] The electronic device (100) can provide content. Here, the content may be video content such as broadcast content, movie content, sports content, etc.
[0053] The electronic device (100) can obtain information related to content. Here, the information related to content may include information related to a person obtained through the captured current screen, image captioning information, and information about text included in the current screen. Additionally, the information related to content may include text information corresponding to voice output from the current screen and information included in metadata.
[0054] The electronic device (100) can identify the type of the current screen (or the type of content) using information related to the content. In one embodiment, the electronic device (100) can identify the type corresponding to the current screen among sports type, movie type, drama type, news type, documentary type, education type, and humor type using information related to the content.
[0055] The electronic device (100) can identify an LLM corresponding to the current screen among a plurality of LLMs based on the type of the identified current screen. That is, each of the plurality of LLMs can correspond to a type of screen. For example, the first LLM may be an LLM corresponding to a sports type and may be a model trained to provide description information of a screen related to sports, the second LLM may be an LLM corresponding to a movie type and may be a model trained to provide description information of a screen related to movies, and the third LLM may be an LLM corresponding to a drama type and may be a model trained to provide description information of a screen related to drama. In addition, each of the plurality of LLMs may be stored in a plurality of servers (200-1, 200-2, 200-3...) shown in FIG. 1.
[0056] Additionally, the electronic device (100) can obtain a prompt to inquire about description information corresponding to the current screen using information related to the content. Here, the electronic device (100) can obtain first description information (hereinafter "second description information") in advance within the electronic device (100) using information related to the content. Then, the electronic device (100) can obtain a prompt based on information related to the content (e.g., captured screen, voice output from the current screen, metadata, etc.) and second description information.
[0057] The electronic device (100) can transmit the acquired prompt to a server corresponding to the identified LLM among a plurality of servers (200-1, 200-2, 200-3...). At this time, the server to which the prompt is transmitted may be a server that stores the LLM corresponding to the identified current screen. Meanwhile, although the above-described embodiment is described as including a plurality of servers (200-1, 200-2, 200-3,...), this is merely one embodiment and can be implemented with a single server. In the case of implementation with a single server, the electronic device (100) can transmit information about the LLM corresponding to the current screen along with the prompt so that the server can identify the LLM corresponding to the current screen among the plurality of LLMs.
[0058] The server can obtain final description information for the current screen (hereinafter referred to as "first description information") by entering a prompt into the stored LLM. At this time, the final description information is more detailed information than the initial description information and may include more specific information than the initial description information (e.g., detailed information about characters included in the current screen, detailed information about places displayed on the current screen, etc.).
[0059] The server can transmit the acquired first description information to the electronic device (100).
[0060] The electronic device (100) may provide the acquired first description information. Here, the electronic device (100) may provide the first description information on a portion of the current screen. In one or more embodiments, the electronic device (100) may first provide the second description information, and then, when the first description information is received, remove the second description information and provide the first description information.
[0061] According to the embodiment described above, the electronic device (100) can provide description information containing various detailed information rather than providing fragmentary description information, so the user experience of the user of the electronic device (100) can be improved.
[0062] Meanwhile, although the above-described embodiment was explained as storing multiple LLMs on an external server, this is merely one embodiment, and it is obvious that multiple LLMs can be stored inside the electronic device (100).
[0063]
[0064] FIG. 2 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure. As shown in FIG. 2, the electronic device (100) may further include a display (110), memory (120), communication interface (130), sensor (140), input / output interface (150), user interface (160), camera (170), microphone (180), and processor (190). However, this is merely one embodiment, and depending on the type of electronic device (100), some components may be removed or added. For example, if the electronic device (100) is implemented as a set-top box, the electronic device (100) may not include a display (110).
[0065] The display (110) may include various types of display panels such as an LCD (Liquid Crystal Display) panel, an OLED (Organic Light Emitting Diodes) panel, an AM-OLED (Active-Matrix Organic Light-Emitting Diode), an LcoS (Liquid Crystal on Silicon), a QLED (Quantum dot Light-Emitting Diode) and DLP (Digital Light Processing), a PDP (Plasma Display Panel) panel, an inorganic LED panel, and a micro LED panel, but is not limited thereto. Meanwhile, the display (110) may form a touchscreen together with a touch panel and may be made of a flexible panel.
[0066] In particular, the display (110) can display content received from various sources (e.g., a communication interface (130), an input / output interface (150), etc.). Additionally, the display (110) can display description information corresponding to the current screen along with the content.
[0067] The memory (120) can store instructions or data related to the components of the electronic device (100) and the operating system (OS) for controlling the overall operation of the components of the electronic device (100). In particular, the memory (120) may include various modules for providing description information corresponding to the current screen. In particular, when an event occurs to provide description information corresponding to the current screen, the electronic device (100) may load data into volatile memory for various modules to perform various operations for providing description information corresponding to the current screen stored in non-volatile memory, as shown in FIG. 3. Here, loading means the operation of bringing data stored in non-volatile memory into volatile memory and storing it so that the processor (190) can access it.
[0068] In one or more embodiments, the memory (120) may include a weight DB that stores information about the weights of the information used when generating the second description information.
[0069] In one or more embodiments, the memory (120) can store at least one LLM.
[0070] Meanwhile, memory (120) can be implemented as non-volatile memory (e.g., hard disk, SSD (Solid state drive), flash memory), volatile memory (memory within the processor (190)), etc.
[0071] The communication interface (130) includes at least one circuit and can perform communication with various types of external devices or servers. In particular, according to one embodiment of the present disclosure, the communication interface (130) may include a plurality of types of communication interfaces. For example, the communication interface (130) may include a Bluetooth communication interface, an IR communication interface, a Wi-Fi communication interface, etc. In addition, the communication interface (130) may include various communication interfaces in addition to the communication interfaces described above (e.g., a cellular communication module, a 3G (3rd generation) mobile communication module, a 4G (4th generation) mobile communication module, a 4th generation LTE (Long Term Evolution) communication module, a 5G (5th generation) mobile communication module, an NFC communication module, etc.).
[0072] In one or more embodiments, the communication interface (130) may transmit a prompt to an external server and receive first description information for the prompt. Additionally, the communication interface (130) may transmit second description information together with the prompt.
[0073] The sensor (140) can detect the state of the electronic device (100) (e.g., movement) or the state of the external environment (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. The sensor (140) may include, for example, a gesture sensor and an accelerometer.
[0074] The input / output interface (150) is configured to input or output at least one of audio and video signals. For example, the input / output interface (150) may be HDMI (High Definition Multimedia Interface), but this is merely an example of an embodiment, and it may be any one of MHL (Mobile High-Definition Link), USB (Universal Serial Bus), DP (Display Port), Thunderbolt, VGA (Video Graphics Array) port, RGB port, D-SUB (D-subminiature), or DVI (Digital Visual Interface). Depending on the implementation example, the input / output interface (140) may include a port for inputting and outputting only audio signals and a port for inputting and outputting only video signals as separate ports, or it may be implemented as a single port for inputting and outputting both audio and video signals.
[0075] In one or more embodiments, the input / output interface (150) can receive video content from an external device.
[0076] The user interface (160) may include a button, a lever, a switch, a touch interface, etc. In this case, the touch interface may be implemented by receiving input through the user's touch on the display (110) screen of the electronic device (100).
[0077] In particular, the user interface (160) can receive various user commands, such as user commands for obtaining description information.
[0078] The camera (170) can capture still images and video. A camera (170) according to various embodiments of the present disclosure may include one or more lenses, an image sensor, an image signal processor, and a flash. One or more lenses may include a telephoto lens, a wide-angle lens, and a super-wide-angle lens disposed on the surface of the electronic device (100), and may also include a three-dimensional depth lens. The camera (170) may be disposed on the surface (e.g., rear or front) of the electronic device (100), but is not limited to such configuration, and various embodiments according to the present disclosure may be implemented through a connection with a camera (170) that exists separately outside the electronic device (100).
[0079] A microphone (180) may refer to a device that detects sound and converts it into an electrical signal. For example, the microphone (180) can detect voice in real time, and by converting the detected voice into an electrical signal, the electronic device (100) can perform an action corresponding to the electrical signal. The microphone (180) may include a TTS module or an STT module.
[0080] The processor (190) can control the electronic device (100) according to at least one instruction stored in memory (120).
[0081] In particular, the processor (190) may include one or more processors. Specifically, one or more processors may include one or more of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), MIC (Many Integrated Core), DSP (Digital Signal Processor), NPU (Neural Processing Unit), hardware accelerator, or machine learning accelerator. One or more processors may control one or any combination of other components of an electronic device and may perform operations or data processing related to communication. One or more processors may execute one or more programs or instructions stored in memory. For example, one or more processors may perform a method according to one embodiment of the present disclosure by executing one or more instructions stored in memory.
[0082] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by a single processor or by a plurality of processors. That is, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an artificial intelligence dedicated processor). For example, according to one embodiment of the present disclosure, operations such as identifying corners within a handwriting image or correcting space within a handwriting image using a neural network model may be performed by a processor that performs parallel operations, such as a GPU or an NPU, and operations such as generating / editing a plan view image or post-processing operations may be performed by a general-purpose processor, such as a CPU.
[0083] One or more processors may be implemented as a single-core processor comprising one core, or as one or more multicore processors comprising multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory such as cache memory or on-chip memory, and a common cache shared by multiple cores may be included in the multicore processor. Additionally, each of the multiple cores included in the multicore processor (or some of the multiple cores) may independently read and execute program instructions for implementing a method according to one embodiment of the present disclosure, or all (or some) of the multiple cores may be linked together to read and execute program instructions for implementing a method according to one embodiment of the present disclosure.
[0084] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one of the plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in a multi-core processor, or the first operation and the second operation may be performed by a first core included in a multi-core processor and the third operation may be performed by a second core included in a multi-core processor.
[0085] In embodiments of the present disclosure, the processor (190) may mean a system-on-chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, GPU, APU, MIC, DSP, NPU, hardware accelerator or machine learning accelerator, etc., but the embodiments of the present disclosure are not limited thereto.
[0086] In particular, the processor (190) executes at least one instruction stored in memory (120) to obtain information related to the content while the content is being provided, identifies an LLM (Large Language Model) corresponding to the current screen among a plurality of LLMs based on the type of the current screen identified using the information related to the content, obtains a prompt to query description information corresponding to the current screen using the information related to the content, transmits the prompt to a server corresponding to the identified LLM to obtain a first description information corresponding to the prompt, and provides the first description information.
[0087] In one or more embodiments, the processor (190) can capture the current screen while the content is provided by executing at least one instruction stored in memory (120), and using the captured current screen, obtain information related to a person included in the current screen, image captioning information for the current screen, and information about text included in the current screen, obtain text information corresponding to the voice output from the current screen through Automatic Speech Recognition (ASR), and obtain metadata related to the content.
[0088] In one or more embodiments, the processor (190) can obtain second description information for the current screen based on information related to a person included in the current screen, image captioning information for the current screen, information about text included in the current screen, text information corresponding to voice, and metadata by executing at least one instruction stored in memory (120).
[0089] In one or more embodiments, the processor (190) can obtain first type information for the current screen by executing at least one instruction stored in memory (120), using information about the content type included in metadata, obtain second type information for the current screen using content description information and a knowledge graph included in metadata, obtain third type information for the current screen using second description information and a knowledge graph, and obtain type information for the current screen based on the first to third type information.
[0090] In one or more embodiments, the processor (190) can obtain third type information through a plurality of screens by executing at least one instruction stored in memory (120), and if the number of the plurality of screens that obtained the third type information is greater than or equal to a threshold value, identify the type of the current screen based on the third type information, and if the number of the plurality of screens that obtained the third type information is less than a threshold value, identify the type of the current screen based on the first type information and the second type information.
[0091] In one or more embodiments, the processor (190) can obtain a prompt using the captured screen, voice output from the current screen, metadata, and second description information by executing at least one instruction stored in memory (120).
[0092] In one or more embodiments, the processor (190) can obtain first description information from the server by executing at least one instruction stored in memory (120) and transmitting a prompt and second description information to a server corresponding to the identified LLM.
[0093] In one or more embodiments, the processor (190) can update the weights of the information for obtaining a second description based on the received first description information by executing at least one instruction stored in the memory (120).
[0094] In one or more embodiments, the processor (190) first provides second description information obtained by the electronic device (100) by executing at least one instruction stored in memory (120), and when first description information is received, the second description can be removed and first description information can be provided.
[0095] FIG. 3 is a diagram illustrating a plurality of modules for providing description information according to one embodiment of the present disclosure. As shown in FIG. 3, an electronic device (100) may include a content information acquisition module (310), a content type identification module (320), an LLM identification module (330), a description generation module (340), a prompt generation module (350), a description acquisition module (360), and a description provision module (370). Here, the electronic device (100) may further include a weight DB (380).
[0096] The content information acquisition module (310) can acquire information related to the content. Specifically, the content information acquisition module (310) can acquire metadata about the currently received content, capture the currently displayed screen, or acquire audio output from the current screen. Additionally, the content information acquisition module (310) can acquire additional information related to the content using the metadata, the captured video, and the audio output from the current screen.
[0097] In one or more embodiments, the content information acquisition module (310) may continuously or periodically capture and store a plurality of screens. The content information acquisition module (310) may acquire information related to content regarding the stored plurality of screens.
[0098] Specifically, the content information acquisition module (310) may include a metadata acquisition module (311), a person recognition module (312), an image captioning module (313), a text detection module (314), and a voice recognition module (315) to acquire information related to various content, as shown in FIG. 3.
[0099] The metadata acquisition module (311) can acquire metadata provided along with the content through a content provider or a service provider. At this time, the metadata may include the title of the content, actors, genre, description related to the content, and other information.
[0100] The person recognition module (312) can recognize a person from the captured current screen. In particular, the person recognition module (312) can obtain information about the person using various machine learning models. Specifically, the person recognition module (312) can extract an area containing a person within the current screen using an object detection model and crop the extracted area. Then, the person recognition module (312) can obtain information about the person by inputting the extracted area into a person recognition engine. Here, the person recognition engine may be stored within the electronic device (100), but this is merely one embodiment, and it is obvious that it may be stored on an external server. At this time, the information about the person may include the person's gender, height, name, and information about the person's appearance.
[0101] The image captioning module (313) can obtain text or phrases describing the current screen captured using image captioning. Image captioning is a technique that describes the content of an image in text. The electronic device (100) can analyze an image through image captioning and express the meaning or context of the image in natural language. Specifically, the electronic device (100) can classify an image through the image captioning module (313), detect objects included within the image, and obtain information about the detected objects in natural language. Therefore, the electronic device (100) can obtain image captioning information as text describing the current screen through the image captioning module (313).
[0102] The text detection module (314) can detect text contained within the captured current screen and obtain information about the text. Here, the text detection module (314) can extract subtitle information configured in the form of an image within the current screen in the form of text through OCR (Optical Character Recognition).
[0103] The voice recognition module (315) can obtain text corresponding to the voice currently displayed on the screen through Automatic Speech Recognition (ASR) technology. Specifically, the voice recognition module (315) can capture voice data currently displayed on the screen and obtain text corresponding to the voice displayed through Automatic Speech Recognition (ASR) technology using the captured voice data.
[0104] The content type identification module (320) can identify a content type based on information related to the content obtained from the content information acquisition module (310). In particular, the content type identification module (320) can identify a type corresponding to the current screen among a plurality of content types.
[0105] Specifically, the content type identification module (320) can identify a type corresponding to the current screen based on information related to a person included in the current screen obtained from the content information acquisition module (310), image captioning information for the current screen, information about text included in the current screen, text information corresponding to voice, and metadata.
[0106] In one or more embodiments, the content type identification module (320) can obtain first type information for the current screen by using information about the content type included in the metadata. For example, if the information about the content included in the metadata includes "movie," the content type identification module (320) can identify that the type corresponding to the current screen is a movie type.
[0107] In one or more embodiments, the content type identification module (320) can obtain second type information for the current screen by using content description information included in metadata and a knowledge graph. Here, a knowledge graph is a data structure that visually represents the relationships between data, and can primarily represent objects (concepts, things, people, etc.) and the relationships between them as nodes (objects) and edges (relationships). For example, if the content description information included in the metadata includes "The plot of this movie is ~~~, and XXX and YYY are the lead actors," the content type identification module (320) can identify that the type corresponding to the current screen is a movie type by using the knowledge graph.
[0108] In one or more embodiments, the content type identification module (320) can obtain third type information for the current screen using the second description information and knowledge graph described below. For example, if the generated second description information is "a pitcher is preparing to throw a ball on the mound," the content type identification module (320) can identify that the type corresponding to the current screen is a sports type using the knowledge graph.
[0109] In particular, the content type identification module (320) can obtain third type information through a plurality of captured screens. And, if the number of multiple screens from which third type information has been obtained is greater than or equal to a threshold value, the content type identification module (320) can identify the type of the current screen based on the third type information. If the number of multiple screens from which third type information has been obtained is less than the threshold value, the content type identification module (320) can identify the type of the current screen based on the first type information and the second type information.
[0110] In addition, the content type identification module (320) can identify the type corresponding to the current screen based on various texts (e.g., text corresponding to subtitles or text corresponding to voice). For example, if the text included in the current screen includes "baseball score," the content type identification module (320) can identify that the type corresponding to the current screen is a sports type.
[0111] The LLM identification module (330) can identify one of a plurality of LLMs based on the type corresponding to the current screen identified by the content type identification module (320). Specifically, the electronic device (100) can store content types that match the plurality of LLMs. That is, each of the plurality of LLMs may be an LLM learned according to the type of content. For example, the first LLM may be an LLM learned based on information about movie content, and the second LLM may be an LLM learned based on information about sports content. That is, the LLM identification module (330) can identify the LLM corresponding to the current screen among the plurality of LLMs and provide more accurate and professional description information about the current screen.
[0112] The description generation module (340) can generate a second description based on information related to the content. Here, the second description may be a description generated by the electronic device (100) and may be distinguished from the first description obtained by the LLM.
[0113] In particular, the description generation module (340) can generate a second description based on image captioning information. Specifically, the description generation module (340) can obtain second description information by adding information about a person appearing on the current screen, text corresponding to subtitles included on the current screen, text corresponding to voice output on the current screen, and content information included in metadata to the image captioning information obtained by the image captioning module (313).
[0114] In one or more embodiments, the description generation module (340) may generate a second description based on weights stored in the weight DB (380). In this case, the weights may be weights for the information used when generating the second description. In particular, the weights stored in the weight DB (380) may initially have the same value. For example, if the information used when generating the second description consists of first information about a person appearing on the current screen, second information including text corresponding to subtitles included on the current screen, third information including text corresponding to voice output on the current screen, and fourth information included in metadata, the weights of the first to fourth information may initially each be 0.25. However, the weights of each piece of information may be updated later by the first description.
[0115] The prompt generation module (350) can generate a prompt for generating a description. Here, the prompt generation module (350) can generate a prompt using a second description along with information related to the content, such as a captured screen, voice output from the current screen, and metadata (e.g., title information, content description information, actor information).
[0116] In one embodiment, the prompt generation module (350) can generate a prompt for generating a description using a previously stored prompt template. In another embodiment, the prompt generation module (350) can generate a prompt by inputting the captured screen, voice output from the current screen, and metadata (e.g., title information, content description information, actor information) among the information related to the content into a second description, along with the information, into a trained neural network model.
[0117] The description acquisition module (360) can transmit the prompt obtained through the prompt generation module (350) to a server corresponding to the LLM identified by the LLM identification module (330). Here, the server corresponding to the identified LLM can obtain first description information for the current screen by inputting the prompt into the LLM. Here, the obtained first description information may include detailed information about the current screen compared to the second description information. When the server corresponding to the identified LLM obtains the first description information for the current screen, the description acquisition module (360) can receive the first description information for the current screen from the server.
[0118] Meanwhile, although the above-described embodiment explains that the first description information for the current screen is obtained using an LLM stored on an external server, this is merely one embodiment, and it is obvious that the first description information for the current screen can be obtained using an LLM stored inside the electronic device (100).
[0119] Additionally, the description acquisition module (360) can update the weights stored in the weight DB (380) based on the first description information for the current screen. In one or more embodiments, the description acquisition module (360) can update the weights corresponding to each of the first information about a person appearing on the current screen, the second information including text corresponding to a subtitle included on the current screen, the third information including text corresponding to a voice output on the current screen, and the fourth information included in the metadata, based on the first description information for the current screen. For example, the description acquisition module (360) can update the weights to increase the weight of the image captioning information when a lot of image captioning information is used when generating the first description information for the current screen.
[0120] The description providing module (370) may provide the first description information obtained by the description acquisition module (360). In one or more embodiments, the description providing module (370) may display the first description information on the display (110) together with the content currently being played. In one or more embodiments, the description providing module (370) may output the first description information through a speaker while the content currently being displayed.
[0121] FIG. 4 is a sequence diagram illustrating a method for an electronic device and a server to provide description information according to one embodiment of the present disclosure.
[0122] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0123] According to one or more embodiments, S410 to S490 may be understood to be performed in a processor (e.g., processor (190) of FIG. 2) of an electronic device (e.g., electronic device (100) of FIG. 1) or a server (e.g., external server of FIG. 1).
[0124] The electronic device (100) can obtain information about the content (S410). Here, information about the content can be obtained through a captured screen, voice capture, and metadata. Specifically, as shown in FIG. 5, the electronic device (100) can obtain information (511) about a person included in the screen based on vision recognition through screen capture (510). In addition, the electronic device (100) can obtain information about the current screen by performing image captioning (512) through screen capture (510). In addition, the electronic device (100) can recognize subtitles (513) using OCR through screen capture (510). In addition, the electronic device (100) can perform ASR-based voice recognition (521) through voice capture. In addition, the electronic device (100) can obtain information about the content (531), such as title, background, and content description information, through metadata (530).
[0125] The electronic device (100) can obtain second description information (S420). Specifically, the electronic device (100) can obtain second description information based on information related to content. More specifically, the electronic device (100) can obtain second description information by adding information about a person appearing on the current screen, text corresponding to subtitles included on the current screen, text corresponding to voice output on the current screen, and content information included in metadata to the image captioning information obtained by the image captioning module (313).
[0126] The electronic device (100) can identify an LLM corresponding to the current screen (S430). Specifically, the electronic device (100) can identify the type of the current screen based on information related to the content. And, the electronic device (100) can identify an LLM corresponding to the type of the current screen among a plurality of LLMs.
[0127] The electronic device (100) can obtain a prompt (S440). Specifically, the electronic device (100) can obtain a prompt using a captured screen, voice output from the current screen, metadata, and second description information.
[0128] The electronic device (100) can transmit the second description information and the prompt to the server (200) (S450). Here, the server (100) may be a server that stores an LLM corresponding to the current screen.
[0129] The server (200) can obtain first description information (S460). Specifically, the server (200) can obtain first description information using the received second description information and a prompt. In particular, the server (200) can obtain first description information for the current screen by inputting the obtained prompt into the LLM. Additionally, the server (200) can modify the obtained first description information based on the second description information.
[0130] The server (200) can transmit the acquired first description information to the electronic device (100) (S470).
[0131] The electronic device (100) may provide first description information (480) (S470). In one or more embodiments, the electronic device (100) may provide second description information (620) for the current screen along with content (610), as shown in FIG. 6a. Here, when the electronic device (100) obtains the second description information (620) at step S420, it may first provide the obtained second description information (620). Then, when the first description is received from the server (200), the electronic device (100) may remove the second description information (620) and provide first description information (630) for the current screen along with content (610), as shown in FIG. 6b. As illustrated in FIGS. 6a and 6b, the first description information (630) may contain more detailed information (e.g., specific information about a person included in the screen and specific information about the current screen) compared to the second description information (620). Meanwhile, when the first description information is received while the screen of FIG. 6a is displayed, the electronic device (100) may provide a UI that asks the user whether to provide the description information containing detailed information, and when user input is received through the UI, it may switch to the screen of FIG. 6b.
[0132] Additionally, the electronic device (100) can provide the second description information to an application or service user who requires the description.
[0133] The electronic device (100) can update the weights of the information for generating the second description information (S490). Specifically, the electronic device (100) can update the weights corresponding to each of the first information about a person appearing on the current screen, the second information including text corresponding to a subtitle included on the current screen, the third information including text corresponding to a voice output on the current screen, and the fourth information included in the metadata, based on the first description information about the current screen.
[0134] FIG. 7 is a flowchart illustrating a control method for an electronic device for providing description information according to one embodiment of the present disclosure.
[0135] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0136] According to one or more embodiments, S710 to S760 may be understood to be performed in a processor (e.g., processor (190) of FIG. 2) of an electronic device (e.g., electronic device (100) of FIG. 1).
[0137] First, the electronic device (100) provides content (S710).
[0138] The electronic device (100) obtains information related to the content (S720). In one or more embodiments, the electronic device (100) may capture the current screen while the content is being provided. Then, the electronic device (100) may use the captured current screen to obtain information related to a person included in the current screen, image captioning information for the current screen, and information about text included in the current screen. Additionally, the electronic device (100) may obtain text information corresponding to the voice output from the current screen through Automatic Speech Recognition (ASR). Additionally, the electronic device (100) may obtain metadata related to the content.
[0139] The electronic device (100) identifies an LLM corresponding to the current screen among a plurality of LLMs based on the type of the current screen identified using information related to the content (S730). In one or more embodiments, the electronic device (100) may obtain second description information for the current screen based on information related to a person included in the current screen, image captioning information for the current screen, information about text included in the current screen, text information corresponding to voice, and metadata.
[0140] In one or more embodiments, the electronic device (100) can obtain first type information for the current screen by using information about the content type included in the metadata. The electronic device (100) can obtain second type information for the current screen by using content description information and a knowledge graph included in the metadata. The electronic device (100) can obtain third type information for the current screen by using second description information and a knowledge graph. Furthermore, the electronic device (100) can obtain type information for the current screen based on the first to third type information. In particular, the electronic device (100) can obtain third type information through a plurality of screens, and if the number of a plurality of screens that have obtained third type information is greater than or equal to a threshold value, the type of the current screen can be identified based on the third type information. If the number of a plurality of screens that have obtained third type information is less than the threshold value, the electronic device (100) can identify the type of the current screen based on the first type information and the second type information.
[0141] The electronic device (100) obtains a prompt to inquire about description information corresponding to the current screen using information related to the content (S740). In one or more embodiments, the electronic device (100) may obtain a prompt using a captured screen, voice output from the current screen, metadata, and second description information.
[0142] The electronic device (100) transmits a prompt to a server corresponding to the identified LLM and provides first description information corresponding to the prompt (S750). In one or more embodiments, the electronic device (100) can obtain first description information from the server by transmitting a prompt and second description information to a server corresponding to the identified LLM.
[0143] In one or more embodiments, the electronic device (100) can update the weights of the information for obtaining a second description based on the received first description information.
[0144] The electronic device (100) provides first description information (S760). In one or more embodiments, the electronic device (100) may first provide second description information obtained by the electronic device (100). When the first description information is received, the electronic device (100) may remove the second description and provide the first description information.
[0145]
[0146] Meanwhile, the method according to various embodiments of the present disclosure may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0147] A method according to various embodiments of the present disclosure may be implemented as software comprising instructions stored on a machine-readable storage medium (e.g., a computer). The machine may include an electronic device (e.g., a TV) according to the disclosed embodiments, which is a device capable of calling instructions stored from the storage medium and operating according to the called instructions.
[0148] Meanwhile, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory storage medium' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.
[0149] When the above instruction is executed by a processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or an interpreter.
[0150] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
Claims
1. In an electronic device, Communication interface; Memory for storing instructions; and It includes at least one processor, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Identifying an artificial intelligence model corresponding to the current screen among a plurality of artificial intelligence models based on the type of the current screen identified using information related to the content, and Using information related to the above content, a prompt is obtained for obtaining description information corresponding to the current screen, and An electronic device that transmits the prompt to a server corresponding to the identified artificial intelligence model and provides a first description information corresponding to the obtained prompt.
2. In Paragraph 1, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Using the current screen captured while the above content is provided, information related to a person included in the current screen, image captioning information for the current screen, and information about text included in the current screen are obtained, and Text information corresponding to the voice output from the current screen is obtained through ASR (Automatic Speech Recognition), and An electronic device for acquiring metadata related to the above content.
3. In Paragraph 2, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device for obtaining second description information for the current screen based on information related to a person included in the current screen, image captioning information for the current screen, information about text included in the current screen, text information corresponding to the voice, and metadata.
4. In Paragraph 3, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Using information about the content type included in the above metadata, first type information for the current screen is obtained, and Using the content description information and knowledge graph included in the above metadata, second type information for the current screen is obtained, and Using the second description information and the knowledge graph, third type information regarding the current screen is obtained, An electronic device that obtains type information for the current screen based on the above first to third type information.
5. In Paragraph 3, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Acquiring the third type information through the aforementioned plurality of screens, and If the number of multiple screens that have acquired the above third type information is greater than or equal to a threshold value, the type of the current screen is identified based on the above third type information, and An electronic device that identifies the type of the current screen based on the first type information and the second type information when the number of multiple screens that have acquired the third type information is less than a threshold value.
6. In Paragraph 3, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that obtains the prompt using the above-described captured screen, the voice output from the above-described current screen, the above-described metadata, and the above-described second description information.
7. In Paragraph 6, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that transmits the prompt and the second description information to a server corresponding to the identified artificial intelligence model to obtain the first description information from the server.
8. In Paragraph 7, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that updates the weights of information for obtaining the second description based on the received first description information.
9. In Paragraph 3, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, First, the second description information obtained by the electronic device is provided, An electronic device that, upon receiving the first description information, removes the second description and provides the first description information.
10. In a method for controlling an electronic device, A step of identifying an artificial intelligence model corresponding to the current screen among a plurality of artificial intelligence models based on the type of the current screen identified using information related to the content; A step of obtaining a prompt for obtaining description information corresponding to the current screen using information related to the above content; A control method comprising the step of transmitting the prompt to a server corresponding to the identified artificial intelligence model and providing a first description information corresponding to the obtained prompt.
11. In Paragraph 10, The above control method is, A step of obtaining information related to a person included in the current screen, image captioning information for the current screen, and information about text included in the current screen using the current screen captured while the above content is provided; A step of obtaining text information corresponding to the voice output from the current screen through ASR (Automatic Speech Recognition); and A control method comprising the step of obtaining metadata related to the above content.
12. In Paragraph 11, The above control method is, A control method further comprising the step of obtaining second description information for the current screen based on information related to a person included in the current screen, image captioning information for the current screen, information about text included in the current screen, text information corresponding to the voice, and metadata.
13. In Paragraph 12, The above identification step is, A step of obtaining first type information for the current screen using information about the content type included in the metadata; A step of obtaining second type information for the current screen using content description information and a knowledge graph included in the above metadata; A step of obtaining third type information for the current screen using the second description information and the knowledge graph; and A control method comprising the step of obtaining type information for the current screen based on the first to third type information.
14. In Paragraph 13, The third type information is obtained through the aforementioned plurality of screens, and The step of obtaining type information for the current screen above is, If the number of multiple screens that have acquired the above third type information is greater than or equal to a threshold value, the type of the current screen is identified based on the above third type information, and A control method for identifying the type of the current screen based on the first type information and the second type information when the number of multiple screens that have acquired the third type information is less than a threshold value.
15. In Paragraph 12, The step of obtaining the above prompt is, A control method for obtaining the prompt using the above-mentioned captured screen, the voice output from the above-mentioned current screen, the above-mentioned metadata, and the above-mentioned second description information.
Citation Information
Patent Citations
Asymmetric epitaxy regions for landing contact plug
KR1020220021386A
Roof airbag for vehicle
KR1020230008460A
Mop for vehicle, mop stick for vehicle and manufacturing method of mop for vehicle
KR1020230070573A
Method and apparatus for producing descriptive video contents
KR102541008B1
Automatic generation of descriptive video service tracks
US20190069045A1