Electronic device and control method therefor
An electronic device processes audio and speech inputs to identify content context and generate responses, addressing the challenge of real-time user interaction with display devices by providing accurate and context-aware responses.
Patent Information
- Application Number
- PCT/KR2025/004974
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-03
- Filing Date
- 2025-04-11
- Publication Date
- 2025-12-11
AI Technical Summary
Conventional technologies struggle to accurately respond to user requests related to content being output on display devices due to the difficulty in grasping the context of the content in real time.
An electronic device equipped with a communication unit, memory, and processor that receives audio information from a display device or external device, acquires content information based on audio and speech input, identifies the context of the content, and generates a response signal to user requests based on this context.
Enables accurate and timely responses to user queries about content by understanding the context, including content information and user emotions, facilitating enhanced interaction with display devices.
Smart Images

Figure KR2025004974_11122025_PF_FP_ABST
Abstract
Description
Electronic device and method of controlling the same
[0001] The present disclosure relates to an electronic device and a control method thereof, and more particularly, to an electronic device and a control method thereof that respond to a user request related to content output from an external display device.
[0002] A display device is an electronic device that outputs images and can be implemented in various electronic devices, such as computers, TVs, smartphones, and tablets. Display devices can perform various actions based on various user interactions.
[0003] For example, while content is being output on a display device, the display device may obtain the content by using Internet searching or Automatic Content Recognition (ACR) technology, which is a technology that automatically recognizes content, and provide the content to the user.
[0004] There is also technology that understands and processes user requests through voice assistants such as AI speakers.
[0005] However, according to conventional technology, it was difficult for the user to grasp the context of the content he was currently watching in real time, and thus there was a problem in not being able to provide an accurate response to the user's question.
[0006] Accordingly, the need for technology that can effectively respond to users' requests regarding content has arisen.
[0007] According to at least one embodiment of the present disclosure, an electronic device includes a communication unit, a memory, and at least one processor, wherein the processor receives audio information input to a display device or an external device within a space in which the display device is arranged through the communication unit, acquires content information about content output from the display device based on the audio information and speech information uttered by a user located within the space in relation to the content, and stores the content information in the memory, and when a user request input to the display device or the external device is received through the communication unit, identifies a context of the content based on the content information and the speech information, generates a response signal corresponding to the user request based on the context, and transmits the response signal to the display device through the communication unit.
[0008] According to at least one embodiment of the present disclosure, a method for controlling an electronic device includes the steps of: receiving audio information input to a display device or an external device within a space in which the display device is arranged; acquiring and storing content information about content output from the display device based on the audio information and speech information uttered by a user located within the space in relation to the content; identifying a context of the content based on the content information and the speech information when a user request input to the display device or the external device is received; generating a response signal corresponding to the user request based on the context; and transmitting the response signal to the display device.
[0009] According to at least one embodiment of the present disclosure, a non-transitory computer-readable recording medium ((medium)) storing a program for executing a control method for controlling an electronic device includes the steps of: receiving audio information input to a display device or an external device within a space in which the display device is arranged; acquiring and storing content information about content output from the display device and speech information uttered by a user located within the space in relation to the content based on the audio information; identifying a context of the content based on the content information and the speech information when a user request input to the display device or the external device is received; generating a response signal corresponding to the user request based on the context; and transmitting the response signal to the display device.
[0010] FIG. 1 is a drawing for explaining the operation of an electronic device according to at least one embodiment of the present disclosure.
[0011] FIGS. 2 and 3 are block diagrams showing the configuration of an electronic device according to at least one embodiment of the present disclosure.
[0012] FIGS. 4 to 6 are drawings for explaining a method of responding to a user request in an electronic device according to various embodiments of the present disclosure.
[0013] FIG. 7 is a diagram illustrating a method for updating emotional information according to at least one embodiment of the present disclosure.
[0014] FIG. 8 is a diagram for explaining a process of responding to a user's request in an electronic device according to at least one embodiment of the present disclosure.
[0015] FIG. 9 is a diagram illustrating a process of acquiring content information and user speech information and learning content context in an electronic device according to at least one embodiment of the present disclosure.
[0016] FIG. 10 is a flowchart illustrating a method for controlling an electronic device according to at least one embodiment of the present disclosure.
[0017] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should not be defined simply as names, but rather based on the meanings of the terms and the overall content of this disclosure.
[0018] In this specification, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), and do not exclude the presence of additional features.
[0019] The expression "at least one of A and / or B" should be understood to mean either "A" or "B" or "A and B".
[0020] As used herein, the expressions “first,” “second,” “first,” or “second,” etc., may describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0021] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).
[0022] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "consist of" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0023] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented as at least one processor, excluding any "modules" or "parts" that need to be implemented as specific hardware.
[0024] In this specification, the term user may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence electronic device).
[0025] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.
[0026] FIG. 1 is a diagram illustrating the operation of an electronic device according to at least one embodiment of the present disclosure. FIG. 1 illustrates an environment in which an electronic device (100), a display device (200), and an external device (300) are used together. Each device (100, 200, 300) may be directly connected to each other through communication or may be connected through various networks. FIG. 1 illustrates a case in which the display device (200) and the external device (300) are used together in the same space. For example, the display device (200) and the external device (300) may be used in various spaces such as a general home, an office, a lobby, or a factory. The electronic device (100) may be separately installed outside the space, or may be installed together within the space. The electronic device (100) may receive or transmit various signals or data from the external device (300). For example, the electronic device (100) may receive audio information input to the external device (300). The electronic device (100) can obtain content information and speech information based on the received audio information and generate a response signal to the user's request information. Here, the audio information may include an audio signal output from the display device (200) and speech information spoken by the user (400). The audio signal output from the display device (200) may include various audio signals such as sound effects, background music, and dialogue of characters of the content output from the display device (200). The user's speech information may include a voice signal actually spoken by the user (400) within the corresponding space.
[0027] Details about audio information, content information, and speech information are described in detail in FIG. 2 and below.
[0028] The external device (300) may be an electronic device that is placed in a space where the display device (200) is placed and can sense sounds generated within the space. The external device (300) can sense sounds generated from the display device (200) and the user (400) through a microphone and transmit the sensed sounds to the electronic device (100). Here, the external device (300) may be implemented as a remote control equipped with a microphone, but is not limited thereto, and may be implemented as a wireless speaker, a mobile phone of the user (400), a tablet PC, various wearable devices worn by the user, other IoT devices, etc. In FIG. 1, the case where the external device (300) and the display device (200) are implemented separately is illustrated, but if the display device (200) is equipped with its own microphone or is connected to a microphone device, the external device (300) may be omitted, and various voice information may be acquired from the display device (200) and transmitted to the electronic device (100).
[0029] The display device (200) may be an electronic device that displays an image. The display device (200) may be any electronic device that outputs visual and audio information, such as a computer, TV, smartphone, or tablet. Although the display device (200) is described in the present disclosure, it is not necessarily limited to a device that has its own display, and may also include a set-top box or other content output device that is connected to an external display, such as a monitor or TV, and outputs content. In addition, a beam projector that projects an image onto an external screen may be included. Alternatively, the display device (200) may be implemented as an audio processing device (e.g., a radio, a speaker, etc.) that outputs only an audio signal without necessarily accompanying an image.
[0030] The display device (200) may be described by various terms, such as a content output device, a content processing device, or a multimedia player. In the various embodiments below, the case where it is implemented as a TV is described as the basis, and its name is described as a display device (200).
[0031] The content output from the display device (200) may be diverse, including a TV broadcast program, a game execution screen of a game console, a content playback screen of a playback device that plays recording media such as a Blu-ray disc, a still image mixed with background music, multimedia content provided through various platforms such as OTT (Over The Top) and VOD (Video on Demand), and multimedia content provided on various websites.
[0032] Content may include video data, audio data, and various additional data. Additional data may include various information, such as the name of the source providing the content, the content title, episode information, dialogue, and text information contained within the content. In this disclosure, such information is referred to as content information. Content information is not limited to the examples described above, and may further include various information related to the content.
[0033] A user (400) may be a person who views content played on a display device (200) within a space where the display device (200) is placed. At this time, the number of users (400) is not limited to one and may be multiple. The user (400) may utter text related to the content, questions related to the content, questions about searches related to the content, questions about content recommendations, etc. At this time, the user's (400) utterance may be sensed by a sensor (i.e., a microphone) included in an external device (300). The external device (300) may transmit the sensed voice of the user (400) to the electronic device (100).
[0034] The electronic device (100) can communicate with an external device (300) and a display device (200). At this time, the electronic device (100) may be driven in the form of an on device from the display device (200) or may be driven in the form of a separate external device such as an AI speaker, and is not limited to the above-described method.
[0035] The electronic device (100) identifies the context of content currently being output on the display device (200) based on audio information received from the external device (300). When the user makes a request in this state, the electronic device (100) generates a response signal to the user's request based on the identified context. The electronic device (100) can transmit the generated response signal to the external device (300) or the display device (200). The specific user request and response process will be described in detail again in the following section.
[0036] FIG. 2 is a block diagram illustrating an electronic device (100) according to one embodiment of the present disclosure.
[0037] Referring to FIG. 2, the electronic device (100) may include a memory (110), at least one processor (120), and a communication unit (130). The electronic device (100) of FIG. 2 may be implemented as a server device. In this case, the electronic device (100) communicates with various display devices (200) and external devices (300) used in a plurality of different spaces, analyzes the context of content based on audio information input from each external device (300), and provides a response based on the analyzed context when a user requests.
[0038] The memory (110) is a configuration for storing various programs, commands, data, etc. used in the operation of the electronic device (100). The memory (110) may be implemented as at least one of various memories such as DRAM (dynamic RAM), SRAM (static RAM), SDRAM (synchronous dynamic RAM), OTPROM (one time programmable ROM), PROM (programmable ROM), EPROM (erasable and programmable ROM), EEPROM (electrically erasable and programmable ROM), mask ROM, flash ROM, flash memory, hard drive, or solid state drive (SSD).
[0039] At least one processor (120) is a configuration for controlling the overall operation of the electronic device (100). The processor (120) may be implemented as a digital signal processor (DSP) for processing digital signals, a microprocessor, but is not limited thereto, and may include one or more of a central processing unit (CPU), a microcontroller unit (MCU), a microprocessor (MPU), a controller, an application processor (AP), a communication processor (CP), an ARM processor, and an artificial intelligence (AI) processor, or may be defined by the relevant terms. In addition, the processor (120) may be implemented as a system on chip (SoC) having a built-in processing algorithm, a large scale integration (LSI), or may be implemented in the form of a field programmable gate array (FPGA). The processor (120) may perform various functions by executing computer executable instructions stored in the memory (110).
[0040] The communication unit (130) can transmit and receive various signals and data to and from external devices through various wired and wireless communication methods such as Bluetooth, AP-based Wi-Fi (Wireless LAN network), Zigbee, wired / wireless LAN (Local Area Network), WAN (Wide Area Network), Ethernet, IEEE 1394, HDMI (High-Definition Multimedia Interface), USB (Universal Serial Bus), MHL (Mobile High-Definition Link), AES / EBU (Audio Engineering Society / European Broadcasting Union), optical, and coaxial.
[0041] According to one embodiment of the present disclosure, the processor (120) can receive audio information input to the display device (200) or an external device (300) within a space where the display device (200) is placed through the communication unit (130).
[0042] The processor (120) can receive audio information through the communication unit (130). According to one example, the display device (200) can record an audio signal output from the display device (200) or a user's speech through a microphone mounted on the display device (200). The display device (200) can transmit the recorded audio information to the electronic device (100).
[0043] According to one example, the display device (200) can transmit audio information of content stored in memory to the electronic device (100).
[0044] According to one example, the external device (300) can record an audio signal of content output from the display device (200) or a user's speech through a microphone and transmit it to the electronic device (100).
[0045] Here, inputting audio information to the display device (200) or external device (300) may mean recording an audio signal through a microphone.
[0046] Audio information may be otherwise called an audio signal, audio data, sound signal, or voice signal, but is referred to as audio information in this specification.
[0047] The external device (300) may be an electronic device that obtains audio information within a space where the display device (200) is placed and transmits the audio information to the electronic device (100). Here, the audio information may be audio output from content played on the display device (200) or a voice signal generated by a user's (400) speech.
[0048] The processor (120) can obtain content information about content output from the display device (200) based on audio information and speech information uttered by a user located within a space in relation to the content and store the information in the memory (110).
[0049] Here, content information and speech information may be information used to identify the context of the content. Content information may be information about the content obtained based on an audio signal output from the display device (200). In this case, content information may include information that can identify the content, such as background music, the content title, the name of the content provider, episode information, and character dialogue.
[0050] Speech information may be information about the user (400) obtained based on the user's voice. The utterance information may include information about the text spoken by the user in relation to the content, the time of the utterance, and the emotion corresponding to the user's utterance.
[0051] For example, let's say that audio information "A" and "B" are received by the electronic device (100) via the communication unit (130). The processor (120) can identify whether the received audio information includes an audio signal output from the display device (200) or a voice signal spoken by the user (400). The context of a content can change from time to time while a single content is played from beginning to end. The processor (120) can determine the context of the content from time to time based on the audio information.
[0052] For example, the processor (120) can identify that audio information “A” is an audio signal output from the display device (200) and voice information “B” is a voice signal of the user (400). At this time, the processor (120) can obtain content based on the audio signal “A” output from the display device (200).
[0053] Here, the processor (120) can identify the content based on the dialogue of the characters included in the audio information of the content, the background music of the content, and the sound effects. Here, the content content can include the theme of the content, the emotions of the characters, the relationships between the characters, the emotions contained in the dialogue, etc.
[0054] In addition, the processor (120) may obtain the audio signal "A" output from the display device (200) and then use a search engine program to obtain the content provision source, content title, and episode information. In addition, the processor (120) may extract the features of the content based on the audio information of the content. The processor (120) may identify the content provision source, content title, episode information, etc. based on the extracted features of the content.
[0055] That is, the processor (120) can obtain information related to the content, such as the content providing source (app) corresponding to the voice information is a content providing platform called AAA OTT, the content title is ABCD, the episode information corresponds to episode 10, and the lines of the actors appearing and the background music, and store the information in the memory (110).
[0056] In addition, the processor (120) can analyze the voice signal "B" uttered by the user to obtain the text, the user's emotion, and the time of utterance. The user's emotion can be information estimating the user's emotion when the corresponding voice signal is uttered. The processor (120) can estimate the emotion based on the text. For example, if the user utters "This is really fun" quickly and strongly, the processor (120) can determine that the user is in a pleasant emotional state or excited state. The processor (120) can also estimate the user's emotion by comprehensively considering not only the text but also the size (amplitude), speed (frequency), etc. of the user's voice signal uttered. For example, if the user utters the exclamation "Wow" loudly and repeatedly, the user can be estimated to be in an excited state watching the corresponding content with enjoyment.
[0057] Speech time is information about when the user uttered the corresponding voice signal. Speech time can be a relative time determined based on the content's playback time.
[0058] For example, if a user utters “Wow, this is really fun” at 9 minutes and 10 seconds after the content is played, the processor (120) can store information that the user uttered an utterance corresponding to the emotion of “fun” at 9 minutes and 10 seconds of the content based on the received voice information.
[0059] The processor (120) can receive a user request input to a display device or an external device through a communication unit.
[0060] A user request may be a request or question made by a user (400) regarding content. A user request may include a question related to content content, a search request related to content, and a recommendation request for related content.
[0061] While watching content, a user (400) may voice a question or request regarding the content. In this case, a condition may be set for a trigger voice for voice recognition to be uttered first, but this is not necessarily the case. Even if the user speaks without a separate trigger voice, the external device (300) may be configured to receive the input.
[0062] A user request uttered by a user is input to an external device (300). The external device (300) can transmit the user request to an electronic device (100).
[0063] In this way, the user (400) can input a user request through natural speech, but is not necessarily limited thereto. For example, the user (400) can directly input a user request through a user interface included in an external device (300), a user interface displayed on a display device (200), a user interface included in an electronic device (100), etc., or can input a user request through a motion recognition method. When implemented through a motion recognition method, when the user (400) takes a predetermined pose, a camera built into or connected to the external device (300) or the display device (200) can capture the pose, analyze the pose, and extract a command corresponding to the pose. The external device (300) or the display device (200) can transmit the extracted command to the electronic device (100).
[0064] The processor (120) can identify the context of the content based on the content information and speech information.
[0065] Here, the context of content can refer to the situational and social background information necessary to understand the content and interpret its meaning. Alternatively, it can be described as "situational information." Content context can include the content's plot, character interactions, background music, visual elements, production background, and social context.
[0066] For example, if the dialogue between characters in a content contains curses or other derogatory words, the context of the content can be identified as a situation involving an argument between the characters. Alternatively, if the dialogue between characters includes expressions of affection for the other person, the context can be identified as a romantic one. The context of the content is not necessarily determined based on the dialogue itself, but can also be determined based on background music or sound effects. For example, if the dialogue includes bursts or screams, the context can be identified as a situation involving physical conflict. Furthermore, if fast-paced background music is playing, the context can be identified as a situation involving a character or a car running. The context of the content can be identified by considering not only the audio signals included in the content but also the user's spoken voice. For example, if the user utters a voice signal such as "You're running well" or "You're fast" while fast-paced background music is playing, the processor (120) can combine the two pieces of information to identify the situation as a situation involving a character or other object running.
[0067] For another example, assume that a user (400) is watching an AAA travel program featuring a character named x. The processor (120) can identify the context of the content based on content information, such as that the character x is traveling to the Uyuni Salt Flats in South America in the summer and that x's current emotional state is joy.
[0068] In addition, if the processor (120) includes an exclamation in the user's speech information, i.e., the voice signal, the processor (120) can identify the context of the content as the user (400) feeling 'surprise' at 10 minutes and 30 seconds into the content based on the exclamation. In addition, if the processor (120) identifies the voice signal 'fun' at 12 minutes and 03 seconds, the processor (120) can identify the context of the content as the user feeling 'fun' at that point.
[0069] The processor (120) can generate a response signal corresponding to the user's request based on the context.
[0070] *64For example, let us assume that the AAA travel program described above is being played on the display device (200). The processor (120) can identify the context of the content based on the content information and speech information.
[0071] At this time, if the user request is 'Tell me how to get there', the processor (120) can generate a response signal based on the context of the content. The processor (120) can identify that 'there' is the Uyuni Salt Flats based on the context of the content, and can search for a way to get to the Uyuni Salt Flats from the current location and generate a response signal. The processor (120) can include content context information, content information, voice signals, etc. in the keywords for search, transmit them to various web servers connected through a network, and then reply with the search results. The web server can be various portal servers or specialized search servers. The processor (120) can select information to respond with based on the reply results.
[0072] Specifically, the processor (120) can generate a response signal containing information such as transportation from Seoul to the Uyuni Salt Flats, as well as information on the timing of visiting the Uyuni Salt Flats, tours, altitude sickness weather, and other tips. When there are many search results, the processor (120) can select a certain number of search results in the order of the search results containing the most keywords and include them in the response signal; however, this is not necessarily limited to this, and all search results can be included.
[0073] The processor (120) can transmit a response signal to the display device through the communication unit (130).
[0074] The response signal may be output in the form of text or voice from the display device (200). However, it is not limited to the above, and may be output as a voice signal from the display device (200), or may be output directly through voice information from the electronic device (100), and may be transmitted to the user's (400) portable terminal device.
[0075] Meanwhile, while the above illustrates and describes only the simple configuration of an electronic device, various additional configurations may be included during implementation. These are described below with reference to FIG. 3.
[0076] FIG. 3 is a block diagram for explaining the configuration of an electronic device implemented in a form having a display. Referring to FIG. 3, the electronic device (100) may include a memory (110), a processor (120), a display (140), a communication unit (130), a microphone (150), and a user interface (160). Among the operations of the memory (110), the processor (120), and the communication unit (130), duplicate descriptions of operations that are the same as those described above will be omitted.
[0077] The display (140) may be implemented as a display including a self-luminous element or a display including a non-luminous element and a backlight. For example, it may be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an LED (Light Emitting Diodes), a micro LED, a Mini LED, a PDP (Plasma Display Panel), a QD (Quantum dot) display, a QLED (Quantum dot light-emitting diodes), etc. The display (140) may also include a driving circuit, a backlight unit, etc., which may be implemented in a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc.
[0078] The processor (120) can play content through the display (140) or output a response signal according to the user's request information to the display (140).
[0079] A user interface (160) is a configuration created for interaction between two or more systems, devices, programs, or users.
[0080] The user interface (160) is a configuration for directly receiving various user commands, etc. from the user. The user interface (160) may be implemented as a touch screen, a touch pad, a button, etc. For example, if implemented as a touch screen, the user interface (160) may include a touch detection sensor built into the display (140). The processor (120) controls the display (140) to display a soft keyboard or a UI screen, and can recognize a user touch on the soft keyboard or a touch or drawing on the UI screen, etc., through the user interface (160). The processor (120) can identify a user request or user query input by the user based on the recognized result.
[0081] The processor (120) may obtain user request information through the user interface (160) or store user emotional information, etc., in the memory (110) when such information is input. In addition, the processor (120) may directly sense the audio signal of the content played on the display device (200) and the user's voice through the microphone (150). Alternatively, as described above, the processor may receive audio information including audio signals and voice signals through the communication unit (130).
[0082] As described above, the processor (120) can identify the context of content based on audio signals and voice signals. Reference data for identifying the context of content may be stored in advance in the memory (110) or may be updated and stored periodically. The reference data may include various words or sound effects that can infer a situation. The processor (120) can identify the context of content by comparing content information and voice signals with the reference data. However, the present invention is not limited thereto, and the processor (120) may also identify the context of content using an artificial intelligence model.
[0083] In one embodiment, the memory (110) can store an artificial intelligence model. In an embodiment utilizing an artificial intelligence model, the processor (120) can input a user's request, content information, and speech information into the artificial intelligence model to generate a response signal based on the context of the content. The artificial intelligence model may be a model trained using training data including dialogue, music, audio signals, etc. labeled with various contexts.
[0084] The AI model can be trained to output information about the context of the content based on content information and spoken voice, or it can be trained to input not only this information but also user requests and output a response signal based on the context of the content.
[0085] FIGS. 4 to 6 are diagrams illustrating generation of a response signal according to a user's request, according to at least one embodiment of the present disclosure.
[0086] FIG. 4 is a diagram illustrating a case where a user's request is a question about content, according to at least one embodiment of the present disclosure.
[0087] According to FIG. 4, if a user's request includes a question about content, the processor (120) can obtain an answer to the question based on the context of the content and transmit a response signal including the answer to the display device through the communication unit.
[0088] Here, content-related questions can refer to user inquiries related to the information, theme, plot, elements, situation, and background contained within the content. Specifically, content-related questions can include questions about character interactions, the content's theme, audiovisual elements such as audio and images, the intended message conveyed through the content, the context in which the content was created, and the structure of the content, such as the introduction, body, and conclusion.
[0089] For example, let's assume that while content A is being played on a display device (200), a user (400) asks a question related to the content, "Why is that man angry?" (700). The user's question may be input through the display device (200) or an external device (300) and then transmitted to the electronic device (100).
[0090] The processor (120) of the electronic device (100) can obtain content information based on the received audio information (710). Here, the processor (120) can obtain and store content information indicating that the content provider of content A is an A OTT service, the content title is A drama, the episode is episode 9, and the dialogue of the corresponding character.
[0091] The processor (120) can determine the context of the content based on the content information, and can obtain information about who "that man" is and "why he is angry" based on the context of the content. Specifically, the processor (120) tracks the previous situation based on the time of the question, and obtains the name of the character displayed at the time the user spoke and the dialogue with surrounding characters. The processor (120) can identify the context of the content based on the obtained information. If the identification result identifies that the son of "that man" referred to by the user caused an accident, the processor (120) can generate a response signal based on the context.
[0092] For example, the processor (120) can obtain a response signal from 'Character A' that 'his son was in a car accident and he is angry.' In this case, the processor (120) can directly generate a response signal according to the context of the content, or can extract previous lines as is to generate a response signal so that the user can understand the situation through those lines. That is, if the line "My son was in a car accident" came out in the previous scene, the processor (120) can include the line identified in the audio signal as is in the response signal. As described above, operations such as identifying the context of the content or generating a response signal can be implemented rule-based, but are not limited thereto and can also be implemented using an artificial intelligence model.
[0093] The processor (120) can transmit a response signal to the display device (200) through the communication unit (130). The display device (200) can output the response signal as text on the display based on the response signal (720).
[0094] FIG. 5 is a diagram illustrating a case where a user's request is a search request related to content, according to at least one embodiment of the present disclosure.
[0095] When a user requests a search related to content, the processor (120) can input the search request into a search engine program to obtain search results corresponding to the search request. The search engine program may be a separately provided application for search purposes, or a browser program for accessing a web server or the like to perform a search. The processor (120) can transmit a response signal containing the search results to a display device via a communication unit.
[0096] Specifically, the processor (120) can generate a search word corresponding to a search request based on content information and speech information, input the search word into a search engine to obtain the result, and then generate a response signal.
[0097] Search requests related to content can be made in various ways. For example, a user may request a search for information, topics, plots, elements, and backgrounds contained in the content. Alternatively, a search request may be entered in the form of a question asking what product an item (e.g., a car, a bag, a cell phone, etc.) appearing in a specific scene is. The processor (120) analyzes each pixel of the image frame displayed at the time the user's search request is entered, extracts the edges of objects contained within the image frame, and then identifies the item based on the shape of the edges. In the case of a bag-shaped object, the processor (120) may transmit a search request for bags of a similar shape to the detected shape to each server and receive the search results. A search engine for such searches may also be implemented using an artificial intelligence model.
[0098] For example, let's assume that a travel program B is being played on a display device (200) and a user (400) asks a question about a search related to the content. The processor (120) can store the content B as a source of the content B, the title of the content B as a travel program, the episode number 5, and the content's statement "Arrived at the Uyuni Desert!".
[0099] At this time, the user (400) can make a request for a search related to content, such as 'the cost of traveling there' (800). The processor (120) can specify a search term based on the context of the content. The processor (120) can identify information corresponding to the travel region, travel period, and travel route of the content being viewed based on the content information. Here, the processor (120) can input the request information for 'the cost of traveling to the Uyuni Desert' into a search engine program and obtain the answer that 'it costs 3 million won per adult for 2 nights and 3 days'. The processor (120) can transmit a response signal including the answer that 'it costs 3 million won per adult for 2 nights and 3 days' to the display device (200) through the communication unit. The display device (200) can output the received response signal on the display (820).
[0100] FIG. 6 is a diagram for explaining a case where a user's request is a content recommendation request according to at least one embodiment of the present disclosure.
[0101] When a user request requesting content recommendation is input (900), the processor (120) can select content that matches the content recommendation request based on the context of the content and obtain a response signal including information about the selected content (910). The processor (120) can transmit the obtained response signal to a display device via a communication unit (920).
[0102] For example, the processor (120) can recommend interesting movies based on information about the content context acquired from multiple display devices. Specifically, rather than recommending content by identifying the genre or characteristics of the content based on information classified by the content provider, the processor (120) can recommend content by identifying the context of the content in real time based on content information and user speech information. Furthermore, the processor (120) can recommend content based on the user's emotional information about a specific scene and the content context.
[0103] For example, let's assume that a user (400) requests "recommendations for interesting movies." The processor (120) may store content contexts received from multiple display devices in the memory (110). At this time, the content contexts may include user (400) speech information regarding the corresponding content. The user's speech information may include emotional information regarding the corresponding content, speech time corresponding to the emotional information, and the spoken text.
[0104] The processor (120) can identify the characteristics of the content based on contextual information about the content. Based on the identified characteristics of the content, the processor (120) can recommend the content to the user (400) if the user feels "fun" about the content at a preset rate or higher. A method for updating the content emotion information is described in detail in FIG. 7.
[0105] Additionally, if users feel that content A is ‘fun’ a lot between 10 and 11 minutes, the processor (120) can recommend the portion corresponding to 10 to 11 minutes of content A to the user.
[0106] According to one embodiment of the present disclosure, audio information may include an audio signal output according to playback of content and a voice signal spoken by a user.
[0107] Additionally, the processor (120) can obtain content information based on an audio signal among audio information and obtain speech information based on a voice signal.
[0108] Additionally, content information may include at least one of the following: the name of the source providing the content, the title of the content, episode information, and text information included in the content. Speech information may include information regarding at least one of the text, utterance time, and emotion uttered by the user in relation to the content.
[0109] According to one embodiment of the present disclosure, if the text spoken by the user is identified as an emotional expression related to content content, the processor (120) can obtain emotional information based on the spoken text.
[0110] The processor (120) identifies whether the text spoken by the user is an expression related to content and an expression of emotion.
[0111] Additionally, the processor (120) can store the acquired emotional information and the speech time corresponding to the emotional information in the memory (110). Furthermore, the processor (120) can identify the content context based on the emotional information and the speech time, and can recommend content to the user based on this.
[0112] The above has described an embodiment of recommending content by taking into account the user's emotional information, but it is not necessarily limited thereto, and content recommendation may be performed based on various content information such as content genre, rating, and number of viewers. For example, if the user utters "Recommend a drama," the processor (120) may randomly recommend content included in the drama genre. Alternatively, if the user utters "Recommend an interesting drama," the processor (120) may recommend a drama with a high rating or a large number of viewers among the content included in the drama genre.
[0113] Details of obtaining emotional information based on text spoken by a user (400) and recommending content to a user based on the content context are described in FIG. 7.
[0114] FIG. 7 is a diagram illustrating a method for updating emotional information according to at least one embodiment of the present disclosure.
[0115] According to FIG. 7, the processor (120) can receive a voice signal uttered by a user from an external device (300) (S410). The processor (120) determines the content of the user's speech contained in the voice signal (S420), and if the content of the user's speech corresponds to a content-related utterance, the processor (120) can identify the user's emotion (S430). Here, the processor (120) can update the content context based on the identified user's emotion.
[0116] For example, let's assume that the following utterance information is input: "Let's meet tomorrow at 3 o'clock in Gangnam" (hereinafter utterance A), "That scene is funny" (hereinafter utterance B), and "I was really surprised yesterday" (hereinafter utterance C). The processor (120) can first identify whether the utterance of the user (400) is a content-related utterance.
[0117] Since utterance A is neither a content-related utterance nor an emotion-related utterance, and utterance C is an emotion-related utterance but not a content-related utterance, the processor (120) does not collect utterances A and C. Here, utterance B is a content-related emotional utterance, so the processor (120) identifies the emotion of the user (400) based on utterance B.
[0118] Here, let us assume that utterance B was uttered at 10 minutes and 30 seconds into content X. The processor (120) can obtain emotional information indicating that the user (400) felt the emotion of 'fun' at 10 minutes and 30 seconds into the content. Here, the processor (120) can store the obtained emotional information in the memory (110).
[0119] In addition, not limited to the above-described speech, the processor (120) can obtain emotional information of the user (400) based on sounds related to emotions, such as the user's (400) laughter or crying sounds.
[0120] Here, user emotional information about the content can be obtained from each of the multiple display devices, as described above. The user emotional information about the content obtained from each of the multiple display devices can be updated in a cloud format.
[0121] Additionally, information regarding real-time reviews of content by multiple users can be obtained based on emotional information acquired from multiple display devices. Here, the electronic device (100) can receive information regarding real-time reviews of content by multiple users and store it in memory (110). Furthermore, content can be recommended based on the real-time reviews of multiple users stored in memory (110).
[0122] For example, let us assume that the processor (120) obtains emotional information indicating that the user (400) felt the emotion of 'sadness' at 9 minutes and 30 seconds of content A and stores it in the memory (110).
[0123] Here, the processor (120) can update the acquired emotional information to various cloud servers. The cloud server can obtain real-time user reviews of content based on the user's emotional information regarding the content acquired from multiple electronic devices.
[0124] That is, the cloud server can identify that multiple users felt the emotion "sadness" around 9 minutes and 30 seconds into content A. The cloud can transmit a review of the acquired content to the electronic device (100). The processor (120) can identify the context of the content based on the received review. Here, the processor (120) can recommend content to the user based on the identified context of the content.
[0125] FIG. 8 is a diagram illustrating generation of a response signal according to a user's request, according to at least one embodiment of the present disclosure.
[0126] According to FIG. 8, the electronic device (100) can receive user request information from an external device (300) (S510). The processor (120) can determine the user's utterance (S515). Here, if the user's utterance is an emotional expression related to content, the content context information can be updated based on the utterance. The details of updating the content context information based on the user's utterance are omitted as they have been described above in FIG. 7.
[0127] Here, the user's utterances may include questions about content, questions about content-related searches, and questions about content recommendations. If the user's request information is a question related to content (S520), response information can be generated based on the content context (S535). If the content context information has not been learned, the processor (120) can input the viewed content information into a search engine program to obtain a response to the user's input request information (S535).
[0128] If the user's utterance is a question about content-related search (S525), the processor (120) can input a question about content-related search into a search engine based on the context of the content to obtain response information to the question (S550).
[0129] If the user's utterance is request information for content recommendation (S530), the processor (120) can obtain content according to the user request information as response information based on the recommendation list.
[0130] FIG. 9 is a diagram for explaining acquisition of content information and user speech information and content context learning according to at least one embodiment of the present disclosure.
[0131] The electronic device (100) can acquire audio information (600) from an external device (300), generate content-related information based on the acquired audio information, and learn content context information (630) (610). The content-related information can include content information (615) and user speech information (620). Here, the electronic device (100) can acquire content-related information received from multiple display devices and learn content context information (640).
[0132] At this time, information about content received from multiple display devices may include not only content information and speech information, but also the number of viewers and the number of recommendations for the content.
[0133] FIG. 10 is a flowchart illustrating a method for controlling an electronic device according to at least one embodiment of the present disclosure.
[0134] According to FIG. 10, an electronic device can receive audio information from a display device or an external device within an external space in which the display device is placed (S1010). In addition, the electronic device can obtain and store content information about content output from the display device and speech information uttered by a user located within the space in relation to the content based on the audio information (S1020). Here, when a user request input to the display device or the external device is received, the context of the content can be identified based on the content information and speech information (S1030). In addition, the electronic device (100) can generate a response signal corresponding to the user request based on the identified context (S1040) and transmit the generated response signal to the display device (S1050).
[0135] Here, the details of identifying the context of the content based on audio information, content information, speech information, and content information and speech information are omitted as they have been described above.
[0136] Meanwhile, the control method of the electronic device described in FIG. 10 can be performed by an electronic device having the configuration of FIGS. 2 and 3, but is not necessarily limited thereto, and can be performed by an electronic device having various other configurations.
[0137] Additionally, the various embodiments described above may be implemented independently, or may be implemented in whole or in part in combination with various other embodiments of the present disclosure. The methods according to the various embodiments of the present disclosure described above may be implemented in the form of applications installable on existing electronic devices.
[0138] The methods according to the various embodiments of the present disclosure described above can be implemented only with a software upgrade or a hardware upgrade for an existing electronic device.
[0139] The various embodiments of the present disclosure described above may also be performed through an embedded server provided in an electronic device, or an external server of at least one of the electronic device and the display device.
[0140] According to an example of the present disclosure, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage media that can be read by a machine (e.g., a computer).
[0141] Specifically, a recording medium storing a program for executing a control method including a step of receiving audio information input to a display device or an external device within a space in which the display device is arranged, a step of acquiring and storing content information about content output from the display device based on the audio information and speech information uttered by a user located within the space in relation to the content, a step of identifying a context of the content based on the content information and the speech information when a user request input to the display device or an external device is received, a step of generating a response signal corresponding to the user request based on the context, and a step of transmitting the response signal to the display device may be provided.
[0142] A device may include an electronic device according to the disclosed embodiments, which is a device that calls a command stored from a storage medium and can operate according to the called command. When the command is executed by a processor, the processor may directly or under the control of the processor perform a function corresponding to the command using other components. The command may include code generated or executed by a compiler or interpreter. A storage medium readable by the device may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish whether data is stored semi-permanently or temporarily in the storage medium.
[0143] According to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0144] Each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of one or more entities, and some of the sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the respective components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0145] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person skilled in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea of the present disclosure.
Claims
1. In electronic devices, Ministry of Communications, memory, comprising at least one processor; At least one processor above, Receives audio information input to a display device or an external device within a space where the display device is placed through the communication unit, Based on the above audio information, content information about content output from the display device and speech information spoken by a user located within the space are acquired and stored in the memory, When a user request input to the display device or the external device is received through the communication unit, the context of the content is identified based on the content information and the speech information, and a response signal corresponding to the user request is generated based on the context. An electronic device that transmits the response signal to the display device through the communication unit.
2. In paragraph 1, The above memory stores the artificial intelligence model, At least one processor above, An electronic device that inputs the user request, the content information, and the speech information into the artificial intelligence model to generate the response signal based on the context of the content.
3. In paragraph 1, At least one processor above, An electronic device that, when the user request includes a question about the content, obtains an answer to the question based on the context of the content and transmits the response signal including the answer to the display device through the communication unit.
4. In paragraph 1, At least one processor above, An electronic device that, when the user request includes a search request related to the content, inputs the search request into a search engine program to obtain a context of the content and a search result corresponding to the search request, and transmits the response signal including the search result to the display device through the communication unit.
5. In paragraph 1, At least one processor above, If the above user request includes a content recommendation request, An electronic device that selects content that matches the content recommendation request based on the context of the content, and transmits the response signal including information of the selected content to the display device through the communication unit.
6. In paragraph 1 The above audio information includes an audio signal output according to the playback of the content and a voice signal spoken by the user. At least one processor above, Obtaining the content information based on the audio signal, Obtaining the speech content information based on the above voice signal, The above content information is, Including at least one of the name of the source providing the content, the title of the content, episode information, and text information included in the content, The above utterance information is, An electronic device comprising information about at least one of text, utterance time, and emotion spoken by the user in relation to the above content.
7. In paragraph 6, At least one processor above, If the text spoken by the user is identified as an emotional expression related to the content, the emotional information is obtained based on the spoken text, The above-mentioned acquired emotional information and the utterance time corresponding to the emotional information are stored in the memory, An electronic device that identifies the content context based on the emotional information and the speech time.
8. In a method for controlling an electronic device, A step of receiving audio information input to a display device or an external device within a space in which the display device is placed; A step of obtaining and storing content information about content output from the display device and speech information spoken by a user located within the space based on the audio information; When a user request input to the display device or the external device is received by the electronic device, a step of identifying the context of the content based on the content information and the speech information; A step of generating a response signal corresponding to the user request based on the above context; and A control method comprising: a step of transmitting the response signal to the display device; 9. In paragraph 8, The step of generating the above response signal is: A control method for inputting the user request, the content information, and the speech information into an artificial intelligence model to generate the response signal based on the context of the content.
10. In paragraph 8, The step of generating the above response signal is: A control method for obtaining an answer to the question based on the context of the content and transmitting a response signal including the answer to the display device when the user request includes a question about the content.
11. In paragraph 8, The step of generating the above response signal is: A control method for inputting the search request into a search engine program to obtain the context of the content and search results corresponding to the search request, if the user request includes a search request related to the content, and transmitting the response signal including the search results to the display device.
12. In paragraph 8, The step of generating the above response signal is: If the above user request includes a request for recommendation of content, A control method for selecting content that matches the content recommendation request based on the context of the content, and transmitting the response signal including information of the selected content to the display device.
13. In paragraph 8, The above audio information includes an audio signal output according to the playback of the content and a voice signal spoken by the user. The above control method is, A step of obtaining the content information based on the audio signal; further comprising: The above content information is, Including at least one of the name of the source providing the content, the title of the content, episode information, and text information included in the content, The above utterance information is, A control method comprising information about at least one of text, utterance time, and emotion spoken by the user in relation to the above content.
14. In paragraph 13, A step of obtaining the emotional information based on the text spoken by the user, when the text spoken by the user is identified as an emotional expression related to the content; A step of storing the acquired emotional information and the utterance time corresponding to the emotional information; and A control method further comprising: a step of identifying the content context based on the emotional information and the speech time; 15. In a non-transitory computer-readable medium storing a program for executing a control method for controlling an electronic device, the control method comprises: A step of receiving audio information input to a display device or an external device within a space in which the display device is placed; A step of obtaining and storing content information about content output from the display device based on the audio information and speech information uttered by a user located within the space in relation to the content; A computer-readable recording medium comprising: when a user request input to the display device or the external device is received, a step of identifying the context of the content based on the content information and the speech information; a step of generating a response signal corresponding to the user request based on the context; and a step of transmitting the response signal to the display device.
Citation Information
Patent Citations
Display device and speech search method thereof
KR1020140028540A
Context based VOD searching system and VOD searching method using same
KR1020150022088A
Intelligent automated assistant for media navigation
KR1020180135884A
image display apparatus and information providing method thereof
KR102185700B1
KR20230130580A