Electronic apparatus and control method thereof
The electronic device addresses the issue of message overload during content consumption by grouping users based on their responses to the content and displaying user-specific objects, thereby enhancing user immersion and reducing interruptions.
Patent Information
- Application Number
- PCT/KR2024/012733
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-08-26
- Publication Date
- 2025-05-22
AI Technical Summary
Users experience interruptions in content viewing due to receiving a large number of unrelated chat messages while using chat services during content consumption.
An electronic device with a camera, display, and processor that identifies a user's response to content and groups users based on context information and user response scores, allowing the device to control the display to output user-specific objects in designated areas.
Enhances user immersion by allowing users to easily identify and interact with other users sharing similar opinions, reducing message overload and minimizing interruptions in content viewing.
Smart Images

Figure KR2024012733_22052025_PF_FP_ABST
Abstract
Description
Electronic device and method of controlling the same
[0001] The present disclosure relates to an electronic device and a control method thereof. The present disclosure relates to an electronic device and a control method thereof that analyzes user responses to content to identify a group including users.
[0002] Recently, various content platforms or content-providing electronic devices have been offering content-related chat services. This allows users viewing content on content platforms or electronic devices to move beyond the traditional one-way viewing experience and engage in real-time chat with other users watching the same content. Notably, these chat services also increase user immersion by allowing users to share their opinions about the content with others while watching.
[0003] However, while using chat services, users may receive numerous chat messages within a short period of time, or messages unrelated to the content or user. Users may experience issues such as interrupting content viewing while reading multiple chat messages or selecting relevant ones.
[0004] According to one embodiment, an electronic device includes a camera, a display, and at least one processor, wherein the at least one processor controls the display to acquire an image including a user through the camera while content is output through the display, identify a first group corresponding to the user among a plurality of groups corresponding to the content based on context information corresponding to the content and the image, and output a first object corresponding to the user to a first area of the display corresponding to the first group.
[0005] The at least one processor may obtain user response information including a user response score corresponding to each of the plurality of groups based on the context information and the image, and identify the first group based on the user response information.
[0006] The at least one processor may obtain context information including a context type and a context-corresponding probability value related to the content, obtain a response type corresponding to the context type based on the image, and obtain the user response information based on the context-corresponding probability value and the response type.
[0007] The at least one processor may identify a group corresponding to the identified user response score as the first group when a user response score greater than or equal to a threshold value is identified among a plurality of user response scores included in the user response information.
[0008] Further comprising a memory storing history information including at least one of content viewing history and group identification history, wherein the at least one processor identifies a group corresponding to the user among the plurality of groups based on the history information, and identifies the first group by applying a first weight to a user response score corresponding to the identified group.
[0009] The electronic device further includes a communication interface, and the at least one processor can receive other user response information of other users corresponding to the first group from a server device through the communication interface, and control the display to output a second object corresponding to the other user together with the first object to the first area based on the other user response information.
[0010] The other user reaction information includes an other user reaction score corresponding to the first group, and the at least one processor can identify a first size of the first object, a first location of the first object, a second size of the second object, and a second location of the second object based on the user reaction score and the other user reaction score, and control the display to output the first object in the first area at the first location with the first size, and control the display to output the second object in the first area at the second location with the second size.
[0011] The electronic device further includes a microphone, and the at least one processor can obtain an audio signal including a user voice through the microphone while displaying the content through the display, obtain first user reaction information based on the image, obtain second user reaction information based on the audio signal, and identify the first group based on the first user reaction information and the second user reaction information.
[0012] The at least one processor can identify the first group by adding the first user response score included in the first user response information and the second user response score included in the second user response information for each group.
[0013] The at least one processor may obtain a first response type at a preset time based on the image, obtain a second response type at the preset time based on the audio signal, and, if the first response type corresponds to the second response type, apply a second weight to a user response score corresponding to the preset time to identify the first group.
[0014] According to one embodiment, a method of controlling an electronic device including a display includes the steps of acquiring an image including a user while content is output through the display, identifying a first group corresponding to the user among a plurality of groups corresponding to the content based on context information corresponding to the content and the image, and outputting a first object corresponding to the user to a first area of the display corresponding to the first group.
[0015] The step of identifying the first group may include obtaining user response information including a user response score corresponding to each of the plurality of groups based on the context information and the image, and identifying the first group based on the user response information.
[0016] The control method may include a step of obtaining context information including a context type and a context-corresponding probability value related to the content, a step of obtaining a response type corresponding to the context type based on the image, and a step of obtaining the user response information based on the context-corresponding probability value and the response type.
[0017] The step of identifying the first group may include identifying a group corresponding to the identified user response score as the first group when a user response score greater than or equal to a threshold value is identified among a plurality of user response scores included in the user response information.
[0018] The electronic device stores history information including at least one of content viewing history and group identification history, and the step of identifying the first group may identify a group corresponding to the user among the plurality of groups based on the history information, and apply a first weight to a user response score corresponding to the identified group to identify the first group.
[0019] The above control method may further include a step of receiving other user response information of other users corresponding to the first group from a server device, and a step of outputting a second object corresponding to the other user together with the first object to the first area based on the other user response information.
[0020] The other user reaction information may include an other user reaction score corresponding to the first group, and the control method may include a step of identifying a first size of the first object, a first location of the first object, a second size of the second object, and a second location of the second object based on the user reaction score and the other user reaction score, a step of outputting the first object in the first area at the first location and the first size, and a step of outputting the second object in the first area at the second location and the second size.
[0021] The control method further includes a step of obtaining an audio signal including a user voice while displaying the content through the display, a step of obtaining first user response information based on the image, and a step of obtaining second user response information based on the audio signal, wherein the step of identifying the first group can identify the first group based on the first user response information and the second user response information.
[0022] The step of identifying the first group may identify the first group by adding the first user response score included in the first user response information and the second user response score included in the second user response information for each group.
[0023] The step of identifying the first group may include obtaining a first response type at a preset time based on the image, obtaining a second response type at the preset time based on the audio signal, and, if the first response type corresponds to the second response type, applying a second weight to a user response score corresponding to the preset time to identify the first group.
[0024] An electronic device is provided. The electronic device includes a camera, a display, and at least one processor, wherein the at least one processor acquires an image including a user through the camera while content is output through the display, identifies a first group corresponding to the user among a plurality of groups corresponding to the content based on context information corresponding to the content and the image, and controls the display to output a first object corresponding to the user to a first area of the display corresponding to the first group.
[0025] The at least one processor obtains user response information including a user response score corresponding to each of the plurality of groups based on the context information and the image, and identifies the first group based on the user response information.
[0026] The at least one processor obtains the context information including a context type and a context-corresponding probability value related to the content, obtains a response type corresponding to the context type based on the image, and obtains the user response information based on the context-corresponding probability value and the response type.
[0027] When the at least one processor identifies a user response score greater than or equal to a threshold value among the plurality of user response scores included in the user response information, the processor identifies a group corresponding to the identified user response score as the first group.
[0028] Further comprising a memory storing history information including at least one of content viewing history and group identification history, wherein the at least one processor identifies a group corresponding to the user among the plurality of groups based on the history information, and identifies the first group by applying a first weight to a user response score corresponding to the identified group.
[0029] The electronic device further includes a communication interface, and the at least one processor receives other user response information of other users corresponding to the first group from a server device through the communication interface, and controls the display to output a second object corresponding to the other user together with the first object to the first area based on the other user response information.
[0030] The other user reaction information includes an other user reaction score corresponding to the first group, and the at least one processor identifies a first size of the first object, a first location of the first object, a second size of the second object, and a second location of the second object based on the user reaction score and the other user reaction score, and controls the display to output the first object in the first area at the first location and the first size, and controls the display to output the second object in the first area at the second location and the second size.
[0031] The electronic device further includes a microphone, and the at least one processor obtains an audio signal including a user voice through the microphone while displaying the content through the display, obtains first user reaction information based on the image, obtains second user reaction information based on the audio signal, and identifies the first group based on the first user reaction information and the second user reaction information.
[0032] The at least one processor identifies the first group by adding the first user response score included in the first user response information and the second user response score included in the second user response information for each group.
[0033] The at least one processor obtains a first response type at a preset time based on the image, obtains a second response type at the preset time based on the audio signal, and if the first response type corresponds to the second response type, applies a second weight to a user response score corresponding to the preset time to identify the first group.
[0034] According to one embodiment, a method of controlling an electronic device including a display includes the steps of acquiring an image including a user while content is output through the display, identifying a first group corresponding to the user among a plurality of groups corresponding to the content based on context information corresponding to the content and the image, and outputting a first object corresponding to the user to a first area of the display corresponding to the first group.
[0035] The step of identifying the first group comprises obtaining user response information including a user response score corresponding to each of the plurality of groups based on the context information and the image, and identifying the first group based on the user response information.
[0036] The control method includes a step of obtaining context information including a context type and a context-corresponding probability value related to the content, a step of obtaining a response type corresponding to the context type based on the image, and a step of obtaining the user response information based on the context-corresponding probability value and the response type.
[0037] The step of identifying the first group includes identifying a group corresponding to the identified user response score as the first group when a user response score greater than or equal to a threshold value is identified among a plurality of user response scores included in the user response information.
[0038] The electronic device stores history information including at least one of content viewing history and group identification history, and the step of identifying the first group identifies a group corresponding to the user among the plurality of groups based on the history information, and applies a first weight to a user response score corresponding to the identified group to identify the first group.
[0039] The subject matter of the present disclosure can best be understood by reference to the accompanying drawings.
[0040] FIG. 1 is an example diagram showing an electronic device of the present disclosure displaying an object corresponding to a user.
[0041] Figure 2 is a schematic diagram of an electronic device of the present disclosure.
[0042] FIG. 3 is a flowchart illustrating the operation of an electronic device that displays an object corresponding to a user based on context information of an image and content of the present disclosure.
[0043] Figure 4 is an example diagram showing multiple groups corresponding to the contents of the present disclosure.
[0044] Figure 5 is an exemplary diagram showing a method for calculating a user response score of the present disclosure.
[0045] FIG. 6 is an exemplary diagram illustrating a method for identifying a first group based on history information of the present disclosure.
[0046] FIG. 7 is an exemplary diagram illustrating a method for identifying a first group based on multiple images and user voices of the present disclosure.
[0047] FIG. 8 is an exemplary diagram illustrating a method for identifying a first group including users by applying weights according to the reaction types corresponding to the images and the user voices, respectively, according to one embodiment of the present disclosure.
[0048] FIG. 9 is an exemplary diagram showing an object corresponding to a user of the present disclosure being displayed in an area of the display set corresponding to the first group.
[0049] FIG. 10 is an exemplary diagram illustrating a method of classifying multiple users of the present disclosure into groups and displaying objects corresponding to users included in each group on a display.
[0050] Figure 11 is a detailed block diagram of an electronic device of the present disclosure.
[0051] FIG. 12 is a sequence diagram showing an electronic device implemented as a server device of the present disclosure.
[0052] Figure 13 is a flowchart schematically showing a method for controlling an electronic device of the present disclosure.
[0053] The present embodiments may be modified and have various embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope to specific embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0054] In describing the present disclosure, if it is determined that a specific description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted.
[0055] The terms used in this disclosure are used only to describe specific embodiments. Singular expressions include plural expressions unless the context clearly indicates otherwise.
[0056] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0057] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to instances where (1) at least one A is included, (2) at least one B is included, or (3) at least one A and at least one B are included.
[0058] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0059] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that said component may be directly coupled to said other component, or may be coupled via another component (e.g., a third component).
[0060] On the other hand, when it is said that a component (e.g., a first component) is “directly connected” or “directly connected” to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between said component and said other component.
[0061] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0062] Instead, in some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing the operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform the operations by executing one or more software programs stored in a memory (190) device.
[0063] In the embodiments, a 'module' or 'part' performs at least one function or operation, and may be implemented by hardware or software, or by a combination of hardware and software. A plurality of 'modules' or 'parts' may be integrated into at least one module and implemented by at least one processor, except for a 'module' or 'part' that needs to be implemented by a specific hardware.
[0064] The various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacings drawn in the attached drawings.
[0065] Hereinafter, with reference to the attached drawings, an embodiment according to the present disclosure will be described in detail so that a person having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement the present disclosure. Electronic device (100) Electronic device (100)
[0066] Meanwhile, the terms described in the description below may be used interchangeably with the terms described in the drawings.
[0067] FIG. 1 is an exemplary diagram showing an electronic device (100) according to one embodiment of the present disclosure displaying an object corresponding to a user.
[0068] Referring to FIG. 1, an electronic device (100) according to an embodiment of the present disclosure can display content through a display. Here, the content can be provided from an external server device (e.g., an OTT (Over The Top) platform, etc.).
[0069] While displaying content, the electronic device (100) can display a graphic object (510) corresponding to a user (1) viewing the content, along with objects (520-1 to 520-9, hereinafter 520) corresponding to multiple other users viewing the same content on the display.
[0070] In particular, the electronic device (100) can classify multiple users viewing content, including the user (1), into multiple groups corresponding to the content, and separately display multiple graphic objects (510, 520) corresponding to multiple users included in each group on the display. For example, if the content is a soccer match (or various sports matches and debates, etc.), the electronic device (100) can set up two groups corresponding to each of the two teams participating in the soccer match, and classify multiple users into multiple groups corresponding to each of the two teams participating in the soccer match. The electronic device (100) can classify multiple users into multiple groups according to the team each user supports. The electronic device (100) separately displays multiple objects (510, 520) corresponding to multiple users included in each of the two groups corresponding to each team on the display. The objects are described as graphic objects.
[0071] Through this, the user (1) can identify multiple other users included in the same group as the user (1). To explain again with the above example, if the content is a soccer match, the user (1) can recognize other users who support the same team as the user (1), and can chat (610, 620) with multiple other users included in the same group (520-1, 520-2) in real time through the electronic device (100).
[0072] Multiple users within the same group may have identical or similar opinions and thoughts regarding content. The electronic device (100) of the present disclosure enables a user (1) viewing content to clearly recognize other users within the same group as the user (1), and, in particular, to accurately recognize only chat messages (610 and 620) (or text corresponding to voice messages) entered by other users within the same group, thereby enhancing the user's (1) sense of immersion in the content.
[0073] To this end, the electronic device (100) of the present disclosure identifies a group including a user (1) among multiple groups corresponding to the content. In particular, the group including the user (1) is identified by analyzing the user's response based on a situation occurring in the content. Hereinafter, embodiments of the present disclosure related to this will be described in detail.
[0074] FIG. 2 is a configuration diagram of an electronic device (100) according to an embodiment of the present disclosure.
[0075] Referring to FIG. 2, an electronic device (100) according to one embodiment of the present disclosure includes a camera (110), a display (120), and a processor (130).
[0076] An electronic device (100) according to one embodiment of the present disclosure (e.g., user 1 of FIG. 1) analyzes a response of a user receiving content and identifies a group including the user among a plurality of groups corresponding to the content. As an example, the electronic device (100) may be a device that displays content composed of a plurality of image frames. For example, the electronic device (100) may be implemented as a device that displays content through a display, such as a TV, a smart TV, a signage, a desktop PC, a laptop, a smart phone, a tablet PC, etc. As another example, the electronic device (100) may also be implemented as a server device that provides content.
[0077] The camera (110) captures an image of an object surrounding the electronic device (100) and acquires an image of the object. The camera (110) captures one or more images of a user (1) located around the electronic device (100). To this end, the camera (110) may include an image capturing device such as a CMOS image sensor (CIS) having a CMOS structure, a CCD (Charge Coupled Device) image sensor, etc. However, the present invention is not limited thereto, and the camera (110) may be implemented with various image sensors capable of capturing a subject and camera (110) modules having various resolutions.
[0078] The camera (110) may include at least one of a depth camera (e.g., an IR depth camera), a stereo camera, or an RGB camera. Accordingly, the image acquired through the camera (110) may further include depth information about an object (e.g., a user (1)).
[0079] The display (120) displays various visual information (e.g., content) under the control of the processor (130). Here, the content may include images in various formats, such as text, still images, moving images, and GUI (Graphical User Interface). The display (120) may be implemented as a touch screen along with a touch panel. The display (120) may function as an output unit that outputs information from the electronic device (100) and, at the same time, as an input unit that provides an input interface for the electronic device (100).
[0080] The display (120) can be implemented as a display in various forms, such as a Liquid Crystal Display Panel (LCD), a light emitting diode (LED), an Organic Light Emitting Diodes (OLED), a Liquid Crystal on Silicon (LCoS), a Digital Light Processing (DLP), etc. The display (120) can additionally include additional components depending on its implementation method. For example, the display (120) can also include a driving circuit, a backlight unit, etc., which can be implemented in forms such as a-Si TFT, LTPS (low temperature poly silicon) TFT, OTFT (organic TFT), etc.
[0081] The processor (130) is electrically connected to the camera (110) and the display (120) and controls the overall operation and / or function of the electronic device (100).
[0082] The processor (130) may include one or more of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an APU (Accelerated Processing Unit), a MIC (Many Integrated Core) processor, a DSP (Digital Signal Processor), an NPU (Neural Processing Unit), a hardware accelerator, or a machine learning accelerator. The processor (130) may control any combination of other components of the electronic device (100) and may perform operations related to communication or data processing. The processor (130) may execute one or more programs or instructions stored in a memory (not shown). For example, the processor (130) may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in the memory.
[0083] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-dedicated processor or another specific processor).
[0084] The processor (130) may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., a homogeneous multicore processor or a heterogeneous multicore processor). When the processor (130) is implemented as a multicore processor, each of the multiple cores included in the multicore processor may include internal memory of the processor (130), such as a cache memory or an on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. Each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.
[0085] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among the plurality of cores included in a multi-core processor, or may be performed by the plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.
[0086] In an embodiment of the present disclosure, the processor (130) may mean a system on a chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but the embodiments of the present disclosure are not limited thereto.
[0087] FIG. 3 is a flowchart illustrating an operation of an electronic device (100) that displays a graphic object (510) corresponding to a user based on context information of an image and content according to one embodiment of the present disclosure.
[0088] According to one embodiment of the present disclosure, the processor (130) acquires one image or multiple user images through the camera (110) while displaying content through the display (120) (S32).
[0089] Multiple user images may represent images containing multiple users or multiple images containing users.
[0090] The processor (130) can control the display (120) to display content stored in the electronic device (100) (e.g., stored in the memory of the electronic device (100)) or content received from an external server device.
[0091] While displaying content, the processor (130) can capture a user (1) using the camera (110) of the electronic device (100) and obtain one image or multiple images of the user (1).
[0092] The processor (130) can analyze content displayed through the display (120) to identify contextual information of the content (S34). For example, the processor (130) can input multiple images constituting the content into a neural network model trained to identify contextual information of the content, thereby identifying contextual information of the content. The neural network model can be implemented as a CNN (Convolutional Neural Network) model, an FCN (Fully Convolutional Networks) model, an RCNN (Regions with Convolutional Neuron Networks features) model, a YOLO model, etc.
[0093] Context information may be situational information of the content. That is, the processor (130) may analyze the content, identify a situation occurring in the content, and obtain contextual information.
[0094] For example, if the content is a soccer match (or other various sports match, debate broadcast, vote counting broadcast, etc.), the context information may identify which team is in an advantageous situation among the first and second teams playing the soccer match. If the processor (130) analyzes the content and identifies that the first team is on the attack (i.e., the second team is on the defense) during the soccer match or that another situation advantageous to the first team has occurred, the context information of the content may be identified as “the first team.” On the other hand, if the processor (130) analyzes the content and identifies that the second team is on the attack (i.e., the first team is on the defense) during the soccer match or that another situation advantageous to the second team has occurred, the context information of the content may be identified as “the second team.” If the situation identified in the soccer match is not advantageous to both the first and second teams, or if the situation occurring in the content is not clearly identified, the processor (130) may identify the context information of the content as “undefined” (or neutral).
[0095] As another example, if the content is a concert video featuring multiple singers, the context information may be information about the singers performing on the concert stage. The processor (130) analyzes the content, and if it identifies that a first singer is performing on the concert stage, the context information for the content may include "Singer 1." If it identifies that a second singer is performing, the context information for the content may include "Singer 21."
[0096] Contextual information can be identified differently depending on the type of content. Therefore, the type (and number) of contextual information can correspond to the type of content.
[0097] The type (and number) of context information may be preset depending on the type of content. The processor (130) can identify the type of content displayed on the display (120), identify a preset type (or number) of context information corresponding to the type of content, and analyze the content to identify the context information.
[0098] The processor (130) can recognize text within the content to identify the type and / or contextual information of the content. For example, if the content is a soccer game, the processor (130) can recognize the score displayed on the display (120) to identify contextual information about the soccer game. If, based on the score within the content, it is determined that the first team scored, the processor (130) can identify a situation that is advantageous to the first team during the soccer game and identify the contextual information of the content as “the first team.” As another example, if the content is a concert video, the processor (130) can recognize the singer introduction text displayed on the display (120) to identify contextual information about the concert video.
[0099] The processor (130) can obtain an audio signal of content and identify context information. To this end, the processor (130) can obtain an audio signal output together with the content, extract a voice included in the audio signal, and identify the content type and / or context information of the content through voice recognition on the extracted voice. For example, the processor (130) can recognize the voice of a soccer game caster or commentator to identify context information of the soccer game. Alternatively, the processor (130) can recognize the voice of a singer performing in a concert video to identify context information of the concert video. To this end, the processor (130) can utilize an automatic speech recognition (ASR) model, a natural language understanding (NLU) model, etc. stored in a memory (not shown) of the electronic device (100).
[0100] The processor (130) can identify contextual information of content based on various methods. For example, it can recognize an object within the content, extract object information, and then identify contextual information of the content.
[0101] In the case of a soccer match video, the processor (130) can recognize a soccer player within the content, extract information about the soccer player (e.g., team, player name, number, etc.), and then identify contextual information about the soccer match. In the case of a concert video, the processor (130) can recognize a singer within the content, extract information about the singer (e.g., singer name), and then identify contextual information about the concert video.
[0102] After identifying context information, the processor (130) identifies a group including the user (1) among multiple groups corresponding to the content based on the identified context information and image (S36). The group including the user is described as a group corresponding to the user.
[0103] The processor (130) can identify multiple groups corresponding to the content. A group is a group comprising multiple users who receive the content. The types and number of groups can be set according to the type of content. In particular, the multiple groups can be classified and set according to the plurality of contextual information identified in each content.
[0104] FIG. 4 is an exemplary diagram showing multiple groups corresponding to content according to one embodiment of the present disclosure.
[0105] Referring to FIG. 4, if the content is a soccer match, the plurality of groups corresponding to the content can be classified into a first group, a second group, and a third group, which correspond to the plurality of contextual information (the first team, the second team, and Undefined) of the soccer match. Here, the first group corresponds to the contextual information “Team 1,” the second group corresponds to the contextual information “Team 2,” and the third group corresponds to the contextual information “Undefined.”
[0106] As another example, if the content is a concert video, multiple groups can be categorized into groups corresponding to each singer, each of which is a set of contextual information (first singer and second singer) in the content. The multiple groups can include a "first group" corresponding to the contextual information "first singer" and a "second group" corresponding to the contextual information "second singer."
[0107] The processor (130) can identify a group (hereinafter, referred to as the first group) that includes the user (1) among a plurality of groups corresponding to the content. In particular, the processor (130) can identify the first group that includes the user (1) based on context information of the content and the user image.
[0108] For example, when the context information of the content is identified or the context information of the content is identified as having changed, the processor (130) may acquire one image or a plurality of images through the camera (110). Based on the acquired image(s), the processor (130) may identify a first group including the user (1) among the plurality of groups. That is, the processor (130) identifies the context information of the content and identifies the response of the user (1) using the plurality of images corresponding to the context information in order to identify the response of the user (1) to the content being displayed. The processor (130) may identify the first group including the user (1) by considering the identified response of the user (1).
[0109] Below, a specific method for identifying a group including a user (1) is described.
[0110] The processor (130) can calculate a user (1) response score for each group of users and identify a group including the user (1) among a plurality of groups.
[0111] In this regard, as an example, the processor (130) can calculate a plurality of user response scores, each corresponding to a plurality of groups, based on the identified context information and images.
[0112] The processor (130) can calculate a user (1) reaction score for each of the plurality of groups corresponding to the content. Here, the user (1) reaction score may be a probability value that the user (1) will be included in each group. In particular, while the content is displayed through the display (120), the processor (130) can determine whether the user (1) reaction related to the identified context information is positive, negative, or neutral, and based on the determined user (1) reaction, can calculate a probability value that the user (1) will be included in each group corresponding to the content as the user (1) reaction score.
[0113] FIG. 5 is an exemplary diagram showing a method for calculating a user (1) response score according to one embodiment of the present disclosure.
[0114] For example, the processor (130) can analyze displayed content to identify context types and context-corresponding probability values as context information. The processor (130) can analyze content displayed through the display (120) in real time to identify the context type of the content.
[0115] The processor (130) can identify a context-corresponding probability value. Here, the context-corresponding probability value may be a probability value for the identified context type. The processor (130) can analyze content displayed through the display (120) in real time to identify multiple context types corresponding to the content. Here, the context types and number may be preset corresponding to the content. The processor (130) can calculate a probability value for each context type, and can identify the context type with the highest probability value as the context type (or context information) of the content.
[0116] Referring to FIG. 5, while a soccer match is displayed on the display (120), the processor (130) can calculate the context type of the soccer match and the probability value for the context type. The processor (130) identifies the context type of the soccer match displayed on the display (120) from t0 to t1 as “Undefined” and calculates 0.8 as the probability value for “Undefined”. The processor (130) identifies the context type of the soccer match displayed on the display (120) from t1 to t2 as “Team 1” and calculates 0.9 as the probability value for “Team 1”. In addition, the processor (130) identifies the context type of the soccer match displayed on the display (120) from t2 to t3 as “Team 2” and calculates 0.8 as the probability value for “Team 2”. The processor (130) identified the context type of the soccer game displayed through the display (120) from t3 to the present as “first team” and calculated 0.8 as the probability value for “first team”.
[0117] As described above, the processor (130) can input a plurality of images constituting a soccer game (i.e., content) into a pre-trained neural network model to obtain a context type of the content and a probability value for the context type.
[0118] The processor (130) can identify a response type corresponding to a context type based on the user's image, and can calculate a plurality of user response scores corresponding to each of a plurality of groups based on the context-corresponding probability value and the response type.
[0119] The processor (130) can identify, based on the image, whether the user's (1) response to the context type is positive, negative, or neutral. In particular, the processor (130) can identify the type of response of the user (1) to the context type based on the user image corresponding to the context type. The user response may include the response type. The user response may represent the user's response to the currently output (or displayed) content.
[0120] For example, referring to FIG. 5, the processor (130) can identify the user's (1) response type to the context type, “Undefined,” based on the images acquired from t0 to t1. The processor can identify the response type to the context type, “Team 1,” based on the images acquired from t1 to t2. Similarly, the processor can identify the response type to the context type, “Team 2,” based on the images acquired from t2 to t3, and can again identify the response type to the context type, “Team 3.”
[0121] The processor (130) can identify the user's (1) reaction type by identifying the user's (1) facial expression and the user's (1) behavioral pattern, etc., based on the image. For example, if the processor (130) identifies the user's (1) as frowning or angry based on the image, the processor (130) can identify the user's (1) reaction type as "negative." On the other hand, if the processor (130) identifies the user's (1) as smiling or dancing based on the image, the processor (130) can identify the user's (1) reaction type as "negative." To this end, the processor (130) can input the image into a neural network model trained to identify the user's (1) reaction type, thereby obtaining a result value regarding the user's (1) reaction type.
[0122] The processor (130) can calculate, based on the image, the response type of the user (1) and also a probability value corresponding to the response type. Here, the probability value corresponding to the response type can be a probability value for the identified response type. The processor (130) can calculate, based on a plurality of images, a probability value corresponding to a positive response type of the user (1) and a probability value corresponding to a negative response type. The processor (130) can identify the response type with the highest probability value as the response type of the user (1). In this way, the processor (130) can identify the response type of the user (1) corresponding to each context type.
[0123] The processor (130) may identify the user's (1) response type repeatedly (e.g., at different times or different periods) in relation to the same contextual information. The user's (1) response type may be identified based on multiple images at preset intervals. Alternatively, if the processor (130) identifies that the user's (1) facial expression and / or behavioral pattern have changed based on multiple images, the processor (130) may identify the user's (1) response type corresponding to the user's (1) facial expression and / or behavioral pattern.
[0124] For example, referring to FIG. 5, the processor (130) identified the response type of the user (1) from t0 to t1 as “positive” and calculated 0.6 as the probability value for “positive”. The processor (130) identified the response type of the user (1) multiple times in relation to the same context information (i.e., the context type of the first team) from t1 to t2. That is, the processor (130) repeatedly identified the response type of the user (1) during △t2 (i.e., from t1 to t2). The processor (130) identified the first response type of the user (1) identified from t1 to t2 as “positive” and calculated 0.9 as the probability value for “positive”, identified the second response type as “positive” and calculated 0.8 as the probability value for “positive”, and identified the third response type as “negative” and calculated 0.4 as the probability value for “negative”. The processor (130) identified the response type corresponding to the context type identified from t2 to t3 (i.e., the second team) as “negative” and calculated the probability value for “negative” as 0.8. The processor (130) identified the response type corresponding to the context type identified from t3 to the present (i.e., the first team) as “positive” and calculated the probability value for “positive” as 0.9.
[0125] The processor (130) can identify the user's (1) reaction type as “neutral” if the probability value for the reaction type is less than a preset value.
[0126] The processor (130) can calculate a plurality of user response scores, each corresponding to a plurality of groups, based on the context-corresponding probability value and response type. The processor (130) can calculate a response score of a user (1) for a context type based on the probability value corresponding to the identified context type while the content is displayed. In particular, the processor (130) can select a group corresponding to a context type from among the plurality of groups, and calculate a user (1) response score for the selected group based on the probability value corresponding to the context type.
[0127] The processor (130) may also calculate a probability value corresponding to a context type identified while the content is displayed, a probability value corresponding to a response type of the user (1), and a plurality of user response scores corresponding to each of a plurality of groups based on the response type.
[0128] The processor (130) can select a group corresponding to a context type from among a plurality of groups, and calculate a user (1) response score for the selected group based on a probability value corresponding to the context type and a probability value corresponding to the response type. For example, the processor (130) can multiply a probability value corresponding to the context type and a probability value corresponding to the response type to obtain a result value, and calculate the user (1) response score for the group selected corresponding to the context type. The processor (130) can determine a sign of the user (1) response score based on the response type of the user (1).
[0129] The processor (130) may determine the sign of the user's (1) response score based on the user's (1) response type. If the response type is "positive," the processor (130) may determine the sign of the response score as positive (+), and if the response type is "negative," the processor (130) may determine the sign of the response score as negative (-). The processor (130) may accumulate and calculate the user's (1) response score for each group while the content is displayed.
[0130] Referring back to FIG. 5, the processor (130) may select a third group from among the plurality of groups, as the type of context of the content displayed from t0 to t1 is identified as “Undefined.” The processor (130) may multiply the probability value of 0.8 for the context type (“Undefined”) by the probability value of 0.6 for the response type (“positive”), and obtain (+) 0.48 as the user (1) response score for the third group.
[0131] The processor (130) may select a first group from among a plurality of groups, as the context type of the content displayed from t1 to t2 is identified as “Team 1.” The processor (130) may multiply the probability value of 0.9 for the context type (“Team 1”) by the probability value of 0.9 for the first identified response type (“positive”) in the context type to produce a result value of (+) 0.81. Similarly, the processor (130) may multiply the probability value of 0.9 for the context type (“Team 1”) by the probability value of 0.8 for the second identified response type (“positive”) in the context type to produce a result value of (+) 0.72. Finally, the processor (130) can multiply the probability value of 0.9 for the context type (“Team 1”) by the probability value of 0.4 for the third identified response type (“Negative”) in the context type to produce a result value of (-) 0.36. The processor (130) can add up the produced result values to produce a user (1) response score for the first group. That is, the processor (130) can produce (+) 1.17 as the user (1) response score for the first group.
[0132] The processor (130) can select a second group from among multiple groups, as the type of context of the content displayed from t2 to t3 is identified as “Team 2.” The processor (130) can multiply the probability value of 0.8 for the context type (“Team 2”) by the probability value of 0.8 for the response type (“negative”), and obtain (-)0.64 as the user (1) response score for the second group.
[0133] The processor (130) can select a first group from among the plurality of groups as the context type of the displayed content from t3 to the present is again identified as “Team 1”. The processor (130) can multiply the probability value of 0.8 for the context type (“Team 1”) by the probability value of 0.9 for the first identified response type (“positive”) in the context type to produce a result value of (+) 0.72. The processor (130) can add (+) 0.72 to the previously produced (i.e., produced from t1 to t2) user (1) response score ((+) 1.17)) corresponding to the first group to produce a user (1) response score corresponding to the first group as (+) 1.89.
[0134] The processor (130) can identify a first group including a user (1) among a plurality of groups corresponding to the content based on a plurality of user response scores calculated corresponding to each group.
[0135] In particular, when a user (1) response score that is greater than or equal to a preset value is identified among the plurality of generated user response scores, the processor (130) can identify a group corresponding to the identified user (1) response score as a first group.
[0136] For example, referring back to FIG. 5, if the preset value is (+)1.5, the processor (130) can identify that the user (1) is included in the first group after t3.
[0137] According to one embodiment of the present disclosure, the processor (130) can identify a first group including the user (1) by further considering history information related to the content of the user (1).
[0138] FIG. 6 is an exemplary diagram illustrating a method for identifying a first group based on history information according to one embodiment of the present disclosure.
[0139] For example, the processor (130) may apply a weight (hereinafter, a first weight) to a user's (1) response score corresponding to any one of a plurality of groups based on history information related to the content of the user (1) stored in the memory of the electronic device (100). Here, the history information may include information about content previously viewed by the user (1) and information about a group to which the user (1) who viewed the content is included. The processor (130) may extract history information about content of the same type as the content from among the history information stored in the memory, or history information about the same content from a source device (or server device) providing the content, and then identify a group to which the user (1) was included in the past based on the extracted history information. The processor (130) may determine a group to which the first weight will be applied from among a plurality of groups corresponding to content currently being displayed through the display (120), based on information about the group to which the user (1) was included in the past. Here, the group to which the first weight is applied may be a group in which the user (1) is predicted to be included among a plurality of groups corresponding to the content currently being displayed through the display (120), based on the user's (1) history information.
[0140] For example, referring to FIG. 6, if the content is a soccer match between the first team and the second team, the processor (130) can extract history information about the soccer match from the history information stored in the memory. The processor (130) can identify a group in which the user (1) is predicted to be included among the groups corresponding to the first team and the second team, respectively. In particular, the processor (130) can identify, based on the extracted history information, that the past user (1) has been repeatedly included in the first group corresponding to the first team, and can also predict that the user (1) will be included in the first group corresponding to the first team in relation to the soccer match currently displayed through the display (120). That is, the processor (130) can identify the first group corresponding to the first team among the first team and the second team as the group to which the first weight will be applied. The processor (130) can apply the first weight to the user (1) response score corresponding to the first team.
[0141] The processor (130) can identify a first group including the user (1) among a plurality of groups based on the reaction score of the user (1) to which the first weight is applied and the reaction scores of the remaining users (1).
[0142] The processor (130) can identify a group having a reaction score greater than a preset value among the reaction scores corresponding to the group to which the first weight is applied and the reaction scores corresponding to the remaining groups to which the first weight is not applied as the first group including the user (1).
[0143] For example, referring to FIG. 6, the processor (130) can identify the first group including the user (1) based on the reaction score corresponding to the first group to which the first weight is applied and the reaction score corresponding to the second group. When the first weight is 2 and the value set for the reaction score for identifying the first group is 2.0, the processor (130) can identify that the user (1) is included in the first group based on the reaction score ((+)2.34) of the user (1) corresponding to the first group calculated from t1 to t2. That is, comparing FIG. 5 and FIG. 6, when the history information of the user (1) is used, the processor (130) can identify the first group including the user (1) more quickly.
[0144] The history information of a user (1) may be classified according to the account set in the electronic device (100) (or the program providing the content). That is, when multiple accounts are set in the electronic device (100) (or the program providing the content), the processor (130) may identify the first group by using only the history information corresponding to the account selected by the user (1) among the multiple accounts.
[0145] According to one embodiment of the present disclosure, the processor (130) can identify a user's (1) reaction to content based on various information as well as multiple images acquired through a camera (110), etc.
[0146] For example, the processor (130) may acquire a user's (1) voice through a microphone of the electronic device (100) while displaying content through the display (120), and, based on the image and the user's (1) voice, may calculate a user's (1) reaction score corresponding to the image and the user's (1) reaction score corresponding to the user's (1) voice regarding the identified context information.
[0147] The processor (130) can identify the user's (1) reaction type (and probability value for the reaction type) based on visual information about the user (1). That is, the processor (130) can identify the user's (1) facial expression and behavioral pattern based on the image acquired through the camera (110), thereby identifying whether the user's (1) reaction to the content is positive, negative, or neutral. In this regard, the description of the present disclosure described above is equally applicable, so a detailed description will be omitted.
[0148] The processor (130) can identify the response type (and probability value for the response type) of the user (1) based on auditory information about the user (1) (e.g., information included in the audio signal).
[0149] When the context information of the content is identified or the context information of the content is identified as having changed, the processor (130) can acquire the voice of the user (1) through the microphone and, based on the acquired voice of the user (1), identify whether the reaction of the user (1) is positive, negative, or neutral. For example, the processor (130) can identify the type of reaction of the user (1) based on the volume of the voice of the user (1) and the content of the speech corresponding to the voice of the user (1). However, the present invention is not limited thereto, and the processor (130) can identify the reaction of the user (1) based on various sounds input through the microphone (for example, cheering sounds that may indicate a positive reaction, etc.).
[0150] The processor (130) uses an audio signal (e.g., a voice signal, data) corresponding to the voice of a user (1) input through a microphone, an automatic speech recognition (ASR) model, a natural language understanding (NLU) model, etc. stored in a memory (not shown) of the electronic device (100), to determine the content of speech corresponding to the voice of the user (1), and based on the determination result, can identify whether the user (1) has made a positive or negative speech about the content.
[0151] The processor (130) can calculate the user's (1) response type and the probability value corresponding to the response type based on the user's (1) voice. That is, the processor (130) can calculate the probability value corresponding to the positive response type of the user (1) and the probability value corresponding to the negative response type of the user (1) based on the user's (1) voice. The processor (130) can identify the response type with the highest probability value as the user's (1) response type.
[0152] The processor (130) can identify the response type of the user (1) (and the probability value for the response type) based on the input user (1) voice whenever the user's (1) voice is input through the microphone or when the user's (1) voice is input more than a preset number of times.
[0153] The processor (130) can calculate a user (1) reaction score corresponding to each of a plurality of groups corresponding to the content based on the reaction types identified based on a plurality of images and probability values for the reaction types. The processor (130) can calculate a user (1) reaction score corresponding to each of a plurality of groups corresponding to the content based on the reaction types identified based on the user (1) voice and probability values for the reaction types. Hereinafter, for the convenience of explanation of the present disclosure, the user (1) reaction score corresponding to the plurality of user images will be referred to as a first user (1) reaction score, and the user (1) reaction score corresponding to the user (1) voice will be referred to as a second user (1) reaction score.
[0154] The processor (130) can identify a first group including the user (1) among the plurality of groups based on the first user (1) reaction score and the second user (1) reaction score for each of the plurality of groups. The processor (130) can add up the first user (1) reaction score and the second user (1) reaction score for each of the plurality of groups, and identify a group in which the reaction score obtained by adding up (hereinafter, the third user (1) reaction score) is greater than or equal to a preset value as the first group.
[0155] FIG. 7 is an exemplary diagram illustrating a method for identifying a first group based on a plurality of images and the voice of a user (1) according to one embodiment of the present disclosure. The response types and probability values for the response types identified based on the plurality of images illustrated in FIG. 7 are identical to those in FIG. 5, and therefore, a detailed description thereof will be omitted.
[0156] Referring to FIG. 7, the processor (130) may calculate a response type corresponding to context information and a probability value for the response type while content is displayed through the display (120) based on the user's (1) voice. The processor (130) did not identify the response type of the user (1) from t0 to t1 based on the voice signal. This may be because the user's (1) voice was not input through the microphone, or the user's (1) voice was input, but the probability value for the response type identified based on the user's (1) voice is less than a preset value.
[0157] The processor (130) identified the response type of the user (1) multiple times based on the voice of the user (1) in response to the same context type (i.e., the first team) from t1 to t2. That is, based on the voice of the user (1), the processor (130) repeatedly identified the response type of the user (1) during △t2 (i.e., from t1 to t2). The processor (130) identified the first response type of the user (1) identified based on the voice of the user (1) from t1 to t2 as “positive” and calculated the probability value for “positive” as 0.8, identified the second response type as “negative” and calculated the probability value for “negative” as 0.8, identified the third response type as “positive” and calculated the probability value for “positive” as 0.9, identified the fourth response type as “positive” and calculated the probability value for “positive” as 0.5.
[0158] The processor (130) identified the response type corresponding to the context type (i.e., the second team) identified from t2 to t3 based on the voice of the user (1) as “positive” and calculated the probability value for “positive” as 0.4. The processor (130) identified the response type corresponding to the context type (i.e., the first team) identified from t3 to the present based on the voice of the user (1) as “negative” and calculated the probability value for “negative” as 0.7.
[0159] The processor (130) may calculate a second reaction score based on the identified reaction type and probability value based on the user's (1) voice. Referring again to FIG. 7, the processor (130) may select a first group from among a plurality of groups, as the context type of the content displayed from t1 to t2 is identified as “Team 1.” The processor (130) may multiply 0.9, which is a probability value for the context type (“Team 1”), by 0.8, which is a probability value for the first identified reaction type (“positive”) corresponding to the context type based on the user's (1) voice, to calculate a result value of (+) 0.72. Similarly, the processor (130) may multiply 0.9, which is a probability value for the context type (“Team 1”), by 0.8, which is a probability value for the second identified reaction type (“negative”) corresponding to the context type based on the user's (1) voice, to calculate a result value of (-) 0.72. The processor (130) can multiply the probability value of 0.9 for the context type (“Team 1”) by the probability value of 0.9 for the third identified response type (“positive”) corresponding to the context type based on the user’s (1) voice to produce a result value of (+) 0.81. The processor (130) can multiply the probability value of 0.9 for the context type (“Team 1”) by the probability value of 0.5 for the third identified response type (“positive”) corresponding to the context type based on the user’s (1) voice to produce a result value of (+) 0.45. The processor (130) can add up the produced result values to produce a response score of the second user (1) for the first group. That is, the processor (130) can produce a response score of the second user (1) for the first group as (+) 1.26.
[0160] Likewise, the processor (130) can select a second group from among multiple groups and calculate a second user (1) response score for the second group as (+) 0.32, as the type of context of the content displayed from t2 to t3 is identified as “second team”.
[0161] The processor (130) can select the first group from among the plurality of groups, as the context type of the content displayed from t3 to the present is again identified as “Team 1.” The processor (130) can calculate the reaction score of the second user (1) corresponding to the first group as (-) 0.56. The processor (130) can calculate the reaction score of the second user (1) corresponding to the first team as (+) 0.7 by adding (-) 0.56 to the reaction score of the second user (1) corresponding to the first team ((+) 1.26)) calculated previously (i.e. calculated from t1 to t2).
[0162] The processor (130) can obtain a third user (1) reaction score for each group by adding up the first user (1) reaction score and the second user (1) reaction score for each group. The processor (130) can identify a third user (1) reaction score ((+)2.43)) corresponding to the first group by adding up the first user (1) reaction score ((+)1.17) and the second user (1) reaction score ((+)1.26) corresponding to the first group calculated from t1 to t2. The processor (130) can identify a third user (1) reaction score ((-)0.32)) corresponding to the second group by adding up the first user (1) reaction score ((-)0.64) and the second user (1) reaction score ((+)0.32) corresponding to the second group calculated from t2 to t3. The processor (130) can identify the third user (1) reaction score ((+)2.59) corresponding to the first group by adding the first user (1) reaction score ((+)1.89) and the second user (1) reaction score ((+)0.7) corresponding to the first group calculated from t3 to the present. If the preset value is 2.5, the processor (130) can identify that the user (1) is included in the first group based on the third user (1) reaction score corresponding to the first group identified after t3.
[0163] For example, the processor (130) may identify a response type corresponding to an image and a response type corresponding to a user's (1) voice, and if each identified response type matches the first type, a weight (hereinafter, a second weight) may be applied to one of the plurality of generated user response scores.
[0164] The processor (130) may apply a second weight to the user's (1) reaction score for the specific situation if the user's (1) reaction identified based on multiple images and the user's (1) reaction identified based on the user's (1) voice are of the same type for a specific situation of the content. For example, if a specific situation of the soccer game is identified as occurring while the user (1) is watching a soccer game, the user's (1) facial expression (or behavioral pattern) for the specific situation is identified as positive, and the input user's (1) utterance is identified as positive, the processor (130) may apply a second weight to the user's (1) reaction score for the specific situation.
[0165] In particular, the processor (130) can identify whether a reaction type identified based on an image and a reaction type identified based on the voice of the user (1) are the same in relation to the same context information. For example, if the signs of the first reaction score and the second reaction score calculated in response to the same context information match, the processor (130) can identify that the reaction type identified based on an image and the reaction type identified based on the voice of the user (1) are the same. The processor (130) can apply a second weight to a third reaction score that is the sum of the first reaction score and the second reaction score.
[0166] FIG. 8 is an exemplary diagram illustrating a method for identifying a first group including a user (1) by applying weights according to the response types corresponding to the image and the user (1) voice, respectively, according to one embodiment of the present disclosure.
[0167] Referring to FIG. 8, the processor (130) can identify that the signs of the first reaction score and the second reaction score calculated from t1 to t2 are positive (+). Accordingly, the processor (130) can identify that the user (1) made a positive facial expression (or behavioral pattern) and made a positive utterance in response to a situation advantageous to the first team that occurred in the soccer game from t1 to t2. Accordingly, the processor (130) can apply a second weight to the third reaction score identified by adding the first reaction score and the second reaction score calculated from t1 to t2. When the second weight is 1.5, the processor (130) can identify the third reaction score from t1 to t2 as 3.645 (e.g., 2 x 2.43).
[0168] On the other hand, the processor (130) can identify that the signs of the first reaction score and the second reaction score calculated from t2 to t3 are different. Accordingly, the processor (130) can identify that the user (1) made a negative facial expression (or behavioral pattern) in response to a situation advantageous to the second team that occurred in the soccer game from t2 to t3, while making a positive utterance. In other words, the processor (130) can identify that the visual reaction type and the auditory reaction type identified for the user (1) are different. Accordingly, the processor (130) may not apply the second weight to the third reaction score identified by adding the first reaction score and the second reaction score calculated from t2 to t3.
[0169] The processor (130) can identify that the signs of the first and second reaction scores calculated from t3 to the present are different. Accordingly, the processor (130) may not apply the second weight to the third reaction score identified by adding the first and second reaction scores calculated from t3 to the present.
[0170] The processor (130) can identify a first group including the user (1) among the plurality of groups based on the user (1) reaction score to which the second weight is applied and the remaining reaction scores to which the second weight is not applied. The processor (130) can identify a third reaction score corresponding to each group calculated based on the user (1) reaction score to which the second weight is applied and the remaining reaction scores to which the second weight is not applied. The processor (130) can identify a group corresponding to a third reaction score that is greater than or equal to a preset value among the third reaction scores corresponding to each group as the first group including the user (1).
[0171] Referring again to FIG. 8, if the value set for the reaction score for identifying the first group is 2.5, the processor (130) can identify that the user (1) is included in the first group based on the reaction score ((+)3.645) of the third user (1) corresponding to the first group calculated from t1 to t2. That is, comparing FIG. 7 and FIG. 8, if the reaction type corresponding to the image and the reaction type corresponding to the voice of the user (1) are compared and the second weight is applied to the reaction score of the user (1), the processor (130) can identify the first group including the user (1) more quickly.
[0172] The processor (130) can identify the user's (1) response based on various information in addition to the plurality of images and the user's (1) voice. For example, based on the text input of the user (1) input through the user (1) interface of the electronic device (100), the processor (130) can identify the user's (1) response type (and the probability value for the response type). The electronic device (100) can provide a chat service to multiple users who receive the same content and a UI that can input chat messages between the multiple users. The processor (130) can identify the user's (1) response type based on the chat message of the user (1) input through the user (1) interface, and can calculate a response score (i.e., a fourth response score) corresponding to the chat message. The processor (130) can identify a first group including the user (1) based on the first and second response scores corresponding to the plurality of images and the user's (1) voice, and the fourth response score corresponding to the chat message.
[0173] According to one embodiment of the present disclosure, when a plurality of users are identified based on an image, the processor (130) calculates a user (1) reaction score of each of the plurality of users based on the identified context information and the image, and identifies a first group including each user (1) among the plurality of groups based on the calculated reaction score.
[0174] If the processor (130) identifies that multiple users are viewing the content based on the acquired images, the processor (130) may calculate a reaction score for each user (1) based on the acquired images. In particular, if the context information (e.g., context type) of the content is identified or the context information (e.g., context type) is identified to have changed, the processor (130) may identify a group corresponding to the context information among multiple groups corresponding to the content based on the context information, and may calculate a user (1) reaction score corresponding to the identified group for each user (1) based on the multiple images acquired corresponding to the context information. The processor (130) may identify a group to which each user (1) is included among the multiple groups based on the reaction scores calculated corresponding to the multiple groups of each user (1). To explain again with the above example, if two users (1) are identified watching a soccer game, the processor (130) may calculate reaction scores corresponding to the first group, the second group, and the third group of each user (1) based on the identified context information and the multiple images acquired. If the preset value is 1.5, the reaction score corresponding to the first group of the first user (1) among two users (1) is 2.0, and the reaction score corresponding to the second group of the remaining user (1) (i.e., the second user (1)) is 1.8, the processor (130) can identify that the first user (1) is included in the first group and the second user (1) is included in the second group.
[0175] The processor (130) can display a graphic object (510) corresponding to the user in an area of the display (120) set corresponding to the first group (S38). The object can be described as at least one of a graphic object, a GUI (Graphic User Interface), and a graphic image.
[0176] The processor (130) can set multiple areas corresponding to multiple groups on the display (120) (e.g., separate areas for each group). When a first group including a user (1) is identified, the processor (130) can identify an area set corresponding to the first group among the multiple areas. The processor (130) can display a graphic object (510) corresponding to the user in the identified area.
[0177] The number and positions of the plurality of areas set on the display (120) may vary depending on the type of content. Accordingly, the processor (130) may identify the type of content, identify a plurality of groups corresponding to the content, and, considering the number of the identified groups and the type of content, set an area in which graphic objects corresponding to each group can be appropriately displayed.
[0178] Graphic objects can be displayed in various forms. For example, the processor (130) can generate an avatar corresponding to the user (1) based on data input by the user and then display the avatar on the display (120).
[0179] The processor (130) can display a graphic object (510) corresponding to the user on the display (120) even before the first group is identified. The processor (130) can set a waiting area on the display (120) and display a graphic object (510) corresponding to the user on the waiting area.
[0180] FIG. 9 is an exemplary diagram showing displaying a graphic object corresponding to a user in an area of a display (120) set corresponding to a first group according to one embodiment of the present disclosure.
[0181] Referring to FIG. 9, when the content is a soccer match, the processor (130) may set areas (410, 420, and 430) corresponding to the first to third groups corresponding to the soccer match on the display (120). The processor (130) may set the area (410) corresponding to the first group (the first group corresponding to the first team) on the right side of the display (120), the area (420) corresponding to the second group (the second group corresponding to the second team) on the left side of the display (120), and the area (430) corresponding to the third group (the third group corresponding to Undefined) on the lower center side of the display (120). The processor (130) may display a graphic object corresponding to the user on the area (430) corresponding to the third group before the first group is identified (t1). That is, the processor (130) may set the waiting area to be the same as the area corresponding to the third group. Thereafter, when the processor (130) identifies that the user (1) is included in the first group (i.e., the first group is identified) (t2), the processor (130) can display a graphic object corresponding to the user in the area (410) corresponding to the first group.
[0182] The processor (130) can display graphic objects corresponding to multiple other users receiving the same content, along with graphic objects corresponding to the user, on the display (120). Through this, the electronic device (100) can increase the user's (1) sense of immersion in the content.
[0183] For example, the processor (130) may receive information about multiple other users who receive content from a server device that provides content through the communication interface of the electronic device (100). Here, the other user information may include information about a group to which the other user belongs. The other user information may include information about other user responses. The processor (130) may also receive other user information from other electronic devices that receive the same content from the server device. In other words, the processor (130) may receive other user information about other users who view content through each of the other electronic devices from multiple other electronic devices.
[0184] The processor (130) can identify groups each including multiple other users based on information about multiple other users received from the server device. That is, the processor (130) can classify multiple other users, along with the user (1), into groups. The processor (130) can display a graphic object corresponding to at least one user included in each group in an area set corresponding to each group.
[0185] FIG. 10 is an exemplary diagram illustrating a method of classifying multiple users into groups and displaying graphic objects corresponding to users included in each group on a display (120) according to one embodiment of the present disclosure.
[0186] Referring to FIG. 10, if the content is a soccer game, the processor (130) can receive information (i.e., information on multiple other users) about multiple other users (2-1 to 2-9) who are watching the same soccer game (or using other electronic devices (100) that provide the same soccer game) from the server device. Based on the information on multiple other users, the processor (130) can classify multiple users (1 and 2-1 to 2-9) who are watching the same soccer game, including the user (1), into a first group, a second group, and a third group.
[0187] If there are 9 other users watching the same soccer game as the user (1), the processor (130) can identify groups that include the 9 other users including the user (1). The processor (130) can identify that the first group includes 2 other users including the user (1), the second group includes 3 other users, and the third group includes 4 other users. Accordingly, the processor (130) can display two graphic objects (520-1 and 520-2) corresponding to the two other users (2-1 and 2-2) included in the first group, along with a graphic object (510) corresponding to the user, on the right area (410) of the display (120) set corresponding to the first group. The processor (130) can display three graphic objects (520-2 to 520-5) corresponding to three other users (2-3 to 2-5) in the left area (420) of the display (120) set corresponding to the second group. The processor (130) can display four graphic objects (520-6 to 520-9) corresponding to four other users (2-6 to 2-9) in the central lower area of the display (120) set corresponding to the third group. The graphic objects corresponding to the other users can be generated based on information of preset graphic objects of the other users included in the other user information.
[0188] The processor (130) can display information input by a user (1) on each object. For example, based on a chat service provided by another electronic device (100) (or an external server device) including the electronic device (100), the processor (130) can display chat messages (610 and 620) input by each user (1) on the object corresponding to each user. Alternatively, the processor (130) can display a message corresponding to the voice of each user (1) input through a microphone on the object corresponding to each user. Alternatively, the processor (130) can change the expression of a graphic object to correspond to the expression of a user (1) identified through a plurality of images. Information about other users' chat messages and voices, and facial expressions of other users can be included in other user information.
[0189] If the number of multiple users included in each group is greater than or equal to a preset number, the processor (130) may display objects corresponding to the multiple users on the display (120) based on a PIP (Picture in Picture) method, or may select only a preset number of users (1) among the multiple users based on reaction scores corresponding to the multiple users, and display only objects corresponding to the selected multiple users.
[0190] For example, in the case of an object corresponding to another user or an object corresponding to another user included in a group other than the group including the user (1), the processor (130) may not display the object according to the user (1) setting.
[0191] According to one embodiment of the present disclosure, other user information may include response score information of other users. The processor (130) may determine the size and location of objects corresponding to multiple users based on the response scores of multiple users included in each group.
[0192] The processor (130) can set the size of the object corresponding to the user (1) to be larger as the user's reaction score or the absolute value of the reaction score is higher.
[0193] Alternatively, the processor (130) may display multiple objects corresponding to multiple users included in each group in the order of highest reaction scores within the area corresponding to each group. Alternatively, the processor (130) may display an object corresponding to a user with the highest reaction score among multiple users included in a group in the center of the area corresponding to the group. The processor (130) may also display objects corresponding to the remaining users in the order of highest reaction scores, centered around the object corresponding to the user with the highest reaction score, and display them on the display (120).
[0194] Here, storing information about a neural network model may mean storing various information related to the operation of the neural network model, for example, information about at least one layer included in the neural network model, information about parameters, biases, etc. used in each of at least one layer, etc. However, it goes without saying that, depending on the implementation form of the processor (130), the information about the neural network model may be stored in the internal memory of the processor (130). For example, when the processor (130) is implemented with dedicated hardware, the information about the neural network model may also be stored in the internal memory of the processor (130).
[0195] FIG. 11 is a detailed block diagram of an electronic device (100) according to one embodiment of the present disclosure.
[0196] Referring to FIG. 11, the electronic device (100) includes a camera (110), a display (120), a communication interface (140), one or more sensors (150), a speaker (160), a user (1) interface (180), a memory (190), and a processor (130). Among the configurations illustrated in FIG. 11, a detailed description of configurations that overlap with those illustrated in FIG. 2 will be omitted.
[0197] The communication interface (140) can transmit and / or receive various types of content. For example, the processor (130) can receive content or other user information through the communication interface (140). The processor (130) can also transmit user information to another electronic device (100) through the communication interface (140). The processor (130) can also receive multiple images through the communication interface (140). That is, the processor (130) can acquire multiple images of the user (1) from an external camera positioned adjacent to the electronic device (100) rather than the camera (110).
[0198] The communication interface (140) can receive or transmit signals in a streaming or download manner from an external device (e.g., a user (1) terminal), an external storage medium (e.g., a USB memory), an external server (e.g., a web hard drive), etc. through a communication method such as AP-based Wi-Fi (Wi-Fi, Wireless LAN network), Bluetooth, Zigbee, wired / wireless LAN (Local Area Network), WAN (Wide Area Network), Ethernet, IEEE 1394, HDMI (High-Definition Multimedia Interface), USB (Universal Serial Bus), MHL (Mobile High-Definition Link), AES / EBU (Audio Engineering Society / European Broadcasting Union), optical, coaxial, etc.
[0199] The electronic device (100) may include one or more sensors (150). The one or more sensors (150) may include sensors that detect objects around the electronic device (100) (e.g., a lidar sensor, a ToF sensor, etc.). In addition, the one or more sensors (150) may further include at least one of a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor (e.g., an RGB (red, green, blue) sensor), a biometric sensor, a temperature / humidity sensor, an illuminance sensor, or a UV (ultra violet) sensor.
[0200] The speaker (160) can output an audio signal to the outside of the electronic device (100). The speaker (160) can output multimedia playback, recording playback, various notification sounds, voice messages, etc. The electronic device (100) may include an audio output device such as the speaker (160), but may also include an output device such as an audio output terminal. In particular, the speaker (160) can provide acquired information, information processed and produced based on acquired information, a response result to a user's (1) voice, and / or an operation result, etc. in voice form.
[0201] The microphone (170) can receive the voice of the user (1). In addition, the microphone (170) can receive various audio signals related to the user (1) viewing the content.
[0202] The user (1) interface (180) is a component used by the electronic device (100) to interact with the user (1), and the processor (130) can receive various information, such as control information of the electronic device (100), through the user (1) interface (180). In particular, a chat message regarding content can be received through the user (1) interface (180). The user (1) interface (180) may include at least one of a touch sensor, a motion sensor, a button, a jog dial, a switch, and a microphone, but is not limited thereto.
[0203] The memory (190) can store data required for various embodiments of the present disclosure. For example, image content displayed on the electronic device (100) can be stored in the memory (190).
[0204] The memory (190) may be implemented in the form of memory embedded in the electronic device (100) or may be implemented in the form of memory that can be attached or detached from the electronic device (100) depending on the purpose of data storage. For example, data for driving the electronic device (100) may be stored in a memory embedded in the electronic device (100), and data for expanding the functions of the electronic device (100) may be stored in a memory that can be attached or detached from the electronic device (100).
[0205] In the case of memory embedded in the electronic device (100), it may be implemented as at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)).
[0206] In the case of a memory that can be attached or detached to an electronic device (100), it can be implemented in the form of a memory card (e.g., CF (compact flash), SD (secure digital), Micro-SD (micro secure digital), Mini-SD (mini secure digital), xD (extreme digital), MMC (multi-media card), etc.), an external memory that can be connected to a USB port (e.g., USB memory), etc.
[0207] For example, the memory (190) may store information regarding a plurality of neural network (or artificial intelligence) models. For example, the memory (190) may store a neural network model trained to identify contextual information of content and a neural network model trained to identify a user's (1) response type (and a probability value corresponding to the response type) based on a plurality of images.
[0208] FIG. 12 is a sequence diagram illustrating an electronic device (100) implemented as a server device according to one embodiment of the present disclosure.
[0209] According to one embodiment of the present disclosure, an electronic device may be implemented as a server device. The server device may be a server device (e.g., an OTT platform server, etc.) that provides content to multiple external electronic devices.
[0210] Referring to FIG. 12, the electronic device provides content to an external electronic device (S1210). While FIG. 12 illustrates content being provided to a single external electronic device, this is merely for convenience of explanation of the invention; the electronic device may provide content to multiple external electronic devices via the communication interface of the electronic device.
[0211] The electronic device (100) can receive an image of a user viewing content via an external electronic device (S1320). In particular, the electronic device (100) can also receive, in addition to the image, user voice information, user chat messages, etc. from the external electronic device.
[0212] The electronic device can identify contextual information of content (S1330). That is, the electronic device can identify contextual information of content provided to an external electronic device in real time.
[0213] The electronic device can identify a type of response according to the user's situation within the content based on the received image and identified context information, and identify a first group including the user among a plurality of groups corresponding to the content projection (S1240).
[0214] The electronic device can transmit user information to an external electronic device (S1250). Here, the user information may include first group information to which the user belongs, and may also include the user's response score information. The processor may also transmit user information from other external electronic devices to the external electronic device. In particular, the electronic device may transmit group information, each of which includes multiple users viewing content provided by the electronic device, to each external electronic device.
[0215] FIG. 13 is a flowchart schematically showing a method for controlling an electronic device (100) according to one embodiment of the present disclosure.
[0216] Referring to FIG. 13, at least one processor (130) may control the display (120) to obtain an image including a user through a camera (110) while content is being output through the display (120) (S1405), identify a first group corresponding to the user among a plurality of groups corresponding to the content based on context information and images corresponding to the content (S1410), and output a first graphic object corresponding to the user to a first area of the display (120) corresponding to the first group (S1415).
[0217] In the above description and the description below, the operations performed by at least one processor (130) may be described as being performed by the electronic device (100).
[0218] At least one processor (130) can output (or display) content through a display (120). At least one processor (130) can acquire an image through a camera (110). The image can be described as a captured image or a user image. At least one processor (130) can acquire at least one image.
[0219] Content can be described as a content image, and images acquired through a camera (110) can be described as a captured image.
[0220] For example, at least one processor (130) can acquire one image corresponding to a specific point in time.
[0221] For example, at least one processor (130) can acquire multiple images corresponding to a specific period of time.
[0222] The image may include a captured user. At least one processor (130) may analyze the user's response based on the user (or user object) included in the image. At least one processor (130) may identify a group that the current user prefers (or represents) based on the user's response.
[0223] At least one processor (130) can identify at least one group based on the content being output. For example, if the content is a soccer match, at least one processor (130) can analyze the content to identify a group corresponding to Team A and a group corresponding to Team B. At least one processor (130) can utilize an OCR function or an ACR function to identify the groups corresponding to the content.
[0224] At least one processor (130) can analyze user responses based on content and images at a specific point in time.
[0225] At least one processor (130) can obtain contextual information corresponding to content at a specific point in time. A detailed description of the contextual information is provided in FIG. 3. At least one processor (130) can obtain a user response corresponding to an image at a specific point in time. At least one processor (130) can compare the contextual information with the user response at a specific point in time. At least one processor (130) can identify a first group corresponding to the user based on the comparison result.
[0226] At least one processor (130) may obtain user response information including a user response score corresponding to each of a plurality of groups based on context information and an image, and identify a first group based on the user response information. The user response information may include a user response score corresponding to each of the plurality of groups.
[0227] For example, if there are multiple groups, the user response score may include multiple user response scores. The user response information may include a first score corresponding to the first group, a second score corresponding to the second group, and a third score corresponding to the third group.
[0228] For example, if the analysis target period comprises multiple periods, at least one processor (130) may obtain a sub-score corresponding to each period. At least one processor (130) may obtain an overall user response score by summing the sub-scores corresponding to each period.
[0229] User response information can be described as user response data, a user response table, a group-response mapping table, a user response score table, etc.
[0230] At least one processor (130) can analyze the content to identify at least one of a first group, a second group, and a third group. The first group may refer to a group associated with the first team recognized in the content. The second group may refer to a group associated with the second team recognized in the content. The third group may refer to a group not recognized as a team in the content.
[0231] At least one processor (130) can identify which group (or team) the user prefers (or cheers for) based on contextual information and images. To identify the group (or preferred group) corresponding to the user, the at least one processor (130) can obtain (or calculate) a user response score for each group. The at least one processor (130) can obtain at least one of a user response score corresponding to a first group, a user response score corresponding to a second group, or a user response score corresponding to a third group. The process of calculating the user response score is described in FIGS. 5 to 8 .
[0232] At least one processor (130) can obtain context information including a context type and a context-corresponding probability value related to content, obtain a response type corresponding to the context type based on an image, and obtain user response information based on the context-corresponding probability value and the response type.
[0233] Reaction types can represent criteria for categorizing user responses into preset categories. Reaction types can include categories indicating whether the user's current response to viewing content is positive, negative, or neutral. Reaction types can be described as response type information.
[0234] The number and definition of response types can be changed depending on the user's settings.
[0235] For example, the number of response types may be two. The response types may include at least one of positive or negative.
[0236] For example, the number of response types may be three. The response types may include at least one of positive, negative, and neutral.
[0237] For example, the number of response types may be four. The response types may include at least one of strong positive, negative, and neutral.
[0238] At least one processor (130) can identify a group corresponding to the identified user response score as a first group when a user response score greater than or equal to a threshold value is identified among a plurality of user response scores included in the user response information.
[0239] At least one processor (130) may determine that a higher user response score indicates a more positive perception of contextual information identified from content. If the user response score is relatively high at a particular point in time, at least one processor (130) may determine that the identified contextual information is more preferred at that particular point in time.
[0240] If a specific user response score is greater than or equal to a threshold value, at least one processor (130) can identify a group corresponding to the specific user response score as a first group.
[0241] It further includes a memory (190) that stores history information including at least one of a content viewing history and a group identification history, and at least one processor (130) can identify a group corresponding to a user among a plurality of groups based on the history information, and identify the first group by applying a first weight to a user response score corresponding to the identified group.
[0242] The history information may include at least one of a past viewing history and / or a past analysis history (e.g., results of a previous analysis such as content analysis). At least one processor (130) may store a viewing history of content viewed by the user. The viewing history of the content may include information about a specific group identified in the content. The information about the specific group may be obtained through metadata of the content. The metadata may include Electronic Program Guide (EPG) data of the content. At least one processor (130) may identify (or analyze) whether a user prefers a certain group through EPG data of the content viewed by the user.
[0243] At least one processor (130) may store an analysis history including the results of identifying a group corresponding to a user. When a group corresponding to a user is identified based on content and images, at least one processor (130) may store the identification results as an analysis history.
[0244] At least one processor (130) may store history information including at least one of a viewing history and an analysis history in the memory (190). The at least one processor (130) may use the history information to identify a user's preferred group. Once the user's preferred group is identified, the at least one processor (130) may apply (or assign) a first weight to the user's preferred group. The at least one processor (130) may use the first weight to obtain (or calculate) a higher user response score for the user's preferred group.
[0245] A description related to the history information is described in Fig. 6.
[0246] The electronic device further includes a communication interface, and at least one processor (130) can receive user reaction information of other users corresponding to the first group from a server device through the communication interface, and control the display (120) to output a second graphic object corresponding to the other users together with the first object to the first area based on the user reaction information.
[0247] The first object and the second object may include at least one of text information or GUI (Graphical User Interface) information representing the user. For example, the first object and the second object may include at least one of an avatar, an icon, and a profile image representing each user.
[0248] At least one processor (130) can identify a group preferred by a user of the electronic device (100) as a first group. At least one processor (130) can output a second object corresponding to another user who also prefers the first group preferred by the user, together with the first object. At least one processor (130) can obtain information related to the other user from a server device (or an external server).
[0249] At least one processor (130) can obtain other user response information corresponding to other users from the server device.
[0250] The other user response information may include other user response scores corresponding to the first group. At least one processor (130) may identify a first size of the first object, a first location of the first object, a second size of the second object, and a second location of the second object based on the user response scores and the other user response scores.
[0251] At least one processor (130) can control the display (120) to output a first object in a first area at a first location and a first size, and can control the display (120) to output a second object in a second area at a second location and a second size.
[0252] At least one processor (130) can identify another user who prefers the same group as the user of the electronic device (100). At least one processor (130) can output a first object corresponding to the user and a second object corresponding to the other user to the first area.
[0253] At least one processor (130) can determine at least one of the size of the object or the location of the object based on the user response score.
[0254] For example, the larger the user response score, the larger the object size can be.
[0255] For example, the higher the user response score, the more likely the object is to be positioned to the left (or upper). As the user response score increases, at least one processor (130) may assign priority in a specific direction to the object corresponding to the user response score. The criteria for a specific direction may be changed based on user settings.
[0256] The user of the electronic device (100) may be described as a first user, and another user may be described as a second user.
[0257] The behavior of displaying objects in relation to users and other users is described in FIGS. 9 and 10.
[0258] The electronic device further includes a microphone (170), and at least one processor (130) can obtain an audio signal including a user voice through the microphone (170) while displaying content through the display (120), obtain first user reaction information based on an image, obtain second user reaction information based on the audio signal, and identify a first group based on the first user reaction information and the second user reaction information.
[0259] At least one processor (130) can identify a first group by adding up the first user response score included in the first user response information and the second user response score included in the second user response information for each group.
[0260] At least one processor (130) may utilize both image and audio signals to analyze the user's response. An example of this is described in FIG. 7.
[0261] At least one processor (130) can obtain a first response type at a preset time based on an image, obtain a second response type at a preset time based on an audio signal, and if the first response type corresponds to the second response type, apply a second weight to a user response score corresponding to the preset time to identify a first group.
[0262] At least one processor (130) can identify whether a first response type obtained from an image at a preset time is the same as a second response type obtained from an audio signal at a preset time. If the first response type and the second response type match, the at least one processor (130) can apply (or assign) a second weight to the user response score obtained at the preset time. The at least one processor (130) can use the second weight so that a higher user response score is obtained (or calculated) for the time when the first response type and the second response type match.
[0263] The methods according to various embodiments of the present disclosure described above may be implemented in the form of applications that can be installed on existing electronic devices. Alternatively, the methods according to various embodiments of the present disclosure described above may be performed using a deep learning-based trained neural network (or a deep learned neural network), i.e., a learning network model. The methods according to various embodiments of the present disclosure described above may be implemented only through a software upgrade or a hardware upgrade of an existing electronic device. The various embodiments of the present disclosure described above may also be performed through an embedded server provided in an electronic device or an external server of the electronic device.
[0264] According to an exemplary embodiment of the present disclosure, the various embodiments described above may be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device may include an electronic device (e.g., an electronic device (100)) according to the disclosed embodiments, which is a device that can call instructions stored in the storage medium and operate according to the called instructions. When an instruction is executed by a processor, the processor may directly or under the control of the processor perform a function corresponding to the instruction using other components. The instruction may include code generated or executed by a compiler or an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.
[0265] According to one embodiment, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0266] Each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the respective components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0267] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made within the scope of the claims.
Claims
1. In electronic devices, camera; display; and comprising at least one processor; At least one processor of the above, An image including a user is acquired through the camera while content is output through the display, Identifying a first group including the user among a plurality of groups corresponding to the content based on context information corresponding to the content and the image, An electronic device that controls the display to output a first graphic object corresponding to the user to a first area of the display corresponding to the first group.
2. In paragraph 1, At least one processor of the above, Obtain user response information including user response scores corresponding to each of the plurality of groups based on the context information and the image, An electronic device that identifies the first group based on the user response information.
3. In paragraph 2, At least one processor of the above, Obtaining the context information including the context type and context response probability value related to the above content, Obtain a response type corresponding to the context type based on the image above, An electronic device that obtains user response information based on the context response probability value and the response type.
4. In paragraph 2, At least one processor of the above, An electronic device that identifies a group corresponding to the identified user response score as the first group when a user response score greater than or equal to a threshold value is identified among a plurality of user response scores included in the user response information.
5. In paragraph 2, Further comprising a memory storing history information including at least one of content viewing history or group identification history; At least one processor of the above, Identifying a group corresponding to the user among the plurality of groups based on the above history information, An electronic device that identifies the first group by applying a first weight to a user response score corresponding to the identified group.
6. In paragraph 1, further comprising a communication interface; At least one processor of the above, Receive other user response information of other users corresponding to the first group from the server device through the above communication interface, An electronic device that controls the display to output a second graphic object corresponding to the other user together with the first graphic object in the first area based on the other user response information.
7. In paragraph 6, The above user response information is: Including other user response scores corresponding to the first group above, At least one processor of the above, Identifying a first size of the first graphic object, a first position of the first graphic object, a second size of the second graphic object, and a second position of the second graphic object based on the user reaction score and the other user reaction score; Controlling the display to output the first graphic object in the first area at the first location and with the first size; An electronic device that controls the display to output the second graphic object in the first area at the second location and with the second size.
8. In paragraph 1, Including Mike; At least one processor of the above, While displaying the content through the display, an audio signal including a user's voice is acquired through the microphone, Obtain first user response information based on the above image, Obtain second user response information based on the above audio signal, An electronic device that identifies the first group based on the first user response information and the second user response information.
9. In paragraph 8, At least one processor of the above, An electronic device that identifies the first group by adding the first user response score included in the first user response information and the second user response score included in the second user response information for each group.
10. In paragraph 8, At least one processor of the above, Obtain the first reaction type at a preset time based on the above image, Obtaining a second response type at the preset time based on the above audio signal, An electronic device that identifies the first group by applying a second weight to the user response score corresponding to the preset time, if the first response type corresponds to the second response type.
11. A method for controlling an electronic device including a display, A step of acquiring an image including a user while content is output through the above display; A step of identifying a first group including the user among a plurality of groups corresponding to the content based on context information corresponding to the content and the image; and A control method, comprising the step of outputting a first graphic object corresponding to the user to a first area of the display corresponding to the first group.
12. In paragraph 11, The step of identifying the first group is: Obtain user response information including user response scores corresponding to each of the plurality of groups based on the context information and the image, A control method for identifying the first group based on the user response information.
13. In paragraph 12, The above control method is, A step of obtaining the context information including a context type and a context corresponding probability value related to the content; A step of obtaining a response type corresponding to the context type based on the image; and A control method, comprising: a step of obtaining user response information based on the context response probability value and the response type.
14. In paragraph 12, The step of identifying the first group is: A control method for identifying a user response score greater than or equal to a threshold value among a plurality of user response scores included in the user response information, wherein a group corresponding to the identified user response score is identified as the first group.
15. In paragraph 12, The above electronic device, Stores history information including at least one of content viewing history or group identification history, The step of identifying the first group is: Identifying a group corresponding to the user among the plurality of groups based on the above history information, A control method for identifying a first group by applying a first weight to a user response score corresponding to the identified group.
Citation Information
Patent Citations
System, server, and program
JP2022034965A
Method and apparatus for real-time viewer interaction with a media presentation
KR101283520B1
Display apparatus, server and control method thereof
KR102112743B1
Operating Method Of Terminal And Server To Recommend Contents
KR102129160B1
Consumption of content with reactions of an individual
KR102180716B1