Electronic device and control method therefor

The electronic device uses AI models to track and display the location of selected objects across frames, addressing the challenge of object visibility outside the frame, ensuring accurate and real-time representation.

WO2026155547A1PCT designated stage Publication Date: 2026-07-23SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2026-01-14
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Users face difficulty in identifying the location of a selected object within content when it is not included in the output frame, especially in real-time applications utilizing AI for dynamic object recognition.

Method used

An electronic device employs a first and second artificial intelligence model to determine the location of a preferred object across multiple frames, utilizing camera field of view information and frame rate to estimate and display the object's location even when it is not present in the current frame, using a display interface to indicate the estimated position.

Benefits of technology

Effectively tracks and displays the location of selected objects across frames, ensuring accurate and real-time representation of the object's position even when it is temporarily out of the camera's view, enhancing user interaction with dynamic content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2026000840_23072026_PF_FP_ABST
    Figure KR2026000840_23072026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. When instructions are individually or collectively executed, one or more processors cause the electronic device to: if a preferred object is identified among a plurality of objects included in content, input, to a trained first artificial intelligence model, a plurality of frames for the content so as to identify a plurality of object positions for each of the plurality of frames; if the preferred object is not included in a first frame output on the display among the plurality of frames, input, to a trained second artificial intelligence model, a position of the preferred object for each of the plurality of frames and camera viewing angle information for each of the plurality of frames so as to obtain position information of the preferred object that is estimated on the basis of the first frame; and on the basis of the position information of the preferred object, display, on the display, a graphic user interface (GUI) indicating an estimated position of the preferred object in the first frame.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method of controlling the same

[0001] The present disclosure relates to an electronic device and a method for controlling the same.

[0002] With the recent advancement of AI technology, technologies are being developed that utilize AI to recognize dynamic objects within content in real time and identify the location of those objects.

[0003] Generally, users can select one of multiple objects within the content via a separate controller and display the location of the selected object in real-time on the output screen. However, if the object selected by the user is not included in the output frame of the content being viewed, the user had difficulty identifying the location of the selected object without a separate controller.

[0004] An electronic device according to one or more embodiments of the present disclosure comprises a display, a memory for storing instructions, and one or more processors including processing circuitry.

[0005] According to one or more embodiments, when the instructions are executed individually or collectively, the electronic device obtains a location corresponding to a preferred object among a plurality of objects included in the content for each of the plurality of frames through a first artificial intelligence model based on a plurality of frames corresponding to the content, and if the preferred object is not included in the first frame output to the display among the plurality of frames, obtains location information of the preferred object through a second artificial intelligence model based on the location corresponding to the preferred object for each of the plurality of frames and camera field of view information for each of the plurality of frames, and controls the display to output a user interface (UI) including an indicator corresponding to the estimated location of the preferred object in the first frame based on the location information of the preferred object obtained through the second artificial intelligence model.

[0006] According to one or more embodiments, the location information of the preferred object is a first location information, and when the instructions are executed individually or collectively by the one or more processors, the electronic device acquires first field of view information of the camera corresponding to the first frame, acquires second location information of the preferred object included in a second frame prior to the first frame and second field of view information of the camera corresponding to the second frame, acquires third location information of the preferred object included in a third frame between the first frame and the second frame and third field of view information of the camera corresponding to the third frame, and acquires the first location information of the preferred object corresponding to the first frame through the second artificial intelligence model based on the first field of view information of the camera, the second location information of the preferred object, the second field of view information of the camera, the third location information of the preferred object, and the third field of view information of the camera.

[0007] According to one or more embodiments, when the instructions are executed individually or collectively by the one or more processors, the electronic device obtains first location information of the preferred object corresponding to the first frame through the learned second artificial intelligence model, wherein if the preferred object is not included in the first frame output to the display among the plurality of frames, the location corresponding to the preferred object for each of the plurality of frames, camera field of view information for each of the plurality of frames, and frame rate information of the plurality of frames.

[0008] According to one or more embodiments, when the instructions are executed individually or collectively by the one or more processors, the electronic device controls the display to output the plurality of object positions and the camera's field of view to a top view image corresponding to the entire space of the content, based on the plurality of frame-by-frame object positions and the camera's field of view information.

[0009] According to one or more embodiments, when the instructions are executed individually or collectively by the one or more processors, the electronic device controls the display to output the plurality of frame-by-frame top-view images, including the plurality of object locations and the field of view of the camera, in a Picture-in-Picture (PIP) format to each of the plurality of frames.

[0010] According to one or more embodiments, when the instructions are executed individually or collectively by the one or more processors, the electronic device controls the display to output a UI related to a location corresponding to the preferred object based on a plurality of object locations included in the first frame to one of the plurality of edge regions of the first frame.

[0011] According to one or more embodiments, the electronic device further includes a communication circuit, and when the instructions are executed individually or collectively by the one or more processors, the electronic device receives data corresponding to a plurality of captured images of the content acquired at a plurality of different shooting angles through the communication circuit, and obtains a position corresponding to the preferred object in the first frame based on the data corresponding to the plurality of captured images.

[0012] According to one or more embodiments, when the instructions are executed individually or collectively by the one or more processors, the electronic device controls the display to output a UI corresponding to the selection of a preferred object among the plurality of objects included in the content, and identifies the selected object as the preferred object through the UI corresponding to the selection of the preferred object.

[0013] According to one or more embodiments, when the instructions are executed individually or collectively by the one or more processors, the electronic device obtains camera field of view information for each of the plurality of frames based on information corresponding to a first object included in the entire space corresponding to the content and information corresponding to a second object included in each of the plurality of frames.

[0014] According to one or more embodiments, when the instructions are executed individually or collectively by the one or more processors, the electronic device obtains first location information of the preferred object corresponding to the first frame when receiving an image corresponding to content in which the location of the preferred object is indicated for each of the plurality of frames from a server through the communication circuit, and controls the display to output a user interface (UI) including an indicator corresponding to the estimated location of the preferred object in the first frame based on the first location information of the preferred object.

[0015] A control method for an electronic device according to one or more embodiments of the present disclosure comprises: an operation of obtaining a location corresponding to a preferred object among a plurality of objects included in the content for each of the plurality of frames through a first artificial intelligence model based on a plurality of frames corresponding to the content; an operation of obtaining location information of the preferred object through a second artificial intelligence model based on the location corresponding to the preferred object for each of the plurality of frames and camera field of view information for each of the plurality of frames if the preferred object is not included in the first frame among the plurality of frames; and an operation of outputting a user interface (UI) including an indicator corresponding to the estimated location of the preferred object in the first frame based on the location information of the preferred object obtained through the second artificial intelligence model.

[0016] A non-transient computer-readable storage medium storing computer instructions that cause the electronic device to perform an operation when executed by a processor of an electronic device according to one or more embodiments of the present disclosure, wherein the operation comprises: an operation of obtaining a location corresponding to a preferred object among a plurality of objects included in the content for each of the plurality of frames through a first artificial intelligence model based on a plurality of frames corresponding to the content; an operation of obtaining location information of the preferred object through a second artificial intelligence model based on the location corresponding to the preferred object for each of the plurality of frames and camera field of view information for each of the plurality of frames if the preferred object is not included in the first frame among the plurality of frames; and an operation of outputting a user interface (UI) including an indicator corresponding to the estimated location of the preferred object in the first frame based on the location information of the preferred object obtained through the second artificial intelligence model.

[0017] FIG. 1 is a drawing for explaining the operation of an electronic device according to one or more embodiments.

[0018] FIG. 2 is a block diagram illustrating the configuration of an electronic device according to one or more embodiments.

[0019] FIG. 3 is a block diagram illustrating the detailed configuration of an electronic device according to one or more embodiments.

[0020] FIG. 4 is a diagram illustrating the process of identifying preferred objects of an electronic device according to one or more embodiments.

[0021] FIG. 5 is a diagram illustrating a process for identifying the locations of multiple objects of an electronic device according to one or more embodiments.

[0022] FIG. 6 is a diagram illustrating the process of acquiring a top-view image of an electronic device according to one or more embodiments.

[0023] FIGS. 7 and FIGS. 8 are drawings for explaining the process of estimating the location of a preferred object of an electronic device according to one or more embodiments.

[0024] FIG. 9 is a diagram illustrating the PIP mode display process of an electronic device according to one or more embodiments.

[0025] FIGS. 10a, FIGS. 10b, and FIGS. 11 are drawings for explaining a GUI display process corresponding to the location of a preferred object of an electronic device according to one or more embodiments.

[0026] FIG. 12 is a diagram illustrating a process for estimating the location of a preferred object based on images captured from various angles of an electronic device according to one or more embodiments.

[0027] FIGS. 13 and FIGS. 14 are drawings for explaining the overall operation process of an electronic device according to one or more embodiments.

[0028] FIG. 15 is a drawing for explaining a method of operation of an electronic device according to one or more embodiments.

[0029] The terms used in the various embodiments of this Disclosure have been selected to be as widely used and general as possible, taking into account their functions within this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the relevant description section of this disclosure. Therefore, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the overall content of this disclosure.

[0030] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.

[0031] The expression "at least one of A or / and B" should be understood as representing either "A" or "B" or "A and B".

[0032] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.

[0033] Where it is stated that a component (e.g., Component 1) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., Component 2), it should be understood that the component may be directly connected to the other component or connected through the other component (e.g., Component 3).

[0034] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this disclosure, terms such as “comprising” or “consisting of” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0035] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of "modules" or a plurality of "parts" may be integrated into at least one module and implemented by at least one processor (not shown), except for a "module" or "part" that needs to be implemented in specific hardware.

[0036] In the present disclosure, the term "user" may refer to a person using an electronic device or a device used by such person.

[0037] An embodiment of the present disclosure will be described in more detail below with reference to the attached drawings.

[0038] FIG. 1 is a drawing for explaining the operation of an electronic device according to one or more embodiments.

[0039] According to one embodiment, the electronic device (100) can display the estimated location of a preferred object preferred by the user through a display (110). Here, the electronic device (100) can be implemented as various types of electronic devices such as a smart TV, digital signage, a monitor, a kiosk, a tablet PC, an electronic photo frame, a mobile phone, a large format display (LFD), a digital information display (DID), a video wall, a projector display, etc. However, in some cases, it may be implemented as an image processing device (e.g., a set-top box, one connected box) that is connected to the electronic device to provide images.

[0040] According to one embodiment, the electronic device (100) may include a display. Specifically, the electronic device (100) may directly display an acquired image or content on the display.

[0041] According to one embodiment, the electronic device (100) may not include a display. The electronic device (100) may be connected to an external display device and may transmit an image or content stored in the electronic device (100) to the external display device.

[0042] The electronic device (100) can transmit an image or content to an external display device along with a control signal for controlling the display of the image or content on the external display device. Here, the external display device may be connected to the electronic device (100) via a communication circuit (110) or an input / output interface (190). For example, the electronic device (100) may not include a display, such as a Set Top Box (STB).

[0043] According to one example, the electronic device (100) may include only a small display capable of displaying simple information such as text information. The electronic device (100) may transmit an image or content to an external display device via a wired or wireless connection through a communication circuit (110) or to an external display device via an input / output interface (190).

[0044] According to one embodiment, the electronic device (100) can receive information from a user that corresponds to a preferred object preferred by the user.

[0045] The preferred object may include an object selected by the user based on user input. For example, if the user selects player number 3 in a soccer match, the preferred object may be a person object corresponding to player number 3. However, the preferred object is not limited thereto and may be referred to as a selected object, main object, target object, or designated object, but in this disclosure, it will be collectively referred to as a preferred object.

[0046] According to one embodiment, the electronic device (100) can identify the location of each of a plurality of objects included in the content displayed through the display (110). The electronic device (100) can identify the location of a plurality of objects included in the content by inputting a frame corresponding to the content into an artificial intelligence model. The electronic device (100) can identify the location of a plurality of objects in each of a plurality of frames for the content.

[0047] According to one embodiment, the electronic device (100) can identify information corresponding to the estimated location of a preferred object when the first frame among a plurality of frames does not contain a preferred object preferred by the user. The electronic device (100) inputs a plurality of frames into an artificial intelligence model to obtain information corresponding to the location of the preferred object, and can estimate the approximate location of the preferred object based on the obtained information.

[0048] According to one embodiment, if the electronic device (100) includes a preferred object preferred by a user in the first frame, it may display a graphic user interface (GUI) for indicating the location of the preferred object through a display (110). For example, the electronic device (100) may display a highlight indication for indicating the location of the preferred object at the bottom of the preferred object through the display (110).

[0049] Referring to FIG. 1, an electronic device (100) can receive information corresponding to a preferred object preferred by the user based on user input. The electronic device (100) can identify the location of an object (10) for each of the multiple frames by inputting multiple frames of content into an artificial intelligence model. When a preferred object (20) is identified in the first frame among the multiple frames, the electronic device (100) can display a GUI through a display (110) to indicate the location of the preferred object.

[0050] Hereinafter, with reference to the drawings, various embodiments will be described in which an electronic device (100) estimates the location of a preferred object when a preferred object is not identified in a first frame, and displays a GUI through a display (110) to indicate the estimated location of the preferred object.

[0051] FIG. 2 is a block diagram illustrating the configuration of an electronic device according to one or more embodiments.

[0052] According to FIG. 2, the electronic device (100) includes a display (110), a memory (120), and one or more processors (130). However, it is not limited thereto, and the electronic device (100) may be implemented with some components excluded or with other components included.

[0053] The display (110) is configured to display the estimated location of content and preferred objects, including a plurality of frames. The display (110) may be implemented as a display including a self-emissive element or as a display including a non-emissive element and a backlight. For example, it may be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an LED (Light Emitting Diodes), a micro LED, a Mini LED, a PDP (Plasma Display Panel), a QD (Quantum dot) display, a QLED (Quantum dot light-emitting diodes), etc. The display (110) may also include a driving circuit, a backlight unit, etc., which may be implemented in the form of an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc.

[0054] The memory (120) can store at least one instruction, data, program, etc. required for the operation of the electronic device (100). For example, the memory (120) can store contour highlighting processing information and location information corresponding to a selected image.

[0055] The memory (120) may be implemented in the form of a memory embedded in the electronic device (100) or in the form of a memory detachable from the electronic device (100), depending on the purpose of data storage. For example, data for operating the electronic device (100) may be stored in a memory embedded in the electronic device (100), and data for the expansion function of the electronic device (100) may be stored in a memory detachable from the electronic device (100).

[0056] In the case of memory embedded in the electronic device (100), it may be implemented as at least one of volatile memory (e.g., DRAM (dynamic RAM), SRAM (static RAM), or SDRAM (synchronous dynamic RAM), non-volatile memory (e.g., OTPROM (one time programmable ROM), PROM (programmable ROM), EPROM (erasable and programmable ROM), EEPROM (electrically erasable and programmable ROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD).

[0057] The memory (120) may be implemented as a single memory that stores data generated in various operations according to the present disclosure, but is not limited thereto, and the memory (120) may be implemented to include a plurality of memories that each store different types of data or each store data generated in different stages.

[0058] One or more processors (130) control the overall operation of the electronic device (100). Specifically, one or more processors (130) may be connected to each component of the electronic device (100) to control the overall operation of the electronic device (100). For example, one or more processors (130) may be electrically connected to the display (110) and memory (120) to control the overall operation of the electronic device (100). One or more processors (130) may include processing circuits and may be composed of one or more processors.

[0059] One or more processors (130) can perform the operation of an electronic device (100) according to various embodiments by executing one or more instructions stored in memory (120).

[0060] One or more processors (130) may include one or more of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), MIC (Many Integrated Core), DSP (Digital Signal Processor), NPU (Neural Processing Unit), hardware accelerator, or machine learning accelerator. One or more processors (130) may control one or any combination of other components of an electronic device and may perform operations or data processing related to communication. One or more processors (130) may execute one or more programs or instructions stored in memory. For example, one or more processors may perform a method according to one or more embodiments of the present disclosure by executing one or more instructions stored in memory.

[0061] When a method according to one or more embodiments of the present disclosure includes a plurality of operations, the plurality of operations may be performed by a single processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to one or more embodiments, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an artificial intelligence dedicated processor).

[0062] One or more processors (130) may be implemented as a single-core processor including one core, or as one or more multicore processors including multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors (130) are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory such as cache memory or on-chip memory, and a common cache shared by multiple cores may be included in the multicore processor. Additionally, each of the multiple cores included in the multicore processor (or some of the multiple cores) may independently read and execute program instructions for implementing a method according to one or more embodiments of the present disclosure, or all (or some) of the multiple cores may be linked together to read and execute program instructions for implementing a method according to one or more embodiments of the present disclosure.

[0063] When a method according to one or more embodiments of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one of the plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to one or more embodiments, the first operation, the second operation, and the third operation may all be performed by a first core included in a multi-core processor, or the first operation and the second operation may be performed by a first core included in a multi-core processor and the third operation may be performed by a second core included in a multi-core processor.

[0064] In the embodiments of the present disclosure, a processor may refer to a system-on-chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, GPU, APU, MIC, DSP, NPU, hardware accelerator, or machine learning accelerator, but the embodiments of the present disclosure are not limited thereto. For convenience of explanation, one or more processors (130) will be referred to as processors (130) below.

[0065] According to one embodiment, when a preferred object is identified among a plurality of objects included in the content, the processor (130) can input a plurality of frames of the content into a learned first artificial intelligence model to identify the location of a plurality of objects per plurality of frames.

[0066] According to one embodiment, the processor (130) can obtain a location corresponding to a preferred object among a plurality of objects included in the content for each frame through a first artificial intelligence model based on a plurality of frames corresponding to the content.

[0067] According to one embodiment, if the processor (130) does not include a preferred object in the first frame output to the display (110) among a plurality of frames, it can obtain first location information of the preferred object corresponding to the first frame through a second artificial intelligence model based on the location corresponding to the preferred object for each of the plurality of frames and camera field of view information for each of the plurality of frames.

[0068] Camera field of view information for multiple frames may be information regarding the spatial range that the camera lens can capture. Camera field of view information may include different information depending on the camera's position and focal length. For example, when photographing a specific space, the further the camera is located from the space and the larger the focal length, the wider the spatial range the camera can capture.

[0069] The first location information of the preferred object may include specific location information of the preferred object in a frame containing the preferred object. For example, the first location information may be approximate location information of the preferred object, or information corresponding to the coordinate values ​​of the preferred object in a frame.

[0070] According to one embodiment, the processor (130) can display a graphic user interface (GUI) through the display (110) to indicate the estimated location of a preferred object in a first frame based on first location information of the preferred object.

[0071] FIG. 3 is a block diagram illustrating the detailed configuration of an electronic device according to one or more embodiments.

[0072] According to FIG. 3, the electronic device (100) includes a display (110), memory (120), one or more processors (130), a communication circuit (140), a microphone (150), a speaker (160), and an input / output interface (170). A detailed description of the configurations shown in FIG. 3 that overlap with the configuration shown in FIG. 2 will be omitted.

[0073] The communication circuit (140) may include wired or wireless input / output interfaces (or input / output terminals) according to various standards. The communication circuit (140) may be configured to communicate with various types of external devices according to various types of communication methods. The communication circuit (140) may include a wireless communication module or a wired communication module. Here, each communication module may be implemented in the form of at least one hardware chip.

[0074] The communication circuit (140) may include various interfaces such as HDMI (High Definition Multimedia Interface), MHL (Mobile High-Definition Link), USB (Universal Serial Bus), DP (Display Port), Thunderbolt, VGA (Video Graphics Array) port, RGB port, D-SUB (D-subminiature), DVI (Digital Visual Interface), Bluetooth, Zigbee, wired / wireless LAN (Local Area Network), WAN (Wide Area Network), Ethernet, IEEE 1394, AES / EBU (Audio Engineering Society / European Broadcasting Union), Optical, Coaxial, etc.

[0075] The microphone (150) is configured to receive user voice or other sounds and convert them into audio data. The processor (130) can identify a preferred object based on the user voice signal received through the microphone (150).

[0076] The speaker (160) can convert and amplify a digital audio signal processed by the processor (130) into an analog audio signal and output it. For example, the speaker (160) may include at least one speaker unit, a D / A converter, an audio amplifier, etc., capable of outputting at least one channel. For example, the speaker (160) may output information corresponding to the caller and the caller's intention regarding a received call.

[0077] The input / output interface (170) may be any one of the following interfaces: HDMI (High Definition Multimedia Interface), MHL (Mobile High-Definition Link), USB (Universal Serial Bus), DP (Display Port), Thunderbolt, VGA (Video Graphics Array) port, RGB port, D-SUB (D-subminiature), DVI (Digital Visual Interface).

[0078] The input / output interface (170) can input and output at least one of audio and video signals. Depending on the implementation example, the input / output interface (170) may include separate ports for inputting and outputting only audio signals and for inputting and outputting only video signals, or it may be implemented as a single port for inputting and outputting both audio and video signals.

[0079] The electronic device (100) can transmit at least one of audio and video signals to an external device (e.g., an external display device or an external speaker) through an input / output interface (170). Specifically, an output port included in the input / output interface (170) may be connected to an external device, and the electronic device (100) can transmit at least one of audio and video signals to the external device through the output port.

[0080] The input / output interface (170) can be connected to the communication circuit (140). The input / output interface (170) can transmit information received from an external device to the communication circuit or transmit information received through the communication interface to the external device.

[0081] FIG. 4 is a diagram illustrating the process of identifying preferred objects of an electronic device according to one or more embodiments.

[0082] According to one embodiment, the electronic device (100) may display a guide UI through a display for selecting a preferred object among a plurality of objects included in the content. The guide UI may be a UI that displays a plurality of person objects included in the content in a list form. For example, the guide UI may include a UI that displays at least one of a person object, a name, and a number at a position by position.

[0083] According to one embodiment, the electronic device (100) can identify an object selected through a guide UI as a preferred object. The electronic device (100) can receive user input for selecting one of a plurality of person objects through a communication circuit (140). For example, the user may select one of a plurality of person objects via a remote control or may select one of a plurality of person objects by touching the display (110).

[0084] According to one embodiment, when a preferred object selected by a user among a plurality of person objects is identified, the electronic device (100) may store information corresponding to the preferred object in memory (120). The information corresponding to the preferred object may include at least one of the face image, number, and name of the preferred object.

[0085] Referring to FIG. 4, when a soccer match begins, the electronic device (100) can receive information corresponding to the soccer match from a server. For example, the electronic device (100) can receive from the server information about the team playing the soccer match, information about the players belonging to the team, information about the match tactics, and information about the players by position.

[0086] The electronic device (100) can display a guide UI (420) for selecting a preferred object among a plurality of person objects based on information received from a server through a display (110). When an object for any one of the plurality of objects is selected based on user input, the electronic device (100) can identify the selected object as a preferred object (410).

[0087] FIG. 5 is a diagram illustrating a process for identifying the locations of multiple objects of an electronic device according to one or more embodiments.

[0088] According to one embodiment, the electronic device (100) can identify multiple object locations per frame by inputting a representative frame among multiple frames of content into an artificial intelligence model.

[0089] According to one example, the electronic device (100) can identify an intra-frame (I-Frame), which is a representative frame, based on header information for each of a plurality of frames. An intra-frame is a frame that contains complete image information within the frame itself and may be an independently compressed frame.

[0090] Referring to FIG. 5, the electronic device (100) can identify a plurality of frames (510) containing motion information of an object. At this time, the electronic device (100) can identify an intra-frame (520) that represents the object while the object is moving. For example, if the object is moving only its arm, the electronic device (100) can identify the frame at the time the arm started moving as the intra-frame (510).

[0091] According to one example, the electronic device (100) can identify an intra-frame (520), which is a representative frame for each different frame interval, based on header information or motion information of a plurality of frames.

[0092] According to one embodiment, the electronic device (100) can input a plurality of intra-frames (530) into a first artificial intelligence model (540) to identify the locations of a plurality of objects for each of the plurality of frames. The first artificial intelligence model (540) may be a model trained to identify objects based on at least one of the image, motion information, and size of each of the plurality of objects included in the frame. That is, the artificial intelligence model may be a model trained to identify the object information and location of each dynamic object for each frame.

[0093] Here, the term "training an artificial intelligence model" means that a basic artificial intelligence model (e.g., an artificial intelligence model containing arbitrary random parameters) is trained by a learning algorithm using multiple training data, thereby creating predefined behavioral rules or an artificial intelligence model configured to perform a desired characteristic (or objective). This learning may be performed via a separate server and / or system, but is not limited thereto, and may also be performed in a cooking device. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples mentioned above.

[0094] Here, the artificial intelligence model may be implemented as, for example, CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), BRDNN (Bidirectional Recurrent Deep Neural Network), or Deep Q-Networks, but is not limited thereto.

[0095] Referring to FIG. 5, the electronic device (100) can input a plurality of intra-frames (530) into a first artificial intelligence model (540). The electronic device (100) can identify the location of each of a plurality of objects per frame through the first artificial intelligence model (540).

[0096] FIG. 6 is a diagram illustrating the process of acquiring a top-view image of an electronic device according to one or more embodiments.

[0097] According to one embodiment, the electronic device (100) may acquire a top view image corresponding to the entire space of the content. The top view image may be an image captured in a vertical direction of the entire space of the content, that is, an image captured from a viewpoint looking down from above. The top view image is not limited thereto and may be referred to as a bird's-eye view image or an overhead view image, but in this disclosure, it will be collectively referred to as a top view image.

[0098] According to one embodiment, the electronic device (100) can display a plurality of object positions and a camera's field of view in a top-view image through a display (110) based on information regarding a plurality of object positions per frame and a camera's field of view. The electronic device (100) can display only the range corresponding to the camera's field of view within the entire space of the content through a top-view image display (110).

[0099] According to one embodiment, the electronic device (100) can obtain information about a first object included in the entire space corresponding to the content, and obtain camera field of view information for each of the multiple frames by comparing the information about the first object with the information about a second object included in each of the multiple frames.

[0100] Information about the first object may be information about static objects in the content space. For example, if the content space is a soccer stadium, the first object may be the center circle, offside line, and goal line displayed on the soccer stadium.

[0101] Information about the second object may be information about static objects (e.g., center circle, offside line, goal line) within the content space included within the camera's field of view.

[0102] According to one example, the electronic device (100) can obtain frame-by-frame viewing angle information by comparing location information for a second object included in each frame with location information for a first object over the entire space of the content. Based on the frame-by-frame viewing angle information, the electronic device (100) can display a camera viewing range in a top-view image.

[0103] Referring to FIG. 6, the electronic device (100) can display a top-view image corresponding to the entire space of the content through the display (110). The electronic device (100) can display a top-view image (610) that displays only the camera field of view range based on the camera's field of view information per frame through the display (110).

[0104] The electronic device (100) can display a top view image (620) with multiple object locations displayed on a top view image (610) that displays a camera field of view through a display (110). The electronic device (100) can acquire a top view image with multiple object locations displayed on each of the multiple frames and store it in memory (120).

[0105] FIGS. 7 and FIGS. 8 are drawings for explaining the process of estimating the location of a preferred object of an electronic device according to one or more embodiments.

[0106] According to one embodiment, the electronic device (100) can acquire first field of view information of a camera corresponding to a first frame. The electronic device (100) can acquire a top-view image corresponding to a first frame.

[0107] According to one embodiment, the electronic device (100) can acquire second location information of a preferred object identified in a second frame prior to the first frame and second field of view information of a camera corresponding to the second frame. The electronic device (100) can acquire a top-view image corresponding to the second frame and display a top-view image (710) with the second location (720) of the preferred object displayed on the top-view image through a display (110).

[0108] According to one embodiment, the electronic device (100) can acquire third location information of a preferred object identified in a third frame between a first frame and a second frame, and third field of view information of a camera corresponding to the third frame. The electronic device (100) can acquire a top-view image corresponding to the third frame and display a top-view image (730) with the third location (740) of the preferred object displayed on the top-view image through a display (110).

[0109] According to one embodiment, the electronic device (100) may input first field of view information of a camera, second location information of a preferred object, second field of view information of a camera, third location information of a preferred object, and third field of view information of a camera into a trained second artificial intelligence model to obtain first location information of a preferred object estimated based on a first frame. The second artificial intelligence model may be a model trained to estimate the location of a preferred object in a frame at a current time based on a plurality of frame-by-frame camera field of view information and location information of a preferred object.

[0110] In FIGS. 7 and 8, the explanation will be based on the assumption that they are the second frame, the third frame, and the first frame in chronological order.

[0111] Referring to FIG. 7, the electronic device (100) can acquire a top-view image (710) corresponding to a second frame and a top-view image (730) corresponding to a third frame. In the top-view image (710) corresponding to the second frame, the second position (720) of the preferred object may be included within the camera's field of view. However, if the camera's field of view information is changed, that is, if the shooting range is changed to the right of the entire space of the content, the third position (740) of the preferred object in the top-view image (730) corresponding to the third frame may not be included within the camera's field of view.

[0112] Referring to FIG. 8, the electronic device (100) can input a top-view image (710) corresponding to a second frame and a top-view image (730) corresponding to a third frame into a second artificial intelligence model (810). The electronic device (100) can identify camera field of view information in the top-view image (820) corresponding to a first frame and identify the camera's field of view range for the entire space. The electronic device (100) can identify that a preferred object is not included within the camera's field of view range of the top-view image (820) corresponding to the first frame.

[0113] The electronic device (100) can estimate the position of the preferred object (830) in the top-view image (820) corresponding to the first frame through the second artificial intelligence model. When the camera's shooting range for the entire space changes from left to right, the preferred object may be included within the camera's field of view in the second frame, but may not be included within the camera's field of view in the first frame.

[0114] According to one embodiment, if the electronic device (100) does not include a preferred object in the first frame output to the display (110) among a plurality of frames, it can input the location of the preferred object for each of the plurality of frames, camera field of view information for each of the plurality of frames, and frame rate information of the plurality of frames into a learned second artificial intelligence model (810) to obtain first location information of the preferred object estimated based on the first frame.

[0115] According to one example, the electronic device (100) can estimate the location of a preferred object in a first frame based on the frame rate for a plurality of frames. For example, if the frame rate is 100Hz, the second frame, the third frame, and the first frame may be frames output at intervals of 1 / 100th of a second. In this case, if the preferred object is included within the camera field of view of the second frame and the preferred object is identified at the boundary outside the camera field of view of the third frame, the electronic device (100) can estimate that the preferred object in the first frame is located outside the camera field of view.

[0116] FIG. 9 is a diagram illustrating the PIP mode display process of an electronic device according to one or more embodiments.

[0117] According to one embodiment, the electronic device (100) can display a plurality of frame-by-frame top-view images including a plurality of object positions and a camera's field of view in a Picture in Picture (PIP) mode for each of the plurality of frames through a display (110).

[0118] According to one example, the electronic device (100) can control the display (110) to display a PIP image including a window of a preset size in a portion of the display (110).

[0119] PIP mode may be a video mode that provides the entire video and the video contained within a window of a preset size by displaying a video of a preset size in a part of the display screen through a window while displaying content in full screen. PIP mode may be a video mode that displays a video smaller in size than the video displayed in full screen in a part of the screen.

[0120] Referring to FIG. 9, the electronic device (100) can display a plurality of frame-by-frame top view images (910) in PIP mode in a portion of the display screen. For example, the electronic device (100) can display the top view images (910) in PIP mode through a window of a preset size at the bottom left of the entire image.

[0121] FIGS. 10a, FIGS. 10b, and FIGS. 11 are drawings for explaining a GUI display process corresponding to the location of a preferred object of an electronic device according to one or more embodiments.

[0122] According to one embodiment, the electronic device (100) can control the display (110) to output a UI related to a position corresponding to a preferred object based on a plurality of object positions included in the first frame to one of the plurality of edge regions of the first frame.

[0123] According to one example, the electronic device (100) may display a GUI corresponding to the location of a preferred object estimated based on the location of a plurality of objects included in the first frame through a display (110) in the left or right edge area of ​​the first frame.

[0124] According to one example, the electronic device (100) can control the display (110) to output a UI related to a position corresponding to a preferred object to at least one of the left / right and up / down regions of the first frame.

[0125] For example, the electronic device (100) can display a UI related to a location corresponding to a preferred object at at least one of the left and right corners or the top and bottom corners of the first frame.

[0126] Referring to FIG. 10, the electronic device (100) can identify the estimated location of a preferred object based on a top-view image (820) corresponding to a first frame. The electronic device (100) can identify whether the preferred object is located to the left or right outside the camera's field of view in the first frame. For example, if the preferred object is located to the left of the camera's field of view in the first frame, the electronic device (100) can display a GUI (1010) corresponding to the location of the preferred object in the left edge area of ​​the first frame through a display (110). The GUI illustrated in FIG. 10 is not limited thereto and can be displayed with various settings for size, length, and thickness.

[0127] Meanwhile, in 10a, the explanation was based on a soccer match case, but in 10b, the explanation will be based on a racing match case.

[0128] Referring to FIG. 10b, the electronic device (100) can identify the location of a preferred object among a plurality of objects included in a first frame. The electronic device (100) may also identify the location of the preferred object based on a top-view image corresponding to the first frame. The electronic device (100) may display a GUI corresponding to the location of the preferred object through a display (110).

[0129] For example, when a preferred object (e.g., a racing car preferred by the user) among a plurality of objects (e.g., a racing car) included in a first frame is identified by the electronic device (100), a GUI corresponding to the location of the preferred object can be displayed through the display (110).

[0130] For example, if the electronic device (100) does not identify a preferred object (e.g., a racing car preferred by the user) in the first frame, it can identify the location of the preferred object based on multiple frames or images taken from different directions.

[0131] In this case, the electronic device (100) can display a GUI corresponding to the estimated position of the preferred object in the first frame on the top / bottom and left / right sides of the display. For example, if the position of the preferred object is located behind the racing car displayed in the first frame, a GUI corresponding to the estimated position of the preferred object in the first frame can be displayed at the top of the display.

[0132] Meanwhile, the present disclosure is not limited to soccer game cases or racing game cases, and can be applied to content including multiple dynamic objects. For example, content including multiple dynamic objects may include sports game content such as basketball and baseball, and bicycle game content. Referring to FIG. 11, the electronic device (100) can display a GUI corresponding to the estimated location of a preferred object and a top-view image showing the location of the preferred object through a display (110). The electronic device (100) can display a top-view image (1110) showing the estimated location (1120) of the preferred object through a display (110) in PIP mode.

[0133] For example, the electronic device (100) can display a top-view image (1110) in which the estimated location (1120) of a preferred object in a star shape in the first frame as a window of a preset size at the top right of the display.

[0134] FIG. 12 is a diagram illustrating a process for estimating the location of a preferred object based on images captured from various angles of an electronic device according to one or more embodiments.

[0135] According to one embodiment, when the electronic device (100) receives a plurality of captured images of content captured at a plurality of different shooting angles through a communication circuit (140), it can identify the location of a preferred object in a first frame based on the plurality of captured images.

[0136] According to one example, when an electronic device (100) receives multiple captured images taken at different shooting angles with respect to the entire space of the content, it can identify the location of a preferred object in the entire space based on the multiple captured images.

[0137] Referring to FIG. 12, an electronic device (100) can receive a plurality of captured images (1210-1 to 1210-3) taken from various angles of the entire space through a communication circuit (140). The electronic device (100) can input the plurality of captured images (1210-1 to 1210-3) into a first artificial intelligence model to identify the locations of a plurality of objects included in each captured image (1210-1 to 1210-3). The electronic device (100) can identify the location of a preferred object from the plurality of captured images (1210-1 to 1210-3). Based on the plurality of captured images (1210-1 to 1210-3), the electronic device (100) can identify the location of a preferred object in the entire space of the content and estimate the location of the preferred object in the first frame. The electronic device (100) can display an image (1220) that displays a GUI corresponding to the position of a preferred object in the first frame through a display (110).

[0138] According to one embodiment, the electronic device (100) may receive a content video from a content streaming server in which the locations of a plurality of frame-by-frame preferred objects are displayed. Based on the content video received from the server, the electronic device (100) may display an image through a display (110) that displays a GUI corresponding to the estimated location of the preferred object in the first frame.

[0139] According to one embodiment, the electronic device (100) can estimate the location of a preferred object by analyzing audio data of the content. For example, the electronic device (100) can estimate the location of a preferred object based on the commentary relay of the content.

[0140] According to one embodiment, the electronic device (100) can control the projection unit (170) based on a user voice signal received through the microphone (150). For example, when a user voice signal for projecting an A UI is received, the electronic device (100) can control the projection unit (170) to display the A UI.

[0141] According to one embodiment, the electronic device (100) can control an external display device connected to the electronic device (100) based on a user voice signal received through a microphone (150). Specifically, the electronic device (100) can generate a control signal to control the external display device so that an operation corresponding to the user voice signal is performed on the external display device, and can transmit the generated control signal to the external display device. Here, the electronic device (100) can store a remote control application for controlling the external display device. The electronic device (100) can transmit the generated control signal to the external display device using at least one communication method among Bluetooth, Wi-Fi, or infrared. For example, when a user voice signal for displaying content A is received, the electronic device (100) can transmit a control signal to the external display device to control the display of content A on the external display device. Here, the electronic device (100) may refer to various terminal devices capable of installing a remote control application, such as a smartphone or an AI speaker.

[0142] According to one embodiment, the electronic device (100) may use a remote control device to control an external display device connected to the electronic device (100) based on a user voice signal received through a microphone (10). Specifically, the electronic device (100) may transmit a control signal to the remote control device to control the external display device so that an operation corresponding to the user voice signal is performed on the external display device. The remote control device may transmit the control signal received from the electronic device (100) to the external display device. For example, when a user voice signal for displaying content A is received, the electronic device (100) transmits a control signal to the remote control device to control the display of content A on the external display device, and the remote control device transmits the received control signal to the external display device.

[0143] According to one embodiment, the communication circuit (110) may use the same communication module (e.g., Wi-Fi module) to communicate with an external device, such as a remote control device, and an external server.

[0144] According to one embodiment, the communication circuit (110) may use different communication modules to communicate with external devices, such as a remote control device and an external server. For example, the communication circuit (110) may use at least one of an Ethernet module or a Wi-Fi module to communicate with an external server, and may use a Bluetooth module to communicate with an external device, such as a remote control device. However, this is merely one embodiment, and the communication circuit (110) may use at least one of various communication modules when communicating with multiple external devices or external servers.

[0145] According to one embodiment, the electronic device (100) can receive a user voice signal through a microphone (150) included in the electronic device (100).

[0146] According to one embodiment, the electronic device (100) may receive a user voice signal from an external device including a microphone. Here, the external device may refer to a remote control device or a smartphone, etc. Here, the received user voice signal may be a digital voice signal, but may be an analog voice signal depending on the implementation example. The electronic device (100) may receive the user voice signal through a wireless communication method such as Bluetooth or Wi-Fi.

[0147] According to one embodiment, the electronic device (100) can obtain text information corresponding to a user voice signal from an external server. Specifically, the electronic device (100) can transmit a user voice signal (audio signal or digital signal) to an external server. Here, the external server may refer to a speech recognition server. Here, the speech recognition server can convert the user voice signal into text information using STT (Speech To Text). Then, the external server can transmit the text information corresponding to the converted user voice signal to the electronic device (100).

[0148] According to one embodiment, the electronic device (100) can independently acquire text information corresponding to a user voice signal. Specifically, the electronic device (100) may directly apply a Speech To Text (STT) function to a digital voice signal to convert it into text information and transmit the converted text information to an external server.

[0149] According to one embodiment, an external server may transmit text information corresponding to a user voice signal to an electronic device (100). Specifically, the external server may be a server that performs a voice recognition function of converting a user voice signal into text information.

[0150] According to one embodiment, an external server may transmit at least one of text information corresponding to a user voice signal or search result information corresponding to text information to an electronic device (100). Specifically, the external server may be a server that performs a search result providing function that provides search result information corresponding to text information in addition to a voice recognition function that converts a user voice signal into text information.

[0151] For example, the external server may be a server that performs both voice recognition and search result provision functions. As another example, the external server may perform only voice recognition functions, while the search result provision function may be performed on a separate server. To obtain search results, the external server may transmit text information to a separate server and obtain search results corresponding to the text information from the separate server.

[0152] According to one embodiment, a communication module for communication with an external device and an external server can be implemented in the same way. For example, the electronic device (100) communicates with the external device using a Bluetooth module, and the external server can also communicate using a Bluetooth module.

[0153] According to one embodiment, a communication module for communication with an external device and an external server may be implemented separately. For example, the electronic device (100) may communicate with an external device using a Bluetooth module and with an external server using an Ethernet modem or a Wi-Fi module.

[0154] FIGS. 13 and FIGS. 14 are drawings for explaining the overall operation process of an electronic device according to one or more embodiments.

[0155] Referring to FIG. 13, an electronic device (100) can receive information corresponding to a preferred object among a plurality of objects through a user interface (1330). When a preferred object is identified, the electronic device (100) can receive a plurality of frames for the content through a RESTful API (1320). The RESTful API (1320) may be an interface for communication between the electronic device (100) and a server over a network. The electronic device (100) can transmit a plurality of frames for the content and information corresponding to the preferred object to a server (1340).

[0156] According to one example, the electronic device (100) can identify multiple objects for each of multiple frames through an object recognition model (1350). The electronic device (100) can estimate the location of a preferred object for each of multiple frames through a camera position estimation model (1360) and a preferred object position estimation model (1370).

[0157] According to one example, when the location of a preferred object is identified in the first frame, the electronic device (100) can display a GUI corresponding to the estimated location of the preferred object through the display (110) via the AR engine (1310).

[0158] Referring to FIG. 14, in operation 1410, the electronic device (100) can obtain data corresponding to a video source from content output through the display (110).

[0159] In operation 1420, the electronic device (100) can identify audio data and video data from a video source.

[0160] In operation 1430, the electronic device (100) can acquire a top-view image corresponding to each of a plurality of frames of content.

[0161] In operation 1440, the electronic device (100) can acquire a top-view image corresponding to the entire space of the content.

[0162] In operation 1450, the electronic device (100) can obtain a top view image in which the positions of a plurality of objects and a preferred object are displayed in a top view image corresponding to the entire space of the content.

[0163] In operation 1460, the electronic device (100) can identify the location of a preferred object from a plurality of frame-by-frame top-view images.

[0164] In operation 1470, the electronic device (100) can estimate the position of the preferred object in the first frame.

[0165] In operation 1480, the electronic device (100) can display a GUI corresponding to the estimated position of the preferred object in the first frame through the display (110).

[0166] FIG. 15 is a drawing for explaining a method of operation of an electronic device according to one or more embodiments.

[0167] Referring to FIG. 15, in operation 1510, when a preferred object among a plurality of objects included in the content is identified, the electronic device (100) can identify the location of a plurality of objects per plurality of frames by inputting a plurality of frames of the content into a learned first artificial intelligence model.

[0168] In operation 1520, if the electronic device (100) does not include a preferred object in the first frame among a plurality of frames, it can input the location of the preferred object for each of the plurality of frames and the camera field of view information for each of the plurality of frames into a learned second artificial intelligence model to obtain the first location information of the preferred object estimated based on the first frame.

[0169] In operation 1530, the electronic device (100) can display a GUI for indicating the estimated position of a preferred object in a first frame based on first position information of the preferred object.

[0170] Since the method for identifying a preferred object and the locations of multiple objects in multiple frames of content and obtaining the first location information of the preferred object has been specifically explained through the embodiments described above, a description thereof will be omitted.

[0171] The control method described in FIG. 15 can be performed by an electronic device (100) having the configuration of FIG. 2 described above, but is not necessarily limited thereto and can be performed by an electronic device having various configurations.

[0172] The various embodiments described above may be implemented as individual embodiments, or at least one embodiment may be combined with one another, either wholly or partially, to be implemented together in a single device.

[0173] According to the various embodiments described above, the electronic device (100) can provide the user with the estimated location of a preferred object by displaying the estimated location of the preferred object preferred by the user on a display while the user is viewing content.

[0174] Meanwhile, the various embodiments described above may be applied to a product as embodiments alone, but at least some of their contents may be combined with other embodiments of the present disclosure to be implemented together.

[0175] The various embodiments described above may be implemented as software containing instructions stored on a machine-readable storage medium (e.g., computer). The machine may include an electronic device (e.g., electronic device (100)) according to the disclosed embodiments, which is a device capable of calling instructions stored from the storage medium and operating according to the called instructions. When instructions are executed by a processor, the processor may perform a function corresponding to the instructions directly or by using other components under the control of the processor. Instructions may include code generated or executed by a compiler or an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory computer-readable storage medium. Here, "non-transitory" means only that the storage medium does not contain a signal and is tangible, and does not distinguish whether data is stored semi-permanently or temporarily in the storage medium.

[0176] In addition, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided by being included in a computer program product.

[0177] Specifically, the method includes: an operation of identifying the locations of multiple objects in each frame by inputting multiple frames of the content into a trained first artificial intelligence model when a preferred object is identified among multiple objects included in the content; an operation of obtaining first location information of the preferred object estimated based on the first frame by inputting the locations of the preferred objects in each frame and camera field of view information in each frame into a trained second artificial intelligence model when the preferred object is not included in the first frame among the multiple frames; and an operation of displaying a graphic user interface (GUI) for indicating the estimated location of the preferred object in the first frame based on the first location information of the preferred object.

[0178] Computer program products may be distributed in the form of device-readable storage media (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0179] In addition, computer instructions or programs for performing control methods of electronic devices according to the various embodiments described above may be stored on a non-transitory computer-readable medium. When computer instructions stored on such a non-transitory computer-readable medium are executed by a processor of a specific device, they cause the specific device to perform processing operations according to the various embodiments described above. A non-transitory computer-readable medium refers to a medium that stores data semi-permanently and is readable by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of a non-transitory computer-readable medium may include CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, ROMs, etc.

[0180] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.

Claims

1. In an electronic device, display; Memory for storing instructions; and One or more processors including processing circuitry; and The above one or more processors, When the above instructions are executed individually or collectively, the electronic device, Based on a plurality of frames corresponding to the content, a first artificial intelligence model obtains a location corresponding to a preferred object among a plurality of objects included in the content for each of the plurality of frames, and If the preferred object is not included in the first frame output to the display among the plurality of frames, the position information of the preferred object is obtained through a second artificial intelligence model based on the position corresponding to the preferred object for each of the plurality of frames and the camera field of view information for each of the plurality of frames, and An electronic device that controls the display to output a UI (user interface) including an indicator corresponding to the estimated position of the preferred object in the first frame, based on the position information of the preferred object obtained through the second artificial intelligence model.

2. In Paragraph 1, The location information of the above-mentioned preferred object is the first location information, and When the above instructions are executed individually or collectively by the one or more processors, the electronic device, Acquire first field of view information of the camera corresponding to the first frame, and Acquiring second position information of the preferred object included in the second frame prior to the first frame and second field of view information of the camera corresponding to the second frame, and Acquiring third position information of the preferred object included in the third frame between the first frame and the second frame and third field of view information of the camera corresponding to the third frame, and An electronic device that obtains the first position information of the preferred object corresponding to the first frame through the second artificial intelligence model based on the first field of view information of the camera, the second position information of the preferred object, the second field of view information of the camera, the third position information of the preferred object, and the third field of view information of the camera.

3. In Paragraph 1, When the above instructions are executed individually or collectively by the one or more processors, the electronic device, An electronic device that, if the preferred object is not included in the first frame output to the display among the plurality of frames, obtains first position information of the preferred object corresponding to the first frame through the learned second artificial intelligence model, the position corresponding to the preferred object for each of the plurality of frames, camera field of view information for each of the plurality of frames, and frame rate information of the plurality of frames.

4. In Paragraph 1, When the above instructions are executed individually or collectively by the one or more processors, the electronic device, An electronic device that controls the display to output the plurality of object positions and the camera's field of view to a top view image corresponding to the entire space of the content, based on the plurality of object positions per frame and the camera's field of view information.

5. In Paragraph 4, When the above instructions are executed individually or collectively by the one or more processors, the electronic device, An electronic device for controlling a display to output a plurality of frame-by-frame top-view images, including the plurality of object positions and the field of view of the camera, in a Picture-in-Picture (PIP) format to each of the plurality of frames.

6. In Paragraph 1, When the above instructions are executed individually or collectively by the one or more processors, the electronic device, An electronic device that controls the display to output a UI related to a position corresponding to the preferred object to one of the multiple edge regions of the first frame based on the multiple object positions included in the first frame.

7. In Paragraph 1, The above electronic device is, Including a communication circuit; further When the above instructions are executed individually or collectively by the one or more processors, the electronic device, An electronic device that, upon receiving data corresponding to a plurality of captured images of the content obtained from a plurality of different shooting angles through the communication circuit, obtains a position corresponding to the preferred object in the first frame based on the data corresponding to the plurality of captured images.

8. In Paragraph 1, When the above instructions are executed individually or collectively by the one or more processors, the electronic device, The display is controlled to output a UI corresponding to the selection of a preferred object among the plurality of objects included in the above content, and An electronic device that identifies a selected object as the preferred object through a UI corresponding to the selection of the preferred object.

9. In Paragraph 1, When the above instructions are executed individually or collectively by the one or more processors, the electronic device, An electronic device that obtains camera field of view information for each of the plurality of frames based on information corresponding to a first object included in the entire space corresponding to the above content and information corresponding to a second object included in each of the plurality of frames.

10. In Paragraph 1, When the above instructions are executed individually or collectively by the one or more processors, the electronic device, When an image corresponding to content in which the location of the preferred object is displayed for each of the plurality of frames is received from the server through the communication circuit, first location information of the preferred object corresponding to the first frame is obtained, and An electronic device that controls the display to output a user interface (UI) including an indicator corresponding to the estimated position of the preferred object in the first frame, based on the first position information of the preferred object.

11. In a method for controlling an electronic device, An operation of obtaining a location corresponding to a preferred object among a plurality of objects included in the content for each of the plurality of frames through a first artificial intelligence model based on a plurality of frames corresponding to the content; If the preferred object is not included in the first frame among the plurality of frames, the operation of obtaining location information of the preferred object through a second artificial intelligence model based on the location corresponding to the preferred object for each of the plurality of frames and camera field of view information for each of the plurality of frames; and A control method comprising: an operation of outputting a UI (user interface) including an indicator corresponding to the estimated position of the preferred object in the first frame based on the position information of the preferred object obtained through the second artificial intelligence model.

12. In Paragraph 11, The location information of the above-mentioned preferred object is the first location information, and An operation to acquire first field of view information of the camera corresponding to the first frame; An operation to obtain second position information of the preferred object included in the second frame prior to the first frame and second field of view information of the camera corresponding to the second frame; The operation of obtaining third position information of the preferred object included in the third frame between the first frame and the second frame and third field of view information of the camera corresponding to the third frame; and A control method comprising: an operation of acquiring the first position information of the preferred object corresponding to the first frame by the second artificial intelligence model based on the first field of view information of the camera, the second position information of the preferred object, the second field of view information of the camera, the third position information of the preferred object, and the third field of view information of the camera.

13. In Paragraph 11, A control method comprising: an operation of obtaining first position information of the preferred object corresponding to the first frame through the learned second artificial intelligence model, wherein if the preferred object is not included in the first frame among the plurality of frames, the position corresponding to the preferred object for each of the plurality of frames, camera field of view information for each of the plurality of frames, and frame rate information of the plurality of frames.

14. In Paragraph 11, An operation to acquire a top view image corresponding to the entire space of the above content; A control method comprising: an operation of outputting the plurality of object positions and the camera's field of view to a top view image corresponding to the entire space of the content, based on the plurality of object positions per frame and the camera's field of view information.

15. A non-transient computer-readable storage medium storing computer instructions that cause said electronic device to perform an operation when executed by a processor of said electronic device, wherein said operation is, An operation of obtaining a location corresponding to a preferred object among a plurality of objects included in the content for each of the plurality of frames through a first artificial intelligence model based on a plurality of frames corresponding to the content; If the preferred object is not included in the first frame among the plurality of frames, the operation of obtaining location information of the preferred object through a second artificial intelligence model based on the location corresponding to the preferred object for each of the plurality of frames and camera field of view information for each of the plurality of frames; and A non-transient computer-readable storage medium comprising: an operation of outputting a user interface (UI) including an indicator corresponding to the estimated position of the preferred object in the first frame based on the position information of the preferred object obtained through the second artificial intelligence model.