Electronic device and control method

The electronic device uses neural networks to identify sports events and generate personalized UIs and questions, addressing the challenge of limited information access during sports broadcasts, thereby improving user understanding and engagement.

WO2025249740A1PCT designated stage Publication Date: 2025-12-04SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/004293
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-28
Filing Date
2025-04-01
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Users watching sports events without prior knowledge struggle to understand the game and distinguish between desired and unwanted information, and are limited to the content provided by the content provider, requiring a way to access additional information during the broadcast.

Method used

An electronic device equipped with neural network models identifies sports events and prompt objects, generates user-customized UIs and questions, and provides additional information through a guide UI and external servers in response to user queries.

Benefits of technology

Enables users to understand sports events better by providing personalized information and answers to questions in real-time during the broadcast, enhancing the viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004293_04122025_PF_FP_ABST
    Figure KR2025004293_04122025_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device comprises a memory, a communication interface, a display, and at least one processor, wherein the at least one processor: inputs an image frame included in image content into a first neural network model to identify a sports event included in the image content; controls the display to provide a UI corresponding to the identified sports event on the image content; identifies, among objects included in the image content, at least one prompt object for generating an expected question about the identified sports event; generates a guide UI for guiding an expected question about the sports event on the basis of information on the identified at least one prompt object; and controls the display to display the generated guide UI with the image content.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices and control methods

[0001] The present invention is an invention that, when watching a sports event based on generative AI, identifies a prompt object on the screen to generate an expected question and provides a user-customized UI for each sports event based on the user's response.

[0002] Nowadays, electronic devices like TVs provide a variety of content and services. For example, electronic devices can receive and present various types of content from various content providers (e.g., broadcasters, OTTs, etc.). Furthermore, electronic devices not only provide content but also provide a variety of information related to that content.

[0003] Meanwhile, traditionally, users who watched sports videos without prior background knowledge of the sport struggled to understand the game. Furthermore, users were unable to distinguish between desired and unwanted information, and were limited to viewing only the information provided by the content provider.

[0004] In addition, in the past, in order to obtain information on a sports event, there was the inconvenience of having to receive information on the sports event through a separate user terminal or having to exit the current sports broadcast screen and search for information on the sports event through a separate screen.

[0005] Therefore, there is a need to find a way to provide information related to sports while watching sports screens.

[0006] According to another exemplary embodiment of the present disclosure for solving the above-described technical problem, an electronic device may be provided, comprising: a memory; a communication interface; a display; and at least one processor; wherein the at least one processor inputs a video frame included in video content into a first neural network model to identify a sport included in the video content, controls the display to provide a UI corresponding to the identified sport on the video content, identifies at least one prompt object for generating an expected question about the identified sport among objects included in the video content, generates a guide UI for guiding an expected question about the sport based on information about the identified at least one prompt object, and controls the display to display the generated guide UI together with the video content.

[0007] In addition, the at least one processor may control the display to display a message asking whether to change to an AI sports mode corresponding to the identified sports event when the sports event is identified, and may control the display to provide a UI corresponding to the identified sports event when a user command to change to the AI ​​sports mode is input.

[0008] In addition, the at least one processor is a second neural network model that is trained to identify at least one prompt object for each of the identified sports items, and the at least one processor can input the image content into the second neural network model to obtain information about the at least one prompt object.

[0009] Additionally, the at least one processor can control the display to display the at least one prompt object distinctly from other objects included in the video content.

[0010] In addition, the at least one processor may input information about the at least one acquired prompt object into a generative LMM model to obtain a prompt text, input the acquired prompt text into a third neural network model to obtain an expected question about the sport, generate a guide UI including the acquired expected question, and control the display to display the generated guide UI together with the video content.

[0011] In addition, the at least one processor can control the display to obtain information responding to the question about the sport included in the user voice when a user voice including a question about the sport is input while the guide UI is displayed, and to provide information responding to the question about the sport included in the user voice.

[0012] Additionally, the at least one processor may control the display to display information responding to the question around a prompt object related to a question about a sport included in the user's voice.

[0013] In addition, the at least one processor may control the display to obtain information on recommended content related to the question about the sport included in the user voice from an external server when a user voice including a question about the sport is input while the guide UI is displayed, and to provide information about the recommended content together with the video content.

[0014] In addition, the at least one processor may store information about a question about a sport included in the user voice by matching it with the identified sport, and when the AI ​​sports mode corresponding to the identified sport is changed again, the processor may control the display to obtain information responding to the question about a sport included in the user voice based on the information about the question included in the stored user voice, and provide information responding to the question about a sport included in the user voice.

[0015] In addition, when the sports event is identified, the method includes: controlling the display to display a message asking whether to change to an AI sports mode corresponding to the identified sports event; and when a user command to change to the AI ​​sports mode is input, controlling the display to provide a UI corresponding to the identified sports event.

[0016] In addition, the second neural network model is a neural network model trained to identify at least one prompt object for each of the identified sports items, and includes a step of inputting the image content into the second neural network model to identify the at least one prompt object.

[0017] A step of displaying at least one identified prompt object to distinguish it from other objects included in the video content;

[0018] The method comprises the steps of: inputting at least one identified prompt object into a third neural network model; and generating a question list for guiding questions corresponding to the identified sport.

[0019] A step of displaying the generated question list together with the video content is included.

[0020] A step of providing information on at least one question when a user requests information on at least one question from the list of questions generated above;

[0021] It includes a step of providing recommended content contextually related to the user's request based on a user's request for at least one question from the above-mentioned generated question list.

[0022] A step of storing information about at least one question requested by a user from the above-mentioned generated question list in correspondence with the identified sport; and a step of displaying information about at least one question requested by a user stored in correspondence with the identified sport when changed to AI sports mode.

[0023] According to an exemplary embodiment of the present disclosure for solving the above-described technical problem, a method for controlling an electronic device may be provided, including: inputting a video frame included in video content into a first neural network model to identify a sports event included in the video content; controlling the display to provide a UI corresponding to the identified sports event on the video content; identifying at least one prompt object for generating a question corresponding to the identified sports event among objects included in the video content; generating a question list for guiding a question corresponding to the sports event based on information about the at least one identified prompt object; and displaying the generated question list together with the video content.

[0024] The solutions to the problems of the present disclosure are not limited to the solutions described above, and solutions that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from this specification and the attached drawings.

[0025] FIG. 1 is a drawing for explaining a block diagram of an electronic device according to one embodiment of the present disclosure.

[0026] FIG. 2 is a diagram illustrating a neural network model stored in a memory of an electronic device according to one embodiment of the present disclosure.

[0027] FIG. 3 is a drawing illustrating a screen displaying a UI for determining a sport included in video content according to one embodiment of the present disclosure.

[0028] FIG. 4 is a diagram illustrating a screen displaying a message asking whether to change to AI sports mode according to one embodiment of the present disclosure.

[0029] FIG. 5 is a diagram illustrating a screen displaying a UI corresponding to a sporting event identified on video content according to one embodiment of the present disclosure.

[0030] FIG. 6 is a drawing illustrating a screen that displays UI corresponding to the type of game, golf hole, golf ball, and description of the game according to one embodiment of the present disclosure.

[0031] FIG. 7 is a diagram illustrating a screen displaying a golfer's name, a maximum height of a ball, data related to a golf ball, a speed of a golf ball, a UI indicating the progress of a game, and a shape such as a circle around a golfer's game score, according to one embodiment of the present disclosure.

[0032] FIG. 8 is a diagram illustrating a screen displaying a UI corresponding to a list of expected questions according to one embodiment of the present disclosure.

[0033] FIG. 9 is a diagram illustrating a screen displaying a guide UI for guiding expected questions about sports events according to one embodiment of the present disclosure.

[0034] FIG. 10 is a diagram illustrating a screen displaying information on one question requested by a user from a list of generated questions according to one embodiment of the present disclosure.

[0035] FIG. 11 is a diagram illustrating a screen that provides recommended content contextually related to a user's request, according to one embodiment of the present disclosure.

[0036] FIG. 12 is a diagram illustrating a screen that displays information about a question matching a sport already stored in a video frame included in video content according to one embodiment of the present disclosure.

[0037] FIG. 13 is a drawing illustrating a screen displaying a UI corresponding to information about distance, speed, heart rate, etc. of players according to one embodiment of the present disclosure.

[0038] FIG. 14 is a diagram illustrating a screen when changed to an AI sports mode for each sport according to one embodiment of the present disclosure.

[0039] FIG. 15 is a flowchart for explaining the operation of an AI sports mode according to one embodiment of the present disclosure.

[0040] Hereinafter, various embodiments of the present invention will be described with reference to the attached drawings. It should be understood that the contents described herein are not intended to limit the scope of the present invention to specific embodiments, but rather include various modifications, equivalents, and / or alternatives of the embodiments. In connection with the description of the drawings, the same or similar reference numerals may be used for similar components.

[0041] Additionally, the terms "first," "second," and the like used herein are used to distinguish various components from each other, regardless of order or importance. Therefore, these terms do not limit the order or importance of the components. For example, the first component could be renamed the second component, and similarly, the second component could be renamed the first component, without departing from the scope of the rights set forth in this document.

[0042] Additionally, when it is stated herein that one component (e.g., a first component) is operatively or communicatively coupled or connected to another component (e.g., a second component), it should be understood that this includes all cases where the components are directly connected or indirectly connected through another component (e.g., a third component). Conversely, when it is stated that a component (e.g., a first component) is "directly coupled" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between the component and the other component.

[0043] The terms used in this disclosure are used to describe certain embodiments and may not be intended to limit the scope of other embodiments. In addition, although singular expressions may be used in this disclosure for convenience of explanation, this may be interpreted to include plural expressions unless the context clearly indicates otherwise. In addition, the terms used in this disclosure may have the same meaning as generally understood by a person of ordinary skill in the relevant technical field. Among the terms used in this disclosure, terms defined in general dictionaries may be interpreted as having the same or similar meaning in the context of the related technology, and shall not be interpreted in an idealized or overly formal meaning unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude the embodiments of this disclosure.

[0044] Hereinafter, various embodiments of the present invention will be described in detail using the attached drawings.

[0045] FIG. 1 is a block diagram illustrating a configuration of an electronic device (100) according to at least one embodiment of the present disclosure. As illustrated in FIG. 1, the electronic device (100) includes a display (110), a memory (120), a user input unit (130), a communication interface (140), a speaker (150), a microphone (160), an input / output interface (170), a camera (180), and at least one processor (190). Meanwhile, the configuration of the electronic device (100) illustrated in FIG. 1 is merely an example, and it is to be understood that some configurations may be added depending on the type of the electronic device (100).

[0046] An electronic device (100) according to one embodiment of the present disclosure may be implemented as a TV, but this is only one embodiment, and may be implemented as a user terminal such as a smart phone, a tablet PC, a notebook PC, etc., and may be implemented as various devices such as home appliances, IoT devices, etc.

[0047] The display (110) can display various information. In particular, the display (110) can display a video frame included in the video content. In particular, the display (110) can display a UI corresponding to a sports event identified in the video content, a guide UI for guiding expected questions about the sports event, a message regarding whether to change to AI sports mode, at least one prompt object in a video frame included in the video content, information responding to a question about the sports event included in the user's voice, etc. For example, if the video content is about 'golf', the display (110) can display a golf broadcast screen, a UI corresponding to golf, a guide UI for guiding expected questions about the golf event, a message regarding whether to change to AI golf mode, and at least one prompt object included in a golf video frame. Here, the prompt object refers to an object that can induce an expected question when generating an expected question related to a sports event in a video frame included in the video content. The extraction method and display method of the UI, video content, and prompt object that can be displayed by the display (110) will be described in detail in FIGS. 3 to 14 below.

[0048] Meanwhile, the display (110) may be implemented as an LCD (Liquid Crystal Display Panel), OLED (Organic Light Emitting Diodes), etc., but is not limited thereto. In addition, the display (110) may be implemented as a flexible display, a transparent display, etc., depending on the case.

[0049] 36 The memory (120) can store an operating system (OS) for controlling the overall operation of components of the electronic device (100) and instructions or data related to components of the electronic device (100). Meanwhile, the memory (120) can be implemented as a non-volatile memory (e.g., hard disk, SSD (Solid state drive), flash memory), volatile memory, etc.

[0050] The memory (120) can store information on prompt objects that can be extracted for each sport. In one embodiment, if the sport is golf, the memory (120) can store information on golfers, holes, balls, number of rounds, etc. as prompt objects related to golf. Accordingly, the electronic device (100) can store prompt objects for each sport.

[0051] The memory (120) can store multiple neural network models. FIG. 2 is a diagram illustrating neural network models stored in the memory of an electronic device according to one embodiment of the present disclosure. As illustrated in FIG. 2, the memory (120) can include a first neural network model (220), a second neural network model (240), a third neural network model (270), and a generative LMM model.

[0052] As illustrated in FIG. 2, the first neural network model (220) is a neural network model trained to acquire information about a sport. When the electronic device (100) inputs a video frame (210) included in the video content into the first neural network model (220), the electronic device (100) can acquire information (230) about the sport. Specifically, when the electronic device (100) inputs a golf broadcast video frame into the first neural network model (220), the electronic device (100) can acquire information (230) that the video frame included in the video content is 'golf' among the sports.

[0053] In addition, the second neural network model (240) is a neural network model for obtaining information (250) about an object that can be a prompt object among the video frames (210). When the electronic device (100) inputs a video frame (210) included in the video content into the second neural network model (240), the electronic device (100) can obtain information (250) about the prompt object. For example, when the electronic device (100) inputs a golf broadcast video frame into the second neural network model (240), the electronic device (100) can obtain information about a UI, such as a golfer, hole, ball, and round number, as a prompt object. However, the electronic device (100) can receive information (250) about a prompt object for each sport from an external device.

[0054] The electronic device (100) can input information (250) about at least one identified prompt object into a generative LMM model to obtain information about prompt text. Specifically, the electronic device (100) can input information about an image or text of 'hole' or 'golf ball' among the at least one identified prompt object of a sport into a generative LMM model to obtain information (260) about prompt texts such as 'hole', 'golf ball', 'slope', 'distance', and 'maximum height of golf ball'.

[0055] Here, the generative LMM model is a type of generative AI model, and the generative LMM model can be implemented as a first neural network model (220), a second neural network model (240), or a third neural network model (270). Specifically, the generative AI model is an open conversational language model based on natural language processing including the GPT (Generative Pretrained Transformer) series, and is characterized in that it is generated by learning a learning dataset that inputs personal information entities and conversation types for specific situations, and outputs conversational text data in which personal information entities and conversation types are combined. Meanwhile, the memory (120) may further include a generative LMM model.

[0056] The third neural network model (270) is a neural network model trained to acquire information (280) about expected questions when information (260) about prompt text is input. Specifically, when information (260) about 'hole', 'golf ball', 'slope', and 'distance' among the acquired prompt texts are input into the third neural network model (270), the electronic device (100) can acquire information (280) about expected questions related to golf. In this case, the information (280) about expected questions related to golf may be, for example, 'What is the distance and slope between the golf ball and the hole?'

[0057] Meanwhile, in the above-described embodiment, it has been described that multiple neural networks are included in the memory (120) of the electronic device (100), but this is only one embodiment, and at least one neural network model among the multiple neural network models illustrated in FIG. 2 may be included in another device (e.g., a server, etc.).

[0058] In addition, the neural network model described in the above-described embodiment is a recognition model implemented in software or hardware that imitates the computational ability of a biological system by using a large number of artificial neurons connected by connection lines. The artificial intelligence model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs the operation of the layer through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), and deep Q-networks, and a transformer, and the neural network in the present disclosure is not limited to the above-described examples except as specified.

[0059] A learning algorithm is a method for training a target device using a large amount of learning data, enabling the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Unless otherwise specified, the learning algorithms in this disclosure are not limited to the aforementioned examples.

[0060] The memory (120) may be implemented in various forms, such as volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)).

[0061] The user input unit (130) may include a remote control receiver, a button, a lever, a switch, a touch interface, etc. In this case, the touch interface may be implemented in a manner of receiving input by the user's touch on the display (110) screen of the electronic device (100).

[0062] In particular, the user input unit (130) may receive a user command for a UI indicating whether to operate in AI sports mode. Alternatively, if a list of expected questions is displayed on the display (110), the user input unit (130) may receive a user command to select at least one question from the list of expected questions. Furthermore, the user input unit (130) may receive a question input when the user asks a question about a sport.

[0063] The communication interface (140) includes at least one circuit and can communicate with various types of external devices or servers. The communication interface (140) can include at least one of a BLE (Bluetooth Low Energy) module, a Wi-Fi communication module, a cellular communication module, a 3G (third generation) mobile communication module, a 4G (fourth generation) mobile communication module, a 4th generation LTE (Long Term Evolution) communication module, and a 5G (fifth generation) mobile communication module.

[0064] In particular, the communication interface (140) can communicate with other electronic devices. The communication interface (140) can receive video content from other devices, or receive information responding to questions about sports included in a user's voice. Alternatively, it can also receive information about prompt objects specific to sports.

[0065] The speaker (150) can output various voice messages and audio. In particular, the speaker (150) can output text regarding information about a sport or a list of expected questions. Specifically, when a user is watching a sports video, the speaker (150) can output text such as, "Ask a question about the sport you are watching," "A list of expected questions has been displayed. Please select a question you are curious about," or the text of the list of expected questions. In this case, the speaker (150) can output the voice of the text guiding the user's question acquired through the TTS module.

[0066] The microphone (160) can acquire the user's voice. In particular, the microphone (160) may be installed inside the electronic device (100), but this is only one embodiment, and the microphone (160) may be installed outside the electronic device (100) and electrically connected to the electronic device (100).

[0067] In particular, the microphone (160) can acquire questions about the user's sports. At this time, the questions the user utters into the microphone (160) can be converted into text and entered.

[0068] A microphone (160) can generate (or convert) voice or sound received from the outside into an electrical signal. The electrical signal generated by the microphone (160) can be stored in the memory (120) or output through a speaker (150). The microphone (160) may be composed of one or more microphones.

[0069] The input / output interface (170) is a configuration for inputting / outputting at least one of audio and video signals. In particular, the input / output interface (170) can receive information about video content from an external device. Among the input / output interfaces (170), the input interface includes a circuit, and at least one processor (190) can receive a user command for controlling the operation of the electronic device (100) through the input interface. Specifically, the input interface may be implemented as a remote control, but this is only one embodiment, and may be configured with a touch screen, a button, a keyboard, a mouse, and the like. For example, the input / output interface (170) may be an HDMI (High Definition Multimedia Interface), but this is only one embodiment, and may be any one interface among an MHL (Mobile High-Definition Link), a USB (Universal Serial Bus), a DP (Display Port), a Thunderbolt (Thunderbolt), a VGA (Video Graphics Array) port, an RGB port, a D-SUB (D-subminiature), and a DVI (Digital Visual Interface). Depending on the implementation example, the input / output interface (170) may include separate ports for inputting and outputting only audio signals and ports for inputting and outputting only video signals, or may be implemented as a single port for inputting and outputting both audio signals and video signals.

[0070] The camera (180) can capture images of the surroundings of the electronic device (100). In particular, the camera (180) can capture images including a user. At least one processor (190) can recognize a user included in the images captured by the camera (180) and determine whether to switch to AI sports mode based on the recognized user.

[0071] Specifically, if a record of user A watching in AI sports mode is stored in the memory (120), and a record of user B watching without switching to AI sports mode is stored in the memory (120), at least one processor (190) can distinguish and recognize users included in an image captured by the camera (180). At this time, if it is recognized that user A is watching video content based on the image captured by the camera (190), it can be controlled to switch to AI sports mode by at least one processor (190). In addition, if it is recognized that user B is watching video content based on the image captured by the camera (180), it can be controlled to maintain the general sports mode by at least one processor (190).

[0072] At least one processor (190) is electrically connected to a display (110), a memory (120), a user input unit (130), a communication interface (140), a speaker (150), a microphone (160), an input / output interface (170), and a camera (180), and controls the overall operation of the electronic device (100) using various commands or programs stored in the memory (120).

[0073] In addition, at least one processor (190) may include one or more processors. Specifically, the one or more processors may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a hardware accelerator, or a machine learning accelerator. The at least one processor (190) may control one or any combination of other components of the electronic device (100) and may perform operations related to communication or data processing. The at least one processor (190) may execute one or more programs or instructions stored in the memory (120). For example, the at least one processor (190) may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in the memory (120).

[0074] At least one processor (190) may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicores or heterogeneous multicores). When one or more processors are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.

[0075] In particular, at least one processor (190) inputs a video frame included in the video content into a first neural network model by executing at least one instruction to identify a sport included in the video content. At least one processor (190) controls the display (110) to provide a UI corresponding to the identified sport on the video content. At least one processor (190) identifies at least one prompt object for generating an expected question about the identified sport among objects included in the video content. At least one processor (190) generates a guide UI for guiding an expected question about the sport based on information about the at least one identified prompt object. At least one processor (190) controls the display (110) to display the generated guide UI together with the video content.

[0076] Hereinafter, the operation of an electronic device (100) controlled by at least one processor (190) will be described in more detail with reference to FIGS. 3 to 14.

[0077] First, the electronic device (100) can acquire video content from various sources. The video content may be sports-related. In one or more embodiments, the video content may be golf-related. However, this is merely an example, and the video content may be related to other sports (e.g., soccer, baseball, etc.).

[0078] The electronic device (100) can store at least one image frame (210) included in the acquired image content in a buffer included in the memory (120).

[0079] The electronic device (100) can identify a sport included in the video content using at least one acquired video frame (210). Specifically, the electronic device (100) can input at least one video frame (210) included in the video content into a first neural network model (220) to identify a sport included in the video content (230). At this time, the first neural network model (220) is a model trained to input a video frame and acquire information corresponding to the video frame, and may be implemented as a CNN, but is not limited thereto. The first neural network model (220) has been described above, and thus a detailed description thereof will be omitted.

[0080] Meanwhile, in the above-described embodiment, the electronic device (100) inputs at least one image frame into the first neural network model (220) to identify a sports event (230) included in the image content. However, this is merely an example, and the sports event (230) included in the image content can be identified in other ways. For example, the electronic device (100) can identify a sports event (230) included in the image content by analyzing metadata of the image content, and the electronic device (100) can identify a sports event (230) included in the image content by performing OCR recognition on the image.

[0081] In one or more embodiments, the electronic device (100) may provide a UI (310) for confirming a sport included in video content to the user. For example, the electronic device (100) may input a video frame (210) included in the acquired video content into a first neural network model (220) to obtain information (230) about a sport called 'golf'. At this time, the electronic device (100) may display a UI (310) for confirming a sport included in the video content to the user, such as "Are you currently watching golf?", as illustrated in FIG. 3 .

[0082] Although FIG. 3 illustrates a case where a UI (310) for requesting confirmation of a sport is displayed on the display (110), it is also possible to display a UI that allows the user to select a sport on the display (110). For example, when a video frame is input to the first neural network model and it is determined to be golf, but in reality it is not golf, it is also possible for the user to directly input the sport through the user input unit (130).

[0083] Additionally, if the first neural network model (220) is not stored in the memory (120), it is also possible for the user to input a sports event through the user input unit (130).

[0084] When a sport is identified, the electronic device (100) may display a message asking whether to change to an AI sports mode corresponding to the identified sport. When a user command to change to the AI ​​sports mode is input, the electronic device (100) may change to the AI ​​sports mode corresponding to the identified sport and display a UI corresponding to the identified sport. Here, the AI ​​sports mode is a mode that displays information about the sport generated by the electronic device along with real-time video content received from an external source.

[0085] Specifically, when information on a sport corresponding to a video frame (210) included in video content is acquired, the electronic device (100) can determine whether to change to AI sports mode. As illustrated in FIG. 4, a UI (410) indicating whether to change to AI sports mode can be displayed on the display (110). At this time, the message (410) regarding whether to change to AI sports mode can be implemented as a UI, but is not limited thereto.

[0086] For example, the display (110) may display a message (410) asking the user whether to change to AI sports mode, such as “Switch to AI golf mode?” as illustrated in FIG. 4. At this time, the display (110) provides the message (410) “Switch to AI sports mode?” However, this is only an example, and if a specific shape is displayed on a part of the display (110) and the user inputs the specific shape using a remote control or the like, the electronic device (100) may change to AI sports mode.

[0087] Specifically, when the electronic device (100) obtains information that the video frame (210) included in the video content is a video frame for 'golf', the message (410) on whether to change to AI sports mode on the display (110) may be "Switch to AI golf mode?" In this case, when the user inputs a refusal to change to AI sports mode through the user input unit (130) while the message (410) on whether to change to AI sports mode is displayed on the display (110), the electronic device (100) may provide video content of the basic sports mode. The basic sports mode refers to a mode in which the electronic device (100) does not generate information on a sports event, but displays video content of a real-time sports broadcast received from an external device.

[0088] If the user inputs a positive response to the message (410) asking whether to change to AI sports mode, the electronic device (100) may provide a screen such as that shown in FIGS. 12 to 14.

[0089] However, the above-described embodiments are merely examples, and when a video frame (210) included in the video content is input into the first neural network model (220), if a sport is identified, it is also possible to automatically change to AI sports mode without user input.

[0090] Additionally, if the user inputs a positive response to the message (410) asking whether to change to AI sports mode, the electronic device (100) may provide a screen including a UI corresponding to the identified sports event, as illustrated in FIG. 5. In this case, the UI corresponding to the sports event refers to a UI that the electronic device (100) provides by default for each sports event.

[0091] For example, the electronic device (100) can input a video frame (210) included in the video content into the first neural network model (220) to obtain information (230) about a sport. At this time, a UI that is basically provided for each sport may be stored in the memory (210). In one embodiment, the video frame (210) included in the video content may be input into the first neural network model (220) to obtain information that the video content is 'golf'. At this time, if the video content is 'golf', the memory (210) may store 'golfer's name', 'congratulatory sound, congratulatory motion, or color change when a specific event occurs' as corresponding basic UIs. If the user inputs a positive response to a message (410) asking whether to change to AI sports mode, the basic UI may be displayed together with the video frame (210) including the video content. Specifically, as illustrated in FIG. 5, a video frame (210) included in the video content may display a golfer's name (510), the distance between the golf ball and the hole (520), and a UI (530) indicating the AI ​​sports mode. In addition, if the identified sport is 'golf', the 'expected trajectory of the golf ball' may also be included in the basic UI.

[0092] The electronic device (100) can obtain information (250) about at least one prompt object included in a video frame to generate a sports-related expected question while operating in AI sports mode.

[0093] Here, a prompt object refers to an object that can elicit expected questions related to a sport from a video frame included in the video content. For example, if the sport included in the video content is identified as "golf," the prompt object could be a "UI displaying the golfer, hole, and number of rounds."

[0094] At this time, the display (110) can be controlled to display at least one prompt object to be distinguished from other objects included in the video content. For example, the electronic device (100) can display an object corresponding to the prompt object in the video frame by distinguishing it with a circle or the like. However, this is only one embodiment, and the display (110) can display the prompt object not only by using a circle or the like, but also by using a method such as making the prompt object blink to be distinguished from other objects included in the video frame.

[0095] Specifically, if the identified sport of the video frame (210) included in the video content is 'golf', the prompt objects may be 'golfer', 'hole', 'golf ball', 'game type', etc. As illustrated in FIG. 6, the electronic device (100) may display the acquired prompt object, which is a golfer (10), a UI (600) displaying the type of game, a golf hole (610), a golf ball (620), and a UI (630) displaying a description of the game. Similarly, in FIG. 7, the electronic device (100) may display a shape such as a circle around the golfer's name (710), the highest height of the ball (720), data related to the golf ball (730), the speed of the golf ball (740), the UI (750) indicating the game progress, and the golfer's game score (760) among the prompt objects included in the video frame (210) including the video content, so as to be displayed in a manner distinguishable from other objects in the video frame (210) included in the video content.

[0096] At this time, the electronic device (100) can input the video frame (210) included in the video content into the second neural network model (240) to obtain information (250) about the prompt object. The method of obtaining information (250) about the prompt object using the second neural network model (240) has been described above in FIG. 2, and a detailed description thereof will be omitted.

[0097] In addition, in order to obtain information (250) about a prompt object, information about the prompt object stored in the memory (120) of the electronic device (100) can be used. Meanwhile, it is also possible to obtain information (250) about the prompt object from another device.

[0098] Meanwhile, FIGS. 6 and 7 illustrate a case where a prompt object is automatically displayed on the display (110), but a user can directly select a prompt object.

[0099] As described above, when the electronic device (100) acquires information (250) about a prompt object, the electronic device (100) inputs the acquired information (250) about at least one prompt object into a generative LMM model to acquire information (260) about a prompt text, inputs the acquired information (260) about the prompt text into a third neural network model (270) to acquire an expected question (280) about a sport, generates a guide UI (810, 910) including the acquired expected question, and displays the generated guide UI (810, 910) together with a video frame (210) included in the video content.

[0100] As illustrated in FIG. 7, the electronic device (100) can input the current game score (760) of golfer AAA and the UI (730) displaying the speed and maximum height of the golf ball, etc., among the prompt objects, into the third neural network model (280) to obtain information (290) regarding the expected question. Specifically, the information (290) regarding the expected question can be, for example, “What is the expected score based on the trajectory of the golf ball hit by this player?”

[0101] As for the method of displaying a guide UI including expected questions, there are possible embodiments of displaying a UI (810) in which a list of expected questions is displayed on a display (110) as illustrated in FIG. 8, as well as an embodiment of displaying a UI that guides a user to freely ask questions on a display (110) as illustrated in FIG. 9.

[0102] At this time, in order for the electronic device (100) to obtain a list of expected questions, information (260) about the prompt text can be input into the third neural network model (270) as illustrated in FIG. 2 to obtain information (280) about the expected questions. This has been described above with reference to FIG. 2, and a detailed description thereof will be omitted. In addition, when the electronic device (100) displays a list of expected questions (810), it can arrange them in order of high probability of selection by the user based on the information (250) about the acquired prompt objects. When displaying the list of expected questions (810) on the display (110), it is of course possible to have an embodiment in which specific questions are highlighted.

[0103] As illustrated in FIG. 9, the electronic device (100) may display a UI (910) containing the message “Feel free to ask questions.” In this case, the user may input a question through a user input unit (130), a microphone (160), or an input / output interface (170).

[0104] A guide UI (810, 910) for guiding expected questions about a sports event may be displayed on a part of a video frame (210) included in the video content as illustrated in FIG. 8 or FIG. 9, but the entire screen of the display (110) may display the guide UI (810, 910) for guiding expected questions.

[0105] When a user voice including a question about a sport is input while a guide UI (810, 910) is displayed, the electronic device (100) can obtain information responding to the question about the sport included in the user voice and display information responding to the question about the sport included in the user voice.

[0106] As described above, when a user inputs a question about a sport according to the guide UI (810, 910) for guiding expected questions about a sport, the electronic device (100) can provide information responding to the question about the sport as illustrated in FIG. 10.

[0107] For example, when a user inputs a question about a sport according to a guide UI (810, 910) for guiding expected questions, the electronic device (100) may display a UI (1010) that displays information responding to the question about the sport.

[0108] Specifically, when a user inputs question 1 from the list of expected questions in FIG. 8, “What is the distance and slope from the ball to the hole?”, the electronic device (100) may display a UI (1010) that displays a message, “The distance from the ball to the hole is 3M, and the slope is -3,” which is information responding to the question about the sport. At this time, the information responding to the question about the sport may be obtained from a content provider (e.g., a broadcaster, etc.) that provides video content or may be obtained from a separate content server. In addition, the information responding to the question about the sport may be obtained by analyzing the video. In addition, the information responding to the question about the sport may be obtained from information stored in the memory (120).

[0109] In addition, as shown in FIG. 10, information responding to a question about a sport is shown as a UI, but the speaker (150) can output a voice responding to a question about a sport obtained through a TTS module.

[0110] In addition, when a user's response is input according to a guide UI for guiding expected questions about a sport, and a user's voice including a question about the sport is input while the guide UI is displayed on the display (110), the electronic device (100) can control the display (110) to obtain information about recommended content related to the question about the sport included in the user's voice from an external server and provide the information about the recommended content together with the video content. As illustrated in FIG. 11, when a user's voice including a question about the sport is input, the electronic device (100) can display recommended content related to the user's question on the display (110).

[0111] For example, if a user inputs “Show me the player hitting the driver again” while a guide UI (810, 910) is displayed on the display (110), the electronic device (100) may determine that the user is interested in “driver hitting posture.” Accordingly, while the guide UI (810, 910) is displayed on the display (110), content (1110) related to “driver hitting posture,” which is information about recommended content related to a question about a sport included in the user’s voice, may be displayed on a part of the display (110). However, this is only one example, and the display (110) may display “a video of the player hitting the driver in another game” as recommended content related to “driver hitting posture.”

[0112] The memory (120) stores information about a question about a sport included in the user's voice by matching it with an identified sport, and when the electronic device (100) is changed back to an AI sports mode corresponding to the identified sport, the electronic device (100) can control the display (110) to obtain information responding to the question about a sport included in the user's voice based on the information about the question included in the stored user's voice, and provide information responding to the question about a sport included in the user's voice.

[0113] As illustrated in FIGS. 12 to 14, when the electronic device (100) inputs a video frame (210) included in the video content into the first neural network model (220), the electronic device (100) can obtain information (230) about a sports event. When the electronic device (100) changes to an AI sports mode corresponding to the identified sports event, the electronic device (100) can display information responding to a question about a sports event included in the user's voice, since information about a question about a sports event included in the user's voice that matches the identified sports event is stored in the memory (120).

[0114] When the electronic device (100) acquires information (230) about a sport, it can change to AI sports mode. At this time, the electronic device (100) can acquire information about questions matching the sport already stored in the memory (120) and display information about questions matching the sport already stored in the video frame (210) included in the video content.

[0115] For example, as illustrated in FIGS. 12 to 14, the display (110) may display information on a question matching a sports event together with a video frame (210) included in the video content. Specifically, according to FIG. 13, when a user watches a soccer video, the user may request information on the players' distance (1310), speed (1320), heart rate (1330), etc. from the electronic device (100). In this case, the memory (120) of the electronic device (100) may store 'players' heart rate', 'players' distance', and 'players' speed' as questions matching the sport of soccer. Specifically, when the video frame (210) included in the video content is identified as a sport of 'soccer', the electronic device (100) may change to the AI ​​soccer mode. The electronic device (100) can display information about the 'players' distance (1310),' 'players' speed (1320),' and 'players' heart rate (1330)' obtained from the memory (120) together with the video frame (210) included in the video content. This also applies to FIGS. 12 and 14. That is, in the case of FIG. 12, the information for the question matching golf stored in the memory (120) is 'the distance and slope from the ball to the hole' (1210), and in the case of FIG. 14, the information for the question matching weightlifting stored in the memory (120) is 'the tilt state of the barbell held by the weightlifter' (1410), 'the probability of the weightlifter's success' (1420), and 'the weightlifter's heart rate' (1430).

[0116] That is, since the memory (120) stores information about a question about a sport included in the user voice matched with the identified sport, the electronic device (100) can obtain information about a question about a sport included in the user voice matched with the identified sport from the memory (120). Accordingly, when the electronic device (100) is subsequently changed to the AI ​​sports mode, the display (110) can display a video frame (210) including video content and information about a question about a sport included in the user voice matched with the identified sport.

[0117] FIG. 15 is a flowchart for explaining the operation of an AI sports mode according to one embodiment of the present disclosure.

[0118] The electronic device (100) can acquire video content (S1510). At this time, the electronic device (100) may acquire video content currently being broadcast, or may acquire video content that has already been broadcast in the past.

[0119] And, the electronic device (100) can identify a sports event by inputting a video frame (210) included in the video content into a first neural network model (220) (S1520).

[0120] After a sport is identified, a UI corresponding to the identified sport can be displayed on the display (S1530). In this case, UIs such as "AI Golf Mode" and "AI Soccer Mode" can be displayed on the display (110) to indicate that the display has switched to AI Sports Mode and is displaying video content on the display (110).

[0121] In addition, the electronic device (100) can identify at least one prompt object for generating a question corresponding to the identified sport among the objects included in the video content (S1540). To identify at least one prompt object, a video frame (210) included in the video content can be input into the second neural network model (240). At this time, a method for identifying the prompt object for each sport may be stored in the memory (120).

[0122] Additionally, a question list can be generated to guide questions corresponding to the sport (S1550). The question list can be generated by inputting information (260) about the prompt text into a third neural network model (270). In this case, the third neural network model (270) is a generative AI model, capable of generating a new question list based on the answers to the acquired question list.

[0123] The electronic device (100) can display the generated question list on the display (110) along with the video frame (210) included in the video content (S1560). Questions regarding sports events are displayed as a UI, as shown in FIG. 8, and the user can input them using a remote control or the like, or select at least one question from the question list using a microphone (160) or the like. Furthermore, in addition to selecting a question from the question list, a method in which the user can freely ask a question directly, as illustrated in FIG. 9, may also be possible.

[0124] In addition, the electronic device (100) stores information about a question about a sport included in the user's voice in the memory (120) by matching it with the identified sport, and when the mode is changed back to an AI sports mode corresponding to the identified sport, the electronic device (100) can obtain information responding to the question about a sport included in the user's voice based on the information about the question included in the stored user's voice, and display the information responding to the question about a sport included in the user's voice on the display (110).

[0125] Specifically, information regarding questions about sports events can be stored in the memory (120). Accordingly, when the electronic device (100) is switched to AI sports mode, information (280) regarding the stored expected questions can be acquired from the memory (120). The acquired information (280) regarding the expected questions can be displayed in a video frame (210) included in the video content.

[0126] In addition, the methods according to various embodiments of the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user (20) devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0127] The methods according to various embodiments of the present disclosure may be implemented as software including commands stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call commands stored from the storage medium and operate according to the called commands, and may include a server device or an electronic device according to the disclosed embodiments.

[0128] Meanwhile, a device-readable storage medium may be provided in the form of a non-transitory readable recording medium. Here, the term "non-transitory readable recording medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is stored semi-permanently in the storage medium and cases where data is stored temporarily. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0129] When the above instruction is executed by the processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter.

[0130] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. In electronic devices, memory; communication interface; display; and comprising at least one processor; At least one processor, By inputting the video frame included in the video content into the first neural network model, the sports event included in the video content is identified, Controlling the display to provide a UI corresponding to the identified sports event on the video content; Identifying at least one prompt object for generating an expected question about the identified sport among the objects included in the video content, Generate a guide UI for guiding expected questions about the sport based on information about at least one prompt object identified above, An electronic device that controls the display to display the generated guide UI together with the video content.

2. In paragraph 1, At least one processor, If the above sports event is identified, control the display to display a message asking whether to change to an AI sports mode corresponding to the identified sports event; An electronic device that controls the display to provide a UI corresponding to the identified sport when a user command to change to the AI ​​sports mode is input.

3. In paragraph 1, The second neural network model is a neural network model trained to identify at least one prompt object for each of the identified sports items. At least one processor, An electronic device that inputs the above video content into the second neural network model to obtain information about the at least one prompt object.

4. In paragraph 3, At least one processor, An electronic device that controls the display to display at least one prompt object distinctly from other objects included in the video content.

5. In paragraph 4, At least one processor, Inputting information about at least one of the above-obtained prompt objects into a generative LLM model to obtain a prompt text, By inputting the above-obtained prompt text into the third neural network model, expected questions about the above-mentioned sports are obtained, Create a guide UI that includes the expected questions obtained above, An electronic device that controls the display to display the generated guide UI together with the video content.

6. In paragraph 5, At least one processor, When a user voice including a question about the sports event is input while the above guide UI is displayed, information responding to the question about the sports event included in the user voice is obtained, An electronic device controlling the display to provide information in response to a question about a sport included in the user's voice.

7. In paragraph 6, At least one processor, An electronic device that controls the display to display information responding to the above question, such that the information is displayed around a prompt object related to a question about a sport included in the user's voice.

8. In paragraph 6, At least one processor, When a user voice including a question about the sports event is input while the above guide UI is displayed, information about recommended content related to the question about the sports event included in the user voice is obtained from an external server, An electronic device that controls the display to provide information about the recommended content together with the video content.

9. In paragraph 7, At least one processor, Information about a question about a sport included in the user's voice is matched with the identified sport and stored, When the AI ​​sports mode corresponding to the identified sports is changed again, information for responding to a question about a sports event included in the user voice is obtained based on information about the question included in the stored user voice, An electronic device controlling the display to provide information in response to a question about a sport included in the user's voice.

10. In a method for controlling an electronic device, A step of inputting a video frame included in video content into a first neural network model to identify a sports event included in the video content; A step of controlling a display to provide a UI corresponding to the identified sports event on the video content; A step of identifying at least one prompt object for generating a question corresponding to the identified sport among the objects included in the video content; A step of generating a list of questions to guide questions corresponding to the sports event based on information about at least one identified prompt object; and A method for controlling an electronic device, comprising: a step of displaying the generated question list together with the video content.

11. In paragraph 10, When the above sports event is identified, a step of controlling the display to display a message asking whether to change to an AI sports mode corresponding to the identified sports event; and A method for controlling an electronic device, comprising: a step of controlling the display to provide a UI corresponding to the identified sport when a user command to change to the AI ​​sports mode is input; 12. In paragraph 10, The second neural network model is a neural network model trained to identify at least one prompt object for each of the identified sports items. A method for controlling an electronic device, comprising: inputting the image content into the second neural network model to identify at least one prompt object.

13. In paragraph 12, A method for controlling an electronic device, comprising: a step of displaying at least one identified prompt object in a manner distinguishable from other objects included in the video content; 14. In paragraph 13, A step of inputting at least one of the identified prompt objects into a third neural network model; and A method for controlling an electronic device, comprising: a step of generating a list of questions for guiding questions corresponding to the identified sports event; 15. In paragraph 13, A method for controlling an electronic device, comprising: a step of displaying the generated question list together with the video content.

Citation Information

Patent Citations

  • Method for Operating Contents

    KR1020090096578A

  • Processing method of rare earth oxide coating layer

    KR1020230049317A

  • Method for pairing wireless apparatus using remote controller and wireless apparatus using the same

    KR102587686B1

  • Interaction Interleaver

    US20190076741A1

  • Dialogue systems using knowledge bases and language models for automotive systems and applications

    US20240095460A1