Electronic device, operation method thereof, and storage medium

By analyzing video metadata and converting voice commands to text, the electronic device effectively executes user commands during playback, addressing inefficiencies in recognizing voice inputs and improving user interaction.

WO2025254446A1PCT designated stage Publication Date: 2025-12-11SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/007638
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-06-04
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing electronic devices struggle to efficiently recognize and execute user commands based on voice inputs during video playback, particularly in complex environments where multiple commands are possible.

Method used

The electronic device employs a processor to analyze metadata from video images and audio inputs, converting voice commands into text and identifying relevant actions within defined playback segments using algorithms like speech-to-text and image recognition, allowing for precise command execution.

Benefits of technology

Enables accurate and timely execution of user commands during video playback, enhancing user interaction and functionality in electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025007638_11122025_PF_FP_ABST
    Figure KR2025007638_11122025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The electronic device comprises a microphone, memory for storing instructions, and at least one processor, wherein the instructions, when executed by the at least one processor, cause the electronic device to: play back a video; define a plurality of playback sections related to the video, including a first playback section and a second playback section; define, from among a set of instructions, a first sub-set of instructions executable in the first playback section; define, from among the set of instructions, a second sub-set of instructions executable in the second playback section; receive a first speech input that is input through the microphone while playing back the video; convert the first speech input into first text; identify a first instruction corresponding to the first text; and when the first speech input is received while playing back the first playback section of the video and the first instruction is included in the first sub-set, execute an operation corresponding to the first instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices and their operating methods and storage media

[0001] The present disclosure relates to an electronic device for confirming a user command, a method of operating the same, and a storage medium.

[0002] Thanks to remarkable advancements in information and communication technology and semiconductor technology, the proliferation and use of various electronic devices is rapidly increasing. Electronic devices are being developed to enable users to carry and communicate with one another. An electronic device can refer to any device that performs a specific function based on its embedded software, such as a mobile communication terminal, tablet PC, audio / video device, desktop / laptop computer, or in-car navigation system.

[0003] Meanwhile, the variety of services and additional features offered through electronic devices, such as smartphones, is steadily increasing. To enhance the utility of these devices and satisfy the diverse needs of users, telecommunications service providers and electronic device manufacturers are competitively developing electronic devices that offer diverse features and differentiate themselves from competitors. Consequently, the various functions offered through electronic devices are also becoming increasingly sophisticated. For example, electronic devices now offer the ability to recognize voice input and perform actions based on user commands related to the recognized voice.

[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0005] An electronic device according to one embodiment of the present disclosure includes a microphone, at least one processor including a processing circuit, and a memory storing instructions, the memory including one or more storage media, wherein the instructions, when individually or collectively executed by the at least one processor, can cause the electronic device to play an image.

[0006] In one embodiment, the instructions may cause the electronic device to define a plurality of playback segments associated with the video, the plurality of playback segments including a first playback segment and a second playback segment.

[0007] In one embodiment, the instructions may cause the electronic device to define a first subset of instructions, among a set of instructions, executable in the first playback section.

[0008] In one embodiment, the instructions may cause the electronic device to define a second subset of instructions, executable in the second playback section, of the set of instructions.

[0009] In one embodiment, the instructions may cause the electronic device to receive a first audio input input through the microphone while playing the video.

[0010] In one embodiment, the instructions may cause the electronic device to convert the first speech input into first text.

[0011] In one embodiment, the instructions may cause the electronic device to identify a first command corresponding to the first text.

[0012] In one embodiment, the instructions may cause the electronic device to execute an action corresponding to the first command when the first voice input is received during playback of the first playback section of the video and the first command is included in the first subset.

[0013] A method of operating an electronic device according to one embodiment of the present disclosure may include an operation of playing an image.

[0014] According to one embodiment, the method may include defining a plurality of playback segments associated with the video, the plurality of playback segments including a first playback segment and a second playback segment.

[0015] In one embodiment, the method may include defining a first subset of instructions, among a set of instructions, that are executable in the first playback section.

[0016] In one embodiment, the method may include defining a second subset of instructions executable in the second playback section among the set of instructions.

[0017] According to one embodiment, the method may include receiving a first voice input input through a microphone while playing the video.

[0018] In one embodiment, the method may include converting the first voice input into first text.

[0019] In one embodiment, the method may include an operation of identifying a first command corresponding to the first text.

[0020] According to one embodiment, the operating method may include an operation of executing an operation corresponding to the first command when the first voice input is received during playback of the first playback section of the video and the first command is included in the first subset.

[0021] In a storage medium storing computer-readable instructions according to one embodiment of the present disclosure, the instructions, when executed by at least one processor of an electronic device, can cause the electronic device to play an image.

[0022] In one embodiment, the instructions may cause the electronic device to define a plurality of playback segments associated with the video, the plurality of playback segments including a first playback segment and a second playback segment.

[0023] In one embodiment, the instructions may cause the electronic device to define a first subset of instructions, from among a set of instructions, that are executable in the first playback section.

[0024] In one embodiment, the instructions may cause the electronic device to define a second subset of instructions, executable in the second playback section, of the set of instructions.

[0025] In one embodiment, the instructions may cause the electronic device to receive a first audio input input through a microphone while playing the video.

[0026] In one embodiment, the instructions may cause the electronic device to convert the first speech input into first text.

[0027] In one embodiment, the instructions may cause the electronic device to identify a first command corresponding to the first text.

[0028] In one embodiment, the instructions may cause the electronic device to execute an operation corresponding to the first command when the first voice input is received during playback of the first playback section of the video and the first command is included in the first subset.

[0029] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0030] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.

[0031] FIG. 2 is a block diagram of configurations of an electronic device according to one embodiment.

[0032] FIG. 3 is a flowchart illustrating an operating method of an electronic device according to one embodiment.

[0033] FIG. 4 is a drawing for explaining a method for checking a playback section of an image according to one embodiment.

[0034] FIG. 5 is a flowchart illustrating a method for verifying a user command according to one embodiment.

[0035] FIG. 6 is a diagram illustrating a method for obtaining metadata according to one embodiment.

[0036] FIG. 7 is a flowchart illustrating a method for confirming a user command based on image information of a video according to one embodiment.

[0037] FIG. 8 is a flowchart illustrating a method for confirming a user command based on audio information of an image according to one embodiment.

[0038] FIG. 9 is a flowchart illustrating a method for confirming a user command based on audio information of an image according to one embodiment.

[0039] FIGS. 10A and 10B are diagrams illustrating a method of providing information through a display according to one embodiment.

[0040] FIG. 11 is a diagram illustrating a method for confirming a user command according to one embodiment.

[0041] FIG. 12 is a diagram illustrating a plurality of modules for verifying a user command according to one embodiment.

[0042] Fig. 13 is a flowchart for explaining an operating method of an electronic device according to one embodiment.

[0043] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0044] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to one embodiment. Referring to FIG. 1 , in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0045] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0046] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0047] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0048] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0049] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0050] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0051] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0052] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0053] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0054] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0055] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0056] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0057] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0058] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0059] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0060] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0061] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0062] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0063] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0064] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0065] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0066] In the detailed description below, reference numerals in the drawings may be used interchangeably or omitted for components that can be easily understood through the preceding embodiments, and their detailed descriptions may also be omitted. An electronic device according to an embodiment disclosed in this document may be implemented by selectively combining components of different embodiments, and components of one embodiment may be replaced by components of another embodiment. For example, it should be noted that the present invention is not limited to specific drawings or embodiments.

[0067] FIG. 2 is a block diagram of electronic device configurations according to one embodiment.

[0068] According to FIG. 2, according to one embodiment, an electronic device (200, e.g., electronic device (101) of FIG. 1) may include a microphone (210), a memory (220, e.g., memory (130) of FIG. 1) for storing instructions, and at least one processor (230, or processor).

[0069] In one embodiment, the microphone (210) may have at least a portion of the same or similar configuration as the input module (150) of FIG. 1. In one embodiment, the microphone (210) may receive an acoustic signal. In one embodiment, the electronic device (200) may include one or more microphones.

[0070] According to one embodiment, the memory (220) may have at least a portion of the same or similar configuration as the memory (130) of FIG. 1. For example, the memory (220) may be configured to temporarily or permanently store digital data and may include at least a portion of the configuration and / or functions of the memory (130) of FIG. 1.

[0071] The memory (220) according to one embodiment can store various instructions that can be executed by at least one processor (230). In addition, the memory (220) can store at least a portion of the program (140) of FIG. 1. Such instructions can include control commands such as logical operations and data input / output that can be recognized and executed by the processor (230). There is no limitation on the type and / or amount of data that the memory (220) can store, but this document will describe the configuration and function of the memory related to the operation of the processor (230) that performs the method and the method of confirming a user command according to various embodiments. The memory (220) can store various information, and the various information stored by the memory (220) will be described in detail below.

[0072] According to one embodiment, at least one processor (230, hereinafter, processor) may have at least a portion of the same or similar configuration as the processor (120) of FIG. 1. According to one embodiment, the processor (230) may include one or more processors.

[0073] According to one embodiment, the processor (230) may perform various operations by executing instructions stored in the memory (220).

[0074] According to one embodiment, the processor (230) may obtain at least one of image information or audio information of a video. According to one example, the video may be video content provided to a user. According to one example, the processor (230) may obtain at least one of image information or audio information of a video from an external electronic device (e.g., the electronic device (102), the electronic device (104), or the server (108) of FIG. 1) through a communication module (e.g., the communication module (190) of FIG. 1).

[0075] According to one embodiment, the processor (230) may obtain metadata corresponding to a playback section of the video based on at least one of image information or audio information of the video.

[0076] For example, metadata may be additional information corresponding to a video. For example, metadata may include information about the total playback time (or running time) of the video, the current playback point, and whether the video is currently playing. For example, metadata may include time information corresponding to the start and end of a video playback section. For example, the time information may indicate a specific point in time within the video playback section.

[0077] Alternatively, according to an example, the metadata may include information related to the video (e.g., information about the video content or text information related to the video content). According to an example, the metadata may have time information mapped to it. According to an example, there may be metadata corresponding to each of at least one playback section of the video. Alternatively, according to an example, the metadata may include identification information for each section of the video.

[0078] For example, the processor (230) may acquire metadata based on image information of a video. For example, if information regarding the video's running time, current playback time, and playback status is included within an image corresponding to a specific point in the video, the processor (230) may acquire metadata using an image analysis algorithm. This will be described in detail with reference to FIG. 4.

[0079] For example, the processor (230) may obtain metadata based on audio information of the video. For example, if audio information corresponding to the entire playback section of the video is obtained, the processor (230) may use a keyword extraction algorithm to obtain metadata including keywords related to the content of the video from the audio information. This will be described in detail with reference to FIG. 8.

[0080] According to one embodiment, the processor (230) may convert the input first sound (or first voice input) into text when the first sound is input through the microphone. According to one example, the first sound may be a user sound for changing the playback status of a video. For example, the user sound for changing the playback status of a video may be a user sound for changing the playback position of a video, such as 'move forward 10 seconds' or 'go to the 1 minute 6 second section', or a user sound for changing the playback environment of a video, such as 'turn up the volume'. According to one example, when the first sound is input through the microphone, the processor (230) may convert the first sound into text using a designated algorithm. For example, the designated algorithm may be a voice recognition algorithm based on an STT (speech to text) algorithm, which will be described in detail with reference to FIG. 12.

[0081] According to one embodiment, the processor (230) may identify a user command related to a playback section based on the converted text and metadata. According to one example, the processor (230) may identify multiple playback sections of the video based on the acquired metadata. According to one example, when multiple playback sections of the video are identified, the processor (230) may identify a first playback section in which a first sound is input among the multiple playback sections. According to one example, the processor (230) may identify a command corresponding to the converted text among the commands corresponding to the first playback section as a user command. A specific method for identifying multiple playback sections of the video will be described in detail with reference to FIG. 4.

[0082] For example, the processor (230) may determine a user command by considering the playback section in which the first sound is input. For example, if, among the commands corresponding to the first playback section, there is no command corresponding to the converted text, the processor (230) may determine that there is no user command related to the playback section. For example, if the command corresponding to the converted text is a command corresponding to the second playback section among the plurality of playback sections, the processor (230) may determine that there is no user command since the first sound is input in the first playback section. A specific method for determining a user command will be described later.

[0083] According to one embodiment, the processor (230) may perform an operation corresponding to the identified user command based on the user command related to the playback section. According to one example, if a user command corresponding to a first playback section among multiple playback sections of a video is identified as a user command for moving the playback time to a first playback time, the processor (230) may move the playback time of the video to the first playback time and provide the video at the first playback time.

[0084] FIG. 3 is a flowchart illustrating an operating method of an electronic device according to one embodiment.

[0085] Hereinafter, an operating method of an electronic device (e.g., an electronic device (200) of FIG. 2) according to various embodiments will be described in detail. According to various embodiments, operations performed by the electronic device described below may be executed by a processor (e.g., at least one processor (230) of FIG. 2) including at least one processing circuitry of the electronic device. According to one embodiment, the operations performed by the electronic device may be stored in a memory (e.g., a memory (220) of FIG. 2) and, when executed, may be executed by instructions that cause the processor (230) to operate. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel. Depending on the implementation, certain operations may be omitted.

[0086] Referring to FIG. 3, according to one embodiment, in operation 301, the operating method may include an operation of obtaining metadata (e.g., metadata of FIG. 2) corresponding to a playback section of a video based on at least one of image information or audio information of the video. According to one example, the electronic device may obtain image information of the video from an external electronic device (e.g., the external electronic device of FIG. 2). According to one example, the electronic device may obtain metadata corresponding to a playback section of the video based on the obtained image information of the video. For example, the electronic device may obtain metadata corresponding to a playback section of the video from a play bar included in an image of the video using a specified image recognition algorithm. According to one example, the metadata may include information about a running time of the video and information about a current playback point in time.

[0087] According to one embodiment, in operation 303, the operating method may include an operation of converting the input first sound (e.g., the first sound of FIG. 2) into text corresponding to the first sound being input into a microphone (e.g., the microphone (210) of FIG. 2). According to one example, when the first sound corresponding to the user's sound is input into the microphone, the electronic device may convert the first sound into text using a designated speech recognition algorithm (e.g., a speech-to-text (STT) algorithm). For example, the electronic device may obtain 'Go to the 1 minute 6 second section' as text information based on the first sound.

[0088] According to one embodiment, in operation 305, the operating method may include an operation of confirming a user command related to a playback section based on the converted text and metadata. According to one example, the electronic device may confirm a plurality of playback sections of the video based on the acquired metadata. According to one example, when the plurality of playback sections of the video are confirmed, the electronic device may confirm a first playback section in which a first sound is input among the plurality of playback sections. According to one example, the electronic device may confirm a command corresponding to the converted text among commands corresponding to the first playback section as a user command. For example, the electronic device may confirm 'Move to the playback time corresponding to 1 minute 6 seconds' among commands corresponding to the first playback section as a command corresponding to 'Go to the 1 minute 6 second section'. Alternatively, according to one example, when the electronic device confirms that 'Go to the 1 minute 6 second section' does not correspond to a command corresponding to the first playback section, the electronic device may determine that there is no user command.

[0089] FIG. 4 is a drawing for explaining a method for checking a playback section of an image according to one embodiment.

[0090] Referring to FIG. 4, according to an embodiment, an electronic device (400, for example, the electronic device (200) of FIG. 2) may display an image through a display (401, for example, the display module (160) of FIG. 1). According to an example, the electronic device (400) may obtain metadata (for example, the metadata of FIG. 2) from image information of the image, and may identify a playback section of the image based on the obtained metadata. According to an example, when an input related to the image is received, the electronic device (400) may obtain playback information including at least one of a total playback time (or running time) of the image, a current playback time, and information on whether the image is being played as metadata from image information corresponding to the time at which the input is received.

[0091] For example, when an input related to a video (e.g., an input related to video playback, such as a touch input or a button input) is received, the electronic device (400) may obtain image information including a play bar (411). For example, the play bar (411) may include text (421) related to the video, such as information about the running time of the video, the current playback time, and whether the video is being played. Alternatively, for example, the text (421) related to the video may include text related to the content of the video. The electronic device (400) may obtain playback information from the image information of the video including the play bar (411) using a designated image recognition algorithm.

[0092] For example, the electronic device (400) may identify multiple playback sections of a video based on the confirmed playback information. For example, if the entire playback section of the video is divided into four playback sections of the same time, the electronic device (400) may identify time information (e.g., time information of FIG. 2) corresponding to the start and end points of each of the multiple playback sections. However, the present invention is not limited thereto, and the sizes of the multiple playback sections of the video may be different. For example, the electronic device (400) may obtain time information corresponding to the start and end points of each playback section as metadata corresponding to each playback section.

[0093] FIG. 5 is a flowchart illustrating a method for verifying a user command according to one embodiment.

[0094] Referring to FIG. 5, according to one embodiment, in operation 501, the operating method may include an operation of identifying a first playback section corresponding to a first sound (e.g., the first sound of FIG. 2) among a plurality of playback sections of the identified video.

[0095] In one example, an electronic device (e.g., the electronic device (200) of FIG. 2) can identify multiple playback sections of a video based on metadata (e.g., the metadata of FIG. 2) including identified playback information (e.g., the playback information of FIG. 4). In one example, when a first sound is input through a microphone (e.g., the microphone (210) of FIG. 2), the electronic device can identify a first playback section among the identified multiple playback sections based on the time at which the first sound was input through the microphone. In one example, the time at which the first sound was input through the microphone may be the time from the time at which playback of the video starts to the time at which the first sound was input. For example, the electronic device can identify a first playback section among the multiple playback sections based on time information of each of the multiple playback sections of the video and the time at which the first sound was input through the microphone.

[0096] According to one embodiment, in operation 503, the method may include an operation of identifying a user command corresponding to converted text among commands corresponding to a first playback section based on a command list including commands corresponding to each of a plurality of playback sections of the video. According to one example, the command list may be stored in a memory (e.g., memory (130) of FIG. 1 ). Alternatively, according to one example, the electronic device may obtain the command list from an external electronic device (e.g., server (108) of FIG. 1 ) through a communication module (e.g., communication module (190) of FIG. 1 ).

[0097] For example, the command list may refer to a set of commands corresponding to each of a plurality of playback sections of a video among the command sets. For example, the command set may refer to a set of a plurality of commands that operate on an electronic device. For example, the command set may be stored in a memory (e.g., the memory (220) of FIG. 2), or the electronic device may obtain the command set from an external electronic device (e.g., the electronic device (102), the electronic device (104), or the server (108) of FIG. 1) through a communication module (e.g., the communication module (190) of FIG. 1).

[0098] For example, the set of commands may include commands related to controlling a video (e.g., a command to increase the volume of a video, a command to decrease the volume, a command to pause a video, a command to set the audio mode of the video to a movie audio mode, a command to change the playback point to the next playback section, a command to play the video again from the beginning, a command to change the playback point to a specific point, a command to end a video, a command to play the next video, a command corresponding to an evaluation of a video, or a command to end a video playback application).

[0099] In one example, the electronic device may define a list of commands (or a subset of commands) corresponding to each of a plurality of playback sections among a set of commands. For example, in one of a plurality of playback sections of a video, a subset of commands corresponding to a starting section may include at least one of a command for changing an audio mode (e.g., a general mode, a music audio mode, or a movie audio mode) and a command for changing a playback time to a next section. In one example, the electronic device may define a subset corresponding to a starting section such that, among the set of commands, a command for changing an audio mode and a command for changing a playback time to a next section are included in the subset corresponding to the starting section.

[0100] Alternatively, for example, a subset of commands corresponding to a highlighted content section (e.g., a content-focused section of FIG. 7) may include a command for changing the playback time to a specified time point. In one example, the specified time point may be, for example, a time point for playing specific content. In one example, the electronic device may identify the command for changing the playback time to the specified time point based on at least one of image information (e.g., image information of FIG. 2) or audio information (e.g., audio information of FIG. 2). Alternatively, in one example, the electronic device may identify a subset of commands based on history information for a user command, which will be described later. Alternatively, in one example, the electronic device may identify a subset of commands based on information about content for each playback section of a video, which will be described later (e.g., information (610) about content of FIG. 6).

[0101] Alternatively, for example, the electronic device may define at least one of the commands for playing the next image and the commands corresponding to evaluating the image as a subset of the commands corresponding to the last section among the set of commands.

[0102] Alternatively, for example, the electronic device may include at least one of a command for increasing the volume of a video, a command for decreasing the volume, or a command for pausing a video, among a set of commands, in a list of commands corresponding to each playback segment included in the plurality of playback segments.

[0103] For example, when a first playback segment is identified, the electronic device can check for a command corresponding to the converted text among the commands corresponding to the identified first playback segment. For example, if the converted text is "Go to the 1 minute 6 second segment," the electronic device can check for a command corresponding to "Go to the 1 minute 6 second segment" among the commands corresponding to the first playback segment.

[0104] For example, if the electronic device determines that the command corresponding to 'Go to the 1 minute 6 second section' among the commands corresponding to the first playback section includes 'Move playback time to the 1 minute 6 second section', the electronic device may provide a video of the corresponding playback time. For example, the electronic device may determine that the user command does not exist based on the determination that the command corresponding to 'Go to the 1 minute 6 second section' among the commands corresponding to the first playback section does not exist. For example, even when the command corresponding to 'Go to the 1 minute 6 second section' among the commands corresponding to the second playback section different from the first playback section exists, the electronic device may determine that the user command does not exist.

[0105] According to one embodiment, the method may include updating a list of commands based on history information about the user command.

[0106] For example, the history information for a user command may include history information for a user command identified based on a sound input through a microphone (e.g., a first sound) and history information for a user command identified based on a user input (e.g., a user's touch input). For example, the history information may be information in which information about a user command and a corresponding playback section is mapped. For example, if a first user command is input in a first playback section, the history information may include information in which the first user command and the first playback section corresponding to the first user command are mapped.

[0107] In one example, the electronic device may update the command list based on history information about the user command. For example, it may be assumed that a first user command is included in the first playback section in the command list. If the electronic device determines that the first user command is frequently input in the second playback section based on the history information, the electronic device may update the command list so that the first user command is included in the command list corresponding to the second playback section. For example, if the electronic device determines that a second user command, which is not included in the command list, is input in the third playback section, the electronic device may determine a command corresponding to the determined second user command and update the command list so that the determined command is included in the command list corresponding to the third playback section.

[0108] For example, if the electronic device determines that a command belonging to multiple playback sections is input only in one playback section, the electronic device may delete the above-described command from the command list corresponding to the remaining playback sections except for the one confirmed playback section.

[0109] According to the above example, the probability of misrecognition of user voices can be reduced by verifying user commands based on the playback section in which the user's voice is input. Furthermore, since the list of commands corresponding to the playback section is updated based on user history information, the recognition rate of user commands corresponding to the user's voice can be improved.

[0110] FIG. 6 is a diagram illustrating a method for obtaining metadata according to one embodiment.

[0111] Referring to FIG. 6, according to one embodiment, an electronic device (600, e.g., the electronic device (200) of FIG. 2) may obtain metadata (e.g., the metadata of FIG. 2) based on information (610, or a time stamp) about content for each playback section of a video, which is information included in image information of a video. According to one example, an image of a video may include information (610) about content for each playback section of the video. According to one example, the electronic device (600) may display the video through a display (601, e.g., the display module (160) of FIG. 1).

[0112] For example, the electronic device (600) may obtain metadata including tag information related to the content of the video and a playback point corresponding to the tag information from information (610) about the content for each playback section of the video using a designated image recognition algorithm (e.g., the image recognition algorithm of FIG. 3). For example, the electronic device (600) may obtain metadata in which the tag 'recent weight status' and the playback point corresponding thereto are mapped with '0:00'. Alternatively, for example, the electronic device (600) may obtain metadata in which the tag 'Coach Cha's PT consultation' and the playback point corresponding thereto are mapped with '1:06'. Alternatively, for example, the electronic device (600) may obtain metadata in which the tag '1_body_building_to_lose_1_weight' and the playback point corresponding thereto are mapped with '1:18'.

[0113] According to one embodiment, the electronic device (600) may update a command list (e.g., the command list of FIG. 5 ) based on the acquired metadata. For example, the electronic device (600) may identify a command for changing the playback time of a video to a playback time corresponding to tag information. For example, the electronic device (600) may update the command list so that the identified command is included in at least one playback section of the video. This will be described in detail with reference to FIG. 7 .

[0114] According to one embodiment, the electronic device (600) can identify a user command related to a playback section based on tag information associated with the acquired content and a playback point corresponding to the tag information. According to one example, when metadata is acquired based on tag information associated with the content and a playback point corresponding to the tag information, and a command list is updated based on the acquired metadata, the electronic device (600) can identify a user command based on the updated command list.

[0115] For example, it may be assumed that the text corresponding to the first sound input in the second playback section (e.g., the converted text in FIG. 2) is 'Go to Coach Cha's PT consultation.' The electronic device (600) may, based on the updated command list, identify a command corresponding to 'change the playback time of the video to the playback time corresponding to 1:06' among the commands corresponding to the second playback section.

[0116] FIG. 7 is a flowchart illustrating a method for confirming a user command based on image information of a video (e.g., image information of FIG. 2) according to one embodiment.

[0117] Referring to FIG. 7, according to one embodiment, in operation 701, the operating method may include an operation of identifying a command corresponding to a second playback section among a plurality of playback sections of a video based on tag information related to the acquired content (e.g., tag information of FIG. 6) and a playback point in time corresponding to the tag information.

[0118] In one example, when an electronic device (e.g., the electronic device (200) of FIG. 2) acquires metadata including tag information related to content and a playback point corresponding to the tag information, the electronic device may acquire a command based on the acquired metadata. For example, the electronic device may acquire a command for changing the playback point of a video to a playback point corresponding to the acquired tag information, based on the acquired metadata. In one example, the electronic device may add the command for changing the playback point corresponding to the acquired tag information to a command corresponding to a designated playback section. For example, the electronic device may update the command list so that the identified command is included in the command list corresponding to a content-intensive section where a user command for moving to a specific playback point is mainly input.

[0119] According to one embodiment, in operation 703, the method may include an operation of identifying a user command corresponding to a converted text (e.g., the converted text of FIG. 2) among commands corresponding to the identified second playback section based on identifying a second playback section corresponding to a first sound (e.g., the first sound of FIG. 2). According to one example, it may be assumed that the first sound is input in a content-focused section. The electronic device may identify a command corresponding to the first sound among the command list corresponding to the content-focused section based on the updated command list.

[0120] FIG. 8 is a flowchart illustrating a method for confirming a user command based on audio information of an image according to one embodiment.

[0121] Referring to FIG. 8, according to one embodiment, in operation 801, the operating method may include an operation of obtaining a keyword related to the content of the video and a playback time corresponding to the keyword based on audio information of the video (e.g., audio information of FIG. 2). According to one example, the electronic device may obtain audio information of the video from an external electronic device (e.g., the electronic device (102), the electronic device (104), or the server (108) of FIG. 1) through a communication module (e.g., the communication module (190) of FIG. 1). According to one example, the audio information of the video may be audio information corresponding to the entire playback section of the video.

[0122] For example, an electronic device may obtain keywords related to the content of a video and corresponding playback points as metadata from audio information using a designated keyword extraction algorithm. For example, the keywords related to the content of the video may be words or phrases that appear repeatedly in a designated playback section or are identified as having high importance in the designated playback section. For example, when a keyword corresponding to a designated playback section is identified, the electronic device may identify either the start point or the end point of the designated playback section as the playback point corresponding to the identified keyword. Alternatively, for example, the playback point corresponding to the identified keyword may be a playback point at which the keyword appears among the entire playback points.

[0123] According to one embodiment, in operation 803, the operating method may include an operation of confirming a user command related to a playback section based on an acquired keyword and a playback time corresponding to the acquired keyword. According to one example, the electronic device may obtain a command based on the acquired keyword and a playback time corresponding thereto. According to one example, the electronic device may obtain a command for changing the playback time of the video to a playback time corresponding to the acquired keyword. According to one example, the electronic device may confirm whether a text corresponding to a first sound (e.g., the first sound of FIG. 2) corresponds to the acquired command, and confirm the user command based thereon. This will be described in detail with reference to FIG. 9 below.

[0124] FIG. 9 is a flowchart illustrating a method for confirming a user command based on audio information of an image (e.g., audio information of FIG. 2) according to one embodiment.

[0125] Referring to FIG. 9, according to one embodiment, in operation 901, the operating method may include an operation of checking a command list (e.g., a command list of FIG. 5) including commands corresponding to each of a plurality of playback sections of an image (e.g., a plurality of playback sections of an image of FIG. 2), based on an acquired keyword (e.g., a keyword of FIG. 8) and a playback point corresponding to the keyword (e.g., a playback point of FIG. 8).

[0126] For example, an electronic device (e.g., electronic device (200) of FIG. 2) may, based on an acquired keyword and a corresponding playback time, identify a command for changing the playback time of a video to a playback time corresponding to the acquired keyword. For example, the electronic device may add a command for changing the playback time of a video to a playback time corresponding to the acquired keyword to commands corresponding to a specified playback section. For example, the electronic device may update a command list such that the identified command is included in a command list corresponding to a content-intensive section where user commands for moving to a specific playback time are mainly input.

[0127] According to one embodiment, at operation 903, the method may include an operation of identifying a user command corresponding to the converted text based on the identified command list.

[0128] For example, when a command list is updated based on an acquired keyword and a corresponding playback point, the electronic device can confirm a user command based on the updated command list. For example, it may be assumed that a first sound (e.g., the first sound of FIG. 2) is input in a content-focused section. Based on the updated command list, the electronic device can confirm a command corresponding to a converted text (e.g., the converted text of FIG. 2) from among the command list corresponding to the content-focused section. For example, if the converted text is 'Move to the first keyword,' the electronic device can confirm a command for moving to a first playback point corresponding to the first keyword from among the command list corresponding to the content-focused section. Or, for example, if the converted text is 'Do you want to see the second keyword again,' the electronic device can confirm a command for moving to a second playback point corresponding to the second keyword from among the command list corresponding to the content-focused section.

[0129] For example, based on the presence of a command corresponding to a first sound, the electronic device may identify the identified command as a user command and perform a corresponding action. For example, if a command to move to a first playback point corresponding to a first keyword is identified, the electronic device may provide a video corresponding to the first playback point.

[0130] FIGS. 10A and 10B are drawings for explaining a method of providing information through a display (1010, e.g., display module (160) of FIG. 1) according to one embodiment.

[0131] Referring to FIG. 10A, according to one embodiment, an electronic device (e.g., the electronic device (200) of FIG. 2) may include a display (1012). According to one example, the electronic device (1010) may display information (1011, or information indicating acquisition of metadata) related to acquisition of metadata (e.g., metadata of FIG. 2) corresponding to a playback section of a video, through the display (1012), based on receiving an input related to a video (e.g., an input related to a video of FIG. 4). According to one example, the electronic device (1010) may display information (1011) related to acquisition of metadata corresponding to a playback section together with the video through the display (1012).

[0132] However, the present invention is not limited thereto, and as an example, the electronic device (1010) may display information (1011) related to acquisition of the above-described metadata through the display (1012) even when an icon of a different type (e.g., a menu icon or an icon corresponding to a play bar) is displayed than the image currently displayed through the display (1012).

[0133] Referring to FIG. 10b, according to one embodiment, the electronic device (1010) can confirm a user command based on the acquired metadata after information (1011) related to acquisition of metadata is displayed through the display (1012).

[0134] For example, the electronic device (1010) can identify a playback section of a video (e.g., a playback section of FIG. 4) based on the acquired metadata. Alternatively, for example, the electronic device (1010) can acquire metadata based on information about content for each playback section of the video (e.g., information about content (610) of FIG. 6, or a time stamp) and update a command list (e.g., a command list of FIG. 5) based thereon. For example, after information (1011) related to acquisition of metadata is displayed through the display (1012), when a first sound (e.g., the first sound of FIG. 2) is input, the electronic device (1010) can identify a user command corresponding to the first sound based on the identified playback section of the video or the updated command list.

[0135] FIG. 11 is a flowchart illustrating a method for verifying a user command (e.g., the user command of FIG. 2) according to one embodiment.

[0136] Referring to FIG. 11, according to one embodiment, in operation 1101, the operating method may include an operation of obtaining metadata (e.g., metadata of FIG. 2) corresponding to a playback section of a video based on at least one of image information (e.g., image information of FIG. 2) or audio information (e.g., audio information of FIG. 2) of the video.

[0137] In one example, an electronic device (e.g., the electronic device (200) of FIG. 2) may obtain image information of a video from an external electronic device (e.g., the external electronic device of FIG. 2). In one example, the electronic device may obtain metadata corresponding to a playback section of the video based on the image information of the obtained video. Alternatively, in one example, the electronic device may obtain audio information of the video from the external electronic device. In one example, the electronic device may obtain keywords and corresponding playback points as metadata from the audio information of the video using a specified voice recognition algorithm.

[0138] In one embodiment, in operation 1103, the method may include an operation of determining whether the first sound corresponds to a wake-up sound when a first sound (e.g., the first sound of FIG. 2) is input to a microphone (e.g., the microphone (210) of FIG. 2). In one example, when the first sound is input through the microphone during playback of a video, the electronic device may determine whether the first sound includes a wake-up sound.

[0139] According to one embodiment, in operation 1105, the method may include an operation of confirming a user command (e.g., a user command of FIG. 2) related to a playback section based on text and metadata corresponding to the first sound, based on determining that the first sound does not correspond to a wake-up sound. According to one example, when it is determined that the first sound does not include a wake-up sound, the electronic device may confirm a first playback section in which the first sound is input among a plurality of playback sections of the video (e.g., a plurality of playback sections of FIG. 2). According to one example, the electronic device may confirm a command corresponding to a converted text among commands corresponding to the first playback section as the user command. For example, the electronic device may confirm 'Move to the playback time corresponding to 1 minute 6 seconds' among commands corresponding to the first playback section as a command corresponding to 'Go to the 1 minute 6 second section'.

[0140] According to the above example, even if the first sound does not include a wake-up sound, the electronic device can identify a user command corresponding to the playback section based on the first sound. This can increase user convenience.

[0141] FIG. 12 is a diagram illustrating a plurality of modules for verifying a user command according to one embodiment.

[0142] Referring to FIG. 12, according to one embodiment, an electronic device (1200, e.g., electronic device (200) of FIG. 2) may include at least one of an image recognition module (1210), a voice recognition module (1220), an image control module (1230, or player control module), an agent module (1240), and a user history storage module (1250).

[0143] According to one embodiment, the image recognition module (1210) may obtain video playback information (e.g., playback information of FIG. 4) based on a play bar or a time stamp included in the image of the video. According to one example, the image recognition module (1210) may obtain video playback information using a specified image recognition algorithm. According to one example, the image recognition module (1210) may transmit the obtained playback information to an agent model. According to one example, when the electronic device (1200) cannot obtain video-related metadata through a background application (e.g., a content playback application), the electronic device (1200) may obtain metadata including playback information (e.g., metadata of FIG. 2) using the image recognition module (1210).

[0144] According to one embodiment, the speech recognition module (1220) can convert sound into text. For example, the speech recognition module (1220) may be a module that performs a speech-to-text (STT) function. For example, the speech recognition module (1220) may convert a first sound (e.g., the first sound of FIG. 2 ) input through a microphone (e.g., the microphone (210) of FIG. 2 ) into text. Alternatively, for example, the speech recognition module (1220) may obtain text corresponding to the sound of an image based on the sound information of the image (e.g., the sound information of FIG. 2 ).

[0145] According to one embodiment, the video control module (1230) may be a module that controls operations related to the playback of video content (e.g., playing, pausing, or adjusting the playback speed of the video). According to one example, the video control module (1230) may also be a module that controls an application related to the playback of the video.

[0146] According to one embodiment, the agent module (1240) may be a module that verifies a user command (e.g., the user command of FIG. 2) and performs an action corresponding to the verified user command. According to one example, the agent module (1240) may obtain metadata (e.g., the metadata of FIG. 2) from the image recognition module (1210). According to one example, the agent module (1240) may obtain text converted from a first sound from the voice recognition module (1220). According to one example, the agent module (1240) may obtain a user command related to a playback section based on the obtained metadata and the converted text.

[0147] Alternatively, according to an example, the agent module (1240) may obtain text corresponding to the audio of the video from the voice recognition module (1220). Alternatively, the agent module (1240) may obtain keywords related to the content of the video and playback points corresponding to the keywords based on the text corresponding to the audio of the video. Alternatively, the agent module (1240) may obtain a command list (e.g., the command list of FIG. 5) based on the obtained keywords and playback points. Alternatively, according to an example, the agent module (1240) may identify a user command related to a playback section based on the obtained keywords and playback points.

[0148] For example, the agent module (1240) can manage the user history storage module (1250). For example, when user history information (e.g., history information of FIG. 5) is updated, the agent module (1240) can update the user history storage module (1250) based on the updated user history information. Alternatively, the agent module (1240) can update the command list (e.g., command list of FIG. 6) based on the user history information obtained from the user history storage module (1250).

[0149] According to one embodiment, the user history storage module (1250) may store history information (1251, 1252, 쪋, 1253) corresponding to at least one user. According to one example, the user history storage module (1250) may update and store history information corresponding to at least one user based on user history information received from the agent module (1240).

[0150] Fig. 13 is a flowchart for explaining an operating method of an electronic device according to one embodiment.

[0151] Referring to FIG. 13, according to one embodiment, the operating method may include, in operation 1301, an operation of playing an image (e.g., the image of FIG. 2). According to one example, an electronic device (e.g., the electronic device (200) of FIG. 2) may play the image.

[0152] According to one embodiment, the method of operation may include, at operation 1303, defining a plurality of playback segments (e.g., the plurality of playback segments of FIG. 2 ) associated with the video, the plurality of playback segments including a first playback segment (e.g., the first playback segment of FIG. 2 ) and a second playback segment (e.g., the second playback segment of FIG. 2 ).

[0153] In one example, an electronic device may define a plurality of playback segments associated with a video based on a specified percentage of the total playback time of the video. In one example, the plurality of playback segments may include a plurality of segments including at least one of a starting segment, a highlight content segment (e.g., a content-focused segment of FIG. 7), and a last segment. In one example, the electronic device may define (or distinguish) the plurality of playback segments using playback information (e.g., the playback information of FIG. 4).

[0154] For example, among the plurality of playback sections, the starting section may be a playback section from the starting point of the video to a point in time when a first time has elapsed from the starting point of the video. For example, the electronic device may determine the first time based on a first ratio to the total playback time of the video (e.g., the total playback time of FIG. 2). For example, the total playback time of the video may be obtained based on metadata (e.g., the metadata of FIG. 2). For example, the first ratio may be 25%, but may also be a different value.

[0155] For example, the last segment of the multiple playback segments may be a segment from the end point of the video to a point before the second time and up to the end point of the video. For example, the electronic device may determine the second time based on a second percentage of the total playback time of the video. For example, the second percentage may be 25%, but may also be a different value.

[0156] In one embodiment, the method of operation may include, at operation 1305, defining a first subset of instructions, from among a set of instructions, executable in a first playback interval.

[0157] For example, a command set may refer to a set of multiple commands that operate on an electronic device. For example, the command set may be stored in a memory (e.g., a memory (220) of FIG. 2), or the electronic device may obtain the command set from an external electronic device (e.g., the electronic device (102), the electronic device (104), or the server (108) of FIG. 1) via a communication module (e.g., the communication module (190) of FIG. 1). For example, the command set may include commands related to controlling a video (e.g., a command for changing a playback time to a specific time, a command for ending a video, or a command for playing the next video).

[0158] In one example, a subset of commands (e.g., the command list of FIG. 5) may be commands executable in one of a plurality of playback sections (e.g., the first playback section of FIG. 2). In one example, the first subset of commands may be commands executable in the first playback section. In one example, the first subset of commands may include at least a portion of the set of commands. In one example, the electronic device may define commands related to a playback section of a video as a subset.

[0159] For example, a subset of commands corresponding to a start section may include at least one of a command for changing a sound mode (e.g., a general mode, a music sound mode, or a movie sound mode) and a command for changing a playback point to a next section. In one example, the electronic device may define a first subset such that, among the set of commands, the command for changing a sound mode and the command for changing a playback point to a next section are included in the first subset.

[0160] Alternatively, for example, a subset of commands corresponding to a highlighted content section may include a command for changing the playback time to a specified time point. In one example, the specified time point may be, for example, a time point for playing a specific content. In one example, the electronic device may determine the command for changing the playback time to a specified time point based on at least one of image information (e.g., image information of FIG. 2 ) or audio information (e.g., audio information of FIG. 2 ). Alternatively, in one example, the electronic device may update a subset of commands based on history information about a user command (e.g., history information about a user command of FIG. 5 ). Alternatively, in one example, the electronic device may update a subset of commands based on information about content for each playback section of a video (e.g., information (610) about content of FIG. 6 ).

[0161] In one embodiment, the method of operation may include, at operation 1307, defining a second subset of instructions executable in a second playback section of the set of instructions.

[0162] In one example, the electronic device may identify a command related to a second playback section from among a set of commands, and define the identified commands as a second subset. For example, the electronic device may define at least one of a command for playing a next video and a command corresponding to an evaluation of the video from among the set of commands as a subset of commands corresponding to the last section. In one example, the command for playing a next video may be a command for playing a video that is scheduled to be played automatically when the currently playing video ends. In one example, the command corresponding to an evaluation of the video may be a command for providing a user interface (UI) for inputting a user's evaluation of the currently playing video.

[0163] In one embodiment, the operating method may include, at operation 1309, receiving a first voice input (e.g., the first sound of FIG. 2 ) input through a microphone (e.g., the microphone (210) of FIG. 2 ) while playing a video. In one example, when a sound is input through the microphone while playing a video, the electronic device may determine whether the input sound is a voice input corresponding to a user.

[0164] According to one embodiment, the method of operation may include, at operation 1311, converting a first voice input into a first text (e.g., the text of FIG. 2 ). According to one example, the electronic device may convert the received first voice input into the first text using a specified algorithm (e.g., a speech-to-text (STT)-based voice recognition algorithm).

[0165] In one embodiment, the method of operation may include, at operation 1313, identifying a first command corresponding to a first text (e.g., a command corresponding to the converted text of FIG. 2 ). In one example, the electronic device may identify, among a set of commands, a first command corresponding to the first text.

[0166] In one embodiment, the method of operation may include, at operation 1315, if a first voice input is received during playback of a first playback segment of a video and the first command is included in the first subset, executing an operation corresponding to the first command (e.g., an operation corresponding to the user command of FIG. 2 ).

[0167] For example, the electronic device may be configured to perform an action corresponding to the first command when a first voice input is received while playing a first playback segment of a video and the identified first command is included in a first subset corresponding to the first playback segment.

[0168] In one example, the electronic device may be controlled not to perform an action corresponding to the first text based on a first voice input being received while playing a second playback segment of the video and the identified first command not being included in a second subset corresponding to the second playback segment.

[0169] An electronic device (101; 200) according to one embodiment of the present disclosure includes a microphone (210), at least one processor (230) including a processing circuit, and a memory (220) storing instructions and including one or more storage media, wherein the instructions, when individually or collectively executed by the at least one processor (230), can cause the electronic device (101; 200) to play an image.

[0170] In one embodiment, the instructions may cause the electronic device (101; 200) to define a plurality of playback segments associated with the video, including a first playback segment and a second playback segment.

[0171] In one embodiment, the instructions may cause the electronic device (101; 200) to define a first subset of instructions, among a set of instructions, executable in the first playback section.

[0172] In one embodiment, the instructions may cause the electronic device (101; 200) to define a second subset of instructions executable in the second playback section of the set of instructions.

[0173] According to one embodiment, the instructions may cause the electronic device (101; 200) to receive a first audio input input through the microphone (210) while playing the video.

[0174] In one embodiment, the instructions may cause the electronic device (101; 200) to convert the first speech input into first text.

[0175] According to one embodiment, the instructions may cause the electronic device (101; 200) to identify a first command corresponding to the first text.

[0176] According to one embodiment, the instructions may cause the electronic device (101; 200) to execute an operation corresponding to the first command when the first voice input is received during playback of the first playback section of the video and the first command is included in the first subset.

[0177] According to one embodiment, the plurality of playback segments, including the first playback segment and the second playback segment, may be defined based on a specified ratio of the total playback time of the video.

[0178] In one embodiment, the instructions may cause the electronic device (101; 200) to control not to execute an action corresponding to the first text based on the first voice input being received during playback of the second playback section of the video and the first command not being included in the second subset.

[0179] In one embodiment, the first playback segment may correspond to a starting segment of the video.

[0180] In one embodiment, the first subset of commands corresponding to the start segment may include at least one of a command for changing the sound mode and a command for changing the playback point to the next segment.

[0181] In one embodiment, the second playback segment may correspond to the last segment of the video.

[0182] According to one embodiment, a second subset of the instructions corresponding to the last section comprises:

[0183] It may include at least one of a command for playing the next video and a command corresponding to an evaluation of the video.

[0184] In one embodiment, the plurality of playback segments may further include a highlight content segment.

[0185] In one embodiment, the instructions may cause the electronic device (101; 200) to define a third subset of instructions executable in the highlighted content section, among the set of instructions.

[0186] In one embodiment, a third subset of commands corresponding to the highlighted content section may include a command for changing the playback point in time to a specified point in time.

[0187] According to one embodiment, the starting section among the plurality of playback sections may be a playback section from the starting point of the video to a point in time when a first time has elapsed from the starting point of the video.

[0188] In one embodiment, the instructions may cause the electronic device (101; 200) to determine the first time based on a first ratio to the total playback time of the video.

[0189] According to one embodiment, the last section of the plurality of playback sections may be a playback section from a point in time before a second time from the end point of the video to the end point of the video.

[0190] In one embodiment, the instructions may cause the electronic device (101; 200) to determine the second time based on a second ratio to the total playback time of the video.

[0191] According to one embodiment, the instructions may cause the electronic device (101; 200) to obtain metadata of the image based on at least one of image information or audio information of the image.

[0192] In one embodiment, the instructions may cause the electronic device (101; 200) to update a first subset of the commands based on the acquired metadata.

[0193] According to one embodiment, the metadata may further include tag information related to the content of the video and a playback point corresponding to the tag information.

[0194] According to one embodiment, the instructions may cause the electronic device (101; 200) to obtain tag information related to the content of the video and a playback point corresponding to the tag information based on image information of the video.

[0195] In one embodiment, the instructions may cause the electronic device (101; 200) to update a first subset of the commands based on tag information associated with the acquired content and a playback point in time corresponding to the tag information.

[0196] According to one embodiment, the metadata may further include keywords related to the content of the video and a playback point corresponding to the keywords.

[0197] The above instructions may cause the electronic device (101; 200) to obtain keywords related to the content of the image and playback points corresponding to the keywords based on the audio information of the image.

[0198] In one embodiment, the instructions may cause the electronic device (101; 200) to update a first subset of the commands based on the acquired keyword and a playback point corresponding to the acquired keyword.

[0199] In one embodiment, the instructions may cause the electronic device (101; 200) to determine whether the first voice input corresponds to a wake up sound.

[0200] In one embodiment, the instructions may cause the electronic device (101; 200) to identify a first command corresponding to the converted first text based on determining that the first voice input does not correspond to a wake-up sound.

[0201] A method of operating an electronic device (101; 200) according to one embodiment of the present disclosure may include an operation of playing an image.

[0202] According to one embodiment, the method may include defining a plurality of playback segments associated with the video, the plurality of playback segments including a first playback segment and a second playback segment.

[0203] In one embodiment, the method may include defining a first subset of instructions, among a set of instructions, that are executable in the first playback section.

[0204] In one embodiment, the method may include defining a second subset of instructions executable in the second playback section among the set of instructions.

[0205] According to one embodiment, the method may include receiving a first voice input input through a microphone (210) while playing the video.

[0206] In one embodiment, the method may include converting the first voice input into first text.

[0207] In one embodiment, the method may include an operation of identifying a first command corresponding to the first text.

[0208] According to one embodiment, the operating method may include an operation of executing an operation corresponding to the first command when the first voice input is received during playback of the first playback section of the video and the first command is included in the first subset.

[0209] According to one embodiment, the plurality of playback segments, including the first playback segment and the second playback segment, may be defined based on a specified ratio of the total playback time of the video.

[0210] In one embodiment, the operating method may include an operation of controlling not to execute an operation corresponding to the first text based on the first voice input being received during playback of the second playback section of the video and the first command not being included in the second subset.

[0211] In one embodiment, the first playback segment may correspond to a starting segment of the video.

[0212] In one embodiment, the first subset of commands corresponding to the start segment may include at least one of a command for changing the sound mode and a command for changing the playback point to the next segment.

[0213] In one embodiment, the second playback segment may correspond to the last segment of the video.

[0214] In one embodiment, the second subset of commands corresponding to the last segment may include at least one of a command for playing a next video and a command corresponding to an evaluation of the video.

[0215] In one embodiment, the plurality of playback segments may further include a highlight content segment.

[0216] In one embodiment, the method may include defining a third subset of commands executable in the highlighted content section among the set of commands.

[0217] In one embodiment, a third subset of commands corresponding to the highlighted content section may include a command for changing the playback point in time to a specified point in time.

[0218] According to one embodiment, the starting section among the plurality of playback sections may be a playback section from the starting point of the video to a point in time when a first time has elapsed from the starting point of the video.

[0219] According to one embodiment, the method may include an operation of determining the first time based on a first ratio to the total playback time of the video.

[0220] In a storage medium (130) storing computer-readable instructions according to one embodiment of the present disclosure, the instructions, when executed by at least one processor (230) of an electronic device (101; 200), can cause the electronic device (101; 200) to play an image.

[0221] In one embodiment, the instructions may cause the electronic device (101; 200) to define a plurality of playback segments associated with the video, including a first playback segment and a second playback segment.

[0222] In one embodiment, the instructions may cause the electronic device (101; 200) to define a first subset of instructions, from among a set of instructions, that are executable in the first playback section.

[0223] In one embodiment, the instructions may cause the electronic device (101; 200) to define a second subset of instructions executable in the second playback section of the set of instructions.

[0224] According to one embodiment, the instructions may cause the electronic device (101; 200) to receive a first audio input input through a microphone (210) while playing the video.

[0225] In one embodiment, the instructions may cause the electronic device (101; 200) to convert the first speech input into first text.

[0226] According to one embodiment, the instructions may cause the electronic device (101; 200) to identify a first command corresponding to the first text.

[0227] According to one embodiment, the instructions may cause the electronic device (101; 200) to execute an operation corresponding to the first command when the first voice input is received during playback of the first playback section of the video and the first command is included in the first subset.

[0228] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains.

[0229] As used herein, the term "if" will be understood to mean "when, upon," "in response to determining," or "in response to detecting," depending on the context. Similarly, "if it is determined to," or "if a stated condition or event is detected," will be understood to mean, optionally, "upon determining," or "in response to determining," "upon detecting a stated condition or event," or "in response to detecting a stated condition or event."

[0230] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. A processing device (or processing circuit) may execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, a single processing device is sometimes described, but one of ordinary skill in the art will recognize that a processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0231] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0232] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include an app store that distributes applications, a site that supplies or distributes various software, and a recording or storage medium managed by a server.

[0233] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components such as the described systems, structures, devices, and circuits are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0234] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

[0235] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0236] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0237] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0238] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0239] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0240] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device (101; 200), Mike (210); At least one processor (230) including a processing circuit, and Includes a memory (220) for storing instructions, The instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Play the video, defining a plurality of playback segments related to the video, including a first playback segment and a second playback segment; Define a first subset of instructions that are executable in the first playback section among the instruction sets, Define a second subset of instructions executable in the second playback section among the above instruction sets, While playing the above video, a first voice input is received through the microphone (210), Converting the above first voice input into a first text, Check the first command corresponding to the first text above, An electronic device (101; 200), characterized in that it includes instructions set to cause an operation corresponding to the first command to be executed when the first voice input is received during playback of the first playback section of the video and the first command is included in the first subset.

2. In paragraph 1, An electronic device (101; 200), characterized in that the plurality of playback sections including the first playback section and the second playback section are defined based on a specified ratio to the total playback time of the video.

3. In paragraph 1 or 2, The instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: An electronic device (101; 200) characterized in that it further includes an instruction set to cause a control not to execute an action corresponding to the first text based on the first voice input being received during playback of the second playback section of the video and the first command not being included in the second subset.

4. In any one of paragraphs 1 to 3, The above first playback section corresponds to the starting section of the video, An electronic device (101; 200), characterized in that the first subset of the commands corresponding to the start section includes at least one of a command for changing the sound mode and a command for changing the playback time to the next section.

5. In any one of paragraphs 1 to 4, The above second playback section corresponds to the last section of the video, The second subset of the above commands corresponding to the last section is, An electronic device (101; 200), characterized in that it comprises at least one of a command for playing a next video and a command corresponding to an evaluation of the video.

6. In any one of paragraphs 1 to 5, The above multiple playback sections further include a highlight content section, The instructions further include instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to define a third subset of instructions executable in the highlighted content section from among the set of instructions, An electronic device (101; 200), characterized in that the third subset of commands corresponding to the above highlighted content section includes a command for changing the playback time to a specified time.

7. In any one of paragraphs 1 to 6, Among the plurality of playback sections, the starting section includes the starting point of the video and the playback section from the starting point of the video to the point at which the first time has elapsed, An electronic device (101; 200), characterized in that the instructions further include instructions that, when executed individually or collectively by the at least one processor, cause the electronic device to determine the first time based on a first ratio to the total playback time of the video.

8. In any one of paragraphs 1 to 7, The last section among the above multiple playback sections is a playback section from the end point of the video to a point before the second time and up to the end point of the video, An electronic device (101; 200), characterized in that the instructions further include instructions that, when executed individually or collectively by the at least one processor, cause the electronic device to determine the second time based on a second ratio to the total playback time of the video.

9. In any one of paragraphs 1 to 8, The instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtaining metadata of the video based on at least one of image information or audio information of the video, An electronic device (101; 200), characterized in that it further comprises instructions configured to cause a first subset of the commands to be updated based on the acquired metadata.

10. In any one of paragraphs 1 to 9, The above metadata further includes tag information related to the content of the video and a playback point corresponding to the tag information, The instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the image information of the above video, tag information related to the content of the above video and a playback point corresponding to the tag information are obtained, An electronic device (101; 200), characterized in that it further includes instructions set to cause a first subset of the command to be updated based on tag information related to the acquired content and a playback time corresponding to the tag information.

11. In any one of paragraphs 1 to 10, The above metadata further includes keywords related to the content of the video and a playback point corresponding to the keywords. The instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the audio information of the above video, a keyword related to the content of the above video and a playback point corresponding to the keyword are obtained, An electronic device (101; 200), characterized in that it further includes instructions set to cause a first subset of the commands to be updated based on the acquired keyword and the playback time corresponding to the acquired keyword.

12. In any one of paragraphs 1 to 11, The instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Verify that the above first voice input corresponds to a wake up sound, An electronic device (101; 200), characterized in that it further includes instructions set to cause a first command corresponding to the converted first text to be identified based on determining that the first voice input does not correspond to a wake-up sound.

13. In the operating method of an electronic device (101; 200), The action of playing a video; An operation of defining a plurality of playback segments related to the video, the plurality of playback segments including a first playback segment and a second playback segment; An operation of defining a first subset of instructions, among a set of instructions, executable in said first playback section; An operation of defining a second subset of instructions executable in the second playback section among the above set of instructions; An action of receiving a first voice input input through a microphone while playing the above video; An operation of converting the first voice input into a first text; An operation of confirming a first command corresponding to the first text, and An operating method of an electronic device, characterized in that it includes an operation of executing an operation corresponding to the first command when the first voice input is received during playback of the first playback section of the video and the first command is included in the first subset.

14. In paragraph 13, An operating method of an electronic device, characterized in that the plurality of playback sections including the first playback section and the second playback section are defined based on a specified ratio of the total playback time of the video.

15. In paragraph 13 or 14, A method of operating an electronic device, characterized in that the method further includes an operation of controlling not to execute an operation corresponding to the first text based on the first voice input being received while playing a second playback section of the video and the first command not being included in the second subset.

Citation Information

Patent Citations

  • Method and apparatus for user interface for multimedia content search

    KR1020140139859A

  • Method, system and non-transitory computer-readable recording medium for automatically improving ai model performance

    KR1020240000225A

  • Electronic device for displaying image and method for controlling the same

    KR1020240036432A

  • Glp-1 receptor agonists, pharmaceutical composition comprising the same and method for preparing the same

    KR102563111B1

  • Media Presentation Device with Voice Command Feature

    US20230217079A1