Electronic device and method for editing video data in electronic device
By integrating AI models to interpret user commands in wearable devices, the solution addresses the challenge of inefficient data capture and editing, enhancing user interaction and functionality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-09-08
- Publication Date
- 2026-05-21
AI Technical Summary
Existing wearable electronic devices lack efficient methods for capturing and editing video and audio data based on user gestures, voice commands, and gaze inputs, leading to suboptimal user interaction and functionality.
Incorporating AI models to identify and execute user commands, such as gestures, voice instructions, and gaze inputs, to control video and audio capture and editing functions, allowing for real-time processing and deletion of unwanted data.
Enhances user interaction by enabling precise control over video and audio capture and editing, improving user experience and functionality in wearable devices.
Smart Images

Figure KR2025013878_21052026_PF_FP_ABST
Abstract
Description
Electronic devices and how to edit image data from electronic devices
[0001] The present disclosure relates to an electronic device and a method for editing image data in an electronic device.
[0002] Various services and additional functions provided through wearable electronic devices, such as AR glasses (augmented reality glasses), VST (video see-through) devices, and HMD (head-mount display) devices, are gradually increasing. To enhance the utility value of these electronic devices and satisfy the needs of diverse users, telecommunications service providers or electronic device manufacturers are competitively developing devices to offer various functions and differentiate themselves from competitors. Accordingly, the various functions provided through wearable electronic devices are also becoming increasingly sophisticated.
[0003] The information described above may be provided as related art for the purpose of aiding understanding of this document. None of the above is to be claimed as prior art related to this document, nor can it be used to determine prior art.
[0004] An electronic device according to one embodiment may include at least one camera, a microphone, at least one processor, and a memory for storing instructions. When the instructions according to one embodiment are executed individually or collectively by the at least one processor, the wearable electronic device may start capturing video data and audio data through the camera and the microphone, respectively. When the instructions according to one embodiment are executed individually or collectively by the at least one processor, the wearable electronic device may, while capturing the video data and audio data, if it identifies at least one instruction among a gesture instruction included in the video data or a designated user's voice instruction included in the audio data using an AI model, perform a function corresponding to the at least one instruction and store information about a first time when the at least one instruction was identified in the memory. When the commands according to one embodiment are executed individually or collectively by the at least one processor, the wearable electronic device confirms the end of shooting of the image data and the audio data, and stores a completed image file including the captured image data and the captured audio data, excluding the data corresponding to the at least one command confirmed at the first time, and the data corresponding to the at least one command can be deleted from the captured image data and the audio data based on information about the first time using the AI model.
[0005] An electronic device according to one embodiment may include a microphone, one or more cameras, at least one processor, and a memory for storing instructions. When the instructions according to one embodiment are executed individually or collectively by the at least one processor, the wearable electronic device may receive control instructions through at least one of the gaze, gesture, and voice of a user wearing the wearable electronic device while capturing or recording at least one of video or audio. When the instructions according to one embodiment are executed individually or collectively by the at least one processor, the wearable electronic device may generate video data or audio data reflecting the control instructions, and if at least one of the gesture and the voice is included in the video data or audio data, it may generate a complete video file or a complete audio file in which the gesture or the voice is removed from the video data or audio data using an AI model.
[0006] A method for editing video data in an electronic device according to one embodiment may include an operation of starting to capture video data and audio data through each of the camera and microphone of the wearable electronic device. The method according to one embodiment may include, while capturing the video data and audio data, an operation of performing a function corresponding to at least one command among a gesture command included in the video data or a designated user's voice command included in the audio data using an AI model, and storing information about a first time when the at least one command was confirmed in the memory of the electronic device. The method according to one embodiment may include an operation of saving a completed video file containing the captured video data and the captured audio data, excluding the data corresponding to the at least one command confirmed at the first time, when the end of capturing the video data and the audio data is confirmed, and the data corresponding to the at least one command may be deleted from the captured video data and the audio data based on information about the first time using the AI model.
[0007] In a non-volatile storage medium storing commands according to one embodiment, the commands are configured to cause the electronic device to perform at least one operation when executed by the electronic device, and may include an operation of starting to capture video data and audio data through the camera and microphone of the electronic device, respectively. The at least one operation according to one embodiment may include, while capturing the video data and audio data, if at least one command among a gesture command included in the video data or a designated user's voice command included in the audio data is identified using an AI model, performing a function corresponding to the at least one command and storing information regarding a first time at which the at least one command was identified in the memory of the electronic device. The at least one operation according to one embodiment may include, if the end of capturing the video data and audio data is confirmed, storing a completed video file containing the captured video data and the captured audio data, excluding the data corresponding to the at least one command identified at the first time, and the data corresponding to the at least one command may be deleted from the captured video data and the audio data based on information regarding the first time using the AI model.
[0008] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.
[0009] FIGS. 2a and 2b are drawings showing the front and rear of a wearable electronic device according to one embodiment.
[0010] FIG. 3a is a block diagram of a wearable electronic device according to one embodiment.
[0011] FIG. 3b is a block diagram illustrating the configuration of a processor and an AI model according to one embodiment.
[0012] FIG. 3c is a diagram illustrating the operation of editing image data in an external electronic device connected to a wearable electronic device according to one embodiment.
[0013] FIG. 3d is a diagram illustrating the operation of generating editing information in a wearable electronic device according to one embodiment.
[0014] FIGS. 4a, FIGS. 4b, and FIGS. 4c are drawings for explaining the operation of gesture commands and voice commands in a wearable electronic device according to one embodiment.
[0015] FIGS. 5A, FIGS. 5B, FIGS. 5C, FIGS. 5D, FIGS. 5E, and FIGS. 5F are drawings for explaining the operation of gesture commands and voice commands in a wearable electronic device according to one embodiment.
[0016] FIGS. 6a and 6b are drawings for explaining the operation of deleting a first command, which is a combination of a gesture command and a voice command, in a wearable electronic device according to one embodiment.
[0017] FIG. 7 is a diagram illustrating the deletion operation of a gesture command in a wearable electronic device according to one embodiment.
[0018] FIGS. 8A and 8B are drawings for explaining the operation of deleting voice commands in a wearable electronic device according to one embodiment.
[0019] FIG. 9 is a diagram illustrating the operation of executing a gesture command in a wearable electronic device according to one embodiment.
[0020] FIG. 10 is a diagram illustrating the operation of confirming a gesture command in a wearable electronic device according to one embodiment.
[0021] FIGS. 11a, FIGS. 11b, and FIGS. 11c are drawings for illustrating the generation of a plurality of new frames containing a new image in a wearable electronic device according to one embodiment.
[0022] FIGS. 12a, FIGS. 12b, FIGS. 12c, FIGS. 12d, and FIGS. 12e are drawings for explaining the operation of editing image data in an electronic device connected to a wearable electronic device according to one embodiment.
[0023] FIGS. 13a, FIGS. 13b, and FIGS. 13c are drawings for illustrating a viewfinder in a wearable electronic device according to one embodiment.
[0024] FIGS. 14a, FIGS. 14b, and FIGS. 14c are drawings for illustrating an AI model in an electronic device according to one embodiment.
[0025] FIG. 15 is a flowchart illustrating the operation of editing video data at the time of completion of shooting video data and audio data in a wearable electronic device according to one embodiment.
[0026] FIG. 16 is a flowchart illustrating the operation of editing video data while capturing video data and audio data in a wearable electronic device according to one embodiment.
[0027] FIG. 17 is a flowchart illustrating the operation of editing gesture commands in image data in a wearable electronic device according to one embodiment.
[0028] FIG. 18 is a flowchart illustrating the operation of editing gesture commands in image data in a wearable electronic device according to one embodiment.
[0029] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to one embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through the server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0030] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0031] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0032] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0033] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0034] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0035] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0036] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0037] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0038] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0039] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0040] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0041] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0042] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0043] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0044] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0045] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0046] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0047] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0048] According to one embodiment, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0049] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0050] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0051] FIGS. 2a and 2b are drawings showing the front and rear of a wearable electronic device according to one embodiment.
[0052] Referring to FIG. 2a and FIG. 2b, in one embodiment, camera modules (211, 212, 213, 214, 215, 216) and / or a depth sensor (217) for acquiring information related to the surrounding environment of the wearable electronic device (200) may be disposed on a first surface (210) of the housing (e.g., the front of the wearable electronic device (200).
[0053] In one embodiment, camera modules (211, 212) can acquire images related to the surrounding environment of the wearable electronic device (200).
[0054] In one embodiment, camera modules (213, 214, 215, 216) can acquire images while the wearable electronic device (200) is worn by a user. The camera modules (213, 214, 215, 216) can be used for hand detection, tracking, and user gesture (e.g., hand movements) recognition. The camera modules (213, 214, 215, 216) can be used for 3DoF, 6DoF head tracking, location (space, environment) recognition, and / or movement recognition. In one embodiment, camera modules (211, 212) may be used for hand detection and tracking and user gestures.
[0055] In one embodiment, the depth sensor (217) may be configured to transmit a signal and receive a signal reflected from a subject, and may be used for determining the distance to an object, such as time of flight (TOF). In place of or additionally to the depth sensor (217), camera modules (213, 214, 215, 216) may determine the distance to an object.
[0056] According to one embodiment, a face recognition camera module (225, 226) and / or a display (221) (and / or a lens) may be disposed on a second surface (220) of the housing (e.g., the rear of the wearable electronic device (200)).
[0057] In one embodiment, a face recognition camera module (225, 226) adjacent to the display is used for recognizing the user's face or can recognize and / or track both of the user's eyes.
[0058] In one embodiment, the display (221) (and / or lens) may be disposed on a second surface (220) of the wearable electronic device (200). In one embodiment, the wearable electronic device (200) may not include camera modules (215, 216) among a plurality of camera modules (213, 214, 215, 216).
[0059] As described above, according to one embodiment, the wearable electronic device (200) may have a form factor for being worn on a user's head. The wearable electronic device (200) may further include a strap and / or a wearing member for being secured on a part of the user's body. The wearable electronic device (200) may provide a user experience based on augmented reality, virtual reality, and / or mixed reality while being worn on the user's head.
[0060] FIG. 3a is a block diagram of a wearable electronic device according to one embodiment, FIG. 3b is a block diagram for explaining the configuration of a processor and an AI model according to one embodiment, FIG. 3c is a diagram for explaining the operation of editing image data in an external electronic device connected to a wearable electronic device according to one embodiment, and FIG. 3d is a diagram for explaining the operation of generating editing information in a wearable electronic device according to one embodiment.
[0061] Referring to FIGS. 3A, 3B, 3C, and 3D, according to one embodiment, the wearable electronic device (301) may be implemented as a wearable electronic device that can be worn on a user's body (e.g., head). For example, the wearable electronic device (301) may be implemented as an augmented reality glass (AR glass), a video see-through (VST) device, or a head-mount display (HMD) device. However, this is an example, and the wearable electronic device (301) may be implemented as various devices. According to one embodiment, in addition to the wearable electronic device, it may be similarly implemented as an electronic device that can be fixed to a user's body or clothing worn by the user (e.g., AI pin) or an electronic device that can be inserted into a pocket by changing its shape (e.g., foldable electronic device).
[0062] Referring to FIGS. 3a and 3b, the wearable electronic device (301) may include a processor (320), a camera (380), a memory (330), a display (360), a microphone (350), a speaker (355), and a communication circuit (390).
[0063] According to one embodiment, the processor (320) can perform overall control operations of the wearable electronic device (301). According to one embodiment, the processor (320) can execute software (e.g., the program (140) of FIG. 1) to control at least one other component (e.g., a hardware or software component) of the electronic device (301) connected to the processor (320), and can perform data processing or operations based on instructions. According to one embodiment, the instructions may include instructions composed of machine language that can be processed by the electronic device (301) or the processor (320). For example, the instructions may include instructions corresponding to operation instructions used in the program.
[0064] According to one embodiment, the processor (320) can start capturing video data and audio data through the camera (380) and the microphone (355), respectively, while the wearable electronic device (301) is worn by a part of the user's body.
[0065] According to one embodiment, the processor (320) can capture audio data through a microphone (355) while capturing image data through a first camera (381) (e.g., camera (211) of FIG. 2a) and a second camera (383) (e.g., second camera (212) of FIG. 2a) which can acquire images related to the surrounding environment of the wearable electronic device (301) among the cameras (380).
[0066] According to one embodiment, the processor (320) can capture audio data while capturing image data through one of a first camera (381) (e.g., camera (211) of FIG. 2a) and a second camera (383) (e.g., second camera (212) of FIG. 2a) that can acquire images related to the surrounding environment.
[0067] According to one embodiment, the capture of video data (video) and audio data (audio) may represent the operation of acquiring video data (video) and audio data (audio) and recording them in memory (330). According to one embodiment, the capture (recording) of audio data (audio) may represent the operation of acquiring audio and recording it in memory (330).
[0068] According to one embodiment, the processor (320) can start shooting video data and audio data when it confirms the selection of a shooting function button among at least one button displayed on the display based on user input.
[0069] According to one embodiment, when the processor (320) receives a voice command requesting a shooting function through a microphone (350), if it identifies the received voice command as a voice command of a designated user, it can interpret the voice command using a first AI model (331a) among a plurality of AI models (331) and start shooting video data and audio data corresponding to the interpreted voice command.
[0070] For example, when the processor (320) confirms a specified user’s utterance such as “hi Bixby, take a picture” or “hi Bixby, I want to record a video”, it generates the user’s utterance in the form of a prompt and inputs it into the first AI model (331a). When the first AI model (331a) interprets the content of the prompt and generates and outputs a command to perform a function corresponding to the content of the prompt, the processor (320) can start capturing video data and audio data corresponding to the command.
[0071] According to one embodiment, the first AI model (331a) is an AI model capable of analyzing text corresponding to audio data received through a microphone (350), and may include a Large Language Model (LLM).
[0072] According to one embodiment, the processor (320) can start the recording function of audio data when it confirms the selection of a recording function button among at least one button displayed through a display based on user input.
[0073] According to one embodiment, when the processor (320) receives a voice command requesting a recording function through a microphone (350), if it identifies the received voice command as a voice command of a designated user, it can interpret the voice command using a first AI model (331a) among a plurality of AI models (331) and start a recording function of audio data corresponding to the interpreted voice command.
[0074] According to one embodiment, when the processor (320) confirms a voice command or gesture command from a user requesting the execution of an editing function for video or recording using an AI model, it can notify a user wearing a wearable electronic device (301) of the execution of an editing-related function to delete a gesture command confirmed during the shooting of video data and audio data or during the recording of audio data, a voice command of a designated user, a first command combining a gesture command and a voice command, or a second command combining an eye tracking command and a voice command of a designated user, through an audio notification, a voice notification, or a UI displayed in a part area of the display (360).
[0075] According to one embodiment, a gesture command may represent a designated gesture (or an object representing the designated gesture) for executing a command among the user's gestures.
[0076] According to one embodiment, the voice command may represent a designated voice for executing the command among the voices following the user's speech.
[0077] According to one embodiment, the gaze command may represent the gaze of a user who stares at a target for command execution for a specified period of time or longer.
[0078] According to one embodiment, the processor (320) receives a control command through at least one of the gaze, gesture, and / or voice of a user wearing a wearable electronic device while capturing video data and audio data or recording only audio data, and can generate video data or audio data by reflecting the control command.
[0079] According to one embodiment, when the processor (320) detects the reception of a gesture command, a voice command of a designated user, a first command combining a gesture command and a voice command, or a second command combining a gaze command and a voice command based on at least one of the first camera (381) or the second camera (383) and a microphone (350) while capturing an image through the first camera (381) and the second camera (383), the processor (320) can generate image data while performing a function corresponding to the detected command.
[0080] According to one embodiment, the image data may include video data and still image data (e.g., photographs).
[0081] According to one embodiment, when the processor (320) detects the reception of a designated user's voice command through the microphone (350) while recording audio through the microphone (350), it can generate audio data while performing a function corresponding to the detected voice command.
[0082] According to one embodiment, the control command may include or be used with the same meaning as a gesture command, a designated user's voice command, a first command combining a gesture command and a voice command, or a second command including a gaze command and a voice command.
[0083] According to one embodiment, a gesture command (gesture) may include various gesture commands indicating shooting start, focus movement, shooting end, shooting pause, zoom in, zoom out, viewfinder ratio setting (e.g., 9:16 and / or 1:1), area cropping, object erasing, object transformation (e.g., person → tree), color transformation of an object, etc., and the gesture command may be specified by the user.
[0084] For example, when the processor (320) detects a clenched fist gesture included in the image data received through at least one of the first camera (381) or the second camera (383) while capturing image data and audio data, it can use the second AI model (331b) to identify the clenched fist gesture as a gesture command for ending the capture and perform the capture end function.
[0085] For example, when the processor (320) detects a hand gesture with the palm spread wide included in the image data received through at least one of the first camera (381) or the second camera (383) while capturing image data and audio data, it can use the second AI model (331b) to identify the hand gesture with the palm spread wide as a gesture command for pausing the shooting and perform a shooting pause function.
[0086] According to one embodiment, the second AI model (331b) is an AI model capable of generating text by analyzing (interpreting) image data (still image data and / or video data) received through a camera (380), and may include a large vision model (LVM).
[0087] According to one embodiment, voice commands (voice) may include various voice commands such as starting shooting, ending shooting, pausing shooting, zooming in, zooming out, and / or setting the viewfinder ratio indicating the shooting area (e.g., 9:16 and / or 1:1), area cropping, and / or object conversion (e.g., person → tree), color conversion of the object, etc., and the voice commands may be specified by the user.
[0088] For example, when the processor (320) receives a voice command such as "stop shooting" or a natural language such as "stop shooting" or "stop saving now" through the microphone (350) while shooting video data and audio data or while recording only audio data, it can use the first AI model (331a) to confirm that it is a voice command for stopping shooting and stop shooting video data and audio data or stop recording audio data.
[0089] For example, when the processor (320) receives a natural language such as “zoom in on the flower” through the microphone (350) while capturing video data and audio data, it can use the first AI model (331a) to confirm that it is a voice command to zoom in on the flower among the subjects being captured, and zoom in on the flower.
[0090] According to one embodiment, a first command combining a gesture command and a designated user's voice command may represent a command combining a gesture command received through at least one camera among the first camera (381) or the second camera (383) and a designated user's voice command received through a microphone (350).
[0091] For example, the processor (320), while capturing video data and audio data, can confirm a hand gesture with the palm open included in the video data received through at least one of the first camera (381) or the second camera (383), and simultaneously or within a specified time, confirm the reception of a specified user's utterance such as "stop shooting" through the microphone (350), and can use the third model (331c) to confirm the hand gesture with the palm open as a command to stop shooting and the user utterance such as "stop shooting" as a voice command to stop shooting, thereby stopping the shooting of the video data and audio data.
[0092] For example, the processor (320), while capturing video data and audio data, checks for a gesture pointing to an object included in the video data received through at least one of the first camera (381) or the second camera (383), and simultaneously or within a specified time checks for the reception of a specified user's utterance, such as "delete it," through the microphone (350), and can use the third model (331c) to identify the gesture pointing to the object as a gesture command pointing to the target that performed the function, and identify the user utterance, such as "delete it," as a voice command to delete the object pointed to by the gesture command, thereby performing the function of deleting an object from the video data being captured.
[0093] According to one embodiment, the third AI model (331c) is an AI model capable of analyzing image data, audio data, and / or text data, and may include a large multimodal model (LMM) such as a vision-language model (VLM).
[0094] According to one embodiment, when the processor (320) confirms a second command combining a gaze command and a voice command based on the gaze tracking camera (385) and microphone (350) while capturing video data and audio data, it performs a function corresponding to the confirmed second command and can delete the second command from the audio data captured during the capture or when confirming the end of the capture.
[0095] According to one embodiment, when the processor (320) confirms a gesture command, a voice command of a designated user, a first command including a gesture command and a voice command of a designated user, or a second command combining a gaze command and a voice command, it performs a function corresponding to the confirmed command and can delete the gesture command confirmed in the video data and the voice command confirmed in the audio data when confirming the end of shooting or during recording.
[0096] According to one embodiment, the processor (320) identifies a command that combines at least one or more of a gesture command, a designated user's voice command, or a gaze command, performs a function corresponding to the identified command, and can delete the gesture command identified in the video data and the voice command identified in the audio data when confirming the end of shooting or during shooting.
[0097] According to one embodiment, when the processor (320) checks a designated voice command received through the microphone (350) while checking an object included in the video data that the user is looking at through the eye-tracking camera (385) while shooting video data and audio data, it can use the first AI model (331a) to perform a function corresponding to the designated user's voice command on the object that the user is looking at, and when checking the end of shooting or while shooting, it can delete the designated user's voice command from the audio data.
[0098] For example, the processor (320), while capturing video data and audio data, while confirming the user’s gaze on an object included in the video captured through the eye-tracking camera (385), when confirming a designated user’s voice command of “delete” received through the microphone (350), can use the first AI model (331a) to perform a deletion function for the object included in the video data, and can delete the designated user’s voice command of “delete” from the audio data when confirming the end of the capture or while capturing.
[0099] According to one embodiment, when a processor (320) receives a first command combined with a gesture command pointing to an object and a voice command requesting to take a picture of the object pointed to by the gesture command (e.g., "I want to take a picture of that") while taking still image data, the processor takes a first still image data at the time of confirming the first command, takes a plurality of still image data from the time of taking the first still image data until a specified time (e.g., the time when the gesture command disappears), and generates a completed image data in which the gesture command is deleted from the first still image data based on the gesture command included in the plurality of still image data.
[0100] According to one embodiment, when the processor (320) receives a first command combining a gesture command and a voice command while capturing still image data, it may execute a deletion process for the gesture command and not execute a deletion process for the unnecessary voice command.
[0101] According to one embodiment, the processor (320) can analyze and verify the object pointed to by the user by performing a calibration operation to compensate for the difference in angle between the user's finger and the camera (380) when receiving a gesture command pointing to an object with the user's finger while capturing video data and audio data, because the scene displayed through the display (360) may not match the actual scene depending on the position (direction and angle) of the camera (380).
[0102] According to one embodiment, when the processor (320) receives a gesture command pointing to an object with a user's finger while capturing video data and audio data, it can analyze and verify the object pointed to by the user's finger by taking into account the error between the position of the user's finger pointing to the object and the position of the user's finger displayed through the display (360).
[0103] According to one embodiment, the processor (320) may perform a calibration operation to compensate for the angle difference between the user's finger and the camera (380) whenever the wearable electronic device is worn by the user or at specified intervals while the wearable electronic device is worn by the user.
[0104] According to one embodiment, the processor (320) can visually or audibly inform the user of the object that is the target of the control command before reflecting the control command while capturing video data and audio data.
[0105] According to one embodiment, when the processor (320) confirms the reception of a first command that combines a gesture command for pointing to an object with a user's finger and a voice command for performing a function on the object while capturing video data and audio data, it may output audio data that audibly indicates information about the object pointed to by the user (e.g., name and / or type, etc.) through the speaker (355) of the electronic device to confirm the object pointed to by the user, and the audio data indicating information about the object (e.g., name and / or type, etc.) may be deleted from the completed video data.
[0106] According to one embodiment, when the processor (320) confirms the reception of a first command that combines a gesture command for pointing to an object with the user's finger and a voice command for performing a function on the object while capturing video data and audio data, it may display a preview screen including the object that the user is pointing to visually through a display (360), or display a preview screen including the object that the user is pointing to visually through a display of an external electronic device connected to the wearable electronic device (301).
[0107] According to one embodiment, the processor (320), while capturing video data and audio data, after applying a control command to an object that is the target of the control command, generates edited video data (e.g., additional information) that can indicate that the control command has been applied to the object or edited video data (e.g., additional information) that can display the content of the control command applied to the video data, and can display the edited video data (e.g., additional information) on the video data through the display (360).
[0108] According to one embodiment, the processor (320) can generate edited video data that indicates that a deletion function has been performed on an object, by displaying the deleted object in a color (e.g., blurry color) that is distinct from the undeleted object (e.g., vivid color) when deleting an object based on confirmation of a gesture command or voice command to delete an object while capturing video data and audio data.
[0109] According to one embodiment, the processor (320) may, while capturing video data and audio data, add a mark to each of at least one object on which a function corresponding to a gesture command, a voice command, a first command combining a gesture command and a voice command, or a second command combining an eye tracking command and a designated user's voice command is executed to display the content of the executed function, or generate edited video data (e.g., additional information) that can display the content of the executed function (e.g., zoom-in or zoom-out ratio, aspect ratio and / or cropping status, etc.) in a certain area of the display (360).
[0110] For example, the processor (320) can generate edited video data that indicates whether the focus is on and / or whether the color is changed by adding a mark on the object that is the focus target and the object that is the color change target.
[0111] According to one embodiment, the processor (320) may wait for the reception of at least one command among a gesture command or a voice command based on the user's first input.
[0112] According to one embodiment, the processor (320) may recognize a specified gesture (e.g., a clapping gesture) as a first input from the user and wait for the reception of at least one command among a gesture command or a voice command.
[0113] According to one embodiment, the processor (320) can determine, through the eye-tracking camera (385) among the cameras (380), that the user is gazing at a button (e.g., an assistant icon) for waiting to receive at least one of a gesture command or a voice command displayed in a part area of the display (360) as a first input of the user, and can wait for the reception of the gesture command or voice command.
[0114] According to one embodiment, the processor (320) may recognize a voice trigger (e.g., "hi Bixby") as a first input from the user and wait for the reception of a gesture command or a voice command.
[0115] According to one embodiment, the processor (320), while capturing video data and audio data, can confirm a voice trigger corresponding to a first input of the user, and then delete the voice trigger corresponding to the first input of the user from the audio data captured during the capture when confirming the end of the capture or while capturing.
[0116] According to one embodiment, the processor (320) can check a gesture command included in a viewfinder that indicates an actual shooting area on the display (360).
[0117] According to one embodiment, the processor (320) can display preview image data received through a camera in a viewfinder that indicates a shooting area where a user wearing a wearable electronic device can recognize the shooting of image data.
[0118] According to one embodiment, the processor (320) can display preview image data received through the camera via the display (360) when a viewfinder is not provided.
[0119] According to one embodiment, the processor (320) may output visual feedback or sound to notify the user that a gesture command has been confirmed.
[0120] For example, the processor (320), while capturing video data and audio data, can use the second AI model (331b) to check a gesture command to adjust the viewfinder ratio received through at least one of the first camera (381) or the second camera (383), and then display the adjusted viewfinder ratio as a color or information to inform the user that a gesture command to adjust the viewfinder ratio has been checked.
[0121] According to one embodiment, the processor (320) may, while capturing video data and audio data through the first camera (381) and the second camera (383), perform a function corresponding to a gesture command, a voice command of a designated user, a first command combined with a gesture command and a voice command, or a second command combined with an eye tracking command and a voice command of a designated user, and while performing the function, store information about a first time in which a gesture command, a voice command of a designated user, a first command combined with a gesture command and a voice command, or a second command combined with an eye tracking command and a voice command of a designated user is confirmed as editing information in memory (330), and when the end of capturing video data and audio data is confirmed, perform an editing operation to delete the gesture command, the voice command of a designated user, the first command combined with a gesture command and a voice command of a designated user, or the second command including an eye tracking command and a voice command of a designated user from the video data based on the editing information.
[0122] According to one embodiment, the first time may represent the time during which a gesture command is displayed in the captured video data while capturing video data and audio data for editing the gesture command.
[0123] According to one embodiment, the first time may represent the time during which the voice command of a designated user is output while video data and audio data are being captured for editing the voice command of a designated user.
[0124] According to one embodiment, the first time may represent a time including the time when the gesture command is displayed in the captured video data and the time when the voice command of the specified user is output while capturing video data and audio data for editing a first command in which the gesture command and the voice command of the specified user are combined.
[0125] According to one embodiment, when the processor (320) checks a voice command of a designated user in audio data received through the microphone (350) while capturing video data and audio data using the first AI model (331a), the time from the start time of output of the checked voice command to the end time of output can be checked as the first time of checking the voice command.
[0126] According to one embodiment, when the processor (320) checks a gesture command included in the captured image data while capturing image data and audio data using the second AI model (331b), the time from the start time of display of the checked gesture command to the end time of display can be identified as the first time of checking the gesture command.
[0127] According to one embodiment, when the processor (320) confirms the end of shooting of video data and audio data, it can use the second AI model (331b) to identify gesture commands included in at least one frame corresponding to the first time in the video data based on editing information, and delete gesture commands included in at least one frame.
[0128] According to one embodiment, the processor (320) can correct a first region (an image corresponding to a gesture command) corresponding to a deleted gesture command among at least one frame based on a previous frame or a next frame of at least one frame.
[0129] According to one embodiment, the processor (320) can use a second AI model (331b) to compare the size of a first area (an image corresponding to a gesture command) corresponding to a gesture command with a first threshold value for determining the ratio of the gesture command to a frame, and based on the comparison result, if the size of the first area is determined to be greater than or equal to the first threshold value, delete at least one frame containing the gesture command, and if the size of the first area is determined to be less than or equal to the first threshold value, delete the gesture command from at least one frame.
[0130] According to one embodiment, the processor (320) can use the second AI model (331b) to determine, based on the comparison result, that the size of the first region is greater than or equal to the first threshold value (e.g., 60%), and when deleting at least one frame containing a gesture command, if it is determined that at least one frame containing a gesture command is included in a frame excluding the middle frame or the last frame of the video data, generate a new frame based on at least one of the previous frame or the next frame of the at least one frame, and insert the new frame at the location of the deleted at least one frame, thereby generating a complete video file that is naturally connected without interruption between frames.
[0131] According to one embodiment, the processor (320) can, by using the second AI model (331b), if it confirms that a control command is included in at least one frame other than the last frame among a plurality of frames of image data, delete the at least one frame containing the gesture command, create a new at least one frame based on at least one of the previous frame or the next frame of the at least one frame, and insert the new at least one frame at the location of the deleted at least one frame among the plurality of frames.
[0132] According to one embodiment, the processor (320) can delete at least one frame containing a gesture command without creating a new frame by using the second AI model (331b) to confirm, based on the comparison result, that the size of the first region is greater than or equal to the first threshold value, and when deleting at least one frame containing a gesture command, if it is confirmed that at least one frame containing a gesture command is included in the last frame of the image data.
[0133] According to one embodiment, the processor (320) can use the second AI model (331b) to confirm the inclusion of a gesture command in at least one frame, including the last frame among a plurality of frames of image data, delete at least one frame from the plurality of frames, and generate a complete image file that does not fill the position corresponding to the at least one frame deleted from the plurality of frames.
[0134] According to one embodiment, the processor (320) can delete the gesture command included in the at least one frame when using the second AI model (331b) to confirm that at least one frame including a gesture command is included in the last frame of the image data, and when confirming that the focus function for the object included in the image data is maintained.
[0135] According to one embodiment, the processor (320) can use the second AI model (331b) to confirm the inclusion of a gesture command in at least one frame, including the last frame among a plurality of frames of image data, delete at least one frame from the plurality of frames, create at least one new frame based on the frame prior to at least one frame or the background image of the plurality of frames, and create a completed image file in which at least one new frame is inserted at the location of the deleted at least one frame.
[0136] According to one embodiment, the processor (320) can delete the gesture command included in the at least one frame by using the second AI model (331b) when confirming that at least one frame including a gesture command is included in the last frame of the image data, and when the object included in the image data is a person, confirming that the person's designated facial expression (e.g., smiling expression) is maintained.
[0137] According to one embodiment, the processor (320) can delete the gesture command included in at least one frame by using the second AI model (331b) when it is confirmed that at least one frame including a gesture command is included in the last frame of the image data, and when it is confirmed that an object included in the image data is located at the center of the image data or is continuously included in the image data.
[0138] According to one embodiment, the processor (320) can use a second AI model (331b) to check the rate of change in the scene (e.g., the value of pixels included in the frame) between at least one frame containing a gesture and the previous frame of at least one frame, compare the rate of change in the scene with a second threshold, and if the rate of change in the scene is found to be greater than or equal to the second threshold based on the comparison result, delete the gesture command included in at least one frame, and if the rate of change in the scene is found to be less than or equal to the second threshold based on the comparison result, delete at least one frame containing the gesture command.
[0139] According to one embodiment, the processor (320) can delete a gesture command included in at least one frame using a second AI model (331b) if the gesture command includes a new movement or a new object that was not identified in the previous frame.
[0140] According to one embodiment, the processor (320) can delete the gesture command included in at least one frame by using the second AI model (331b) to identify the object included in at least one frame containing the gesture command as the main object based on the object included in all frames of the image data.
[0141] According to one embodiment, when the processor (320) confirms the end of shooting of video data and audio data, it can delete the voice command of a designated user confirmed at the first time using the first AI model (331a).
[0142] According to one embodiment, the processor (320) can use the first AI model (331a) to check audio data confirmed at the first time and separate and delete the voice command of a designated user included in the confirmed audio data.
[0143] According to one embodiment, the processor (320) can correct the audio data confirmed at the first time in which the voice command of a designated user is deleted or generate new audio data when the voice command of a designated user included in the audio data confirmed at the first time is deleted using the first AI model (331a) and the previous audio data and the next audio data are unnatural based on the time of the audio data confirmed at the first time.
[0144] According to one embodiment, when the processor (320) confirms the end of shooting of video data and audio data, it may store in memory (330) a completed video file or a completed audio file generated after performing an editing operation to delete a gesture command included in at least one frame corresponding to a first time in the video data, a voice command of a designated user confirmed at the first time, a first command combining a gesture command included in at least one frame corresponding to the first time and a voice command of a designated user confirmed at the first time, or a second command combining an eye tracking command confirmed at the first time and a voice command of a designated user.
[0145] According to one embodiment, when the processor (320) confirms the end of shooting of video data and audio data, it may store in memory (330) a completed video file or a completed audio file in which data corresponding to a gesture command, a designated user's voice command, a first command combining a gesture command and a designated user's voice command, or a second command including an eye tracking command and a designated user's voice command is excluded.
[0146] According to one embodiment, when the processor (320) confirms the end of shooting of video data and audio data, it performs an editing operation to generate a completed video file or a completed audio file, and can store at least one of an original video file in which control commands are not reflected based on user input, an edited video file generated based on edited video data, a completed video file and / or a completed audio file in memory (330).
[0147] According to one embodiment, the processor (320) may perform an editing operation to delete a gesture command, a designated user's voice command, a first command combined with a gesture command and a designated user's voice command, or a second command combined with an eye tracking command and a voice command, after or while performing a function corresponding to a gesture command, a designated user's voice command, a gesture command and a designated user's voice command, while performing the shooting of video data and audio data through the first camera (381) and the second camera (383).
[0148] According to one embodiment, the processor (320), while capturing video data and audio data, checks for a gesture command included in the captured video data received through at least one camera among the first camera (381) or the second camera (383) using the second AI model (331b), checks for at least one frame including the gesture command, and can delete the gesture command included in at least one frame.
[0149] According to one embodiment, the processor (320) can correct a first area (an image corresponding to a gesture command) corresponding to a deleted gesture command among at least one frame based on a previous frame or a next frame of at least one frame while capturing video data and audio data.
[0150] According to one embodiment, when the processor (320) checks a gesture command included in the captured image data while capturing image data and audio data, it can use a second AI model (331b) to compare the size of a first area corresponding to the gesture command and a first threshold value for checking the ratio of the gesture command to at least one frame.
[0151] According to one embodiment, the processor (320) can delete at least one frame containing a gesture command when it is determined that the size of the first area is greater than or equal to the first threshold value based on the comparison result, and can delete the gesture command from at least one frame when it is determined that the size of the first area is less than or equal to the first threshold value.
[0152] According to one embodiment, the processor (320), while capturing video data and audio data, uses a second AI model (331b) to determine that the size of a first region corresponding to a gesture command based on a comparison result is greater than or equal to a first threshold value (e.g., 60%), and when deleting at least one frame containing a gesture command, it can determine that at least one frame containing a gesture command is included in a frame excluding an intermediate frame or the last frame of the video data.
[0153] According to one embodiment, when the processor (320) confirms that at least one frame containing a gesture command is included in a frame excluding an intermediate frame or a last frame of the image data, it can generate a new frame based on at least one of the previous frame or the next frame of the at least one frame, and insert the new frame at the location of the deleted at least one frame so that the frames can be connected naturally without interruption.
[0154] According to one embodiment, the processor (320), while capturing video data and audio data, uses a second AI model (331b) to determine that the size of a first region corresponding to a gesture command is greater than or equal to a first threshold value based on a comparison result, and when deleting at least one frame containing a gesture command, it can determine that at least one frame containing a gesture command is included in the last frame of the video data.
[0155] According to one embodiment, if the processor (320) identifies at least one frame containing a gesture command as the last frame of the image data, it may delete at least one frame containing a gesture command without creating a new frame.
[0156] According to one embodiment, the processor (320), while capturing video data and audio data, can delete the gesture command included in at least one frame by using the second AI model (331b) when confirming that at least one frame including a gesture command is included in the last frame of the video data and confirming that the focus function for the object included in the video data is maintained.
[0157] According to one embodiment, the processor (320), while capturing video data and audio data, can delete the gesture command included in at least one frame by using the second AI model (331b) when confirming that at least one frame including a gesture command is included in the last frame of the video data, and when the object included in the video data is a person, confirming that the person's designated facial expression (e.g., smiling expression) is maintained.
[0158] According to one embodiment, the processor (320), while capturing video data and audio data, can delete the gesture command included in at least one frame by using the second AI model (331b) when confirming that at least one frame including a gesture command is included in the last frame of the video data, or when confirming that an object included in the video data is located at the center of the video data or is continuously included in the video data.
[0159] According to one embodiment, the processor (320), while capturing video data and audio data, uses a second AI model (331b) to check the rate of change of scene between at least one frame containing a gesture and the previous frame of at least one frame, compares the rate of change of scene with a second threshold value, and if the rate of change of scene is confirmed to be greater than or equal to the second threshold value based on the comparison result, deletes the gesture command included in at least one frame, and if the rate of change of scene is confirmed to be less than or equal to the second threshold value based on the comparison result, deletes at least one frame containing the gesture command.
[0160] According to one embodiment, the processor (320) may delete a gesture command included in at least one frame if, while capturing video data and audio data, a new movement or new object not identified in the previous frame is included in at least one frame using a second AI model (331b).
[0161] According to one embodiment, the processor (320), while capturing video data and audio data, can delete the gesture command included in at least one frame by using the second AI model (331b) to identify the object included in at least one frame containing the gesture command as the main object based on the object included in all frames of the captured video data.
[0162] According to one embodiment, the processor (320), while capturing video data and audio data, can delete the voice command of a designated user received through the microphone (350) using the first AI model (331a).
[0163] According to one embodiment, the processor (320), while capturing video data and audio data, uses the first AI model (331a) to check whether a designated user's voice command is included, and if it checks whether the audio data includes a designated user's voice command, it can separate and delete the designated user's voice command from the audio data.
[0164] According to one embodiment, the processor (320), while capturing video data and audio data, can use the first AI model (331a) to delete a designated user's voice command from audio data received through the microphone (350), and when the audio data received through the microphone (350) is unnatural based on the previous audio, correct the audio data from which the designated user's voice command has been deleted or generate new audio data.
[0165] According to one embodiment, the processor (320) can perform an editing operation while capturing video data and audio data, and generate a completed video file or a completed audio file and store it in memory (330).
[0166] According to one embodiment, the processor (320) performs an editing operation while capturing video data and audio data, and can store a completed video file or a completed audio file in memory (330) from which data corresponding to a gesture command, a voice command of a designated user, a first command combining a gesture command and a voice command of a designated user, or a second command including an eye tracking command and a voice command of a designated user is excluded.
[0167] According to one embodiment, when the processor (320) confirms the end of shooting of video data and audio data, it may store at least one of an original video file in which control commands are not reflected based on user input, an edited video file generated based on edited video data, a finished video file and / or a finished audio file in memory (330).
[0168] According to one embodiment, the processor (320) can generate a completed video file or a completed audio file by using an AI model (331) to delete (remove) at least some of the gesture commands from a plurality of frames containing gesture commands after the end of shooting video data and audio data or while shooting video data and audio data.
[0169] According to one embodiment, the processor (320) can use an AI model (331) to delete multiple frames containing gesture commands after the end of shooting video data and audio data or while shooting video data and audio data, and replace the locations of the deleted multiple frames containing gesture commands with new multiple frames containing newly generated images to create a completed video file or a completed audio file.
[0170] According to one embodiment, when deleting multiple frames containing gesture commands, if the multiple frames containing gesture commands are consecutive for a certain period, the processor (320) may use an AI model (331) to generate new multiple frames that add (include) a newly generated image corresponding to a specific audio signal (first audio signal) among audio other than voice included within the certain period. According to one embodiment, the processor (320) may delete multiple frames containing gesture commands and generate a completed image file by replacing the locations of the deleted multiple frames containing gesture commands with new multiple frames containing a newly generated image.
[0171] For example, the processor (320) can identify a gesture command for zooming in on the first object (e.g., person) among the first object (e.g., person) and the second object (e.g., car) included in the video data, and perform a zoom-in function on the first object (person). When the processor (320) deletes multiple frames containing a gesture command for zooming in on the first object (person), it can use an AI model (331) to generate new multiple frames containing a newly generated video based on an audio signal of the second object (e.g., car) recorded at a time corresponding to the multiple frames containing the gesture command. The processor (320) can generate a completed video file by replacing (inserting) the new multiple frames at the locations of the deleted multiple frames containing the gesture command.
[0172] According to one embodiment, the processor (320) may use an AI model (331) to generate new multiple frames containing a newly generated image based on a previous frame or a next frame of multiple frames containing a gesture command. According to one embodiment, when the processor (320) inserts (includes) new multiple frames containing a newly generated image at the location of multiple frames containing a deleted gesture command, the processor (320) may generate a completed image file or a completed audio file in which a specific audio signal (a first audio signal) unrelated to the newly generated image is deleted from the audio, excluding the voice recorded at the time corresponding to the multiple frames containing the gesture command.
[0173] For example, the processor (320) can identify a gesture command for zooming in on the first object (e.g., person) among the first object (e.g., person) and the second object (e.g., car) included in the image data, and perform a zoom-in function on the first object (person). When the processor (320) deletes multiple frames containing a gesture command for zooming in on the first object (person), it can use an AI model (331) to generate new multiple frames containing a newly generated image based on the previous frame or the next frame of the multiple frames containing the gesture command. When the processor (320) includes new multiple frames containing a newly generated image at the location of the multiple frames containing the deleted gesture command, if it confirms that the audio of the second object (e.g., car) recorded at the time corresponding to the multiple frames containing the gesture command is not related to the newly generated image, it can generate a completed image file or a completed audio file with the audio signal of the second object (e.g., car) deleted.
[0174] According to one embodiment, when deleting multiple frames containing gesture commands, the processor (320) may use an AI model (331) to generate new multiple frames containing newly generated images based on audio recorded at a time corresponding to the multiple frames containing gestures and information related to the multiple frames containing gestures (e.g., type, color and / or shape of an object included in the frame, etc.), and generate a completed image file containing new multiple frames at the locations of the multiple frames containing deleted gestures.
[0175] According to one embodiment, when deleting multiple frames containing gesture commands, the processor (320) can identify audio recorded at a time corresponding to the multiple frames containing gestures and information related to the multiple frames containing gestures (e.g., type, color and / or shape of an object included in the frame). According to one embodiment, when the processor (320) cannot generate new multiple frames containing a new image based on the identified audio and information related to the multiple frames containing gestures, the processor (320) can generate new multiple frames containing a base image. According to one embodiment, the processor (320) can generate a completed image file containing new multiple frames containing a base image at the location of the multiple frames containing deleted gesture commands.
[0176] According to one embodiment, when deleting multiple frames containing gesture commands, the processor (320) uses an AI model (331) to identify audio recorded at a time corresponding to the multiple frames containing gestures and information related to the multiple frames containing gestures (e.g., type, color and / or shape of an object included in the frame, etc.), and when it is not possible to generate new multiple frames containing new images based on the identified audio and information related to the multiple frames containing gestures, the control command can be deleted from multiple frames containing control commands.
[0177] According to one embodiment, when the processor (320) generates a completed video file or a completed audio file, if it confirms an unclear gesture command, an unclear voice command of a designated user, a first command including an unclear gesture command and a voice command of a designated user, or a second command combining an unclear eye tracking command and a voice command of a designated user while capturing video data and audio data, it may generate a completed video file or a completed audio file with the unclear command deleted based on the user's confirmation.
[0178] According to one embodiment, the processor (320) can identify an unclear gesture command, an unclear voice command, and / or an unclear eye tracking command when it cannot be determined by using the AI model (331). For example, the processor (320) can identify an unclear gesture command, an unclear voice command, and / or an unclear eye tracking command when it is unclear to use the AI model (331) the object corresponding to the target for executing the command indicated by the gesture command, voice command, or eye tracking command.
[0179] For example, the processor (320) can use an AI model (331) to identify a recognized voice based on a user's speech as an unclear voice command if it is not recognized as a perfect voice command.
[0180] According to one embodiment, when generating a completed video file or a completed audio file, the processor (320) analyzes video data using a third AI model (331c), and if it identifies audio data unrelated to the video data (e.g., voices of people other than the voice of a person included in the recorded audio data, or noise) based on the analyzed video data, it can generate a completed video file or a completed audio file by deleting the audio data unrelated to the video data based on user confirmation.
[0181] A memory (330) according to one embodiment may be implemented substantially identically or similarly to the memory (130) of FIG. 1.
[0182] According to one embodiment, an on-device AI model (331) may be stored in the memory (330), and the on-device AI model (331) may be configured in various ways, such as including a first AI model (331a), a second AI model (331b), and a third AI model (331c), or including at least one AI model among the first AI model (331a), the second AI model (331b), and the third AI model (331c), or including at least one of the first AI model (331a), the second AI model (331b), and the third AI model (331c) as one identical AI model, or including the first AI model (331a), the second AI model (331b), and the third AI model (331c) as one identical AI model.
[0183] At least some of the on-device AI models (331) according to one embodiment can be operated as external AI models.
[0184] An on-device AI model (331) according to one embodiment performs processing of sensitive information such as personal information, and an external AI model can perform processing of information excluding sensitive information such as personal information.
[0185] The on-device AI model (331) according to one embodiment is an AI model implemented in an electronic device (301) and can provide various functions without a network.
[0186] According to one embodiment, a plurality of AI models can be stored in the memory (330).
[0187] According to one embodiment, each of the plurality of AI models may be a model trained based on a specified type of learning algorithm, and may be an AI model implemented to receive various types of data (or content) as input, perform calculations, and output (or obtain) result data. According to one embodiment, the plurality of AI models may include a generative AI model.
[0188] According to one embodiment, a plurality of AI models may include a large vision model (LVM) capable of analyzing images, a large language model (LLM) capable of analyzing text and generating text or images from input text, and a large multimodal model (LMM) capable of analyzing video data, audio data, and / or text data and generating audio, images, and text from input audio data and / or text data.
[0189] A generative AI model according to one embodiment can generate and output new content (e.g., text, images, and / or computer code, etc.) based on learned content in response to an input prompt. For example, in an electronic device (601), learning is performed to output specific types of result data as output data using data of specified types based on a machine learning algorithm or a deep learning algorithm, thereby generating multiple AI models (e.g., machine learning models and deep learning models) that are stored in the electronic device (301), or AI models learned from an external electronic device (e.g., an external server) can be transmitted to and stored in the electronic device (201). For example, the electronic device (301) can output input data (input data) as output data of a model learned through artificial intelligence of specified types based on a machine learning algorithm or a deep learning algorithm. Machine learning algorithms include supervised learning algorithms such as linear regression and logistic regression, unsupervised learning algorithms such as clustering, visualization and dimensionality reduction, and association rule learning, and reinforcement learning algorithms, and deep learning algorithms may include Artificial Neural Networks (ANN), Deep Neural Networks (DNN), and Convolutional Neural Networks (CNN), and may further include various learning algorithms not limited to those described.The trained AI model includes at least one operation (e.g., a convolution layer or a pooling layer) for processing input data, and can be implemented to output result data by performing operations on the input data based on at least one operation.
[0190] According to one embodiment, an application that can be connected to an external AI model (381) may be stored in the memory (330).
[0191] According to one embodiment, the external AI model (381) may each include a first AI model (331a), a second AI model (331b), and a third AI model (331c), or include at least one AI model among the first AI model (331a), the second AI model (331b), and the third AI model (331c), or include the first AI model (331a), the second AI model (331b), and the third AI model (331c) as one identical AI model.
[0192] According to one embodiment, the memory (330) may include one or more programs for processing data obtained from a sensor (e.g., sensor module (178) of FIG. 1) and a camera (e.g., camera module (180) of FIG. 1 and / or camera (480) of FIG. 3a), and one or more programs may include a gesture tracker, STT (speech to text), TTS (text to speech) and / or an eye tracker, etc.
[0193] According to one embodiment, the display (360) may be implemented substantially identically or similarly to the display module (160) of FIG. 1.
[0194] According to one embodiment, the microphone (350)) can be implemented substantially identically or similarly to the input module (150) of FIG. 1.
[0195] According to one embodiment, the microphone (350) can receive ambient audio data including voice commands while capturing audio data.
[0196] According to one embodiment, the speaker (355)) can be implemented substantially identically or similarly to the acoustic output module (155) of FIG. 1.
[0197] According to one embodiment, the camera (380) may be implemented substantially identically or similarly to the camera module (180) of FIG. 1.
[0198] According to one embodiment, the camera (380) may include a first camera (381), a second camera (383), and an eye-tracking camera (385).
[0199] According to one embodiment, the first camera (381) (e.g., the first camera (211) of FIG. 2a) and the second camera (383) (e.g., the second camera (212) of FIG. 2a) can acquire images related to the surrounding environment of the wearable electronic device (301).
[0200] According to one embodiment, the first camera (381) (e.g., the first camera (211) of FIG. 2a) and the second camera (383) (e.g., the second camera (212) of FIG. 2a) can acquire an image including an external user (e.g., the external user (411) of FIG. 3) located around the wearable electronic device.
[0201] According to one embodiment, at least one camera among the first camera (381) (e.g., the first camera (211) of FIG. 2A) or the second camera (383) (e.g., the second camera (212) of FIG. 2A) can receive a gesture command included in the image data while capturing image data.
[0202] According to one embodiment, the eye-tracking camera (385) may represent an eye-tracking camera that tracks the user's gaze. The eye-tracking camera (385) may include an infrared (IR) camera.
[0203] According to one embodiment, the eye-tracking camera (385) may be placed on a second surface of the housing (e.g., the second surface (220) of FIG. 2b).
[0204] According to one embodiment, the eye-tracking camera (385) may be implemented in the same or similar manner as the camera modules (225, 226) of FIG. 2b. However, the number and arrangement structure of the eye-tracking cameras (385) included in the wearable electronic device (301) may be the same or similar as the camera modules (225, 226) of FIG. 2b or different from each other.
[0205] According to one embodiment, the eye-tracking camera (385) can acquire iris information of a user wearing a wearable electronic device (301).
[0206] According to one embodiment, the communication circuit (390) can form a communication connection with an external electronic device (e.g., another electronic device, or a server) using various types of communication methods and transmit and / or receive data. As described above, the communication methods may include communication methods that establish a direct communication connection, such as Bluetooth and Wi-Fi Direct, communication methods that use an access point (AP) (e.g., Wi-Fi communication), or communication methods that use cellular communication using a base station (e.g., 3G, 4G / LTE, 5G). Since the communication circuit (390) can be implemented as described above in the communication module (190) in FIG. 1, a redundant description is omitted.
[0207] Referring to FIGS. 3c to 3d, according to one embodiment, when a wearable electronic device (301) confirms the end of shooting of video data and audio data, if it confirms a control command (e.g., a gesture command, a voice command, a first command combining a gesture command and a voice command, or a second command combining a gaze command and a voice command), it can generate video data that performs a function corresponding to the control command and store information about the first time when the control command was confirmed in the video data as editing information.
[0208] According to one embodiment, the wearable electronic device (301) can determine the time from when the control command is displayed in the image data until when the control command is not displayed in the image data as the first time for checking the control command.
[0209] As shown in FIG. 3d, the wearable electronic device (301), for example, when one frame is 1 / 30 second, can use an AI model (331) to identify the time from when the display of the control command starts at time 00:03:12.24 (t1) until when the control command is not displayed at time 00:03:12.32 (t2) as the first time for checking the control command, and can store editing information for the first time.
[0210] According to one embodiment, the wearable electronic device (301) can generate a prompt for editing control commands included in image data by using the AI model (331) of the wearable electronic device (301).
[0211] For example, when the wearable electronic device (301) generates editing information based on FIG. 3D, it can use the AI model (331) of the wearable electronic device (301) to generate prompts such as "delete the hand displayed from 00:03:12.24 to 00:03:12.32 in the video data" or "delete the voice recorded from 00:03:14.12 to 00:03:15.14 in the video data."
[0212] According to one embodiment, the wearable electronic device (301) can generate a prompt for editing control commands included in image data by using an AI model (441) of an external electronic device (401).
[0213] For example, when the wearable electronic device (301) generates editing information based on FIG. 3D, it can use the AI model (441) of the external electronic device (401) to generate prompts such as "Delete the hand displayed from 00:03:12.24 to 00:03:12.32 in the video data" or "Delete the voice recorded from 00:03:14.12 to 00:03:15.14 in the video data."
[0214] According to one embodiment, the wearable electronic device (301) can generate a prompt that can delete the plurality of control commands at a first time when each of the plurality of control commands is checked, when the image data includes a plurality of control commands.
[0215] According to one embodiment, the wearable electronic device (301) can transmit a first message requesting editing of the image data, including image data, editing information and / or a prompt, to an external electronic device (401) through a communication circuit (390).
[0216] According to one embodiment, the wearable electronic device (301) can transmit the captured video to an external electronic device (401) while capturing video data and audio data.
[0217] According to one embodiment, the wearable electronic device (301) can, when it confirms the reception of a control command while capturing video data and audio data, perform a function corresponding to the control command and transmit a second message requesting an edit of the control command included at the time when the function corresponding to the control command was performed and at the time when the function corresponding to the control command was performed to an external electronic device (401) through a communication circuit (390).
[0218] According to one embodiment, the external electronic device (401) may include a first processor (420), an AI model (441) contained in the memory of the electronic device, a first display (460), and a first communication circuit (490).
[0219] According to one embodiment, when the first processor (420) receives a first message requesting editing of video data from a wearable electronic device (301) through a first communication circuit (490), including video data, editing information and / or a prompt, it can perform an editing operation to delete control commands included in the video data using an AI model (441) based on the first message to create a completed video file or a completed audio file, and display that the completed video file or a completed audio file has been created by performing an editing operation to delete control commands included in the video data.
[0220] According to one embodiment, the first processor (420) can receive a first message from an external electronic device other than the wearable electronic device (301) through the first communication circuit (490).
[0221] According to one embodiment, the first processor (420) can store video data and a completed video file or a completed audio file in the memory of an external electronic device, respectively, or delete the selected video data, completed video file, or completed audio file based on user input.
[0222] According to one embodiment, when the first processor (420) receives a first message from a wearable electronic device (301), it notifies the user of the external electronic device of the receipt of the first message, and when it confirms the selection of editing based on the user's input, it can generate a completed video file or a completed audio file by performing an editing operation to delete control commands included in the video data using an AI model (441) based on the first message.
[0223] According to one embodiment, the first processor (420) can perform an editing operation to delete control commands included in video data using an AI model (441) based on a first message to generate a completed video file or a completed audio file, and then perform additional editing or saving of the completed video file or the completed audio file based on a user's selection.
[0224] According to one embodiment, the first processor (420) can confirm the user's selection for executing an editing operation to delete a control command included in the image data based on the first message.
[0225] According to one embodiment, the first processor (420) can delete at least one frame containing the selected control command by using an AI model (441) when the user confirms the user's selection for executing an editing operation and the user confirms the user's selection for a control command included in the image data, or delete at least one frame containing the control command by decreasing or increasing the number of frames to be deleted.
[0226] According to one embodiment, the first processor (420) can use an AI model (441) to generate at least one new frame based on a previous frame or a subsequent frame (e.g., a background image included in multiple frames of video data) of at least one deleted frame, and insert at least one new frame at the location of at least one deleted frame.
[0227] According to one embodiment, the first processor (420) generates a prompt based on at least a portion of the first message and executes an editing operation to edit the prompt based on the user's selection, thereby allowing the user to see the prompt edited by the user.
[0228] According to one embodiment, the first processor (420) can distinguish between editable items and non-editable items in an editing operation for editing a prompt.
[0229] For example, the first processor (420) can enable the user to change the object that is the target of the control command through the third user interface, such as changing the control "change hand to tree" to "change tree to maple tree."
[0230] For example, the first processor (420) can enable the user to add a new object that is not included in the image data, such as "create a maple tree" through the third user interface.
[0231] According to one embodiment, the first processor (420) can receive user input for editing image data as a gesture command or a voice command.
[0232] According to one embodiment, the first processor (420) can receive user input for adding or subtracting additional images, images, or objects from video data as a gesture command or voice command.
[0233] According to one embodiment, when the first processor (420) receives a second message from the wearable electronic device (301) while confirming the reception of an image from the wearable electronic device (301) through the first communication circuit (490), it may delete a control command (gesture command) included in at least one frame corresponding to the time when a function corresponding to a control command is performed based on the second message, or delete a control command (voice command) recorded at the time when a function corresponding to a control command is performed, and then save a completed audio file.
[0234] According to one embodiment, the first processor (420) can transmit the first communication circuit (490) to a wearable electronic device (301) a completed video file or a completed audio file.
[0235] According to one embodiment, the first display (460) may be implemented substantially identically or similarly to the display module (160) of FIG. 1.
[0236] According to one embodiment, the AI model (441) can perform the same function as the AI model (331) of a wearable electronic device.
[0237] According to one embodiment, the first communication circuit (490) can form a communication connection with an external electronic device (e.g., another electronic device, or a server) using various types of communication methods and transmit and / or receive data. As described above, the communication methods may include a communication method that establishes a direct communication connection, such as Bluetooth and Wi-Fi Direct, a communication method using an access point (AP) (e.g., Wi-Fi communication), or a communication method using cellular communication using a base station (e.g., 3G, 4G / LTE, 5G). Since the communication circuit (490) can be implemented as described above in the communication module (190) in FIG. 1, a redundant description is omitted.
[0238] FIGS. 4a, FIGS. 4b, and FIGS. 4c are drawings for explaining the operation of gesture commands and voice commands in a wearable electronic device according to one embodiment.
[0239] Referring to FIG. 4a, according to one embodiment, an electronic device (301) (e.g., the electronic device (101) of FIG. 1) can identify a gesture command (411) included in the image data received through at least one camera of the electronic device (301), either the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a), by using a second AI model (e.g., the second AI model (331b) of FIG. 3b) while capturing image data and audio data. For example, the gesture command (411) may include a gesture representing the ratio of an area corresponding to the gesture of the user's two fingers to set the viewfinder ratio to 16:9.
[0240] Referring to FIG. 4b, according to one embodiment, an electronic device (301) (e.g., the electronic device (101) of FIG. 1) can identify a gesture command (413) included in the image data received through at least one camera of the electronic device (301), either the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a), by using a second AI model (e.g., the second AI model (331b) of FIG. 3b) while recording image data and audio data. For example, the gesture command (413) may include a gesture representing the ratio of an area corresponding to a gesture of the user's two fingers to set the viewfinder ratio to 4:3.
[0241] Referring to FIG. 4c, according to one embodiment, an electronic device (301) (e.g., the electronic device (101) of FIG. 1) can identify a gesture command (415) included in the image data received through at least one camera of the electronic device (301), either the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a), by using a second AI model (e.g., the second AI model (331b) of FIG. 3b) while capturing image data and audio data. For example, the gesture command (415) may include a gesture representing the ratio of an area corresponding to the gesture of the user's two fingers to set the viewfinder ratio to 1:1.
[0242] According to one embodiment, an electronic device (301) (e.g., the electronic device (101) of FIG. 1) can use a second AI model (e.g., the second AI model (331b) of FIG. 3b) to identify a gesture command that specifies the horizontal and vertical orientations of a viewfinder (e.g., horizontal mode or vertical mode of a viewfinder) included in image data received through at least one camera of the electronic device (301), such as the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a).
[0243] FIGS. 5A, FIGS. 5B, FIGS. 5C, FIGS. 5D, FIGS. 5E, and FIGS. 5F are drawings for explaining the operation of gesture commands and voice commands in a wearable electronic device according to one embodiment.
[0244] Referring to FIG. 5a, according to one embodiment, an electronic device (301) (e.g., the electronic device (101) of FIG. 1) can identify a user’s clenched fist gesture included in the image data received through at least one camera of the electronic device (301), such as the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a), as a gesture command (511) for ending the shooting, by using a third AI model (e.g., the third AI model (331c) of FIG. 3b) while capturing image data and audio data.
[0245] According to one embodiment, the electronic device (301) can identify a designated user’s utterance of “stop recording” received through a microphone (e.g., microphone (350) in FIG. 3a) while recording video data and audio data as a voice command (513) for ending recording.
[0246] According to one embodiment, the electronic device (301) can identify a first command combining a gesture command (511) and a voice command (513) as a first command for ending the shooting while shooting video data and audio data.
[0247] Referring to FIG. 5b, according to one embodiment, an electronic device (301) can use a third AI model (e.g., the third AI model (331c) of FIG. 3b) while capturing video data and audio data to identify a gesture of a user’s open palm included in the video data received through at least one camera of the electronic device (301), such as the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a), as a gesture command (531) for pausing the recording.
[0248] According to one embodiment, the electronic device (301) can identify a designated user’s utterance of “pause” received through a microphone (e.g., the microphone (350) in FIG. 3a) while capturing video data and audio data as a voice command (533) for pausing the capture.
[0249] According to one embodiment, the electronic device (301) can identify a first command, which is a combination of a gesture command (531) and a voice command (533), as a first command for pausing the recording of video data and audio data while recording video data and audio data.
[0250] Referring to FIG. 5c, according to one embodiment, an electronic device (301) can identify a gesture command (551) indicating the ratio of an area with a user's finger included in the image data received through at least one camera of the electronic device (301), such as the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a), by using a third AI model (e.g., the third AI model (331c) of FIG. 3b) while capturing image data and audio data.
[0251] According to one embodiment, the electronic device (301) can identify the gesture command (551) as a gesture command for setting the viewfinder ratio to a specific ratio (e.g., 1:1).
[0252] According to one embodiment, the electronic device (301) can identify a designated user's utterance of "1:1" received through a microphone (e.g., microphone (350) in FIG. 3a) while capturing video data and audio data as a voice command (553) for setting the viewfinder ratio to a specific ratio (e.g., 1:1).
[0253] According to one embodiment, the electronic device (301) can identify a first command, which is a combination of a gesture command (551) and a voice command (553), as a first command for setting the viewfinder ratio to a specific ratio (e.g., 1:1) while capturing image data and audio data.
[0254] According to one embodiment, the electronic device (301) may provide the adjusted ratio of the viewfinder as visual feedback (e.g., color (555)) to inform the user that the input of a first command to set the viewfinder ratio to a specific ratio (e.g., 1:1) has been confirmed.
[0255] Referring to FIG. 5d, according to one embodiment, an electronic device (301) can use a third AI model (e.g., the third AI model (331c) of FIG. 3b) while capturing video data and audio data to identify a gesture of widening the gap between the user's fingers included in the video data received through at least one camera of the electronic device (301), such as the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a), as a gesture command (571) for zooming in.
[0256] According to one embodiment, the electronic device (301) can identify a designated user’s utterance of “zoom in” received through a microphone (e.g., microphone (350) in FIG. 3a) while capturing video data and audio data as a voice command (573) for zooming in.
[0257] According to one embodiment, the electronic device (301) can identify a first command, which is a combination of a gesture command (571) and a voice command (573), as a first command for zooming in on an area where the user's finger is located while capturing video data and audio data.
[0258] According to one embodiment, the electronic device (301) may provide various effects as user feedback corresponding to the area where the user's finger is located for zooming in, in order to notify the user that the input of a first command for zooming in on the area where the user's finger is located has been confirmed. For example, the electronic device (301) may include various multimodal effects as user feedback, such as vibration, image effects (e.g., display of color (575)), and / or animation effects.
[0259] Referring to FIG. 5e, according to one embodiment, an electronic device (301) can use a third AI model (e.g., the third AI model (331c) of FIG. 3b) while capturing video data and audio data to identify a gesture of narrowing the gap between the user's fingers included in the video data received through at least one camera of the electronic device (301), such as the first camera (e.g., the first camera (381) of FIG. 3a) or the second camera (e.g., the second camera (383) of FIG. 3a), as a gesture command (591) for zooming out.
[0260] According to one embodiment, the electronic device (301) can identify a designated user’s utterance of “zoom out” received through a microphone (e.g., the microphone (350) in FIG. 3a) while capturing video data and audio data as a voice command (593) for zooming in.
[0261] According to one embodiment, the electronic device (301) can identify a first command, which is a combination of a gesture command (591) and a voice command (593), as a first command for zooming out of an area where the user's finger is located while capturing video data and audio data.
[0262] According to one embodiment, the electronic device (301) may provide various effects as user feedback corresponding to the area where the user's finger is located for zooming out, in order to notify the user that the input of a first command for zooming out to the area where the user's finger is located has been confirmed. For example, the electronic device (301) may include various multimodal effects as user feedback, such as vibration, image effects (e.g., display of color (595)), and / or animation effects.
[0263] Fig. 5f <511> Referring to the above, according to one embodiment, the electronic device (301)) can identify a gesture corresponding to a user's finger pointing to an object (585) (e.g., a person) in the image data received through at least one camera of the electronic device (301), such as the first camera (e.g., the first camera (381) of FIG. 3A) or the second camera (e.g., the second camera (383) of FIG. 3A), as a gesture command (581) pointing to an object by using a third AI model (e.g., the third AI model (331c) of FIG. 3B) while capturing image data and audio data.
[0264] According to one embodiment, when the electronic device (301) confirms a designated user’s utterance of “object delete” received through a microphone (e.g., microphone (350) of FIG. 3a) while capturing video data and audio data as a voice command (583) for deleting an object (585) pointed to by a gesture command, it can confirm a first command combining the gesture command (581)) and the voice command (583) as a first command for deleting an object (585) pointed to by the gesture command (581).
[0265] Fig. 5f <513> Referring to the above, according to one embodiment, the electronic device (301) can be identified by a first command to delete an object (585) pointed to by a gesture command (581). According to one embodiment, the electronic device (101) can delete an object (585) pointed to by a gesture command (581) included in the viewfinder at the time of identification.
[0266] According to one embodiment, when the electronic device (301) identifies the gesture command (581) and the voice command (583) as a first command to delete the object (585) pointed to by the gesture command (581), it can mark the object (585) and delete the marked object (585) at the time of ending the recording of the video data and audio data.
[0267] According to one embodiment, the electronic device (301) can provide natural image data by generating and adding a new partial image based on a background image to an area containing a deleted object (585) using a third AI model (331c).
[0268] FIGS. 6a and 6b are drawings for explaining the operation of deleting a first command, which is a combination of a gesture command and a voice command, in a wearable electronic device according to one embodiment.
[0269] Referring to FIG. 6a, according to one embodiment, while capturing image data and audio data through an electronic device (301) (e.g., the electronic device (101) of FIG. 1), a first camera (e.g., the first camera (211) of FIG. 2a and / or the first camera (381) of FIG. 3a) and a second camera (e.g., the second camera (212) of FIG. 2a and / or the second camera (383) of FIG. 3a), a third AI model (e.g., the third AI model (331c) of FIG. 3b) is used to identify a first command combining a gesture command (611) pointing to an object (651) (e.g., a flower) included in image data (671) received through at least one of the first or second cameras and a designated user's voice command (631) called "focus" received through a microphone (e.g., the microphone (350) of FIG. 3a), thereby focusing on the object (651) (e.g., a flower). It can perform functions.
[0270] According to one embodiment, when the electronic device (301) is set to perform an editing operation for a gesture command, a designated user's voice command, a first command combined with a gesture command and a designated user's voice command, or a second command combined with an eye tracking command and a designated user's voice command when the shooting of video data and audio data is finished, the device may store information about a first time when the first command combined with a gesture command (611) and a designated user's voice command (631) called "focus" received through a microphone (e.g., the microphone (350) in FIG. 3a) is confirmed as editing information.
[0271] Referring to FIG. 6b, according to one embodiment, when the electronic device (301) confirms the end of shooting of video data and audio data, it can use a third AI model (e.g., the third AI model (331c) of FIG. 3b) to confirm that the size of a first area (image corresponding to the gesture command) corresponding to a gesture command (611) is greater than or equal to a first threshold value for confirming the ratio of the gesture command to at least one frame.
[0272] According to one embodiment, the electronic device (301) can identify a gesture command (611) included in at least one frame (671a, 671b, 671c and 671d) corresponding to a first time in the video data and a voice command (631) identified at the first time based on editing information.
[0273] According to one embodiment, the electronic device (301) may store at least one frame (671a, 671b, 671c and 671d) containing a gesture command (611) in the image data and a completed image file or completed audio file with the voice command (631) deleted.
[0274] According to one embodiment, when an electronic device (301) is configured to perform an editing operation for a gesture command, a designated user's voice command, or a first command combining a gesture command and a designated user's voice command while capturing video data and audio data, it can verify a first command combining a gesture command (611) and a designated user's voice command (631) by using a third AI model (e.g., the third AI model (331c) of FIG. 3b) while capturing video data and audio data. For example, the electronic device (301) can perform a focusing function for an object (651) (e.g., a flower) and verify that the size of a first area corresponding to the gesture command (611) is greater than or equal to a first threshold value for verifying the ratio that the gesture command occupies in at least one frame.
[0275] According to one embodiment, the electronic device (301) can delete at least one frame (671a, 671b, 671c and 671d) containing a gesture command (611) and a voice command (631) while capturing video data and audio data, and when the capture of video data and audio data is confirmed to be finished, save a completed video file or a completed audio file.
[0276] FIG. 7 is a diagram illustrating the deletion operation of a gesture command in a wearable electronic device according to one embodiment.
[0277] Referring to FIG. 7, according to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1) can identify at least one frame (713) containing a gesture command (713a) for adjusting the viewfinder ratio to a specific ratio (e.g., 1:1) by using a second AI model (e.g., the second AI model (331b) of FIG. 2b) at the time when the end of shooting of video data and audio data is confirmed or while shooting video data and audio data is confirmed. According to one embodiment, if the electronic device confirms that the size of a first area (713b) (an image corresponding to the gesture command) corresponding to the gesture command (713a) is greater than or equal to a first threshold value for confirming the ratio occupied by the gesture command (713a) in at least one frame (713), the electronic device can delete at least one frame (713) containing the gesture command (713a).
[0278] According to one embodiment, the electronic device can confirm that at least one frame (713) containing a gesture command (713a) is included in a frame excluding an intermediate frame or the last frame of the image data.
[0279] According to one embodiment, the electronic device can generate new frames (715a, 715b, and / or 715c) based on at least one of the previous original frame (711) or the next original frame (717) of at least one frame (713), and insert the new frames (715a, 715b, and 715c) at the location of the deleted at least one frame (713) to generate a complete video file that is naturally connected without interruption between frames.
[0280] FIGS. 8a and 8b are drawings for explaining the operation of deleting voice commands in a wearable electronic device according to one embodiment.
[0281] Referring to FIG. 8a, according to one embodiment, a first audio received through a microphone (350) of an electronic device (e.g., the microphone (350) of FIG. 3a) can be transmitted to a processor of an electronic device (e.g., the processor (120) of FIG. 1) or a first AI model (e.g., the AI model (331a) of FIG. 3b) (811).
[0282] According to one embodiment, the processor or the first AI model checks whether the first audio data received through the microphone (350) contains a voice command of a designated user, and if it checks whether the first audio data contains a voice command of a designated user, it can generate a second audio data in which the voice command of a designated user is deleted from the first audio data. (813)
[0283] According to one embodiment, the processor or the first AI model can store a completed audio file that is synchronized between the captured video data and the second audio data from which the user's voice command specified in the first audio data, which is the original audio, has been deleted (815).
[0284] Referring to Fig. 8b, <831> As such, the processor or the first AI model can check whether the first audio data received through the microphone (350) contains a user's voice command, and if it checks whether the first audio data contains a designated user's voice command, it can separate the surrounding audio (a1) received by the microphone (350) of the electronic device while capturing video data and audio data from the first audio data and the designated user's voice command (a2), and then generate a second audio data in which the designated user's voice command (a2) is deleted.
[0285] <833> ...represents first audio data corresponding to original audio including voice commands received through the microphone (350), and <835> The processor or first AI model represents a completed audio file that is synchronized between the second audio data, from which the user's voice command (a2) specified in the first audio data, which is the original audio, has been deleted, and the captured video data.
[0286] FIG. 9 is a diagram illustrating the operation of executing a gesture command in a wearable electronic device according to one embodiment.
[0287] Referring to Fig. 9, <911> As such, according to one embodiment, a wearable electronic device (e.g., the electronic device (101) of FIG. 1) can display a border area (913a) of a viewfinder (913) representing an actual shooting area on the display of the wearable electronic device in a first color while capturing image data and audio data.
[0288] <931> As shown above, according to one embodiment, when a wearable electronic device detects the start of a gesture command (933) for area cropping while capturing video data and audio data, it may display a line (935) at a location indicating the start of setting the area to be cropped, and provide image effects (e.g., color, image size, motion, etc.) to the border area (913b) of the viewfinder (913).
[0289] According to one embodiment, the wearable electronic device can signal the start of a cropping action by providing not only image effects but also other effects (e.g., vibration or sound). <951> As shown, according to one embodiment, the wearable electronic device may provide an image effect (e.g., opaque black) to the cropped area to indicate that the cropped area (955) will not be included in the image data when the completion of a gesture command (933) for area cropping (e.g., moving to the left to select the area to be cropped) is confirmed while capturing image data and audio data.
[0290] <971> As shown, according to one embodiment, the wearable electronic device can display the color of the border area (913c) of the viewfinder (913), excluding the cropped area (955) that provides an image effect (e.g., opaque black), as a first color when checking the crop setting for the area based on a gesture command.
[0291] According to one embodiment, the electronic device can intuitively divide a cropped area using a gesture command while capturing video data and audio data, and can generate a completed video file or a completed audio file by deleting the gesture command for cropping the area from the video data.
[0292] According to one embodiment, when an area cropping operation is set based on a gesture command, the electronic device can display a viewfinder (913) excluding the cropped area (955) and adjust the position of the viewfinder in real time according to the position of the viewing angle.
[0293] FIG. 10 is a diagram illustrating the operation of confirming a gesture command in a wearable electronic device according to one embodiment.
[0294] Referring to FIG. 10, a wearable electronic device (e.g., the electronic device (101) of FIG. 1) can acquire a gesture command (1011) of pointing to an object (1013) with a user's finger while capturing video data and audio data.
[0295] According to one embodiment, a wearable electronic device can recognize a gesture command corresponding to a user’s gesture included in image data and acquire actions corresponding to the form of the user’s gesture.
[0296] According to one embodiment, <1031> In this example, the position of the actual user's fingertip pointing to the object (1013) and / or the direction of the finger (1011a) is indicated, and <1035> In this case, the position of the user's fingertip pointing at an object (1013) identified through the camera of the electronic device and / or the direction of the finger (1011a) may be indicated.
[0297] According to one embodiment, the electronic device can identify an error between the position and / or direction of the fingertip of the actual user pointing to the object (1011a) and the position and / or direction of the fingertip of the user confirmed through the camera of the electronic device (1011b), and perform a calibration operation to compensate for the angle difference between the position and / or direction of the fingertip of the user (1011a) and the camera of the wearable electronic device (e.g., camera (380) of FIG. 3a) and analyze and identify the object (1013) pointed to by the user.
[0298] FIGS. 11a, FIGS. 11b, and FIGS. 11c are drawings for illustrating the generation of a plurality of new frames containing a new image in a wearable electronic device according to one embodiment.
[0299] Referring to FIG. 11a, according to one embodiment, a wearable electronic device (e.g., the electronic device (101) of FIG. 1)) can delete a plurality of frames (1115) containing a gesture command (1111) by using an AI model (e.g., the AI model (331) of FIG. 3b).
[0300] According to one embodiment, a plurality of frames (1115) may include a first object (e.g., a1, a2, and a3) (e.g., a passing vehicle) only in the plurality of frames, and a specific audio signal (e.g., the sound of a vehicle passing) may be recorded in a certain section corresponding to the plurality of frames (1115).
[0301] Referring to FIG. 11b, according to one embodiment, when a wearable electronic device deletes a plurality of frames (1115) containing a gesture command (1111), it can generate a new plurality of frames (1119a) containing a newly generated image based on a previous frame (1113) or a next frame (1117) of the plurality of frames (1115) containing the gesture command (111).
[0302] According to one embodiment, when a wearable electronic device includes new multiple frames (1119a) containing a newly generated image at the location of multiple frames (1115) containing a deleted gesture command, it can identify a specific audio signal (e.g., the sound of a vehicle passing) that is not related to the newly generated image among audio, excluding voice recorded at the time corresponding to the multiple frames (1115) containing the gesture command (111). According to one embodiment, the electronic device can generate a completed image file or a completed audio file in which the specific audio signal (e.g., the sound of a vehicle passing) that is not related to the newly generated image is deleted.
[0303] Referring to FIG. 11c, according to one embodiment, when a wearable electronic device deletes a plurality of frames (1115) containing a gesture command (1111), if the plurality of frames containing the gesture command are consecutive for a certain period, it can use an AI model (e.g., the AI model (331) of FIG. 3b) to generate a new plurality of frames (1119b) containing a newly generated image (b1, b2 and b3) (e.g., a passing vehicle) corresponding to a specific audio signal (e.g., the sound of a vehicle passing) among audio other than voice included within the certain period.
[0304] According to one embodiment, when a wearable electronic device deletes a plurality of frames (1115) containing a gesture command (1111), if the plurality of frames containing the gesture command are consecutive for a certain period, it can use an AI model (e.g., the AI model (331) of FIG. 3b) to analyze the deleted plurality of frames (1115), and use the results of the analysis to restore an object included in the plurality of frames (1115) to generate a new plurality of frames (1119b) containing a newly generated image (b1, b2 and b3) (e.g., a passing vehicle).
[0305] According to one embodiment, the wearable electronic device may generate a complete video file or a complete audio file by including new multiple frames (1119b) containing a newly generated image at the location of multiple frames (1115) containing a deleted gesture command (1111).
[0306] FIGS. 12a, FIGS. 12b, FIGS. 12c, FIGS. 12d, and FIGS. 12e are drawings for explaining the operation of editing image data in an electronic device connected to a wearable electronic device according to one embodiment.
[0307] Referring to FIG. 12a, according to one embodiment, an external electronic device (401) (e.g., the electronic device (102) of FIG. 1) can confirm receipt of a first message from a wearable electronic device (e.g., the electronic device (101) of FIG. 1).
[0308] According to one embodiment, an external electronic device (401) can generate a completed video file or a completed audio file by performing an editing operation to delete control commands included in the video data using an AI model of the external electronic device (401) (e.g., AI model (441) of FIG. 3c) based on a first message requesting editing of the video data, which includes video data, editing information and / or a prompt.
[0309] According to one embodiment, an external electronic device (401) can display a completed video file or a completed audio file (1213) by performing an editing operation to delete control commands using an AI model and video data (1211) containing control commands through the display of the external electronic device (401) (e.g., the first display (460) of FIG. 3c).
[0310] Referring to FIG. 12b, according to one embodiment, when an external electronic device (401) confirms receipt of a first message from a wearable electronic device, it may display, through a display (e.g., the first display (460) in FIG. 3c), image data (1231) received from the wearable electronic device, text (1233) notifying of the receipt of image data from the wearable electronic device, a first button (1235) that can edit the image data received from the wearable electronic device based on user input, and / or a second button (1237) that can automatically modify and save the image data received from the wearable electronic device.
[0311] Referring to FIG. 12c, according to one embodiment, an external electronic device (401) can confirm the reception of a first message from a wearable electronic device.
[0312] According to one embodiment, an external electronic device (401) can generate a completed video file or a completed audio file by performing an editing operation to delete control commands included in the video data using an AI model of the external electronic device (401) (e.g., AI model (441) of FIG. 3c) based on a first message requesting editing of the video data, which includes video data, editing information and / or a prompt.
[0313] According to one embodiment, an external electronic device (401) may display, through a display (e.g., a first display (460) of FIG. 3c), a completed video file or completed audio file (1251) that has been edited from video data received from a wearable electronic device, text (1253) indicating that the editing of video data from the wearable electronic device is complete, a third button (1255) that can further edit the completed video file or completed audio file based on user input, and / or a fourth button (1257) that can save the completed video file or completed audio file.
[0314] Referring to FIG. 12d, according to one embodiment, an external electronic device (401) can receive a first message from a wearable electronic device.
[0315] According to one embodiment, in an editing operation to delete control commands included in image data, an external electronic device (401) may display image data (1271) received from a wearable electronic device, information (1273) of a first time in which control commands can be checked in the image data, a fifth button (1275) that can delete control commands based on user input, and / or a button (1277) that can return to before editing.
[0316] According to one embodiment, the external electronic device (401) can move to at least one frame containing a control command based on the user's selection of information (1273) of a first time in which a control command can be confirmed in the image data, delete at least one frame containing a control command, or delete the frame containing the control command by decreasing or increasing the number of frames to be deleted.
[0317] According to one embodiment, when an external electronic device (401) deletes at least one frame containing a control command, it can generate a new plurality of frames containing a new image based on the previous frame or the next frame of the at least one frame containing the control command, and generate a completed image file including the new plurality of frames at the location of the at least one frame containing the deleted control command.
[0318] Referring to FIG. 12e, according to one embodiment, an external electronic device (401) receives a first message from a wearable electronic device and, when executing an editing operation of a prompt included in the first message based on user input, can display a frame (1291) containing a control command (1291a) in image data and / or a prompt (1293) for editing the control command (1291a) included in the frame (1291) through a display (e.g., the first display (460) of FIG. 3c).
[0319] According to one embodiment, when an external electronic device (401) confirms the editing of a prompt based on user input, it can edit (delete) the control command (1291a) included in the frame (1291) based on the edited prompt.
[0320] According to one embodiment, the external electronic device (401) can distinguish between editable items and non-editable items in an editing operation for editing a prompt.
[0321] According to one embodiment, the external electronic device (401) can receive video data captured from the wearable electronic device in real time and store it in the memory of the external electronic device (401).
[0322] According to one embodiment, while receiving image data from a wearable electronic device, an external electronic device (401) may receive notification information regarding the execution of a control command from the wearable electronic device (e.g., a voice command of a designated user, a first command combining a gesture command and a voice command, or a second command including a gaze command and a voice command).
[0323] According to one embodiment, an external electronic device (401) can receive notification information regarding the execution of a control command through an AI model of a wearable electronic device, and can store information regarding a second time at which the notification information regarding the execution of a control command was received in the memory of the external electronic device (401).
[0324] According to one embodiment, if the external electronic device (401) does not receive image data from the wearable electronic device due to the termination of shooting of image data from the wearable electronic device, it may check at least one frame corresponding to a second time among a plurality of frames of image data received from the wearable electronic device, delete at least one frame or a gesture command included in at least one frame, or delete some audio data recorded at the second time among the audio data and generate a completed image file or a completed audio file.
[0325] According to one embodiment, the external electronic device (401) may store a completed video file or a completed audio file, or store original video data or original audio data received from a wearable electronic device and a completed video file or a completed audio file, respectively.
[0326] According to one embodiment, when an external electronic device (401) generates a completed video file or a completed audio file, it notifies the user of the external electronic device (401) of the generation of the completed video file or the completed audio file, and can perform additional editing operations on the completed video file or the completed audio file based on the user's selection.
[0327] According to one embodiment, the external electronic device (401) may provide a user interface that allows the user of the external electronic device (401) to select an edited part when performing an additional editing operation.
[0328] FIGS. 13a, FIGS. 13b, and FIGS. 13c are drawings for illustrating a viewfinder in a wearable electronic device according to one embodiment.
[0329] As shown in FIG. 13a, according to one embodiment, an electronic device (e.g., the electronic device (131) of FIG. 1)) can display a viewfinder that captures an area of the user's field of view through a camera of the wearable electronic device (e.g., the first camera (381) and the second camera (383) of FIG. 3a) in a first area (an image corresponding to a gesture command) (1301) of the display (360) of the wearable electronic device (e.g., the first camera (381) and the second camera (383) of FIG. 3a)) and can display in real time the type of object newly included in the captured image data while capturing image data and audio data in a second area (1303) of the display that does not overlap with the viewfinder.
[0330] For example, when the electronic device detects that a cloud object (1303a) is newly included in the captured image data while capturing image data and audio data, it can display the newly included cloud object (1303a) in a second area (1303) of the display that is not overlapped with the viewfinder.
[0331] As shown in FIG. 13b, according to one embodiment, the electronic device may display a viewfinder (1311) that captures an area of the user's field of view through a camera of the wearable electronic device (e.g., a first camera (e.g., a first camera (381) and a second camera (383) of FIG. 3a)) in a first area (1311) of the display (360) of the wearable electronic device (e.g., the first camera (381) and the second camera (383) of FIG. 3a). According to one embodiment, the electronic device may display a timeline (1331) of functions performed while capturing video data and audio data, such as the time when a voice command is confirmed while capturing video data and audio data, the time when a crop operation is performed on an area, or the time when an object is marked for deletion, in a second area (1313) of the display that does not overlap with the viewfinder.
[0332] As shown in FIG. 13c, according to one embodiment, when the electronic device captures image data with a camera located at the bottom of the wearable electronic device rather than a camera located at the top of the wearable electronic device (e.g., a first camera (e.g., the first camera (381) and the second camera (383) of FIG. 3a)), the electronic device may display a viewfinder in a second area (1335) of the display (360) of the wearable electronic device (e.g., the display (360) of FIG. 3a) that captures an area outside the user's field of view through the camera located at the bottom of the wearable electronic device. According to one embodiment, the electronic device may display image data displayed in the viewfinder of the second area (1335) as preview image data (1331a) in a first area (1331) of the display that does not overlap with the viewfinder.
[0333] FIGS. 14a, FIGS. 14b, and FIGS. 14c are drawings for illustrating an AI model in an electronic device according to one embodiment.
[0334] Referring to FIG. 14a, according to one embodiment, an electronic device (301) (e.g., electronic device (101) of FIG. 1, electronic device (200) of FIG. 2a to 2b and / or electronic device (301) of FIG. 3a to 3b) can convert audio received from an input device (1411) of the electronic device (301) (e.g., microphone (350) of FIG. 3a) into text by performing automatic speech recognition (ASR) (1413) (1413), and transmit the converted text to a first AI model (1401a) included in an external server (1401).
[0335] According to one embodiment, the first AI model (1401a) may include an on-device AI model included in the electronic device (301).
[0336] According to one embodiment, a first AI model (1401a) (e.g., a large language model (LLM)) included in an external server (1401) can identify the intent and / or content of a user's utterance based on text received from an electronic device (301) (1415), generate an answer (1417) that matches the intent, and transmit it to the electronic device (301) in the form of text.
[0337] According to one embodiment, a first AI model (1401a) (e.g., a large language model (LLM)) can analyze audio received through a microphone of an electronic device.
[0338] According to one embodiment, the electronic device (301) converts text received from an external server (1401) into audio using a text-to-speech (TTS) (1409), outputs the converted audio through an output device (1421) (e.g., a speaker of the electronic device), and while outputting the converted audio through the output device (1421) (e.g., a speaker of the electronic device), the text can be displayed through the display of the electronic device (301).
[0339] According to one embodiment, the operation of the first AI model (1401a) included in the external server (1401) (e.g., large language model (LLM)) can be performed in the same way in the first AI model included as an on-device model in the electronic device (301) (e.g., the first AI model (331a) of FIG. 3a).
[0340] According to one embodiment, by using a first AI model (1401a) included in an external server (1401) or an on-device first AI model of an electronic device, the speech of a designated user received through the microphone of a wearable electronic device (301) can be analyzed to perform the recording of video data and audio data, and the video data and audio data being recorded can be terminated.
[0341] Referring to FIG. 14b, according to one embodiment, an electronic device (301) (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b and / or the electronic device (301) of FIG. 3a to 3b) can transmit image data received through an input device (1431) (e.g., a camera of the electronic device) to an external server (1401).
[0342] According to one embodiment, a second AI model (1401b) (e.g., Vision model) included in an external server (1401) can determine the type of image received from an electronic device (301) as a label, or, if there are multiple objects in one image, determine a label for each object (1433), and then transmit the result of the label (1435) to the electronic device (301).
[0343] According to one embodiment, a second AI model (1401b) (e.g., Vision model) can analyze image data (still image data and / or video data) captured through a camera of an electronic device (301).
[0344] According to one embodiment, the second AI model (1401b) may include an on-device AI model included in the electronic device (301).
[0345] According to one embodiment, when the electronic device (301) receives a result (1435) of a label from an external server (1401), it can process and use it according to an application or service (1437).
[0346] According to one embodiment, the operation of the second AI model (1401b) (e.g., Vision model) included in the external server (1401) can be performed in the same way in the second AI model (e.g., the second AI model (331b) of FIG. 3b) included as an on-device model in the electronic device (301).
[0347] According to one embodiment, by using a second AI model (1401b) included in an external server (1401) or an on-device second AI model of an electronic device, when a first gesture (e.g., a hand gesture with the palm spread wide) included in the captured video data is identified while capturing video data and audio data, the first gesture (e.g., a hand gesture with the palm spread wide) is identified as a gesture command for ending the capture, and the capture of video data and audio data can be ended.
[0348] Referring to FIG. 14c, according to one embodiment, a structure for a method of processing multiple input data in a third AI model (e.g., a large multimodal model (LMM)) capable of analyzing image data, audio data, and / or text data can be described.
[0349] According to one embodiment, a third AI model (e.g., a large multimodal model (LMM)) can extract all features (1451b, 1453b, and 1455b) for audio (1451a), visual data (1453a), and / or text data (1455a) from data collected according to each sensor of the electronic device.
[0350] According to one embodiment, a third AI model (e.g., LMM (large multimodal model)) can fuse features extracted from an Encoder (1457) and a Dense layer (1459) into one to generate a final output (label) (1461).
[0351] For example, while the wearable electronic device is capturing video data and audio data, it can use a third AI model (e.g., LMM (large multimodal model)) to identify a gesture in which a user's finger included in the captured video data points to an object included in the video data, and simultaneously identify a user's utterance of "Focus," thereby identifying the gesture pointing to an object included in the video data as a gesture command and identifying the user's utterance of "Focus" as a voice command, and thus perform a focus function on the object pointed to by the user's finger.
[0352] An electronic device according to one embodiment (101 of FIG. 1; 200 of FIG. 2a to 2b; 301 of FIG. 3a to 3c) may include at least one camera (180 of FIG. 1; 380 of FIG. 3a), a microphone (150 of FIG. 1; 350 of FIG. 3a), at least one processor (120 of FIG. 1; 320 of FIG. 3a), and a memory for storing instructions (130 of FIG. 1; 330 of FIG. 3a). When the instructions according to one embodiment are executed individually or collectively by the at least one processor, the electronic device may start capturing video data and audio data through the camera and the microphone, respectively. According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may, while capturing the image data and the audio data, perform a function corresponding to the at least one command when it checks, using an AI model, at least one command among a gesture command included in the image data or a designated user's voice command included in the audio data, and store information regarding the first time when the at least one command was checked in the memory. According to one embodiment, when the commands are executed individually or collectively by the at least one processor, when the electronic device checks the end of capturing the image data and the audio data, it may store a completed image file including the captured image data and the captured audio data, excluding the data corresponding to the at least one command checked at the first time, and the data corresponding to the at least one command may be deleted from the captured image data and the audio data based on the information regarding the first time using the AI model.
[0353] According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may use the AI model to compare the size of a first region corresponding to a gesture command included in at least one frame among a plurality of frames corresponding to the image data with a first threshold value. According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may use the AI model to delete the at least one frame containing the gesture command if the size of the first region is greater than or equal to the first threshold value. According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may use the AI model to delete the gesture command included in the at least one frame if the size of the first region is less than or equal to the first threshold value.
[0354] According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may delete the at least one frame containing the gesture command by using the AI model if it confirms that the control command is included in at least one frame other than the last frame among the plurality of frames. According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may generate a new at least one frame based on at least one of the previous frame or the next frame of the at least one frame by using the AI model. According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may insert the new at least one frame at the location of the deleted at least one frame among the plurality of frames by using the AI model.
[0355] When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device can delete the at least one frame from the plurality of frames by using the AI model to confirm the inclusion of the gesture command in the at least one frame, including the last frame among the plurality of frames. When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device can generate the completed image file that does not fill the position corresponding to the deleted at least one frame in the plurality of frames by using the AI model.
[0356] When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device can use the AI model to check the rate of change of the scene between the at least one frame containing the gesture command and the previous frame of the at least one frame. When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device can use the AI model to delete the gesture command included in the at least one frame if the rate of change of the scene is greater than or equal to a second threshold value. When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device can use the AI model to delete the at least one frame containing the gesture command if the rate of change of the scene is less than or equal to the second threshold value.
[0357] When the commands according to one embodiment are executed individually or collectively by the at least one processor, the electronic device can use the AI model to delete the gesture command included in the at least one frame, and correct a first region corresponding to the deleted gesture command based on at least one of the previous frame or the next frame of the at least one frame.
[0358] An electronic device according to one embodiment (101 of FIG. 1; 200 of FIG. 2a to 2b; 301 of FIG. 3a to 3c) may include at least one camera (180 of FIG. 1; 380 of FIG. 3a), a microphone (150 of FIG. 1; 350 of FIG. 3a), at least one processor (120 of FIG. 1; 320 of FIG. 3a), and a memory for storing instructions (130 of FIG. 1; 330 of FIG. 3a). According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may receive control commands through at least one of the gaze, gesture, and voice of a user wearing the electronic device while capturing or recording at least one of video or audio. When the instructions according to one embodiment are executed individually or collectively by the at least one processor, the electronic device generates video data or audio data by reflecting the control instructions, and when at least one of the gesture and the voice is included in the video data or the audio data, the AI model can be used to generate a complete video file or a complete audio file in which the gesture or the voice is removed from the video data or the audio data.
[0359] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may generate the completed video file by removing at least some of the gestures from a plurality of frames containing the gestures through the AI model, or by replacing the plurality of frames containing the gestures with new plurality of frames containing a newly generated video by the AI model.
[0360] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may generate the completed image file by replacing the plurality of frames containing the gesture with new plurality of frames containing the gesture with new images newly generated by the AI model, wherein the plurality of frames containing the gesture are continuous for a certain period, and an image corresponding to a specific audio signal among the audio other than the voice included within the certain period may be generated through the AI model and added to the newly generated image.
[0361] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device can generate the completed image file by removing a first region corresponding to the gesture from the image data if the image is a still image.
[0362] According to one embodiment, when the commands are executed individually or collectively by the at least one processor, the electronic device may, when the electronic device further includes at least one of a display and a speaker, or transmits the image being captured to an external electronic device including a display, visually display an object to be the target of the control command within the image on a display included in the electronic device or a display included in the external electronic device while capturing the image data and audio data, or audibly output sound of an object to be the target of the control command through the speaker.
[0363] According to one embodiment, when the electronic device further includes a display or transmits the image data being captured to an external electronic device including a display, the instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to display edited image data in which the control instruction is reflected, or edited image data in which the content of the control instruction reflected in the image data is displayed, on a display included in the electronic device or a display included in the external electronic device.
[0364] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may store at least one of an original image file in which the control instructions are not reflected, an edited image file generated based on the edited image data, and a completed image file.
[0365] FIG. 15 is a flowchart illustrating an operation to edit video data at the time of completion of capturing video data and audio data in a wearable electronic device according to one embodiment. The operations for editing video data may include operations 1501 to 1513. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, at least two operations may be performed in parallel, or other operations may be added.
[0366] In operation 1501, a wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can start capturing video data and audio data.
[0367] According to one embodiment, the wearable electronic device can perform a shooting function for image data through a first camera (381) (e.g., camera (211) of FIG. 2a and / or the first camera (381) of FIG. 3)) and a second camera (e.g., the second camera (212) of FIG. 2a and / or the second camera (383)) among the cameras of the wearable electronic device (e.g., camera (380) of FIG. 3a) that can acquire image data related to the surrounding environment of the wearable electronic device while the wearable electronic device is worn by a part of the user's body.
[0368] According to one embodiment, the wearable electronic device can perform a shooting function for image data through one of a first camera and a second camera capable of acquiring image data related to the surrounding environment.
[0369] According to one embodiment, the wearable electronic device can start a shooting function for video data and audio data when it confirms the selection of a shooting function button among at least one button displayed through the display of the electronic device (e.g., the display (360) of FIG. 3a) based on user input.
[0370] According to one embodiment, when a wearable electronic device receives a voice command requesting a shooting function through the microphone of the electronic device (e.g., the microphone (350) of FIG. 3a), if the received voice command is identified as a voice command of a designated user, the device can interpret the voice command using at least one artificial intelligence model (e.g., the first AI model (331a)) among at least one AI model (331) and perform a shooting function corresponding to the interpreted voice command.
[0371] According to one embodiment, a wearable electronic device can interpret a voice command using an artificial intelligence model and perform a shooting function corresponding to the interpreted voice command.
[0372] According to one embodiment, a wearable electronic device can interpret a voice command by individually utilizing a plurality of artificial intelligences and perform a shooting function corresponding to the interpreted voice command.
[0373] In operation 1503, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to FIG. 2b, and / or the wearable electronic device (301) of FIG. 3a to FIG. 3b) can check at least one command among gesture commands or voice commands while capturing video data and audio data.
[0374] According to one embodiment, while the wearable electronic device is performing the capture of image data through the first camera and the second camera of the electronic device, it can confirm the reception of a gesture command, a voice command of a designated user, a first command combining a gesture command and a voice command, or a second command combining a gaze command and a voice command based on at least one of the first camera or the second camera and the microphone of the electronic device (e.g., the microphone (350) of FIG. 3a).
[0375] According to one embodiment, the wearable electronic device can confirm the input of a control command through at least one of the gaze, gesture, and / or voice of a user wearing the wearable electronic device while capturing video data and audio data or recording only audio data.
[0376] According to one embodiment, the control command may include a designated user's voice command, a first command combining a gesture command and a voice command, and / or a second command combining a gaze command and a voice command.
[0377] In operation 1505, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to FIG. 2b, and / or the wearable electronic device (301) of FIG. 3a to FIG. 3b) can perform a function corresponding to at least one of a gesture command or a voice command while capturing video data and audio data.
[0378] According to one embodiment, the wearable electronic device can, while capturing video data and audio data, reflect a control command to an object that is the target of the control command, and then generate edited video data that can indicate that the control command has been reflected to the object or edited video data that can display the content of the control command reflected in the video data, and display the edited video data through a display (e.g., the display (360) of FIG. 3).
[0379] In operation 1507, a wearable electronic device (e.g., electronic device (101) of FIG. 1, electronic device (200) of FIG. 2a to 2b, and / or wearable electronic device (301) of FIG. 3a to 3b) may store information about a first time when at least one of a gesture command or a voice command is confirmed while capturing video data and audio data as editing information.
[0380] According to one embodiment, a wearable electronic device may store information about a first time in which at least one of a plurality of commands is identified as editing information while performing a function corresponding to a gesture command, a voice command of a designated user, or a first command combining a gesture command and a voice command of a designated user, while performing a function while capturing image data through a first camera and a second camera of the electronic device.
[0381] According to one embodiment, at least one command may include a gesture command, a voice command of a designated user, a first command combining the gesture command and the voice command of the designated user, and / or a second command combining a gaze command and a voice command.
[0382] According to one embodiment, the first time may represent the time during which a gesture command is displayed in the captured video data while capturing video data and audio data for editing the gesture command.
[0383] According to one embodiment, the first time may represent the time during which the voice command of a designated user is output while video data and audio data are being captured for editing the voice command of a designated user.
[0384] According to one embodiment, the first time may represent a time including the time when the gesture command is displayed in the captured video data and the time when the voice command of the specified user is output while capturing video data and audio data for editing a first command in which the gesture command and the voice command of the specified user are combined.
[0385] According to one embodiment, a wearable electronic device can identify a designated user's voice command in audio received through the microphone of the electronic device while capturing video data and audio data using a first AI model (e.g., the first AI model of FIG. 3b (331a)), and identify the time from the start time of output of the identified voice command to the end time of output as the first time of identifying the voice command.
[0386] According to one embodiment, a wearable electronic device can identify a gesture command included in the captured image data while capturing image data and audio data using a second AI model (e.g., the second AI model of FIG. 2b (331b)), and identify the time from the start time of display of the identified gesture command to the end time of display as the first time of identifying the gesture command.
[0387] In operation 1509, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can check whether the recording of video data and audio data has ended.
[0388] In operation 1509, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to FIG. 2b, and / or the wearable electronic device (301) of FIG. 3a to FIG. 3b) can perform an editing operation for at least one of a gesture command or a voice command based on editing information in operation 1511 when it confirms that the recording of video data and audio data has ended.
[0389] According to one embodiment, a wearable electronic device can perform an editing operation to delete at least one of a plurality of commands from image data based on editing information.
[0390] According to one embodiment, at least one command may include a gesture command, a voice command of a designated user, a first command combining the gesture command and the voice command of the designated user, and / or a second command combining a gaze command and a voice command.
[0391] According to one embodiment, a wearable electronic device can use a second AI model (e.g., the second AI model of FIG. 3b (331b)) to identify a gesture command included in at least one frame corresponding to a first time in the image data based on editing information, and delete the gesture command included in at least one frame.
[0392] According to one embodiment, a wearable electronic device can correct a first region (an image corresponding to a gesture command) corresponding to a deleted gesture command among at least one frame based on a previous frame or a next frame of at least one frame using a second AI model (e.g., the second AI model of FIG. 3b (331b)).
[0393] According to one embodiment, a wearable electronic device may use a second AI model to determine the deletion of at least one frame containing a gesture command or the deletion of a gesture command included in at least one frame based on the result of comparing the size of a first area corresponding to a gesture command and a first threshold value for determining the ratio of a gesture command to a frame in at least one frame.
[0394] According to one embodiment, an operation to edit a gesture command based on the comparison result of a first threshold value for determining the size of a first area corresponding to a gesture command and the ratio of the gesture command to at least one frame can be described in detail in FIG. 17.
[0395] According to one embodiment, a wearable electronic device may use a second AI model to determine the deletion of at least one frame containing a gesture command or the deletion of a gesture command included in at least one frame based on the result of comparing the scene change rate between at least one frame containing a gesture and the previous frame of at least one frame with a second threshold value.
[0396] According to one embodiment, an operation to edit a gesture command based on the result of comparing the scene change rate between at least one frame containing a gesture and the previous frame of at least one frame and a second threshold value can be described in detail in FIG. 18.
[0397] According to one embodiment, a wearable electronic device can use a first AI model (e.g., the first AI model of FIG. 3b (331a)) to separate and delete a designated user's voice command included in audio data identified at a first time.
[0398] According to one embodiment, when a wearable electronic device deletes a voice command of a designated user from audio data identified at a first time using a first AI model, if the audio data identified at the first time is unnatural based on previous audio data and next audio data, the audio data from which the voice command of the designated user identified at the first time has been deleted can be corrected or new audio data can be generated.
[0399] In operation 1513, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) may store a completed video file or a completed audio file.
[0400] According to one embodiment, a wearable electronic device may create a completed video file or a completed audio file and store it in the memory of the electronic device (e.g., memory (330) of FIG. 3a) after performing an editing operation to delete a gesture command included in at least one frame corresponding to a first time in the video data, a voice command of a designated user identified at the first time, a first command combined with a gesture command included in at least one frame corresponding to the first time and a voice command of a designated user identified at the first time, or a second command combined with a gaze command identified at the first time and a voice command.
[0401] According to one embodiment, when a wearable electronic device performs an editing operation to generate a finished video file, it may store in memory at least one of an original video file in which control commands are not reflected based on user input, an edited video file generated based on edited video data, a finished video file, and / or a finished audio file.
[0402] FIG. 16 is a flowchart illustrating an operation for editing video data while capturing video data and audio data in a wearable electronic device according to one embodiment. The operations for editing video data may include operations 1601 to 1611. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, at least two operations may be performed in parallel, or other operations may be added.
[0403] According to one embodiment, in operation 1601, a wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) may start capturing video data and audio data.
[0404] According to one embodiment, the wearable electronic device can perform a shooting function for image data through a first camera (381) (e.g., camera (211) of FIG. 2a and / or the first camera (381) of FIG. 3)) and a second camera (e.g., the second camera (212) of FIG. 2a and / or the second camera (383)) among the cameras of the electronic device (e.g., camera (380) of FIG. 3a) that can acquire image data related to the surrounding environment of the wearable electronic device while the wearable electronic device is worn by a part of the user's body.
[0405] According to one embodiment, the wearable electronic device can perform a shooting function for image data through one of a first camera and a second camera capable of acquiring image data related to the surrounding environment.
[0406] According to one embodiment, the wearable electronic device can start a shooting function for video data and audio data when it confirms the selection of a shooting function button among at least one button displayed through the display of the electronic device (e.g., the display (360) of FIG. 3a) based on user input.
[0407] According to one embodiment, when a wearable electronic device receives a voice command requesting a shooting function through the microphone of the electronic device (e.g., the microphone (350) of FIG. 3a), if the received voice command is identified as a voice command of a designated user, the device can interpret the voice command using at least one artificial intelligence model (e.g., the first AI model (331a)) among at least one AI model (331) and perform a shooting function corresponding to the interpreted voice command.
[0408] In operation 1603, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to FIG. 2b, and / or the wearable electronic device (301) of FIG. 3a to FIG. 3b) can check at least one command among gesture commands or voice commands while capturing video data and audio data.
[0409] According to one embodiment, while the wearable electronic device is taking video data through the first camera and the second camera of the electronic device, it can confirm the reception of a gesture command, a voice command of a designated user, or a second command combined with a gaze command and a voice command based on at least one of the first camera or the second camera and the microphone of the electronic device (e.g., the microphone (350) of FIG. 3a).
[0410] According to one embodiment, the wearable electronic device can confirm the input of a control command through at least one of the gaze, gesture, and voice of a user wearing the wearable electronic device while capturing video data and audio data or recording only audio data.
[0411] According to one embodiment, the wearable electronic device can confirm the reception of a specified user's voice command based on the microphone of the electronic device (e.g., the microphone (350) of FIG. 3a) while performing the capture of video data and audio data.
[0412] According to one embodiment, the control command may include a voice command of a designated user, a first command combining a gesture command and a voice command, and / or a second command combining a gaze command and a voice command.
[0413] In operation 1605, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to FIG. 2b, and / or the wearable electronic device (301) of FIG. 3a to FIG. 3b) can perform a function corresponding to at least one of a gesture command or a voice command while capturing video data and audio data.
[0414] According to one embodiment, the wearable electronic device can perform a function corresponding to a voice command of a designated user of the electronic device while capturing video data and audio data. According to one embodiment, while capturing video data and audio data, the wearable electronic device can generate edited video data that can indicate that the control command has been reflected on the object, or edited video data (e.g., additional information) that can display the content of the control command reflected in the video data, after reflecting the control command on the object that is the target of the control command, and can display the edited video data (e.g., additional information) on the video data through a display (e.g., the display (360) of FIG. 3).
[0415] In operation 1607, a wearable electronic device (e.g., electronic device (101) of FIG. 1, electronic device (200) of FIG. 2a to 2b, and / or wearable electronic device (301) of FIG. 3a to 3b) can perform an editing operation for at least one of a gesture command or a voice command.
[0416] According to one embodiment, a wearable electronic device can perform an editing operation to delete a gesture command, a first command combining a gesture command and a voice command of a designated user, or a second command combining a gaze command and a voice command from video data and audio data.
[0417] According to one embodiment, the wearable electronic device may perform an editing operation to delete a designated user's voice command from audio data. According to one embodiment, the wearable electronic device may use a second AI model (e.g., the second AI model of FIG. 3b (331b)) to identify a gesture command included in at least one frame in image data, delete the gesture command included in at least one frame, and correct a first region (an image corresponding to the gesture command) corresponding to the deleted gesture command in at least one frame based on the previous frame or the next frame of at least one frame.
[0418] According to one embodiment, a wearable electronic device may use a second AI model to determine the deletion of at least one frame containing a gesture command or the deletion of a gesture command included in at least one frame based on the result of comparing the size of a first area corresponding to a gesture command and a first threshold value for determining the ratio of a gesture command to a frame in at least one frame.
[0419] According to one embodiment, an operation to edit a gesture command based on the comparison result of a first threshold value for checking the size of a first area (an image corresponding to a gesture command) corresponding to a gesture command and the ratio of the gesture command to at least one frame can be described in detail in FIG. 14.
[0420] According to one embodiment, a wearable electronic device may use a second AI model to determine the deletion of at least one frame containing a gesture command or the deletion of a gesture command included in at least one frame based on the result of comparing the scene change rate between at least one frame containing a gesture and the previous frame of at least one frame with a second threshold value.
[0421] According to one embodiment, an operation to edit a gesture command based on the result of comparing the scene change rate between at least one frame containing a gesture and the previous frame of at least one frame and a second threshold value can be described in detail in FIG. 15.
[0422] According to one embodiment, a wearable electronic device can delete a voice command of a designated user received through a microphone of the electronic device (e.g., a microphone (350) in FIG. 3a) by using a first AI model (e.g., the first AI model (331a) of FIG. 3b).
[0423] According to one embodiment, a wearable electronic device can use a first AI model to check whether a designated user's voice command is included, and if it checks whether the audio data includes the designated user's voice command, it can separate and delete the designated user's voice command from the audio data.
[0424] According to one embodiment, a wearable electronic device can use a first AI model to correct the audio data received through the microphone or generate new audio data when the audio data received through the microphone is unnatural based on previous audio data when deleting a designated user's voice command from audio data received through the microphone of the electronic device.
[0425] In operation 1609, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can check whether the recording of video data and audio data has ended.
[0426] In operation 1609, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to FIG. 2b, and / or the wearable electronic device (301) of FIG. 3a to FIG. 3b) can save the completed video file or the completed audio file in operation 1611 when it confirms the end of shooting of the video data and audio data.
[0427] According to one embodiment, the wearable electronic device performs an editing operation while capturing video data and audio data to generate a finished video file, and when the end of capturing the video data and audio data is confirmed, it can store at least one of an original video file in which control commands are not reflected based on user input, an edited video file generated based on edited video data, a finished video file and / or a finished audio file in memory (e.g., memory (330) of FIG. 3a).
[0428] FIG. 17 is a flowchart illustrating an operation for editing gesture commands in image data in a wearable electronic device according to one embodiment. The operations for editing image data may include operations 1701 to 1715. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, at least two operations may be performed in parallel, or other operations may be added.
[0429] In operation 1701, a wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can identify at least one frame containing a gesture command (an object representing a specified gesture for performing a command).
[0430] According to one embodiment, a gesture command may represent an object representing a specified gesture for executing a command.
[0431] According to one embodiment, a wearable electronic device can identify at least one frame including an object representing a specified gesture for performing a command corresponding to a gesture command.
[0432] In operation 1703, a wearable electronic device (e.g., electronic device (101) of FIG. 1, electronic device (200) of FIG. 2a to 2b, and / or wearable electronic device (301) of FIG. 3a to 3b) can compare a first area (an image corresponding to a gesture command) with a first threshold value.
[0433] According to one embodiment, a wearable electronic device can use a second AI model (e.g., the second AI model of FIG. 2b (331b)) to compare the size of a first area corresponding to a gesture command and a first threshold value for determining the ratio of the gesture command to at least one frame.
[0434] In operation 1703, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can delete at least one frame containing a gesture command in operation 1705 if, based on the comparison result, the size of the first region is found to be greater than or equal to a first threshold value.
[0435] For example, if the size of a first area corresponding to a gesture command is small in a frame (image) containing a gesture command, the wearable electronic device may determine that it is smaller than a first threshold value and not delete the gesture command.
[0436] According to one embodiment, the first threshold value may be set by a user, or an electronic device may use a value that is set or pre-set based on situational judgment.
[0437] According to one embodiment, the first threshold value can be changed according to the scene by an AI model, and the size of the eleventh area corresponding to the gesture command can also be changed in real time according to the gesture.
[0438] According to one embodiment, the wearable electronic device can variably determine the first area based on a change in size and / or a change in shape of the first area corresponding to a gesture command.
[0439] For example, the wearable electronic device may exclude from the first area if it is determined to be a specific gesture (e.g., a gesture that does not correspond to a gesture command), or change the value of the first threshold so that it is not determined to be the first area.
[0440] According to one embodiment, a wearable electronic device can delete at least one frame containing a gesture command when it is confirmed that the size of a first area corresponding to a gesture command is greater than or equal to a first threshold value using a second AI model.
[0441] In operation 1707, a wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can determine whether at least one frame containing a gesture command is included in the last frame.
[0442] According to one embodiment, a wearable electronic device can use a second AI model to determine whether at least one frame containing a gesture command is included in the last frame.
[0443] In operation 1707, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to FIG. 2b, and / or the wearable electronic device (301) of FIG. 3a to FIG. 3b) can, if it is determined that at least one frame containing a gesture command is included in a frame excluding an intermediate frame or the last frame of the image data, in operation 1709, generate a new plurality of frames including a newly generated image based on at least one of the previous frame or the next frame of the at least one frame.
[0444] According to one embodiment, a wearable electronic device can generate a plurality of new frames including a newly generated image based on at least one of the previous frame or the next frame of the image data when it confirms, using a second AI model, that at least one frame including a gesture command is included in a frame excluding an intermediate frame or the last frame of the image data.
[0445] In operation 1711, a wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can generate a complete image file by inserting a plurality of new frames containing a newly generated image.
[0446] According to one embodiment, the wearable electronic device can create a complete video file that is naturally connected without interruption between frames by inserting a new plurality of frames, including a newly generated image, at the location of at least one deleted frame using a second AI model.
[0447] In operation 1707, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can, upon confirming that at least one frame containing a gesture command is included in the last frame of the image data, delete at least one frame containing a gesture command in operation 1713 and generate a completed image file in which at least one frame containing a gesture command is deleted in operation 1711.
[0448] According to one embodiment, a wearable electronic device can generate a complete video file with the last frame deleted by using a second AI model when it confirms that at least one frame containing a gesture command is included in the last frame of the video data.
[0449] According to one embodiment, when a wearable electronic device uses a second AI model to confirm that at least one frame containing a gesture command is included in the last frame of the image data, it can generate a complete image file in which at least one frame containing a gesture command is deleted without generating a new plurality of frames containing a newly generated image.
[0450] According to one embodiment, a wearable electronic device can generate a complete image file that does not fill in the position corresponding to at least one frame deleted from a plurality of frames when it confirms, using a second AI model, that at least one frame including a gesture command is included in the last frame of the image data.
[0451] In operation 1703, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can delete the gesture command included in at least one frame in operation 1715 and generate a completed image file with the gesture command deleted in operation 1711 when the size of the first area is confirmed to be less than or equal to the first threshold value based on the comparison result.
[0452] According to one embodiment, a wearable electronic device can delete a gesture command included in at least one frame when it is confirmed that the size of a first area corresponding to a gesture command is less than or equal to a first threshold value using a second AI model.
[0453] FIG. 18 is a flowchart illustrating an operation for editing gesture commands in image data in a wearable electronic device according to one embodiment. The operations for editing image data may include operations 1801 to 1815. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, at least two operations may be performed in parallel, or other operations may be added.
[0454] In operation 1801, a wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can identify at least one frame containing a gesture command.
[0455] In operation 1803, a wearable electronic device (e.g., electronic device (101) of FIG. 1, electronic device (200) of FIG. 2a to 2b, and / or wearable electronic device (301) of FIG. 3a to 3b) can compare a scene change rate (e.g., values of pixels included in a frame) between at least one frame containing a gesture and a previous frame of at least one frame with a second threshold value.
[0456] According to one embodiment, a wearable electronic device can use a second AI model (e.g., the second AI model of FIG. 2b (331b)) to check the rate of change of scene between at least one frame including a gesture and the previous frame of at least one frame, and compare the rate of change of scene with a second threshold value.
[0457] According to one embodiment, the wearable electronic device may decide whether to delete a gesture command included in at least one frame or to delete at least one frame containing a gesture command based on an analysis of the context and movement of the image data.
[0458] For example, a wearable electronic device may use an AI model to delete a gesture command included in at least one frame if it is determined that the frame containing the gesture command is important and meaningful in the context of the image data, and delete at least one frame containing the gesture command if it is determined that the frame containing the gesture command is not important or meaningful in the context of the image data.
[0459] For example, a wearable electronic device may use an AI model to delete a gesture command included in at least one frame when a new movement is detected in a frame containing a gesture command, and delete at least one frame containing a gesture command when a new movement is not detected in a frame containing a gesture command.
[0460] In operation 1803, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can delete at least one frame containing a gesture command in operation 1805 if, based on the comparison result, it is confirmed that the rate of change of the scene is below a second threshold value.
[0461] According to one embodiment, a wearable electronic device may use a second AI model to delete a gesture command included in at least one frame if the gesture command includes a new movement or a new object that was not identified in the previous frame.
[0462] According to one embodiment, a wearable electronic device can delete a gesture command included in at least one frame by using a second AI model to identify an object included in at least one frame containing a gesture command as a main object based on an object included in all frames of image data.
[0463] In operation 1807, a wearable electronic device (e.g., electronic device (101) of FIG. 1, electronic device (200) of FIG. 2a to 2b, and / or wearable electronic device (301) of FIG. 3a to 3b) can determine whether at least one frame containing a gesture command is included in the last frame.
[0464] According to one embodiment, a wearable electronic device can use a second AI model to determine whether at least one frame containing a gesture command is included in the last frame.
[0465] In operation 1807, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can, if it is determined that at least one frame containing a gesture command is included in a frame excluding the middle frame or the last frame of the image data, in operation 1809, generate a new plurality of frames including a newly generated image based on at least one of the previous frame or the next frame of the at least one frame.
[0466] According to one embodiment, a wearable electronic device can generate a plurality of new frames including a newly generated image based on at least one of the previous frame or the next frame of the image data when it confirms, using a second AI model, that at least one frame including a gesture command is included in a frame excluding an intermediate frame or the last frame of the image data.
[0467] In operation 1811, a wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can generate a complete image file by inserting a plurality of new frames containing a newly generated image.
[0468] According to one embodiment, the wearable electronic device can create a complete video file that is naturally connected without interruption between frames by inserting a new plurality of frames, including a newly generated image, at the location of at least one deleted frame using a second AI model.
[0469] In operation 1807, the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) can, upon confirming that at least one frame containing a gesture command is included in the last frame of the image data, delete at least one frame containing a gesture command in operation 1813 and generate a completed image file in which at least one frame containing a gesture command is deleted in operation 1811.
[0470] According to one embodiment, the wearable electronic device can delete the last frame when it confirms, using a second AI model, that at least one frame containing a gesture command is included in the last frame of the image data.
[0471] According to one embodiment, a wearable electronic device can delete at least one frame containing a gesture command without generating a new plurality of frames containing a newly generated image when it confirms, using a second AI model, that at least one frame containing a gesture command is included in the last frame of the image data.
[0472] In operation 1803, if the wearable electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (200) of FIG. 2a to 2b, and / or the wearable electronic device (301) of FIG. 3a to 3b) confirms that the rate of change of the scene is below a second threshold value based on the comparison result, in operation 1815, the gesture command included in at least one frame can be deleted, and in operation 1811, a completed video file with the gesture command deleted can be generated.
[0473] According to one embodiment, a wearable electronic device can delete a gesture command included in at least one frame when it is confirmed that the size of a first area (an image corresponding to a gesture command) corresponding to a gesture command is less than or equal to a first threshold value by using a second AI model.
[0474] A method for editing image data in an electronic device according to one embodiment (101 in FIG. 1; 200 in FIG. 2a to 2b; 301 in FIG. 3a to 3c) may include an operation of starting to capture image data and audio data through each of the camera and microphone of the electronic device. The method according to one embodiment may include, while capturing the image data and audio data, an operation of performing a function corresponding to at least one command among a gesture command included in the image data or a designated user's voice command included in the audio data using an AI model, and storing information about a first time when the at least one command was confirmed in the memory of the electronic device. The method according to one embodiment includes, upon confirming the end of shooting of the video data and the audio data, an operation of saving a completed video file including the captured video data and the captured audio data, excluding the data corresponding to the at least one command confirmed at the first time, wherein the data corresponding to the at least one command may be deleted from the captured video data and the audio data based on information regarding the first time using the AI model.
[0475] The method according to one embodiment may include an operation of comparing the size of a first area corresponding to the gesture command included in at least one frame among a plurality of frames corresponding to the image data with a first threshold value. The method according to one embodiment may include an operation of deleting the at least one frame containing the gesture command if the size of the first area is greater than or equal to the first threshold value. The method according to one embodiment may further include an operation of deleting the gesture command included in the at least one frame if the size of the first area is less than or equal to the first threshold value.
[0476] The method according to one embodiment may include an operation of deleting the at least one frame containing the gesture command when it is confirmed that the control command is included in at least one frame other than the last frame among the plurality of frames. The method according to one embodiment may include an operation of creating a new at least one frame based on at least one of the previous frame or the next frame of the at least one frame. The method according to one embodiment may include an operation of inserting the new at least one frame at the location of the deleted at least one frame among the plurality of frames.
[0477] The method according to one embodiment may include an operation of deleting the at least one frame from the plurality of frames when confirming the inclusion of the gesture command in the at least one frame including the last frame among the plurality of frames. The method according to one embodiment may include an operation of generating the completed image file that does not fill the position corresponding to the deleted at least one frame among the plurality of frames.
[0478] The method according to one embodiment may include an operation of checking the rate of change of a scene between the at least one frame containing the gesture command and the previous frame of the at least one frame. The method according to one embodiment may include an operation of deleting the gesture command included in the at least one frame if the rate of change of the scene is greater than or equal to a second threshold value. The method according to one embodiment may include an operation of deleting the at least one frame containing the gesture command if the rate of change of the scene is less than or equal to the second threshold value.
[0479] The method according to one embodiment may include, when the gesture command included in the at least one frame is deleted, an operation of correcting a first area corresponding to the deleted gesture command based on at least one of the previous frame or the next frame of the at least one frame.
[0480] The electronic device according to one embodiment disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance. The electronic device according to the embodiment of this document is not limited to the aforementioned devices.
[0481] One embodiment of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, each of phrases such as "A or B", "at least one of A and B", "at least one of A or B", "A, B or C", "at least one of A, B and C", and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., first) component is referred to as “coupled” or “connected” to another (e.g., second) component, with or without the terms “functionally” or “communicationally,” it means that said component may be connected to said other component directly (e.g., wired), wirelessly, or through a third component.
[0482] The term "module" as used in an embodiment of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to an embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0483] One embodiment of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101) or electronic device (301)). For example, a processor (e.g., processor (520)) of the machine (e.g., electronic device (301)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0484] According to one embodiment, the method according to one embodiment disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0485] According to one embodiment, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to one embodiment, one or more of the components or operations among the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to one embodiment, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101 of FIG. 1; 200 of FIG. 2a to 2b; 301 of FIG. 3a to 3c), At least one camera (180 in FIG. 1; 380 in FIG. 3a); Microphone (150 in Fig. 1; 350 in Fig. 3a); At least one processor (120 in FIG. 1; 320 in FIG. 3a); and It includes memory for storing instructions (130 in FIG. 1; 330 in FIG. 3a), and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Starting to record video data and audio data through the camera and microphone respectively, and While capturing the above video data and the above audio data, if at least one command among a gesture command included in the above video data or a designated user's voice command included in the above audio data is identified using an AI model, a function corresponding to the at least one command is performed, and information regarding the first time at which the at least one command was identified is stored in the memory. When the end of recording of the above video data and the above audio data is confirmed, a completed video file including the captured video data and the captured audio data is saved, excluding the data corresponding to the at least one command confirmed at the first time. An electronic device in which data corresponding to at least one of the above commands is deleted from the captured video data and the audio data based on information about the first time using the AI model.
2. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device uses the AI model, Comparing the size of a first region corresponding to a gesture command included in at least one frame among a plurality of frames corresponding to the above image data with a first threshold value, If the size of the first area is greater than or equal to the first threshold value, the at least one frame containing the gesture command is deleted, and An electronic device that deletes the gesture command included in the at least one frame when the size of the first area is less than or equal to the first threshold value.
3. In any one of paragraphs 1 to 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device uses the AI model If it is confirmed that the above control command is included in at least one frame other than the last frame among the plurality of frames, the at least one frame containing the gesture command is deleted, and A new at least one frame is generated based on at least one of the previous frame or the next frame of the above at least one frame, and An electronic device that inserts the new at least one frame at the location of the deleted at least one frame among the plurality of frames.
4. In any one of paragraphs 1 to 3, When the above instructions are executed individually or collectively by the at least one processor, the electronic device uses the AI model If the inclusion of the gesture command is confirmed in the at least one frame including the last frame among the plurality of frames, the at least one frame is deleted from the plurality of frames, and An electronic device that generates the completed image file that does not fill the position corresponding to at least one deleted frame among the plurality of frames.
5. In any one of paragraphs 1 through 4, When the above instructions are executed individually or collectively by the at least one processor, the electronic device uses the AI model, Checking the rate of change of the scene between the at least one frame containing the above gesture command and the previous frame of the at least one frame, and If the rate of change of the above scene is greater than or equal to a second threshold value, the gesture command included in the at least one frame is deleted, and An electronic device that deletes at least one frame including the gesture command if the rate of change of the above scene is less than or equal to the above 2 threshold value.
6. In any one of paragraphs 1 through 5, When the above instructions are executed individually or collectively by the at least one processor, the electronic device uses the AI model, An electronic device that, when a gesture command included in at least one frame is deleted, corrects a first region corresponding to the deleted gesture command based on at least one of the previous frame or the next frame of the at least one frame.
7. In any one of paragraphs 1 through 6, The electronic device comprising at least one of the above-mentioned command, a gesture command, a voice command of a designated user, a first command combining a gesture command and a voice command, or a second command combining an eye-tracking command and a voice command of a designated user.
8. A method for editing image data in an electronic device (101 of FIG. 1; 200 of FIG. 2a to 2b; 301 of FIG. 3a to 3c), The operation of starting to capture video data and audio data through the camera and microphone of the electronic device, respectively; While capturing the above video data and the above audio data, if at least one command among a gesture command included in the above video data or a designated user's voice command included in the above audio data is identified using an AI model, a function corresponding to the at least one command is performed, and information regarding a first time at which the at least one command was identified is stored in the memory of the electronic device; and When the end of recording of the above video data and the above audio data is confirmed, the operation includes saving a completed video file comprising the above-captured video data and the above-captured audio data, excluding the data corresponding to the at least one command confirmed at the first time. A method in which data corresponding to at least one of the above commands is deleted from the captured video data and the audio data based on information about the first time using the AI model.
9. In a non-volatile storage medium storing instructions, said instructions are configured to cause said electronic device to perform at least one operation when executed by said electronic device, said at least one operation being, The operation of starting to capture video data and audio data through the camera and microphone of the electronic device, respectively; While capturing the above video data and the above audio data, if at least one command among a gesture command included in the above video data or a designated user's voice command included in the above audio data is identified using an AI model, a function corresponding to the at least one command is performed, and information regarding a first time at which the at least one command was identified is stored in the memory of the electronic device; and When the end of recording of the above video data and the above audio data is confirmed, the operation includes saving a completed video file comprising the above-captured video data and the above-captured audio data, excluding the data corresponding to the at least one command confirmed at the first time. The data corresponding to at least one of the above commands is a storage medium that deletes the captured video data and the audio data based on information about the first time using the AI model.
10. In an electronic device, mike; At least one camera; At least one processor; and Includes memory for storing instructions; and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, While shooting or recording at least one of video or audio, control commands are received through at least one of the gaze, gestures, and voice of a user wearing the electronic device, and Generate video data or audio data by reflecting the above control command, An electronic device that, when at least one of the gesture and the voice is included in the video data or the audio data, uses an AI model to generate a complete video file or a complete audio file in which the gesture or the voice is removed from the video data or the audio data.
11. In Paragraph 10, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, If the above video is a video, Generating the completed video file by removing at least some of the gesture from a plurality of frames containing the gesture through the above AI model, or An electronic device that generates a completed video file by replacing a plurality of frames containing the above gesture with a plurality of new frames containing a newly generated video by the AI model.
12. In any one of paragraphs 10 to 11, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Multiple frames containing the above gesture are continuous for a certain period, and When generating the completed video file by replacing the multiple frames containing the above gesture with the new multiple frames containing a video newly generated by the AI model, An electronic device that generates an image corresponding to a specific audio signal among audio data other than the voice included within the above-mentioned fixed interval through the AI model and adds it to the newly generated image.
13. In any one of paragraphs 10 to 12, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, If the above image is a still image, An electronic device that generates the completed image file by removing a first region corresponding to the gesture from the above image data.
14. In any one of paragraphs 10 through 13, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, The above electronic device further includes at least one of a display and a speaker, or The above image data being captured is transmitted to an external electronic device including a display, and An electronic device that, while capturing the above video, controls to visually display an object to be the target of the control command within the above video on a display included in the electronic device or a display included in the external electronic device, or to audibly output a sound representing the object to be the target of the control command through the speaker.
15. In any one of paragraphs 10 through 14, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, The above electronic device is, Includes additional displays, or The above image data being captured is transmitted to an external electronic device including a display, and Edited video data indicating that the above control command has been reflected, or edited video data indicating the content of the above control command reflected in the above video data, is displayed on a display included in the electronic device or a display included in the external electronic device, and An electronic device that stores at least one of an original video file in which the above control command is not reflected, an edited video file generated based on the above edited video data, and the above completed video file.