Electronic device for correcting image and control method therefor
Patent Information
- Application Number
- PCT/KR2025/002552
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-28
- Filing Date
- 2025-02-24
- Publication Date
- 2025-10-02
AI Technical Summary
Conventional image stabilization techniques require auxiliary devices or manual selection of anti-shake modes, which are inconvenient for users.
An electronic device uses a learning model to identify overlapping areas between frames, generate missing parts of frames using a prompt, and synthesize them to correct image shake through a simple user interface.
Enhances user convenience by automatically compensating for image shake without the need for auxiliary devices or manual settings.
Smart Images

Figure KR2025002552_02102025_PF_FP_ABST
Abstract
Description
Electronic device for correcting images and method for controlling the same
[0001] The present disclosure relates to an electronic device for correcting an image and a method for controlling the same.
[0002] The variety of services and additional features offered through electronic devices, such as smartphones, is steadily increasing. To enhance the utility of these devices and satisfy the diverse needs of users, telecommunications service providers and electronic device manufacturers are competitively developing electronic devices that offer diverse features and differentiate themselves from competitors. Accordingly, the various functions offered through wearable devices are also becoming increasingly sophisticated.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0004] According to related technologies for compensating for image shake, conventional techniques have required the use of auxiliary devices such as gimbals to prevent image shake, or the selection of an anti-shake mode before shooting, which has been inconvenient. According to one embodiment of the present disclosure, image shake can be compensated for using a learning model with a simple user interface manipulation, thereby providing an electronic device with enhanced user convenience.
[0005] An electronic device according to one embodiment of the present disclosure includes a display, at least one processor, and a memory, wherein the memory may be configured to store instructions that, when executed by the at least one processor, cause the electronic device to identify a plurality of frames of at least one selected video based on a user's selection input for at least one video, determine an overlapping area between a first frame and a second frame among the identified plurality of frames, generate a first prompt expressed to generate a part of the first frame that is not included in the second frame based on the determined overlapping area, generate a part of the first frame using the generated first prompt and a learning model, generate a third frame by synthesizing the part of the generated first frame with the second frame, and display the generated third frame through the display.
[0006] A method according to one embodiment of the present disclosure may include: identifying a plurality of frames of at least one selected video based on a user's selection input for at least one video; determining an overlapping area between a first frame and a second frame among the identified plurality of frames; generating a first prompt expressed to generate a part of the first frame that is not included in the second frame based on the determined overlapping area; generating a part of the first frame using the generated first prompt and a learning model; generating a third frame by synthesizing a part of the generated first frame with the second frame; and displaying the generated third frame through the display.
[0007] A computer-readable non-transitory recording medium according to one embodiment of the present disclosure is configured to store a plurality of instructions, wherein the plurality of instructions, when executed by at least one processor of an electronic device, may include instructions that cause the electronic device to: identify a plurality of frames of at least one selected video based on a user's selection input for at least one video; determine an overlapping area between a first frame and a second frame among the identified plurality of frames; generate a first prompt expressed to generate a part of the first frame that is not included in the second frame based on the determined overlapping area; generate a part of the first frame using the generated first prompt and a learning model; generate a third frame by synthesizing the part of the generated first frame with the second frame; and display the generated third frame through the display.
[0008] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments of the present disclosure.
[0009] FIG. 2 is an exemplary block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0010] FIG. 3 is an exemplary drawing for explaining a function or operation of an electronic device according to one embodiment of the present disclosure to correct a shaken image using a learning model.
[0011] FIGS. 4A, 4B, 4C, and 4D are exemplary drawings for explaining a function or operation of an electronic device according to one embodiment of the present disclosure to identify an overlapping area of frames (e.g., a first frame and a second frame) to compensate for shaking of an image.
[0012] FIG. 5 is an exemplary drawing for explaining a function or operation of an electronic device according to one embodiment of the present disclosure to generate a part to be synthesized into a second frame using a learning model.
[0013] FIG. 6 is an exemplary drawing for explaining a function or operation of an electronic device according to one embodiment of the present disclosure to generate a third frame (e.g., a corrected second frame).
[0014] FIG. 7 is an exemplary drawing for explaining a function or operation of an electronic device according to one embodiment of the present disclosure for correcting shaking of at least one object included in an image and then outputting an image including at least one corrected object.
[0015] FIG. 8 is an exemplary drawing for explaining a case where an image captured by an electronic device according to one embodiment of the present disclosure includes shaking.
[0016] FIG. 9 is an exemplary drawing for explaining a function or operation of an electronic device according to one embodiment of the present disclosure to generate a shake-compensated frame using a plurality of frames (e.g., a fourth frame, a fifth frame, and a sixth frame).
[0017] FIG. 10 is an exemplary drawing for explaining a function or operation of an electronic device according to one embodiment of the present disclosure to generate a shake-compensated frame using a plurality of frames (e.g., a seventh frame, an eighth frame, and a ninth frame).
[0018] FIG. 11 is an exemplary drawing for explaining a function or operation of an electronic device according to an embodiment of the present disclosure to correct shaking of an image by replacing some frames with frames generated using a learning model when it is determined that the ratio of overlapping areas between a plurality of frames is less than a threshold ratio.
[0019] Figure 12 is an example drawing for explaining the function or operation illustrated in Figure 11 from a user interface perspective.
[0020] FIG. 13 is an exemplary drawing for explaining a function or operation of an electronic device according to one embodiment of the present disclosure to move a display location of an object of interest included in an image into a region of interest and generate and display another part of the image based on the movement using a learning model.
[0021] FIG. 14a, FIG. 14b, FIG. 14c, FIG. 14d, FIG. 14e, and FIG. 14f are exemplary drawings for explaining the function or operation illustrated in FIG. 13 from a user interface perspective.
[0022] FIG. 15 is an exemplary drawing for explaining a function or operation of an electronic device according to one embodiment of the present disclosure to generate a reference image (e.g., a second spatial map) for performing out-painting and to perform out-painting using the reference image generated for a plurality of frames.
[0023] FIGS. 16A, 16B, 16C, 16D, and 16E are exemplary drawings illustrating a plurality of frames required for an electronic device to generate a first spatial map according to an embodiment of the present disclosure.
[0024] FIGS. 16f, 16g, 16h, 16i, and 16j are exemplary drawings for explaining an area in which outpainting should be performed when an electronic device according to one embodiment of the present disclosure moves and displays the display position of at least one object included in an image.
[0025] FIG. 17A is an exemplary drawing for explaining a first spatial map generated by an electronic device according to one embodiment of the present disclosure.
[0026] FIG. 17b is an exemplary drawing for explaining a second spatial map generated by an electronic device according to one embodiment of the present disclosure.
[0027] FIG. 18 is an exemplary diagram for explaining a generative artificial intelligence system according to one embodiment of the present disclosure.
[0028] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.
[0029] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0030] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0031] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0032] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0033] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0034] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0035] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0036] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0037] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0038] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0039] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0040] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0041] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0042] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0043] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0044] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0045] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0046] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0047] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0048] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0049] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0050] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0051] FIG. 2 is an exemplary block diagram for explaining the configuration of an electronic device (101) according to one embodiment of the present disclosure.
[0052] Referring to FIG. 2, an electronic device (101) according to an embodiment of the present disclosure may include a video analysis unit (210), an out-painting model unit (220), a motion image generation model unit (230), a loss rate setting unit (240), and a UI management unit (250). The video analysis unit (210) according to an embodiment of the present disclosure may include a video frame extraction unit (212), an inter-frame overlapping area determination unit (214), an extended area analysis model unit (216), and a frame characteristic analysis unit (218).
[0053] The video frame extraction unit (212) according to one embodiment of the present disclosure can extract or identify at least one frame included in a designated video. The designated video according to one embodiment of the present disclosure may include, for example, an video selected by a user. The inter-frame overlapping area determination unit (214) according to one embodiment of the present disclosure can determine an overlapping area between at least two frames. The inter-frame overlapping area determination unit (214) according to one embodiment of the present disclosure can determine a ratio that the overlapping area occupies in a designated frame (e.g., the first frame (410)). The extended region analysis model unit (216) according to one embodiment of the present disclosure can use at least some of the frames included in the designated video as input for outputting or generating a prompt. The extended region analysis model unit (216) according to one embodiment of the present disclosure can determine an overlapping area between neighboring frames, or an overlapping area (IoU; Intersection of Union) between a reference frame and other frames (e.g., non-neighboring) in at least some of the frames used as input. The extended area analysis model unit (216) according to one embodiment of the present disclosure can determine an extended area required in a frame of the image (e.g., the second frame (420)) to reduce shaking errors in the image when the determination of overlapping areas for at least some frames is completed, and can generate and output a prompt based on the determination result. The frame characteristic analysis unit (218) according to one embodiment of the present disclosure can identify properties of a specified frame (e.g., resolution, number of objects included in the frame, type of objects).
[0054] The out-painting model unit (220) according to one embodiment of the present disclosure can use at least some frames of an image containing shaking errors and a prompt output from the extended region analysis model unit (216) as inputs for performing out-painting. The out-painting according to one embodiment of the present disclosure can be performed by a learning model (e.g., a generative artificial intelligence model). The out-painting model unit (220) according to one embodiment of the present disclosure can output at least one frame on which out-painting has been performed.
[0055] The motion image generation model unit (230) according to one embodiment of the present disclosure generates a motion image to replace an image (e.g., a frame) in which a shaking error is greater than or equal to a threshold value (e.g., the ratio of overlapping areas between neighboring frames is less than a threshold value, the displayed area of a major object is less than a threshold value, and the major area loss area within a major object area is greater than or equal to a threshold value). According to one embodiment of the present disclosure, the criteria for shaking errors described above can be equally applied to various functions or operations described in the present disclosure in addition to the function or operation of generating a motion image. The motion image generation model unit (230) according to one embodiment of the present disclosure can use as input the frames before and after a frame (e.g., a frame section) in which a shaking error occurs (e.g., the 10th frame (1210) and the 13th frame (1240)). The motion image generation model unit (230) according to one embodiment of the present disclosure can compare the previous and subsequent frames to determine an overlapping area between the previous and subsequent frames. According to one embodiment of the present disclosure, the motion image generation model unit (230) may generate at least one frame between previous and subsequent frames based on the determined overlapping region, using the learning model and previous and subsequent frames. In this case, the number of at least one generated frame may correspond to the number of frames to be deleted according to generation, or may correspond to the number of frames within the frame section determined to be a shaking error frame. Alternatively, the number of at least one generated frame may correspond to the number of at least one generated frame when at least one frame is generated using the previous and / or subsequent frames. According to one embodiment of the present disclosure, when a section determined to have a shaking error is determined, image videos before / after the corresponding shaking error section may be transmitted as input to an AI model for motion generation.In this case, in a step prior to transmitting as input, the electronic device (101) may identify an area where the largest difference occurs between the previous and next frames, and perform a function or operation to identify an object corresponding to the area. For example, in FIG. 12, the electronic device (101) according to an embodiment of the present disclosure may identify an object corresponding to an area where a difference occurs between the 10th frame (1210) and the 13th frame (1240), which are previous and next frames, as a user's arm. When the electronic device (101) according to an embodiment of the present disclosure inputs the previous and next frames to an AI model (e.g., the generative AI model (1830) of FIG. 18), the electronic device may transmit a prompt indicating "generate a motion image (frame) of a person's arm moving between two images" as an input to the AI model based on the analyzed object information.
[0056] The loss rate setting unit (240) according to one embodiment of the present disclosure may crop an image according to a specified loss rate for at least one frame generated by the outpainting model unit (220). The specified loss rate according to one embodiment of the present disclosure may include a loss rate specified by a user. According to one embodiment of the present disclosure, when outpainting is performed, an additional area is created in the frame, so the size of the data increases and areas not intended by the user may be created. To solve this problem, a crop area may be specified by the user. An interface for specifying such a crop area may be provided through a display (e.g., a display module (160)) of the electronic device (101). The shape and / or display method of the interface for specifying the crop area according to one embodiment of the present disclosure may be determined by the UI management unit (250). The electronic device (101) according to one embodiment of the present disclosure may provide a screen according to the user's crop as a preview screen through a display (e.g., a display module (160)).
[0057] FIG. 3 is an exemplary diagram for explaining a function or operation of an electronic device (101) according to an embodiment of the present disclosure to correct a shaken image using a learning model (510). FIGS. 4A to 4D are exemplary diagrams for explaining a function or operation of an electronic device according to an embodiment of the present disclosure to identify an overlapping area of frames (e.g., a first frame and a second frame) to correct shake in an image.
[0058] Referring to FIG. 3, an electronic device (101) (e.g., processor (120)) according to an embodiment of the present disclosure may, in operation 310, identify a plurality of frames of at least one selected video based on a user's selection input for at least one video. The at least one video according to an embodiment of the present disclosure may include an image including a plurality of frames (e.g., a first frame (410) and a second frame (420)) having shaking errors, as illustrated in FIGS. 4A and 4B . The electronic device (101) according to an embodiment of the present disclosure may determine that the image has shaking errors, for example, if the ratio of an overlapping area (430) exceeds a predetermined ratio based on a designated frame (e.g., the first frame (410)). Operation 310 according to an embodiment of the present disclosure may be performed, for example, by a video frame extraction unit (212).
[0059] An electronic device (101) (e.g., processor (120)) according to an embodiment of the present disclosure may, in operation 320, determine an overlapping area (430) between a first frame (410) and a second frame (420) among a plurality of frames (e.g., all frames included in an image) identified according to operation 310. The first frame and the second frame described in the present disclosure are expressed as neighboring frames for the purpose of explaining various embodiments of the present disclosure, but the first frame and the second frame may not be neighboring each other, and a plurality of frames may exist between the first frame and the second frame. The first frame (410) according to an embodiment of the present disclosure may include, for example, a start frame (e.g., a first frame) of an image, and the second frame (420) may include a frame immediately after the first frame. An electronic device (101) according to an embodiment of the present disclosure can determine an overlapping area (430) between a first frame (410) and a second frame (420), as illustrated in FIGS. 4C and 4D . An electronic device (101) according to an embodiment of the present disclosure can identify a ratio of a portion remaining in the first frame (410) excluding the overlapping area (430). Operation 320 according to an embodiment of the present disclosure can be performed, for example, by an inter-frame overlapping area determination unit (214).
[0060] An electronic device (101) (e.g., processor (120)) according to an embodiment of the present disclosure may generate a first prompt expressed to generate an area corresponding to a portion of the first frame (410) that is not included in the second frame (420) based on the overlapping area (430) determined according to operation 320 in operation 330. FIG. 5 is an exemplary drawing for explaining a function or operation of an electronic device (101) according to an embodiment of the present disclosure to generate at least a portion of a third frame (e.g., a modified second frame) (e.g., using a learning model (510), generating a portion (520) to be synthesized into the second frame (420) and / or using a learning model (510), generating at least a portion of an object that is included in the first frame (410) but not included in the second frame (420). An electronic device (101) according to an embodiment of the present disclosure may generate a prompt expressed as, for example, “generate about 15% of the lower area of the second frame (420) based on the first frame (410)” as illustrated in FIG. 5 . The electronic device (101) according to an embodiment of the present disclosure may input the generated prompt as input data of the learning model (510). The electronic device (101) according to an embodiment of the present disclosure may also input the first frame (410) together with the generated prompt as input data of the learning model (510) into the learning model (510). The learning model (510) according to an embodiment of the present disclosure may be stored in the electronic device (101) or may be stored in an external electronic device (e.g., an artificial intelligence server) that is operable with the electronic device (101).
[0061] According to an embodiment of the present disclosure, the electronic device (101) (e.g., the processor (120)) may, in operation 340, use the first prompt and learning model (520) generated according to operation 330 to generate an extended area of the second frame based on a portion of the first frame (410) (e.g., a portion (520) to be synthesized into the second frame (420)). In order to generate the extended area of the second frame, the electronic device (101) (e.g., the processor (120)) according to an embodiment of the present disclosure may identify at least a portion of at least one object included in the first frame and not included in the second frame. The electronic device (101) (e.g., the processor (120)) according to an embodiment of the present disclosure may generate at least one object using the learning model. When generating at least one object, the electronic device (101) (e.g., the learning model) according to an embodiment of the present disclosure may generate at least one object by considering a time difference between frames. For example, an electronic device (101) (e.g., a learning model) according to one embodiment of the present disclosure may generate at least one object based on a position to which the at least one object moves over time and / or a shape that changes over time when generating at least a portion of an object that moves over time.
[0062] An electronic device (101) (e.g., processor (120)) according to one embodiment of the present disclosure can generate a third frame (610) by synthesizing an extended area of a second frame (420) generated based on a partial area of a first frame (410) generated according to operation 340, in operation 350. A learning model (510) according to one embodiment of the present disclosure can generate a portion (520) to be synthesized into the second frame (420) using an input prompt and the first frame (410). FIG. 6 is an exemplary drawing for explaining a function or operation of an electronic device (101) according to one embodiment of the present disclosure generating a third frame (e.g., a corrected second frame) (610) by synthesizing a portion generated by the learning model (520) and the second frame (420). An electronic device (101) according to one embodiment of the present disclosure may generate a third frame (610) as illustrated in FIG. 6 by synthesizing a portion (520) to be synthesized into a second frame (420) and the second frame (420) (e.g., by performing out-painting on the second frame). According to one embodiment of the present disclosure, when generating the third frame (610), a portion (e.g., an upper portion) of the second frame (420) may be deleted (e.g., cropped) and the remaining area may be synthesized with the portion to be synthesized into the second frame, thereby generating a third frame corresponding to the resolution of the original image. According to one embodiment of the present disclosure, the portion (520) to be synthesized into the second frame (420) may be obtained from an external electronic device (e.g., an artificial intelligence server) or generated by the electronic device (101).
[0063] An electronic device (101) (e.g., processor (120)) according to an embodiment of the present disclosure may display the third frame generated according to operation 350 through a display (e.g., display module (160)) in operation 360. An electronic device (101) according to an embodiment of the present disclosure may replace the second frame (420) with the generated third frame (610) and display the same through a display (e.g., display module (160)). Accordingly, an image with shake errors removed may be output through the electronic device (101).
[0064] FIG. 7 is an exemplary drawing for explaining a function or operation of an electronic device (101) according to one embodiment of the present disclosure for correcting shaking of at least one object included in an image and then outputting an image including at least one corrected object.
[0065] Referring to FIG. 7, an electronic device (101) (e.g., a processor (120)) according to an embodiment of the present disclosure may determine, in operation 710, whether shaking has occurred for at least one object (e.g., a main object) included in a plurality of frames. FIG. 8 is an exemplary diagram for explaining a case in which an image captured by an electronic device (101) according to an embodiment of the present disclosure is captured including shaking. Referring to FIG. 8, an electronic device (101) according to an embodiment of the present disclosure may identify that an image having a shaking error has been captured based on a plurality of frames (e.g., a fourth frame (820), a fifth frame (830), and a sixth frame (840)) included in a specified image.
[0066] According to an embodiment of the present disclosure, the electronic device (101) (e.g., the processor (120)) may, in operation 720, correct shaking for at least one object (e.g., a main object) by using information about at least one object (e.g., a main object) included in a plurality of frames (e.g., a fourth frame (820), a fifth frame (830), and a sixth frame (840)), based on determining that shaking has occurred according to operation 710. FIG. 9 is an exemplary diagram for explaining a function or operation of the electronic device (101) according to an embodiment of the present disclosure to generate a frame (850) in which shaking is corrected by using a plurality of frames (e.g., a fourth frame (820), a fifth frame (830), and a sixth frame (840)). Referring to FIG. 9, the electronic device (101) according to an embodiment of the present disclosure may identify at least one object included in a specified image. For example, an electronic device (101) according to an embodiment of the present disclosure may identify that a first object (e.g., a person) and a second object (e.g., a cloud) are included in a specified image. An electronic device (101) according to an embodiment of the present disclosure may identify that a smile-shaped object is included in a fourth frame (820). An electronic device (101) according to an embodiment of the present disclosure may identify that a first object (e.g., a person) and a second object (e.g., a cloud) are included in a fifth frame (830). An electronic device (101) according to an embodiment of the present disclosure may identify that a first object (e.g., a person), a smile-shaped object, and a sea-shaped object are included in a sixth frame (840). An electronic device (101) according to one embodiment of the present disclosure may generate a prompt for generating a corrected frame (850) by using information about at least one object included in a plurality of frames (e.g., a fourth frame (820), a fifth frame (830), a sixth frame (840)).A prompt for generating a corrected frame (850) according to one embodiment of the present disclosure may include a prompt expressed as "Add clouds on top of the first frame and insert a sun shape under the smile shape of the person's clothes." The electronic device (101) according to one embodiment of the present disclosure may transmit the generated prompt and / or at least one frame (e.g., the fourth frame (820), the fifth frame (830), and the sixth frame (840)) to an external electronic device (e.g., an artificial intelligence server) or input the generated prompt and / or at least one frame as input data of a learning model stored in the electronic device (101). The electronic device (101) according to one embodiment of the present disclosure may obtain information about the corrected frame (850) from the external electronic device or obtain the information as output data of a learning model stored in the electronic device (101). A corrected frame (850) according to one embodiment of the present disclosure may include frames including a first object (e.g., a person), a second object (e.g., a cloud), and objects having various shapes (e.g., a smile-shaped object, a sea-shaped object). A corrected frame (850) according to one embodiment of the present disclosure may include frames that have the same effect as if they were captured with a wider angle of view than frames included in an image captured by the electronic device (101) (e.g., the fourth frame (820), the fifth frame (830), and the sixth frame (840)). In FIG. 9, the corrected frame (850) is exemplarily illustrated as one frame, but according to one embodiment of the present disclosure, the corrected frame (850) may be generated for each of the fourth frame (820), the fifth frame (830), and the sixth frame (840). An electronic device (101) according to one embodiment of the present disclosure can generate a corrected image based on a corrected frame generated for each of the fourth frame (820), the fifth frame (830), and the sixth frame (840).An electronic device (101) according to one embodiment of the present disclosure can output a corrected image automatically or according to a user's input.
[0067] FIG. 10 is an exemplary drawing for explaining a function or operation of an electronic device (101) according to one embodiment of the present disclosure to generate a shake-corrected frame (1010) using a plurality of frames (e.g., a seventh frame (1020), an eighth frame (1030), and a ninth frame (1040)).
[0068] Referring to FIG. 10, an electronic device (101) according to an embodiment of the present disclosure may determine a reference frame from among a plurality of frames. For example, the electronic device (101) according to an embodiment of the present disclosure may determine a first frame as a reference frame (e.g., the seventh frame (1020)). The electronic device (101) according to an embodiment of the present disclosure may identify an overlapping area between the reference frame and at least one subsequent frame (e.g., the eighth frame (1030) and the ninth frame (1040)). The electronic device (101) according to an embodiment of the present disclosure may generate a prompt expressed to add a portion (e.g., a smile-shaped object included in the seventh frame (1020)) that is not included in the subsequent frame (e.g., the eighth frame (1030)) and delete a portion (e.g., a cloud included in the eighth frame (1030)) that is not included in the reference frame (e.g., the seventh frame (1020)). According to one embodiment of the present disclosure, a corrected frame (1010) may be generated for each of at least one subsequent frame (e.g., the eighth frame (1030) and the ninth frame (1040)). In FIG. 10, at least one object (e.g., a cloud in the eighth frame (1030) and a sun in the ninth frame (1040)) to be deleted from the frames (e.g., the eighth frame (1030) and the ninth frame (1040)) is illustrated as a dotted line to generate the corrected frame. The electronic device (101) according to one embodiment of the present disclosure may generate a prompt expressed to add a portion (e.g., a part of a person's face included in the seventh frame (1020)) that is not included in the subsequent frame (e.g., the ninth frame (1040)) and delete a portion (e.g., a sun-shaped object included in the ninth frame (1040)) that is not included in the reference frame (e.g., the seventh frame (1020)). An electronic device (101) according to one embodiment of the present disclosure can transmit the generated prompt to an external electronic device or input it as input data of a learning model stored in the electronic device (101).An electronic device (101) according to an embodiment of the present disclosure may transmit at least one frame (e.g., the seventh frame (1020), the eighth frame (1030), and / or the ninth frame (1040)) to an external electronic device. An electronic device (101) according to an embodiment of the present disclosure may obtain at least one corrected frame (e.g., the corrected eighth frame and / or the corrected ninth frame) from the external electronic device. Alternatively, an electronic device (101) according to an embodiment of the present disclosure may obtain at least one corrected frame (e.g., the corrected eighth frame and / or the corrected ninth frame) as output data of a learning model. The at least one corrected frame (e.g., the corrected eighth frame and / or the corrected ninth frame) may correspond to the corrected frame (1010) illustrated in FIG. 10. An electronic device (101) according to one embodiment of the present disclosure can replace at least one frame (e.g., the eighth frame (1030) and / or the ninth frame (1040)) containing a shaking error with at least one corrected frame. The electronic device (101) according to one embodiment of the present disclosure can output an image including the replaced frame through the electronic device (101). Accordingly, an image with the shaking error removed can be output through the electronic device (101).
[0069] FIG. 11 is an exemplary drawing for explaining a function or operation of correcting image shake by replacing some frames (e.g., the 11th frame (1220) and the 12th frame (1230)) with frames generated using a learning model (e.g., the 11th-1st frame (1220a), the 12th-1st frame (1230a)) when it is determined that the ratio of overlapping areas between a plurality of frames (e.g., the 10th frame (1210), the 11th frame (1220), the 12th frame (1230)) is less than a threshold ratio, by an electronic device (101) according to one embodiment of the present disclosure. FIG. 12 is an exemplary drawing for explaining the function or operation illustrated in FIG. 11 from a user interface perspective.
[0070] Referring to FIG. 11, an electronic device (101) according to an embodiment of the present disclosure may determine, in operation 1110, whether a ratio of an overlapping area between a plurality of frames (e.g., a 10th frame (1210), an 11th frame (1220), a 12th frame (1230), and a 13th frame (1240)) exceeds a predetermined threshold ratio. The electronic device (101) according to an embodiment of the present disclosure may determine a ratio of an overlapping area between the 10th frame (1210) and the 11th frame (1220), which is a frame following the 10th frame (1210), based on the 10th frame (1210). The electronic device (101) according to an embodiment of the present disclosure may determine a ratio of an overlapping area between the 10th frame (1210) and the 12th frame (1230), which is a frame following the 11th frame (1220), based on the 10th frame (1210). An electronic device (101) according to an embodiment of the present disclosure may determine the ratio of an overlapping area between a 10th frame (1210) and a 13th frame (1240), which is a frame following a 12th frame (1230), based on the 10th frame (1210). In FIG. 12, an example is shown in which the ratio of the overlapping area between the 11th frame (1220) and the 12th frame (1230) is determined to be less than a threshold ratio. Although not illustrated in FIG. 11, an electronic device (101) according to an embodiment of the present disclosure may perform the "extended area outpainting function or operation for shaken frame correction" described in the present disclosure when the ratio exceeds the threshold ratio as determined by operation 1110, and may perform operation 1120 when the ratio is determined to be less than the threshold ratio.
[0071] An electronic device (101) according to an embodiment of the present disclosure may, in operation 1120, generate at least one new frame using a learning model based on the ratio of overlapping areas being less than a predetermined threshold ratio. An electronic device (101) according to an embodiment of the present disclosure may transmit frames before and after at least one frame (e.g., the 10th frame (1220) and the 13th frame (1230)) of which the ratio of overlapping areas is less than the threshold ratio to an external electronic device (e.g., an artificial intelligence server). An electronic device (101) according to an embodiment of the present disclosure may transmit a prompt (e.g., "Generate a motion image between the 10th frame and the 13th frame") expressed to generate at least one frame between the 10th frame (1220) and the 13th frame (1230) together with previous / next frames (e.g., the 10th frame (1220) and the 13th frame (1230)) to an external electronic device. Alternatively, the previous / next frames (e.g., the 10th frame (1220) and the 13th frame (1230)) and the prompt may be input as input data of a learning model stored in the electronic device (101). An electronic device (101) according to an embodiment of the present disclosure may obtain at least one frame (e.g., the 11th-1st frame (1220a), the 12th-1st frame (1230a)) generated by the external electronic device from the external electronic device. Alternatively, the electronic device (101) according to one embodiment of the present disclosure may obtain at least one frame (e.g., the 11-1 frame (1220a), the 12-1 frame (1230a)) as output data of the learning model. The at least one frame (e.g., the 11-1 frame (1220a), the 12-1 frame (1230a)) according to one embodiment of the present disclosure may include at least one frame that can be naturally connected to previous / next frames (e.g., the 10th frame (1220) and the 13th frame (1230)).For example, as illustrated in FIG. 12, if the preceding / following frames (e.g., the 10th frame (1220) and the 13th frame (1230)) include a motion of a person lowering an arm, at least one generated frame (e.g., the 11th-1st frame (1220a), the 12th-1st frame (1230a)) may include a frame including a designated motion to which the motion of a person lowering an arm can be naturally connected.
[0072] An electronic device (101) according to an embodiment of the present disclosure may, in operation 1130, replace some of the frames (e.g., the 11th frame (1220) and the 12th frame (1230)) among a plurality of frames (e.g., the 10th frame (1210), the 11th frame (1220), the 12th frame (1230), and the 13th frame (1240)) with at least one frame (e.g., the 11-1st frame (1220a), the 12-1st frame (1230a)) generated using a learning model. In operation 1140, an electronic device (101) according to an embodiment of the present disclosure may output (e.g., display through a display) the replaced at least one frame (e.g., the 11-1st frame (1220a), the 12-1st frame (1230a)) through the electronic device (101). Accordingly, even for images with excessive shaking, an image with shaking errors removed can be output through the electronic device (101).
[0073] According to one embodiment of the present disclosure, if it is determined that an area of an object identified as a main object in a video (e.g., identified as a main object by a user or through image analysis on a device) displayed within a frame image is below a threshold ratio (e.g., a person, which is a main object, is displayed only in a face area that is smaller than other comparison frames), a frame section that is below the threshold ratio may be determined, and a motion image to replace the determined section may be generated. At this time, the reference frame may include frames before and after the section in which it is determined that a shaking error has occurred. Alternatively, if it is determined that a main area of an object identified as a main object in the video (e.g., a face area of a person or a dog, a display area of a mobile device, a keyboard area of a piano, an area where flowers are in bloom on a tree, etc.) is not displayed within a frame or is an area that is not displayed within a frame or is an area that is not displayed within a frame or is an area that is not displayed within a frame or is an area that is above a threshold ratio, a frame section in which the area that is not displayed within a frame or is an area that is above a threshold ratio may be determined, and a motion image to replace the determined section may be generated. At this time, the reference frame may include frames before and after the section in which it is determined that a shaking error has occurred.
[0074] FIG. 13 is an exemplary drawing for explaining a function or operation of an electronic device (101) according to an embodiment of the present disclosure, which moves the display position of an object of interest (1430) included in a first image (1410) (e.g., at least one frame) into a region of interest (1420), and generates and displays another part (e.g., 1440) of the first image (1410) based on the movement using a learning model. FIGS. 14A to 14F are exemplary drawings for explaining the function or operation illustrated in FIG. 13 from a user interface perspective.
[0075] Referring to FIG. 13, an electronic device (101) according to an embodiment of the present disclosure may, in operation 1310, obtain a first user input for setting a region of interest (1420) for moving a display position of an object (e.g., an object of interest (1430)) included in at least one frame. In operation 1320, the electronic device (101) according to an embodiment of the present disclosure may display the set region of interest through a display based on the acquisition of the first user input. The electronic device (101) according to an embodiment of the present disclosure may display a user interface (1410a) for setting a region of interest (1420) within a first image (1410), as illustrated in FIG. 14A. The electronic device (101) according to an embodiment of the present disclosure may obtain a user input for the user interface (1410a). An electronic device (101) according to one embodiment of the present disclosure can set a region of interest (1420) to have a specified shape, as illustrated in FIG. 14B. An electronic device (101) according to one embodiment of the present disclosure can determine the size of the region of interest (1420) based on a user input or a pre-specified size.
[0076] An electronic device (101) according to an embodiment of the present disclosure may obtain a second user input for selecting a first object (e.g., an object of interest (1430)) in operation 1330. As illustrated in FIG. 14C , the electronic device (101) according to an embodiment of the present disclosure may select (e.g., determine) the object of interest (1430) based on a user input (e.g., a touch, a long touch gesture, a circle gesture, etc.) for a designated object. The electronic device (101) according to an embodiment of the present disclosure may extract the shape of the selected object of interest (1430) from an image (e.g., extract an area of the object by separating the background from the original image). The electronic device (101) according to an embodiment of the present disclosure may input the selected object of interest (1430) as input data of a learning model (510). To this end, the electronic device (101) according to one embodiment of the present disclosure may transmit information about the selected object of interest (1430) to an external electronic device (e.g., an artificial intelligence server). As illustrated in FIG. 14d, the electronic device (101) according to one embodiment of the present disclosure may obtain information about objects having various shapes (e.g., a first-first object of interest (1430a), a first-second object of interest (1430b), a first-third object of interest (1430c)) generated by the learning model (510) from the external electronic device. Alternatively, the electronic device (101) according to one embodiment of the present disclosure may obtain information about objects having various shapes (e.g., a first-first object of interest (1430a), a first-second object of interest (1430b), a first-third object of interest (1430c)) as output data of the learning model (510) stored in the electronic device (101).An electronic device (101) according to one embodiment of the present disclosure can accurately identify a display location of an object of interest (1430) in a plurality of frames included in a first image (1430) by using objects having various shapes (e.g., a first-first object of interest (1430a), a first-second object of interest (1430b), a first-third object of interest (1430c)) generated.
[0077] An electronic device (101) according to an embodiment of the present disclosure may, in operation 1340, identify motion vector information for a first object (e.g., an object of interest (1430)) selected according to a second user input. For example, if the first object (e.g., an object of interest (1430)) is an object moving to the right, the electronic device (101) according to an embodiment of the present disclosure may identify the direction of the motion vector as being in the right direction.
[0078] According to an embodiment of the present disclosure, the electronic device (101) may, in operation 1350, use the identified motion vector information to move the display position of the first object (e.g., the object of interest (1430)) so that the first object is located within the region of interest (1420). If the region of interest (1420) is located at the center of the screen and the object of interest (1430) is an object moving from the right side to the right (e.g., the motion vector points to the right), the electronic device (101) may move the position of the object of interest (1430) to the left and display the direction of the object so that it points to the right, which is the direction in which the motion vector points. Alternatively, the electronic device (101) may display the first object (e.g., the object of interest (1430)) within the region of interest in response to a user input for positioning the first object (e.g., the object of interest (1430)) within the region of interest. In this case, action 1350 may be omitted.
[0079] An electronic device (101) according to an embodiment of the present disclosure may output (e.g., display) a first object (e.g., object of interest (1430)) whose position has been moved and a result of performing out-painting (e.g., an image including a background image generated using a learning model) through the electronic device (101) in operation 1360. The electronic device (101) according to an embodiment of the present disclosure may perform out-painting on a right area (1440) of the image when a motion vector of the object is directed to the right and the object has been moved to the left. The out-painting according to an embodiment of the present disclosure may also be performed by the learning model (510). The electronic device (101) according to an embodiment of the present disclosure may correct at least one frame included in the first image so that the display position of the first object (e.g., object of interest (1430)) whose position has been moved is reflected in at least one frame. An electronic device (101) according to one embodiment of the present disclosure can generate corrected frames based on a frame correction function or operation according to one embodiment of the present disclosure, and generate an edited video in which the generated corrected frames are reflected.
[0080] FIG. 15 is an exemplary drawing for explaining a function or operation of an electronic device (101) according to one embodiment of the present disclosure generating a reference image (e.g., a second spatial map (1720)) for performing out-painting and performing out-painting using the reference image generated for a plurality of frames (e.g., a 14th frame (1610), a 15th frame (1620), a 16th frame (1630), a 17th frame (1640), an 18th frame (1650)).
[0081] Referring to FIG. 15, an electronic device (101) according to an embodiment of the present disclosure may generate a first spatial map (1710) by overlapping a plurality of frames (e.g., a 14th frame (1610), a 15th frame (1620), a 16th frame (1630), a 17th frame (1640), and / or an 18th frame (1650)) in operation 1510. FIGS. 16A to 16E are exemplary drawings for explaining a plurality of frames (e.g., a 14th frame (1610), a 15th frame (1620), a 16th frame (1630), a 17th frame (1640), and / or an 18th frame (1650)) required for the electronic device (101) according to an embodiment of the present disclosure to generate the first spatial map (1710). FIG. 17A is an exemplary diagram for explaining a first spatial map (1710) according to one embodiment of the present disclosure. As illustrated in FIG. 17A, an electronic device (101) according to one embodiment of the present disclosure may generate a first spatial map (1710) that includes all non-overlapping areas based on overlapping areas of a plurality of neighboring frames (e.g., a 14th frame (1610), a 15th frame (1620), a 16th frame (1630), a 17th frame (1640), and / or an 18th frame (1650)). The function or operation of generating the first spatial map (1710) according to one embodiment of the present disclosure may be performed by the electronic device (101) or an external electronic device (e.g., an artificial intelligence server).
[0082] An electronic device (101) according to an embodiment of the present disclosure may perform out-painting on the generated first spatial map (1710) in operation 1520 to generate a second spatial map (1720). FIG. 17B is an exemplary diagram for explaining the second spatial map (1720) according to an embodiment of the present disclosure. An electronic device (101) according to an embodiment of the present disclosure may perform out-painting on a spatial region (e.g., the first spatial region (1710a), the second spatial region (1710b), and / or the third spatial region (1730c)) using a learning model. An electronic device (101) according to an embodiment of the present disclosure may generate a background image excluding an identified main object (e.g., a moving person) as the second spatial map (1720). Alternatively, the electronic device (101) according to one embodiment of the present disclosure may generate an image including an identified key object (e.g., a moving person) as a second spatial map (1720). The function or operation of generating the second spatial map (1720) according to one embodiment of the present disclosure may be performed by the electronic device (101) or an external electronic device (e.g., an artificial intelligence server).
[0083] An electronic device (101) according to an embodiment of the present disclosure may perform out-painting on a plurality of frames (e.g., a 14th frame (1610), a 15th frame (1620), a 16th frame (1630), a 17th frame (1640), and / or an 18th frame (1650)) using the generated second spatial map (1720) in operation 1530. FIGS. 16F to 16J are exemplary drawings for explaining an area in which out-painting should be performed when an electronic device (101) according to an embodiment of the present disclosure moves and displays a display position of at least one object included in an image. An electronic device (101) according to an embodiment of the present disclosure may identify an area requiring out-painting (e.g., a first out-painting area (1610a), a second out-painting area (1620a), a third out-painting area (1630a), a fourth out-painting area (1640a), and a fifth out-painting area (1650a)) in each frame to correct at least one main object (e.g., a moving person) to be located in a central area of an image (e.g., a region of interest (1420)). An electronic device (101) according to an embodiment of the present disclosure may generate a prompt to perform out-painting for an area requiring out-painting (e.g., a first out-painting area (1610a), a second out-painting area (1620a), a third out-painting area (1630a), a fourth out-painting area (1640a), and a fifth out-painting area (1650a)). An electronic device (101) according to an embodiment of the present disclosure may input the generated prompt and / or the second spatial map (1720) as input data of a learning model. An electronic device (101) according to an embodiment of the present disclosure may obtain an out-painting performance result based on the second spatial map (1720) as an output result of the learning model. An electronic device (101) according to an embodiment of the present disclosure may also obtain an out-painting performance result from an external electronic device (e.g., an artificial intelligence server).By following these functions or actions, consistent out-painting performance results can be obtained.
[0084] FIG. 18 is an exemplary diagram for explaining a generative artificial intelligence system according to one embodiment of the present disclosure.
[0085] Referring to FIG. 18, a generative artificial intelligence system according to one embodiment of the present disclosure may include a user query / response interface (1810), an AI framework (1820), and / or a generative AI model (1830).
[0086] A user query / response interface (1810) according to one embodiment of the present disclosure can receive a user's input. The user's input according to one embodiment of the present disclosure can include forms such as natural language, images, and / or videos. Furthermore, context information can also be transmitted when the received user's input is transmitted. The context information according to one embodiment of the present disclosure can include various additional information at the time of user input. For example, the context information according to one embodiment of the present disclosure can include information about an application currently being used by the user or information about the user's location. Furthermore, the user's input according to one embodiment of the present disclosure can also be in a form that combines the aforementioned natural language, images, sounds, and context information. Furthermore, the user's input can also be in a non-natural language form, such as a gesture for selecting a menu. The user query / response interface (1810) according to one embodiment of the present disclosure can output a result of a generative artificial intelligence system. The output according to one embodiment of the present disclosure can be output in a natural language form or in a specific content form. The output according to one embodiment of the present disclosure may be provided in the form of an action requested by the user.
[0087] An AI framework (1820) according to one embodiment of the present disclosure can receive user input and control each component necessary to perform the user's intent based on the user's query. The AI framework (1820) according to one embodiment of the present disclosure can include a prompt design component (1822), an API / plugin management component (1824), and / or an output adjustment component (1826).
[0088] User input received in a user query / response interface (1810) according to one embodiment of the present disclosure may be transmitted to a prompt design component (1822). The prompt design component (1822) according to one embodiment of the present disclosure may generate a prompt suitable for inputting the user input into an LLM or LMM. The prompt design component (1822) according to one embodiment of the present disclosure may include an AI component using a machine learning algorithm or a neural network. The prompt design component (1822) according to one embodiment of the present disclosure may access a knowledge repository including user preference data, a prompt library, and prompt examples based on the user input to generate a prompt, and transmit the generated prompt to the LLM or LMM.
[0089] The API / plugin management component (1824) according to one embodiment of the present disclosure may communicate with an external device when a request for additional information is made when transmitting user input as input of the generative model to the generative AI model (1830). The API / plugin management component (1824) according to one embodiment of the present disclosure establishes a channel for communicating with the outside of the AI interface through an API, and may access various data sources through the established channel. The API / plugin management component (1824) according to one embodiment of the present disclosure may request information about the specified action from the application or service through the API when a specified action must be performed in the application or service. The API / plugin management component (1824) according to one embodiment of the present disclosure may also use information acquired from an external source together with user input to generate a prompt in the prompt design component (1822). Alternatively, the API / plugin management component (1824) according to one embodiment of the present disclosure may transmit information obtained from an external source to the generative AI model (1830) as input to the generative AI model (1830).
[0090] An output tuning component (1826) according to one embodiment of the present disclosure can tune the output from the generative AI model (1830). For example, an output tuning component (1826) according to one embodiment of the present disclosure can verify whether the content generated through the LLM and / or LMM is relevant to the user's input, does not contain biased content, or does not contain harmful content. An output tuning component (1826) according to one embodiment of the present disclosure can determine how closely the output result matches the result desired by the user. An output tuning component (1826) according to one embodiment of the present disclosure can also perform additional processes if additional processes are required (e.g., if the matching ratio exceeds a specified error range). An output tuning component (1826) according to one embodiment of the present disclosure can configure hints related to the results and provide them to the user.
[0091] A generative AI model (1830) according to one embodiment of the present disclosure may include an artificial intelligence neural network capable of generating new types of data based on user input information. A generative AI model (1830) according to one embodiment of the present disclosure may include a model for generating images and / or a model for generating language. According to one embodiment of the present disclosure, a model for generating images typically includes a generative adversarial network (GAN) and a variational auto encoder (VAE), and may include a VAE and a Diffusion-based generative model using a Transformer structure. According to one embodiment of the present disclosure, a model for generating language may include a model trained to output the most statistically appropriate output value based on an input value, and may include models such as CHAT-GPT 3 and CHAT-GPT 4, for example. In addition, a large multimodal model (LMM) capable of recognizing various types of data input, such as text, images, and voice, and generating new data corresponding thereto may also be included.
[0092] An electronic device according to one embodiment of the present disclosure includes a display, at least one processor, and a memory, wherein the memory may be configured to store instructions that, when executed by the at least one processor, cause the electronic device to identify a plurality of frames of at least one selected video based on a user's selection input for at least one video, determine an overlapping area between a first frame and a second frame among the identified plurality of frames, generate a first prompt expressed to generate a part of the first frame that is not included in the second frame based on the determined overlapping area, generate a part of the first frame using the generated first prompt and a learning model, generate a third frame by synthesizing the part of the generated first frame with the second frame, and display the generated third frame through the display.
[0093] A method according to one embodiment of the present disclosure may include: identifying a plurality of frames of at least one selected video based on a user's selection input for at least one video; determining an overlapping area between a first frame and a second frame among the identified plurality of frames; generating a first prompt expressed to generate a part of the first frame that is not included in the second frame based on the determined overlapping area; generating a part of the first frame using the generated first prompt and a learning model; generating a third frame by synthesizing a part of the generated first frame with the second frame; and displaying the generated third frame through the display.
[0094] A computer-readable non-transitory recording medium according to one embodiment of the present disclosure is configured to store a plurality of instructions, wherein the plurality of instructions, when executed by at least one processor of an electronic device, may include instructions that cause the electronic device to: identify a plurality of frames of at least one selected video based on a user's selection input for at least one video; determine an overlapping area between a first frame and a second frame among the identified plurality of frames; generate a first prompt expressed to generate a part of the first frame that is not included in the second frame based on the determined overlapping area; generate a part of the first frame using the generated first prompt and a learning model; generate a third frame by synthesizing the part of the generated first frame with the second frame; and display the generated third frame through the display.
[0095] Electronic devices according to various embodiments disclosed in the present disclosure may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to embodiments of the present disclosure are not limited to the aforementioned devices.
[0096] The various embodiments of the present disclosure and the terminology used therein are not intended to limit the technical features described in the present disclosure to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In the present disclosure, each of the phrases "A or B," "at least one of A and B," "at least one of A or B," "A, B, or C," "at least one of A, B, and C," and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among the phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0097] The term "module" used in various embodiments of the present disclosure may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0098] Various embodiments of the present disclosure may be implemented as software (e.g., a program (2540)) including one or more instructions stored in a storage medium (e.g., an internal memory (2536) or an external memory (2538)) readable by a machine (e.g., an electronic device (2501)). For example, a processor of the machine (e.g., the electronic device (2501)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0099] According to one embodiment, the method according to various embodiments disclosed in the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0100] According to one embodiment of the present disclosure, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separately arranged in other components. According to one embodiment of the present disclosure, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment of the present disclosure, operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, display, At least one processor, and A memory configured to store instructions, said instructions, when executed by said at least one processor, causing said electronic device to: Receive user input for a graphic object corresponding to a video, Based on at least a portion of the user input, identifying a first frame and a second frame from the video, Identify at least one object included in the first frame and not included in the second frame, To create a modified second frame, a first prompt is created using the second frame and the at least one object, Based on the first prompt and the second frame, the modified second frame including at least one object generated is generated using a learning model, Generating a modified video including the first frame and the modified second frame, and An electronic device characterized in that it is set to store instructions set to display the modified video through the display.
2. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device, characterized in that it further comprises an instruction set to crop the modified second frame based on a specified loss rate.
3. In paragraph 1 or 2, The above instructions, when executed by the at least one processor, cause the electronic device to: Identifying multiple frames from the above video, Generate a second prompt based on an object contained in an area other than an overlapping area of at least two of the identified plurality of frames, and An electronic device further comprising instructions configured to display the modified video, which includes a third frame generated based on the second prompt, on the display.
4. In any one of paragraphs 1 to 3, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device characterized in that it includes an instruction set to store a fourth frame generated based on the learning model by replacing the second frame based on a ratio of the overlapping area being less than a predetermined threshold ratio.
5. In any one of paragraphs 1 to 4, The above instructions, when executed by the at least one processor, cause the electronic device to: Changing the position of at least one first object within the first frame based on a user input for moving at least one first object included in the first frame; Using the above learning model, a background image is generated according to the position movement of at least one first object, An electronic device characterized in that it further includes instructions set to display the generated background image through the display.
6. In any one of paragraphs 1 to 5, The above instructions, when executed by the at least one processor, cause the electronic device to: Generating a first spatial map using non-overlapping areas of the first frame and the second frame and at least one second object included in the first frame and the second frame, An electronic device characterized in that it further includes an instruction set to generate the third frame using the generated first spatial map and the learning model.
7. In any one of paragraphs 1 to 6, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device, characterized in that it further includes instructions configured to perform outpainting on the first spatial map to generate a second spatial map.
8. In a computer-readable non-transitory recording medium, the non-transitory recording medium is configured to store a plurality of instructions, and the plurality of instructions, when executed by at least one processor of an electronic device, cause the electronic device to: Receive user input for a graphic object corresponding to a video, Based on at least a portion of the user input, identifying a first frame and a second frame from the video, Identify at least one object included in the first frame and not included in the second frame, To create a modified second frame, a first prompt is created using the second frame and the at least one object, Based on the first prompt and the second frame, the modified second frame including at least one object generated is generated using a learning model, Generating a modified video including the first frame and the modified second frame, and A computer-readable non-transitory recording medium comprising instructions set to display the modified video through the display.
9. In paragraph 8, The above instructions, when executed by the at least one processor, cause the electronic device to: A computer-readable non-transitory recording medium, characterized in that it further comprises instructions set to crop the modified second frame based on a specified loss rate.
10. In paragraph 8 or 9, The above instructions, when executed by the at least one processor, cause the electronic device to: Identifying multiple frames from the above video, Generate a second prompt based on an object contained in an area other than an overlapping area of at least two of the identified plurality of frames, A computer-readable non-transitory recording medium, characterized in that it further includes instructions set to display a third frame generated based on the second prompt on the display.
11. In any one of paragraphs 8 to 10, The above instructions, when executed by the at least one processor, cause the electronic device to: A computer-readable non-transitory recording medium, characterized in that it includes an instruction set to store a fourth frame generated based on the learning model by replacing the second frame based on a ratio of the overlapping area exceeding a predetermined threshold ratio.
12. In any one of paragraphs 8 to 11, The above instructions, when executed by the at least one processor, cause the electronic device to: Changing the position of at least one first object within the first frame based on a user input for moving at least one first object included in the first frame; Using the above learning model, a background image is generated according to the position movement of at least one first object, and A computer-readable non-transitory recording medium, characterized in that it further includes instructions set to display the generated background image through the display.
13. In any one of paragraphs 8 to 12, The above instructions, when executed by the at least one processor, cause the electronic device to: Generating a first spatial map using non-overlapping areas of the first frame and the second frame and at least one second object included in the first frame and the second frame, A computer-readable non-transitory recording medium, characterized in that it further includes instructions set to generate the third frame using the generated first spatial map and the learning model.
14. In any one of paragraphs 8 to 13, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device, characterized in that it further includes instructions configured to perform outpainting on the first spatial map to generate a second spatial map.
15. In any one of paragraphs 8 to 14, The above instructions, when executed by the at least one processor, cause the electronic device to: A computer-readable non-transitory recording medium, characterized in that it further includes instructions set to determine a first frame section based on the determination that an area in which at least one object is displayed within a frame image is less than a threshold ratio, and to generate a motion image to replace the determined first frame section.