Electronic apparatus and method for processing video

The integration of a generative model in electronic devices allows for dynamic video playback by generating acoustic content based on visual object processing, addressing the limitations of existing systems and enhancing user interaction.

WO2025263770A1PCT designated stage Publication Date: 2025-12-26SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/004602
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-11
Filing Date
2025-04-04
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing video processing systems lack the ability to seamlessly integrate visual and acoustic content modifications based on user input, limiting the dynamic and interactive experience in video playback.

Method used

An electronic device equipped with a generative model that processes visual objects from video frames to generate corresponding acoustic content, allowing for interactive video playback by modifying characteristics and adding sound content based on user input.

Benefits of technology

Enables dynamic and user-driven integration of visual and acoustic elements in video playback, enhancing the interactive experience by generating acoustic content that matches user intent and modifying video characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004602_26122025_PF_FP_ABST
    Figure KR2025004602_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an electronic apparatus and method for processing a video. An operation method of the electronic apparatus according to an embodiment may comprise an operation of acquiring a visual object included in one frame from among a plurality of frames of an original video. The operation method may comprise an operation of generating sound content corresponding to the visual object, on the basis of at least one characteristic of the visual object. The operation method may comprise an operation of outputting the sound content by including the sound content in frames including the visual object from among the plurality of frames. Various other embodiments may be possible.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices and methods for processing video

[0001] Embodiments of the present invention relate to electronic devices and methods for processing video.

[0002] A generative model is an artificial intelligence model that can generate new data based on input data. Generative models can also generate new content (e.g., acoustic content, images) that matches the user's intent based on input data (e.g., prompts).

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.

[0004] An electronic device according to one embodiment may include a display, at least one processor including processing circuitry, and a memory storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a visual object included in one of a plurality of frames of an original video.

[0005] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate acoustic content corresponding to the first object based on at least one property of the first object.

[0006] The above instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to play the sound content by including the first object in frames among the plurality of frames.

[0007] A method of operating an electronic device according to one embodiment may include obtaining a visual object included in one frame of a plurality of frames of an original video.

[0008] The above method of operation may include an operation of generating acoustic content corresponding to the visual object based on at least one property of the visual object.

[0009] The above operating method may include an operation of outputting the audio content by including it in frames among the plurality of frames that include the visual object.

[0010] According to one embodiment, a computer-readable recording medium storing one or more computer programs may include instructions for causing a processor to perform the method.

[0011] An electronic device according to one embodiment may include a display, at least one processor including processing circuitry, and a memory storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display an original video on a first portion of the display and to display an image on a second portion of the display.

[0012] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to modify at least one characteristic of a first object included in one of a plurality of frames of the original video based on a second object included in the image.

[0013] The above instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate acoustic content corresponding to the changed characteristic.

[0014] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to play the changed object and the sound content by including the first object in the frames among the plurality of frames.

[0015] A method of operating an electronic device according to one embodiment may include displaying an original video on a first portion of a display and displaying an image on a second portion of the display.

[0016] The method may include changing at least one characteristic of a first object included in one of a plurality of frames of an original video based on a second object included in the image.

[0017] The above method of operation may include an operation of generating acoustic content corresponding to the changed characteristic.

[0018] The above operating method may include an operation of including the changed object and the sound content in frames including the first object among the plurality of frames and outputting them.

[0019] According to one embodiment, a computer-readable recording medium storing one or more computer programs may include instructions for causing a processor to perform the method.

[0020] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0021] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.

[0022] FIG. 2 is a diagram for explaining a generative artificial intelligence system according to one embodiment.

[0023] FIG. 3 is a diagram illustrating a video processing system according to one embodiment.

[0024] FIG. 4A and FIG. 4B are diagrams for explaining a method for video processing according to one embodiment.

[0025] FIGS. 5A to 5C are diagrams illustrating an interface for video processing according to one embodiment.

[0026] FIG. 6 is a diagram illustrating an interface for providing a user with processed video according to one embodiment.

[0027] FIGS. 7A to 7C are diagrams for explaining a segmentation operation of an electronic device according to one embodiment.

[0028] FIG. 8 is a diagram illustrating an operation of adding an object to a video according to one embodiment.

[0029] FIG. 9 is a diagram illustrating a video processing guide interface provided to a user according to one embodiment.

[0030] FIGS. 10A to 10C are diagrams illustrating an interface for video processing according to one embodiment.

[0031] FIG. 11 is a diagram illustrating an interface for video processing of an electronic device according to one embodiment.

[0032] Fig. 12 shows a flowchart of an operating method of an electronic device according to one embodiment.

[0033] Fig. 13 shows a flowchart of an operating method of an electronic device according to one embodiment.

[0034] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.

[0035]

[0036] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100), according to one embodiment.

[0037] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0038] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may be implemented as a circuit (e.g., a processing circuit) such as a system on chip (SoC) or an integrated circuit (IC). The processor (120) may include one or more processors. For example, the processor (120) may include a combination of one or more processors, such as a CPU, a GPU, an MPU, an AP, and a CP. Instructions stored in the memory (130) may be executed by one processor to cause the electronic device (101) to perform and / or control the operations of the electronic device (101) to be described with reference to FIGS. 2 to 13. Instructions stored in the memory (130) may be executed by multiple processors to cause the electronic device (101) to perform and / or control the operations of the electronic device (101) to be described with reference to FIGS. 2 to 13.

[0039] According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0040] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0041] The memory (130) can store various data used by at least one component (e.g., the processor (120) or the sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., the program (140)) and input data or output data for commands related thereto. The memory (130) can include a volatile memory (132) or a non-volatile memory (134). According to one embodiment, the memory (130) can include one or more memories. The instructions stored in the memory (130) can be stored in one memory. The instructions stored in the memory (130) can be divided and stored in multiple memories. The instructions stored in the memory (130) can be executed by the processor (120) to cause the electronic device (101) to perform and / or control the operations of the electronic device (101) to be described with reference to FIGS. 2 to 13.

[0042] According to one embodiment, the instructions stored in the memory (130) may cause the electronic device (101) to perform one or more operations when individually or collectively executed by at least one processor (e.g., the main processor (121) and / or the auxiliary processor (123)). For example, the instructions stored in the memory (130) may be executed by one processor (e.g., the main processor (121) or an auxiliary processor (123) such as a communication processor) or by a plurality of processors operating cooperatively (e.g., the main processor (121) and the auxiliary processor (123)).

[0043] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0044] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0045] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0046] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0047] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0048] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0049] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0050] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0051] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0052] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0053] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0054] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0055] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0056] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0057] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0058] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0059] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0060] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0061]

[0062] An electronic device according to an embodiment disclosed in this document may take various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance. The electronic device according to an embodiment of this document is not limited to the aforementioned devices.

[0063] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0064] The term "module" used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0065] One embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0066] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0067] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0068]

[0069] FIG. 2 is a diagram for explaining a generative artificial intelligence system according to one embodiment.

[0070] Referring to FIG. 2, according to one embodiment, the generative artificial intelligence system (200) may be a program (e.g., a software module) implemented on an electronic device (e.g., an electronic device (101) of FIG. 1) and / or a server (e.g., a server (108) of FIG. 1).

[0071] According to one embodiment, a user query / response interface (210) may receive user input. The user input may be an input of a type (or modality) such as natural language, image, audio, and / or video. Additionally, context information may also be transmitted when the user input is transmitted. The context information may include various side information related to the time when the user input is input to the artificial intelligence system (200). For example, the side information may include information such as application information currently being used by the user or location information of the user. Additionally, the user input may be a mixed type of input of the above-described natural language, image, audio, video, and / or context information. Additionally, the user input may include non-natural language input, such as selecting a menu.

[0072] According to one embodiment, a user query / response interface (210) may provide output from a generative artificial intelligence system to a user. The output may include a natural language-based response and / or specific content. The output may also include an action requested by the user.

[0073] In one embodiment, an AI framework (220) may receive user input. Based on the user input (e.g., a user's query), the AI ​​framework (220) may coordinate and / or control one or more components necessary to perform an action corresponding to the user's intent.

[0074] According to one embodiment, user input received from the user query / response interface (210) may be transmitted to a prompt design component (221). The prompt design component (221) may be used to generate a prompt suitable as input to a generative model (e.g., a large language model (LLM) and / or a large multimodal model (LMM)) based on the user input.

[0075] In one embodiment, the prompt design component (221) may be an AI component that utilizes a machine learning algorithm or a neural network. The prompt design component (221) may generate improved prompts over time through learning. The prompt design component (221) may access a knowledge repository (230) to generate prompts based on user input. The knowledge repository (230) may include user preference data, a prompt library, and / or prompt examples. The prompt design component (221) may provide the generated prompts to a generative model (e.g., an LLM and / or an LMM).

[0076] According to one embodiment, the APIs / Plugins management component (223) can communicate with an external information source based on a request for additional information when user input is transmitted to the generative model.

[0077] In one embodiment, the APIs / Plugins management component (223) can establish a communication channel for communication with the outside of the system (200) via the API. The APIs / Plugins management component (223) can enable access to various data sources via the communication channel. The acquired information can be used to generate prompts by the prompt design component (221) along with user input, or can be used as input to the generative model (250).

[0078] According to one embodiment, the APIs / Plugins management component (223) may request the application / service component (240) via the API for a final action in response to user input, rather than an intermediate action, if the final action must be performed by the application or service.

[0079] In one embodiment, the refiner component (225) can fine-tune the output of the generative model (250). For example, the refiner component (225) can determine the relevance (e.g., a score) between the output (e.g., content) of the generative model and the user input. For example, the refiner component (225) can determine whether the output contains biased information (e.g., selective information). For example, the refiner component (225) can determine whether the output contains harmful information (e.g., violent content or profanity).

[0080] In one embodiment, the refinement component (225) may determine the degree of matching (e.g., a score) between the output of the generative model (250) and the user input (e.g., the intent of the user input). If the refinement component (225) determines that the output of the generative model (250) does not correspond to the user input, the refinement component (225) may modify the output to correspond to the user input.

[0081] In one embodiment, the refinement component (225) may provide hints to the user (e.g., hints for prompt generation) to enable the user to obtain information that matches the user's intent from the generative model (250).

[0082] According to one embodiment, a generative model (250) may refer to an artificial intelligence neural network that generates new data (e.g., text, images, audio, or video) based on user input (e.g., user utterance). The generative model (250) may include an image generation model and / or a language generation model.

[0083] In one embodiment, the image generation model may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). An example of an image generation model is a diffusion-based generative model having the structure of a VAE and a transformer.

[0084] In one embodiment, a language generation model (e.g., ChatGPT) may be a model trained to generate statistically most appropriate output based on input. The language generation model may include an LMM. The LMM can identify various types of input, such as text, images, audio (e.g., speech), and / or video, and generate new data corresponding to the input.

[0085]

[0086] FIG. 3 is a diagram illustrating a video processing system according to one embodiment.

[0087] Referring to FIG. 3, according to one embodiment, a video processing system (300) may include one or more components (310-320 and 330-375). The components (310-320 and 330-375) may be software modules implemented on an electronic device (e.g., the electronic device (101) of FIG. 1) or a server (e.g., the server (108) of FIG. 1). The components (310-320 and 330-375) are illustrated as examples for describing the video processing system (300). Accordingly, the video processing system (300) may include various variations of the components (310-320 and 330-375) as long as the operations of the video processing system (300) described in the present disclosure can be implemented. For example, two or more components may be combined, or one or more components may be added or omitted. Alternatively, the video processing system (300) may further include one or more components (e.g., components (210-250) of FIG. 2) of a generative artificial intelligence system (e.g., the generative artificial intelligence system (200) of FIG. 2).

[0088] According to one embodiment, the video processing system (300) may include a frame extraction module (310), an audio extraction module (320), an object processing module (330), an image selection module (340), an audio preprocessing module (350), a category matching module (360), a generative model (370) (e.g., the generative model (250) of FIG. 2), and a post processing module (375). The video processing system (300) may receive an original video (301), process the original video (301), and output a synthesized video (381). The original video (301) may include a video stored in the electronic device (101) and / or a video being played on the electronic device (101).

[0089] According to one embodiment, the frame extraction module (310) can extract a plurality of frames (315) from the original video (301). The plurality of frames (315) can include a plurality of still images constituting the original video (301). The plurality of frames (315) can be passed to the generative model (370). The audio extraction module (320) can extract audio (e.g., extracted audio (325)) from the original video (301). The extracted audio (325) can include audio of the original video (301). For example, the extracted audio (325) can include sound data included in the original video (301). The extracted audio (325) can be passed to the audio preprocessing module (350).

[0090] According to one embodiment, the object processing module (330) can process objects (e.g., objects such as people, animals other than people, walls, and objects) included in the original video (301). For example, the object processing module (330) can recognize objects (e.g., visual objects, first objects) included in the original video (301) and classify the recognized objects. The object processing module (330) can include an object recognition module (332) and an object classification module (334).

[0091] According to one embodiment, the object recognition module (332) can recognize an object (e.g., a visual object, a first object) included in one frame among a plurality of frames (315) extracted from the original video (301). The object recognition module (332) can recognize at least one object included in at least one multimedia content. The at least one multimedia content can include audio. The at least one multimedia content can include an audio source and / or video content different from the original video (301). The object recognition module (332) can recognize an object and output information about the recognized object. For example, the object recognition module (332) can output information such as a location (e.g., coordinates), a size, a category (e.g., a category such as a person, an animal other than a person, a wall, an object), a mask image, and confidence of the recognized object. The location of the object can be expressed in the form of a location box. For example, the location of an object can be expressed in a format such as the coordinates of a location box including the coordinates of the initial point of the location box and the coordinates of the end point of the location box (e.g., x_init: 30, y_init: 40 / x_end: 145, y_end: 187). The mask image can include information that can identify the location of the object in pixel units. The object recognition module (332) can output information about the object in the form of metadata. The object recognition module (332) can recognize at least one object included in a plurality of frames (315) using an artificial intelligence (AI) model (e.g., a segmentation AI model).The segmentation AI model may be a model trained through a training data pair containing information about an image and an object included in the image. The object recognition module (332) may adjust the degree of object recognition (e.g., intensity of segmentation) from a plurality of frames (315) to adjust the objects recognized (e.g., type and number of objects). By adjusting the degree of object recognition (hereinafter referred to as the degree of object recognition), the object recognition module (332) may more accurately reflect the user's intention or improve the recognition speed. For example, when the degree of object recognition is set high, the number and / or types of objects recognized from a plurality of frames (315) may increase, resulting in a high degree of accuracy; however, the time required for recognition processing may increase. For example, if the object recognition level is set low, the time required for recognition processing may be reduced, allowing for faster processing. However, the number and / or types of objects recognized may be reduced, so the user's intention may not be accurately reflected. The strength of the segmentation may vary based on the configuration parameter values ​​used to execute the segmentation AI model and / or information about the object. The information about the object may include information about the box used to recognize the object (e.g., a location box). The information about the location box may include information such as the size of the box and the location of the box (e.g., a location expressed in absolute coordinates). The strength of the segmentation may vary depending on a threshold value regarding the size of the object. The threshold value regarding the size of the object may include a value that serves as a standard for the size of the object, and a value that prevents objects below the threshold value from being recognized (e.g., detected). For example, if the threshold value is set low, smaller objects may be detected, and the strength of the segmentation may increase.For example, if the threshold value is set high, the strength of the segmentation may be lowered because small objects are excluded from the detection target. The object recognition module (332) may determine one object among at least one object included in at least one multimedia content as an object candidate for changing an object (e.g., a visual object) included in one of the multiple frames (315) of the original video (301). For example, the object recognition module (332) may determine at least one object recognized from at least one multimedia content as an object candidate for changing a visual object. The object candidate may then be used by the generative model (370) to generate audio content to replace the audio corresponding to the visual object if the visual object and the category are determined to match by the category matching module (360). The object recognition module (332) may transmit the result of recognizing the object and / or information about the object to the object classification module (334). The object recognition module (332) can transmit the result of recognizing an object and / or information about the object to the category matching module (360). The object recognition module (332) can transmit an object candidate to the category matching module (360).

[0092] According to one embodiment, the object classification module (334) can classify an object (e.g., a visual object, a first object) included in one frame among a plurality of frames (315) extracted from the original video (301). For example, the object classification module (334) can classify an object based on information about the object (e.g., information such as location, size, category, mask image, and reliability). The object classification module (334) can adjust the type of object to be classified by adjusting the degree of classifying the object from the plurality of frames (315) (e.g., the degree of classification such as the number of classified kinds, types, and quantity). The object classification module (334) can transmit the result of classifying the object (e.g., the result of classifying the category) to the category matching module (360).

[0093] According to one embodiment, the image selection module (340) may determine a second object to modify an object (e.g., a first object) included in one of a plurality of frames (315) of an original video (301) based on an image stored in the electronic device (101). The image selection module (340) may receive a user's selection of an image and determine an object included in the selected image as the second object to modify the first object. The second object may be transmitted to the category matching module (360).

[0094] According to one embodiment, the category matching module (360) can determine whether a category of an object (e.g., a visual object, a first object) included in one frame among a plurality of frames (315) of the original video (301) matches with at least one object included in content (e.g., content such as an image or multimedia content) different from the original video (301). The at least one object included in the multimedia content may include a visual object and / or an audio object. The category matching module (360) can output, as a matching result, an object whose category matches an object (e.g., a visual object, a first object) included in one frame among a plurality of frames (315) of the original video (301) among at least one object included in content (e.g., content such as an image or multimedia content) different from the original video (301). The category matching module (360) can determine whether the categories match based on whether the correlation level (e.g., similarity) between the categories satisfies a predetermined threshold. The threshold may vary depending on the matching result. For example, if there are no objects (e.g., visual objects, first objects) included in one of the multiple frames (315) of the original video (301) according to the matching result and the category matches, or there are too few objects, the category matching module (360) can lower the threshold to flexibly determine the correlation between the categories, thereby detecting more objects that match the categories. For example, if there are too many objects (e.g., visual objects, first objects) included in one of the multiple frames (315) of the original video (301) according to the matching result and the category matches, the category matching module (360) can increase the threshold to strictly determine the correlation between the categories, thereby detecting fewer objects that match the categories.

[0095] According to one embodiment, the category matching module (360) can determine whether the categories of objects match based on specified categories. Whether the categories match can be determined based on the correlation between the categories. The types and / or number of the specified categories can be the same as the types and / or number of categories of objects that the object recognition module (332) can recognize from a plurality of frames (315). A dataset in which the categories of objects recognized by the object recognition module (332) are stored can be managed by the category matching module (360). The category matching module (360) can prevent objects (e.g., visual objects, first objects) of the original video (301) from being unrealistically edited by determining whether the categories between objects match. For example, the category matching module (360) can determine whether the categories of the first object and the second object match so that a first object in the food category cannot be replaced with a second object in a category such as a vehicle or an animal. The category matching module (360) can determine whether to use multimedia content to generate audio content by determining whether the categories of a visual object and an object candidate match. The category matching module (360) can determine whether the categories match based on the relationship between the categories of a first object and a second object even without a designated category, and it should be noted that the operation of determining whether a category matches performed by the category matching module (360) is not limited to the above example.

[0096] According to one embodiment, the audio preprocessing module (350) can preprocess audio (325) extracted from the original video (301). For example, the audio preprocessing module (350) can preprocess the extracted audio (325) by removing noise. The audio preprocessing module (350) can preprocess the extracted audio (325) and transfer it to the generative model (370). The generative model (370) can use the extracted audio (325) as information about the original video (301) (e.g., information such as frequency band and amplitude information), thereby generating sound content that is not incongruous with the original video (301).

[0097] According to one embodiment, the generative model (370) can generate audio content. The generative model (370) can generate audio content corresponding to a visual object based on at least one characteristic of a visual object included in one frame among a plurality of frames (315) of the original video (301). The generative model (370) can include the audio content in frames including the visual object among the plurality of frames of the original video (301). The generative model (370) can generate a frame in which audio corresponding to a visual object among audio of the original video (301) (e.g., extracted audio (325)) is replaced with audio content. The generative model (370) can generate audio content based on at least one multimedia content including audio. The audio content can have a correlation with at least one characteristic of the visual object that is greater than a predetermined threshold value.

[0098] According to one embodiment, the generative model (370) can modify at least one characteristic of a first object included in one of a plurality of frames (315) of an original video (301) based on a second object. The generative model (370) can generate an object (e.g., an image corresponding to the modified object) in which at least one characteristic of the first object is modified based on the second object. The generative model (370) can generate acoustic content corresponding to the modified characteristic. The generative model (370) can include the modified object and acoustic content in frames including the first object among the plurality of frames (315). For example, the generative model (370) can render such that the modified object replaces the first object and the acoustic content is output to replace audio corresponding to the first object. The video generated by the generative model (370) can be transmitted to a post-processing module (375). The video generated by the generative model (370) can be post-processed in the post-processing module (375) and ultimately output as a synthetic video (381).

[0099] According to one embodiment, the post-processing module (375) can post-process the video generated by the generative model (370). For example, the post-processing module (375) can perform post-processing such as tone-mapping and denoising on the video generated by the generative model (370). The post-processing module (375) can output a synthetic video (381). The synthetic video (381) is the final result of the video processing process and can be displayed on the electronic device (101) through the display module (160). The electronic device (101) can store the synthetic video (381) in association with the original video (301) in the electronic device (101).

[0100]

[0101] FIG. 4A and FIG. 4B are diagrams for explaining a method for video processing according to one embodiment.

[0102] Referring to FIGS. 4A and 4B , according to one embodiment, the electronic device (101) can output (e.g., display) a screen (401) to a screen (405) for video processing. The electronic device (101) can display the screen (401) to the screen (405) through a display module (e.g., the display module (160) of FIG. 1 ). The screen (401) to the screen (405) can include a user interface (UI) for a user (e.g., a user of the electronic device (101)) to process (e.g., edit) a video (e.g., the original video (301) of FIG. 3 ).

[0103] According to one embodiment, the electronic device (101) may display an original video (301) on a first portion of a display (e.g., a display module (160)). Note that the first portion does not mean a specific area of ​​the display module (160) and that the location and / or shape may vary depending on the embodiment. The electronic device (101) may receive a scene that the user wants to process (e.g., edit) in the original video (301) through a user input. For example, the user may pause the original video (301) and select a frame (411) displayed on the screen (401) as a frame corresponding to the scene that the user wants to edit. The frame (411) may be one of a plurality of frames (e.g., a plurality of frames (315) of FIG. 3). The frame (411) may include a frame corresponding to the scene that the user wants to edit. The user can select a frame to be edited from among multiple frames (315) of the original video (301) in various ways, and it should be noted that pausing the video being played and selecting the displayed frame is only one example.

[0104] According to one embodiment, the screen (401) may include an icon (410). The icon (410) may be an icon that allows a user to edit a frame (411) using an image (e.g., an image including an object). For example, the icon (410) may be an icon that allows a user to edit a frame (411) by adding an extra object to the frame (411) using an image, changing at least one object (e.g., an object (412)) within the frame (411), or deleting at least one object (e.g., an object (412)) within the frame (411). The electronic device (101) may display the screen (402) in response to a user's input (e.g., a touch input) to the icon (410).

[0105] According to one embodiment, the screen (402) may include an interface for editing the frame (411) using an image (e.g., an image including an object). The screen (402) may include an icon (e.g., icon (413)) for allowing a user to retrieve an image (e.g., an image including an object) that can be used to edit the frame (411). The image may include at least one image. The image may include an image stored in the electronic device (101) and / or an image stored in a server. The image may include an image in sticker format. The sticker format image is an image including an object depicted on a transparent background, and may be created by a user using the electronic device (101) to separate a specific area (e.g., an object) from an original image (e.g., an image stored in a gallery). The user may place the sticker format image in a specific location of an image and / or a video (e.g., at least one frame of a video). The icon (413) may be an icon for enabling loading of a sticker-type image created using an image stored in the electronic device (101). When the electronic device (101) creates a sticker-type object image by separating an object from a stored image, if the separated object is not in a complete form (e.g., only a portion of the object exists), the electronic device (101) may create the remaining part of the object and store it as an image containing the object in a complete form.

[0106] According to one embodiment, the electronic device (101) may display an image (414) on a second portion of the display in response to a user input for the icon (413). The second portion may be different from the first portion where the original video (301) is displayed. The image (414) may include a sticker-type image (e.g., a sticker image (415)) generated using an image stored in the electronic device (101). For example, the image (414) may include a sticker-type image (e.g., a sticker image (415) for a yellow umbrella, a sticker image such as a sticker image for a piano) generated using an image stored in the electronic device (101). The electronic device (101) may receive a selection input from the user for an image to be used for editing a frame (411) among the images (414). The electronic device (101) may receive a user's selection for an image (414) and determine an object included in the selected image as a second object for changing at least one characteristic of a first object (e.g., object (412)) included in a frame (411). For example, the electronic device (101) may determine a yellow umbrella, which is an object included in a sticker image (415) corresponding to the user's selection input, as the second object. The electronic device (101) may display at least one object (e.g., object (412)) included in the frame (411). The electronic device (101) may display the object (412) on the display module (160) together with a user interface component (UI) indicating information about the object (412). For example, the electronic device (101) may display UI elements such as a keyword representing the object (412) (e.g., a keyword expressed in the form of a hashtag (e.g., #umbrella)) and a dotted line (e.g., a yellow dotted line) for indicating an object area, on the display module (160) together with the object (412).The electronic device (101) can obtain a first object included in one frame (e.g., frame (411)) among a plurality of frames (315) of the original video (301) by using the object processing module (330). The object (412) may be output by processing the frame (411) by the object processing module (e.g., object processing module (330) of FIG. 3). For example, the object (412) may be recognized by the object processing module (330) by using a segmentation artificial intelligence model. By outputting at least one object (e.g., object (412)) included in the frame (411) by using the object processing module (330), the electronic device (101) can reduce the inconvenience of a user having to manually select an object to be edited, and can enable more accurate selection of an editing target.

[0107] According to one embodiment, a user may select a first object to be processed (e.g., edited) from among at least one object included in a frame (411). The first object may include an object to be edited (e.g., object (412)). The electronic device (101) may determine the first object based on a user's selection input for an image (414), even without a user's selection input for at least one object included in the frame (411). For example, the electronic device (101) may analyze characteristics of an object (e.g., a yellow umbrella) included in an image (e.g., a sticker image (415)) selected by the user as an image to be used for editing the frame (411), and determine the object (412) from among the objects included in the frame (411) as the first object based on the analyzed characteristics. The electronic device (101) may determine the first object based on characteristics of a second object. For example, the electronic device (101) may analyze the characteristics of the second object (e.g., a yellow umbrella) included in the selected sticker image (415), such as category, size, shape, color, pattern, style, texture, use, associated sound (e.g., a sound associated with the object, such as a crying sound), and interaction with other objects, and determine an object (412) having the same or similar characteristics as the second object among at least one object included in the frame (411) as the first object. The electronic device (101) may determine an object of which at least one characteristic can be changed using the second object as the first object based on the characteristics of the second object. The electronic device (101) may analyze the characteristics of the second object and the characteristics of at least one object included in the frame (411) using an object classification module (e.g., the object classification module (334) of FIG. 3). If the electronic device (101) determines that the categories of the second object and the first object do not match, it can provide a notification to the user.For example, if the electronic device (101) determines that it is impossible to change at least one characteristic of the first object using the second object because the categories of the second object and the first object are completely different, the electronic device (101) may provide the user with a notification including the corresponding content (e.g., content expressed in text such as “You cannot edit a video using the selected object”, “There is no image that can edit the selected object”). In response to the user’s selection of the object (412), the electronic device (101) may search for an object that can change at least one characteristic of the object (412) from the stored images and suggest the object to the user. For example, the electronic device (101) may analyze the characteristics of at least one object included in the frame (411), and display at least one object among the stored images that has the same or similar characteristics as the object (412) in the second portion and suggest the object to the user.

[0108] According to one embodiment, the electronic device (101) may display an interface (e.g., interface (416), interface (426)) in response to a user's selection input for a sticker image (415) and / or a user's selection input for an object (412). The interface may include an interface for providing an editing guide to the user. The editing guide (e.g., an editing guide implemented as an AI assistant function) may be for suggesting things that the user may consider when editing an object (412) included in one of a plurality of frames of the original video (301). The electronic device (101) may provide the editing guide to the user based on the characteristics of the second object through the interface (e.g., interface (416), interface (426)). The electronic device (101) may generate a guide to be provided to the user through the interface (e.g., interface (416), interface (426)) by using an artificial intelligence model (e.g., generative model (370) of FIG. 3). The artificial intelligence model may include an artificial intelligence model trained to generate an editing guide. The electronic device (101) may analyze the difference between the characteristics of the first object and the characteristics of the second object using the artificial intelligence model, and analyze the characteristics suitable for reflection in the frame (411) to generate a guide for editing. For example, if the object (412) determined as the first object is a transparent umbrella in a widely spread shape, and the object of the sticker image (415) determined as the second object is a yellow umbrella in a dome shape, the electronic device (101) may derive that there is a difference in color and shape between the characteristics of the object (412) and the object of the sticker image (415), and generate a guide for editing.The electronic device (101) may display an interface (e.g., interface (416), interface (426)) that includes text asking the user for their editing intent (e.g., text such as "Should the umbrella shape remain original?", "Should only the color of the umbrella change?", "Should the shape change along with the color of the umbrella change?").

[0109] In one embodiment, the electronic device (101) may perform different operations based on a user response to an interface (e.g., interface (416), interface (426)). For example, in response to a user input indicating that an edit will be performed according to an editing guide included in the interface (e.g., a user selection input for a “Yes” icon), the electronic device (101) may generate an edited image according to the content of the provided guide. For example, in response to a user input indicating that an edit will not be performed according to an editing guide included in the interface (e.g., a user selection input for a “No” icon), the electronic device (101) may suggest a new editing guide to the user (e.g., an editing guide including text such as “Should I change the shape as well as the color of the umbrella?”).

[0110] According to one embodiment, the electronic device (101) can change at least one characteristic of a first object (e.g., object (412)) included in one of a plurality of frames (315) of an original video (301) based on a second object (e.g., object of a sticker image (415)) included in an image (414). The electronic device (101) can change at least one characteristic of the first object based on the characteristic of the second object and the relationship between the second object and the first object. For example, the electronic device (101) can change at least one characteristic of the object (412) based on the characteristic of the object of the sticker image (415) (e.g., yellow color, dome shape, umbrella category) and whether the category matches between the object of the sticker image (415) and the object (412). The electronic device (101) can change at least one characteristic of the first object using a generative model (e.g., generative model (370)). For example, the generative model (370) can change the characteristics (e.g., color, shape, pattern, size, style, material, etc.) of the first object based on the characteristics (e.g., color, shape, pattern, size, style, material, etc.) of the second object to generate a changed object (e.g., changed object (417), changed object (427)). The changed object (417) may be an object in which only the color is changed while maintaining the shape of the object (412). The changed object (427) may be an object in which both the shape and the color of the object (412) are changed. When a margin is generated by changing at least one characteristic of an object, the electronic device (101) can perform in-painting on the margin. For example, when the electronic device (101) changes the shape of an object, it can perform in-painting on the margin according to the difference in shape. In-painting can be an image processing technique that restores damaged areas in an image or fills in deleted parts by creating them.The electronic device (101) can perform in-painting using a generative model (e.g., generative model (370)). The generative model (370) can analyze a frame (e.g., frame (411)) and / or an object (e.g., object (412), an object of a sticker image (415)) and perform in-painting on the blank space to generate a natural result. For example, when the shape of the object (412) is changed based on the shape of the object of the sticker image (415), the generative model (370) can perform in-painting on the blank space generated by changing the widely spread shape into a dome shape to generate a natural result.

[0111] According to one embodiment, the electronic device (101) may generate a prompt including a first prompt portion corresponding to a first object and a second prompt portion corresponding to a second object, and input the prompt into a model (e.g., a generative model (370)) to generate a changed object (417). The electronic device (101) may generate the prompt so that objects other than the first object included in the frame are not changed. For example, the electronic device (101) may generate a prompt (e.g., “Change object A to object B and do not change other elements in the original”) to change the first object A included in one of the plurality of frames (315) of the original video (301) to the second object B. The electronic device (101) can generate a prompt in a natural language format (e.g., a human-readable prompt such as “Change object A to object B and do not change other elements in the original”), tokenize the prompt, and binarize it into a format that the generative model (370) can recognize. The prompt may be generated internally by the generative model (370). The generative model (370) can receive the prompt and generate a changed object (e.g., a changed object (417), a changed object (427)). The changed object may be generated by changing at least one characteristic of the first object based on the second object, rather than a second object simply being overlaid on the first object. The electronic device (101) can output a preview image including the changed object (e.g., a changed object (417), a changed object (427)).The electronic device (101) may generate a composite video (e.g., composite video (381) of FIG. 3) in response to a user input for the icon (418). The composite video (381) may be generated by including a modified object (e.g., modified object (417), modified object (427)) in frames containing a first object. For example, the composite video (381) may be generated by replacing a first object with a modified object (e.g., modified object (417), modified object (427)) in frames containing the first object.

[0112] According to one embodiment, when there is an object interacting with a first object (hereinafter, “interaction object”) in a frame, the electronic device (101) may change at least one characteristic of the interaction object based on a second object. When changing at least one characteristic of the first object based on a second object, the electronic device (101) may change the interaction object based on at least one characteristic of the first object and the characteristic of the second object. The electronic device (101) may input a prompt including information about the first object, information about the second object, and information about the interaction object into a model (e.g., a generative model (370)) to change the interacting object. The information about the first object included in the prompt may include audio information corresponding to the first object. The information about the second object included in the prompt may include audio information corresponding to the second object. The generative model (370) may receive a prompt including information about the first object, information about the second object, and information about the interaction object, and generate an interaction object having at least one characteristic changed. The generative model (370) can generate audio for a changed interactive object based on a prompt.

[0113]

[0114] FIGS. 5A to 5C are diagrams illustrating an interface for video processing according to one embodiment.

[0115] FIG. 5A is a drawing for explaining sound content according to one embodiment.

[0116] Referring to FIG. 5A, according to one embodiment, the electronic device (101) may display an original video (e.g., the original video (301) of FIG. 3) on a first portion of the display, and may display an image (515) (e.g., the image (415) of FIG. 4) on a second portion of the display. The electronic device (101) may obtain a first object included in one frame (e.g., the frame (512)) among a plurality of frames (e.g., the plurality of frames (315) of FIG. 3) of the original video. The electronic device (101) may change at least one characteristic of the first object based on a second object included in the image (515). The image (515) may include at least one image (e.g., a sticker image (516)). The operation of the electronic device (101) changing at least one characteristic of the first object based on the second object included in the image is substantially the same as that described with reference to FIG. 4, and thus, a redundant description thereof will be omitted.

[0117] According to one embodiment, when at least one characteristic of a first object included in a frame (512) is changed based on a second object, the electronic device (101) may generate acoustic content corresponding to the changed characteristic. The acoustic content may include sound information corresponding to a specific object. The electronic device (101) may determine whether generation of acoustic content associated with the changed characteristic is necessary and may generate acoustic content corresponding to the changed characteristic. For example, when audio of the original video (301) and audio corresponding to the first object are inconsistent (e.g., do not match) due to change in at least one characteristic (e.g., characteristics such as texture and shape) of the first object, the electronic device (101) may determine that generation of acoustic content associated with the changed characteristic is necessary. The electronic device (101) may output the acoustic content by including it in frames including the first object among a plurality of frames. The electronic device (101) may cause the audio content to replace the audio corresponding to the first object among the audio of the original video (301). For example, the electronic device (101) may cause the audio corresponding to the object (513) among the audio of the original video (301) (e.g., a character's speech such as "I brought an umbrella", audio including the sound of rain falling on an umbrella) to be replaced by the generated audio content (e.g., a character's speech such as "I brought a yellow umbrella", the sound of rain changed according to a change in the characteristics of the umbrella (e.g., characteristics such as texture and shape)). The electronic device (101) may include the generated audio content in frames including the first object among a plurality of frames and output them, thereby producing a more natural result (e.g., the composite video (381) of FIG. 3) without a sense of incongruity. The electronic device (101) can generate sound content corresponding to the changed characteristics using a generative model (e.g., the generative model (370) of FIG. 3).The electronic device (101) can generate a prompt including a first prompt portion corresponding to a first object and a second prompt portion corresponding to a second object. The electronic device (101) can input the prompt into a model (e.g., a generative model (370)) to generate audio content. The electronic device (101) can provide the user with an interface (e.g., interface (511)) that confirms whether to include audio content in a frame including the first object. For example, the electronic device (101) can confirm the user's intention by providing the user with an interface (511) that includes text (e.g., text such as "Do you want to change the voice to 'I brought a yellow umbrella'?") that asks the user for an intention to generate the audio content.

[0118]

[0119] FIG. 5b is a drawing for explaining an interface according to the size of a display according to one embodiment.

[0120] Referring to FIG. 5B, according to one embodiment, the electronic device (101) may include an electronic device that provides a large screen (e.g., an electronic device such as a foldable electronic device or a rollable electronic device). The foldable electronic device may be an electronic device that includes a foldable housing and a foldable display disposed within a space formed by the foldable housing. FIG. 5B illustrates an interface provided when the foldable display of the electronic device (101) is unfolded. The electronic device (101) may display an original video (e.g., the original video (301) of FIG. 3) on a first portion of the display, and may display an image (535) (e.g., the image (415) of FIG. 4) on a second portion of the display. The electronic device (101) may acquire a first object included in one frame (e.g., the frame (512)) among a plurality of frames (e.g., the plurality of frames (315) of FIG. 3) of the original video. The electronic device (101) can change at least one characteristic of the first object based on a second object included in the image (535). The image (515) can include at least one image (e.g., a sticker image (536)). The operation of the electronic device (101) changing at least one characteristic of the first object based on a second object included in the image is substantially the same as that described with reference to FIG. 4, and any overlapping descriptions thereof will be omitted.

[0121] According to one embodiment, the shape of the electronic device (101) may include a shape corresponding to an unfolded shape of the foldable display and a folded shape corresponding to a folded shape of the foldable display. The electronic device (101) may acquire a first object included in one frame (e.g., frame (532)) among a plurality of frames (315) of the original video (301). The first object may be determined by being segmented differently depending on the size of the portion where the original video (301) is displayed. For example, when the electronic device (101) is implemented as a large screen rather than a case where the electronic device (101) is implemented as a normal-sized display (e.g., the electronic device (101) illustrated in FIG. 5A), the size of the portion where the original video (301) is displayed is larger, so that the strength of segmentation can be increased to allow more and more diverse objects to be detected. The electronic device (101) can additionally acquire an object (534) (e.g., a person) in addition to the object (513) (e.g., an umbrella) as the output of segmentation. The screen (503) can include an interface (531) for providing a guide for editing to the user. The electronic device (101) can provide the guide for editing to the user in greater detail on a large screen than on a screen displayed on a standard-sized display. The interface for the editing guide, which is provided differently depending on the screen size displayed on the display, will be described in detail later with reference to FIG. 9.

[0122]

[0123] FIG. 5c is a drawing for explaining a changed object according to one embodiment.

[0124] Referring to FIG. 5C, according to one embodiment, the electronic device (101) may output the changed object (555) by including it in frames (e.g., frames (550)) that include a first object (e.g., a transparent umbrella) among a plurality of frames (315). The frames (550) may include at least one frame that includes the first object among the plurality of frames (315) of the original video (301). For example, the frames (550) may include a frame corresponding to a scene in which a main character is using or holding an umbrella, which is the first object. The frames (550) illustrated in FIG. 5C may be examples of only some of the frames that include the first object. The changed object (555) may replace the first object in the frames (550) that include the first object. The electronic device (101) may generate a modified object (555) by changing at least one characteristic of the first object based on a second object included in the sticker image (516), rather than simply overlaying the sticker image (516) on the location of the first object. The modified object (555) may be one in which at least one characteristic of the first object is changed based on the characteristic of the second object. For example, the modified object (555) may be one in which the color of the first object (e.g., a transparent umbrella) is changed to yellow, which is the color of the second object (e.g., a yellow umbrella).

[0125]

[0126] FIG. 6 is a diagram illustrating an interface for providing a user with processed video according to one embodiment.

[0127] Referring to FIG. 6, according to one embodiment, the electronic device (101) may output an edited video (e.g., the composite video (381) of FIG. 3) through a display (e.g., the display module (160) of FIG. 1). The electronic device (101) may generate a video (e.g., a preview video, a collection view video) by extracting only frames containing a changed object. The video (e.g., the preview video, the collection view video) may be generated by including the changed object in frames containing the first object, which is the original object, among a plurality of frames (315) of the original video (301). The electronic device (101) may provide the preview video to the user. Before finally generating the composite video (381), the electronic device (101) may provide an interface on a screen (e.g., the screen 601) that asks the user whether to play the preview video (e.g., a video including only frames containing the changed object). For example, the electronic device (101) can receive the user's intent by displaying an interface (610) including text asking the user's intent (e.g., text such as "Do you want to continue playing only the edited video?"). In response to the user's input to the interface (610) (e.g., a user selection input for a "Yes" icon), the electronic device (101) can output a preview video through the display module (160). For example, the electronic device (101) can output frames including a first object (e.g., frames (611 to 613)) including a changed object. The user can check the preview video to determine whether at least one characteristic of the first object has been properly changed as intended. The frames (611 to 613) may be generated by a generative model (e.g., the generative model (370)). The electronic device (101) can use the generative model (370) to generate an image (e.g., frames (611 to 613)) in which a changed object replaces the original object, the first object.The electronic device (101) can provide a user with a collection video. After finally generating the composite video (381), the electronic device (101) can provide the user with a collection video (e.g., a video including only frames containing changed objects). The electronic device (101) can provide the user with a collection video that allows the user to quickly check only the edited content along with the composite video (381). For example, the electronic device (101) can provide the user with a collection video that includes only frames containing changed objects (e.g., frames 611 to 613). The user can quickly check only the edited portion within the composite video (381) through the collection video. When the user checks the edited content through the final result video (e.g., the composite video (381)), it may take a lot of time to check, so the electronic device (101) can output a collection video that includes only frames containing changed objects (e.g., frames 611 to 613). A user can check the preview video to see whether at least one characteristic of the first object has been changed as intended. Frames (611 to 613) may be a portion of a frame containing the first object among a plurality of frames (315) of the original video (301). In order to reduce the time required, the electronic device (101) may generate a video (e.g., a preview video, a preview video) by including the changed object only for some frames, rather than all frames containing the first object.

[0128]

[0129] FIGS. 7A to 7C are diagrams for explaining a segmentation operation of an electronic device according to one embodiment.

[0130] Referring to FIGS. 7A to 7C , according to one embodiment, the electronic device (101) can differently recognize (e.g., detect) an object included in one frame among a plurality of frames (e.g., the plurality of frames (315)) of an original video (e.g., the original video (301) of FIG. 3 ). For example, the electronic device (101) can differently detect objects included in the plurality of frames (315) by varying the type and number of objects. The electronic device (101) can recognize and classify objects included in the plurality of frames (315) by using an object processing module (e.g., the object processing module (330)). The electronic device (101) can differently detect a plurality of objects by differently setting the intensity of segmentation according to the size of a portion where the original video (301) is displayed on a display (e.g., the display module (160) of FIG. 1 ). For example, the electronic device (101) can set the segmentation level to be high when the size of the portion where the original video is displayed is large, and can set the segmentation level to be low when the size of the portion where the original video is displayed is small. Since the user's visibility can be improved as the size of the portion where the original video (301) is displayed on the display is large, the electronic device (101) can perform segmentation in detail to output more objects. The electronic device (101) can perform segmentation on one frame among a plurality of frames (315) of the original video (301) to output information about the object (e.g., information about a box used to detect the object (e.g., information such as the size of the box, the position of the box)), and can detect a plurality of objects differently based on the output information about the object (e.g., information about the object such as the size of the object).For example, the electronic device (101) may be set to detect only objects whose size is greater than or equal to a predetermined threshold value among the segmented and output objects, so that objects smaller than a predetermined size may not be detected. The electronic device (101) may set a different threshold value depending on the size of the portion of the display where the original video (301) is displayed. For example, the electronic device (101) may set a smaller threshold value as the portion of the display where the original video (301) is displayed becomes larger, so that objects of smaller sizes may also be detected. As another example, the electronic device (101) may set a larger threshold value as the portion of the display where the original video (301) is displayed becomes smaller, so that only objects larger than or equal to a predetermined size may be detected. The electronic device (101) may perform segmentation to obtain a first object included in one frame among a plurality of frames of the original video. The first object may be determined by segmenting multiple objects contained in one frame among multiple frames of the original video differently according to the size of the portion of the display where the original video is displayed. The first object may include an object that can be the target of editing.

[0131] FIG. 7A illustrates a case where an electronic device (101) according to one embodiment is an electronic device including a normal-sized display, and FIGS. 7B and 7C may illustrate a case where an electronic device (101) according to one embodiment is an electronic device providing a large screen (e.g., a foldable electronic device, a multi-foldable electronic device). The foldable electronic device may include a folding axis. The foldable electronic device may include a display about twice the size of a normal-sized display. The multi-foldable electronic device may include a plurality of folding axes. The multi-foldable electronic device may include a display about n times the size of a normal-sized display (where n is a natural number greater than or equal to 3). Since it is common for the size of the portion where a video (e.g., the original video (301)) is displayed to increase as the size of the display increases, the electronic device (101) may determine the first object by segmenting the frame differently depending on the size of the display.

[0132] According to one embodiment, the electronic device (101) may detect the first object differently depending on the size of the portion (e.g., region (702), region (706), region (711), and region (731)) on the display where the original video (301) is displayed. Region (702) may include a portion where the original video (301) is displayed in a portrait mode (701) of a normal-sized display. Region (706) may include a portion where the original video (301) is displayed in a landscape mode (705) of a normal-sized display. Region (711) may include a portion where the original video (301) is displayed in a large screen (710) displayed on a display that is about twice the size of a normal-sized display. Region (731) may include a portion where the original video (301) is displayed in a large screen (730) displayed on a display that is about three times the size of a normal-sized display. The electronic device (101) can detect more objects as the area where the original video (301) is displayed becomes larger. For example, the electronic device (101) can detect only one object (703) (e.g., an umbrella) for the original video (301) displayed in the area (702), and can detect more objects (713) (e.g., a person) for the original video (301) displayed in the area (711) larger than the area (702). The electronic device (101) can further detect the object (713) (e.g., a person) by segmenting it into objects (733) (e.g., a face) and objects (734) (e.g., clothes) for the original video (301) displayed in the area (731).

[0133] According to one embodiment, since the size of the portion where the original video (301) is displayed may vary depending on the usage mode even in the same electronic device (101), the electronic device (101) may detect objects differently depending on the usage mode. For example, since the sizes of the portions (e.g., area (702), area (706)) where the original video (301) is displayed may be different in portrait mode (701) and landscape mode (705) of an electronic device (101) including a normal-sized display, the electronic device (101) may detect more various objects when the original video (301) is displayed larger (e.g., landscape mode (705)) than when the original video (301) is displayed smaller (e.g., portrait mode (701)). For another example, if the electronic device (101) is an electronic device including a foldable display (e.g., such as the electronic device illustrated in FIGS. 7b and 7c), when the foldable display corresponds to an unfolded form, a greater number of objects can be detected from the original video (301) than when the display corresponds to a folded form. The electronic device (101) can detect the folded form of the foldable display based on the sensing result of the sensor module (e.g., the sensor module (176) of FIG. 1) and can detect the form of the electronic device according to the folded form of the foldable display. The electronic device (101) can increase the strength of the segmentation or set a smaller threshold value to be applied to the size of the detected object so that more types of objects can be detected. The electronic device (101) can reduce the number of detected objects (e.g., types of objects, number of objects) by decreasing the strength of the segmentation or setting a larger threshold value to be applied to the size of the detected object. The electronic device (101) can display the detected objects on the screen.

[0134]

[0135] FIG. 8 is a diagram illustrating an operation of adding an object to a video according to one embodiment.

[0136] Referring to FIG. 8, according to one embodiment, the electronic device (101) may display a second object in one frame (e.g., frame 810) among a plurality of frames (e.g., frames 315) of an original video (e.g., original video (301) of FIG. 3). The electronic device (101) may not only change at least one characteristic of a first object included in one frame among the plurality of frames (315) of the original video (301) based on the second object, but may also add the second object to the frame (810) and display it as a separate additional object. The second object may include a type of object not included in the frame (810). The electronic device (101) may output the second object by including it in the frame (810) without replacing the first object with an object changed based on the second object (e.g., changed object (417) of FIG. 4). For example, the electronic device (101) may output the second object as a separate, independent object within the frame (810). In response to a user input for an icon (811) for adding an object, the electronic device (101) may display an image (e.g., image (812)) in a second portion of the display. The second portion may be a different portion from the first portion where the original video (301) is displayed. The user may select an image (e.g., image (812)) to be added to the frame (810). The electronic device (101) may determine the second object to be added to the frame (810) based on the user's selection input. Before adding the second object to the frame (810), the electronic device (101) may determine whether the second object is appropriate to be added to the frame (810) (e.g., a frame to be edited). For example, the electronic device (101) can analyze information about the scene corresponding to the frame (810) (e.g., information such as atmosphere, situation, and location) to determine whether it matches a second object selected by the user.The second object can be placed at a location selected by the user within the frame (810) (e.g., a location of a piano selected by the user within the screen (802). The user can drag and drop an image corresponding to the second object to position it at a desired location. The user's action of selecting a location for the second object can be implemented in various ways, and it should be noted that the drag-and-drop action is merely an example. Even if the user does not select a location for the second object, the electronic device (101) can analyze a location within the frame (810) suitable for placing the second object and display the second object at that location. The electronic device (101) can generate video contents corresponding to the second object. For example, if the electronic device (101) determines that the second object is suitable for addition to the frame (810), it can generate video contents corresponding to the second object. The electronic device (101) can generate video content based on the characteristics of the second object (e.g., characteristics such as category, size, purpose, and interaction with other objects). The electronic device (101) can display an interface that asks the user whether to generate the video content based on the characteristics of the second object. For example, the electronic device (101) can display an interface (813) including text asking whether to generate video content corresponding to the second object (e.g., text such as "I placed a piano. Would you like to generate a video of it playing?") on the screen (803) through the display module (160). The electronic device (101) can generate the video content in response to a user input to the interface (813). For example, the electronic device (101) can generate the video content in response to a user input (e.g., a user selection input for a "Yes" icon).The video content may be generated by a generative model (e.g., generative model (370)) based on the characteristics of the second object and the relationship between the second object and the frame (810). The electronic device (101) may use the generative model (370) to generate the video content based on information such as the interaction between the second object and other objects included in the frame (810) (e.g., a person object included in the frame (810)) and the relationship with the situation of the scene corresponding to the frame (810). For example, the electronic device (101) may generate video content composed of frames about a person playing the piano based on the relationship between the person object included in the frame (810) and the second object (e.g., a piano). When a second object is added to the original video (301), the number of frames increases, and thus the overall playback time may become longer. For example, as video content is generated, the playback time of the original video (301) may increase from the existing playback time (e.g., 10:00) to a playback time (e.g., playback time (815) (e.g., 10:28)) that is the amount of time of the generated video content added to the original video. The electronic device (101) may provide the user with an interface (814) that asks whether to add the generated video to the original video. In response to the user input to the interface (814), the electronic device (101) may output the final result (e.g., composite video (381)). The electronic device (101) may output the generated video content by including it in the original video. The electronic device (101) may store the generated video content in association with the original video (301).

[0137] According to one embodiment, the electronic device (101) may delete an object included in one frame (e.g., frame (810)) among a plurality of frames (e.g., frames (315)) of the original video (301). The electronic device (101) may delete an object included in the frame (810) in response to a user input to delete the object included in the frame (810). For example, the electronic device (101) may delete a person object included in the frame (810) in response to a user input to delete the person object included in the frame (810) and perform in-painting on an area where the person object existed. The electronic device (101) may perform in-painting on an area where the deleted object existed using a generative model (e.g., generative model (370) of FIG. 3). The electronic device (101) may delete audio corresponding to the deleted object from among audio of the original video (301). For example, the electronic device (101) can identify and delete audio (e.g., audio of a person walking) corresponding to a deleted human object from the audio of the original video (301). When deleting an object within a frame (810), the electronic device (101) can also generate new audio content based on the context of the frame (810). For example, when deleting a human object from a frame of a scene in which a dog barks at a person, the sound of the dog barking cheerfully can be generated to replace the existing audio. The electronic device (101) can generate new audio content according to the context of the frame (810) using a generative model (e.g., the generative model (370)).

[0138]

[0139] FIG. 9 is a diagram illustrating a video processing guide interface provided to a user according to one embodiment.

[0140] Referring to FIG. 9, according to one embodiment, the electronic device (101) may provide different video processing (e.g., editing) guides depending on the size of the display. The interface (901) and the interface (911) may be examples of editing guides that the electronic device (101) may provide to the user when a second object is added to the original video described with reference to FIG. 8 (e.g., the original video (301) of FIG. 3). The electronic device (101) may provide different interfaces depending on the size of the display. For example, the electronic device (101) may provide an interface (901) including text (e.g., text such as "I've placed the piano. Should I generate a video of it playing?") on a normal-sized display, and may provide an interface (911) including relatively long sentences of text (e.g., text such as "I've placed the piano. Should I generate a video of it playing the song "Everything happens to me" that I've been listening to a lot lately?") on a display capable of providing a large screen (e.g., a display such as a foldable display).

[0141] According to one embodiment, the interface (902) and the interface (912) may be examples of editing guides that can be provided when changing at least one characteristic of a first object included in one of a plurality of frames (e.g., the plurality of frames (315)) of the original video (301) described with reference to FIGS. 4 to 7C based on a second object. The electronic device (101) may provide different interfaces depending on the size of the display. For example, the electronic device (101) may provide an interface (902) that includes text (e.g., text such as "I replaced the violin with a piano. Should I create a video of the violin being played?") on a normal-sized display, and may provide an interface (912) that includes a preview image (e.g., an image such as an image of a person playing the piano) along with a relatively long sentence of text (e.g., text such as "I replaced the violin in the video with a piano. Should I create a video of the violin being played in a close-up, as in the image below?") on a display capable of providing a large screen (e.g., a display such as a foldable display). The electronic device (101) can generate a preview image using a generative model (e.g., the generative model (370) of FIG. 3).

[0142]

[0143] FIGS. 10A to 10C are diagrams illustrating an interface for video processing according to one embodiment.

[0144] Referring to FIGS. 10A to 10C , according to one embodiment, an electronic device (101) may include a foldable display. The electronic device (101) may acquire a visual object (e.g., an object (1012)) included in one frame (e.g., a frame (1011)) among a plurality of frames (e.g., a plurality of frames (315)) of an original video (e.g., an original video (301) of FIG. 3 ), and may generate audio content corresponding to the visual object based on at least one characteristic of the visual object. The visual object may include an object (e.g., a graphic object) having a visually distinguishable shape and / or structure included in an image and / or video. The electronic device (101) may generate audio content based on at least one characteristic of the visual object. The audio content may include audio different from audio corresponding to the visual object among audio included in the original video (301). For example, the electronic device (101) may generate audio corresponding to the object (1012) (e.g., a hot dog) and other audio content (e.g., other audio content associated with a hot dog) based on at least one characteristic of the object (1012) (e.g., a hot dog) which is a visual object included in the frame (1011). The electronic device (101) may output the audio content by including it in frames including the visual object among the plurality of frames (315). The electronic device (101) may generate the audio content based on at least one multimedia content including audio. The at least one multimedia content may include an audio source and / or video content different from the original video (301). The electronic device (101) may determine one object among at least one object included in the at least one multimedia content as an object candidate for changing the visual object.The electronic device (101) may generate audio content based on the object candidate when the category correlation between the object candidate and the visual object satisfies a predetermined threshold (e.g., when the categories match). The electronic device (101) may determine whether the multimedia content is to be used to generate audio content by determining whether the categories between the object candidate and the visual object match. The operation of the electronic device (101) to determine whether the categories match is substantially the same as the operation of the category matching module (360) described with reference to FIG. 3, and thus, any redundant description thereof will be omitted. At least one multimedia content may be associated with at least one characteristic of the visual object. The electronic device (101) may generate audio content using a generative model (e.g., the generative model (370) of FIG. 3). The electronic device (101) may generate a prompt including a first prompt portion and a second prompt portion. The first prompt portion may correspond to the visual object. The second prompt portion may correspond to audio corresponding to the object candidate among audio included in the multimedia content. An electronic device (101) can input a prompt into a model (e.g., a generative model (370)) to generate audio content. The audio content can have a correlation level greater than a predetermined threshold value with at least one characteristic of a visual object.

[0145] According to one embodiment, the electronic device (101) may generate a plurality of audio contents based on at least one characteristic of a visual object. For example, the electronic device (101) may generate a plurality of audio contents associated with an object (1012) (e.g., a plurality of audio contents associated with at least one characteristic of a hot dog). The electronic device (101) may provide an interface for receiving a user's selection for the plurality of audio contents. For example, the electronic device (101) may provide an interface (e.g., interface (1031)) for receiving a pre-listening and selection for each of the plurality of audio contents. Based on the user's selection, the electronic device (101) may determine an audio content to be included in frames including a first object among the plurality of audio contents and output. The electronic device (101) may cause the audio content to replace audio corresponding to the visual object among audio of the original video (301). The electronic device (101) can generate video content (e.g., composite video (381) of FIG. 3) in which audio corresponding to a visual object is replaced with sound content in frames containing a visual object among a plurality of frames of the original video (301). The electronic device (101) can store the video content (e.g., composite video (381)) in association with the original video (301).

[0146] According to one embodiment, the electronic device (101) can display an original video (301) (e.g., one frame (1011) of a plurality of frames (315) of the original video (301)) on a first portion of the display, and display an image on a second portion of the display. The electronic device (101) can display an image (e.g., image (1014)) stored in the electronic device (101) on the second portion of the display. For example, the electronic device (101) can display the original video (301) and an interface for editing the original video (e.g., interface (1013)) on the first portion of the display, and display an image stored in a gallery on the second portion. The electronic device (101) can change at least one characteristic of a first object included in one of the plurality of frames (315) of the original video (301) based on a second object included in the image (1014). The electronic device (101) can generate acoustic content corresponding to the changed characteristic. The electronic device (101) can output (play) the changed object and acoustic content by including them in frames including the first object among the plurality of frames (315). The operation of the electronic device (101) changing at least one characteristic of the first object based on the second object included in the image (1014) is substantially the same as the operation of the electronic device (101) described with reference to FIGS. 1 to 9, and thus, any overlapping descriptions will be omitted. The electronic device (101) can determine the first object included in the frame (1011) based on a user's input before or without performing segmentation. For example, the electronic device (101) can determine the selected object (1012) as the first object based on a user's selection input for the frame (1011). The first object may include an object to be edited.The electronic device (101) may receive a user's selection for an image displayed in the second portion and determine a second object included in the selected image. For example, the electronic device (101) may receive a user's selection for an object (1015) included in the image (1014) and determine the object (1015) as the second object. The electronic device (101) may display the selected first object and second object by distinguishing them from other objects. For example, the electronic device (101) may display the first object and the second object by distinguishing them with a dotted line. It should be noted that the method by which the electronic device (101) displays the first object and the second object by distinguishing them from other objects may vary and is not limited to the above example.

[0147] According to one embodiment, the electronic device (101) may change at least one characteristic of the first object based on the relationship between the second object and the first object. For example, the electronic device (101) may determine whether the category (e.g., food) of the second object (e.g., object (1015) (e.g., hamburger)) matches the category (e.g., food) of the first object (e.g., object (1012) (e.g., hot dog)), and may determine whether at least one characteristic of the first object can be changed using the second object. The electronic device (101) may change at least one characteristic of the first object based on the characteristic of the second object. If the original video (301) is streaming video content, the electronic device (101) may display the second object by overlaying it on top of the first object, since it is impossible to edit the video received in real time. The electronic device (101) may analyze information about the original video (301) being streamed (e.g., information such as topic, situation, location, and included objects) when the original video (301) is streaming video content, and provide customized content to the user. The electronic device (101) may display an interface (1013) on the screen (1010) to provide an editing guide to the user. The electronic device (101) may change at least one characteristic of a first object and at the same time confirm the user's intention whether to generate audio content corresponding to the changed characteristic. For example, the electronic device (101) may confirm the user's intention by displaying an interface (1013) on the screen (1010) that includes text for guiding the user in editing (e.g., text such as "Would you like to change the object image selected in the currently playing video to the object image selected in the gallery with sound?"). The electronic device (101) can display a screen (1030) through a display module (e.g., display module (160) of FIG. 1) in response to a user input to the interface (1013).The electronic device (101) can generate a plurality of audio contents based on the changed characteristics. The electronic device (101) can generate the plurality of audio contents using a generative model (e.g., the generative model (370) of FIG. 3). The electronic device (101) can display the plurality of audio contents on a screen (1030) using a corresponding interface (1031). The electronic device (101) can generate a first prompt portion corresponding to a first object (e.g., object (1012)) and a prompt corresponding to a second object (e.g., object (1015)). The prompt corresponding to the second object can include information about the second object. For example, the information about the second object can include additional information that can be obtained by identifying the object (1015), such as the shape of the object (1015) (e.g., a hamburger) and the brand (e.g., a hamburger franchise brand). The electronic device (101) can input a prompt into a model (e.g., a generative model (370)) to generate audio content (e.g., a plurality of audio contents). The generative model (370) can generate audio content that reflects information that identifies a second object based on the prompt. For example, based on information about the brand of the second object (1015) (e.g., a hamburger), the model can generate audio content such as “I’ll eat a Burger King hamburger today” or “I’ll eat a Shake Shack burger today.” The electronic device (101) can display an interface (1033) on the screen (1030) that asks whether to proceed with editing with an audio content recommended by an artificial intelligence model among the plurality of audio contents. In response to a user input on the interface (1033), the electronic device (101) can generate a video (e.g., a composite video (381)) or provide a screen (1050) so that the user can directly select audio content.For example, in response to a user input (e.g., a user selection input for a "Yes" icon) indicating that the electronic device (101) wishes to edit audio content recommended by an artificial intelligence model, the electronic device (101) may determine audio content with the highest relevance to be output by including the audio content in frames containing the first object. For example, in response to a user input (e.g., a user selection input for a "No" icon), the electronic device (101) may display a screen (1050) through the display module (160).

[0148] According to one embodiment, the electronic device (101) may display an interface on the screen (1050) for receiving a selection input for a plurality of audio contents from a user. The electronic device (101) may display an interface (1051) including text guiding the user to make a selection (e.g., text such as "Would you like to change the object image selected in the gallery to the sound of the currently playing video?") on the screen (1050). The electronic device (101) may provide the user with a preview of the plurality of audio contents. For example, the electronic device (101) may play the corresponding audio contents in response to the user's selection input for the interface (1031). The electronic device (101) may provide the preview of the contents together with a frame (1052) including the changed object. The user may preview each audio contents and determine which audio contents to output by including them in the frames including the first object. The electronic device (101) can receive a user's selection of at least one audio content among a plurality of audio contents and a user's selection of an icon (1053) to generate a composite video (381).

[0149]

[0150] FIG. 11 is a diagram illustrating an interface for video processing of an electronic device according to one embodiment.

[0151] Referring to FIG. 11, according to one embodiment, an electronic device (101) may include a slidable display. The slidable display may include a display that can expand or contract depending on a user's gesture. The slidable display may be implemented as a flexible display. The flexible display may be disposed on the electronic device (101) in a bendable, foldable, or rollable form. FIG. 11 may illustrate an interface provided to an electronic device (101) including a slidable display. The electronic device (101) may expand the display based on a user's input (e.g., an input for expanding the screen). For example, the electronic device (101) may expand the display based on a user's gesture (1110) (e.g., a gesture such as a downward swiping gesture). It should be noted that the user's input for expanding the screen of the electronic device (101) is not limited to the form of a gesture input. For example, the electronic device (101) can expand the display in response to an input containing the user's intention to expand the screen, regardless of the format, such as a user's selection input for an interface included in the screen, a user's input for a physical switch for expanding the screen, etc. The electronic device (101) can expand the display differently depending on the form of the display. For example, an electronic device including a different form of slider display may expand the display to the right in response to a gesture of sliding to the right. The screen (1101) may be displayed on the display in the form before the user's gesture (1110). The electronic device (101) can expand the display in response to the user's gesture (1110) to display a screen such as the screen (1102). The expanded display (1125) may include an interface for taking pictures.The display (1125) can display an icon (1120) for taking a picture and data received from a camera (e.g., the camera module (180) of FIG. 1). The electronic device (101) can change an object (1130) included in an original image using an object (e.g., a bicycle) included in an image captured by the camera module (180). The operation of the electronic device (101) changing the object (1130) included in the original image is substantially the same as the operation of the electronic device (101) changing at least one characteristic of the first object described with reference to FIGS. 4 to 7C, and any overlapping description thereof will be omitted. When the electronic device (101) receives a captured image, the electronic device (101) can reduce the display to its original size. The electronic device (101) can display a screen (1103) including the changed object (1140) on the reduced display.

[0152]

[0153] Fig. 12 shows a flowchart of an operating method of an electronic device according to one embodiment.

[0154] Referring to FIG. 12, according to one embodiment, operations 1410 to 1440 may be substantially identical to operations performed by the electronic device (e.g., the electronic device (101) of FIG. 1) described with reference to FIGS. 1 to 11. Therefore, any redundant description will be omitted.

[0155] According to one embodiment, operations 1210 to 1250 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).

[0156] In operation 1210, the electronic device (101) may acquire a visual object included in one frame among a plurality of frames (e.g., the plurality of frames (315)) of an original video (e.g., the original video (301) of FIG. 3). The visual object may be determined by segmenting a plurality of objects included in one frame among a plurality of frames (e.g., the plurality of frames (315)) of the original video (e.g., the original video (301)) differently according to the size of a portion where the original video (e.g., the original video (301)) is displayed on a display (e.g., the display module (160) of FIG. 1).

[0157] In operation 1230, the electronic device (101) may generate acoustic content corresponding to a visual object based on at least one characteristic of the visual object. The electronic device (101) may input a prompt into a model (e.g., the generative model (250) of FIG. 2) to generate the acoustic content.

[0158] In operation 1250, the electronic device (101) may output audio content by including it in frames containing visual objects among a plurality of frames. The electronic device (101) may cause the audio content to replace audio corresponding to the visual object among the audio of the original video.

[0159] In one embodiment, operations 1210 through 1250 may be performed sequentially, but are not limited thereto. For example, two or more operations may be performed in parallel.

[0160]

[0161] Fig. 13 shows a flowchart of an operating method of an electronic device according to one embodiment.

[0162] Referring to FIG. 13, operations 1310 to 1370 may be substantially the same as operations performed by the electronic device (e.g., the electronic device (101) of FIG. 1) described with reference to FIGS. 1 to 11. Therefore, redundant descriptions will be omitted.

[0163] According to one embodiment, operations 1310 to 1370 may be understood to be performed in a processor (e.g., processor (120) of FIG. 1) of an electronic device (e.g., electronic device (101) of FIG. 1).

[0164] In operation 1310, the electronic device (101) may display an original video (e.g., original video (301) of FIG. 3) on a first portion of the display and an image (e.g., image (414) of FIG. 4) on a second portion of the display.

[0165] In operation 1330, the electronic device (101) can change at least one characteristic of a first object included in one of a plurality of frames (e.g., the plurality of frames (315) of FIG. 3) of an original video (e.g., the original video (301)) based on a second object included in an image (e.g., the image (414)).

[0166] In operation 1350, the electronic device (101) can generate acoustic content corresponding to the changed characteristic.

[0167] In operation 1370, the changed object and sound content can be output by including them in frames containing the first object among a plurality of frames (e.g., a plurality of frames (315)).

[0168] In one embodiment, operations 1310 through 1370 may be performed sequentially, but are not limited thereto. For example, two or more operations may be performed in parallel.

[0169]

[0170] An electronic device (101) according to one embodiment may include a display (160), at least one processor (120), and a memory (130) for storing instructions.

[0171] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain a visual object included in one of a plurality of frames (315) of an original video (301).

[0172] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate acoustic content corresponding to the visual object based on at least one property of the visual object.

[0173] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to output (play) the audio content by including it in frames among the plurality of frames (315) that include the visual object.

[0174] The above audio content may be generated based on at least one multimedia content including audio.

[0175] The above-mentioned audio content may have a correlation level greater than a defined threshold value with at least one characteristic of the above-mentioned visual object.

[0176] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to determine one object among at least one object included in the at least one multimedia content as an object candidate for changing the visual object.

[0177] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate the audio content based on the object candidate when the correlation between the category of the object candidate and the visual object satisfies a predetermined threshold value.

[0178] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a prompt including a first prompt portion and a second prompt portion.

[0179] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to input the prompt to the model (250; 370) to generate the acoustic content.

[0180] The above first prompt portion may correspond to the visual object.

[0181] The second prompt portion may correspond to audio corresponding to the object candidate among audio included in the multimedia content.

[0182] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to output the audio content by replacing audio corresponding to the visual object among the audio of the original video (301).

[0183] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate video content in which audio corresponding to the visual object in frames containing the visual object among the plurality of frames is replaced with the sound content.

[0184] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to store the video content in association with the original video in the electronic device (101).

[0185] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a plurality of audio contents based on at least one characteristic of the visual object.

[0186] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to provide an interface (1031) for receiving a user's selection of the plurality of audio contents.

[0187] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to determine, based on the user's selection, which audio content to output by including the visual object in frames among the plurality of audio contents.

[0188] The above visual object may be determined by segmenting multiple objects included in one frame among multiple frames (315) of the original video (301) differently according to the size of the portion where the original video is displayed on the display (160).

[0189] An electronic device (101) according to one embodiment may include a display (160), at least one processor (120), and a memory (130) for storing instructions.

[0190] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display an original video (301) on a first portion of the display (160) and to display an image (414; 515; 535; 1014) on a second portion of the display.

[0191] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate acoustic content corresponding to the changed characteristic.

[0192] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to output (play) the changed object and the sound content by including them in frames including the first object among the plurality of frames.

[0193] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a prompt including a first prompt portion corresponding to the first object and a second prompt portion corresponding to the second object.

[0194] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to input the prompt to the model (250; 370) to generate the acoustic content.

[0195] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to receive a user's selection of the image (414; 515; 535; 1014) and determine an object included in the selected image as the second object.

[0196] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to change at least one characteristic of the first object based on a characteristic of the second object and a relationship between the second object and the first object.

[0197] The relationship between the second object and the first object may include a matching status of the category of the second object and the category of the first object.

[0198] The first object may be determined by segmenting multiple objects included in one frame among multiple frames of the original video (301) differently according to the size of the portion where the original video (301) is displayed on the display (160).

[0199] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to provide an interface for the user to confirm whether to include the audio content in a frame containing the first object.

[0200] The second part may be different from the first part.

[0201] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display the second object in one of a plurality of frames (315) of the original video (301).

[0202] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate video content corresponding to the second object.

[0203] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to output the generated video content by including it in the original video (301).

[0204] A method of operating an electronic device (101) according to one embodiment may include an operation of obtaining a visual object included in one frame among a plurality of frames (315) of an original video (301).

[0205] The above method of operation may include an operation of generating acoustic content corresponding to the visual object based on at least one characteristic of the visual object.

[0206] The above operating method may include an operation of outputting the audio content by including it in frames among the plurality of frames that include the visual object.

[0207] The above audio content may be generated based on at least one audio source different from the original video (301).

[0208] The above-mentioned acoustic content may have a correlation level greater than a predetermined threshold value with at least one characteristic of the above-mentioned visual object.

[0209] The act of generating the above audio content may include an act of generating a prompt including a first prompt portion corresponding to the visual object and a second prompt portion corresponding to the at least one sound source.

[0210] The action of generating the above-described acoustic content may include an action of generating the above-described acoustic content by inputting the above-described prompt into the model (250; 370).

[0211] The above outputting operation may include an operation of outputting the audio content by replacing the audio corresponding to the visual object among the audio of the original video (301).

[0212] An electronic device according to one embodiment may include a display (160), at least one processor (120), and a memory (130) for storing instructions.

[0213] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display an original video including a visual object and audio corresponding to a designated object through the display (160).

[0214] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to receive a user input selecting the visual object while at least one image frame including the visual object among a plurality of image frames of the original video is displayed.

[0215] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain another acoustic object corresponding to the visual object in response to the user input, based at least in part on external acoustic information about the original video.

[0216] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to store the other audio object in association with the original video so that the other audio object is played instead of the audio when an image frame including the visual object among the plurality of image frames is displayed.

[0217] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain a first set of one or more image frames from among the plurality of image frames.

[0218] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain segment information for distinguishing a plurality of image objects including the visual object from the first set.

[0219] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to, in response to the user input, distinguish the visual object among a plurality of image objects included in the at least one image frame based at least in part on the segment information.

[0220] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain at least a portion of the external audio information from at least one additional video content different from the original video.

[0221] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to, in response to the user input, display through the display (160) one or more indications, each representing one or more audio objects, including the other audio object.

[0222] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to perform the operation of acquiring the other acoustic object based on another user input received for an indication representing the other acoustic object among the one or more indications displayed through the display (160).

[0223] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to select the other audio object from the one or more audio objects based at least in part on the association between category information corresponding to the visual object and category information corresponding to the other audio object satisfying a specified condition.

[0224] The one or more acoustic objects may include, respectively, corresponding first acoustic objects and second acoustic objects, each corresponding to different first category information and second category information.

[0225] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to provide the first acoustic object as the other acoustic object instead of the second acoustic object, at least in part based on a first correlation between the category information and the first category information being higher than a second correlation between the category information and the second category information.

[0226] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a frame set including the other acoustic object and an image object associated with the other acoustic object.

[0227] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to store a frame set including the other acoustic object and an image object associated with the other acoustic object in association with the video content.

[0228] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a prompt, based at least in part on the user input, that includes at least a portion of the original video, a first prompt portion corresponding to the visual object, and a second prompt portion corresponding to the other acoustic object.

[0229] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the prompt to the model (250; 370).

[0230] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to receive, in response to the prompt, from the model (250; 370), one or more image frames synthesized by the model (250; 370) in association with the other acoustic object.

[0231] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to provide, through the display (160), one or more other image objects obtained from at least one additional video content different from the original video.

[0232] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to store the selected other image object from among the one or more other image objects in association with the original video so that the selected other image object is played in place of the visual object when a frame including the visual object from among the plurality of image frames is displayed.

[0233] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to perform an operation of synthesizing an area corresponding to the visual object on a frame including the visual object among the plurality of image frames with the selected other image object, as part of a method of storing the selected other image object in association with the original video.

[0234] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to select the selected other image object from the one or more other image objects, at least in part based on the association between category information corresponding to the visual object and category information corresponding to the other image object satisfying a specified condition.

[0235] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate a prompt, based at least in part on the user input, that includes at least a portion of the original video, a first prompt portion corresponding to the visual object, and a second prompt portion corresponding to the selected other image object.

[0236] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to transmit the prompt to the model (250; 370).

[0237] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to, in response to the prompt, receive from the model (250; 370) the visual object contained in one or more image frames synthesized by the model (250; 370) synthesized with the other selected image object.

[0238] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to select an acoustic object associated with the selected other image object as the other acoustic object.

[0239] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to play at least a portion of the other sound object together with the other selected image object when the other sound object is played.

[0240] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from this document.

[0241]

[0242] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0243] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0244] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0245] Various embodiments of the present document may be implemented as software (e.g., a program (1740)) including one or more instructions stored in a storage medium (e.g., an internal memory (1736) or an external memory (1738)) readable by a machine (e.g., an electronic device (1701)). For example, a processor (e.g., a processor (1720)) of the machine (e.g., an electronic device (1701)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0246] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0247] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separately arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device (101), display (160); At least one processor (120) comprising processing circuitry; and Memory for storing instructions (130) Including, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Obtain a visual object contained in one frame among multiple frames (315) of the original video (301), Generating acoustic content corresponding to the visual object based on at least one property of the visual object, To output (play) the above audio content by including it in frames among the plurality of frames that include the visual object, Electronic device (101).

2. In paragraph 1, The above audio content is, It is generated based on at least one multimedia content including audio, The above audio content is, At least one characteristic of the visual object has a correlation level greater than a defined threshold value, Electronic device (101).

3. In any one of paragraphs 1 and 2, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: determining one object among at least one object included in the at least one multimedia content as an object candidate to change the visual object; If the correlation between the category of the object candidate and the visual object satisfies a predetermined threshold value, the sound content is generated based on the object candidate. Electronic device (101).

4. In any one of paragraphs 1 to 3, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Create a prompt that includes a first prompt part and a second prompt part, By inputting the above prompt into the model (250; 370), the above sound content is generated, The first prompt part above is, Corresponds to the above visual object, The second prompt part above is, Among the audio included in the above multimedia content, the audio corresponding to the object candidate, Electronic device (101).

5. In any one of paragraphs 1 to 4, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: The above audio content is output by replacing the audio corresponding to the visual object among the audio of the original video (301). Electronic device (101).

6. In any one of paragraphs 1 to 5, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Generating video content in which audio corresponding to the visual object is replaced with the sound content in frames containing the visual object among the plurality of frames, To store the above video content in the electronic device (101) in association with the above original video, Electronic device (101).

7. In any one of paragraphs 1 to 6, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Generating a plurality of audio contents based on at least one characteristic of the visual object, Provides an interface (1031) for receiving a user's selection for the above multiple audio contents, Based on the user's selection, the audio content to be output by including the visual object in the frames among the plurality of audio contents is determined. Electronic device (101).

8. In any one of paragraphs 1 to 7, The above visual object is, A plurality of objects included in one frame among a plurality of frames (315) of the original video (301) are segmented differently and determined according to the size of the portion where the original video (301) is displayed on the display (160). Electronic device (101).

9. In the electronic device (101), display (160); At least one processor (120) comprising processing circuitry; and Memory for storing instructions (130) Including, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Displaying an original video (301) on a first part of the display (160), and displaying an image (414; 515; 535; 1014) on a second part of the display, Modify at least one characteristic of a first object included in one of a plurality of frames (315) of the original video (301) based on a second object included in the image (414; 515; 535; 1014), Generate acoustic content corresponding to the changed characteristics, To output (play) the changed object (417; 555) and the sound content by including them in the frames containing the first object among the plurality of frames (315). Electronic device (101).

10. In paragraph 9, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Generate a prompt including a first prompt portion corresponding to the first object and a second prompt portion corresponding to the second object, By inputting the above prompt into the model (250; 370), the above sound content is generated. Electronic device (101).

11. In any one of paragraphs 9 to 10, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Receiving a user's selection for the above image (414; 515; 535; 1014) and determining an object included in the selected image (414; 515; 535; 1014) as the second object, To change at least one characteristic of the first object based on the characteristics of the second object and the relationship between the second object and the first object. Electronic device (101).

12. In any one of paragraphs 9 to 11, The relationship between the second object and the first object is Whether the category of the second object matches the category of the first object (matching status) including, Electronic device (101).

13. In any one of paragraphs 9 to 12, The above first object is, A plurality of objects included in one frame among a plurality of frames (315) of the original video (301) are segmented differently and determined according to the size of the portion where the original video (301) is displayed on the display (160). Electronic device (101).

14. In any one of paragraphs 9 to 13, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Providing an interface for the user to confirm whether to include the audio content in a frame containing the first object; Electronic device (101).

15. In any one of paragraphs 9 to 14, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Displaying the second object in one frame among the plurality of frames (315) of the original video (301), Generate video content corresponding to the second object, To output the generated video content by including it in the original video (301). Electronic device (101).

Citation Information

Patent Citations

  • Personalizing a video

    KR101348521B1

  • Apparatus for learning semantic segmentation and method thereof

    KR1020250076935A

  • System and method for automatic video editing using utilization of auto labeling and insertion of design elements

    KR102523438B1

  • Portable blower

    KR102587734B1

  • KR20220122217A