Electronic device for adjusting audio volume of video, operating method thereof, and non-transitory computer-readable storage medium
The electronic device uses a generative AI model to analyze video content and adjust audio volumes based on object size and location, enhancing the audio experience by generating new audio for unmapped objects, addressing the inconsistency in existing devices.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-13
- Publication Date
- 2026-04-23
AI Technical Summary
Existing electronic devices lack the capability to dynamically adjust audio volumes based on the objects and movements within a video, leading to inconsistent audio experiences.
An electronic device equipped with a processor and memory that utilizes a generative AI model to analyze video content, identify objects, and adjust audio volumes accordingly based on object size and location, generating new audio for unmapped objects using user inputs.
Enhances audio experience by dynamically adjusting volumes based on video content, providing a more immersive and tailored audio output for each object within the video.
Smart Images

Figure KR2025016030_23042026_PF_FP_ABST
Abstract
Description
Electronic device for adjusting the audio volume of a video, method of operation thereof, and non-transient computer-readable storage medium
[0001] The present disclosure relates to an electronic device for adjusting the audio volume of a video, a method of operating the same, and a non-transient computer-readable storage medium.
[0002] Driven by the remarkable advancements in information and communication technology and semiconductor technology, the distribution and use of various electronic devices are increasing rapidly. Electronic devices are being developed to allow users to carry them around and communicate. The term "electronic device" may refer to a device that performs specific functions according to an installed program, such as mobile communication terminals, tablet PCs (personal computers), wearable electronic devices, video / audio devices, desktop / laptop computers, or vehicle navigation systems.
[0003] Electronic devices may utilize artificial intelligence (AI) models to provide services for specific purposes. For example, AI models are used in various fields such as content streaming, translation, photo editing, finance, new drug development, law, and the military. At least some of the various AI models for specific services may be implemented as generative AI models. Depending on the implementation, the AI model may operate in a connected form.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] According to one embodiment, the electronic device may include a display, at least one processor, and a memory for storing instructions. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may analyze the video in response to a command to edit the audio of the video displayed on the display. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may identify, based on the analysis of the video, a first audio corresponding to a first object among a plurality of objects displayed on a first screen of the video and a second audio corresponding to a second object among the plurality of objects. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may display a second screen on the display in which a first portion of the first screen is enlarged or cropped based on a first user input regarding the first screen. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may determine at least one of the shape, attribute, or movement of the third object based on analyzing the third object when a third object not corresponding to the first audio and the second audio is identified on the second screen. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may determine whether to generate a third audio corresponding to the third object based on at least one of the shape, attribute, or movement of the third object.According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may generate the third audio corresponding to the third object based on providing a generative AI model with first information regarding at least a portion of the first screen and second information regarding at least a portion of the second screen including the third object, in response to confirming that the electronic device generates the third audio corresponding to the third object. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may adjust the volume of the first audio, the second audio, and the third audio, respectively, based on the area size and location of each of the first object, the second object, and the third object included in the second screen.
[0006] According to one embodiment, the method of operation of the electronic device may include an operation of analyzing the video in response to a command to edit the audio of the video displayed on the display. According to one embodiment, the method of operation of the electronic device may include an operation of identifying, based on the analysis of the video, a first audio corresponding to a first object among a plurality of objects displayed on a first screen of the video and a second audio corresponding to a second object among the plurality of objects. According to one embodiment, the method of operation of the electronic device may include an operation of displaying a second screen on the display in which a first portion of the first screen is enlarged or cropped based on a first user input regarding the first screen. According to one embodiment, the method of operation of the electronic device may include an operation of identifying at least one of the shape, attributes, or movement of the third object based on the analysis of the third object when a third object not corresponding to the first audio and the second audio is identified on the second screen. According to one embodiment, the method of operation of the electronic device may include an operation of determining whether to generate a third audio corresponding to the third object based on at least one of the shape, attribute, or movement of the third object. According to one embodiment, the method of operation of the electronic device may include an operation of generating the third audio corresponding to the third object based on providing a generative AI model with first information regarding at least a portion of the first screen and second information regarding at least a portion of the second screen including the third object, in response to determining to generate the third audio corresponding to the third object.According to one embodiment, the method of operation of the electronic device may include an operation of adjusting the volume of the first audio, the second audio, and the third audio, respectively, based on the area size and position of each of the first object, the second object, and the third object included in the second screen.
[0007] According to one embodiment, in a non-transient storage medium for storing instructions, when the instructions are executed collectively or individually by at least one processor, the electronic device analyzes the video in response to a command to edit the audio of the video displayed on the display, and based on the analysis of the video, identifies a first audio corresponding to a first object among a plurality of objects displayed on a first screen of the video and a second audio corresponding to a second object among the plurality of objects, and based on a first user input for the first screen, displays a second screen on the display in which a first part of the first screen is enlarged or cropped, and if a third object not corresponding to the first audio and the second audio is identified on the second screen, based on the analysis of the third object, identifies at least one of the shape, attribute, or movement of the third object, and based on at least one of the shape, attribute, or movement of the third object, determines whether to generate a third audio corresponding to the third object, and confirms to generate the third audio corresponding to the third object. In response to this, the third audio corresponding to the third object is generated based on providing the first information regarding at least a portion of the first screen and the second information regarding at least a portion of the second screen including the third object to the generative AI model, and the volume of the first audio, the second audio, and the third audio can be adjusted based on the area size and location of each of the first object, the second object, and the third object included in the second screen.
[0008] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0009] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.
[0010] FIG. 2 is a schematic block diagram of an electronic device according to one embodiment.
[0011] FIG. 3 is a flowchart illustrating a method for an electronic device to adjust the volume of audio in a video according to one embodiment.
[0012] FIGS. 4a and FIGS. 4b are drawings for explaining a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0013] FIG. 5 is a diagram illustrating a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0014] FIGS. 6a, 6b, and 6c are drawings for explaining a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0015] FIGS. 7a and 7b are drawings for explaining a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0016] FIG. 8 is a diagram illustrating a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0017] FIGS. 9a, 9b, 9c, 9d, and 9e are drawings for illustrating a method of generating audio for an object that is not mapped to audio, according to one embodiment.
[0018] FIG. 10 is a flowchart illustrating a method for generating audio for an object that is not mapped to audio, according to one embodiment.
[0019] FIGS. 11a and FIGS. 11b are flowcharts for explaining how an electronic device adjusts the volume of audio of a video according to one embodiment.
[0020] FIG. 12 is a diagram illustrating a generative artificial intelligence system according to one embodiment.
[0021] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0022] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.
[0023] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0024] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0025] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0026] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0027] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0028] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0029] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0030] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0031] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0032] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0033] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0034] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0035] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0036] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0037] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0038] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0039] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0041] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0042] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0043] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0044] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0045] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as a CPU (central processing unit), AP (application processor), or DSP (digital signal processor), graphics-dedicated processors such as a GPU (graphic processing unit) or VPU (vision processing unit), or artificial intelligence-dedicated processors such as an NPU (neural processing unit). The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0046] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined rules of operation or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0047] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained by the artificial intelligence model during the learning process is reduced or minimized. The artificial neural network may include a deep neural network (DNN), and examples include, but are not limited to, convolutional neural networks (CNN), deep neural networks (DNN), recurrent neural networks (RNN), restricted Boltzmann machines (RBM), deep belief networks (DBN), bidirectional recurrent deep neural networks (BRDNN), or deep Q-networks.
[0048] FIG. 2 is a schematic block diagram of an electronic device according to one embodiment.
[0049] Referring to FIG. 2, an electronic device (201) according to one embodiment may include a processor (220), memory (230), speaker (240), display (260), and communication circuit (290). For example, the electronic device (201) may be implemented in the same or similar manner as the electronic device (101) of FIG. 1.
[0050] According to one embodiment, the processor (220) can control the overall operation of the electronic device (201). For example, the processor (220) may be implemented identically or similarly to the processor (120) of FIG. 1. According to one embodiment, the processor (220) can control at least one other component (e.g., hardware or software component) of the electronic device (201) connected to the processor (220) by executing software (e.g., program (140) of FIG. 1), and can perform data processing or operations based on instructions. According to one embodiment, the instructions may include instructions composed of machine language that can be processed by the electronic device (201) or the processor (220). For example, the instructions may include instructions corresponding to operation instructions used in the program.
[0051] Meanwhile, although FIG. 2 illustrates that the electronic device (201) includes one processor (220), this is exemplary and the technical concept of the present invention may not be limited thereto. For example, the electronic device (201) may include at least one processor. For example, the processor (220) may be implemented as at least one processor.
[0052] According to one embodiment, the memory (230) (e.g., the memory (130) of FIG. 1) may store at least one instruction (or instruction) that causes at least one operation of the electronic device (201). When executed by the processor (220), the at least one instruction may cause the electronic device (201) to perform the corresponding operation.
[0053] According to one embodiment, the memory (230) can store content. For example, the content may include text, audio, sound, images, and / or videos.
[0054] According to one embodiment, the processor (220) can modify, change, edit, or generate content stored in memory (230) using a generative AI model. For example, the generative AI model may be stored in memory (230). Alternatively, the generative AI model may be stored in an external electronic device (e.g., the server (108) of FIG. 1). For example, the processor (220) may use the generative AI model to generate or obtain content (e.g., audio or sound) based on a prompt. For example, the prompt may include a command (e.g., text) for generating new content (e.g., audio or sound) or for modifying, changing, or editing at least one element or at least one object included in the content (e.g., audio or sound). For example, the prompt may include a command in the form of text that the generative AI model can recognize.
[0055] According to one embodiment, when a generative AI model is stored in an external electronic device (e.g., a server), the processor (220) can transmit a prompt to the external electronic device through a communication circuit (290) (e.g., a communication module (190) of FIG. 1). Additionally, the processor (220) can obtain content based on the prompt from the external electronic device through the communication circuit (290).
[0056] For convenience of explanation, the following description will focus on how the electronic device (201) acquires content (e.g., audio or sound) when a generative AI model is stored in memory (230). However, the technical concept of the present invention is not limited thereto, and the electronic device (201) may acquire content (e.g., audio or sound) through an external electronic device (e.g., server (108) of FIG. 1) according to the same or similar method.
[0057] According to one embodiment, the processor (220) can adjust the volume of the audio (or sound) of the video (e.g., volume up or volume down). For example, the processor (220) can separate the audio (or sound) included in the video into individual audio (or sound) according to a plurality of objects displayed in the video. The processor (220) can map the separated audio to each object displayed in the video. Additionally, the processor (220) can generate new audio for objects among the objects displayed in the video for which audio is not mapped. For example, when generating new audio, a generative AI model may be used.
[0058] According to one embodiment, the processor (220) can individually adjust the volume of each audio mapped to each object based on user input. For example, the processor (220) can determine the area size and location of each object included in the screen based on zooming in, zooming out, or cropping a screen of a video. The processor (220) can adjust the volume of the audio corresponding to each object based on the area size and location of each object in the screen.
[0059] According to one embodiment, the processor (220) can output sound according to an adjusted audio volume through a speaker (240) (e.g., the sound output module (155) of FIG. 1) when playing a video.
[0060] According to one embodiment, the processor (220) can receive video from an external source in real time (e.g., video call or video streaming) through a communication circuit (290). The processor (220) can adjust the audio of the video received in real time. For example, the processor (220) can separate the audio included in the video according to the objects displayed in the video. The processor (220) can map the separated audio to each object displayed in the video. Additionally, the processor (220) can generate new audio for objects among the objects displayed in the video for which the audio is not mapped.
[0061] According to one embodiment, the processor (220) can individually adjust the volume of each audio mapped to each object based on user input. For example, the processor (220) can determine the area size and location of each object included in the screen based on zooming in, zooming out, or cropping a screen of a video. The processor (220) can adjust the volume of the audio corresponding to each object according to the area size and location of each object in the screen. According to one embodiment, the processor (220) can play a video received in real time while outputting sound according to the adjusted audio volume.
[0062] At least some of the operations of the electronic device (201) described below may be performed by the processor (220). However, for the convenience of explanation, the operations below will be described as being performed by the electronic device (201).
[0063] FIG. 3 is a flowchart illustrating a method for an electronic device to adjust the volume of audio in a video according to one embodiment.
[0064] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0065] Referring to FIG. 3, according to one embodiment, in operation 301, an electronic device (e.g., the electronic device (201) of FIG. 2) can analyze a video corresponding to a screen displayed on a display (e.g., the display (260) of FIG. 2) based on confirming a command to edit the audio of a video. The electronic device (201) can confirm a command to edit the audio of a video based on confirming user input (e.g., a tab input) for a specific object (e.g., an object requesting audio editing of a video) displayed on the display (260). Based on analyzing the video, the processor (220) can confirm the type of video, the playback time of the video, audio information contained in the video, scene information in the video, objects appearing in the video, and / or the place and time where the video was filmed. Depending on the implementation, the video corresponding to the screen may already be analyzed before the command to edit the audio of the video is confirmed. At this time, the electronic device (201) can check information about the pre-analyzed video when it confirms a command to edit the audio of the video. According to one embodiment, the electronic device (201) can identify objects appearing in the video based on analyzing the video. For example, objects appearing in the video may be distinguishable (or identifiable) on the screen and may be objects that occupy a relatively large area. For example, objects that occupy a large area may represent objects that appear frequently in the video and are displayed in an area larger than a designated area on at least one screen of the video. For example, the objects may be objects that have attributes related to sound. For example, the objects may include objects (e.g., objects, musical instruments, rivers, seas), living things (e.g., people, animals), and / or backgrounds.
[0066] According to one embodiment, in operation 303, the electronic device (201) can separate and identify a first audio corresponding to a first object and a second audio corresponding to a second object among a plurality of objects displayed on a first screen (e.g., a representative screen) of the video, based on analyzing the video (e.g., a set of sounds generated from various objects or targets). For example, the first screen may represent a frame displayed on the display (260) among a plurality of frames of the video. For example, the electronic device (201) can separate the first audio (e.g., a violin sound) and the second audio (e.g., a flute sound) from the entire audio of the video. Additionally, the electronic device (20!) can identify what kind of sound (e.g., a violin sound) the first audio is and what kind of sound (e.g., a flute sound) the second audio is based on analyzing the separated first audio (e.g., a violin sound) and second audio (e.g., a flute sound). According to one embodiment, the electronic device (201) may map the separated first audio and second audio to corresponding objects. For example, the electronic device (201) may map the first audio to the first object and map the second audio to the second object. For example, the first audio may represent audio corresponding to a sound generated by the first object. The second audio may represent audio corresponding to a sound generated by the second object. For example, when the first object is a violin or a violin player, the first audio may include a sound that the first object can generate (e.g., a violin playing sound). For example, when the second object is a flute or a flute player, the second audio may include a sound that the second object can generate (e.g., a flute playing sound). For example, when the second object is a background that cannot be specifically distinguished, the second audio may represent audio corresponding to sounds other than those generated by a specific object (e.g., the first object).For example, the electronic device (201) can map a violin sound to a violin or violin player when the first object is a violin or violin player and the first audio is a violin sound. For example, the electronic device (201) can map a flute sound to a flute or flute player when the second object is a flute or flute player and the second audio is a flute sound.
[0067] According to one embodiment, in operation 305, the electronic device (201) may display on the display (260) a second screen in which a first portion of the first screen is enlarged / reduced or cropped based on a first user input (e.g., a pinch zoom gesture or a crop (or cut) gesture) for the first screen. For example, a pinch zoom gesture may include a gesture of narrowing or widening the distance between two fingers using said fingers. For example, a crop gesture may include a gesture of selecting a portion of the entire screen with a finger (or input means). For example, a first portion of the first screen may be enlarged or cropped in proportion to the distance at which the first user input is entered. For example, when the first user input is a pinch zoom gesture, the second screen may be a screen in which the first portion of the first screen at which the first user input is entered is enlarged by said distance. For example, when the first user input is a crop gesture, the second screen may be a screen corresponding to the portion remaining on the first screen after being cropped by the first user input. Alternatively, when the first user input is a gesture of clicking or touching an object (or object) that enlarges the screen, the second screen may be a screen in which a first part of the first screen is enlarged in proportion to the number of times the first user input is entered or the touch time. That is, the second screen may represent the screen displayed on the display (260) after the first screen has been enlarged, reduced, or cropped by the first user input.
[0068] According to one embodiment, in operation 307, if a third object that does not correspond to the first audio and the second audio is identified on the second screen, the electronic device (201) can determine at least one of the shape, attributes, or movement of the third object based on analyzing the third object. For example, the electronic device (201) can determine whether the third object is an object or a person. Additionally, when the third object is a person, the electronic device (201) can determine whether it possesses an object. The electronic device (201) can determine whether the third object has attributes related to sound (e.g., whether the third object can generate sound). Additionally, when the third object possesses an object, the electronic device (201) can determine whether the object has attributes related to sound. The electronic device (201) can also determine whether the third object is moving. For example, the electronic device (201) can also check whether sound can be generated by the movement when the third object is moving.
[0069] According to one embodiment, in operation 309, the electronic device (201) may determine whether to generate a third audio corresponding to the third object based on at least one of the shape, attribute, or movement of the third object. For example, if the electronic device (201) determines that the third object has an attribute related to sound, it may decide to generate a third audio to be mapped to the third object. For example, if the electronic device (201) determines that the third object does not have an attribute related to sound, it may decide not to generate a third audio.
[0070] According to one embodiment, in operation 311, the electronic device (201) may generate a third audio corresponding to a third object based on providing a generative AI model with first information regarding at least a portion of a first screen and second information regarding at least a portion of a second screen containing a third object, in response to confirming to generate a third audio corresponding to a third object. For example, the first information may include an image corresponding to at least a portion of the first screen, first context information analyzing the first screen, and / or a prompt based on the first context information. For example, the second information may include an image corresponding to at least a portion containing a third object on the second screen, second context information analyzing the third object (or a portion containing the third object) on the second screen, and / or a prompt based on the second context information. For example, the electronic device (201) may generate a third audio corresponding to a third object based on providing at least a portion of a first screen and at least a portion of a second screen (e.g., an area containing a third object) to a generative AI model. Alternatively, the electronic device (201) may generate a third audio corresponding to a third object based on providing a first context information (or prompt information based on the first context information) of a third object including at least one of the shape, attributes, or movement of the third object and a second context information (or prompt information based on the second context information) of the first screen to a generative AI model.
[0071] According to one embodiment, the electronic device (201) can generate a third audio corresponding to a third object based on providing an image corresponding to at least a portion of a first screen and an image corresponding to at least a portion of a second screen (e.g., an area containing a third object) to a generative AI model. In this case, the operation of analyzing the context of the first screen and the third object may be omitted.
[0072] According to one embodiment, the electronic device (201) can map the generated third audio to a third object. For example, when the third object is an audience, the third audio may include sounds that the audience can make (e.g., applause). For example, the first context information may refer to information indicating the results of an analysis of the third object (e.g., the shape, type, role, attributes, or movement of the third object). The second context information may refer to information indicating the results of an analysis of the first screen (e.g., the shape of objects included in the first screen, the type of said objects, the role of said objects, the place shown on the first screen, the time shown on the first screen, and / or the atmosphere shown on the first screen). For example, the electronic device (201) may acquire or extract the first context information and the second context information using an artificial intelligence model (e.g., LLM).
[0073] According to one embodiment, the electronic device (201) may generate a first prompt based on first context information and generate a second prompt based on second context information. At this time, the electronic device (201) may generate a third audio based on providing the first prompt and the second prompt to a generative AI model. For example, the first prompt may include text representing the first context information. The second prompt may include text representing the second context information. For example, the electronic device (201) may acquire or extract the first prompt and the second prompt using an artificial intelligence model (e.g., LLM).
[0074] According to one embodiment, in operation 313, the electronic device (201) can adjust the volume of each of the first audio, the second audio, and the third audio. For example, the electronic device (201) can adjust the volume of each of the first audio, the second audio, and the third audio based on the area size and position of each of the first object, the second object, and the third object included in the second screen. The electronic device (201) can display information indicating the adjusted volume (e.g., a change in the size of a number or image) on the display (260).
[0075] According to one embodiment, the electronic device (201) can output sound according to an adjusted volume when playing a video. For example, the electronic device (201) can output a first audio, a second audio, and a third audio of an adjusted volume through a speaker (e.g., speaker (240) in FIG. 2).
[0076] Through the method described above, the electronic device (201) can provide an environment for distinguishing and adjusting the audio of a video by object through a visual user interface. Through this, the user of the electronic device (201) can intuitively and easily adjust or edit the volume of the audio distinguished by object of the video.
[0077] FIGS. 4a and FIGS. 4b are drawings for explaining a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0078] Referring to FIG. 4a, according to one embodiment, an electronic device (e.g., the electronic device (201) of FIG. 2) may display a first screen (410) of a video on a display (e.g., the display (260) of FIG. 2). A first object (415) for editing the audio of the video may be displayed on the first screen (410). For example, when user input (e.g., a tab input) for the first object (415) is detected, the electronic device (201) may start an operation to edit the audio of the video.
[0079] According to one embodiment, the electronic device (201) can perform an operation to analyze a video when user input (e.g., tap input) for the first object (415) is confirmed. Based on analyzing the video, the electronic device (201) can identify at least one object included in the first screen (410). For example, the electronic device (201) can identify at least one object that occupies an area larger than a designated area included in the first screen (410). For example, the electronic device (201) can identify objects included in the first screen (410) using an artificial intelligence model (e.g., LLM). For example, the electronic device (201) can identify a first object (421) corresponding to a woman, a second object (422) corresponding to a man, and a background (423) included in the first screen (410). For example, the first object (421) and the second object (422) may be objects having attributes related to sound.
[0080] According to one embodiment, the electronic device (201) can distinguish the audio of the video into a plurality of audios corresponding to identified objects based on analyzing the video, and acquire each distinguished audio. For example, if the audio includes sounds related to an orchestral ensemble, the electronic device (201) can identify a first audio corresponding to the sound of a performance by a first object (421) (e.g., female), a second audio corresponding to the sound of a performance by a second object (e.g., male), and a third audio corresponding to other sounds (e.g., audio excluding the first and second audios from the total audio). The electronic device (201) can map the first audio to the first object (421), map the second audio to the second object (422), and map the third audio to the third object (423).
[0081] According to one embodiment, the electronic device (201) can determine or adjust the volume of each of the first audio, the second audio, and the third audio based on the area size (e.g., the size of the area occupied by the object on the screen) and / or position (e.g., the relative position where the object is located on the screen) of each of the first object (421), the second object (422), and the third object (423) included (or displayed) on the first screen (410) or the second screen (420).
[0082] According to one embodiment, the electronic device (201) can display a second screen (420) on the display (260) with highlights (e.g., visual effects on the outlines of the objects) applied to identified objects (e.g., a first object (421) and a second object (422).
[0083] According to one embodiment, the electronic device (201) may include thumbnail images (426, 427, 428) representing identified objects (421, 422, 423), respectively, on a second screen (420). Each of the thumbnail images (426, 427, 428) may include an image representing the corresponding object. For example, the electronic device (201) may adjust or change the size of the first thumbnail image (426) corresponding to the first object (421), the size of the second thumbnail image (427) corresponding to the second object (422), and the size of the third thumbnail image (428) corresponding to the third object (423) according to the volume of the first audio, the volume of the second audio, and the volume of the third audio. Additionally, the electronic device (201) may display information (e.g., volume value) indicating the volume of audio corresponding to the area around the thumbnail images (426, 427, 428) (e.g., below). For example, the electronic device (201) may adjust or change the information (e.g., volume value) indicating the volume of audio according to the volume of the first audio, the volume of the second audio, and the volume of the third audio. For example, the volume value may represent the volume of each audio as a relative value when the total volume of the audio is assumed to be 100%.
[0084] According to one embodiment, the electronic device (201) may display an object (450) for storing the volume of audios on a second screen (420). The electronic device (201) may store the volume value of the audios set in response to user input regarding the object (450).
[0085] According to one embodiment, the electronic device (201) can detect a first user input (430) for a second screen (420). For example, the first user input (430) may be an input (e.g., pinch zoom input) that causes the second screen (420) (or a first part of the second screen (420)) to be enlarged.
[0086] According to one embodiment, the electronic device (201) can display a third screen (440) in which a first portion of a second screen (420) is enlarged on a display (260) in response to a first user input.
[0087] According to one embodiment, the electronic device (201) can determine or adjust the volume of the first audio, the second audio, and the third audio, respectively, based on the area size and / or location of each of the first object (421), the second object (422), and the third object (423) included (or displayed) on the third screen (440).
[0088] Referring to FIG. 4b (a), according to one embodiment, the electronic device (201) may increase the volume (460) of the first audio mapped to the first object (421) based on an increase in the area size of the first object (421) included (or displayed) in the third screen (440). The electronic device (201) may decrease the volume (470) of the second audio mapped to the second object (422) and the volume (480) of the third audio mapped to the third object (423) based on a decrease in the area size of the second object (422) and the third object (423) included (or displayed) in the third screen (440). For example, the volume (or degree of volume change) of each of the first audio, second audio, and third audio may be proportional to the area size (or degree of area size change) of each of the first object (421), second object (422), and third object (423) included (or displayed) on the third screen (440).
[0089] Referring to (b) of FIG. 4b, according to one embodiment, the electronic device (201) can change the position of the volume of each of the first audio mapped to the first object (421), the second audio mapped to the second object (422), and the third audio mapped to the third object (423) based on the change in the relative positions of the first object (421), the second object (422), and the third object (423) included (or displayed) in the third screen (440) (e.g., a virtual position where the user may perceive that the corresponding sound is output at the corresponding position). For example, if the electronic device (201) confirms that the first object (421) is located in the center part of the third screen (440), it can adjust the position of the first audio corresponding to the first object (421) to a position corresponding to the center part. Likewise, the electronic device (201) can also change the position of the second audio corresponding to the second object (422) and the third audio corresponding to the third object (423).
[0090] According to one embodiment, the electronic device (201) may adjust or change the size of the first thumbnail image (426), the size of the second thumbnail image (427), and the size of the third thumbnail image (428) according to the volume of the first audio, the volume of the second audio, and the volume of the third audio adjusted (or changed) on the third screen (440). For example, the electronic device (201) may increase the size of the first thumbnail image (426) as the volume of the first audio increases. Additionally, the electronic device (201) may decrease the size of the second thumbnail image (426) and the third thumbnail image (427) as the volumes of the second audio and the third audio decrease. Additionally, the electronic device (201) can also change information indicating the volume of audio corresponding to the area around the thumbnail images (426, 427, 428) (e.g., below).
[0091] According to one embodiment, the electronic device (201) can store the volume of each adjusted audio. Subsequently, the electronic device (201) can output audio of the adjusted volume during video playback.
[0092] Meanwhile, the thumbnail image illustrated in FIG. 4a is exemplary, and embodiments of the present invention may apply various types of indicators capable of indicating objects identified on the screen and the volume of said objects.
[0093] FIG. 5 is a diagram illustrating a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0094] Referring to FIG. 5, according to one embodiment, an electronic device (e.g., the electronic device (201) of FIG. 2) may display one screen (510) of a video on a display (e.g., the display (260) of FIG. 2) based on starting an operation to edit the volume of the audio of the video. For example, one screen (510) may be a screen in which a first part of a first screen (e.g., the first screen (410) of FIG. 4a) is enlarged. For example, one screen (510) may be a screen in which a second object (422) corresponding to a male is prominently displayed.
[0095] According to one embodiment, the electronic device (201) may determine or adjust the volume of the first audio mapped to the first object (421), the second audio mapped to the second object (422), and the third audio mapped to the third object (423), respectively, based on the area size and location of each of the first object (421), the second object (422), and the third object (423) included in one screen (510). For example, the volume size of the second audio corresponding to the second object (422) (e.g., 68%) may be the largest. At this time, thumbnail images (426, 427, 428) and volume information (e.g., volume values) representing the volumes of the first audio, the second audio, and the third audio may be displayed. For example, the size of the thumbnail images (426, 427, 428) may be proportional to the volumes of the first audio, the second audio, and the third audio. For example, among the thumbnail images (426, 427, 428), the second thumbnail image (427) corresponding to the second object (422) may have the largest size.
[0096] According to one embodiment, the electronic device (201) may display another screen (520) on the display (260) in which a second portion of the first screen (410) is enlarged, based on a second user input (530) for a first screen (510). For example, the second user input may include a panning input or a panning gesture for moving the portion of the entire first screen displayed on the display (260). For example, the other screen (520) may be a screen in which a second portion of the first screen (410) is enlarged, corresponding to the distance moved by the second user input (530). For example, the other screen (520) may be a screen in which a first object (421) corresponding to a woman is prominently displayed.
[0097] According to one embodiment, the electronic device (201) may determine or adjust the volume of the first audio mapped to the first object (421), the second audio mapped to the second object (422), and the third audio mapped to the third object (423), respectively, based on the area size and location of the first object (421), the second object (422), and the third object (423) included in another screen (520). For example, the volume of the first audio corresponding to the first object (421) (e.g., 73%) may be increased. Additionally, the volume of the second audio and the third audio may be decreased based on the area size of the objects (422, 423) displayed on the other screen (520). At this time, thumbnail images (426, 427, 428) and volume information (e.g., volume values) representing the volumes of the first audio, the second audio, and the third audio may be displayed. For example, among the thumbnail images (426, 427, 428), the size of the first thumbnail image (426) corresponding to the first object (421) may be the largest.
[0098] According to one embodiment, the electronic device (201) can store the volume of each adjusted audio. Subsequently, the electronic device (201) can output audio of the adjusted volume during video playback.
[0099] As described above, the screen displayed on the display (260) can be changed according to various forms of user input, and the electronic device (201) can adjust the volume of each audio of the video based on the area sizes of the objects displayed on the screen displayed on the display (260).
[0100] FIGS. 6a, 6b, and 6c are drawings for explaining a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0101] Referring to FIG. 6a, according to one embodiment, an electronic device (e.g., the electronic device (201) of FIG. 2) can display a first screen (610) of a video on a display (e.g., the display (260) of FIG. 2) based on starting an operation to edit the volume of the audio of the video. The electronic device (201) can identify a first object (421) corresponding to a woman, a second object (422) corresponding to a man, and a third object corresponding to a background displayed on the first screen (610).
[0102] According to one embodiment, the electronic device (201) may display a second screen (620) on the display (260) in which a first portion of the first screen (610) is cropped, based on a first user input (630) for the first screen (610). For example, the first user input (630) may be a gesture (e.g., a tap input or a drag input) for cropping a portion of the first screen (610). For example, the second screen (620) may be a screen in which a significant portion of the first object (421) and the second object (422) is cropped.
[0103] According to one embodiment, the electronic device (201) can determine or adjust the volume of the first audio mapped to the first object (421), the second audio mapped to the second object (422), and the third audio mapped to the third object (423), respectively, based on the area size and location of each of the first object (421), the second object (422), and the third object (423) included in the second screen (620). Referring to FIG. 6b, the electronic device (201) can decrease the volume (662) of the first audio corresponding to the first object (421) and the volume (672) of the second audio corresponding to the second object (422). The electronic device (201) can increase the volume (682) of the third audio corresponding to the third object (423). Additionally, the electronic device (201) can adjust the position of each audio based on the relative positions of the objects (421, 422, 423) included in the second screen (620).
[0104] According to one embodiment, the electronic device (201) may display a third screen (640) on the display (260) in which a second portion of the first screen (610) is cropped, based on a second user input for the second screen (620). For example, the second user input may be an input or gesture (e.g., a tap input or a drag input) for panning the screen displayed on the display (260).
[0105] According to one embodiment, the electronic device (201) can determine or adjust the volume of the first audio mapped to the first object (421), the second audio mapped to the second object (422), and the third audio mapped to the third object (423), respectively, based on the area size and location of each of the first object (421), the second object (422), and the third object (423) included in the third screen (640). Referring to FIG. 6c, the electronic device (201) can increase the volume (663) of the first audio corresponding to the first object (421) and the volume (673) of the second audio corresponding to the second object (422). The electronic device (201) can decrease the volume (683) of the third audio corresponding to the third object (423). Additionally, the electronic device (201) can adjust the position of each audio based on the relative positions of the objects (421, 422, 423) included in the second screen (620).
[0106] According to one embodiment, the electronic device (201) can store the volume of each adjusted audio. Subsequently, the electronic device (201) can output audio of the adjusted volume during video playback.
[0107] As described above, the screen displayed on the display (260) can be changed according to various forms of user input, and the electronic device (201) can adjust the volume of each audio of the video based on the area sizes of the objects displayed on the screen displayed on the display (260).
[0108] FIGS. 7a and 7b are drawings for explaining a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0109] Referring to FIG. 7a, according to one embodiment, an electronic device (e.g., the electronic device (201) of FIG. 2) can receive a video from an external communication circuit (e.g., the communication circuit (290) of FIG. 2) and play the received video in real time. For example, the electronic device (201) can run a video call application or a video conferencing application to play at least one video received from an external communication circuit (e.g., the communication circuit (290) of FIG. 2) in real time.
[0110] According to one embodiment, the electronic device (201) may display a first screen including a first object (711), a second screen including a second object (712), a third screen including a third object (713), and a fourth screen including a fourth object (714) on a display (e.g., the display (260) of FIG. 2). The electronic device (201) may detect the first audio, the second audio, the third audio, and the fourth audio generated from each object.
[0111] According to one embodiment, the electronic device (201) can adjust the first audio, the second audio, the third audio, and the fourth audio by adjusting the size of each screen. For example, the electronic device (201) can change the size of at least one of the first screen, the second screen, the third screen, or the fourth screen in response to a user input or user gesture (e.g., a tap input or a drag input) for changing the size of at least one screen. For example, the electronic device (201) can increase the size of the first screen and reduce the size of the remaining screens. Referring to FIG. 7b, the electronic device (201) can increase the volume (760) of the first audio corresponding to the first object (711) based on the increased size of the first screen. The electronic device (201) can reduce the volume of the second audio (770) corresponding to the second object (712), the volume of the third audio (780) corresponding to the third object (713), and the volume of the fourth audio (790) corresponding to the fourth object (714) based on the size of the reduced second screen, third screen, and fourth screen.
[0112] According to one embodiment, the electronic device (201) can store the volume of each adjusted audio. Subsequently, the electronic device (201) can output audio of the adjusted volume during video playback.
[0113] As described above, the screen displayed on the display (260) can be changed according to various forms of user input, and the electronic device (201) can adjust the volume of each audio of the video based on the area sizes of the objects displayed on the screen displayed on the display (260).
[0114] FIG. 8 is a diagram illustrating a method for an electronic device to adjust the volume of audio of a video according to one embodiment.
[0115] Referring to FIG. 8, according to one embodiment, an electronic device (e.g., the electronic device (201) of FIG. 2) may display a volume bar (or volume control bar) (840) for adjusting the volume of any one of the first audio, the second audio, and the third audio on a screen (520) of a video. For example, when a specified input (830) (e.g., a long press input) for the screen (520) is detected, the electronic device (201) may display a volume bar (840) for adjusting the volume of the audio (e.g., the first audio) corresponding to an object (e.g., the first object (421)) located at the point where the input is detected. When the volume bar (840) is displayed, the electronic device (201) can apply a visual effect to a thumbnail image (e.g., 426) representing an object (e.g., a first object (421)) among the thumbnail images (426, 427, 428) corresponding to the volume bar (840).
[0116] According to one embodiment, the electronic device (201) can adjust the volume of the corresponding audio (e.g., first audio) based on a specified input (850) for the volume bar (840) (e.g., a tap input or a drag input in the up and down direction) after the volume bar (840) is displayed.
[0117] As described above, the volume of each audio in the video can be adjusted according to various forms of user input.
[0118] FIGS. 9a, FIGS. 9b, FIGS. 9c, FIGS. 9d, and FIGS. 9e are drawings for illustrating a method of generating audio for an object that is not mapped to audio, according to one embodiment.
[0119] Referring to FIG. 9a, according to one embodiment, an electronic device (e.g., the electronic device (201) of FIG. 2) can perform an operation to adjust the volume of the audio of a video. The electronic device (201) can display a first screen (910) of the video on a display (e.g., the display (260) of FIG. 2). The electronic device (201) can identify a first object (421), a second object (422), and a third object (423) appearing in the video. Additionally, the electronic device (201) can identify a first audio, a second audio, and a third audio corresponding to the first object (421), the second object (422), and the third object (423), respectively, of the audio of the video. The electronic device (201) can map the first object (421), the second object (422), and the third object (423) to the first audio, the second audio, and the third audio, respectively.
[0120] According to one embodiment, the electronic device (201) can detect a first user input (915) for a first screen (910) including a first object (421), a second object (422), and a third object (423). For example, the first user input (915) may be an input that causes the first screen (910) to be magnified (e.g., pinch zoom input).
[0121] According to one embodiment, the electronic device (201) can display a second screen (920) in which a first portion of a first screen (910) is enlarged on a display (260) in response to a first user input.
[0122] According to one embodiment, the electronic device (201) can identify a fourth object (923) that is not mapped to the first audio, the second audio, and the third audio on the second screen (920). For example, the fourth object (923) may have an area size larger than a specified size. The electronic device (201) can analyze the fourth object (923). For example, the electronic device (201) may perform an operation to analyze the fourth object (923) using an artificial intelligence model.
[0123] According to one embodiment, the electronic device (201) can determine at least one of the shape, type, role, attribute, or movement of the fourth object (923) based on analyzing the fourth object (923). For example, the electronic device (201) can identify the fourth object (923) as an audience based on the result of analyzing the fourth object.
[0124] According to one embodiment, the electronic device (201) may determine or decide whether to generate a fourth audio corresponding to the fourth object (923) based on at least one of the shape, type, role, attribute, or movement of the fourth object (923). For example, the electronic device (201) may not generate the fourth audio if it is determined that the fourth object (923) has no attribute related to sound or is not of a type related to sound. Alternatively, the electronic device (201) may decide to generate a fourth audio related to the type or attribute if it is determined that the fourth object (923) has an attribute related to sound or is of a type related to sound. For example, the electronic device (201) may not generate the fourth audio if it is determined that the fourth object (923) does not exhibit a movement related to sound. The electronic device (201) may decide to generate a fourth audio related to the movement if it is determined that the fourth object (923) exhibits a movement related to sound. For example, the electronic device (201) may decide to (automatically) generate a fourth audio related to the fourth object (923) (e.g., audience) based on the result of analyzing the fourth object. Afterward, the electronic device (201) may perform the operation of generating the fourth audio.
[0125] Referring to FIG. 9b, according to one embodiment, the electronic device (201) may display an indicator (925) on a fourth object (923) that is not mapped to the first audio, the second audio, and the third audio when the fourth object (923) is identified on the second screen (920). For example, the indicator (925) may indicate an object that is not mapped to audio. For example, the electronic device (201) may decide to generate a fourth audio to be mapped to the fourth object (923) when user input (e.g., touch input) to the indicator (925) is identified. Subsequently, the electronic device (201) may perform the operation of generating the fourth audio. Meanwhile, the shape, size, or display position of the indicator (925) in FIG. 9b is exemplary and the technical features of the present invention may not be limited thereto.
[0126] According to one embodiment, the electronic device (201) may decide to generate a fourth audio related to a fourth object (923) (e.g., audience) according to either embodiment of FIG. 9a or FIG. 9b described above.
[0127] According to one embodiment, the electronic device (201) can generate a fourth audio corresponding to a fourth object (924) using a generative AI model (e.g., the generative AI model (990) of FIG. 9d and FIG. 9e).
[0128] According to one embodiment, with reference to FIG. 9d, the electronic device (201) may provide to a generative AI model (990) an image corresponding to at least a portion of the first screen (910) and an image corresponding to at least a portion of the second screen (920) including the fourth object (923), in response to confirming or deciding to generate a fourth audio corresponding to a fourth object (923). At this time, the generative AI model (990) may be an AI model trained to generate audio corresponding to an object included in the image using at least one image. The electronic device (201) may generate a fourth audio corresponding to the fourth object (923) based on providing the images to the generative AI model (990). For example, the fourth audio may include sounds generated by the "audience" (e.g., applause or cheering).
[0129] According to one embodiment, with reference to FIG. 9e, the electronic device (201) may provide the generative AI model (990) with first context information of the fourth object (924) including at least one of the shape, type, role, attribute, or movement of the fourth object (923) and second context information of the first screen (910) (e.g., information about objects (421, 422, 423) of the first screen (910), scene information, mood information, place information, time information) in response to confirming or deciding to generate a fourth audio corresponding to the fourth object (923). Alternatively, the electronic device (201) may provide the generative AI model (990) with prompted information (e.g., first prompt and second prompt) of the first context information and the second context information. In this case, the generative AI model (990) may be an AI model trained to generate audio using context information or prompts based on context information. The electronic device (201) can generate a fourth audio corresponding to a fourth object (923) based on providing first context information and second context information to a generative AI model (990). For example, the fourth audio may include sounds generated by an "audience" (e.g., applause or cheering).
[0130] According to one embodiment, the electronic device (201) can generate a fourth audio related to a fourth object (923) (e.g., audience) according to either the embodiment of FIG. 9d or FIG. 9e described above.
[0131] According to one embodiment, the electronic device (201) may map the fourth audio to the fourth object (923). Additionally, the electronic device (201) may display a fourth thumbnail image (429) representing the fourth object (923) on the second screen (920) instead of the third thumbnail image (428). For example, the electronic device (201) may not identify the audio for the object because the third object (423) is displayed over an area smaller than the specified size.
[0132] According to one embodiment, the electronic device (201) can adjust the volume of the corresponding audio based on the area size and location of each of the objects displayed or included in the second screen (920). Referring to FIG. 9c, the electronic device (201) can decrease the volume of the first audio (960) and the volume of the second audio (970) based on the area size of each of the objects included in the second screen (920). The electronic device (201) can increase the volume of the fourth audio (980) based on the area size of each of the objects included in the second screen (920).
[0133] As described above, the electronic device (201) can generate new audio that is not included in the existing audio and map the generated audio to an object included in the screen. Through this, the electronic device (201) can also provide a method to easily adjust the volume of the generated audio.
[0134] FIG. 10 is a flowchart illustrating a method for generating audio for an object that is not mapped to audio, according to one embodiment.
[0135] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0136] Referring to FIG. 10, according to one embodiment, in operation 1001, an electronic device (e.g., the electronic device (201) of FIG. 2) may obtain first context information representing a third object based on at least one of the shape, type, or attribute of the third object to which audio is not mapped. Additionally, the electronic device (201) may obtain a first prompt based on the first context information. For example, the electronic device (201) may obtain the first context information and / or the first prompt using an artificial intelligence model (e.g., LLM).
[0137] According to one embodiment, in operation 1003, the electronic device (201) may obtain second context information corresponding to the result of scene analysis of a first screen (e.g., representative screen) of a video, and obtain a second prompt based on the second context information. For example, the electronic device (201) may obtain the second context information and / or the second prompt using an artificial intelligence model (e.g., LLM).
[0138] According to one embodiment, in operation 1005, the electronic device (201) may generate a third audio by providing a first prompt and a second prompt to a generative AI model. Depending on the implementation, the electronic device (201) may generate the third audio by providing first context information and second context information to a generative AI model instead of the first prompt and the second prompt. In this case, the operation of generating the first prompt and the second prompt in operations 1001 and 1003 may be omitted.
[0139] According to one embodiment, in operation 1007, the electronic device (201) may adjust the volume of the third audio based on the area (and location) of the third object on the screen. For example, the electronic device (201) may adjust the volume of each audio by considering the area sizes of the first and second objects to which the audio was previously mapped and the area sizes of the third object to which the audio was newly mapped. According to another embodiment, operation 1007 may be omitted. For example, the electronic device (201) may determine the volume of the generated third audio as a default value. In this case, the operation of separately adjusting the volume of the third audio may be omitted.
[0140] As described above, the electronic device (201) can generate new audio that is not included in existing audio and map the generated audio to an object included on the screen. Through this, the user of the electronic device (201) can also easily adjust the volume of the generated audio.
[0141] FIGS. 11a and FIGS. 11b are flowcharts for explaining how an electronic device adjusts the volume of audio of a video according to one embodiment.
[0142] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0143] Referring to FIG. 11a, according to one embodiment, in operation 1101, the electronic device (201) can perform an operation to analyze the video in response to a command to edit the audio of the video (e.g., adjust the volume of the audio). For example, the electronic device (201) can analyze objects appearing in the video. Additionally, the electronic device (201) can analyze the audio included in the video.
[0144] According to one embodiment, in operation 1103, the electronic device (201) can determine whether the video contains audio. For example, if it is determined that the video does not contain audio (No in operation 1103), in operation 1109, the electronic device (201) can generate audio for each of the major objects displayed or appearing in the video. For example, the electronic device (201) can generate audio that can be mapped to the objects using a generative AI model and map each generated audio to each object.
[0145] According to one embodiment, if it is confirmed that audio is included in the video (example of operation 1103), in operation 1105, the electronic device (201) can determine whether the audio is separable. For example, the electronic device (201) can determine whether the audio can be separated into audio that can be mapped to objects appearing in the video.
[0146] According to one embodiment, if it is determined that the audio is not separable (No in 1105), in operation 1109, the electronic device (201) can generate audio for each of the major objects that are displayed or appearing in the video. For example, the electronic device (201) can use a generative AI model to generate audio that can be mapped to the objects and map each generated audio to each object.
[0147] According to one embodiment, if it is confirmed that the audio is separable (e.g., 1105), in operation 1107, the electronic device (201) can map the separated audio to the objects displayed in the video. For example, the electronic device (201) can use an artificial intelligence (AI) model to identify audio originating from a specific object and map the identified audio to the corresponding specific object.
[0148] Referring to FIG. 11b, according to one embodiment, in operation 1111, the electronic device (201) can determine whether the screen displaying the video has been changed based on user input in a mode for editing the audio of the video. For example, the screen being changed may include the screen displayed on the display (e.g., the display (260) of FIG. 2) being enlarged, reduced, or cropped, or the enlarged portion of the screen being changed.
[0149] According to one embodiment, if it is confirmed that the screen displaying the video has not changed (No in operation 1111), in operation 1113, the electronic device (201) can check whether a volume control bar is displayed on the screen. For example, the volume control bar may be displayed in response to a specified input (e.g., a long press input).
[0150] According to one embodiment, if it is determined that the volume control bar is not displayed on the screen (No of operation 1115), the electronic device (201) may terminate the operation of editing the volume of the audio or perform the operation of waiting until additional user input is confirmed.
[0151] According to one embodiment, when it is confirmed that a volume control bar is displayed on the screen (example of operation 1113), in operation 1115, the electronic device (201) can control the volume of the corresponding audio mapped to the corresponding object according to user input to the volume control bar.
[0152] According to one embodiment, when it is confirmed that the screen displaying the video has changed (example of operation 1111), in operation 1117, the electronic device (201) may determine or confirm whether an object not mapped to any of the previously organized audios is identified in the video screen. For example, an object not mapped to the previously organized audio may be an object displayed with an area size larger than a specified size (due to the enlargement of the existing screen).
[0153] According to one embodiment, if no object not mapped to audio is identified in the video screen (No in operation 1117), in operation 1123, the electronic device (201) can adjust the volume of the corresponding audio based on the area size and position of each of the objects on the screen.
[0154] According to one embodiment, if an object not mapped to audio is identified in the video screen (e.g., operation 1117), in operation 1119, the electronic device (201) may determine whether to generate audio for said object. For example, the electronic device (201) may determine whether to generate audio for said object based on the result of analyzing said object.
[0155] According to one embodiment, if it is decided not to generate audio for the object (No of 1119), the electronic device (201) can adjust the volume of the audio based on the area size and position of each of the objects on the screen without generating additional audio.
[0156] According to one embodiment, if it is determined to generate audio for the object (e.g., 1119), in operation 1121, the electronic device (201) can generate audio for the object using a generative AI model and map the generated audio to the object. In operation 1123, the electronic device (201) can adjust the volume of the audio based on the area size and position of each of the objects on the screen.
[0157] As described above, the electronic device (201) can map audio to each object included in the video among existing audio, and if necessary, generate new audio and map it to the object. The electronic device (201) can easily adjust the volume of individual audio corresponding to each object included in the video based on user input regarding the video screen.
[0158] FIG. 12 is a diagram illustrating a generative artificial intelligence system according to one embodiment.
[0159] According to one embodiment, a user query / response interface (1210) may receive user input. User input may be in the form of natural language, images, and / or videos, but is not limited thereto. Additionally, context information may be transmitted along with the user input. Context information may include various additional information at the time of user input. For example, additional information may include information about the application currently being used by the user or the user's location information. Additionally, user input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Furthermore, user input may be in a non-natural language form, such as selecting a menu. The user query / response interface (1210) may output results from a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user query / response interface (1210) may output results from a generative artificial intelligence system to the user. The output can be in the form of natural language or specific content, and it can also be provided in a form such as the action requested by the user.
[0160] The AI framework (1240) can receive input from the user and coordinate and control each component necessary to perform the user's intent based on the user's query.
[0161] User input received from the user query / response interface (1210) can be transmitted to a prompt design component (1241). The prompt design component (1241) can be used to generate prompts suitable for inputting user input into a large language model (LLM) or a large multimodal model (LMM). The prompt design component (1241) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design component (1241) can generate prompts by accessing a knowledge component containing user preference data, a prompt library, and prompt examples based on user input, and can transmit the generated prompts to the LLM or LMM.
[0162] The API / Plug-in management component (1242) can perform the role of communicating with external information when there is a request for additional information when user input is passed as input to a generative model. The API / Plug-in management component (1242) establishes a channel to communicate with the outside of the AI Interface via API, and can enable access to various data sources (e.g., knowledge repository (1220)) through the established channel. Additionally, if the API / Plug-in management component (1242) needs to perform an action that executes the user input as a final step rather than an intermediate result in an application or service, it can request that action from the application / service component (1230) via API. The information obtained from the outside may be used to generate a prompt in the prompt design component (1241) along with the user input, or it may be passed as input to the generative model.
[0163] The output modification component (or refiner component) (1243) can fine-tune the output of the generative model. For example, the output modification component (1243) can verify whether the content generated through LLM and / or LMM is irrelevant, contains biased content, or contains harmful content. Additionally, the output modification component (1243) can determine the extent to which the output matches the desired result and, if necessary, proceed with the additional process. Furthermore, the output modification component (1243) can configure and provide hints to the user to avoid unwanted output.
[0164] A generative AI model (1260) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A generative AI model (1260) may include a model that generates images and / or a model that generates language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational auto encoders), and examples include Diffusion-based generative models that use VAEs and Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also LMMs (large multimodal models) that can recognize various forms of data input, such as text, images, and voice, and generate new data corresponding to them.
[0165] According to one embodiment, the electronic device may include a display, at least one processor, and a memory for storing instructions. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may analyze the video in response to a command to edit the audio of the video displayed on the display. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may identify, based on the analysis of the video, a first audio corresponding to a first object among a plurality of objects displayed on a first screen of the video and a second audio corresponding to a second object among the plurality of objects. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may display a second screen on the display in which a first portion of the first screen is enlarged or cropped based on a first user input regarding the first screen. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may determine at least one of the shape, attribute, or movement of the third object based on analyzing the third object when a third object not corresponding to the first audio and the second audio is identified on the second screen. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may determine whether to generate a third audio corresponding to the third object based on at least one of the shape, attribute, or movement of the third object.According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may generate the third audio corresponding to the third object based on providing a generative AI model with first information regarding at least a portion of the first screen and second information regarding at least a portion of the second screen including the third object, in response to confirming that the electronic device generates the third audio corresponding to the third object. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may adjust the volume of the first audio, the second audio, and the third audio, respectively, based on the area size and location of each of the first object, the second object, and the third object included in the second screen.
[0166] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may obtain a first prompt based on first context information representing the third object based on at least one of the shape or attribute of the third object. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may obtain a second prompt based on second context information corresponding to the analysis result of the scene of the first screen. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may provide the first prompt and the second prompt to the generative AI model to generate the third audio.
[0167] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may display a first thumbnail image representing the first object corresponding to the first audio and a second thumbnail image representing the second object corresponding to the second audio on the first screen based on analyzing the video.
[0168] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may display the first thumbnail image, the second thumbnail image, and a third thumbnail image representing the third object corresponding to the third audio on the second screen based on generating the third audio.
[0169] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may change the size of the first thumbnail image, the size of the second thumbnail image, and the size of the third thumbnail image according to the volume of the first audio, the volume of the second audio, and the volume of the third audio.
[0170] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may display a volume bar on the second screen for adjusting the volume of any one of the first audio, the second audio, and the third audio. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may adjust the volume of any one of the audio based on user input to the volume bar.
[0171] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may apply a visual effect to a thumbnail image representing an object corresponding to the volume bar among the first object, the second object, and the third object when the volume bar for any one of the audio is displayed.
[0172] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may display on the display a third screen in which a second portion of the first screen corresponding to the distance of the second user input is enlarged or cropped, based on the second user input for the second screen. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may adjust the volume of the first audio, the second audio, and the third audio, respectively, based on the area size and location of each of the first object, the second object, and the third object included in the third screen.
[0173] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may adjust the volume of the first audio based on the difference between the area of the first object displayed on the first screen and the area of the first object displayed on the second screen. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may adjust the volume of the second audio based on the difference between the area of the second object displayed on the first screen and the area of the second object displayed on the second screen.
[0174] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, if the electronic device determines that the area of the first object displayed on the second screen has increased compared to the area of the first object displayed on the first screen, the volume of the first audio may be increased based on the increased area. According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, if the electronic device determines that the area of the second object displayed on the second screen has decreased compared to the area of the second object displayed on the first screen, the volume of the second audio may be decreased based on the decreased area.
[0175] According to one embodiment, when the instructions are executed collectively or individually by the at least one processor, the electronic device may adjust the positions of the first audio, the second audio, and the third audio, respectively, based on the relative positions of the first object, the second object, and the third object included in the second screen.
[0176] According to one embodiment, the method of operation of the electronic device may include an operation of analyzing the video in response to a command to edit the audio of the video displayed on the display. According to one embodiment, the method of operation of the electronic device may include an operation of identifying, based on the analysis of the video, a first audio corresponding to a first object among a plurality of objects displayed on a first screen of the video and a second audio corresponding to a second object among the plurality of objects. According to one embodiment, the method of operation of the electronic device may include an operation of displaying a second screen on the display in which a first portion of the first screen is enlarged or cropped based on a first user input regarding the first screen. According to one embodiment, the method of operation of the electronic device may include an operation of identifying at least one of the shape, attributes, or movement of the third object based on the analysis of the third object when a third object not corresponding to the first audio and the second audio is identified on the second screen. According to one embodiment, the method of operation of the electronic device may include an operation of determining whether to generate a third audio corresponding to the third object based on at least one of the shape, attribute, or movement of the third object. According to one embodiment, the method of operation of the electronic device may include an operation of generating the third audio corresponding to the third object based on providing a generative AI model with first information regarding at least a portion of the first screen and second information regarding at least a portion of the second screen including the third object, in response to determining to generate the third audio corresponding to the third object.According to one embodiment, the method of operation of the electronic device may include an operation of adjusting the volume of the first audio, the second audio, and the third audio, respectively, based on the area size and position of each of the first object, the second object, and the third object included in the second screen.
[0177] According to one embodiment, in a non-transient storage medium for storing instructions, when the instructions are executed collectively or individually by at least one processor, the electronic device analyzes the video in response to a command to edit the audio of the video displayed on the display, and based on the analysis of the video, identifies a first audio corresponding to a first object among a plurality of objects displayed on a first screen of the video and a second audio corresponding to a second object among the plurality of objects, and based on a first user input for the first screen, displays a second screen on the display in which a first part of the first screen is enlarged or cropped, and if a third object not corresponding to the first audio and the second audio is identified on the second screen, based on the analysis of the third object, identifies at least one of the shape, attribute, or movement of the third object, and based on at least one of the shape, attribute, or movement of the third object, determines whether to generate a third audio corresponding to the third object, and confirms to generate the third audio corresponding to the third object. In response to this, the third audio corresponding to the third object is generated based on providing the first information regarding at least a portion of the first screen and the second information regarding at least a portion of the second screen including the third object to the generative AI model, and the volume of the first audio, the second audio, and the third audio can be adjusted based on the area size and location of each of the first object, the second object, and the third object included in the second screen.
[0178] The embodiment(s) and the terms used in this document are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or any combination thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0179] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0180] Various embodiments of this document may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium (e.g., internal memory or external memory) readable by a machine (e.g., an electronic device). For example, a processor (e.g., a processor) of the machine (e.g., an electronic device) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.
[0181] According to one embodiment, the method according to various embodiments of the present disclosure may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0182] According to embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device, display; At least one processor; and The electronic device includes a memory for storing instructions, and when the instructions are executed collectively or individually by the at least one processor, the electronic device, In response to a command to edit the audio of a video displayed on the above display, the video is analyzed, and Based on analyzing the above video, a first audio corresponding to a first object among a plurality of objects displayed in the first screen of the above video and a second audio corresponding to a second object among the plurality of objects are identified among the above audio of the above video, and Based on a first user input for the first screen, a second screen in which a first portion of the first screen is enlarged or cropped is displayed on the display, and If a third object that does not correspond to the first audio and the second audio is identified on the second screen, at least one of the shape, attribute, or movement of the third object is identified based on the analysis of the third object, and Based on at least one of the shape, attribute, or movement of the third object, determine whether to generate a third audio corresponding to the third object, and In response to confirming to generate the third audio corresponding to the third object, the third audio corresponding to the third object is generated based on providing the generative AI model with first information regarding at least a portion of the first screen and second information regarding at least a portion of the second screen including the third object. An electronic device for adjusting the volume of the first audio, the second audio, and the third audio, respectively, based on the area size and position of each of the first object, the second object, and the third object included in the second screen.
2. In paragraph 1, when the instructions are executed collectively or individually by the at least one processor, the electronic device, A first prompt is obtained based on first context information representing the third object based on at least one of the shape or attribute of the third object, and A second prompt is obtained based on second context information corresponding to the analysis result of the scene of the first screen, and An electronic device that provides the first prompt and the second prompt to the generative AI model to generate the third audio.
3. In any one of paragraphs 1 to 2, when the instructions are executed collectively or individually by the at least one processor, the electronic device, Based on analyzing the above video, a first thumbnail image representing the first object corresponding to the first audio and a second thumbnail image representing the second object corresponding to the second audio are displayed on the first screen, and An electronic device that, based on generating the third audio, displays the first thumbnail image, the second thumbnail image, and a third thumbnail image representing the third object corresponding to the third audio on the second screen.
4. In any one of paragraphs 1 to 3, when the instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that changes the size of the first thumbnail image, the size of the second thumbnail image, and the size of the third thumbnail image according to the volume of the first audio, the volume of the second audio, and the volume of the third audio.
5. In any one of claims 1 to 4, when the instructions are executed collectively or individually by the at least one processor, the electronic device, A volume bar for adjusting the volume of any one of the first audio, the second audio, and the third audio is displayed on the second screen, and An electronic device that adjusts the volume of any one of the above based on user input to the volume bar.
6. In any one of claims 1 to 5, when the instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that applies a visual effect to a thumbnail image representing an object corresponding to the volume bar among the first object, the second object, and the third object when the volume bar for any one of the above audio is displayed.
7. In any one of claims 1 to 6, when the instructions are executed collectively or individually by the at least one processor, the electronic device, Based on the second user input for the second screen, a third screen is displayed on the display in which the second portion of the first screen corresponding to the distance of the second user input is enlarged or cropped, and An electronic device for adjusting the volume of the first audio, the second audio, and the third audio, respectively, based on the area size and position of each of the first object, the second object, and the third object included in the third screen.
8. In any one of claims 1 to 7, when the instructions are executed collectively or individually by the at least one processor, the electronic device, The volume of the first audio is adjusted based on the difference between the area of the first object displayed on the first screen and the area of the first object displayed on the second screen, and An electronic device that adjusts the volume of the second audio based on the difference between the area of the second object displayed on the first screen and the area of the second object displayed on the second screen.
9. In any one of claims 1 through 8, when the instructions are executed collectively or individually by the at least one processor, the electronic device, If it is confirmed that the area of the first object displayed on the second screen has increased compared to the area of the first object displayed on the first screen, the volume of the first audio is increased based on the increased area, and An electronic device that reduces the volume of the second audio based on the reduced area when it is confirmed that the area of the second object displayed on the second screen is reduced compared to the area of the second object displayed on the first screen.
10. In any one of claims 1 to 9, when the instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device for adjusting the positions of the first audio, the second audio, and the third audio, respectively, based on the relative positions of the first object, the second object, and the third object included in the second screen.
11. In a method of operating an electronic device, An operation to analyze the video in response to a command to edit the audio of the video displayed on the above display; Based on analyzing the above video, an operation to identify a first audio corresponding to a first object among a plurality of objects displayed on the first screen of the above video and a second audio corresponding to a second object among the plurality of objects, among the audio of the above video; An operation of displaying a second screen on the display in which a first portion of the first screen is enlarged or cropped based on a first user input for the first screen; If a third object not corresponding to the first audio and the second audio is identified in the second screen, an action of identifying at least one of the shape, attribute, or movement of the third object based on analyzing the third object; An operation to determine whether to generate a third audio corresponding to the third object based on at least one of the shape, attribute, or movement of the third object; An operation to generate the third audio corresponding to the third object based on providing a generative AI model with first information regarding at least a portion of the first screen and second information regarding at least a portion of the second screen including the third object, in response to confirming to generate the third audio corresponding to the third object; and A method of operation of an electronic device comprising the operation of adjusting the volume of each of the first audio, the second audio, and the third audio based on the area size and position of each of the first object, the second object, and the third object included in the second screen.
12. In Paragraph 11, An operation of obtaining a first prompt based on first context information representing the third object based on at least one of the shape or attribute of the third object; An operation to obtain a second prompt based on second context information corresponding to the analysis result of the scene of the first screen; and A method of operation of an electronic device further comprising the operation of generating the third audio by providing the first prompt and the second prompt to the generative AI model.
13. In any one of paragraphs 11 to 12, Based on analyzing the above video, the operation of displaying a first thumbnail image representing the first object corresponding to the first audio and a second thumbnail image representing the second object corresponding to the second audio on the first screen; and A method of operation of an electronic device further comprising the operation of displaying the first thumbnail image, the second thumbnail image, and a third thumbnail image representing the third object corresponding to the third audio on the second screen based on generating the third audio.
14. In any one of paragraphs 11 through 13, A method of operation of an electronic device further comprising the operation of changing the size of the first thumbnail image, the size of the second thumbnail image, and the size of the third thumbnail image according to the volume of the first audio, the volume of the second audio, and the volume of the third audio.
15. In a non-transient storage medium for storing instructions, When the above instructions are executed collectively or individually by at least one processor, the electronic device, In response to a command to edit the audio of a video displayed on the above display, the video is analyzed, and Based on analyzing the above video, a first audio corresponding to a first object among a plurality of objects displayed in the first screen of the above video and a second audio corresponding to a second object among the plurality of objects are identified among the above audio of the above video, and Based on a first user input for the first screen, a second screen in which a first portion of the first screen is enlarged or cropped is displayed on the display, and If a third object that does not correspond to the first audio and the second audio is identified on the second screen, at least one of the shape, attribute, or movement of the third object is identified based on the analysis of the third object, and Based on at least one of the shape, attribute, or movement of the third object, determine whether to generate a third audio corresponding to the third object, and In response to confirming to generate the third audio corresponding to the third object, the third audio corresponding to the third object is generated based on providing the generative AI model with first information regarding at least a portion of the first screen and second information regarding at least a portion of the second screen including the third object. A storage medium that causes the volume of each of the first audio, the second audio, and the third audio to be adjusted based on the area size and location of each of the first object, the second object, and the third object included in the second screen.
Citation Information
Patent Citations
Playback device, playback method, and program
JP2023075334A
Display device and controlling method thereof
KR1020170002119A
Method and apparatus for automatically segmenting ground-glass opacity and consolidation region using deep learning
KR1020220143185A
Method and apparatus for detecting lesion and determining disease symptom by using fruit disease symptom filter
KR1020240057895A
Buoyancy device
KR102509362B1