Method for conditioning diffusion model for generative expansion of video and computer-readable recording medium thereof
The method uses AI models and a diffusion model to generate expanded background images for music broadcasts, addressing the limitations of traditional video footage and enhancing user immersion in VR and MR devices.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KOREAN BROADCASTING SYST
- Filing Date
- 2024-12-04
- Publication Date
- 2026-04-30
AI Technical Summary
Existing music broadcast videos lack the ability to provide expanded backgrounds or different viewing angles, limiting user immersion and engagement, especially with the rise of VR and MR devices.
A method utilizing AI models and a diffusion model to generate expanded background images by processing vectors from reference and current frames, allowing for 3D viewing experiences.
Enables users to watch music broadcasts in a 360-degree panoramic view with immersive background expansion, enhancing user engagement and providing cinematic effects.
Smart Images

Figure KR2024019740_30042026_PF_FP_ABST
Abstract
Description
Conditioning method of a diffusion model for generative expansion of video and computer-readable recording medium thereof
[0001] The present disclosure relates to a method for controlling an electronic device and a computer-readable recording medium. Specifically, it relates to a method for controlling an electronic device to acquire expanded image data based on AI and a computer-readable recording medium thereof.
[0002] Domestic terrestrial broadcasters provide music programs.
[0003] Previously, when broadcasters provided music programs, they only provided video footage acquired through cameras filming the programs; consequently, there was a limitation in that additional filming was required to provide music videos with expanded backgrounds or those including different backgrounds.
[0004] In addition, when a user watches videos such as music broadcasts, it is necessary to explore ways to expand the viewing angle by zooming in on the stage background included in the music broadcast video.
[0005] Meanwhile, various open-source models such as Stable Diffusion have recently emerged, enabling the generation of images based on inputs such as text or images. Furthermore, the widespread adoption of VR (Virtual Reality) or MR (Mixed Reality) devices is expected.
[0006] Accordingly, there is a need for a method that allows users to watch three-dimensional music broadcasts by wearing VR or MR devices, rather than being limited to watching music broadcast videos filmed with a camera.
[0007] A method for controlling an electronic device according to one embodiment of the present disclosure comprises: a step of acquiring image data; a step of downscaling the image data; a step of acquiring a plurality of frames including the downscaled image data; a step of acquiring first and second vectors using a reference frame among the plurality of frames, and acquiring third and fourth vectors using a previous frame or a current frame among the plurality of frames; and a step of acquiring an adjusted frame by inputting the first to fourth vectors into a diffusion model for acquiring an adjusted frame.
[0008] It may include the step of removing at least one object included in the acquired image data.
[0009] The method includes the step of obtaining the first vector containing characteristic information of the reference frame by inputting it into a first artificial intelligence model based on the reference frame; wherein the characteristic information of the reference frame may include information regarding the shape, pose, or arrangement of an object included in the reference frame.
[0010] The method includes the step of inputting the reference frame into a second artificial intelligence model to obtain the second vector containing information about the contour of the reference frame.
[0011] The above previous frame is input into a third artificial intelligence model to obtain the third vector containing information about the above previous frame or the above current frame.
[0012] The method includes the step of inputting the current frame into a fourth artificial intelligence model to obtain the fourth vector containing information about the current frame.
[0013] At least one of the first to third vectors, the fourth vector, and the current frame are input into a diffusion model to obtain the adjusted frame.
[0014] The method includes the step of obtaining a frame upscaled from the above-mentioned adjusted frame and setting the above-mentioned adjusted frame as the previous frame.
[0015] It includes the step of obtaining an output frame by synthesizing the upscaled frame and the current frame.
[0016] FIG. 1 is a drawing for explaining an electronic device according to one embodiment of the present disclosure, and
[0017] FIG. 2 is a flowchart for illustrating a method for controlling an electronic device according to at least one embodiment of the present disclosure, and
[0018] FIGS. 3 to 6 are drawings for explaining a process of obtaining a vector according to an embodiment of the present disclosure, and
[0019] FIG. 7 is a diagram illustrating the process of obtaining a generated frame by inputting a plurality of vectors and a current frame into a diffusion model according to one embodiment of the present disclosure.
[0020] FIG. 8 is a drawing for explaining the configuration of an electronic device according to one embodiment of the present disclosure.
[0021] The embodiments described herein are subject to various modifications and may have various forms; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals may be used for similar components.
[0022] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is omitted.
[0023] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concept of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to make the present disclosure more faithful and complete and to fully convey the technical concept of the present disclosure to those skilled in the art.
[0024] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of the rights. The singular expression includes the plural expression unless the context clearly indicates otherwise.
[0025] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.
[0026] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.
[0027] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.
[0028] Where it is stated that a certain component (e.g., a first component) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., a second component), it should be understood that the said certain component may be directly connected to the said other component or connected through another component (e.g., a third component).
[0029] On the other hand, when it is stated that a certain component (e.g., a first component) is "directly connected" or "directly coupled" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between said certain component and said other component.
[0030] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware.
[0031] Instead, in some situations, the expression “device configured to do something” may mean that the device is “capable of doing something” together with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing those operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) capable of performing those operations by executing one or more software programs stored in a memory device.
[0032] In the embodiments, a 'module' or 'part' performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented by at least one processor, except for the 'module' or 'part' that needs to be implemented in specific hardware.
[0033] Meanwhile, the various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0034] Hereinafter, various embodiments of the present invention will be described in detail using the attached drawings.
[0035] FIG. 1 is a drawing for explaining a method of controlling an electronic device according to one embodiment of the present disclosure. In FIG. 1, the electronic device (100) is shown as a VR device, but this is merely one embodiment and can be implemented as a wearable device such as an AR (augmented reality) headset, MR (mixed reality) device, smart glasses, or a server.
[0036] The electronic device (100) conventionally displayed only music broadcast video captured by a camera as shown in the upper drawing of FIG. 1. However, according to the present disclosure, as shown in the lower drawing of FIG. 1, the background image of the stage included in the music broadcast video can be enlarged to display a music broadcast video including the enlarged stage.
[0037] As illustrated below in FIG. 1, a user can wear a VR device or an MR device and watch a broadcast video generated from a server device, a VR device, or an MR device as described above. At this time, the user can watch the music broadcast video in a 360-degree panoramic view. When the user wears an MR device and watches the music broadcast, the user can see real-world objects and virtual objects simultaneously, so the electronic device (100) provides the effect of allowing the user to immerse themselves in the music broadcast video while maintaining a sense of reality. In other words, when the user wears a VR device or an MR device and watches the music broadcast video, the user can watch the music broadcast video even if they turn their head 360 degrees, so the user can immerse themselves in the music broadcast video.
[0038] Furthermore, the user can wear a VR device or an MR device and use a joystick to select one of several music broadcast videos. Additionally, the user can use the joystick to zoom in or out of the music broadcast video screen. The user can zoom in on the music broadcast video by pulling the trigger button included on the joystick with their index finger and watch the zoomed-in music broadcast video. Alternatively, the user can zoom out on the music broadcast video by pressing the trigger button included on the joystick and watch the music broadcast video on the zoomed-out screen.
[0039] Watching a music broadcast video while wearing a VR device or an MR device is just one example, and it is obvious that an example in which the user watches a music broadcast video without wearing a VR device or an MR device is also possible.
[0040] Meanwhile, after generating music broadcast video data including an enlarged stage on a server device, the generated music broadcast video data can be transmitted to a VR device, etc. At this time, the VR device, etc. can upscale the received music broadcast video data.
[0041] FIG. 2 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.
[0042] In particular, the electronic device (100) may acquire image data, downscale the acquired image data, and then acquire a plurality of frames including the downscaled image data and an expanded background. Additionally, a first or second vector (330, 430) may be acquired using a reference frame (310) among the plurality of frames, and a third and fourth vector (530, 630) may be acquired using a previous frame (510) or a current frame (610) among the plurality of frames. At this time, the acquired first to fourth vectors (330, 430, 530, 630) may be input into a diffusion model for acquiring an adjusted frame to acquire an adjusted frame. Specific details regarding each step will be described later.
[0043] First, the electronic device (100) can acquire video data (S210). The electronic device (100) may receive video data from an external device or acquire video data stored in memory (120). At this time, the video data may be a music broadcast video, a music video, or a live performance video.
[0044] At this time, the electronic device (100) can remove at least one object included in the acquired image data. Here, the object included in the image data may be a person or a stage device, etc.
[0045] In one embodiment, the electronic device (100) can manually remove objects using video editing software. Specifically, the video editing software can recognize and remove objects included in the frame.
[0046] In one or more embodiments, the electronic device (100) can use a video restoration tool to remove objects included in the acquired image data, automatically remove objects in specific parts of the image, and automatically fill them with the surrounding background.
[0047] Afterward, the electronic device (100) can obtain a plurality of frames with the image data downscaled and the background expanded (S220). At this time, the plurality of frames can be determined based on at least one of a preset time, a scene unit, and a preset frame unit.
[0048] The electronic device (100) can obtain first and second vectors (330, 430) using a reference frame (310) among a plurality of frames, and obtain third and fourth vectors (530, 630) using a previous frame (510) or a current frame (610) among a plurality of frames (S230). This will be described later with reference to FIGS. 3 to 6.
[0049] Meanwhile, for the convenience of explanation, the conditioning vector obtained through the first artificial intelligence model (320) is defined as the first vector (330), the conditioning vector obtained through the second artificial intelligence model (420) is defined as the second vector (430), the conditioning vector obtained through the third artificial intelligence model (520) is defined as the third vector (530), and the conditioning vector obtained through the fourth artificial intelligence model (620) is defined as the fourth vector (630) for further description. At this time, a conditioning vector refers to a vector that provides additional information to control the output of an artificial intelligence model.
[0050] First, the electronic device (100) may set any frame among the plurality of frames obtained by the method described above as a reference frame (310). At this time, the reference frame (310) may be the clearest frame among the plurality of frames, and may be a frame containing important elements (e.g., the angle or lighting of an object included in the frame).
[0051] The electronic device (100) can input the first artificial intelligence model (320) based on the reference frame (310) to obtain a first vector (330) containing characteristic information of the reference frame (310), and the characteristic information of the reference frame (310) may include information about the shape, pose, or arrangement of an object included in the reference frame (310). This will be described later with reference to FIG. 3.
[0052] The first artificial intelligence model (320) can be an artificial intelligence model trained to generate a first vector (330) based on reference information using a ControlNet structure.
[0053] ControlNet is a deep learning-based control structure that can finely adjust generated frames. Specifically, ControlNet can finely adjust detailed elements while maintaining the overall structure of objects included in the frames.
[0054] At this time, ControlNet may include a function to control the generation of a frame similar in form to the reference frame (310) or the input frame.
[0055] As illustrated in FIG. 3, when a reference frame (310) is input into a first artificial intelligence model (320), a first vector (330) can be obtained. At this time, the first vector (330) may include information about a reference image (Reference information). For example, the first vector (330) may include information about detailed information such as the style, shape, and texture of the reference frame (310). At this time, the style of the reference frame (310) may include information about the color palette or pattern of the reference frame (310), the shape may include information about the position or size of the reference frame (310), and the texture may include information about the texture of the reference frame (310).
[0056] That is, the first artificial intelligence model (320) can be a model for giving a consistent concept while maintaining a shape similar to the reference frame (310) during the process of generating frames.
[0057] For example, the electronic device (100) can set one of a plurality of frames including a downscaled image of a music broadcast video and an expanded background as a reference frame (310), and input the reference frame (310) into a first artificial intelligence model (320) to obtain a first vector (330).
[0058] When a reference frame (310) is input into the first artificial intelligence model (320), the first artificial intelligence model (320) may output a first vector (330) containing information about the concert video. At this time, the first vector (330) may include information related to the structure, arrangement, or size of the music broadcast stage. Additionally, the first vector (330) may include information reflecting the lighting effects of the video included in the music broadcast stage, the style of the stage design, or information regarding the background texture and design elements of the stage.
[0059] The electronic device (100) can obtain a frame in which the style or texture of the frame is adjusted to be similar to the reference frame (310) based on the acquired first vector (330).
[0060] The electronic device (100) can input the reference frame (310) into the second artificial intelligence model (420) to obtain a second vector (430) containing information about the contour of the reference frame (310). This will be described later with reference to FIG. 4.
[0061] At this time, the second artificial intelligence model (420) may be an artificial intelligence model that generates a second vector (430) based on a lineart by utilizing a ControlNet structure. A lineart refers to a frame in which the shape of an object consists only of lines, emphasizing the outline or main boundary of the frame.
[0062] The second artificial intelligence model (420) can utilize Lineart information included in the reference frame (310) based on the ControlNet structure. The second artificial intelligence model (420) is a model designed to solve the problem of the shape appearing different for each frame during continuous image generation. That is, the second artificial intelligence model (420) is a model designed to maintain a similar surrounding composition.
[0063] As described above, in order to utilize the Lineart information included in the reference frame (310), the second artificial intelligence model (420) can use various methods.
[0064] In one embodiment, the Canny (Canny Edge Detection) algorithm is an algorithm for finding boundary lines (edges) included in a reference frame (310). The Canny algorithm can obtain information about the main shapes and boundaries included in the reference frame (310) by highlighting the main structures and contours of the frame.
[0065] In one or more embodiments, the Depth Map may provide information about the depth (distance) of each pixel included in the reference frame (310). By using the Depth information, information can be obtained to distinguish between the background and the object and to adjust the perspective.
[0066] However, this is only one example, and the second artificial intelligence model (420) may use the Normal method, the Multi-Line Segment Detection (MLSD) method, or the Softedge method to obtain information about the line art or outline included in the frame.
[0067] When structural information or contours included in the reference frame (310) are extracted, the second artificial intelligence model (420), as shown in FIG. 4, can obtain a second vector (430) containing information about the Lineart.
[0068] For example, according to the present disclosure, when one of a plurality of frames including a downscaled image of a music broadcast video and an expanded background is set as a reference frame (310) and input into a second artificial intelligence model (420), the second artificial intelligence model (420) can obtain a second vector (430). At this time, the second vector (430) may include information regarding the structure of the music broadcast stage in the frame, equipment included in the music broadcast, the outline of the stage equipment, the boundaries and outlines of the singer, and the arrangement of stage equipment (e.g., lighting, speakers, stage equipment).
[0069] The electronic device (100) can be adjusted to accurately represent the outline of the frame using the acquired second vector (430).
[0070] The electronic device (100) may include the step of inputting a previous frame (510) into a third artificial intelligence model (520) to obtain a third vector (530) containing information about the previous frame (510) or the current frame (610). This will be described later with reference to FIG. 5.
[0071] At this time, the third artificial intelligence model (520) can be trained to generate a vector that induces the current frame (610) to maintain temporal continuity with the previous frame (510) when the previous frame (510) is input. The third artificial intelligence model (520) (e.g., TemporalNet model) can be adjusted so that the previously input frame and the subsequently input frame are connected naturally in time.
[0072] As illustrated in FIG. 5, when the current frame (610) or the previous frame (510) is input into the third artificial intelligence model (520), the third artificial intelligence model (520) can obtain a third vector (530) containing information about the temporal pattern, the movement and location of an object (e.g., a person or stage device) or context information.
[0073] Among the information included in the third vector (530), the temporal pattern may include the temporal pattern between frames and information about the movement of the object included in the frame. By using the information about the temporal pattern, the current position of the object can be estimated by calculating the distance the object has moved in consecutive frames. In addition, the amount of movement of the object can be determined through the pixel movement vector in consecutive frames.
[0074] Among the information included in the third vector (530), context information may include information about changes in the frame's background or lighting.
[0075] For example, among multiple frames including a downscaled video of a music broadcast video and an expanded background, the current frame (610) or the previous frame (510) can be input into the third artificial intelligence model (520) to obtain a third vector (530) containing information for maintaining the temporal consistency of the music broadcast video.
[0076] At this time, the current frame (610) or previous frame (510) input to the third artificial intelligence model (520) may include information about visual elements such as lighting included on the stage and the movement of the singer. Additionally, the current frame (610) and previous frame (510) may include information about the temporal continuity of objects included in the frame based on a camera moving over time, the dance of the singer performing on the stage, the movement of lighting, and changes in the intensity of the lighting.
[0077] Alternatively, the third vector (530) can adjust the color of the light included in the next frame to blue if the color of the light included in the previous frame (510) is blue and the color of the light included in the current frame (610) is yellow.
[0078] The electronic device (100) can input the current frame (610) into the fourth artificial intelligence model (620) to obtain a fourth vector (630) containing information or context information about the current frame (610). This will be described later with reference to FIG. 6.
[0079] At this time, the fourth artificial intelligence model (620) can be implemented as a model used to extract features inherent in the currently input frame for the purpose of eliminating temporal inconsistency.
[0080] As illustrated in FIG. 6, when the current frame (610) is input into the fourth artificial intelligence model (620), the fourth artificial intelligence model (620) can obtain a fourth vector (630). Since the fourth vector (630) may contain information necessary for performing the frame restoration operation, it enables adjustment so that the frames are connected naturally in time.
[0081] For example, a frame containing stage lighting can be input to the fourth artificial intelligence model (620), and the color of the stage lighting included in the frame may be red. At this time, the fourth artificial intelligence model (620) can generate a fourth vector (630) containing information that the color of the edge area of the stage is red. That is, the fourth vector (630) may include context information included in the current frame.
[0082] Returning to FIG. 2, the electronic device (100) can obtain an adjusted frame using the first to fourth vectors (330, 430, 530, 630) obtained according to the method described above. Specifically, the electronic device (100) can obtain an adjusted frame by inputting the first to fourth vectors (330, 430, 530, 630) and the current frame (610) into a Diffusion Model (S240). This will be described later with reference to FIG. 7.
[0083] When the electronic device (100) inputs the first to fourth vectors (330, 430, 530, 630) into a diffusion model, it can generate a single conditioning vector (not shown) in which the first to fourth vectors (330, 430, 530, 630) are merged. The merged conditioning vector may include consistent characteristic information for the entire image of the first vector (330), information that helps to properly reflect the structure and shape of the line art of the second vector (430), information about the temporal continuity of the third vector (530), and information about the currently input frame of the fourth vector (630).
[0084] Meanwhile, the electronic device (100) will be described later for the case where it generates a single conditioning vector by merging the first to fourth vectors (330, 430, 530, 630), but this is only one example, and the electronic device (100) can generate a single conditioning vector based on one of the first to third vectors (330, 430, 530) and the fourth vector (630).
[0085] When a diffusion model generates a frame using a merged conditioning vector, it can adjust the frame using the information contained in the merged conditioning vector. Specifically, the diffusion model can acquire an adjusted frame by filling the current input frame with noise and then removing the noise step by step using the information contained in the conditioning vector.
[0086] The process of the electronic device (100) acquiring an adjusted frame while removing noise will be described in detail below.
[0087] First, the electronic device (100) can generate a frame in which the current frame (610) is filled with random noise. At this time, the electronic device (100) can remove the noise contained in the current frame (610) by utilizing information contained in the merged conditioning vector. The noise contained in the current frame (610) can be adjusted based on information contained in the merged conditioning vector, for example, information about style, information about boundaries, direction information about the movement of objects, or information about damaged areas, to obtain a frame adjusted.
[0088] The electronic device (100) can obtain a frame that has been upscaled from an adjusted frame and set the adjusted frame as a previous frame (510).
[0089] In this context, upscaling refers to the process of increasing the resolution of an image or video to add larger size and detail.
[0090] In one embodiment, the electronic device (100) can perform WebUI Upscaling during generative upscaling. Generative upscaling (WebUI Upscaling) is a function that increases the resolution of an image or frame through a web UI.
[0091] In one or more embodiments, the electronic device (100) can perform ControlNet Tile and Ultimate SD Upscale Script-based upscaling. The resolution of the frame can be increased through ControlNet Tile and Ultimate SD Upscale Script-based upscaling. ControlNet Tile-based upscaling is a method of processing frames by dividing them into tiles, and the image can be upscaled by performing individual processing on each tile. Ultimate SD Upscale Script is an upscaling method in a Stable Diffusion model. In this case, the upscaling method in a Stable Diffusion model is to obtain a frame with increased resolution by dividing the frame into tiles of a small resolution (e.g., 512*512), applying the Stable Diffusion method to each tile, and then recombining them.
[0092] However, this is only one embodiment, and the electronic device (100) may use a commercial upscaling method (e.g., Photoshop Super-resolution method, Topaz Video Enhance method or Davinci Resolve method) for upscaling the adjusted frame.
[0093] The electronic device (100) can efficiently process calculations in the diffusion model and enhance the details of the frame by inputting the merged conditioning vector and the current frame (610) into the diffusion model and upscaling the acquired adjusted frame.
[0094] Meanwhile, the electronic device (100) has been described in detail as having a step of upscaling based on the merged conditioning vector and the current frame (610), but this is only one embodiment and it is obvious that the upscaling step can be omitted.
[0095] The electronic device (100) can obtain an output frame (710) by synthesizing an upscaled frame and a current frame (610).
[0096] Specifically, the electronic device (100) inputs the conditioning vector, in which the first to fourth vectors (330, 430, 530, 630) are merged, and the current frame (610) into a diffusion model to upscale the obtained frame, and synthesizes the upscaled frame and the current frame (610) to obtain an output frame (710) as shown in FIG. 7. The electronic device (100) can output the output frame (710).
[0097] Meanwhile, the process of the electronic device (100) synthesizing the upscaled frame and the current frame (610) to obtain an output frame (710) will be described later, but this is merely one embodiment, and it is obvious that the electronic device (100) can omit the upscaled step and output the current frame (610).
[0098] The electronic device (100) can acquire a plurality of frames in which the background image included in the video frame is expanded according to the method described above, and output them. Accordingly, the user can immerse themselves in the video by watching the video with the expanded background image. In addition, cinematic effects can be provided by displaying the video with the expanded background image. That is, when the background is expanded, the video fills all parts of the screen, so the user can watch with a wider field of view when watching the video.
[0099] FIG. 8 is a drawing for explaining an electronic device (100) according to one embodiment of the present disclosure. The configuration shown in FIG. 8 is merely an example of various embodiments, and some configurations may be omitted and new configurations may be added.
[0100] As illustrated in FIG. 8, the electronic device (100) may include a communication interface (110), a memory (120), a display (130), and a processor (140). The configuration illustrated in FIG. 8 is merely one embodiment, and it is understood that some components may be omitted or added depending on the configuration of the electronic device (100).
[0101] The communication interface (110) is configured to communicate with various types of external devices according to various types of communication methods. In particular, the electronic device (100) can receive video data from an external device through the communication interface (110).
[0102] A wireless communication module may be a module that communicates wirelessly with an external device. For example, the wireless communication module may include at least one module among a Wi-Fi module, a Bluetooth module, an infrared communication module, an Ultra Wide-Band (UWB) module, or other communication modules.
[0103] Wi-Fi modules and Bluetooth modules can perform communication via Wi-Fi and Bluetooth methods, respectively. When using a Wi-Fi module or a Bluetooth module, various connection information, such as the SSID (service set identifier) and session key, is transmitted and received first; after establishing a communication connection using this information, various types of information can be transmitted and received.
[0104] The infrared communication module performs communication according to infrared communication (IrDA, Infrared Data Association) technology, which uses infrared rays located between visible light and millimeter waves to wirelessly transmit data over short distances.
[0105] Other communication modules may include at least one communication chip that performs communication according to various wireless communication standards such as Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), and 5G (5th Generation), in addition to the communication method described above.
[0106] A wired communication module may be a module that communicates with an external device via a wire. For example, a wired communication module may include at least one of a Local Area Network (LAN) module, an Ethernet module, a pair cable, a coaxial cable, or a fiber optic cable.
[0107] The memory (120) can store an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and instructions or data related to the components of the electronic device (100).
[0108] In particular, the memory (120) can store image data and multiple artificial intelligence models.
[0109] Memory (120) can be implemented in various forms such as volatile memory (e.g., DRAM (dynamic RAM), SRAM (static RAM), or SDRAM (synchronous dynamic RAM), non-volatile memory (e.g., OTPROM (one time programmable ROM), PROM (programmable ROM), EPROM (erasable and programmable ROM), EEPROM (electrically erasable and programmable ROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD).
[0110] The display (130) can display various information. Specifically, the display (130) can display an image in which the background image is expanded under the control of the processor (140).
[0111] The display (130) can be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, and a PDP (Plasma Display Panel). The display may also include a driving circuit, a backlight unit, etc., which can be implemented in forms such as an a-si TFT (amorphous silicon thin film transistor), an LTPS (low temperature poly silicon) TFT, and an OTFT (organic TFT). The display can be implemented as a touch screen combined with a touch sensor, a flexible display, a 3D display, a three-dimensional display, etc. According to various embodiments of the present disclosure, the display (130) may include not only a display panel that outputs an image, but also a bezel that houses the display panel.
[0112] The processor (140) may include one or more processors. Specifically, one or more processors may include one or more of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), MIC (Many Integrated Core), DSP (Digital Signal Processor), NPU (Neural Processing Unit), hardware accelerator, or machine learning accelerator. The processor (150) may control one or any combination of other components of the electronic device and may perform operations or data processing related to communication. One or more processors may execute one or more programs or instructions stored in memory. For example, one or more processors may perform a method according to one embodiment of the present disclosure by executing one or more instructions stored in memory.
[0113] One or more processors may be implemented as a single-core processor comprising one core, or as one or more multicore processors comprising multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory such as cache memory or on-chip memory, and a common cache shared by multiple cores may be included in the multicore processor. Additionally, each of the multiple cores included in the multicore processor (or some of the multiple cores) may independently read and execute program instructions for implementing a method according to one embodiment of the present disclosure, or all (or some) of the multiple cores may be linked together to read and execute program instructions for implementing a method according to one embodiment of the present disclosure.
[0114] In particular, the processor (140) can acquire image data, downscale the image data, and acquire multiple frames included in the image data. At this time, the first and second vectors can be acquired using a reference frame among the acquired multiple frames, and the third and fourth vectors (530, 630) can be acquired using the previous frame and the current frame among the multiple frames. Additionally, the first to fourth vectors can be input into a diffusion model for acquiring an adjusted frame to acquire an adjusted frame.
[0115] According to the present disclosure, when a user watches a music broadcast video, a music video, or a live concert video, the user can adjust the background image of the video to watch an enlarged video, or modify the background image of the video to watch a video of a different concept.
[0116] Additionally, methods according to various embodiments of the present disclosure may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user (20) devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a service provider's server, a manufacturer's server, an application store's server, or a relay server.
[0117] A method according to various embodiments of the present disclosure may be implemented as software comprising instructions stored on a machine-readable storage medium (e.g., a computer). The machine may include a server device or an electronic device according to the disclosed embodiments, which is a device capable of calling instructions stored from the storage medium and operating according to the called instructions.
[0118] Meanwhile, a device-readable storage medium may be provided in the form of a non-transitory readable recording medium. Here, 'non-transitory readable recording medium' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.
[0119] When the above instruction is executed by a processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or an interpreter.
[0120] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
Claims
1. In a method for controlling an electronic device, Step of acquiring image data; A step of downscaling the above image data; A step of acquiring a plurality of frames including the downscaled image; A step of obtaining a first and second vector using a reference frame among the plurality of frames, and obtaining a third and fourth vector using a previous frame or a current frame among the plurality of frames; and A control method comprising the step of acquiring an adjusted frame by inputting at least one of the first to third vectors and the fourth vector into a diffusion model for acquiring an adjusted frame.
2. In Paragraph 1, A control method comprising the step of removing at least one object included in the acquired image data.
3. In Paragraph 1, The method includes the step of inputting the reference frame into a first artificial intelligence model based on the reference frame to obtain the first vector containing characteristic information of the reference frame; The characteristic information of the above reference frame is, A control method comprising information on the shape, pose, or placement of an object included in the above reference frame.
4. In Paragraph 1, A control method comprising the step of inputting the above reference frame into a second artificial intelligence model to obtain the second vector containing information about the contour of the above reference frame.
5. In Paragraph 1, A control method comprising the step of inputting the aforementioned previous frame into a third artificial intelligence model to obtain the aforementioned third vector containing information about the aforementioned previous frame or the aforementioned current frame.
6. In Paragraph 1, A control method comprising the step of inputting the current frame into a fourth artificial intelligence model to obtain the fourth vector containing information about the current frame.
7. In Paragraph 1, A control method comprising the step of setting the above-mentioned adjusted frame to the above-mentioned previous frame.
8. In Paragraph 1, A control method comprising the step of upscaling the adjusted frame and synthesizing the current frame to obtain an output frame.
9. A computer-readable recording medium comprising a method for controlling an electronic device, The control method of the above electronic device is, Step of acquiring image data; A step of downscaling the above image data; A step of acquiring a plurality of frames including the downscaled image; A step of obtaining a first and second vector using a reference frame among the plurality of frames, and obtaining a third and fourth vector using a previous frame or a current frame among the plurality of frames; and A computer-readable recording medium comprising the step of acquiring an adjusted frame by inputting at least one of the first to third vectors and the fourth vector into a diffusion model for acquiring an adjusted frame.