System and method for automatically generating view-expanding content for multi-channel synchronization based on artificial intelligence

KR103025533B1Active Publication Date: 2026-09-29VRUNCH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
KR1020250164295
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-09-29
Estimated Expiration
2045-11-04

Smart Images

  • Figure 112025123014103-PAT00002_ABST
    Figure 112025123014103-PAT00002_ABST
Patent Text Reader

Abstract

According to an embodiment of the present application, a method for generating content that expands the field of view is provided. The method may include: a step of obtaining editing reference information corresponding to an input video sequence; a step of dividing the input video sequence into a plurality of cut units based on the editing reference information; a step of selecting a keyframe from each of the divided cuts; a step of generating an anchor still, which is an expanded still of the outer region of the keyframe, based on each of the keyframes; a step of generating an expanded video sequence for the entire frame of the cut to which each of the keyframes belongs, using the anchor still as a constraint; and a step of combining the expanded video sequences generated in the cut units to generate a final expanded video.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present application relates to a system and method for automatically generating field-of-view expansion content for AI-based multi-channel synchronization, and more specifically, to a system and method for automatically generating field-of-view expansion content suitable for immersive display environments such as multi-channel theaters, media art exhibitions, XR experience centers, and immersive advertisements, and providing video output synchronized between channels. Background Technology

[0002] Recently, as illustrated in Fig. 1, there has been an increasing trend of multi-screen theaters or immersive exhibition spaces that provide a high level of immersion to the audience by extending the field of view beyond the central screen to the left and right walls and even to the ceiling. Extended video content used in such environments has technical requirements to naturally generate an area outside the field of view without any discrepancy with the original video when converting existing video or producing new video, and to be synchronized among multiple channels.

[0003] Conventional generative imaging technologies tend to rely on a single text prompt or a single reference image input, or focus on post-processing the quality of already generated images. These approaches have limitations in that it is difficult to finely control scene elements by simultaneously reflecting complex user conditions such as text, reference images, masks, and motion hints, and they fail to guarantee consistent reproducibility of results under identical conditions.

[0004] In addition, many conventional technologies generate or process each frame of a video sequence independently. This leads to the accumulation of issues such as flicker between frames, color drift where object colors are inconsistent or change, or positional instability where specific objects shake slightly from frame to frame. This degradation of temporal consistency can increase visual fatigue for the audience and severely impair immersion in multi-screen environments.

[0005] The production of content for multi-screen projection involves the process of placing and compositing various 3D objects, such as 3D special effects (VFX), lighting, and particle effects, into the video. Traditionally, this work relies heavily on manual labor by artists. This causes quality inconsistencies during repetitive tasks and frequently leads to rework and re-rendering, increasing overall production lead time and costs. In particular, operational stability is compromised when processing large volumes of long-form content or numerous derivative content.

[0006] Ultimately, problems also arise during the process of delivering multi-channel (e.g., left / center / right) video to projection devices. Video from each channel may be output in non-standard formats, such as different resolutions, codecs, and frame rates, and synchronization to match playback times between channels is often performed manually. This can lead to serious delivery risks, including reduced compatibility with projection devices and playback failures caused by synchronization errors between channels.

[0007] Accordingly, there is a demand for technology capable of providing high-quality images by automating the entire process of generating expanded images while precisely controlling complex input conditions. The problem to be solved

[0008] The present application aims to provide a system and method for automatically generating content with expanded field of view for artificial intelligence-based multi-channel synchronization. means of solving the problem

[0009] According to an embodiment of the present application, a method for generating content that expands the field of view is provided. The method may include: a step of obtaining editing reference information corresponding to an input video sequence; a step of dividing the input video sequence into a plurality of cut units based on the editing reference information; a step of selecting a keyframe from each of the divided cuts; a step of generating an anchor still, which is an expanded still of the outer region of the keyframe, based on each of the keyframes; a step of generating an expanded video sequence for the entire frame of the cut to which each of the keyframes belongs, using the anchor still as a constraint; and a step of combining the expanded video sequences generated in the cut units to generate a final expanded video.

[0010] Additionally, the method further includes the step of using a multi-conditioning node graph to separate multiple input conditions for image generation into node units and generating composite condition information through a weighted merging rule between the nodes, and at least one of the step of generating the anchor still and the step of generating the extended image sequence can be performed based on the composite condition information.

[0011] Additionally, the above input conditions may include at least one of text, a reference image, a mask, and a motion hint.

[0012] Additionally, the step of generating the expanded image sequence may include: providing a preview of the anchor still or candidate expanded image sequence through a lightweight preview path of low resolution or low computational amount; and generating a final expanded image sequence through a refinement path of high resolution or high computational amount based on the approval of the preview.

[0013] In addition, the method may further include a step of performing temporal consistency correction on the generated extended image sequence.

[0014] Additionally, the above-mentioned time consistency correction can be performed using at least one of sliding window-based interpolation or re-synthesis, or an inter-frame feedback path.

[0015] Additionally, the step of selecting the keyframe can be performed by selecting the first frame of the cut as the keyframe, but if the quality of the first frame of the cut is below a preset standard, replacing it with the median frame of the cut as the keyframe.

[0016] Additionally, the step of generating the extended image sequence may include embedding a global style token representing at least one of the global look, tone, and lighting characteristics of the input image sequence to maintain tone consistency among the plurality of extended image sequences.

[0017] Additionally, the method includes the step of placing 3D special effect objects in each of the expanded video sequences generated in cut units through scenario analysis on a 3D graphics engine based on rules; and the step of outputting a multi-channel high-resolution image sequence including the placed 3D special effect objects in conjunction with a time-series automation module of the 3D graphics engine, and the step of generating the final expanded video can be performed by combining the multi-channel high-resolution image sequences generated in cut units.

[0018] Additionally, the step of combining the extended image sequence may include a step of performing correction to mitigate discontinuities between cuts at the boundaries of adjacent cuts.

[0019] In addition, based on the above-mentioned final expanded video, the method may further include the step of generating a master package for compatibility with a multi-channel screening system.

[0020] A computer program is provided according to an embodiment of the present application. The program may be stored on a recording medium to execute a method according to an embodiment of the present application.

[0021] In an embodiment of the present application, a system for generating content that expands the field of view is provided according to an embodiment of the present application. The device comprises at least one processor; and a memory that stores a program executable by the processor. The processor, by executing the program, obtains editing reference information corresponding to an input image sequence, divides the input image sequence into a plurality of cut units based on the editing reference information, selects a keyframe from each of the divided cuts, generates an anchor still which is an expanded still of the outer area of ​​the keyframe based on each of the keyframes, generates an expanded image sequence for the entire frame of the cut to which each of the keyframes belongs using the anchor still as a constraint, and combines the expanded image sequences generated in the cut units to generate a final expanded image. Effects of the invention

[0022] According to an embodiment of the present application, by using generative artificial intelligence to multi-dimensionally expand the original video, the production time can be shortened by significantly reducing the repetitive work of the editing and compositing process compared to the existing manual-based VFX production process, and delivery stability can be improved through multi-channel synchronization and automation of quality control.

[0023] Furthermore, according to the embodiments of the present application, commercialization efficiency can be enhanced by automatically generating high-quality video synchronized in various immersive display environments (e.g., ScreenX, LED Cube, media wall, etc.). Additionally, it is applicable to various fields of realistic content production, such as the video industry, exhibition industry, and performance industry.

[0024] Furthermore, according to the embodiments of the present application, it can be widely applied to screening environments where multiple display devices are combined, such as movie theaters, immersive performance venues, XR studios, and media facades. Therefore, it has very high industrial practicality as a commercialization platform for AI video generation technology and an automation solution for multi-screen screening systems.

[0025] In addition, according to an embodiment of the present application, by dividing a sequence into cut units based on Edit Reference Lines (EDL) and utilizing anchor stills generated from keyframes of each cut as constraints, initial condition drift within the cut can be suppressed. Through this, the consistency of shape, color, and composition is stabilized during frame generation, thereby reducing compositional instability or color distortion that frequently occurred in conventional frame-independent processing.

[0026] In addition, according to an embodiment of the present application, multiple input conditions, such as text, reference images, and motion hints, can be separated into nodes through a multiconditioning node graph and combined using weighted merging rules. Through this, scene elements such as background, foreground, character, and effect can be separated and precisely controlled, and the reproducibility of results under the same conditions can be ensured.

[0027] In addition, according to an embodiment of the present application, a dual pipeline consisting of a low-latency preview path and an offline purification path can be operated. Users can quickly review results and shorten the approval cycle through the lightweight preview path, and ensure final delivery quality by performing a high-quality purification path only on approved cuts.

[0028] In addition, according to an embodiment of the present application, a time consistency correction layer including sliding window-based interpolation or an inter-frame feedback path may be applied. By doing so, flickering, color drift, and object position instability occurring throughout the entire sequence can be structurally suppressed, thereby improving the synchronization quality between multi-channels and visual immersion.

[0029] In addition, according to an embodiment of the present application, special effect objects can be automatically placed based on rules and high-resolution sequences can be automatically output by linking with a 3D graphics engine. Furthermore, through a synchronization master packaging that includes setting information for standardized encoding and multi-channel synchronization, the delivery specifications of a multi-channel projection device can be complied with and the risk of on-site playback synchronization errors can be prevented in advance.

[0030] The effects obtainable from the embodiments of the present application are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present application pertains from the description below. Brief explanation of the drawing

[0031] A brief description of each drawing is provided to help to better understand the drawings cited in this application. Figure 1 illustrates a multi-screen theater as an example. FIG. 2 is a flowchart of a method for generating content with expanded field of view according to an embodiment of the present application. FIG. 3 is a flowchart illustrating an example of step S240 of FIG. 2. FIG. 4 is a flowchart of a method for generating expanded field of view content according to an embodiment of the present application. FIG. 5 is a flowchart of a method for generating content with expanded field of view according to an embodiment of the present application. FIG. 6 is a diagram illustrating the process of generating an anchor still in a method for generating expanded viewing content according to an embodiment of the present application. FIGS. 7a and 7b exemplarily illustrate expanded content generated according to the method for generating expanded content according to an embodiment of the present application. FIG. 8 is a block diagram showing the configuration of a field-of-view expansion content generation system according to an embodiment of the present application. Specific details for implementing the invention

[0032] The technical concept of the present application is subject to various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the technical concept of the present application to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the scope of the technical concept of the present application.

[0033] In explaining the technical concept of the present application, detailed descriptions of related prior art are omitted if it is determined that such descriptions may unnecessarily obscure the essence of the present application.

[0034] The terms used herein are for describing embodiments and are not intended to limit or / or restrict the present application. Singular expressions include plural expressions unless the context clearly indicates otherwise. Additionally, numbers used herein (e.g., First, Second, etc.) are merely identifiers to distinguish one component from another.

[0035] In this specification, when it is stated that a part is connected to another part, this includes not only cases where they are directly connected, but also cases where they are indirectly connected with other components in between. Furthermore, when it is stated that a part includes a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0036] Furthermore, in this application, the term “or” is intended to mean an implied “or” rather than an exclusive “or.” That is, unless otherwise specified or evident from the context, “X uses A or B” is intended to mean one of the natural implied substitutions. In other words, if X uses A; if X uses B; or if X uses both A and B, “X uses A or B” may apply to any of these cases. Additionally, the term “and / or” as used herein should be understood to refer to and include all possible combinations of one or more of the enumerated related configurations.

[0037] In addition, terms such as “~part,” “~device,” “~device,” and “~module” described in this application refer to a unit that processes at least one function or operation, and this may be implemented as hardware or software or a combination of hardware and software, such as a processor, microprocessor, microcontroller, CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerate Processor Unit), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), and FPGA (Field Programmable Gate Array).

[0038] Furthermore, it is intended to clarify that the classification of the components in this application is merely based on the primary function each component is responsible for. That is, two or more components described below may be combined into a single component, or a single component may be divided into two or more components based on more subdivided functions. Additionally, each component described below may additionally perform some or all of the functions performed by other components in addition to its own primary function, and it is obvious that some of the primary functions performed by each component may be exclusively performed by other components.

[0040] The method according to the embodiment of the present application may be performed on a personal computer, workstation, server computer device, etc., equipped with computing power, or on a separate device for this purpose.

[0041] Additionally, the method may be performed on one or more computing devices. For example, at least one step of the method according to an embodiment of the present application may be performed on a client device, and other steps may be performed on a server device. In this case, the client device and the server device may be connected via a network to transmit and receive computation results. Alternatively, the method may be performed by distributed computing technology.

[0043] In this specification, the term "artificial intelligence learning model" may be used interchangeably with "artificial intelligence model," "computational model," "machine learning model," etc. An artificial intelligence learning model may be trained by various algorithms, such as, for example, decision tree, random forest, Gaussian naive bayes, k-nearest neighbor, Ada Boost, support vector machine, voting, bagging, neural network, and deep learning. However, it is not limited thereto.

[0044] An artificial intelligence learning model can be trained using at least one of supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. The training of an artificial intelligence learning model may be a process of applying knowledge to the model to perform a specific action.

[0045] When algorithms such as neural networks or deep learning are applied to an artificial intelligence learning model, the AI ​​learning model may be referred to as a network function. The term "network function" can be used interchangeably with "neural network." A neural network can generally be composed of a set of interconnected computational units referred to as nodes. These nodes may also be referred to as neurons. A neural network is composed of at least one node, and the nodes may be interconnected by one or more links.

[0046] Neural networks may include deep neural networks (DNNs). Deep neural networks may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q networks, U networks, Siamese networks, and Generative Adversarial Networks (GANs), but are not limited thereto.

[0048] Hereinafter, embodiments of the present application will be described in detail in turn.

[0050] FIG. 2 is a flowchart of a method for generating content with expanded field of view according to an embodiment of the present application, and FIG. 3 is a flowchart showing an embodiment of step S240 of FIG. 2.

[0051] In step S210, editing reference information corresponding to the input video sequence can be obtained.

[0052] Here, the input video sequence may refer to original video data that is the subject of the expanded video generation process and may consist of multiple video frames. For example, the input video sequence may be video content with a standard aspect ratio, such as a 16:9 ratio, and the embodiment may generate an image with an expanded field of view by generating an outer area of ​​such an input video sequence. Additionally, the editing reference information may refer to data that defines the editing structure of the video sequence. For example, the editing reference information may include an Edit Decision List (EDL) or storyboard information. Such editing reference information may include, but is not limited to, time-series information such as the start point, end point, and transition effects between cuts of each cut.

[0053] In step S220, based on the acquired editing reference information, the input video sequence can be divided into multiple cut units.

[0054] In the embodiment, the input video sequence can be separated into physical or logical cut segments by parsing the EDL or storyboard information obtained in step S210. For example, one cut may refer to a video unit with a length of 5 to 10 seconds. However, it is not limited thereto. By dividing the video into cut units in this way, parallel processing of cut units may be possible in subsequent steps.

[0055] In step S230, keyframes can be selected from each cut divided through step S220, and anchor stills, which are extended stills of the outer areas of the keyframes, can be generated.

[0056] A keyframe may be a reference frame representing the cut. For example, the first frame of the cut may be selected as a keyframe. However, if the quality of the first frame of the cut is below a preset standard, the median frame of the cut may be selected as a replacement keyframe. Here, the quality standard may include, but is not limited to, motion blur or exposure instability.

[0057] Subsequently, based on the selected keyframe, an extended still image can be generated for the outer area of ​​the keyframe, for example, the left, right, or top area. This extended still image functions as an anchor still image and can be used as a constraint or guide when generating an extended video sequence for the entire frame of the corresponding cut in a subsequent step (S240). The generation of the anchor still image can be performed using a pre-trained artificial intelligence model. For example, an anchor still image can be generated by performing context-aware fill or outpainting using generative AI-based image generation techniques, such as a latent diffusion model. However, such techniques are exemplary and are not limited thereto.

[0058] In step S240, an extended video sequence for the entire frame of the cut to which each keyframe belongs can be generated by using anchor stills as constraints.

[0059] Similar to step S230, the generation of the expanded video sequence can be performed using a pre-trained artificial intelligence model. For example, generative AI models of the Latent Diffusion, DiT (Diffusion Transformer), or Video-VAE series can consistently generate expanded regions for the remaining frames within a cut by anchoring the anchor still generated in step S230 as a boundary or constraint condition. By using the anchor still as a constraint, the consistency of the shape, color, or composition within the cut can be stably maintained, and the phenomenon of initial conditions drifting can be suppressed.

[0060] In an embodiment, step S240 may include steps S241 to S243 as illustrated in FIG. 3. Step S240 may be performed via a dual path.

[0061] In step S241, a preview of an anchor still or candidate extended image sequence can be provided through a lightweight preview path with low resolution or low computational power.

[0062] A lightweight preview path may mean providing a low-latency response by reducing the amount of computation compared to the final refinement path described below. For example, the preview path may use a lower resolution than the final result and may apply a low-steps method that reduces the number of inference steps of the AI ​​learning model. Additionally, the amount of computation may be reduced by using at least one of a partial attention mechanism, a lightweight sampler, or a lightweight decoding method. This can minimize perceived latency by providing the user with fast feedback (low latency) of about 3 to 6 seconds.

[0063] In step S242, based on the approval of the preview, a final expanded image sequence can be generated through a high-resolution or high-computation refinement path. In the embodiment, the refinement path can ensure the quality of the final result by applying a higher resolution or many high steps compared to the preview path.

[0064] This dual-path structure ensures rapid user review and approval while guaranteeing the quality of the final deliverable.

[0065] In step S243, temporal consistency correction can be performed on the generated final expanded image sequence. This can reduce flicker, color drift, or object position instability issues that may occur due to independent generation of each frame. In the embodiment, temporal consistency correction can be performed using at least one of sliding window-based interpolation or re-synthesis, or an inter-frame feedback path.

[0066] In an embodiment, the step of generating an extended video sequence (S240) may include the step of embedding a global style token representing at least one of the global look, tone, and lighting characteristics of the input video sequence. By injecting this global style token throughout the cut generation, tone consistency between multiple extended video sequences, i.e., between cuts, can be maintained.

[0067] In step S250, the expanded video sequences generated at the cut level can be combined to generate the final expanded video.

[0068] That is, in step S250, the extended video sequences generated in parallel for each cut in step S240 can be combined again in chronological order to generate a single complete extended video. In the embodiment, step S250 may include a step of performing correction to mitigate discontinuities between cuts at the boundaries of adjacent cuts. For example, a guard-band overlap can be applied to a specific frame interval (e.g., ±4 to 8 frames) of the cut boundary, and cross-blending or dissolve effects can be applied to ensure a natural transition between cuts.

[0069] In step S260, a master package for compatibility with a multi-channel projection device can be generated based on the final extended video.

[0070] This may refer to the step of formatting the generated final extended video so that it is immediately compatible with multi-projection devices or multi-channel digital signage. For example, standard encoding tools such as FFmpeg can be used to batch standardize the resolution, codec, or frame rate per channel. Additionally, a synchronization master package containing frame indices or timeline metadata (e.g., JSON, XML) can be generated to ensure playback synchronization between channels. However, such methods are exemplary and are not limited thereto.

[0071] The configurations of FIGS. 2 and FIGS. 3 are exemplary, and various configurations may be applied according to embodiments of the present application.

[0073] FIG. 4 is a flowchart of a method for generating expanded field of view content according to an embodiment of the present application.

[0074] The method (400) of FIG. 4 may further include steps S410 and S420 in addition to the method (200) described above with reference to FIG. 2. Meanwhile, although steps S410 and S420 are shown in FIG. 4 as being performed before step S210, this is not limited thereto, and can be varied in various ways, such as steps S410 and / or S420 being performed after steps S210 and / or S220.

[0075] In step S410, a plurality of input conditions for controlling image generation may be obtained. In an embodiment, the input conditions may include at least one of text (e.g., a prompt), a reference image, a mask, and a motion hint. The motion hint may include, for example, optical flow or keypoint information, but is not limited thereto.

[0076] In step S420, a multiconditioning node graph can be used to separate multiple input conditions into node units and generate complex condition information through weighted merging rules between nodes. Specifically, the node graph architecture can model different types of input conditions, such as text, reference images, masks, and motion hints, by separating them into individual nodes. Subsequently, these can be combined through weighted merging rules or priority rules that regulate the influence between each node to generate final complex condition information. This node graph-based approach facilitates the separation and control of scene elements (e.g., background, foreground, character, effects), thereby improving the quality and repeatability of the output.

[0077] The generated composite condition information can be used as foundational information when performing at least one of the anchor still generation in step S230 and the extended video sequence generation in step S240. For example, in step S230, the composite condition information can be injected as input conditioning for an artificial intelligence learning model. This allows the model to generate the outer region of a keyframe by reflecting all conditions of the text, reference image, and mask. Additionally, in step S240, the composite condition information can serve as a guide that is consistently applied across all frames within a cut. This enables precise control of the spatiotemporal characteristics of the generated video sequence, such as inducing an object to move following motion hints or maintaining a specific global style throughout the cut.

[0078] The configuration of FIG. 4 is exemplary, and various configurations may be applied according to embodiments of the present application.

[0080] FIG. 5 is a flowchart of a method for generating content with expanded field of view according to an embodiment of the present application.

[0081] The method (500) of FIG. 5 may further include steps S510 and S520 in addition to the method (200) described above with reference to FIG. 2.

[0082] In step S510, 3D special effect objects can be placed rule-based on each of the extended video sequences generated in cut units in step S240 through scenario analysis on a 3D graphics engine. For example, the 3D graphics engine may include Unity or Unreal Engine, and 3D special effect objects such as Visual Effects (VFX), light sources, flares, or particles can be automatically placed by analyzing the scene's scenario or timing. This can improve the inefficiency of the existing method that relied on manual compositing.

[0083] In step S520, a multi-channel high-resolution image sequence containing the placed 3D special effect objects can be output by linking with a time-series automation module of the 3D graphics engine. For example, through a time-series automation module such as Unity’s Timeline or Unreal’s Sequencer, a multi-channel (e.g., left, center, right) high-resolution image sequence containing the 3D objects placed in S510 can be automatically output (rendered).

[0084] In this case, the final expanded image generation of step S250 can be performed by combining multi-channel high-resolution image sequences generated in cut units.

[0085] The configuration of FIG. 5 is exemplary, and various configurations may be applied according to embodiments of the present application.

[0087] FIG. 6 is a diagram illustrating the process of generating an anchor still in a method for generating expanded viewing content according to an embodiment of the present application.

[0088] Referring to FIG. 6, it can be seen that based on the keyframe (Original resource) of the input image, an extended still image is generated for the outer regions of the keyframe, namely the left, right, and top regions. Here, the center region may correspond to the region of the original keyframe. The extended still image of the left, right, and top regions generated in this way, along with the entire still image including the center region, may be an anchor still generated in step S230 described above with reference to FIG. 2. This anchor still can be used as a constraint or boundary condition when generating an extended video sequence for the entire frame of the cut.

[0090] FIGS. 7a and 7b exemplarily illustrate expanded content generated according to the method for generating expanded content according to an embodiment of the present application.

[0091] Specifically, FIG. 7a is an original video sequence (or keyframe) before expansion, and FIG. 7b may be an image of an expanded field of view in which an outer area is generated based on this original video through an expansion content generation method according to an embodiment of the present application.

[0092] At this time, 3D special effect objects (VFX) may be additionally placed in the expanded field of view image. The scenario or timing of the scene is analyzed on a 3D graphics engine, and based on the results of this analysis, 3D special effect objects (VFX), such as the floating island, upper structure, or energy beam exemplified in FIG. 7b, can be automatically placed in a rule-based manner. This process may include automatically specifying the position, scale, rotation values, etc., of the 3D objects.

[0094] FIG. 8 is a block diagram showing the configuration of a field-of-view expansion content generation system according to an embodiment of the present application.

[0095] The system (800) may be implemented with components performing each function integrated into a single physical device (or server) or distributed and operated across multiple physical devices. In this case, the multiple physical devices may exchange data with each other and perform a series of processes as described in the embodiment of the present application.

[0096] The communication unit (810) can receive or transmit data from inside or outside. The communication unit (810) may include a wired or wireless communication unit. If the communication unit (810) includes a wired communication unit, the communication unit (810) may include one or more components that enable communication through a Local Area Network (LAN), a Wide Area Network (WAN), a Value Added Network (VAN), a mobile radio communication network, a satellite communication network, and combinations thereof. Additionally, if the communication unit (810) includes a wireless communication unit, the communication unit (810) may transmit or receive data or signals wirelessly using cellular communication, a wireless LAN (e.g., Wi-Fi), etc. In an embodiment, the communication unit (810) may transmit or receive data or signals to or from an external device or an external server under the control of a processor (840).

[0097] The input unit (820) can receive various user commands through external operation. To this end, the input unit (820) may include or be connected to one or more input devices. For example, the input unit (820) may receive user commands by being connected to an interface for various inputs, such as a keypad or a mouse. To this end, the input unit (820) may include an interface such as a USB port as well as a Thunderbolt. Additionally, the input unit (820) may receive external user commands by including or being combined with various input devices such as a touchscreen or a button.

[0098] The memory (830) can store programs and / or program instructions for the operation of the processor (840) and can temporarily or permanently store input / output data. The memory (830) may include at least one type of storage medium among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory, etc.), RAM, SRAM, ROM, EEPROM, PROM, magnetic memory, magnetic disk, and optical disk.

[0099] Additionally, the memory (830) can store various artificial intelligence learning models, network functions and algorithms, and can store various data, programs (one or more instructions), applications, software, commands, code, etc. for operating and controlling the system (800).

[0100] The processor (840) can control the overall operation of the system (800). The processor (840) can execute one or more programs or software stored in memory (830). The processor (840) may mean a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), or a dedicated processor (840) on which the methods according to embodiments of the present application are performed.

[0101] In an embodiment, the processor (840) can perform the methods (200, 400, 500) described above with reference to FIGS. 2 to 5 by executing one or more programs or software stored in memory (830).

[0102] In an embodiment, the processor (840) can obtain editing reference information corresponding to an input video sequence by executing a program stored in memory (830). Based on the obtained editing reference information, the processor (840) can divide the input video sequence into multiple cut units. Additionally, the processor (840) can select a keyframe from each of the divided cuts. Based on each keyframe, the processor (840) can generate an anchor still, which is an extended still of the outer area of ​​the keyframe. Then, the processor (840) can generate an extended video sequence for the entire frame of the cut to which each keyframe belongs by using the anchor still as a constraint. Finally, the processor (840) can generate a final extended video by combining the extended video sequences generated by cut units.

[0103] In an embodiment, the processor (840) can use a multiconditioning node graph to separate multiple input conditions for image generation into node units and generate composite condition information through weighted merging rules between nodes. In this case, the processor (840) can generate an anchor still or an extended image sequence based on the composite condition information.

[0104] In an embodiment, the input condition may include at least one of text, a reference image, a mask, and a motion hint.

[0105] In an embodiment, when the processor (840) generates an extended image sequence, it may provide a preview of an anchor still or a candidate extended image sequence through a lightweight preview path of low resolution or low computational amount. Subsequently, based on the approval of the preview, the processor (840) may generate a final extended image sequence through a refinement path of high resolution or high computational amount.

[0106] In an embodiment, the processor (840) can perform temporal consistency correction on the generated extended image sequence.

[0107] In the embodiment, time consistency correction can be performed using at least one of sliding window-based interpolation or re-synthesis, or an inter-frame feedback path.

[0108] In an embodiment, when the processor (840) selects a keyframe, it may select the first frame of the cut as the keyframe. However, if the quality of the first frame of the cut is below a preset standard, the processor (840) may select the median frame of the cut as the keyframe instead.

[0109] In an embodiment, when the processor (840) generates an extended image sequence, it can maintain tone consistency between multiple extended image sequences by embedding a global style token representing at least one of the global look, tone, and lighting characteristics of the input image sequence.

[0110] In an embodiment, the processor (840) can place 3D special effect objects on each cut-unit expanded video sequence generated through scenario analysis on a 3D graphics engine based on rules. Additionally, the processor (840) can output a multi-channel high-resolution image sequence containing the placed 3D special effect objects by linking with the time-series automation module of the 3D graphics engine. In this case, the processor (840) can generate a final expanded video by combining the cut-unit multi-channel high-resolution image sequences.

[0111] In an embodiment, when combining extended image sequences, the processor (840) can perform a correction to mitigate discontinuities between cuts at the boundaries of adjacent cuts.

[0112] In an embodiment, the processor (840) can generate a master package for compatibility with a multi-channel screening system based on the final expanded image.

[0113] The configuration of FIG. 8 is exemplary, and various configurations may be applied according to embodiments of the present application.

[0115] The method according to an embodiment of the present application may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the present application or may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.

[0116] Additionally, the method according to the disclosed embodiments may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product.

[0117] A computer program product may include a software program and a computer-readable storage medium on which the software program is stored. For example, a computer program product may include a product in the form of a software program (e.g., a downloadable app) that is electronically distributed through a manufacturer of an electronic device or an electronic market (e.g., Google Play Store, App Store). For electronic distribution, at least a portion of the software program may be stored on a storage medium or temporarily created. In this case, the storage medium may be a server of the manufacturer, a server of the electronic market, or a storage medium of a relay server that temporarily stores the software program.

[0118] A computer program product may include a storage medium of a server or a storage medium of a client device in a device composed of a server and a client device. Alternatively, if a third device (e.g., a smartphone) is communicationally connected to the server or client device, the computer program product may include a storage medium of the third device. Alternatively, the computer program product may include the S / W program itself, which is transmitted from the server to the client device or the third device, or from the third device to the client device.

[0119] In this case, one of the server, the client device, and the third device may execute the computer program product to perform the method according to the disclosed embodiments. Alternatively, two or more of the server, the client device, and the third device may execute the computer program product to perform the method according to the disclosed embodiments in a distributed manner.

[0120] For example, a server (e.g., a cloud server or an artificial intelligence server, etc.) can execute a computer program product stored on the server to control a client device connected to the server in communication to perform the method according to the disclosed embodiments.

[0122] Although the embodiments have been described in detail above, the scope of the present application is not limited thereto, and various modifications and improvements by those skilled in the art using the basic concept of the present application as defined in the following claims also fall within the scope of the present application.

Claims

Claim 1 A method for generating expanded field of view content, comprising: a step of obtaining editing reference information corresponding to an input video sequence; a step of dividing the input video sequence into a plurality of cut units based on the editing reference information; a step of selecting a keyframe from each of the divided cuts; a step of generating an anchor still, which is an expanded still of the outer area of ​​the keyframe, based on each of the keyframes; a step of generating an expanded video sequence for the entire frame of the cut to which each of the keyframes belongs, using the anchor still as a constraint; a step of performing temporal consistency correction on the generated expanded video sequence; and a step of combining the expanded video sequences generated in the cut units to generate a final expanded video. Claim 2 The method of claim 1 further comprises the step of using a multi-conditioning node graph to separate a plurality of input conditions for image generation into node units and generating composite condition information through a weighted merging rule between the nodes, wherein at least one of the step of generating the anchor still and the step of generating the extended image sequence is performed based on the composite condition information. Claim 3 In claim 2, the input condition comprises at least one of text, a reference image, a mask, and a motion hint. Claim 4 The method of claim 1, wherein the step of generating the expanded image sequence comprises: providing a preview of the anchor still or candidate expanded image sequence through a lightweight preview path of low resolution or low computational amount; and generating a final expanded image sequence through a refinement path of high resolution or high computational amount based on the approval of the preview. Claim 5 delete Claim 6 A method according to claim 1, wherein the time consistency correction is performed using at least one of sliding window-based interpolation or re-synthesis, or an inter-frame feedback path. Claim 7 A method according to claim 1, wherein the step of selecting the keyframe is performed by selecting the first frame of the cut as the keyframe, and if the quality of the first frame of the cut is less than a preset standard, replacing the median frame of the cut with the keyframe. Claim 8 The method of claim 1, wherein the step of generating the extended image sequence comprises embedding a global style token representing at least one of the global look, tone, and lighting characteristics of the input image sequence to maintain tone consistency among a plurality of the extended image sequences. Claim 9 A method according to claim 1, comprising: a step of placing 3D special effect objects in each of the expanded image sequences generated in cut units through scenario analysis on a 3D graphics engine based on rules; and a step of outputting a multi-channel high-resolution image sequence including the placed 3D special effect objects in conjunction with a time-series automation module of the 3D graphics engine, wherein the step of generating the final expanded image is performed by combining the multi-channel high-resolution image sequences generated in cut units. Claim 10 A method according to claim 1, wherein the step of combining the extended image sequence comprises the step of performing a correction to mitigate inter-cut discontinuity at the boundaries of adjacent cuts. Claim 11 A method according to claim 1, further comprising the step of generating a master package for compatibility with a multi-channel projection device based on the final expanded image. Claim 12 A computer program stored on a recording medium to execute a method according to any one of claims 1 through 4 and claims 6 through 11. Claim 13 A system for generating content for expanding a field of view, comprising: at least one processor; and a memory for storing a program executable by said processor, wherein the processor, by executing said program, obtains editing reference information corresponding to an input video sequence, divides said input video sequence into a plurality of cut units based on said editing reference information, selects a keyframe from each of said divided cuts, generates an anchor still which is an expanded still of an outer area of ​​said keyframe based on each of said keyframes, generates an expanded video sequence for the entire frame of the cut to which each of said keyframes belongs by using said anchor still as a constraint, performs temporal consistency correction on said expanded video sequence, and combines said expanded video sequences generated in said cut units to generate a final expanded video.

Citation Information

Patent Citations

  • Event-based motion estimation for multi-frame stacking

    KR1020250055411A

  • Image processing apparatus and method for generating information beyond image area

    KR102492430B1

  • Methods for generating video and multiple still images simultaneously and apparatuses using the same

    US20140078343A1

  • User interface for just-in-time image processing

    US20210142542A1

  • Automated key frame selection

    US20230260251A1