Method and system for reconstructing 3D spaces from content videos

The method and system for reconstructing 3D spaces from content videos address the inefficiencies of conventional methods by segmenting and classifying video shots, extracting visual features, and applying 3D structuring techniques, resulting in efficient and cost-effective 3D reconstruction of places/spaces for use in virtual and augmented reality platforms.

WO2025135498A1PCT designated stage expired Publication Date: 2025-06-26CJ OLIVENETWORKS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017533
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-11-07
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Conventional methods for 3D reconstruction of spaces from video content are inefficient and costly, particularly when dealing with places/spaces that were not originally intended for 3D reconstruction, as they require expensive equipment and studio environments, and struggle with isolating specific spaces from edited video content.

Method used

A method and system for reconstructing 3D spaces from content videos, involving shot segmentation, scale classification, close-up shot removal, visual feature extraction, similar location grouping, 3D structuring, camera motion extraction, frame sampling, additional frame 3D structuring, and rendering, allowing for the 3D reconstruction of places/spaces from edited video content without the need for expensive equipment or studio environments.

Benefits of technology

Enables efficient and cost-effective 3D reconstruction of places/spaces from edited video content, allowing for immediate use in virtual and augmented reality platforms, reducing the time and cost of creating new 3D models, and facilitating the creation of a database of 3D reconstructed spaces for viewing from different viewpoints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017533_26062025_PF_FP_ABST
    Figure KR2024017533_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for reconstructing 3D spaces from content videos, which allows the 3D reconstruction of places / spaces that appear within captured video content, and the method for reconstructing 3D spaces from content videos according an embodiment of the present invention includes a shot segmentation step, a shot scale classification step, a close-up shot removal step, a visual feature extraction step, a similar location grouping step, a 3D structuring step, a camera motion extraction step, a frame sampling step, an additional frame 3D structuring step, and a rendering step.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR RECONSTRUCTING 3D SPACES FROM CONTENT VIDEOS

[0001] The present invention relates to a method and system for reconstructing 3D spaces from content videos, which allows the 3D reconstruction of places / spaces that appear within captured video content.

[0002] Recently, 3D models for objects, people, spaces, etc. have been rapidly developed and are being utilized in various fields such as movies, entertainment shows, dramas, advertisements, education, news, games, sports, performances, etc. to provide a high level of immersion and presence.

[0003] Developing 3D models from scratch requires a lot of time and cost, and creating all objects anew without reference materials is a very challenging process. Therefore, there has been extensive research on 3D reconstruction to restore the actual 3D shapes of objects / spaces / people from images / videos. Moreover, efforts are being made to reconstruct 3D models of objects that appear in existing content such as photos, movies, dramas, entertainment shows, performances, etc.

[0004] In the past, creating 3D models required the use of expensive equipment such as 3D scanners or the acquisition of data in studio environments where multiple cameras were fixedly installed in various locations. According to this conventional method, it is possible to acquire data for subjects such as people and objects in such environments, but it is impossible to construct data for places / spaces in this way. Furthermore, obtaining multi-view data for the places / spaces typically requires the use of equipment such as drones; however, it is extremely difficult to achieve 3D reconstruction based on video content that was not originally intended for 3D reconstruction. Specifically, the video content that was not originally intended for 3D reconstruction is taken and edited with multiple cameras, resulting in numerous shots (in units of screen transition). In addition, 3D reconstruction requires multi-view data for a specific space, but it is nearly impossible to isolate a specific space from such video content and then acquire multi-view data.

[0005] The present invention is proposed to solve the problems of the conventional method and relates to a method for reconstructing 3D places and spaces from content videos taken at various scales with motions in multiple spaces, and a system therefor.

[0006] Accordingly, the present invention has been made to solve the above-described problems, and an object of the present invention is to provide a method for reconstructing a 3D space from a content video, which allows the 3D reconstruction of places / spaces from a content video taken and edited at various scales with motions in multiple spaces.

[0007] Another object of the present invention is to provide a method for reconstructing a 3D space, which allows the 3D reconstruction of places / spaces for immediate use in the metaverse and similar platforms, thereby reducing the time and cost of creating new places / spaces, resulting in improved productivity.

[0008] Still another object of the present invention is to provide a method for reconstructing a 3D space, which can build a database of 3D reconstructed places / spaces that appear within the content and allow for viewing these 3D reconstructed places / spaces stored in that database from different viewpoints.

[0009] The objects of the present invention are not limited to those mentioned above, and other objects not mentioned will be clearly understood by those skilled in the art from the following description.

[0010] To accomplish the above objects of the present invention, there is provided a method for reconstructing a 3D space from a content video, the method comprising: a shot segmentation step; a shot scale classification step; a close-up shot removal step; a visual feature extraction step; a similar location grouping step; a 3D structuring step; a camera motion extraction step; a frame sampling step; an additional frame 3D structuring step; and a rendering step.

[0011] According to the present invention, it is possible to group shots taken in the same space to allow the 3D reconstruction of places / spaces from a content video taken and edited at various scales with motions in multiple spaces and to also allow the 3D reconstruction of places / spaces that appear within the video content taken without any intention for 3D reconstruction.

[0012] Moreover, according to the present invention, it is possible to allow the 3D reconstruction of places / spaces in existing content for immediate use in the virtual and augmented reality, the metaverse, and similar platforms, thereby reducing the time and cost of creating new places / spaces, resulting in improved productivity.

[0013] Furthermore, according to the present invention, it is possible to build a database of 3D reconstructed places / spaces that appear within the content and allow for viewing these 3D reconstructed places / spaces stored in that database from different viewpoints. Therefore, the present invention offers the advantage of allowing a user to draw a captured image of a similar environment without needing to revisit the same location.

[0014] The effects of the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description.

[0015] FIG. 1 is a block diagram illustrating a system according to an embodiment of the present invention.

[0016] FIG. 2 is a block diagram illustrating an AI module according to an embodiment of the present invention.

[0017] FIG. 3 is a flowchart illustrating a method for reconstructing a 3D space from a content video according to an embodiment of the present invention.

[0018] FIG. 4 is a diagram illustrating a scale classification step according to an embodiment of the present invention.

[0019] FIG. 5 is a diagram illustrating a similar location grouping step according to an embodiment of the present invention.

[0020] FIGS. 6 to 8 are diagrams illustrating a Structure from Motion (SfM) algorithm according to an embodiment of the present invention.

[0021] FIGS. 9 to 12 are diagrams illustrating a camera motion extraction step according to an embodiment of the present invention.

[0022] FIGS. 13 and 14 are diagrams showing rendered images produced according to an embodiment of the present invention.

[0023] The advantages and features of the present invention, as well as the methods for achieving them will become clear with reference to the embodiments described in detail below along with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and can be implemented in various forms. These embodiments are provided to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art to which the present invention pertains about the scope of the present invention, and the present invention is defined only by the scope of the claims. Throughout the specification, the same reference numerals refer to the same components.

[0024] When a component is referred to as being "connected to" or "coupled with" another component, it encompasses both cases where it is directly connected to or coupled with another component and where there is an intervening component. On the contrary, when a component is referred to as being "directly connected to" or "directly coupled with" another component, it indicates that there is no component between them. The term "and / or" includes each of the mentioned items and all possible combinations thereof.

[0025] The terms used herein are intended to describe the embodiments and are not meant to limit the present invention. As used herein, singular forms also include plural forms, unless specifically stated otherwise in the context. The terms "comprises" and / or "comprising" as used herein do not exclude the presence of addition of one or more other components, steps, operations and / or elements.

[0026] Although the terms "first", "second", or the like are used to describe various components, these components are not limited by these terms. These terms are merely used to distinguish one component from another. Therefore, it goes without saying that a component referred to as "first" in the following description may also be a "second" component within the technical spirit of the present invention.

[0027] Throughout the specification, when a part is said to "include" a certain component, it means that it may further include other components, rather than excluding other components, unless specifically stated otherwise. Furthermore, throughout the specification, the term "on" refers to being positioned above or below the target, and does not necessarily refer to being positioned above relative to the direction of gravity.

[0028] As used herein, the terms related to direction, such as "front," "rear," "left," "right," "vertical," and "horizontal" are used in a relative sense for the convenience of explanation and may vary depending on the viewing direction .

[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein may have the same meaning as those commonly understood by those of ordinary skill in the art to which the present disclosure pertains. In addition, terms defined in commonly used dictionaries are not interpreted ideally or excessively unless explicitly specifically defined.

[0030] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the attached drawings. Here, it should be noted that the same components are denoted by the same reference numerals in the accompanying as much as possible. Detailed descriptions of well-known functions and configurations that may obscure the gist of the present invention will be omitted. For the same reason, some components are exaggerated, omitted, or schematically shown in the accompanying drawings.

[0031] FIG. 1 is a block diagram illustrating a system according to an embodiment of the present invention, FIG. 2 is a block diagram illustrating an AI module according to an embodiment of the present invention, FIG. 3 is a flowchart illustrating a method for reconstructing a 3D space from a content video according to an embodiment of the present invention, FIG. 4 is a diagram illustrating a scale classification step according to an embodiment of the present invention, FIG. 5 is a diagram illustrating a similar location grouping step (S500) according to an embodiment of the present invention, FIGS. 6 to 8 are diagrams illustrating a Structure from Motion (SfM) algorithm according to an embodiment of the present invention, FIGS. 9 to 12 are diagrams illustrating a camera motion extraction step (S700) according to an embodiment of the present invention, and FIGS. 13 and 14 are diagrams showing rendered images produced according to an embodiment of the present invention.

[0032] The system illustrated in FIG. 1 may comprise an input unit 120, a control unit 110, a first memory 140, an output unit 130, etc. The components illustrated in FIG. 1 are not essential for implementing an apparatus, and the system or apparatus described herein may have more or fewer components than those listed above.

[0033] The first memory 140 stores data supporting various functions of the system. The first memory 140 may store a plurality of application programs or applications running on the system, as well as data and instructions for the operations of the system. Meanwhile, the application programs may be stored in the first memory 140, installed on the system, and executed by the control unit 110 to perform the operations (or functions) of the system.

[0034] The control unit 110 typically controls the overall operation of the system as well as the operations related to the application programs. The control unit 110 may provide or process appropriate information or functions by processing signals, data, information, etc., which are input or output through the components discussed above, or running the application programs stored in the first memory 140. Furthermore, the control unit 110 may control at least some of the components described with reference to FIG. 1 to run the application programs stored in the first memory 140. In addition, the control unit 110 may operate at least two of the components included in the system in combination to run the application programs.

[0035] At least some of the above components may cooperate with each other to implement the operation, control, or control method of the system according to various embodiments described below. Additionally, the operation, control, or control method of the system can be implemented on an electronic device by running at least one application program stored in the first memory 140.

[0036] FIG. 2 is a block diagram illustrating an AI module according to an embodiment of the present invention. In an embodiment of the present invention, the AI module performs the operations and calculations required in the shot segmentation step (S100) through the rendering step (S1000) of the present invention. The AI module may comprise an electronic device capable of performing AI processing or a server including the AI module. Moreover, the AI module may be included as at least part of the components of the system illustrated in FIG. 1 to perform at least part of the AI processing together. The AI module may comprise an AI processor 210, a second memory 230, and / or a communication unit 220. In an embodiment of the present invention, the AI module can be implemented as a computing device capable of learning a neural network and can be implemented as various electronic devices such as servers, desktop PCs, laptop PCs, tablet PCs, etc. The AI processor 210 can learn a neural network using a program stored in the second memory 230.

[0037] FIG. 3 is a flowchart specifying the steps of a method for reconstructing a 3D space from a content video according to an embodiment of the present invention. This is merely a preferred embodiment for achieving the objects of the present invention, and some steps may be added or deleted as needed, and further, one step may be included in another step.

[0038] Referring to FIG. 3, the method for reconstructing a 3D space from a content video according to an embodiment of the present invention may comprise a shot segmentation step (S100), a shot scale classification step (S200), a close-up shot removal step (S300), a visual feature extraction step (S400), a similar location grouping step (S500), a 3D structuring step (S600), a camera motion extraction step (S700), a frame sampling step (S800), an additional frame 3D structuring step (S900), and a rendering step (S1000).

[0039] The present invention relates to a method for constructing a database of 3D reconstructed places / spaces that appear within a video based on the video and creating an image from a new viewpoint for the places / spaces in the video from the database.

[0040] Meanwhile, it is understood that the type of video used for the 3D reconstruction in the present invention can be varied, but in a preferred embodiment of the present invention, the 3D reconstruction is performed using video content such as movies, dramas, entertainment shows, etc. This is aimed at building a database of 3D reconstructed places and spaces using the characteristics of video content such as movies, dramas, entertainment shows, etc., which are taken and edited with multiple cameras in various ways across different spaces, of which a single shot is taken with various types of screen scales and camera motions, and of which the space is divided into multiple viewpoints. In this case, the video may be one that was not originally intended for 3D reconstruction.

[0041] Referring again to FIG. 3, the shot segmentation step (S100) is to segment a video into shots. More specifically, the shot segmentation step (S100) is to segment the video into shots based on screen transitions. In this case, a shot may be a collection of frames from a certain segment of the video. There may be various criteria for segmentation depending on the content or composition of the video, but preferably, one criterion may be the place / space as shown in the video, where the video was taken.

[0042] The above-described shot segmentation step (S100) is to make it easier for the AI model to analyze and understand the video. More specifically, in cases where videos such as movies, dramas, and entertainment content videos are taken at various scales with multiple cameras across various places / spaces and then combined into a single video through editing, it may be difficult for a single AI model to analyze and understand the video. However, the shot segmentation step (S100) is first performed to segment a video into shot units so as to make it easier for the AI model to analyze the video.

[0043] Next, the shot scale classification step (S200) may be executed. Referring to FIG. 4, the shot scale classification step (S200) is to classify the shots segmented in the shot segmentation step (S100) according to their scales. The shot scale classification step (S200) is to classify the shots based on the scales for the video containing the shots taken at various scales. At this time, the scale criteria for classifying the shots can be defined in various ways. For example, as shown in FIG. 4, if the scale is defined based on the degree to which a person is closed up or the area occupied by the person within the frame, the scales can be classified into extreme close-up shot, close-up shot, medium shot, full shot, long shot, etc. The scale is not necessarily defined as the degree to which a person is closed up; and the scale may be defined based on the degree to which an object is closed up, that is, the area it occupies within the frame, or based on the proportion of the background within the frame. The shot scale classification step (S200) is performed to classify the shots taken at close-up scales, which will be removed in the close-up shot removal step which will be described later

[0044] After the shot scale classification step, if there is any shot classified as a close-up shot in the shot scale classification step (S200), the close-up shot removal step (S300) is performed to delete the shot(s) classified as the close-up shot(s). In the case of a close-up shot, it is difficult to obtain information about the place / space within a video because it is a scene that focuses on a single object within the video; however, this step is performed to remove the close-up shot so as to facilitate the AI model's ability to obtain information about the place / space, i.e., so as to reduce the amount of calculation of the AI model and increase the efficiency of the calculation. Meanwhile, in the present invention, if there is no shot classified as the close-up shot, it is also possible to omit this close-up shot removal step (S300).

[0045] Referring briefly to FIG. 4, the shot scale classification step (S200) and the close-up shot removal step (S300) will now be described as an example. In an embodiment of the present invention, the shot scale classification step can classify the shots included in the video into five scales: an extreme close-up shot, a close-up shot; a medium shot; a full shot, and a long shot. In the close-up shot removal step (S300), the shots corresponding to the extreme close-up shot and the close-up shot can be removed.

[0046] The visual feature extraction step (S400) is to extract visual features for each shot classified in the shot scale classification step (S200). The visual feature extraction step (S400) is performed to reduce the amount of calculation required for a feature point matching algorithm performed in the subsequent Structure from Motion (SfM) step. The feature point matching algorithm, which is performed for each image pair, requires a large amount of calculation if performed on all shots except for those removed in the close-up shot removal step (S300). Therefore, in an embodiment of the present invention, the classification of a candidate set of the most similar shots is first performed, and then the feature point matching algorithm is performed on the classified shots. The visual feature extraction step (S400) is to extract features of the shots to classify the most similar shots.

[0047] Meanwhile, in an embodiment of the present invention, the visual feature extraction step (S400) can be implemented to extract visual features by selecting only the first frame of each shot. This approach is based on the premise that a single shot is likely taken at the same location, and even if the shot is taken at various angles, the background information taken at the same location remains visually similar. Extracting visual features from every frame within a shot could increase the calculation time and require significant calculation resources. Therefore, the present invention can be implemented to extract visual features by selecting only the first frame of each shot.

[0048] After the visual feature extraction step, the similar location grouping step (S500) is performed to group the shots taken at similar locations based on the visual features for each shot extracted in the visual feature extraction step. More specifically, in the similar location grouping step (S500), a clustering algorithm is performed based on the visual features for each shot extracted in the visual feature extraction step to group the shots taken at similar places.

[0049] Meanwhile, according to an embodiment of the present invention, the similar location grouping step (S500) can use an agglomerative clustering algorithm, a type of hierarchical clustering, to perform the grouping. More specifically, in an embodiment of the present invention, if the agglomerative clustering algorithm is applied based on the visual features extracted for each shot in the visual feature extraction step (S400), it can initiate the clusters with the smallest units as shown in FIG. 5, and then hierarchically create the connections between clusters based on the similarity of visual features of each shot. In FIG. 5, the horizontal axis represents the index of small unit clusters, and the vertical axis represents the degree of similarity between clusters.

[0050] Referring again to FIG. 3, the 3D structuring step (S600) is to extract 3D information about the shots grouped in the similar location grouping step. More specifically, this step applies the Structure from Motion (SfM) algorithm to the shots classified as being taken at the same location in the shot segmentation step (S100) through the similar location grouping step (S500). The Structure from Motion (SfM) algorithm is a method for extracting 3D information of an object from 2D images taken at various angles of the same object. This method makes it possible to calculate camera intrinsic parameters such as the camera's focal length, camera external parameters including camera 3D position and camera rotation information, and object 3D location information.

[0051] The Structure from Motion (SfM) algorithm will now be described in more detail with reference to FIGS. 6 to 8. The SfM algorithm reconstructs 3D information from images of the same location as shown in FIG. 6. As shown in FIG. 7, the feature point matching algorithm is performed to match the correspondences between the visual features (feature points) extracted from one shot and those from another shot taken at the same location (the lines connect the feature points in the left image with the feature points in the right image). As shown in FIG. 8, the 3D information about the place / space is obtained by calculating the geometric relationships of the place / space on the shots from the correspondences of multiple shots of the same place / space.

[0052] Meanwhile, in an embodiment of the present invention, the first frame of the shot is used instead of the entire shot in the visual feature extraction step (S400) through the 3D structuring step (S600). This approach is intended to reduce the amount of calculation required in the visual feature extraction step (S400) through the 3D structuring step (S600), as typical videos consist of a very large number of frames.

[0053] The camera motion extraction step (S700), which is performed after the 3D structuring step, is to calculate a camera motion for each shot based on the 3D information obtained in the 3D structuring step (S600). The video used in an embodiment of the present invention may include various types of camera motions. For example, in an embodiment of the present invention, a video may include shots taken by a fixed camera where only the subject or object moves while the background remains stationary, shots taken by a camera moving simply up, down, left or right, shots taken by a camera mounted on a drone covering a wide area or moving dynamically, etc. In this context, the camera motion extraction step (S700) applies a simultaneous location and mapping (SLAM) algorithm based on the camera intrinsic parameters for the shots extracted in the 3D structuring step (S600) to perform the camera motion extraction for each shot.

[0054] The camera motion extraction step (S700) according to an embodiment of the present invention will now be described in more detail with reference to FIGS. 9 to 12. In an embodiment of the present invention, the camera motion displayed in three dimensions can be represented as shown in FIGS. 9 and 11. FIG. 10 is a diagram showing the camera motion of FIG. 9 in two dimensions, and FIG. 12 is a diagram showing the camera motion of FIG. 11 in two dimensions. Meanwhile, FIGS. 9 and 10 illustrate a case where an image is taken by a rotating and moving camera, and the line shown in FIG. 10 represents the motion of the camera. Moreover, FIGS. 11 and 12 illustrate a case where an image is taken by a fixed camera, and the point shown in FIG. 12 indicates the image taken by the fixed camera. In the present invention, the camera motion was quantified using the most commonly used metric for calculating the camera trajectory error. When the ground truth (GT) is set to 0, the motion values can be quantified. The terms used in the graphs of FIGS. 9 and 11 can be explained as follows: ATE (Absolute Trajectory Error) is a simple metric to accumulate errors of differences in x, y, and z values from the GT, not reflecting the gradual deviation along the sequence; and RPE (Relative Pose Error) is a metric to measure translation and rotation errors separately and accumulate the relative position errors, compensating for the shortcomings of ATE.

[0055] Meanwhile, in the present invention, if the camera captures the shots included in the video while remaining stationary, it is also possible to omit the camera motion extraction step (S700). Furthermore, in this embodiment, it is also possible to perform each step of the present invention using only the first frame of the shot.

[0056] Meanwhile, in the present invention, for the shot taken by a fixed camera, it also possible to perform each step of the present invention using only the first frame; however, for the shot where there is a motion, it is possible to extract further 3D information by further sampling the frames.

[0057] The frame sampling step (S800) is to sample and add frames of a shot where there is a camera motion. In this case, the frame sampling step (S800) is performed based on the camera motion for the shot extracted in the camera motion extraction step (S700).

[0058] Meanwhile, in an embodiment of the present invention, the frame sampling step (S800) performs 3D structuring on the sampled and added frames to extract 3D information. Moreover, in an embodiment of the present invention, the frame sampling step (S800) can be implemented to perform 3D structuring only on the sampled and added frames to extract 3D information.

[0059] Meanwhile, the additional frame 3D structuring step (S900) may be omitted if the frame sampling step (S800) or the camera motion extraction step (S700) is omitted.

[0060] The additional frame 3D structuring step (S900) is to perform 3D structuring on the shots where there a camera motion using the frames sampled in the frame sampling step (S800). The additional frame 3D structuring step (S900) further utilizes the images of various angles / views added by the frame sampling step (S800), thereby obtaining more accurate information than the 3D information extracted in the 3D structuring step (S600). In this context, the 3D information may include camera intrinsic parameters, camera external parameters, camera 3D position, camera rotation information, and object 3D location information.

[0061] Meanwhile, in an embodiment of the present invention, if the additional frame 3D structuring step (S900) is performed for one video, it is possible to restore each grouped space / place in the video in the form of a point cloud.

[0062] The rendering step (S1000) performs rendering based on the 3D information for each shot. In an embodiment of the present invention, neural rendering, represented by NeRF (Neural Radiance Field), is applied to output a rendered result image as if the place / space for the shot has been reconstructed in 3D on a typical display when viewed from any direction. The rendered result images can be represented as shown in FIGS. 13 and 14.

[0063] Although the preceding detailed description has been primarily focused on the field of video content production, the 3D space reconstruction method and system described earlier can also be applied in various fields, not necessarily limited to video content production and consumption.

[0064] For example, in the field of multi-screen cinemas where multiple screens installed on the front and side walls of the cinema are used as projection surfaces to show film content or in the fiend of creating film content for multi-screen cinemas, the present invention can be utilized to create side videos (often referred to as "Screen X wing content" in this field). Specifically, it allows for rapid exploration of shots taken at a specific location from a large number of captured videos before editing and facilitates 3D reconstruction of the space corresponding to that location. This not only aids in understanding the space but also allows the use of information obtained from the reconstructed space when creating side videos, thereby improving the efficiency of the side video production process.

[0065] Furthermore, the present invention can be used for purposes such as revisiting a specific location where content was filmed, or creating production references for other video content. For example, if there is a need to recall a location that appeared in existing video content during the production of a variety show, drama, or film, or if a return visit or preliminary visit is necessary, or if a 3D image of a specific location is needed, the 3D reconstructed space can be used as a production reference, and this reference can be archived in a database and utilized as a reference for future production.

[0066] In addition, the present invention can also be used to reproduce a stadium space. In particular, by performing 3D reconstruction of the space based on sports event videos filmed abroad or videos of past sports events, it is possible to provide users watching the sports event with a three-dimensional sense of presence.

[0067] The embodiments of the present invention disclosed in the specification and the drawings are merely specific examples provided to easily explain the technical features of the invention and to aid in understanding the present invention, and are not intended to limit the scope of the present invention. It is obvious to those skilled in the art that other modifications based on the technical idea of the present invention can be made beyond the embodiments disclosed herein.

Claims

1.A method for reconstructing a 3D space from a content video, the method comprising:(a) a shot segmentation step for segmenting a video into shots;(b) a shot scale classification step for classifying the shots segmented in the shot segmentation step according to their scales;(d) a visual feature extraction step for extracting visual features for each shot classified in the shot scale classification step;(e) a similar location grouping step for grouping the shots taken at similar locations based on the visual features for each shot; and(f) a 3D structuring step for extracting 3D information about the shots grouped in the similar location grouping step.2.The method for reconstructing a 3D space from a content video according to claim 1, further comprising, after step (e),(j) a rendering step for performing rendering based on the 3D information for each shot.3.The method for reconstructing a 3D space from a content video according to claim 1, further comprising, if there is any shot classified as a close-up shot in step (b),(c) a close-up shot removal step for deleting the shot classified as the close-up shot.4.The method for reconstructing a 3D space from a content video according to claim 1, wherein step (d) through step (f) use a first frame of the shot instead of the entire shot.5.The method for reconstructing a 3D space from a content video according to claim 1, wherein step (e) is to perform the grouping by an agglomerative clustering algorithm.6.The method for reconstructing a 3D space from a content video according to claim 1, wherein step (f) is to extract 3D information including camera intrinsic parameters for each shot.7.The method for reconstructing a 3D space from a content video according to claim 6, further comprising, after step (f),(g) a camera motion extraction step for calculating a camera motion for each shot based on the 3D information.8.The method for reconstructing a 3D space from a content video according to claim 7, further comprising, after step (g),(h) a frame sampling step for sampling and adding frames of a shot where there is a camera motion.9.The method for reconstructing a 3D space from a content video according to claim 8, further comprising, after step (h),(i1) a 3D structuring step for performing 3D structuring on the shots where there a camera motion using the frames sampled in step (h).10.The method for reconstructing a 3D space from a content video according to claim 8, further comprising, after step (h),(i2) an additional frame 3D structuring step for performing 3D structuring on the sampled and added frames in step (h) to extract 3D information.11.The method for reconstructing a 3D space from a content video according to claim 9, further comprising, after step (i1),(j) a rendering step for performing rendering based on the 3D information for each shot.12.The method for reconstructing a 3D space from a content video according to claim 10, further comprising, after step (i2),(j) a rendering step for performing rendering based on the 3D information for each shot.13.A system for reconstructing a 3D space from a content video, the system comprising a control unit and a memory,wherein the control unit executes instructions for executing a method for reconstructing a 3D space from a content video stored in the memory,wherein the method comprising:(a) a shot segmentation step for segmenting a video into shots;(b) a shot scale classification step for classifying the shots segmented in the shot segmentation step according to their scales;(d) a visual feature extraction step for extracting visual features for each shot classified in the shot scale classification step;(e) a similar location grouping step for grouping the shots taken at similar locations based on the visual features for each shot; and(f) a 3D structuring step for extracting 3D information about the shots grouped in the similar location grouping step.

Citation Information

Patent Citations

  • Method and system for converting two-dimensional video into three-dimensional video

    CN102724531A

  • Shot segmentation method

    CN103578094A

  • Content-based three-dimensional CG animation search method and device

    CN111666447A

  • Camera offset detection method and device, electronic equipment and storage medium

    CN112652021A

  • Stereoscopic video generation method based on 3D convolution neural network

    US20190379883A1