Method and system for reconstructing 3D space from content video

CN122743522APending Publication Date: 2026-09-11CJ OLIVENETWORKS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480087676.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-11-07
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

根据该传统方法,在这样的环境中有可能获取诸如人和物体的对象的数据,但是以这种方式无法构造用于地点/空间的数据

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122743522A_ABST
    Figure CN122743522A_ABST
Patent Text Reader

Abstract

This invention relates to a method for reconstructing 3D space from content video, which allows for 3D reconstruction of locations / spaces appearing within captured video content, and the method for reconstructing 3D space from content video according to an embodiment of the invention includes: a shot segmentation step, a shot scale classification step, a close-up removal step, a visual feature extraction step, a similar location grouping step, a 3D construction step, a camera motion extraction step, a frame sampling step, an additional frame 3D construction step, and a rendering step.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for reconstructing 3D space from video content, which allows for 3D reconstruction of locations / spaces appearing within captured video content. Background Technology

[0002] Recently, 3D models for objects, people, spaces, etc. have been rapidly developed and are being used in various fields such as film, entertainment, theater, advertising, education, news, games, sports, performances, etc., to provide a high level of immersion and presentation.

[0003] Developing 3D models from scratch requires significant time and resources, and recreating all objects without reference materials is a highly challenging process. Therefore, extensive research has been conducted on 3D reconstruction of objects / spaces / people from images / videos. Furthermore, efforts are underway to reconstruct 3D models of objects appearing in existing content such as photographs, films, plays, entertainment performances, and shows.

[0004] In the past, creating 3D models required expensive equipment, such as 3D scanners, or data acquisition in a studio environment where multiple cameras were fixed in different locations. This traditional approach allowed for the acquisition of object data, such as people and objects, but it was impossible to construct location / space data in this way. Furthermore, obtaining multi-view data for location / space typically required equipment such as targets; however, achieving 3D reconstruction based on video content not initially intended for 3D reconstruction is extremely difficult. Specifically, video content not initially intended for 3D reconstruction was captured and edited by multiple cameras, resulting in a large number of shots (in frames per second). Moreover, 3D reconstruction requires multi-view data for a specific space, but it is nearly impossible to isolate that specific space from such video content and then acquire multi-view data.

[0005] This invention addresses the problems of conventional methods, and relates to a method for 3D position and space reconstruction based on content video with motion in multiple spaces and obtained on various scales, as well as a system for using this method. Summary of the Invention

[0006] Technical issues

[0007] Therefore, the present invention was made to solve the above-mentioned problems, and the object of the present invention is to provide a method for reconstructing 3D space from content video, which allows for location / space 3D reconstruction from content video with motion in multiple spaces, acquired and edited at various scales.

[0008] Another object of the present invention is to provide a method for reconstructing 3D space that allows the 3D reconstruction of location / space to be immediately used in the metaverse and similar platforms, thereby reducing the time and cost of creating new location / space and resulting in increased productivity.

[0009] Another object of the present invention is to provide a method for reconstructing 3D space, which can construct a database of 3D reconstructed locations / spaces appearing in content and allow viewing of these 3D reconstructed locations / spaces stored in the database from different viewpoints.

[0010] The objectives of this invention are not limited to those mentioned above, and those skilled in the art will clearly understand from the following description other objectives not mentioned.

[0011] Technical solution

[0012] To achieve the above-mentioned objectives of the present invention, a method for reconstructing 3D space from content video is provided, the method comprising: a shot segmentation step; a shot scale classification step; a close-up removal step; a visual feature extraction step; a similar position grouping step; a 3D construction step; a camera motion extraction step; a frame sampling step; an additional frame 3D construction step; and a rendering step.

[0013] Beneficial effects

[0014] According to the present invention, shots acquired in the same space can be grouped to allow 3D reconstruction of location / space based on content videos acquired and edited at various scales with motion in multiple spaces, and also allows 3D reconstruction of location / space appearing in acquired video content without any intention for 3D reconstruction.

[0015] Furthermore, according to the present invention, it is possible to 3D reconstruct locations / spaces in existing content for immediate use in virtual and augmented reality, metaverse, and similar platforms, thereby reducing the time and cost of creating new locations / spaces and thus improving productivity.

[0016] Furthermore, according to the present invention, a database of 3D reconstructed locations / spaces appearing in content can be constructed, and these 3D reconstructed locations / spaces stored in the database can be viewed from different viewpoints. Therefore, the present invention provides the advantage of allowing users to draw captured images of similar environments without needing to revisit the same locations.

[0017] The effects of the present invention are not limited to those described above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. Attached Figure Description

[0018] Figure 1This is a block diagram illustrating a system according to an embodiment of the present invention.

[0019] Figure 2 This is a block diagram illustrating an AI module according to an embodiment of the present invention.

[0020] Figure 3 This is a flowchart illustrating a method for reconstructing 3D space from content video according to an embodiment of the present invention.

[0021] Figure 4 This is a diagram illustrating the scaling classification steps according to an embodiment of the present invention.

[0022] Figure 5 This is a diagram illustrating a similar position grouping step according to an embodiment of the present invention.

[0023] Figures 6 to 8 This is a diagram illustrating the Structure from Motion (SfM) algorithm according to an embodiment of the present invention.

[0024] Figures 9 to 12 This is a diagram illustrating the camera motion extraction steps according to an embodiment of the present invention.

[0025] Figure 13 and Figure 14 This is a diagram illustrating a rendered image produced according to an embodiment of the present invention. Detailed Implementation

[0026] The advantages and features of the present invention, as well as the methods for implementing them, will become clear with reference to the following detailed description of the embodiments and the accompanying drawings. However, the invention is not limited to the embodiments disclosed below and can be implemented in various forms. These embodiments are provided to ensure that the disclosure of the invention is complete and to fully inform those skilled in the art about the scope of the invention, which is defined only by the scope of the claims. Throughout the specification, the same reference numerals denote the same parts.

[0027] When a component is referred to as "connected to" or "linked to" another component, it covers both cases where it is directly connected to another component or linked to another component with intermediate components present. Conversely, when a component is referred to as "directly connected to" or "directly linked to" another component, it indicates that there are no components in between. The term "and / or" includes each of the items mentioned and all possible combinations thereof.

[0028] The terminology used herein is intended to describe these embodiments and is not intended to limit the invention. As used herein, singular forms also include plural forms unless specifically stated otherwise in the context. The terms “comprising” and / or “including” as used herein do not exclude the addition of one or more other components, steps, operations, and / or elements.

[0029] Although the terms "first," "second," etc., are used to describe various components, these components are not limited by these terms. These terms are only used to distinguish one component from another. Therefore, it is self-evident that a component referred to as "first" in the following description can also be a "second" component within the technical scope of this invention.

[0030] Throughout the specification, when a section is referred to as "comprising" a component, it means that it may exclude other components, but does not exclude other components, unless otherwise specifically stated. Furthermore, throughout the specification, the term "on" refers to being located above or below the target, and does not necessarily mean being above relative to the direction of gravity.

[0031] As used here, directional terms such as “front,” “back,” “left,” “right,” “vertical,” and “horizontal” are used in a relative sense for ease of explanation and can vary depending on the viewing direction.

[0032] Unless otherwise defined, all terms used herein (including technical and scientific terms) are to have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Furthermore, terms defined in commonly used dictionaries are not to be ideally or over-interpreted unless explicitly and specifically defined.

[0033] In the following, preferred embodiments of the invention will be described in detail with reference to the accompanying drawings. It should be noted that, wherever possible, the same components are represented by the same reference numerals in the drawings. Detailed descriptions of well-known functions and configurations that may obscure the essence of the invention will be omitted. For the same reason, some components are exaggerated, omitted, or shown schematically in the drawings.

[0034] Figure 1 This is a block diagram illustrating a system according to an embodiment of the present invention. Figure 2 This is a block diagram illustrating an AI module according to an embodiment of the present invention. Figure 3 This is a flowchart illustrating a method for reconstructing 3D space from content video according to an embodiment of the present invention. Figure 4 This is a diagram illustrating the scaling classification steps according to an embodiment of the present invention. Figure 5 This is a diagram illustrating a similar position grouping step S500 according to an embodiment of the present invention. Figures 6 to 8This is a diagram illustrating the structure derived from the motion (SfM) algorithm according to an embodiment of the present invention. Figures 9 to 12 This is a diagram illustrating the camera motion extraction step S700 according to an embodiment of the present invention, and Figure 13 and Figure 14 This is a diagram illustrating a rendered image produced according to an embodiment of the present invention.

[0035] Figure 1 The system illustrated may include an input unit 120, a control unit 110, a first memory 140, an output unit 130, etc. Figure 1 The components illustrated herein are not required for the implementation of the device, and the system or device described herein may have more or fewer components than those listed above.

[0036] The first memory 140 stores data that supports various functions of the system. The first memory 140 may store multiple applications or programs running on the system, as well as data and instructions used for system operation. In addition, applications may be stored in the first memory 140, installed on the system, and executed by the control unit 110 to perform system operations (or functions).

[0037] Control unit 110 typically controls the overall operation of the system and operations related to the application. Control unit 110 can provide or process appropriate information or functions by processing signals, data, information, etc., input or output through the components discussed above, or by running the application stored in the first memory 140. Furthermore, control unit 110 can control references... Figure 1 At least some of the components described are used to run the application stored in the first memory 140. Furthermore, the control unit 110 can operate in combination at least two of the components included in the system to run the application.

[0038] At least some of the aforementioned components can cooperate with each other to implement the operation, control, or control method of the system according to the various embodiments described below. Furthermore, by running at least one application program stored in the first memory 140, the operation, control, or control method of the system can be implemented on an electronic device.

[0039] Figure 2 This is a block diagram illustrating an AI module according to an embodiment of the present invention. In this embodiment, the AI ​​module performs the operations and calculations required in the shot segmentation step S100 to the rendering step S1000 of the present invention. The AI ​​module may include an electronic device capable of performing AI processing or a server including the AI ​​module. Furthermore, the AI ​​module may be included as... Figure 1At least a portion of the components of the illustrated system are used together to perform at least a portion of the AI ​​processing. The AI ​​module may include an AI processor 210, a second memory 230, and / or a communication unit 220. In embodiments of the invention, the AI ​​module may be implemented as a computing device capable of learning neural networks, and may be implemented as various electronic devices such as servers, desktop PCs, laptop PCs, tablet PCs, etc. The AI ​​processor 210 may use a program stored in the second memory 230 to learn the neural network.

[0040] Figure 3 This is a flowchart illustrating the steps of a method for reconstructing 3D space from content video according to an embodiment of the present invention. It is merely a preferred embodiment for achieving the objectives of the present invention, and some steps may be added or deleted as needed, and furthermore, one step may be included in another step.

[0041] Reference Figure 3 According to embodiments of the present invention, a method for reconstructing 3D space from content video may include a shot segmentation step S100, a shot scaling classification step S200, a close-up shot removal step S300, a visual feature extraction step S400, a similar position grouping step S500, a 3D construction step S600, a camera motion extraction step S700, a frame sampling step S800, an additional frame 3D construction step S900, and a rendering step S1000.

[0042] This invention relates to a database for constructing 3D reconstructions of locations / spaces appearing within a video based on a video, and a method for creating images from new viewpoints for locations / spaces in the video from the database.

[0043] Furthermore, it should be understood that the type of video used for 3D reconstruction in this invention can vary, but in a preferred embodiment, video content such as movies, plays, and entertainment performances is used to perform 3D reconstruction. This aims to construct a database of locations and spaces for 3D reconstruction using features of video content such as movies, plays, and entertainment performances. These locations and spaces are acquired and edited by multiple cameras in various ways across different spaces, where individual shots are acquired using various types of frame scales and camera movements, and where space is divided into multiple viewpoints. In this case, the video may be video that was not initially intended for 3D reconstruction.

[0044] Refer again Figure 3Shot segmentation step S100 involves dividing the video into shots. More specifically, shot segmentation step S100 divides the video into shots based on frame transitions. In this case, a shot can be a collection of frames from a segment of the video. Depending on the content or composition of the video, various criteria can exist for segmentation, but preferably, a criterion can be a location / space shown in the video where the video is taken.

[0045] The shot segmentation step S100 described above is designed to make it easier for the AI ​​model to analyze and understand the video. More specifically, in cases where videos, such as movies, dramas, and entertainment content, are captured by multiple cameras at different scales in different locations / spaces and then combined into a single video through editing, analyzing and understanding the video can be difficult for a single AI model. However, by first performing the shot segmentation step S100 to divide the video into shot units, it makes it easier for the AI ​​model to analyze the video.

[0046] Next, the lens scale classification step S200 can be performed. (Refer to...) Figure 4 The lens scale classification step S200 classifies the lenses segmented in the lens segmentation step S100 according to the lens scale. The lens scale classification step S200 classifies the lenses based on the scale of the video containing lenses obtained at various scales. At this time, the scale standard used for classifying the lenses can be defined in various ways. For example, such as... Figure 4 As shown, if the scale is defined based on the degree of proximity to a person or the area occupied by a person within the frame, the scale can be classified into extreme close-ups, close-ups, medium shots, full shots, long shots, etc. The scale does not necessarily have to be defined as the degree of proximity to a person; it can also be defined based on the degree of proximity to an object (i.e., the area it occupies within the frame) or based on the proportion of the background within the frame. The shot scale classification step S200 is performed to classify shots obtained under the close-up scale, which will be removed in the close-up removal step described later.

[0047] Following the shot scaling classification step, if any shot S200 is classified as a close-up in the shot scaling classification step, then the close-up removal step S300 is performed to remove the shot classified as a close-up. In the case of close-ups, it is difficult to obtain information about location / space within the video because it is a scene focused on a single object within the video; however, performing this step to remove close-ups enhances the AI ​​model's ability to obtain information about location / space, i.e., thereby reducing the computational load of the AI ​​model and improving computational efficiency. Alternatively, in this invention, if no shot is classified as a close-up, the close-up removal step S300 can also be omitted.

[0048] Brief reference Figure 4 The shot scale classification step S200 and the close-up shot removal step S300 will now be described as examples. In an embodiment of the invention, the shot scale classification step can classify shots included in the video into five scales: extreme close-up, close-up, medium shot, full shot, and long shot. In the close-up shot removal step S300, shots corresponding to extreme close-ups and close-ups can be removed.

[0049] The visual feature extraction step S400 extracts visual features for each shot classified in the shot scaling classification step S200. Performing the visual feature extraction step S400 reduces the computational cost required for the feature point matching algorithm executed in the subsequent Structure from Motion (SfM) step. If the feature point matching algorithm were performed on all shots except those removed in the close-up removal step S300, the feature point matching algorithm performed on each image pair would require a large amount of computation. Therefore, in this embodiment of the invention, a candidate set of the most similar shots is first classified, and then the feature point matching algorithm is performed on the classified shots. The visual feature extraction step S400 extracts features from the shots to classify the most similar shots.

[0050] Furthermore, in embodiments of the present invention, a visual feature extraction step S400 can be performed to extract visual features by selecting only the first frame of each shot. This approach is based on the premise that a single shot may be taken from the same location, and that even if shots are taken from different angles, the background information taken from the same location remains visually similar. Extracting visual features from each frame within a shot increases computation time and requires significant computational resources. Therefore, the present invention can be implemented to extract visual features by selecting only the first frame of each shot.

[0051] Following the visual feature extraction step, a similar location grouping step S500 is performed to group the shots taken at similar locations based on the visual features extracted for each shot in the visual feature extraction step. More specifically, in the similar location grouping step S500, a clustering algorithm is performed based on the visual features extracted for each shot in the visual feature extraction step to group the shots taken at similar locations.

[0052] Furthermore, according to an embodiment of the present invention, the similar location grouping step S500 can be performed using a clustering algorithm (a type of hierarchical clustering). More specifically, in an embodiment of the present invention, if the clustering algorithm is applied based on the visual features extracted for each shot in the visual feature extraction step S400, it can be performed as follows: Figure 5 The smallest unit shown initiates clustering, and then connections between clusters are created hierarchically based on the similarity of visual features of each shot. Figure 5 In the diagram, the horizontal axis represents the index of the small unit cluster, and the vertical axis represents the similarity between clusters.

[0053] Refer again Figure 3 The 3D construction step S600 extracts 3D information about the shots grouped in the similar position grouping step. More specifically, this step applies the Structure from Motion (SfM) algorithm to shots classified as taken at the same location in the shot segmentation step S100 to the similar position grouping step S500. The Structure from Motion (SfM) algorithm is a method for extracting 3D information about an object from 2D images taken from various angles of the same object. This method enables the calculation of camera intrinsic parameters such as camera focal length, camera extrinsic parameters including camera 3D position and camera rotation information, and object 3D position information.

[0054] Now refer to Figures 6 to 8 A more detailed description of the structure comes from the motion (SfM) algorithm. For example... Figure 6 As shown, the SfM algorithm reconstructs 3D information from images at the same location. Figure 7 As shown, a feature point matching algorithm is performed to match the correspondence between visual features (feature points) extracted from one shot and visual features (feature points) extracted from another shot taken at the same location (the lines connecting feature points in the left image to feature points in the right image). Figure 8 As shown, 3D information about a location / space is obtained by calculating the geometric relationship of the location / space on the lens from the correspondence of multiple shots at the same location / space.

[0055] Furthermore, in an embodiment of the present invention, the first frame of the shot is used instead of the entire shot in the visual feature extraction steps S400 to the 3D construction steps S600. This approach aims to reduce the computational load required in the visual feature extraction steps S400 to the 3D construction steps S600, since typical videos consist of a very large number of frames.

[0056] The camera motion extraction step S700, performed after the 3D construction step, calculates camera motion for each shot based on the 3D information obtained in the 3D construction step S600. The video used in embodiments of the invention can include various types of camera motion. For example, in embodiments of the invention, the video can include: shots taken by a fixed camera, where only the object or body moves while the background remains stationary; shots taken by a camera that simply moves up, down, left, or right; shots taken by a camera mounted on a drone covering a wide area or in dynamic motion; and so on. In this case, the camera motion extraction step S700 applies a Simultaneous Localization and Mapping (SLAM) algorithm based on the camera-specific parameters extracted in the 3D construction step S600 to perform camera motion extraction for each shot.

[0057] Now refer to Figures 9 to 12 The camera motion extraction step S700 according to an embodiment of the present invention is described in more detail. In embodiments of the present invention, camera motion displayed in three dimensions can be as follows: Figure 9 and Figure 11 As shown in the figure. Figure 10 This is an example Figure 9 A diagram showing the camera's movement in two dimensions. Figure 12 This is an example Figure 11 The image shows the camera's movement in two dimensions. Additionally, Figure 9 and Figure 10 An example is given of acquiring images by rotating and moving the camera, and Figure 10 The lines shown represent the movement of the camera. Furthermore, Figure 11 and Figure 12 An example is given of an image acquired by a fixed camera, and... Figure 12 The points shown represent images captured by a fixed camera. In this invention, camera motion is quantified using the most commonly used metric for calculating camera trajectory error. Motion values ​​can be quantized when the ground truth (GT) is set to 0. Figure 9 and Figure 11 The terms used in the graph can be explained as follows: ATE (Absolute Trajectory Error) is a simple measure of the error accumulated from the differences in x, y, and z values ​​of GT, and does not reflect gradual deviations along the sequence; and RPE (Relative Pose Error) is a measure that separately measures translation and rotation errors and accumulates relative position errors, compensating for the shortcomings of ATE.

[0058] Furthermore, in this invention, if the camera captures shots included in the video while remaining stationary, the camera motion extraction step S700 can be omitted. Additionally, in this embodiment, only the first frame of the shot can be used to perform the various steps of this invention.

[0059] In addition, in this invention, for shots captured by a fixed camera, only the first frame can be used to perform the various steps of the invention; however, for shots with motion, further 3D information can be extracted by further sampling the frames.

[0060] Frame sampling step S800 involves sampling and adding frames of shots where camera motion exists. In this case, frame sampling step S800 is performed based on the camera motion for the shot extracted in camera motion extraction step S700.

[0061] Furthermore, in an embodiment of the present invention, the frame sampling step S800 performs 3D construction on the sampled and added frames to extract 3D information. Additionally, in an embodiment of the present invention, the frame sampling step S800 can be implemented to perform 3D construction only on the sampled and added frames to extract 3D information.

[0062] Additionally, if the frame sampling step S800 or the camera motion extraction step S700 is omitted, the additional frame 3D construction step S900 can be omitted.

[0063] The additional frame 3D construction step S900 performs 3D construction on a lens where camera motion exists, using the frames sampled in the frame sampling step S800. The additional frame 3D construction step S900 also utilizes images from various angles / views added by the frame sampling step S800 to obtain more accurate 3D information than that extracted in the 3D construction step S600. In this context, the 3D information may include camera intrinsic parameters, camera extrinsic parameters, camera 3D position, camera rotation information, and object 3D position information.

[0064] Furthermore, in an embodiment of the present invention, if the additional frame 3D construction step S900 is performed on a video, the various grouped spaces / locations in the video can be recovered in the form of a point cloud.

[0065] Rendering step S1000 performs rendering based on 3D information for each shot. In embodiments of the invention, neural rendering, represented by NeRF (Neural Radiance Field), is applied to output the rendered image, as if the position / space of the shot were reconstructed in 3D on a typical display when viewed from any direction. The rendered image can be as follows: Figure 13 and Figure 14 As shown in the figure.

[0066] Although the detailed description above focuses primarily on the field of video content production, the 3D spatial reconstruction methods and systems described above can also be applied to various fields, not limited to video content production and consumption.

[0067] For example, in the field of multi-screen cinemas, where multiple screens mounted on the front and side walls of a cinema are used as projection surfaces to display film content, or in the field of creating film content for multi-screen cinemas, this invention can be used to create side-view videos (often referred to in this field as "screen-to-side content"). Specifically, it allows for the rapid detection of shots taken at specific locations from a large volume of captured video before editing, and facilitates 3D reconstruction of the space corresponding to that location. This not only aids in understanding the space but also allows for the use of information obtained from the reconstructed space when creating side-view videos, thereby improving the efficiency of the side-view video production process.

[0068] Furthermore, this invention can be used for purposes such as revisiting specific locations where content was acquired, or creating production references for other video content. For example, if it is necessary to recall locations appearing in existing video content during the production of various performances, plays, or films, or if a revisit or initial visit is required, or if 3D images of specific locations are needed, the 3D reconstruction space can be used as a production reference, and this reference can be archived in a database and used as a reference for future productions.

[0069] Furthermore, this invention can also be used to recreate stadium spaces. Specifically, by performing 3D reconstruction of the space based on videos of sporting events obtained abroad or videos of past sporting events, a three-dimensional sense of presence can be provided to users watching sporting events.

[0070] The embodiments of the invention disclosed in the specification and drawings are merely specific examples provided to facilitate the explanation of the technical features of the invention and to aid in understanding the invention, and are not intended to limit the scope of the invention. It will be apparent to those skilled in the art that other modifications based on the technical concept of the invention can be made in addition to the embodiments disclosed herein.

Claims

1. A method for reconstructing 3D space from content video, the method comprising: (a) Shot segmentation step, used to segment the video into shots; (b) Lens scale classification step, used to classify the lenses segmented in the lens segmentation step according to the lens scale; (d) Visual feature extraction step, used to extract visual features for each lens classified in the lens scaling step; (e) A similar location grouping step, used to group shots taken at similar locations based on the visual features for each shot; and (f) A 3D construction step for extracting 3D information about the lenses grouped in the similar location grouping step.

2. The method for reconstructing 3D space from content video according to claim 1, the method further comprising, after step (e): (j) Rendering step, used to perform rendering for each shot based on the 3D information.

3. The method for reconstructing 3D space from content video according to claim 1, further comprising, if any shot classified as a close-up exists in step (b), (c) Close-up shot removal step, used to remove the shot that is classified as the close-up shot.

4. The method for reconstructing 3D space from content video according to claim 1, wherein, Steps (d) through (f) use the first frame of the shot instead of the entire shot.

5. The method for reconstructing 3D space from content video according to claim 1, wherein, Step (e) is to perform the grouping using an aggregation clustering algorithm.

6. The method for reconstructing 3D space from content video according to claim 1, wherein, Step (f) involves extracting 3D information, including the inherent parameters of the camera, for each lens.

7. The method for reconstructing 3D space from content video according to claim 6, the method further comprising, after step (f): (g) Camera motion extraction step, used to calculate camera motion for each lens based on the 3D information.

8. The method for reconstructing 3D space from content video according to claim 7, the method further comprising, after step (g), (h) Frame sampling step, used to sample and add frames of the shot when there is camera movement.

9. The method for reconstructing 3D space from content video according to claim 8, the method further comprising, after step (h), (i1) 3D construction step, used to perform 3D construction on the lens with camera motion using the frames sampled in step (h).

10. The method for reconstructing 3D space from content video according to claim 8, the method further comprising, after step (h), (i2) Additional frame 3D construction step, used to construct 3D on the frames sampled and added in step (h) to extract 3D information.

11. The method for reconstructing 3D space from content video according to claim 9, the method further comprising, after step (i1), (j) Rendering step, used to perform rendering for each shot based on the 3D information.

12. The method for reconstructing 3D space from content video according to claim 10, the method further comprising, after step (i2): (j) Rendering step, used to perform rendering for each shot based on the 3D information.

13. A system for reconstructing 3D space from content video, the system comprising a control unit and a memory, in, The control unit executes instructions for performing a method for reconstructing 3D space from video content stored in the memory. The method includes the following steps: (a) Shot segmentation step, used to segment the video into shots; (b) Lens scale classification step, used to classify the lenses segmented in the lens segmentation step according to the lens scale; (d) Visual feature extraction step, used to extract visual features for each lens classified in the lens scaling step; (e) A similar location grouping step, used to group shots taken at similar locations based on the visual features for each shot; and (f) A 3D construction step for extracting 3D information about the lenses grouped in the similar location grouping step.