Method and apparatus for reproducing dynamic scene as 3D model using video

The method and apparatus convert dynamic scenes from videos into 3D models with dynamic movement, addressing the limitations of existing technologies by eliminating the need for specialized personnel and enabling cost-effective, time-efficient creation of immersive 3D content.

WO2025116689A1PCT designated stage expired Publication Date: 2025-06-05CJ OLIVENETWORKS +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/096385
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-10-22
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing image-based neural rendering technologies can only convert static objects to 3D models, requiring additional work and specialized personnel to animate and add effects, which is costly and time-consuming.

Method used

A method and apparatus that capture a dynamic scene as a video and directly convert it into a 3D model with dynamic movement, classifying elements within the video into static and dynamic areas, and converting each element into a 3D model based on its characteristics.

Benefits of technology

Enables the creation of 3D content with dynamic movement from videos without the need for specialized personnel, reducing costs and time while providing a new type of video content that can be viewed from any viewpoint.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096385_05062025_PF_FP_ABST
    Figure KR2024096385_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for reproducing a dynamic scene as a 3D model using a video, and an apparatus therefor, and more specifically, to a method and apparatus for reproducing a dynamic scene as a 3D model, which can directly convert such a model to 3D content with a dynamic movement without the need for specialized personnel.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR REPRODUCING DYNAMIC SCENE AS 3D MODEL USING VIDEO

[0001] The present invention relates to a method for reproducing a dynamic scene as a 3D model using a video, and an apparatus therefor, and more specifically, to a method and apparatus for reproducing a dynamic scene as a 3D model, which can directly convert such a model to 3D content with a dynamic movement without the need for specialized personnel.

[0002] With the advancement of the metaverse and VR / AR industries, 3D models for objects, people, and spaces are being developed rapidly, and these models are utilized across various fields such as movies, entertainment shows, dramas, advertisements, education, news, games, sports, performances, etc. to provide a high level of immersion and a vivid sense of realism. Creating a 3D model typically involves securing the design of a subject and then producing the 3D model using 3D modeling software in a labor-intensive manner by skilled professionals. Recently, technologies utilizing deep learning have been actively researched to convert images of motionless objects, people, spaces, and scenes to 3D models, respectively.

[0003] Meanwhile, in industries such as metaverse, VR / AR, gaming, etc. that utilize 3D models, there is a need not only for static objects but also for dynamic ones; and existing image-based neural rendering (3D conversion technology using deep learning) is limited to converting static objects to 3D models. To apply such static 3D models created in this way to content with various movements, additional works are required to animate and add effects to these models. However, these additional works require specialized personnel, which incur significant costs, and a lot of time is required to achieve natural movements.

[0004] The present invention has been made to address these challenges and provides a method and apparatus for reproducing a dynamic scene as a 3D model using a video, which can directly convert such a model to 3D content with a dynamic movement without the need for specialized personnel.

[0005] An object of the present invention is to capture a dynamic situation as a video and convert the dynamic scene to a 3D model using the captured video.

[0006] Another object of the present invention is to directly convert such a model to 3D content with a dynamic movement without the need for specialized personnel.

[0007] Still another object of the present invention is to classify elements contained in a scene within a video.

[0008] Yet another object of the present invention is to convert each element to a 3D model based on the characteristics of each classified element.

[0009] A further object of the present invention is to generate a video by combining 3D models for each element.

[0010] A still further object of the present invention is to segment elements contained in a video into a static area and a dynamic area and convert them to a 3D model.

[0011] A yet further object of the present invention is to provide a new type of video content and service that can be viewed from a free viewpoint, rather than the camera's direction when a desired scene was captured.

[0012] To accomplish the above objects of the present invention, there is provided a method for reproducing a dynamic scene as a 3D model using a video, the method comprising the steps of: receiving a first video that is to be reproduced as a 3D model from a user; based on elements contained in a scene within the first video, classifying the elements using a predetermined classification method; based on the characteristics of each element, converting each element to a 3D model; generating a second video by combining the 3D models for each element; and providing the second video to the user.

[0013] According to the method for reproducing a dynamic scene as a 3D model using a video, the step of classifying the elements may comprise the steps of: classifying the elements in the scene within the first video into a first element containing a static element and a second element containing a dynamic element; and if the second element includes a third element containing a lighting element, classifying the second element into the third element and a fourth element containing a moving element other than the third element.

[0014] Moreover, according to method for reproducing a dynamic scene as a 3D model using a video, the first element may be an element that does not move within the scene, including a non-moving object and background, and the second element may be an element whose position or brightness changes within the scene, including at least one of the third element containing a changing lighting element and a fourth element containing a moving object.

[0015] Furthermore, according to the method for reproducing a dynamic scene as a 3D model using a video, the step of classifying the elements in the scene within the first video into a first element containing a static element and a second element containing a dynamic element may comprise the steps of: segmenting the first video into a first area corresponding to a static area without a movement, including the first element; and segmenting the areas other than the first area within the first video into a second area corresponding to a dynamic area with a movement, including the second element.

[0016] In addition, according to the method for reproducing a dynamic scene as a 3D model using a video, the step of classifying the elements may comprise the step of segmenting the areas within the first video using a data learning-based method, including an optical flow extraction method based on deep learning, or using information within the first video that includes a change in frames.

[0017] Additionally, according to the method for reproducing a dynamic scene as a 3D model using a video, the step of classifying the second element into the third element and a fourth element containing a moving element other than the third element may comprise the steps of: segmenting the first video into the third area corresponding to the fourth element for every frame; and segmenting the areas other than the third area within the first video into the fourth area corresponding to the third area for every frame.

[0018] Moreover, according to the method for reproducing a dynamic scene as a 3D model using a video, the step of converting each element to a 3D model may comprise the step of: converting each of the first area corresponding to the first element, the fourth area corresponding to the third element, and the third area corresponding to the fourth element to a 3D model based on the characteristics of each element.

[0019] Furthermore, according to the method for reproducing a dynamic scene as a 3D model using a video, the step of generating a second video by combining the 3D models for each element may comprise the step of: generating the second video corresponding to the 3D model of the first video by combining the 3D models corresponding to each element.

[0020] In addition, according to the method for reproducing a dynamic scene as a 3D model using a video, the step of providing the second video to the user may comprise the step of: providing the second video to the user as video content that allows the first video to be viewed from a free viewpoint.

[0021] Meanwhile, an apparatus for reproducing a dynamic scene as a 3D model using a video according to an embodiment of the present invention may comprise: a video reception unit that receives a video from a user; a video classification unit that classifies elements contained in the video using a predetermined classification method; a 3D conversion unit that converts each element to a 3D model; a video generation unit that generates a video by combining the 3D models converted for each element; and a video provision unit that provides the video to the user.

[0022] Moreover, according to the apparatus for reproducing a dynamic scene as a 3D model using a video, the video reception unit may receive a first video that is to be reproduced as a 3D model from the user; based on the elements contained in a scene within the first video, the video classification unit may classify the elements using a predetermined classification method; based on the characteristics of each element, the 3D conversion unit may convert each element to a 3D model; the video generation unit may generate a second video by combining the 3D models for each element; and the video provision unit may provide the second video to the user.

[0023] Furthermore, according to the apparatus for reproducing a dynamic scene as a 3D model using a video, the video classification unit may classify the elements in the scene within the first video into a first element containing a static element and a second element containing a dynamic element, and if the second element includes a third element containing a lighting element, classify the second element into the third element and a fourth element containing a moving element other than the third element.

[0024] In addition, according to the apparatus for reproducing a dynamic scene as a 3D model using a video, the first element may be an element that does not move within the scene, including a non-moving object and background, and the second element may be an element whose position or brightness changes within the scene, including at least one of the third element containing a changing lighting element and a fourth element containing a moving object.

[0025] A system for reproducing a dynamic scene as a 3D model using a video according to another embodiment of the present invention may comprise: a memory in which at least one program is recorded; and a processor for executing the program, and the program may comprise instructions for carrying out the steps of: receiving a first video that is to be reproduced as a 3D model from a user; based on elements contained in a scene within the first video, classifying the elements using a predetermined classification method; based on the characteristics of each element, converting each element to a 3D model; generating a second video by combining the 3D models for each element; and providing the second video to the user.

[0026] According to the present invention, it is possible to capture a dynamic situation as a video and convert the dynamic scene to a 3D model using the captured video.

[0027] Moreover, according to the present invention, it is possible to directly convert such a model to 3D content with a dynamic movement without the need for specialized personnel.

[0028] Furthermore, according to the present invention, it is possible to classify elements contained in a scene within a video.

[0029] In addition, according to the present invention, it is possible to convert each element to a 3D model based on the characteristics of each classified element.

[0030] Additionally, according to the present invention, it is possible to generate a video by combining 3D models for each element.

[0031] Moreover, according to the present invention, it is possible to segment elements contained in a video into a static area and a dynamic area and convert them to a 3D model.

[0032] Furthermore, according to the present invention, it is possible to provide a new type of video content and service that can be viewed from a free viewpoint, rather than the camera's direction when a desired scene was captured.

[0033] Meanwhile, the effects of the present invention are not limited to those mentioned above, and other technical effects not mentioned will be clearly understood by those skilled in the art from the following description.

[0034] FIG. 1 is a diagram illustrating the overall flow of a method for reproducing a 3D model according to the present invention for a comprehensive understanding of the present invention.

[0035] FIG. 2 illustrates the steps of a method for reproducing a dynamic scene as a 3D model using a video according to the present invention.

[0036] FIG. 3 illustrates the steps of classifying elements according to the present invention.

[0037] FIG. 4 illustrates the steps of segmenting the areas for static elements and dynamic elements according to the present invention.

[0038] FIG. 5 illustrates the steps of segmenting the areas for dynamic elements according to the present invention.

[0039] FIG. 6 is a diagram illustrating an example of the results of segmenting the elements contained in scenes in a video and combining 3D models according to the present invention.

[0040] FIG. 7 is a diagram illustrating an example of a method for controlling the color of lighting as an exemplary application of the present invention.

[0041] FIG. 8 illustrates the detailed components of an apparatus for reproducing a dynamic scene as a 3D model using a video according to the present invention.

[0042] Details regarding the objects and technical features of the present invention and the resulting effects will be more clearly understood from the following detailed description based on the drawings attached to the specification of the present invention. Preferred embodiments according to the present invention will be described in detail with reference to the accompanying drawings.

[0043] The embodiments disclosed in this specification should not be construed or used to limit the scope of the present invention. It is obvious to those skilled in the art that the description, including the embodiments, of this specification has various applications. Therefore, any embodiments described in the detailed description of the present invention are illustrative to better illustrate the present invention and are not intended to limit the scope of the present invention to these embodiments.

[0044] The functional blocks shown in the drawings and described below are only examples of possible implementations. In other implementations, different functional blocks may be used without departing from the spirit and scope of the detailed description. Moreover, although one or more functional blocks of the present invention are shown as individual blocks, one or more of the functional blocks of the present invention may be a combination of various hardware and software components that perform the same functions.

[0045] Furthermore, the term "comprising" certain components, which is an "open-ended" term, simply refers to the presence of the corresponding components and should not be understood as excluding the presence of additional components.

[0046] In addition, if a specific component is referred to as being "connected to" or "coupled with" another component, it should be understood that it may be directly connected or coupled to another other component, but there may be other components therebetween.

[0047] FIG. 1 is a diagram illustrating the overall flow of a method for reproducing a 3D model for a dynamic scene using a video according to the present invention.

[0048] Referring to FIG. 1, the method for reproducing a dynamic scene as a 3D model using a video according to the present invention may first comprise the step of segmenting the elements in a scene within a video into static elements and dynamic elements. When any video is captured, the video may contain an element that moves over time (i.e., a dynamic element) and an element that do not move over time (i.e., a static element). For example, if a dancer dancing on a stage is captured using a fixed camera, the dancing dancer can be classified as a dynamic element, and the stage on which the dancer is dancing can be classified as a static element.

[0049] Meanwhile, the dynamic element within a video may contain various types of things. For example, in a video featuring a dancer, the dancing dancer and the movement of lighting directed to the dancer can be classified into the dynamic elements. Specifically, in the method for reproducing a 3D model according to the present invention, if a lighting element is contained as a dynamic element in the video, the method may comprise the step of segmenting the dynamic elements into a moving element and a lighting element. Since the respective elements that constitute a scene possess different properties, the step of distinguishing these elements is critical for creating an accurate 3D model that represents the combined scene of these elements well.

[0050] Next, each of the static element, the moving element, and the lighting element may be subject to a 3D modelling conversion. The 3D modelling conversion may refer to creating a 3D model that realistically represents the entire scene by reflecting the characteristics of the segmented static elements, moving elements, and lighting elements.

[0051] Lastly, the 3D models converted for each element can be combined to reproduce the 3D model of the video. Referring back to the example of the dancer video mentioned earlier, the stage, which is a static element, the dancing dancer, which is a dynamic element, and the lighting, which is another dynamic element, can each be converted to a 3D model and then combined to create a complete 3D modeling video.

[0052] In the foregoing, the overall flow of the method for reproducing a dynamic scene as a 3D model using a video according to the present invention has been discussed. Based on the overall flow, the method for reproducing a dynamic scene as a 3D model using a video will now be discussed in detail with reference to the drawings.

[0053] FIG. 2 illustrates the steps of the method for reproducing a dynamic scene as a 3D model using a video according to the present invention.

[0054] Referring to FIG. 2, the method for reproducing a dynamic scene as a 3D model using a video according to the present invention may first comprise the step of receiving a first video that is to be reproduced as a 3D model from a user (S210). Here, the first video may refer to a video featuring a dynamic scene. There may also be a video that captures only a static element; however, such a video will not be discussed in this detailed description, and it should be understood that the first video refers to a video containing a moving object, preferably a video containing a dynamic element.

[0055] After step S210, the method may comprise the step of, based on the elements contained in a scene within the first video, classifying the elements using a predetermined classification method (S220). This step is to classify the elements into static elements and dynamic elements in the video, as previously mentioned, which will be discussed in more detail with reference to FIG. 3.

[0056] FIG. 3 illustrates the steps of classifying elements according to the present invention. Referring to FIG. 3, the step of classifying elements in the method for reproducing a dynamic scene as a 3D model using a video according to the present invention may first comprise the step of classifying the elements in the scene within the first video into a first element containing a static element and a second element containing a dynamic element (S310).

[0057] Here, the first element may be an element that does not move within the scene, including a non-moving object and background, i.e., a static element. On the other hand, the second element may be an element whose position or brightness changes within the scene, including at least one of a third element containing a changing lighting element and a fourth element containing a moving object. That is, the second element may be a dynamic element with a movement of a subject (such as an object, a person, etc.). For segmenting the elements into static elements and dynamic elements, a data learning-based method, such as a deep learning-based optical flow extraction, or a method of using information in the video, such as using a change in frames, can be utilized.

[0058] FIG. 4 illustrates the above-described step S310 in more detail.

[0059] Referring to FIG. 4, the step of segmenting the areas for static elements and dynamic elements in the method for reproducing a dynamic scene as a 3D model using a video according to the present invention may first comprise the step of segmenting the first video into a first area corresponding to a static area without a movement, including the first element (S410).

[0060] Next, the method may comprise the step of segmenting the areas other than the first area within the first video into a second area corresponding to a dynamic area with a movement, including the second element (S420). The three-dimensional information for conversion to a 3D model is obtained from the static area, and if a dynamic element is contained in the static area, inaccurate three-dimensional information can be obtained. Therefore, it is necessary to clearly distinguish between static and dynamic areas.

[0061] In the foregoing description, it has been mentioned that the step of classifying the elements may comprise the step of segmenting the areas within the first video using a data learning-based method, including an optical flow extraction method based on deep learning, or using information within the first video that includes a change in frames.

[0062] In this case, the term "optical flow" refers to the movement pattern of an object in a video, indicating the pattern of an object movement that appears due to the movement of a camera or object between the previous frame and the next frame. It can be used to predict the movement of the object in the video.

[0063] The optical flow can also be used to store a video by compensating for hand shake, and when compressing the video using motion information derived from the optical flow, it is possible to compress a high-quality video using fewer bits.

[0064] For reference, in the optical flow, two assumptions are made: i) the pixel intensities of a moving object do not change between consecutive frames, and ii) neighboring pixels have similar motion. Moreover, there are two types of optical flow: one is called Sparse Optical Flow, which calculates only some pixels, and the other is called Dense Optical Flow, which calculates all pixels in the entire video.

[0065] Representative algorithms for optical flow include the Lucas-Kanade algorithm and the Farneback algorithm. The Lucas-Kanade algorithm utilizes the assumption that "neighboring pixels moves similarly." It calculates motion using a small window (3x3 patch) based on the assumption that neighboring pixels have similar motion. However, a drawback of this method is that it can encounter problems when the object motion is large due to the small window size. To overcome this drawback, an image pyramid can be used. The image pyramid makes it possible that a small motion is less noticeable as it goes toward the top (as the image gets smaller) and a large motion appears looks like a small motion. This allows for the detection of a large motion as well. The Lucas-Kanade algorithm typically calculates optical flow vectors at specific pixels by calculating the optical flow for user-specified feature points. In this case, the algorithm returns motion information for a few user-specified feature points, allowing the detection of where the object has moved in the next frame. Here, the feature points, often corners, can be specified and utilized for this purpose.

[0066] On the contrary, the Farneback algorithm calculates the optical flow for all pixels and uses all pixels in the entire video. Because the Farneback algorithm uses all pixels in the entire video, there is no need for the user to specify feature points. Therefore, using the Farneback algorithm allows for the detection of motion across the entire video.

[0067] Meanwhile, referring back to FIG. 3, after step S310, the method may comprise the step of reclassifying the dynamic elements for each type (S320). Specifically, in step S320, the previously classified second element can be further classified into a third element containing a lighting element and a fourth element containing a moving element other than the third element. Referring back to the dancer video mentioned earlier, the human movement and the change in lighting in the video continuously change over time, but the movement patterns of these objects have entirely different characteristics. These movement patterns can be a crucial criterion for distinguishing between the change in lighting (third element) and the human movement (fourth element).

[0068] FIG. 5 illustrates the step S320 divided into detailed steps. Referring to FIG. 5, the method of segmenting the areas for dynamic elements may comprise the steps of: segmenting the first video into the third area corresponding to the fourth element for every frame (S510); and then segmenting the areas other than the third area within the first video into the fourth area corresponding to the third area for every frame (S520).

[0069] For example, if the dynamic area contains two different types of elements, the dynamic element distinguished in the previous step can be further segmented. For instance, if a video contains both the human movement and the change in lighting, these two elements are the same in that both involves changes, but they have entirely different characteristics and thus need to be distinguished from each other. Therefore, by segmenting the video into the human area for each frame and excluding the human area from the dynamic area, the remaining area can be regarded as being affected by the change in lighting.

[0070] With reference to FIGS. 3 to 5, the methods for classifying the elements in a video into static element and dynamic element, as well as the methods for further classifying different types of dynamic elements have been discussed.

[0071] Referring back to FIG. 2, after step S220, the method may comprise the step of converting each element to a 3D model (S230). This step may comprise the step of converting each of the first area corresponding to the first element, the fourth area corresponding to the third element, and the third area corresponding to the fourth element to a 3D model based on the characteristics of each element.

[0072] Meanwhile, after step S230, the method may comprise the step of generating a second video by combining the 3D models for each element (S240). This step may comprise the step of generating the second video corresponding to the 3D model of the first video by combining the 3D models corresponding to each element.

[0073] For example, after the image-by-image segmentation of each element has been performed, the individual elements can be combined into a single scene while reflecting their individual characteristics, allowing the scene to be restored identical to the input video. Additionally, all elements can be rendered to be well-represented from a new camera direction or at any given timestamp.

[0074] FIG. 6 is a diagram illustrating an example of the results of segmenting the elements contained in scenes in a video and combining 3D models. In FIG. 6, the first row represents dynamic elements with a movement, the second row represents static elements without a movement, and the third row shows the results of combining the elements from the first and second rows in each column.

[0075] In FIG. 6, the four columns from the left show the results obtained using a conventional 3D model reproduction method, from which it can be observed that the dynamic elements and static elements are segmented using the conventional method, there are issues where each element appears blurred or is not accurately segmented.

[0076] On the contrary, the fifth column from the left in FIG. 6 shows the results of the 3D model reproduction method according to the present invention, from which it can be seen that the person, which is the dynamic element, and the background, which is the static element, are clearly segmented. For reference, the last column on the right in FIG. 6 shows the ground truth for the segmentation of each element.

[0077] Referring back to FIG. 2, after step S240, the method may comprise the step of providing the second video, in which the 3D models for each element have been combined (S250). The step of providing the second video may comprise the step of providing the second video to the user as video content that allows the first video to be viewed from a free viewpoint.

[0078] For example, converting to a 3D model means that the scene can be viewed from any direction. Therefore, when it is applied to previously recorded movies, dramas, entertainment shows, sports, and performance videos, it is possible to provide a new type of video content and service that can be viewed from a free viewpoint, rather than the camera's direction when a desired scene was captured.

[0079] Moreover, the converted 3D models can be used in new environments such as metaverse and VR / AR, and the second video can be used as-is within these new environments, or the individually separated elements can be used separately in different contexts.

[0080] Once converted to a 3D model, various combinations with other 3D models become possible. For example, if there is a specific product created as a 3D model, the corresponding product can be added to the scene. Furthermore, if a space model is created by capturing a manufacturing environment, it can be used for simulations of hazardous situations.

[0081] FIG. 7 is a diagram illustrating an example of a method for controlling the color of lighting as an exemplary application of the present invention.

[0082] Referring to FIG. 7, in the case of (a), it shows changing the color of a separated lighting element. By segmenting the lighting element from the dynamic elements, changing only the color of the lighting, and then recombining it with other elements, it can be reproduced as a 3D model.

[0083] In the case (b), it shows representing the dynamic movement with respect to a person while keeping the color of the lighting fixed. By segmenting the lighting element from the dynamic elements, fixing the color of the lighting, and then combining only the moving element of the person for each frame, it can be reproduced as a 3D model.

[0084] In the case (c), it shows changing the color of the lighting while keeping the person's movement fixed. By keeping the person's movement fixed in the dynamic elements, changing only the color of the lighting, and then recombining it with other elements for each frame, it can be reproduced as a 3D model.

[0085] FIG. 8 illustrates the detailed components of an apparatus for reproducing a dynamic scene as a 3D model using a video according to the present invention.

[0086] Referring to FIG. 8, a 3D model reproduction apparatus 100 may comprise a video reception unit 110, a video classification unit 120, a 3D conversion unit 130, a video generation unit 140, and a video provision unit 150.

[0087] The video reception unit 110 is configured to receive a video from a user and may receive a first video that is to be reproduced as a 3D model from the user. Here, the first video may refer to a video featuring a dynamic scene.

[0088] The video classification unit 120 is configured to classify elements contained in the video using a predetermined classification method and, based on the elements contained in a scene within the first video, may classify the elements contained in the video using the predetermined classification method.

[0089] The 3D conversion unit 130 is configured to convert each element to a 3D model and, based on the characteristics of each element, may convert each element to a 3D model.

[0090] The video generation unit 140 is configured to generate a video by combining the 3D models converted for each element and may generate a second video by combining the 3D models for each element.

[0091] The video provision unit 150 is configured to provide a video to a user and may provide the second video to the user.

[0092] To convert a video featuring a dynamic scene to a 3D model, the first step is to distinguish between static elements (such as a non-moving object, background, etc.) and dynamic elements (such as a moving object, a changing lighting element, etc.).

[0093] If the dynamic area contains two different types of elements, these two elements are further distinguished. For example, if there are other dynamic elements along with lighting, they are distinguished into lighting and other moving elements. Since the respective elements that constitute a scene possess different properties, the step of distinguishing these elements is critical for creating an accurate 3D model that represents the combined scene of these elements well. The 3D modelling conversion may refer to creating a 3D model that realistically represents the entire scene by reflecting the characteristics of the segmented static elements, moving elements, and lighting elements.

[0094] According to an embodiment, the video classification unit 120 may classify the elements in the scene within the first video into a first element containing a static element and a second element containing a dynamic element. Here, the first element may be an element that does not move within the scene, including a non-moving object and background. Moreover, the second element may be an element whose position or brightness changes within the scene, including at least one of a third element containing a changing lighting element and a fourth element containing a moving object. Furthermore, if the second element includes a third element containing lighting element, the video classification unit 120 may classify the second element into the third element and a fourth element containing a moving element other than the third element.

[0095] According to an embodiment, the video classification unit 120 may segment the areas for static elements and dynamic elements. When classifying the static elements and dynamic elements within the first video, the video classification unit 120 may segment the first video into a first area corresponding to a static area without a movement, including the first element. Moreover, the video classification unit 120 may segment the areas other than the first area within the first video into a second area corresponding to a dynamic area with a movement, including the second element.

[0096] The three-dimensional information for conversion to a 3D model is obtained from the static area, and if a dynamic element is contained in the static area, inaccurate three-dimensional information can be obtained. Therefore, it is necessary to clearly distinguish between static and dynamic areas.

[0097] According to an embodiment, the video classification unit 120 may segment the areas for each of the dynamic elements. The video classification unit 120 may segment the first video into the third area corresponding to the fourth element for every frame. Moreover, the video classification unit 120 may segment the areas other than the third area within the first video into the fourth area corresponding to the third area for every frame. For example, if the dynamic area contains two different types of elements, the dynamic element distinguished in the previous step can be further segmented. For instance, if a video contains both the human movement and the change in lighting, these two elements are the same in that both involves changes, but they have entirely different characteristics and thus need to be distinguished from each other. Therefore, by segmenting the video into the human area for each frame and excluding the human area from the dynamic area, the remaining area can be regarded as being affected by the change in lighting.

[0098] According to an embodiment, when converting each element into a 3D model, the 3D conversion unit 130 may convert each of the first area corresponding to the first element, the fourth area corresponding to the third element, and the third area corresponding to the fourth element to a 3D model based on the characteristics of each element.

[0099] According to an embodiment, when generating a second video by combining the 3D models for each element, the video generation unit 140 may generates the second video corresponding to the 3D model of the first video by combining the 3D models corresponding to each element. For example, after the image-by-image segmentation of each element has been performed, the individual elements can be combined into a single scene while reflecting their individual characteristics, allowing the scene to be restored identical to the input video. Additionally, all elements can be rendered to be well-represented from a new camera direction or at any given timestamp.

[0100] Finally, a calculation unit 160, as the entity that executes and controls the tasks performed by the preceding components of the 3D model reproduction apparatus 100, essentially corresponds to the central processing unit. The central processing unit can also be referred to as a controller, microcontroller, microprocessor, microcomputer, etc. Moreover, the central processing unit can be implemented in hardware, firmware, software, or a combination thereof. When it is implemented using hardware, it can take the form of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), or a field programmable gate array (FPGA), and when it is implemented using firmware or software, the firmware or software can be configured to include modules, procedures, or functions that perform the above-mentioned functions or operations. Furthermore, the 3D model reproduction apparatus 100 may also include a storage unit 170. The storage unit can be implemented using a memory such as a Read Only Memory (ROM), Random Access Memory (RAM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory, Static RAM (SRAM), Hard Disk Drive (HDD), or Solid State Drive (SSD).

[0101] As above, the method and apparatus for reproducing a dynamic scene as a 3D model using a video according to the present invention has been discussed. Meanwhile, the present invention is not limited to the specific embodiments and applications described above, and various modifications can be made by those skilled in the art without departing from the gist of the present invention as claimed in the claims. These modified implementations should not be understood as being separate from the technical spirit or scope of the present invention.

[0102] In particular, the components that implement the technical features of the present invention included in the block diagrams and flowcharts shown in the drawings attached to this specification signify the logical boundaries between these components. However, according to embodiments of software or hardware, the depicted components and their functions are implemented as standalone software modules, monolithic software structures, codes, services, or combinations thereof, and can be implemented by being stored on a medium executable by a computer equipped with a processor capable of executing the stored program codes and instructions. Therefore, all such embodiments should also be considered as falling within the scope of the present invention.

[0103] Therefore, while the accompanying drawings and their descriptions illustrate the technical features of the present invention, a specific arrangement of software for implementing these technical features should not be simply inferred unless clearly stated. In other words, various embodiments as described above may exist, and such embodiments, while retaining the same technical features as the present invention, may have some modification, and thus these should also be considered as falling within the scope of the present invention.

[0104] Moreover, although the flowchart describes the operations in a specific sequence in the drawings, this is illustrated to achieve the most desirable results and should not be understood as requiring such operations to be executed in the specific order or sequential order as shown, or all illustrated operations to be necessarily executed. In certain cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments as described above should not be understood as being required in all embodiments, and it should be understood that the program components and systems as described can generally be integrated together into a single software product or packaged into multiple software products.

Claims

1.A method for reproducing a dynamic scene as a 3D model using a video, the method comprising the steps of:receiving a first video that is to be reproduced as a 3D model from a user;based on elements contained in a scene within the first video, classifying the elements using a predetermined classification method;based on the characteristics of each element, converting each element to a 3D model;generating a second video by combining the 3D models for each element; andproviding the second video to the user.2.The method for reproducing a dynamic scene as a 3D model using a video according to claim 1, characterized in that the step of classifying the elements comprises the steps of:classifying the elements in the scene within the first video into a first element containing a static element and a second element containing a dynamic element; andif the second element includes a third element containing a lighting element, classifying the second element into the third element and a fourth element containing a moving element other than the third element.3.The method for reproducing a dynamic scene as a 3D model using a video according to claim 2, characterized in that the first element is an element that does not move within the scene, including a non-moving object and background, andthe second element is an element whose position or brightness changes within the scene, including at least one of the third element containing a changing lighting element and a fourth element containing a moving object.4.The method for reproducing a dynamic scene as a 3D model using a video according to claim 3, characterized in that the step of classifying the elements in the scene within the first video into a first element containing a static element and a second element containing a dynamic element comprises the steps of:segmenting the first video into a first area corresponding to a static area without a movement, including the first element; andsegmenting the areas other than the first area within the first video into a second area corresponding to a dynamic area with a movement, including the second element.5.The method for reproducing a dynamic scene as a 3D model using a video according to claim 4, characterized in that the step of classifying the elements comprises the step of:segmenting the areas within the first video using a data learning-based method, including an optical flow extraction method based on deep learning, or using information within the first video that includes a change in frames.6.The method for reproducing a dynamic scene as a 3D model using a video according to claim 5, characterized in that the step of classifying the second element into the third element and a fourth element containing a moving element other than the third element comprises the steps of:segmenting the first video into the third area corresponding to the fourth element for every frame; andsegmenting the areas other than the third area within the first video into the fourth area corresponding to the third area for every frame.7.The method for reproducing a dynamic scene as a 3D model using a video according to claim 6, characterized in that the step of converting each element to a 3D model comprises the step of:converting each of the first area corresponding to the first element, the fourth area corresponding to the third element, and the third area corresponding to the fourth element to a 3D model based on the characteristics of each element.8.The method for reproducing a dynamic scene as a 3D model using a video according to claim 7, characterized in that the step of generating a second video by combining the 3D models for each element comprises the step of:generating the second video corresponding to the 3D model of the first video by combining the 3D models corresponding to each element.9.The method for reproducing a dynamic scene as a 3D model using a video according to claim 8, characterized in that the step of providing the second video to the user comprises the step of:providing the second video to the user as video content that allows the first video to be viewed from a free viewpoint.10.An apparatus for reproducing a dynamic scene as a 3D model using a video, the apparatus comprising:a video reception unit that receives a video from a user;a video classification unit that classifies elements contained in the video using a predetermined classification method;a 3D conversion unit that converts each element to a 3D model;a video generation unit that generates a video by combining the 3D models converted for each element; anda video provision unit that provides the video to the user.11.The apparatus for reproducing a dynamic scene as a 3D model using a video according to claim 10, characterized in that the video reception unit receives a first video that is to be reproduced as a 3D model from the user;based on the elements contained in a scene within the first video, the video classification unit classifies the elements using a predetermined classification method;based on the characteristics of each element, the 3D conversion unit converts each element to a 3D model;the video generation unit generates a second video by combining the 3D models for each element; andthe video provision unit provides the second video to the user.12.The apparatus for reproducing a dynamic scene as a 3D model using a video according to claim 11, characterized in that the video classification unit classifies the elements in the scene within the first video into a first element containing a static element and a second element containing a dynamic element,and if the second element includes a third element containing a lighting element, classifies the second element into the third element and a fourth element containing a moving element other than the third element.13.The apparatus for reproducing a dynamic scene as a 3D model using a video according to claim 12, characterized in that the first element is an element that does not move within the scene, including a non-moving object and background, andthe second element is an element whose position or brightness changes within the scene, including at least one of the third element containing a changing lighting element and a fourth element containing a moving object.14.A system for reproducing a dynamic scene as a 3D model using a video, the system comprising:a memory in which at least one program is recorded; anda processor for executing the program, characterized in that the program comprises instructions for carrying out the steps of:receiving a first video that is to be reproduced as a 3D model from a user;based on elements contained in a scene within the first video, classifying the elements using a predetermined classification method;based on the characteristics of each element, converting each element to a 3D model;generating a second video by combining the 3D models for each element; andproviding the second video to the user.

Citation Information

Patent Citations

  • Systems and methods for 3D scene augmentation and reconstruction

    JP2022508674A

  • Method and apparatus for actively tracking a target using radar

    KR1020250072273A

  • Methods circuits devices systems and associated computer executable code for video feed processing

    US20180012366A1

  • Image rendering method and apparatus

    US20220309740A1

  • Appratus and method with 3D modeling

    US20230186580A1