Monocular dynamic scene reconstruction method and system based on self-supervised flow matching
Patent Information
- Application Number
- CN202511324049.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing dynamic scene reconstruction methods rely on accurate external motion estimation models, which leads to system complexity and reconstruction results that are limited by the accuracy of external motion estimation, especially in extreme environments such as deep space exploration.
A self-supervised flow matching mechanism is designed to learn 3D motion by directly utilizing the temporal changes between video frames, construct a self-supervised loss function, and achieve high-fidelity dynamic scene reconstruction without external priors, including motion constraints of camera streams and complete streams.
It simplifies the reconstruction process, improves the accuracy and robustness of motion learning, and significantly enhances reconstruction quality, especially in complex environments such as deep space exploration.
Smart Images

Figure CN120833442B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional reconstruction, in particular to a monocular dynamic scene reconstruction method and system based on self-supervised flow matching. BACKGROUND
[0002] Dynamic scene reconstruction is one of the core tasks in the field of computer vision, aiming to recover the three-dimensional appearance and structure of dynamic scenes from two-dimensional video sequences. This technology has broad application potential in various frontier fields, especially in deep space exploration, virtual reality, autonomous driving simulation, and digital content creation. In particular, in deep space exploration, how to accurately reconstruct dynamic three-dimensional scenes from remote sensing images or image sequences collected by probes has become a key technology to improve task accuracy and completeness.
[0003] Unlike static scene reconstruction, dynamic scene reconstruction not only needs to focus on the accurate expression of geometric shape, but more importantly, it needs to accurately capture the dynamic characteristics of objects and scenes. In deep space exploration, there is a long delay in signal transmission, the view angle and shooting conditions of the probe are different, and the imaging quality is affected by atmospheric interference, low-texture areas, and occlusion. Therefore, existing dynamic reconstruction methods based on external motion priors (such as optical flow field or feature point trajectory) usually rely on accurate external motion estimation models to provide constraints, which makes the system complex and the reconstruction result is directly limited by the accuracy of external motion estimation, especially in extreme or complex environments. It is particularly evident. It is worth noting that the dynamic information of the scene itself is contained in the continuous image frames of the video. The change of pixels in the time dimension is the direct projection of real motion in three-dimensional space. Therefore, directly establishing the constraint relationship between the motion of the three-dimensional model prediction and the change of the two-dimensional image observation can fundamentally simplify the reconstruction process and improve the accuracy and robustness of motion learning. This idea has more outstanding application value in dynamic scene reconstruction in deep space environments. Therefore, the present application designs a novel self-supervised flow matching mechanism to directly learn three-dimensional motion from the temporal changes between video frames, achieving high-fidelity dynamic scene reconstruction without external priors. SUMMARY
[0004] The present application aims to design a dynamic scene reconstruction framework that learns three-dimensional motion from monocular dynamic video in a self-supervised manner. The main content of the present application includes: first, constructing a complete regular space covering all dynamic and static elements in the scene; then explicitly separating the dynamic and static parts of the space to support differential modeling; finally, through a self-supervised mechanism consisting of complete flow matching and camera flow matching, precise motion constraints are applied to the separated dynamic and static parts, thereby achieving high-fidelity dynamic reconstruction.
[0005] The application adopts the technical solutions below.
[0006] In a first aspect, the application provides a monocular dynamic scene reconstruction method based on self-supervised flow matching, comprising:
[0007] The input video is divided into multiple segments and key video frames are selected, camera parameters and depth maps of each video frame are obtained based on the key video frames and through a hierarchical alignment strategy, pixels of each video frame are back-projected to a three-dimensional space in combination with a dynamic mask, separate static point clouds and dynamic point clouds are generated, and a static three-dimensional Gaussian model and a dynamic three-dimensional Gaussian model decoded in space-time are initialized respectively.
[0008] Two two-dimensional flow fields rendered from the dynamic scene are defined, including a camera flow caused by camera motion and a complete flow caused by camera and object motion together; a self-supervised loss function is designed: motion consistency constraints and cross-time rendering constraints are applied to the complete flow to form a complete flow matching loss, and motion consistency constraints and cross-time rendering constraints are applied to the camera flow in the static region to form a camera flow matching loss; the static three-dimensional Gaussian model and the dynamic three-dimensional Gaussian model are continuously optimized through the complete flow matching loss and the camera flow matching loss to obtain the final static three-dimensional Gaussian model and dynamic three-dimensional Gaussian model.
[0009] In one of the embodiments, the camera parameters and depth maps of each video frame are obtained based on the key video frames and through a hierarchical alignment strategy, specifically comprising:
[0010] The key video frames are coarsely aligned among the segments of the video based on a preset geometric basic model to establish a globally consistent geometric structure, so that the camera pose, camera intrinsic parameter and depth map of the key video frames are optimized. The coarse alignment result of the key video frames is taken as an initial value, and fine alignment is performed on all video frames in each video segment , so that the camera parameters and depth information of each video frame are recovered; the camera parameters include the camera pose and the camera intrinsic parameter.
[0011] In one of the embodiments, the static three-dimensional Gaussian model and the dynamic three-dimensional Gaussian model decoded in space-time are initialized respectively, specifically comprising:
[0012] The parameters of the static three-dimensional Gaussian model are decoded only by three orthogonal spatial feature planes, for representing the static background irrelevant to time; the parameters of the dynamic three-dimensional Gaussian model are decoded by the spatial feature planes and an additional temporal feature plane, capable of expressing the dynamic content changing with time in the scene.
[0013] In one of the embodiments, the camera flow caused by the camera motion is calculated in the following manner:
[0014] for Moment Camera For any pixel in the scene, firstly, the depth value of that pixel is rendered from the 3D dynamic scene. Then, combined with camera intrinsics and camera pose, the pixel is back-projected to 3D world coordinates. Finally, the resulting 3D point is projected onto... The camera of time In the process, new two-dimensional coordinates are obtained, and the new two-dimensional coordinates are compared with... Moment Camera The difference in the two-dimensional coordinates of the pixels is the camera flow.
[0015] In one embodiment, the calculation method for the complete flow caused by the combined motion of the camera and the object is as follows:
[0016] First, according to the recipient The static and dynamic 3D Gaussian models are each affected by any pixel on the time-mapping image, resulting in... The two-dimensional projection parameters at time t are then used to obtain the three-dimensional Gaussian at the time of deformation. The projection parameters at each moment are used to obtain the motion displacement of the pixel caused by the three-dimensional Gaussian. Finally, the complete flow of the pixel is obtained by weighted mixing of the motion displacements corresponding to all relevant three-dimensional Gaussians. .
[0017] In one embodiment, the application of motion consistency constraints and time-series rendering constraints to the complete stream constitutes the complete stream matching loss, specifically including:
[0018] Motion Consistency Constraints Use the complete stream Transformation Real images of moments and synthesize A true image of the moment, and Real images of moments Towards consensus;
[0019] ;
[0020] For use of the complete stream Transform real image The resulting image is synthesized A true image of the moment. This represents the loss of photometric consistency established between the two images;
[0021] Time-based rendering constraints Use the complete stream distortion a rendered image of the moment and synthesizing a rendered image of the moment, and a real image of the moment are consistent:
[0022] ;
[0023] for matching the complete flow based on using the complete flow a warped rendered image a synthesized image of the moment after the rendered image;
[0024] and together constitute a complete flow matching loss : , and are weights of and respectively;
[0025] The camera flow imposes a motion consistency constraint for ensuring background stability and a cross-time rendering constraint on the static region to constitute a camera flow matching loss, specifically comprising:
[0026] The same constraint mode as the complete flow is adopted for the camera flow in the static region to obtain a motion consistency constraint for ensuring background stability and a cross-time rendering loss , and together constitute a camera flow matching loss : ; and are weights of and respectively;
[0027] The static three-dimensional Gaussian model and the dynamic three-dimensional Gaussian model are continuously optimized through the complete flow matching loss and the camera flow matching loss, specifically comprising:
[0028] ;
[0029] is a total loss, is a complete flow matching loss, is a camera flow matching loss, is a weight of , is a weight of .
[0030] Secondly, the present invention provides a monocular dynamic scene reconstruction system based on self-supervised flow matching, comprising:
[0031] Scene representation module: Divide the input video into multiple segments and select key video frames. Based on the key video frames, obtain the camera parameters and depth map of each video frame through a hierarchical alignment strategy. Combine dynamic masking to back-project the pixels of each video frame to three-dimensional space, generate separate static point clouds and dynamic point clouds, and initialize the static three-dimensional Gaussian model and the spatiotemporally jointly decoded dynamic three-dimensional Gaussian model respectively.
[0032] The self-supervised flow matching module defines two types of 2D flow fields rendered from a dynamic scene: camera flow caused by camera motion and complete flow caused by the combined motion of the camera and objects. It designs a self-supervised loss function: applying motion consistency constraints and time-series rendering constraints to the complete flow constitutes the complete flow matching loss; applying motion consistency constraints and time-series rendering constraints to the camera flow in static regions to ensure background stability constitutes the camera flow matching loss. The module continuously optimizes the static and dynamic 3D Gaussian models using the complete flow matching loss and the camera flow matching loss, resulting in the final static and dynamic 3D Gaussian models.
[0033] In one embodiment, obtaining the camera parameters and depth map of each video frame based on key video frames and through a hierarchical alignment strategy specifically includes:
[0034] Based on a pre-defined geometric model, key video frames are coarsely aligned between video segments to establish a globally consistent geometric structure, thereby optimizing the camera pose of the key video frames. Camera internal parameters and depth map Using the coarse alignment result of the key video frames as initial values, apply the following to each video segment: All video frames within the frame are finely aligned to recover the camera parameters and depth information of each video frame; the camera parameters include camera pose and camera intrinsic parameters.
[0035] The initialization of the static 3D Gaussian model and the spatiotemporally jointly decoded dynamic 3D Gaussian model specifically includes:
[0036] The parameters of the static 3D Gaussian model are decoded by only three orthogonal spatial feature planes, which are used to characterize a time-independent static background; the parameters of the dynamic 3D Gaussian model are decoded by both spatial feature planes and an additional temporal feature plane, which can express the dynamic content of the scene that changes over time.
[0037] In one embodiment, the camera flow caused by camera motion is calculated as follows:
[0038] For any pixel point at time t, first render the depth value of the pixel point from the three-dimensional dynamic scene, then combine the camera intrinsic and camera pose to back-project the pixel point to the three-dimensional world coordinate, and then project the obtained three-dimensional point to the two-dimensional coordinate at time t to obtain a new two-dimensional coordinate, the difference between the two-dimensional coordinate at time t and the two-dimensional coordinate of the pixel point at time t is the camera flow. The camera at time t The camera at time t The camera at time t The camera at time t The camera at time t The camera at time t
[0039] The calculation method of the complete flow caused by the camera and the object motion is as follows:
[0040] First, according to the static three-dimensional Gaussian model and the dynamic three-dimensional Gaussian model affected by any pixel point on the image at time t, obtain each three-dimensional Gaussian at time t. At time t, the two-dimensional projection parameters are obtained, and then the deformation field is used to obtain the projection parameters of the three-dimensional Gaussian at time t. .
[0041] In one embodiment, the motion consistency constraint and the cross-time rendering constraint on the complete flow constitute a complete flow matching loss, which specifically includes:
[0042] Motion consistency constraint : use the complete flow transform the real image at time t and synthesize the real image at time t, and the real image at time t tend to be consistent.
[0043] ;
[0044] For the real image at time t synthesized by using the complete flow transform the real image , the real image at time t obtained after the transformation, indicates the photometric consistency loss established between the two images. Cross-time rendering constraint
[0045] : use the complete flow warp the rendered image generated by the three-dimensional Gaussian model at time t . and synthesizing rendered images at the time instant, with real images at the time instant consistent:
[0046] ;
[0047] for based on using complete flow distorted rendered images post-image synthesis rendered images at the time instant;
[0048] and together constitute complete flow matching loss : , and are weights of and respectively;
[0049] The pair of camera streams applies motion consistency constraints for ensuring background stability and cross-time rendering constraints in the static area to constitute camera stream matching loss, and specifically comprises:
[0050] The camera stream in the static area adopts the same constraint mode as the complete flow to obtain motion consistency constraints for ensuring background stability and cross-time rendering loss , and together constitute camera stream matching loss : ; and are weights of and respectively;
[0051] The static three-dimensional Gaussian model and the dynamic three-dimensional Gaussian model are continuously optimized through the complete flow matching loss and the camera stream matching loss, and specifically comprise:
[0052] ;
[0053] is the total loss, is the complete flow matching loss, is the camera stream matching loss, is the weight of , and is the weight of .
[0054] Compared with the prior art, the beneficial technical effects of the present application are:
[0055] This invention proposes a novel, fully self-supervised motion constraint method for monocular dynamic scene reconstruction tasks. This method directly establishes a constraint relationship between the predicted 3D motion and the observed 2D image changes in the video, without relying on external motion priors such as optical flow estimators. This not only simplifies the reconstruction process but also fundamentally improves the accuracy and robustness of motion learning. Attached Figure Description
[0056] Figure 1 This is a flowchart of the method in an embodiment of the present invention;
[0057] Figure 2 This is an overall flowchart of an embodiment of the present invention;
[0058] Figure 3 This is a schematic diagram of the complete stream and camera stream in an embodiment of the present invention;
[0059] Figure 4 This is a schematic diagram of the self-supervised flow matching constraint in an embodiment of the present invention. Detailed Implementation
[0060] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0061] like Figure 1 As shown, a monocular dynamic scene reconstruction method based on self-supervised flow matching in this invention includes the following steps:
[0062] S1. Divide the input video into multiple segments and select key video frames. Based on the key video frames and through a hierarchical alignment strategy, obtain the camera parameters and depth map of each video frame. Combine dynamic masking to back-project the pixels of each video frame to three-dimensional space to generate separate static point clouds and dynamic point clouds. Then initialize the static three-dimensional Gaussian model and the spatiotemporally co-decoded dynamic three-dimensional Gaussian model respectively.
[0063] S2 defines two types of two-dimensional flow fields rendered from a dynamic scene: camera flow caused by camera motion and complete flow caused by the combined motion of the camera and objects. A self-supervised loss function is designed: motion consistency constraints and time-series rendering constraints are applied to the complete flow to form the complete flow matching loss; motion consistency constraints and time-series rendering constraints are applied to the camera flow in static regions to ensure background stability to form the camera flow matching loss. The static 3D Gaussian model and the dynamic 3D Gaussian model are continuously optimized using the complete flow matching loss and the camera flow matching loss to obtain the final static 3D Gaussian model and the dynamic 3D Gaussian model.
[0064] The application provides a dynamic scene reconstruction method for learning three-dimensional motion from monocular videos in a self-supervised manner, mainly consisting of two core modules: (1) a scene representation module customized for flow matching; and (2) a self-supervised flow matching module.
[0065] The overall technical process is as shown in Figure 2 , and the detailed technical scheme is described as follows.
[0066] (1) The scene representation module customized for flow matching.
[0067] This module aims to build a scene representation that can completely cover and clearly distinguish static and dynamic components, providing a basis for subsequent application of fine self-supervised motion constraints.
[0068] Specifically, first, the input video frame sequence is divided into non-overlapping video segments , and the first frame of each segment is selected as the key video frame. The i-th video frame is denoted as . Subsequently, the application adopts a geometric basis model to initialize through a coarse-to-fine hierarchical alignment strategy:
[0069] First, the key video frames are coarsely aligned between the video segments to establish a globally consistent geometric structure, thereby optimizing the camera pose , camera intrinsic parameters and depth map Figure 2 of the key video frames. , where is the key video frame of the video segment , and the first video frame in the video segment is generally selected as the key video frame; are the camera pose, camera intrinsic parameters and depth map of , respectively, and are uniformly denoted as .
[0070] Second, the alignment results of the key video frames are used as initial values to finely align all video frames within each video segment , thereby restoring the accurate camera parameters and depth information of each video frame in the video.
[0071] Based on the finally obtained camera parameters, depth map and dynamic mask provided by the geometric basis model, the explicit separation of static point cloud set and dynamic point cloud set can be achieved by back-projecting the pixels of each video frame to the three-dimensional space.Finally, the point cloud is used to initialize three-dimensional Gaussian models with different characteristics: the parameters of the static Gaussian model are decoded by three orthogonal spatial feature planes, which represent the static background independent of time; and the parameters of the dynamic Gaussian model are decoded by the spatial feature plane and an additional temporal feature plane, which can express the dynamic content in the scene that changes over time.
[0072] (2) Self-supervised flow matching module. To make full use of the scene representation constructed by the previous module, which explicitly separates static and dynamic components, this module designs a self-supervised loss function without external motion labels, directly using the inter-frame changes of the video to supervise the learning of three-dimensional motion, the core mechanism of which is shown in Figure 3 and Figure 4 .
[0073] To achieve accurate constraints on different regions, this method defines two two-dimensional flow fields rendered from the dynamic three-dimensional scene.
[0074] The first is the camera flow , which represents the pixel displacement caused only by the change in camera pose on the static background. Its calculation method is as follows: for any pixel point under the camera at time , first render the depth value of the point from the three-dimensional scene, then combine the camera intrinsic and camera pose to back-project it to the three-dimensional world coordinate, and then project the three-dimensional point to the camera at time , to get the new two-dimensional coordinate, and the difference between the two two-dimensional coordinates is .
[0075] The second is the complete flow , which represents the pixel displacement caused by the motion of the scene objects and the camera motion on the entire image. Its calculation method is as follows: for any pixel point on the image at time , first get the two-dimensional projection parameters of each three-dimensional Gaussian that affects it at time (two-dimensional Gaussian), then use the deformation field to get the projection parameters of the three-dimensional Gaussian at time , to get the motion displacement of the three-dimensional Gaussian to the pixel point, and finally get the complete flow of the pixel point by weighting and mixing the displacements of all related three-dimensional Gaussians. The self-supervised principle of the present application is to use the above flow field to transform the image, requiring the transformed image to align with the real observation in the video, thereby realizing the use of the internal motion information of the video to supervise the motion modeling.
[0076] Specifically, for the complete flow , two constraints are designed: motion consistency constraint , which uses transform real image at time t post-synthesis real image at time t, and real image at time t as consistent as possible; and cross-time rendering constraints even if warp rendered image at time t post-synthesis rendered image at time t, and real image at time t consistent, to reinforce the temporal coherence of the model itself.
[0077] ;
[0078] ;
[0079] for based on using the complete flow transform real image post-synthesis of the resulting image real image at time t, for based on using the complete flow warp rendered image post-synthesis of the resulting image rendered image at time t.
[0080] The two constraints together constitute a complete flow matching loss .
[0081] photometric consistency loss established between the two images, to for example:
[0082] ;
[0083] is a weight parameter; respectively or height and width, respectively the pixel index in the height direction and the width direction, is the structural similarity.
[0084] The rendered image is an image rendered by the reconstructed three-dimensional Gaussian model, and the rendered image and the real image represent images under the same time, the same camera pose.
[0085] At the same time, for the camera flow This invention employs the exact same constraint method, but applies it only to image regions identified as static, thereby obtaining motion consistency constraints specifically designed to ensure background stability. and rendering loss across time These together constitute the camera stream matching loss. Specifically:
[0086] Motion consistency constraints used to ensure background stability ,use Transformation Real still images of moments Post-synthesis A true still image of the moment, and with Real still images of moments Maintaining consistency as much as possible; time-series rendering constraints to ensure background stability. ,use distortion Rendering static images at different times Post-synthesis Render static images at specific times, and with Real still images of moments Consistent.
[0087] ;
[0088] ;
[0089] For using camera stream Transform real static image The resulting image is synthesized A true still image of a moment; For using camera stream Distort rendering of static images The resulting image is synthesized Rendering static images at any given moment.
[0090] ;
[0091] and They are respectively and The weight.
[0092] These two flow matching losses complement each other and together constitute the self-supervised signal of this invention for learning 3D motion without the need for external priors:
[0093] .
[0094] In the present application, the basic dynamic three-dimensional Gaussian sputtering model and its dynamic modeling method using the space-time feature plane can be regarded as prior art. The core innovation of the present application is to design a novel self-supervised flow matching mechanism composed of complete flow matching and camera flow matching based on this framework, and to propose a matching dynamic and static separation scene representation construction process. This design enables the model to learn motion information only from the video itself, reduces the dependence on external motion priors, and significantly improves the quality and robustness of dynamic scene reconstruction.
[0095] The present application can be applied to deep space exploration, robot navigation and digital scene modeling. Through the dynamic scene video collected by the probe or robot, combined with the dynamic scene reconstruction method of the present application, a high-precision four-dimensional scene model can be generated. These models not only can be used for environmental analysis and task planning in deep space exploration, but also can provide accurate data support for robot autonomous navigation, path planning and virtual environment simulation. In addition, the generated digital scene assets can be widely used in virtual reality, augmented reality and other digital content creation fields, promoting the development of virtual simulation and automation control technology.
[0096] The terms used herein are merely intended to describe specific embodiments and are not intended to limit the present application. The terms "comprise", "include" and the like as used herein indicate the presence of the stated features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0097] In one embodiment, the present application also provides a computer system, which can be a server. The computer system includes a processor, a memory and a network interface connected by a system bus. Among them, the processor of the computer system is used to provide computing and control capability. The memory of the computer system includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer system is used to store the data used in the above method. The network interface of the computer system is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to realize the above method.
[0098] The technical features of the above embodiments can be combined in any way. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0099] It will be obvious to a person skilled in the art that the application is not limited to the details of the foregoing exemplary embodiments and can be implemented in other concrete forms without departing from the spirit or essential characteristics of the application. The embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein and no
[0100] Furthermore, it should be understood that although the description is made on the basis of the embodiments, not every embodiment contains only one independent technical solution, and the description is made in this way only for the sake of clarity, and a person skilled in the art should consider the description as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by a person skilled in the art.
Claims
1. A monocular dynamic scene reconstruction method based on self-supervised flow matching, characterized in that, include: The input video is divided into multiple segments and key video frames are selected. Based on the key video frames, the camera parameters and depth map of each video frame are obtained through a hierarchical alignment strategy. The pixels of each video frame are back-projected into three-dimensional space using dynamic masks to generate separate static point clouds and dynamic point clouds. The static three-dimensional Gaussian model and the spatiotemporally co-decoded dynamic three-dimensional Gaussian model are initialized respectively. Define two types of two-dimensional flow fields rendered from dynamic scenes, including camera flow caused by camera motion and complete flow caused by the combined motion of camera and objects; Design a self-supervised loss function: Apply motion consistency constraints and time-series rendering constraints to the complete flow to form the complete flow matching loss. Motion Consistency Constraints Use the complete stream Transformation Real images of moments and synthesize A true image of the moment, and Real images of moments Towards consensus; ; For use of the complete stream Transform real image The resulting image is synthesized A true image of the moment. This represents the photometric consistency loss established between the two images; time-series rendering constraints. Use the complete stream distortion Rendered image generated by the 3D Gaussian model at time step and synthesize The rendered image at any given moment, and Real images of moments Consistency: ; For use of the complete stream Distorted rendering image The resulting image is synthesized The rendered image at any given moment; and Together they constitute the complete flow matching loss : , and They are respectively and The weights; The camera flow matching loss is formed by applying motion consistency constraints and time-series rendering constraints to the camera flow in static regions to ensure background stability. Applying the same constraint method to the camera flow as to the complete flow in the static region yields motion consistency constraints to ensure background stability. and rendering loss across time , and Together they constitute the camera stream matching loss : ; and They are respectively and The weights; The static and dynamic 3D Gaussian models are continuously optimized using full flow matching loss and camera flow matching loss to obtain the final static and dynamic 3D Gaussian models: ; For the total loss, For complete stream matching loss, For camera stream matching loss, for The weight, for The weight.
2. The monocular dynamic scene reconstruction method based on self-supervised flow matching according to claim 1, characterized in that, The process of obtaining camera parameters and depth maps for each video frame based on key video frames and using a hierarchical alignment strategy specifically includes: Based on a pre-defined geometric model, key video frames are coarsely aligned between video segments to establish a globally consistent geometric structure, thereby optimizing the camera pose of the key video frames. Camera internal parameters and depth map Using the coarse alignment result of the key video frames as initial values, apply the following to each video segment: All video frames within the frame are finely aligned to recover the camera parameters and depth information of each video frame; the camera parameters include camera pose and camera intrinsic parameters.
3. The monocular dynamic scene reconstruction method based on self-supervised flow matching according to claim 1, characterized in that, The initialization of the static 3D Gaussian model and the spatiotemporally jointly decoded dynamic 3D Gaussian model specifically includes: The parameters of the static 3D Gaussian model are decoded by only three orthogonal spatial feature planes, which are used to characterize a time-independent static background; the parameters of the dynamic 3D Gaussian model are decoded by both spatial feature planes and an additional temporal feature plane, which can express the dynamic content of the scene that changes over time.
4. The monocular dynamic scene reconstruction method based on self-supervised flow matching according to claim 1, characterized in that, The calculation method for the camera flow caused by camera motion is as follows: for Moment Camera For any pixel in the scene, firstly, the depth value of that pixel is rendered from the 3D dynamic scene. Then, combined with camera intrinsics and camera pose, the pixel is back-projected to 3D world coordinates. Finally, the resulting 3D point is projected onto... The camera of time In the process, new two-dimensional coordinates are obtained, and the new two-dimensional coordinates are compared with... Moment Camera The difference in the two-dimensional coordinates of the pixels is the camera flow.
5. The monocular dynamic scene reconstruction method based on self-supervised flow matching according to claim 1, characterized in that, The calculation method for the complete flow caused by the combined motion of the camera and the object is as follows: First, according to the recipient The static and dynamic 3D Gaussian models are each affected by any pixel on the time-mapping image, resulting in... The two-dimensional projection parameters at time t are then used to obtain the three-dimensional Gaussian at the time of deformation. The projection parameters at each moment are used to obtain the motion displacement of the pixel caused by the three-dimensional Gaussian. Finally, the complete flow of the pixel is obtained by weighted mixing of the motion displacements corresponding to all relevant three-dimensional Gaussians. .
6. A monocular dynamic scene reconstruction system based on self-supervised flow matching, characterized in that, include: Scene representation module: Divide the input video into multiple segments and select key video frames. Based on the key video frames, obtain the camera parameters and depth map of each video frame through a hierarchical alignment strategy. Combine dynamic masking to back-project the pixels of each video frame to three-dimensional space, generate separate static point clouds and dynamic point clouds, and initialize the static three-dimensional Gaussian model and the spatiotemporally jointly decoded dynamic three-dimensional Gaussian model respectively. Self-supervised flow matching module: Defines two types of 2D flow fields rendered from dynamic scenes, including camera flow caused by camera motion and complete flow caused by the combined motion of camera and objects; Design a self-supervised loss function: Apply motion consistency constraints and time-series rendering constraints to the complete flow to form the complete flow matching loss: Motion consistency constraints Use the complete stream Transformation Real images of moments and synthesize A true image of the moment, and Real images of moments Towards consensus; ; For use of the complete stream Transform real image The resulting image is synthesized A true image of the moment. This represents the photometric consistency loss established between the two images; time-series rendering constraints. Use the complete stream distortion Rendered image generated by the 3D Gaussian model at time step and synthesize The rendered image at any given moment, and Real images of moments Consistency: ; For use of the complete stream Distorted rendering image The resulting image is synthesized The rendered image at any given moment; and Together they constitute the complete flow matching loss : , and They are respectively and The weights; The camera flow matching loss is formed by applying motion consistency constraints and time-series rendering constraints to the camera flow in static regions to ensure background stability: applying the same constraints to the camera flow as to the complete flow in static regions yields motion consistency constraints to ensure background stability. and rendering loss across time , and Together they constitute the camera stream matching loss : ; and They are respectively and The weights; The static 3D Gaussian model and the dynamic 3D Gaussian model are continuously optimized by the full flow matching loss and the camera flow matching loss to obtain the final static 3D Gaussian model and the dynamic 3D Gaussian model. ; For the total loss, For complete stream matching loss, For camera stream matching loss, for The weight, for The weight.
7. A monocular dynamic scene reconstruction system based on self-supervised flow matching according to claim 6, characterized in that, The process of obtaining camera parameters and depth maps for each video frame based on key video frames and using a hierarchical alignment strategy specifically includes: Based on a pre-defined geometric model, key video frames are coarsely aligned between video segments to establish a globally consistent geometric structure, thereby optimizing the camera pose of the key video frames. Camera internal parameters and depth map Using the coarse alignment result of the key video frames as initial values, apply the following to each video segment: All video frames within the frame are finely aligned to recover the camera parameters and depth information of each video frame; the camera parameters include camera pose and camera intrinsic parameters. The initialization of the static 3D Gaussian model and the spatiotemporally jointly decoded dynamic 3D Gaussian model specifically includes: The parameters of the static 3D Gaussian model are decoded by only three orthogonal spatial feature planes, which are used to characterize a time-independent static background; the parameters of the dynamic 3D Gaussian model are decoded by both spatial feature planes and an additional temporal feature plane, which can express the dynamic content of the scene that changes over time.
8. A monocular dynamic scene reconstruction system based on self-supervised flow matching according to claim 6, characterized in that, The calculation method for the camera flow caused by camera motion is as follows: for Moment Camera For any pixel in the scene, firstly, the depth value of that pixel is rendered from the 3D dynamic scene. Then, combined with camera intrinsics and camera pose, the pixel is back-projected to 3D world coordinates. Finally, the resulting 3D point is projected onto... The camera of time In the process, new two-dimensional coordinates are obtained, and the new two-dimensional coordinates are compared with... Moment Camera The difference in the two-dimensional coordinates of the pixels below is the camera flow; The calculation method for the complete flow caused by the combined motion of the camera and the object is as follows: First, according to the recipient The static and dynamic 3D Gaussian models are each affected by any pixel on the time-mapping image, resulting in... The two-dimensional projection parameters at time t are then used to obtain the three-dimensional Gaussian at the time of deformation. The projection parameters at each moment are used to obtain the motion displacement of the pixel caused by the three-dimensional Gaussian. Finally, the complete flow of the pixel is obtained by weighted mixing of the motion displacements corresponding to all relevant three-dimensional Gaussians. .
Citation Information
Patent Citations
Dynamic scene reconstruction method and device, equipment, medium and product
CN119379907A
Dynamic Gaussian scene reconstruction method based on depth regularization
CN119991974A