Dynamic scene reconstruction method and system based on dynamic object and static scene separation
By separating dynamic objects and static scenes, and utilizing multi-view images and a 3D Gaussian splashing algorithm, local and global object representations are constructed, and pose parameters are optimized. This solves the problem of inconsistent dynamic object reconstruction in autonomous driving systems and achieves high-precision dynamic scene reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SUZHOU TANGYUAN TECHNOLOGY CO LTD
- Filing Date
- 2025-04-21
- Publication Date
- 2026-08-04
AI Technical Summary
Existing technologies struggle to effectively handle dynamic objects in autonomous driving systems, resulting in motion blur or missing objects in the reconstruction results. Furthermore, existing dynamic modeling schemes disrupt temporal correlations, making it difficult to guarantee the consistency of motion of dynamic objects over time.
By acquiring continuous frame images and multi-view images, and using instance segmentation and 3D Gaussian splashing algorithms, dynamic objects and static scenes are separated, local and global object representations are constructed, pose parameters are optimized, and high-precision reconstruction of dynamic objects is achieved. Dynamic objects and static backgrounds are then fused through interpolation pose.
It achieves high-precision separation and reconstruction of dynamic objects and static scenes, enhances the accuracy of 3D reconstruction, ensures the consistency of motion of dynamic objects in the time dimension, and generates a complete dynamic scene with spatiotemporal consistency.
Smart Images

Figure CN120411367B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, specifically relating to a dynamic scene reconstruction method and system based on the separation of dynamic objects and static scenes. Background Technology
[0002] Current autonomous driving systems rely heavily on manually designed rules for their planning modules, making it difficult to cover complex and ever-changing road scenarios. While data-driven end-to-end planning models show promise, their training depends on large-scale, highly generalizable multimodal data. Therefore, this paper proposes constructing digitally processable 3D models to support the generation of new perspectives and scene understanding and interaction in the field of autonomous driving.
[0003] However, existing methods have many limitations for temporal scene modeling. Traditional 3D reconstruction techniques focus on static scene modeling. However, they assume the scene is static and cannot handle dynamic objects (such as moving vehicles), resulting in motion blur or missing objects in the reconstruction results. In existing dynamic modeling schemes, most methods decouple continuous motion processes into discrete, frame-by-frame independent processing. This paradigm severs temporal correlation and makes it difficult to guarantee the consistency of motion of dynamic objects in the time dimension. The final reconstruction results may present non-physically realistic motion trajectory breaks or abrupt surface deformations. Summary of the Invention
[0004] This invention provides a method and system for dynamic scene reconstruction based on the separation of dynamic objects and static scenes.
[0005] The technical solution of the present invention is as follows: This invention provides a dynamic scene reconstruction method based on the separation of dynamic objects and static scenes, comprising: S1: Acquire continuous frame images and corresponding multi-view images of the frames, and select one frame in the continuous frame images as a reference frame. Use the reference frame to perform instance segmentation on the continuous frame multi-view images to obtain the multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the reference frame is processed by the visual shell algorithm to obtain the reference frame object point cloud data. S2: Based on the multi-view image sequence of the target object, the point cloud data of the reference frame object, and the camera parameters, the appearance parameters of the current frame are optimized by fixing the relative pose parameters; The relative pose parameters are optimized by fixing the optimized appearance parameters of the current frame and calculating the loss of the multi-view images of the adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame. The optimized relative pose parameters are multiplied by the absolute pose parameters to obtain the updated absolute pose parameters, and the adjacent frames of the current frame are used as optimized frames; return to execute S2 until all frames have been processed to obtain the set of optimized frames; Based on the optimized frame set, a 3D dynamic object representation is obtained by calculating the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object. S3: Based on the representation of 3D dynamic objects, the interpolated pose is calculated using the rotation matrix and translation vector of adjacent frames. The interpolated pose is multiplied by the global 3D appearance parameters of the object and then superimposed on the static scene to generate the reconstructed dynamic scene.
[0006] S2 operates as follows: S21: The appearance reconstructed from the current frame data is the appearance parameter of the current frame, and the relative change from the current frame to the adjacent frame is the relative pose parameter; The appearance reconstructed from fused multi-frame data is used as the global 3D appearance parameter of the object, and the change from the reference frame to the current frame is used as the absolute pose parameter. S22: Initialize the appearance parameters of the current frame based on the object point cloud data and absolute pose parameters of the reference frame; S23: With fixed relative pose parameters, optimize the appearance parameters of the current frame by calculating the loss between the current frame multi-view image and the corresponding rendered image obtained by 3D Gaussian splashing under the initialized appearance parameters of the current frame. S24: Fix the optimized appearance parameters of the current frame, and optimize the relative pose parameters by calculating the loss of the multi-view images of the adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame. S25: Multiply the optimized relative pose parameters by the absolute pose parameters to obtain the updated absolute pose parameters, and take the adjacent frames of the current frame as optimized frames; execute S21 until all frames have been processed to obtain the set of optimized frames; S26: Based on the optimized frame set, a 3D dynamic object representation is obtained by calculating the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object.
[0007] S2 further includes using three-dimensional Gaussian splashing to represent the three-dimensional appearance and geometry of dynamic objects, as shown by the formula: , ,accomplish; In the formula, For pose parameters, , Let be a rotation matrix. It is a translation vector; For appearance parameters; The total number of Gaussian ellipsoids; The index is the Gaussian ellipsoid number; , where represents the mean of the Gaussian distribution; , Represents the rotation parameters in quaternion form; where This represents the transformation from a rotation matrix to a quaternion; Representing the rotation matrix The corresponding quaternion; Representing quaternion multiplication; , represents the anisotropic scaling factor; , representing the transparency of each Gaussian ellipsoid; This represents the spherical harmonic coefficients.
[0008] S2 optimizes the relative pose parameters by calculating the loss between the multi-view images of adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the optimized appearance parameters of the current frame and the relative pose parameters. This optimization is achieved through the formula: ,accomplish; In the formula, To optimize the relative pose parameters during the process; These are relative pose parameters; These are the appearance parameters for the current frame. The rendered image obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame; For the current frame, the adjacent frames in terms of view The image below, The objects in the adjacent frames of the current frame in terms of view The mask below.
[0009] S2, based on the optimized frame set, calculates the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object, to obtain the 3D dynamic object representation, which is expressed by the formula: ,accomplish; In the formula, This is the loss value; To optimize the frame set The sampling frames in; These are absolute pose parameters; These are the global 3D appearance parameters of the object; The corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object; For sampling frames From the perspective The image below, For sampling frames The object in the view The mask below.
[0010] S3, based on the representation of a three-dimensional dynamic object, calculates the interpolated pose using the rotation matrix and translation vector of adjacent frames, specifically as follows: Based on the representation of three-dimensional dynamic objects, rotation interpolation is calculated by using the rotation matrix of adjacent frames and spherical linear interpolation of quaternions. Based on the translation vectors of adjacent frames, linear weighting is applied to calculate the translation interpolation; According to the formula: The interpolated pose is obtained; In the formula, , , These are interpolation pose, rotation interpolation, and translation interpolation, respectively. These are the interpolation parameters.
[0011] In step S3, the interpolated pose is multiplied by the global 3D appearance parameters of the object, and then superimposed on the static scene to generate a reconstructed dynamic scene. Specifically: According to the formula: This generates a reconstructed dynamic scene; In the formula, , , These represent the reconstructed dynamic scene, static scene, and appearance representation after dot product of interpolated pose and global 3D appearance parameters of the object, respectively. The static scene is obtained by reconstructing images that do not contain dynamic objects.
[0012] In step S22, the appearance parameters of the current frame are initialized based on the object point cloud data and absolute pose parameters of the reference frame, specifically as follows: According to the formula: Generate the point cloud of the current frame, and use the point cloud of the current frame to initialize the appearance parameters of the current frame; In the formula, For the current frame point cloud, These are absolute pose parameters. The reference frame contains object point cloud data.
[0013] In step S1, instance segmentation is performed on the multi-view images of consecutive frames using reference frames to obtain a multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the reference frame is processed by the visual shell algorithm to obtain the point cloud data of the object in the reference frame, specifically: After target detection, the multi-view images of the reference frame are segmented into instances to generate a multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the multi-view image of the reference frame is back-projected into the three-dimensional space to obtain the contour cone. The intersection of the contour cones of each view is taken as the visible shell of the target object. The visible shell is voxelized and sampled to obtain the point cloud data of the object in the reference frame.
[0014] This invention also provides a dynamic scene reconstruction system based on the separation of dynamic objects and static scenes, comprising: Separation module: used to acquire continuous frame images and corresponding multi-view images of the frame, and select one frame in the continuous frame images as a reference frame. The reference frame is used to perform instance segmentation on the continuous frame multi-view images to obtain the multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the reference frame is processed by the visual shell algorithm to obtain the point cloud data of the reference frame object. 3D dynamic object generation module: Based on the multi-view image sequence of the target object, the point cloud data of the reference frame object and camera parameters, the appearance parameters of the current frame are optimized by fixing the relative pose parameters; The relative pose parameters are optimized by fixing the optimized appearance parameters of the current frame and calculating the loss of the multi-view images of the adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame. The optimized relative pose parameters are multiplied by the absolute pose parameters to obtain the updated absolute pose parameters, and the adjacent frames of the current frame are used as optimized frames; return to the 3D dynamic object generation module until all frames have been processed to obtain the optimized frame set; Based on the optimized frame set, a 3D dynamic object representation is obtained by calculating the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object. Dynamic scene reconstruction module: Based on the representation of 3D dynamic objects, the module calculates the interpolated pose using the rotation matrix and translation vector of adjacent frames. The interpolated pose is then multiplied by the global 3D appearance parameters of the object and superimposed on the static scene to generate the reconstructed dynamic scene.
[0015] Beneficial effects This invention achieves dynamic scene reconstruction by separating dynamic objects from static scenes based on 3D Gaussian splashing. First, multimodal perception data is used to decouple the dynamic scene into dynamic objects and static scenes. In the dynamic object reconstruction stage, by constructing local object representations and global object representations, the motion parameters of the objects are explicitly modeled as differentiable optimization variables, achieving high-precision modeling of dynamic objects in continuous spatiotemporal dimensions. To address the information limitations of single-frame observation data, cross-frame temporal correlation modeling is used to fuse the observation data of dynamic objects in multiple frames, effectively enhancing the accuracy of 3D reconstruction. Finally, interpolation pose is used to fuse dynamic objects with static backgrounds, ultimately obtaining a complete dynamic scene with spatiotemporal consistency. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating step S1 of the present invention.
[0017] Figure 2 This is a flowchart illustrating step S2 of the present invention. Detailed Implementation
[0018] The following examples are intended to illustrate the present invention, and not to further limit the invention.
[0019] This invention provides a dynamic scene reconstruction method based on the separation of dynamic objects and static scenes, comprising the following steps: S1: Acquire continuous frame images and corresponding multi-view images of the frames, and select one frame in the continuous frame images as a reference frame. Use the reference frame to perform instance segmentation on the continuous frame multi-view images to obtain the multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the reference frame is processed by the visual shell algorithm to obtain the point cloud data of the reference frame object.
[0020] This stage involves preprocessing the raw data collected by sensors (multiple cameras) to decouple dynamic objects from the static background in the raw data. For example... Figure 1 As shown, the specific processing steps are as follows: Step 1: Acquire continuous frame images and multi-view images of corresponding frames, and select a reference frame.
[0021] In order to associate and identify target objects in different frames and from different perspectives, the following operations are then performed: real-time target detection is performed on the multi-view images of the reference frame based on target detection methods (such as YOLO) to obtain a set of two-dimensional bounding boxes of target objects in each perspective.
[0022] Step 2: Based on the camera calibration parameters, perform 3D back-projection calculations on the center point of each 2D bounding box to generate back-projection rays pointing from the optical center of each camera to the center of the target space. All back-projection rays form a 3D ray cluster.
[0023] Step 3: Calculate the spatial distance of the back-projected rays under different viewpoints, determine the same target object observed from multiple viewpoints based on the preset threshold, and obtain the target bounding box of the target object under each viewpoint, thereby realizing dynamic target instance-level matching across cameras.
[0024] Step 4: Based on the target bounding box of the target object in each viewpoint, perform instance segmentation on the continuous frame multi-view images (such as processing with the Segment Anything Model 2 model, abbreviated as SAM2) to generate a multi-view image sequence of the target object and the corresponding target object mask.
[0025] Step 5: Back-project the target object mask of the multi-view image of the reference frame into 3D space to obtain the contour cone. Take the intersection of the contour cones of each view as the visual hull of the target object. Perform voxel sampling on the visual hull to obtain the dynamic object point cloud, i.e., the object point cloud data of the reference frame.
[0026] S2: Based on the multi-view image sequence of the target object, the point cloud data of the reference frame object, and the camera parameters, the three-dimensional dynamic object is reconstructed through three-dimensional Gaussian splashing (3DGS).
[0027] In this stage, during the reconstruction of dynamic objects, the multi-view image sequence obtained from S1 is utilized. Corresponding object mask and reference frame object point cloud data Combining multiple camera intrinsic parameter matrices extrinsic matrix ,in Indicates the camera serial number. This indicates the frame number. This data is used for dynamic object reconstruction, enabling geometric modeling, appearance modeling, and motion trajectory capture of the object.
[0028] The operation is as follows: S21: The appearance representation reconstructed from the current frame data is the appearance parameter of the current frame. The relative change from the current frame to the adjacent frame is the relative pose parameter. This leads to the construction of local object representations. This allows for accurate estimation of the relative motion of dynamic objects between adjacent frames.
[0029] The appearance representation reconstructed from fused multi-frame data is used as the global 3D appearance parameter of the object. The change from the reference frame to the current frame is the absolute pose parameter. This leads to the construction of a global object representation. To achieve cross-frame 3D information fusion.
[0030] S21 further includes using three-dimensional Gaussian splashing to represent the three-dimensional appearance and geometric structure of dynamic objects, as shown by the formula: , ,accomplish; In the formula, For pose parameters, , Let be a rotation matrix. It is a translation vector; For appearance parameters; The total number of Gaussian ellipsoids; The index is the Gaussian ellipsoid number; , where represents the mean of the Gaussian distribution; , Represents the rotation parameters in quaternion form; where This represents the transformation from a rotation matrix to a quaternion; Representing the rotation matrix The corresponding quaternion; Representing quaternion multiplication; , represents the anisotropic scaling factor; , representing the transparency of each Gaussian ellipsoid; This represents the spherical harmonic coefficients.
[0031] After completing S21, a progressive optimization strategy starting from the reference frame is adopted, and motion estimation and global model optimization are achieved through frame-by-frame iteration. For the first... For each frame, the local 3DGS is first optimized to obtain the relative motion between the current local 3DGS and the local 3DGS of adjacent frames, and the absolute motion of the global 3DGS is updated. Finally, the global 3DGS is optimized to achieve cross-frame information fusion. See [link to relevant documentation]. Figure 2 The details are as follows: S22: Initialize the appearance parameters of the current frame based on the object point cloud data and absolute pose parameters of the reference frame, specifically: According to the formula: Generate the point cloud of the current frame, and initialize the appearance parameters of the current frame using the point cloud of the current frame. ; In the formula, For the current frame point cloud, These are absolute pose parameters. The reference frame contains object point cloud data.
[0032] S23: With fixed relative pose parameters, optimize the appearance parameters of the current frame by calculating the loss between the multi-view image of the current frame and the corresponding rendered image obtained by 3D Gaussian splashing under the initialized appearance parameters of the current frame.
[0033] It can be done through the formula: ,accomplish; In the formula, To optimize the appearance parameters of the current frame during the optimization process; These are the appearance parameters for the current frame. This is the corresponding rendered image obtained by 3D Gaussian splashing under the initialized appearance parameters of the current frame; For the current frame in the viewpoint The image below, The object in the current frame in terms of view The mask below.
[0034] S24: Fix the optimized appearance parameters of the current frame, and optimize the relative pose parameters by calculating the loss of the multi-view images of the adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame.
[0035] This optimization process can obtain accurate inter-frame relative motion parameters, which facilitates inter-frame relative motion estimation, and is based on the formula: ,accomplish; In the formula, To optimize the relative pose parameters during the process; These are relative pose parameters; These are the appearance parameters for the current frame. The rendered image obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame; For the current frame, the adjacent frames in terms of view The image below, The objects in the adjacent frames of the current frame in terms of view The mask below.
[0036] S25: Multiply the optimized relative pose parameters by the absolute pose parameters to obtain the updated absolute pose parameters, and take the adjacent frames of the current frame as optimized frames; return to execute S21 until all frames have been processed to obtain the optimized frame set.
[0037] S26: Based on the optimized frame set, a 3D dynamic object representation is obtained by calculating the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object.
[0038] This optimization process improves the accuracy of global 3DGS, as shown by the formula: ,accomplish; In the formula, This is the loss value; To optimize the frame set The sampling frames in; These are absolute pose parameters; These are the global 3D appearance parameters of the object; The corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object; For sampling frames From the perspective The image below, For sampling frames The object in the view The mask.
[0039] In the dynamic object reconstruction stage, this invention constructs local and global object representations to explicitly model the motion parameters of objects as differentiable optimization variables, thereby achieving high-precision modeling of dynamic objects in continuous spatiotemporal dimensions. To address the information limitations of single-frame observation data, cross-frame temporal correlation modeling is used to fuse observation data of dynamic objects in multiple frames, effectively enhancing the accuracy of 3D reconstruction.
[0040] S3: Based on the rotation matrix and translation vector of adjacent frames, calculate the interpolated pose, multiply the interpolated pose by the global 3D appearance parameters of the object, and then superimpose it with the static scene to generate the reconstructed dynamic scene.
[0041] When merging dynamic objects with static scenes at this stage, the 3D dynamic object representation generated by S2 was utilized. This includes the global 3D appearance representation of objects. and frame-by-frame motion representation To obtain continuous motion of objects, interpolation of the object's motion is performed between adjacent frames to achieve dynamic scene rendering. Specifically: First, during the dynamic object fusion process, given time... Let adjacent frames be respectively and ,in , These are the rotation matrix and translation vector, and the interpolation parameters, respectively. The interpolated pose is generated through the following steps. : Based on the rotation matrices of adjacent frames, the rotation interpolation is calculated using quaternion-based spherical linear interpolation (Slerp) to smoothly transition the rotation state. The formula is as follows: ; in, , representing the rotation matrix The corresponding quaternion; , representing the rotation matrix The corresponding quaternion; This represents the transformation from a quaternion to a rotation matrix; , represents the angle between two quaternions.
[0042] Based on the translation vectors of adjacent frames, linear weighting is applied to calculate the translation interpolation. The formula is as follows: ; According to the formula: The interpolated pose is obtained; In the formula, , , These are interpolation pose, rotation interpolation, and translation interpolation, respectively. These are the interpolation parameters.
[0043] Finally, the interpolated pose is multiplied by the object's global 3D appearance parameters and then superimposed on the static scene to generate the reconstructed dynamic scene, specifically: According to the formula: This generates a reconstructed dynamic scene; In the formula, , , These represent the reconstructed dynamic scene, static scene, and appearance representation after dot product of interpolated pose and global 3D appearance parameters of the object, respectively. The static scene is obtained by reconstructing (e.g., using 3DGS) an image that does not contain dynamic objects.
[0044] This invention achieves dynamic scene reconstruction by separating dynamic objects from static scenes based on 3D Gaussian splashing. First, multimodal perception data is used to decouple the dynamic scene into dynamic objects and static scenes. In the dynamic object reconstruction stage, by constructing local object representations and global object representations, the motion parameters of the objects are explicitly modeled as differentiable optimization variables, achieving high-precision modeling of dynamic objects in continuous spatiotemporal dimensions. To address the information limitations of single-frame observation data, cross-frame temporal correlation modeling is used to fuse the observation data of dynamic objects in multiple frames, effectively enhancing the accuracy of 3D reconstruction. Finally, interpolation pose is used to fuse the dynamic objects with the static background, resulting in a complete dynamic scene with spatiotemporal consistency.
[0045] This invention also provides a dynamic scene reconstruction system based on the separation of dynamic objects and static scenes, comprising: Separation module: used to acquire continuous frame images and corresponding multi-view images of the frame, and select one frame in the continuous frame images as a reference frame. The reference frame is used to perform instance segmentation on the continuous frame multi-view images to obtain the multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the reference frame is processed by the visual shell algorithm to obtain the point cloud data of the reference frame object. 3D dynamic object generation module: Based on the multi-view image sequence of the target object, the point cloud data of the reference frame object and camera parameters, the appearance parameters of the current frame are optimized by fixing the relative pose parameters; The relative pose parameters are optimized by fixing the optimized appearance parameters of the current frame and calculating the loss of the multi-view images of the adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame. The optimized relative pose parameters are multiplied by the absolute pose parameters to obtain the updated absolute pose parameters, and the adjacent frames of the current frame are used as optimized frames; return to the 3D dynamic object generation module until all frames have been processed to obtain the optimized frame set; Based on the optimized frame set, a 3D dynamic object representation is obtained by calculating the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object. Dynamic scene reconstruction module: Based on the representation of 3D dynamic objects, the module calculates the interpolated pose using the rotation matrix and translation vector of adjacent frames. The interpolated pose is then multiplied by the global 3D appearance parameters of the object and superimposed on the static scene to generate the reconstructed dynamic scene.
Claims
1. A method for dynamic scene reconstruction based on dynamic object and static scene separation, characterized in that, include: S1: Acquire continuous frame images and corresponding multi-view images of the frames, and select one frame in the continuous frame images as a reference frame. Use the reference frame to perform instance segmentation on the continuous frame multi-view images to obtain the multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the reference frame is processed by the visual shell algorithm to obtain the reference frame object point cloud data. S2: Based on the multi-view image sequence of the target object, the point cloud data of the reference frame object, and the camera parameters, the appearance parameters of the current frame are optimized by fixing the relative pose parameters; The relative pose parameters are optimized by fixing the optimized appearance parameters of the current frame and calculating the loss of the multi-view images of the adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame. The optimized relative pose parameters are multiplied by the absolute pose parameters to obtain the updated absolute pose parameters, and the adjacent frames of the current frame are used as the optimized frames. Return to execute S2 until all frames have been processed, resulting in an optimized frame set; Based on the optimized frame set, a 3D dynamic object representation is obtained by calculating the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object. S3: Based on the representation of 3D dynamic objects, the interpolated pose is calculated using the rotation matrix and translation vector of adjacent frames. The interpolated pose is multiplied by the global 3D appearance parameters of the object and then superimposed on the static scene to generate the reconstructed dynamic scene.
2. The method of claim 1, wherein, S2 operates as follows: S21: The appearance reconstructed from the current frame data is the appearance parameter of the current frame, and the relative change from the current frame to the adjacent frame is the relative pose parameter; The appearance reconstructed from fused multi-frame data is used as the global 3D appearance parameter of the object, and the change from the reference frame to the current frame is used as the absolute pose parameter. S22: Initialize the appearance parameters of the current frame based on the object point cloud data and absolute pose parameters of the reference frame; S23: With fixed relative pose parameters, optimize the appearance parameters of the current frame by calculating the loss between the current frame multi-view image and the corresponding rendered image obtained by 3D Gaussian splashing under the initialized appearance parameters of the current frame. S24: Fix the optimized appearance parameters of the current frame, and optimize the relative pose parameters by calculating the loss of the multi-view images of the adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame. S25: Multiply the optimized relative pose parameters by the absolute pose parameters to obtain the updated absolute pose parameters, and use the adjacent frames of the current frame as the optimized frames. Execute S21 until all frames have been processed, resulting in an optimized frame set; S26: Based on the optimized frame set, a 3D dynamic object representation is obtained by calculating the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object.
3. The dynamic scene reconstruction method based on the separation of dynamic objects and static scenes according to claim 1, characterized in that, The S2, further comprising representing the three-dimensional appearance and geometry of the dynamic object with a three-dimensional Gaussian splat, for implementing by the formula: , , In the formula, For pose parameters, , For rotation matrix, It is a translation vector; For appearance parameters; The total number of Gaussian ellipsoids; The index is the Gaussian ellipsoid number; , where represents the mean of the Gaussian distribution; , Represents the rotation parameters in quaternion form; where This represents the transformation from a rotation matrix to a quaternion; Representing the rotation matrix The corresponding quaternion; Representing quaternion multiplication; , represents the anisotropic scaling factor; , representing the transparency of each Gaussian ellipsoid; This represents the spherical harmonic coefficients.
4. The dynamic scene reconstruction method based on the separation of dynamic objects and static scenes according to claim 3, characterized in that, S2 optimizes the relative pose parameters by calculating the loss between the multi-view images of adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the optimized appearance parameters of the current frame and the relative pose parameters. This optimization is achieved through the formula: ,accomplish; In the formula, To optimize the relative pose parameters during the process; These are relative pose parameters; These are the appearance parameters for the current frame. The rendered image obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame; For the current frame, the adjacent frames in terms of view The image below, The objects in the adjacent frames of the current frame in terms of view The mask below.
5. The dynamic scene reconstruction method based on the separation of dynamic objects and static scenes according to claim 3, characterized in that, S2, based on the optimized frame set, calculates the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object, to obtain the 3D dynamic object representation, which is expressed by the formula: ,accomplish; In the formula, This is the loss value; To optimize the frame set The sampling frames in; These are absolute pose parameters; These are the global 3D appearance parameters of the object; The corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object; For sampling frames From the perspective The image below, For sampling frames The object in the view The mask below.
6. The dynamic scene reconstruction method based on the separation of dynamic objects and static scenes according to claim 1, characterized in that, S3, based on the representation of a three-dimensional dynamic object, calculates the interpolated pose using the rotation matrix and translation vector of adjacent frames, specifically as follows: Based on the representation of three-dimensional dynamic objects, rotation interpolation is calculated by using the rotation matrix of adjacent frames and spherical linear interpolation of quaternions. Based on the translation vectors of adjacent frames, linear weighting is applied to calculate the translation interpolation; According to the formula: The interpolated pose is obtained; In the formula, , , These are interpolation pose, rotation interpolation, and translation interpolation, respectively. These are the interpolation parameters.
7. The dynamic scene reconstruction method based on the separation of dynamic objects and static scenes according to claim 1, characterized in that, In step S3, the interpolated pose is multiplied by the global 3D appearance parameters of the object, and then superimposed on the static scene to generate a reconstructed dynamic scene. Specifically: According to the formula: This generates a reconstructed dynamic scene; In the formula, , , These represent the reconstructed dynamic scene, static scene, and appearance representation after dot product of interpolated pose and global 3D appearance parameters of the object, respectively. The static scene is obtained by reconstructing images that do not contain dynamic objects.
8. The dynamic scene reconstruction method based on the separation of dynamic objects and static scenes according to claim 2, characterized in that, In step S22, the appearance parameters of the current frame are initialized based on the object point cloud data and absolute pose parameters of the reference frame, specifically as follows: According to the formula: Generate the point cloud of the current frame, and use the point cloud of the current frame to initialize the appearance parameters of the current frame; In the formula, For the current frame point cloud, These are absolute pose parameters. The reference frame contains object point cloud data.
9. The dynamic scene reconstruction method based on the separation of dynamic objects and static scenes according to claim 1, characterized in that, In step S1, instance segmentation is performed on the multi-view images of consecutive frames using reference frames to obtain a multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the reference frame is processed by the visual shell algorithm to obtain the point cloud data of the object in the reference frame, specifically: After target detection, the multi-view images of the reference frame are segmented into instances to generate a multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the multi-view image of the reference frame is back-projected into the three-dimensional space to obtain the contour cone. The intersection of the contour cones of each view is taken as the visible shell of the target object. The voxelized sampling of the visible shell is used to obtain the point cloud data of the object in the reference frame.
10. A dynamic scene reconstruction system based on the separation of dynamic objects and static scenes, characterized in that, include: Separation module: used to acquire continuous frame images and corresponding multi-view images of the frame, and select one frame in the continuous frame images as a reference frame. The reference frame is used to perform instance segmentation on the continuous frame multi-view images to obtain the multi-view image sequence of the target object and the corresponding target object mask. The target object mask of the reference frame is processed by the visual shell algorithm to obtain the point cloud data of the reference frame object. 3D dynamic object generation module: Based on the multi-view image sequence of the target object and camera parameters, the appearance parameters of the current frame are optimized by fixing the relative pose parameters; The relative pose parameters are optimized by fixing the optimized appearance parameters of the current frame and calculating the loss of the multi-view images of the adjacent frames of the current frame and the corresponding rendered images obtained by 3D Gaussian splashing under the relative pose parameters acting on the optimized appearance parameters of the current frame. The optimized relative pose parameters are multiplied by the absolute pose parameters to obtain the updated absolute pose parameters, and the adjacent frames of the current frame are used as the optimized frames. Return to the 3D dynamic object generation module until all frames have been processed, resulting in an optimized frame set; Based on the optimized frame set, a 3D dynamic object representation is obtained by calculating the loss between the multi-view image of each optimized frame and the corresponding rendered image obtained by 3D Gaussian splashing under the global 3D appearance parameters and absolute pose parameters of the object. Dynamic scene reconstruction module: Based on the representation of 3D dynamic objects, the module calculates the interpolated pose using the rotation matrix and translation vector of adjacent frames. The interpolated pose is then multiplied by the global 3D appearance parameters of the object and superimposed on the static scene to generate the reconstructed dynamic scene.