Image reconstruction model generation method and system, image reconstruction method and system, equipment and medium

By training local dynamic neural radiation fields, adjusting camera attitude and parameters, and generating image reconstruction models, the problem of inaccurate camera attitude estimation in dynamic scenes is solved, and high-quality image reconstruction and viewing angle synthesis are achieved.

CN120472028APending Publication Date: 2025-08-12SHANGHAI TECH UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510563123.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In video image reconstruction in dynamic scenes, the prior art has the problem of inaccurate camera posture estimation, which affects the three-dimensional reconstruction and viewing angle synthesis effects.

Method used

By acquiring the video frame sequence and the camera pose sequence, using local dynamic neural radiation fields for training, adjusting the camera pose and parameters, generating an image reconstruction model, and optimizing image reconstruction with static and dynamic modules.

Benefits of technology

It significantly improves the accuracy and robustness of camera pose estimation, and improves the image reconstruction quality and rendering effect in dynamic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472028A_ABST
    Figure CN120472028A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an image reconstruction model training method, an image reconstruction method, a system, equipment and a medium. The method comprises the following steps: acquiring a to-be-processed video frame sequence and a to-be-adjusted camera attitude sequence corresponding to the to-be-processed video frame sequence; selecting a corresponding local dynamic nerve radiation field according to the to-be-adjusted camera attitude sequence; training the selected local dynamic neural radiation field according to the video frame sequence and the to-be-adjusted camera attitude sequence, and adjusting the parameters of the corresponding camera attitude and the local dynamic neural radiation field according to the difference degree between the reconstructed images generated by the video frame and the local dynamic neural radiation field. Adjusting a corresponding camera attitude range according to the adjusted camera attitude and / or offset degree so as to complete training of the local dynamic nerve radiation field; and generating an image reconstruction model by using the plurality of trained local dynamic neural radiation fields in parallel. According to the invention, the accuracy of image reconstruction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image reconstruction model training and image reconstruction method, system, device and medium. Background Art

[0002] With the advancement of technology, people are increasingly using handheld videos to record their daily activities, such as strolling, hiking, and cycling. However, such videos often feature significant camera pose variations and numerous dynamic objects, posing challenges for subsequent 3D understanding. Dynamic view synthesis can recreate video content from different perspectives, enhancing the user's immersive viewing experience. Dynamic neural radiance field methods, based on neural radiance fields, have demonstrated excellent performance in synthesizing images from new perspectives. However, the fundamental premise of neural radiance fields is that they require a set of images with accurate camera poses as input. Especially in dynamic scenes, this pose information relies on traditional 3D reconstruction methods based on multi-frame feature matching to estimate and recover the camera motion trajectory and scene structure. However, in practical applications, handheld videos often contain numerous moving objects, resulting in inaccurate or even complete failure of camera pose estimation, severely impacting subsequent 3D reconstruction and view synthesis. Therefore, there is a need for training image reconstruction models and providing methods, systems, devices, and media for image reconstruction. Summary of the Invention

[0003] In view of the above shortcomings of the prior art, the purpose of the present invention is to provide a method, system, device and medium for generating an image reconstruction model, so as to improve the problem of low accuracy of video image reconstruction in dynamic scenes in the prior art.

[0004] To achieve the above-mentioned purpose and other related purposes, the present invention provides a method for generating an image reconstruction model, comprising: obtaining a video frame sequence to be processed and a corresponding camera posture sequence to be adjusted; wherein the video frame sequence is derived from a dynamic scene, and each video frame has a corresponding time identifier; selecting a corresponding local dynamic neural radiation field according to the camera posture sequence to be adjusted; wherein each local dynamic neural radiation field is preset with a corresponding camera posture range; according to the video frame sequence and the camera posture sequence to be adjusted, training the selected local dynamic neural radiation field, adjusting the corresponding camera posture and parameters of the local dynamic neural radiation field according to the difference between the reconstructed image generated by the video frame and the local dynamic neural radiation field, and adjusting the corresponding camera posture range according to the adjusted camera posture and / or offset to complete the training of the local dynamic neural radiation field; wherein the offset is the deviation between the camera posture range and the adjusted camera posture; generating an image reconstruction model by parallelizing the trained multiple local dynamic neural radiation fields.

[0005] In one embodiment of the present invention, for each selected local dynamic neural radiation field, the selected local dynamic neural radiation field is trained according to the video frame sequence and the camera posture sequence to be adjusted, and the corresponding camera posture and the parameters of the local dynamic neural radiation field are adjusted according to the difference between the video frame and the reconstructed image generated by the local dynamic neural radiation field, and the corresponding camera posture range is adjusted according to the adjusted camera posture and / or offset to complete the training of the local dynamic neural radiation field, including: selecting a number of video frames in sequence from the video frame sequence; inputting each selected video frame and the corresponding camera posture into the current local dynamic neural radiation field to generate a corresponding reconstructed image; calculating the difference between each reconstructed image and the corresponding video frame, and adjusting the corresponding camera posture and the parameters of the local dynamic neural radiation field based on the difference; calculating the camera posture range and the offset between the camera posture adjusted based on the difference, and adjusting the camera posture range according to the adjusted camera posture and / or offset to complete the training of the local dynamic neural radiation field.

[0006] In one embodiment of the present invention, when the video frame input to the current local dynamic neural radiation field does not include the last video frame of the video frame sequence, the offset between the camera posture range and the camera posture adjusted based on the difference is calculated, and the camera posture range is adjusted according to the adjusted camera posture and / or the offset to complete the training of the local dynamic neural radiation field, including: calculating the offset between the adjusted camera posture and the corresponding camera posture range, and judging whether the offset is greater than a preset offset threshold: if so, adjusting the corresponding camera posture range based on the adjusted camera posture to complete the training of the current local dynamic neural radiation field; otherwise, reselecting several video frames from the video frame sequence and training the current local dynamic neural radiation field again.

[0007] In one embodiment of the present invention, when the video frame input to the local dynamic neural radiation field includes the last video frame of the video frame sequence, the offset between the camera posture range and the camera posture adjusted based on the difference is calculated, and the camera posture range is adjusted according to the adjusted camera posture and / or the offset to complete the training of the local dynamic neural radiation field, including: adjusting the camera posture range based on the adjusted camera posture to complete the training of the current local dynamic radiation field.

[0008] In one embodiment of the present invention, for each video frame, the video frame and the corresponding camera posture are input into the corresponding local dynamic neural radiation field, and the process of generating the corresponding reconstructed image includes: based on the camera posture corresponding to the video frame and the preset camera intrinsic parameters, the video frame is back-projected to obtain a ray sequence corresponding to the video frame; for each ray in the ray sequence: the ray is uniformly sampled according to a preset depth range to obtain multiple spatial sampling points along the ray direction; the spatial sampling points of the ray sequence are arranged in sequence according to the sampling order to obtain a spatial sampling point sequence; the line of sight direction sequence corresponding to the spatial sampling point sequence is determined according to the camera posture; the spatial sampling point sequence, the corresponding line of sight direction sequence and the time stamp corresponding to the video frame are input into the local dynamic neural radiation field to generate a reconstructed image corresponding to the video frame.

[0009] In one embodiment of the present invention, a spatial sampling point sequence and a corresponding gaze direction sequence are input into a local dynamic neural radiation field to generate a reconstructed image corresponding to a video frame, including: inputting the spatial sampling point sequence and the corresponding gaze direction sequence into a static module of the local dynamic neural radiation field to obtain the static color and static density of each spatial sampling point; inputting the spatial sampling point sequence and the corresponding gaze direction sequence into a dynamic module of the local dynamic neural radiation field to obtain the dynamic color, dynamic density and dynamic mask of each spatial sampling point based on the time stamp of the corresponding video frame; for each spatial sampling point: performing weighted fusion on the dynamic color, dynamic density, static color, static density and dynamic mask corresponding to the spatial sampling point to obtain the fused color and fused density corresponding to the spatial sampling point; and accumulating and weighting the fused colors and fused densities of each spatial sampling point based on their depth order in the corresponding ray to generate a reconstructed image corresponding to the video frame.

[0010] In one embodiment of the present invention, an image reconstruction method is also provided, the method comprising: selecting any time marker from a video frame sequence to be processed, and presetting a corresponding camera posture for the time marker; determining a target video frame corresponding to the time marker from the video frame sequence; determining an adjusted camera posture range that matches the camera posture based on the camera posture, and determining the corresponding local dynamic neural radiation field in the image reconstruction model accordingly; wherein the image reconstruction model is obtained by the method for generating an image reconstruction model according to any one of claims 1-6; inputting the target video frame and the camera posture into the local dynamic neural radiation field corresponding to the image reconstruction model to generate a reconstructed image under the current time marker and camera posture.

[0011] In one embodiment of the present invention, a training system for an image reconstruction model is further provided, the system comprising: a data acquisition module for acquiring a video frame sequence to be processed and a corresponding camera pose sequence to be adjusted; wherein the video frame sequence is derived from a dynamic scene, and each video frame has a corresponding time stamp; a radiation field determination module for selecting a corresponding local dynamic neural radiation field based on the camera pose sequence to be adjusted; wherein each local dynamic neural radiation field is preset with a corresponding camera pose range;

[0012] The radiation field training module is used to train the selected local dynamic neural radiation field according to the video frame sequence and the camera posture sequence to be adjusted, adjust the corresponding camera posture and the parameters of the local dynamic neural radiation field according to the difference between the video frame and the reconstructed image generated by the local dynamic neural radiation field, and adjust the corresponding camera posture range according to the adjusted camera posture and / or offset to complete the training of the local dynamic neural radiation field; wherein the offset is the deviation between the camera posture range and the adjusted camera posture; the model generation module is used to generate an image reconstruction model for multiple local dynamic neural radiation fields that have been trained in parallel.

[0013] In one embodiment of the present invention, an electronic device is also provided, including: one or more processors; a storage device for storing one or more programs, which, when the one or more programs are executed by one or more processors, enables the electronic device to implement any of the above-mentioned image reconstruction model generation methods or image reconstruction methods.

[0014] In one embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a computer processor, the computer executes any of the above-mentioned methods for generating an image reconstruction model or image reconstruction methods.

[0015] As described above, the training of an image reconstruction model and the image reconstruction method, system, device and medium of the present invention have the following beneficial effects: by processing the video frame sequence and the corresponding camera posture sequence to be adjusted in the dynamic scene, combined with the gradual training and optimization of multiple local dynamic neural radiation fields, high-quality modeling and image reconstruction of complex three-dimensional dynamic scenes are achieved. Among them, during the radiation field training process, the camera posture is adjusted according to the difference between the reconstructed image and the original video frame, thereby significantly improving the accuracy and robustness of the posture estimation. Furthermore, by setting the camera posture range of the local dynamic neural radiation field and dynamically updating the camera posture range in combination with the offset, each local dynamic neural radiation field can be more accurately adapted to the current viewing angle, greatly improving the overall rendering quality, so that the final reconstructed image can be more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic flow chart of a method for training an image reconstruction model provided by an embodiment of the present invention;

[0017] Figure 2 Shown is a system architecture diagram of the image reconstruction model of the present invention;

[0018] Figure 3 A schematic flow chart of an image reconstruction method provided by an embodiment of the present invention;

[0019] Figure 4 Shown is a structural block diagram of a training system for an image reconstruction model provided by an embodiment of the present invention;

[0020] Figure 5 Shown is a structural schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0022] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0023] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0024] The present invention provides a training method for an image reconstruction model. By introducing a progressive optimization strategy, it is possible to efficiently optimize dynamic objects, static background radiation fields, and camera postures in videos of different trajectory lengths, thereby achieving high-quality 3D reconstruction in complex scenes. The present invention gradually refines the representation of dynamic objects and static backgrounds during the optimization process, and adaptively adjusts the optimization strategy according to changes in the video trajectory, so that long-trajectory videos can still obtain a stable view synthesis effect. In addition, the present invention can also effectively distinguish between dynamic and static elements in the scene, while ensuring the consistency of the scene, avoiding the interference of dynamic content on the reconstruction of geometric structures. The present invention performs well in dynamic scenes, not only outperforming existing methods in reconstruction accuracy, but also achieving significant advantages in long-term trajectory optimization and dynamic content processing.

[0025] like Figure 1 As shown, the training method of the image reconstruction model includes the following steps:

[0026] S11. Obtain a video frame sequence to be processed and a corresponding camera posture sequence to be adjusted; wherein the video frame sequence is derived from a dynamic scene, and each video frame has a corresponding time stamp.

[0027] The original video can be shot by a mobile shooting device such as a handheld camera or a mobile phone, and the shot video can be decomposed into continuous video frames in chronological order to obtain a video frame sequence. Among them, the video content is a dynamic scene, which includes but is not limited to dynamic objects such as people and vehicles that change over time, so there is a large change in perspective between video frames. Each video frame carries a unique time identifier, which is used to indicate the position of the video frame in the entire video frame sequence. In addition, each video frame also has an initial camera pose, which is used to indicate the rotation state and spatial position of the corresponding video frame at the time of shooting, so that the motion trajectory of the mobile shooting device in three-dimensional space can be characterized. It can be understood that the initial camera pose can be obtained by various traditional methods and adjusted during the subsequent training of the local dynamic neural radiation field.

[0028] S12. Selecting a corresponding local dynamic neural radiation field according to the camera posture sequence to be adjusted; wherein each local dynamic neural radiation field is preset with a corresponding camera posture range.

[0029] When training the local dynamic neural radiation field, the local dynamic neural radiation field that matches the camera posture corresponding to each video frame in the camera posture sequence to be adjusted will be dynamically selected for training. Specifically, each local dynamic neural radiation field will be preset with a camera posture range during initialization to characterize the applicable viewing angle range of the local dynamic neural radiation field. It should be noted that all local dynamic neural radiation fields have the same network structure, but their corresponding camera posture ranges are different, and the initial parameters are not shared with each other. Therefore, each local dynamic neural radiation field can independently adapt to different camera motion trajectories. This structure makes the entire image reconstruction model have good scalability and adaptability, and is particularly suitable for processing handheld videos with a large range of viewing angle changes and complex dynamic content.

[0030] S13. Train the selected local dynamic neural radiation field according to the video frame sequence and the camera posture sequence to be adjusted, adjust the corresponding camera posture and parameters of the local dynamic neural radiation field according to the difference between the video frame and the reconstructed image generated by the local dynamic neural radiation field, and adjust the corresponding camera posture range according to the adjusted camera posture and / or offset to complete the training of the local dynamic neural radiation field; wherein the offset is the deviation between the camera posture range and the adjusted camera posture.

[0031] For each selected local dynamic neural radiation field, the following process is performed: a number of video frames are selected in sequence from the video frame sequence, the selected video frames and their corresponding camera poses are used as input, a reconstructed image corresponding to the video frame is generated through the local dynamic radiation field, and the reconstructed image is compared with the video frame, and the difference between the two is calculated. The difference is used to adjust the parameters of the current local dynamic neural radiation field and the camera pose corresponding to the video frame through back propagation, thereby improving the accuracy of view synthesis and the geometric consistency of three-dimensional modeling. After completing a local training, the deviation (i.e., offset) between the adjusted camera pose and the camera pose range preset by the current local dynamic neural radiation field is calculated, and whether the camera pose range needs to be adjusted is determined based on the adjusted camera pose or the adjusted camera pose and the difference. After adjusting the camera pose, the training of the current local dynamic neural radiation field is terminated, and a new local dynamic neural radiation field is started for modeling, thereby achieving segmented and stable modeling of the entire dynamic video.

[0032] Specifically, in one embodiment of the present invention, for each selected local dynamic neural radiation field, the processing steps of S13 include S131 to S134 (not shown in the figure):

[0033] S131. Select a number of video frames in sequence from a video frame sequence.

[0034] For the first selected local dynamic neural radiance field, during its initial training, a preset number of consecutive video frames can be selected from the first video frame of the video frame sequence as the first video frame subsequence, and the first local dynamic neural radiance field can be trained in combination with the corresponding camera pose. In subsequent iterations, unused video frames are sequentially selected from the video frame sequence and sequentially added to the first video frame subsequence to form a new first video frame subsequence to train the first local dynamic neural radiance field again.

[0035] For a selected non-first (Nth) local dynamic neural radiance field: During its initial training, there are two ways to select the first video frame from the video frame sequence to be input into the local dynamic neural radiance field: one is to select the video frame following the last video frame (denoted as the kth video frame) used at the end of training the previous (i.e., N-1th) local dynamic neural radiance field, and use it as the first video frame of the currently selected local dynamic neural radiance field; the other is to start from the kth video frame, backtrack a preset first number of video frames, and use the video frame at the end of the backtracking as the first video frame of the currently selected local dynamic neural radiance field. Starting from the selected first video frame, a preset number of consecutive video frames are reselected from the video frame sequence to form a new video frame subsequence, and the current local dynamic neural radiance field is trained based on its corresponding camera pose. In subsequent iterations, unused video frames are sequentially selected from the video frame sequence and added to the video frame subsequence to form a new video frame subsequence for retraining the current local dynamic neural radiance field.

[0036] S132. Input each selected video frame and the corresponding camera posture into the current local dynamic neural radiation field to generate a corresponding reconstructed image.

[0037] The following process is performed for each local dynamic neural radiation field: the selected video frame subsequence and the corresponding camera pose are input into the current local dynamic neural radiation field to generate a reconstructed image corresponding to the input video frame. The reconstructed image refers to an image synthesized after volume rendering of the spatial sampling points by the corresponding local dynamic neural radiation field based on the camera pose and time stamp corresponding to the input video frame, which is used to simulate the observation results of the original scene at that perspective and time. For example, in a handheld video of a city walk, the time stamp corresponding to the 120th frame is the 4th second, when the camera is facing the street directly in front. If the time stamp (4th second) and a new camera pose (for example, looking diagonally downward from the upper left) are input, the corresponding local dynamic neural radiation field will generate an image seen from this new perspective based on the existing training data. This image is the reconstructed image, which simulates the appearance of the scene at the same time but a different perspective.

[0038] Specifically, in one embodiment of the present invention, for each video frame, the process of step S132 includes:

[0039] S1321. Based on the camera posture corresponding to the video frame and the preset camera internal parameters, perform back projection processing on the video frame to obtain a ray sequence corresponding to the video frame.

[0040] The pixel coordinates are normalized to the camera coordinate system based on the camera intrinsic parameters, and then transformed to the world coordinate system based on the camera pose, thereby calculating a unique 3D ray for each pixel. These rays constitute the ray sequence corresponding to the current video frame.

[0041] S1322. For each ray in the ray sequence: uniformly sample the ray according to a preset depth range to obtain multiple spatial sampling points along the ray direction.

[0042] For each ray, the following operations are performed: according to the preset depth range, the ray is sampled uniformly at a fixed interval or using a layered sampling strategy, thereby obtaining multiple 3D spatial sampling points distributed along the ray direction. These spatial sampling points represent the possible voxel regions at different depth positions from the camera's perspective, serving as the input to the corresponding local dynamic neural radiation field.

[0043] S1323 , arranging the spatial sampling points of the ray sequence in sequence according to a sampling order to obtain a spatial sampling point sequence.

[0044] Each ray corresponds to a set of spatial sampling points. The spatial sampling points corresponding to all rays are combined in sequence to form a spatial sampling point sequence for the current video frame, which is used as the input of subsequent static and dynamic modules.

[0045] S1324. Determine a sight direction sequence corresponding to the spatial sampling point sequence according to the camera posture.

[0046] For each pixel in each video frame, the following processing is performed: the pixel's position in the video frame is back-projected using the previously acquired camera intrinsic parameters to convert the pixel from the pixel coordinate system to the camera coordinate system, and the converted value is used as the viewing direction of the pixel. After all pixels are processed, the spatial sampling point sequence corresponding to the entire video frame is obtained, and the viewing direction sequence of the display screen is obtained.

[0047] S1325. Input the spatial sampling point sequence, the corresponding sight direction sequence, and the time stamp corresponding to the video frame into the local dynamic neural radiation field to generate a reconstructed image corresponding to the video frame.

[0048] The local dynamic neural radiation field includes a static module and a dynamic module. The static module is used to model information that does not change over time, and the dynamic module is used to model information that changes dynamically over time, such as moving vehicles or people. The spatial sampling point sequence, the corresponding gaze direction sequence, and the time stamp corresponding to the video frame are input into the dynamic module, and the time stamp is used to generate the dynamic color, dynamic density, and dynamic mask of each pixel in the video frame. The spatial sampling point sequence and the corresponding gaze direction sequence are input into the static module, and the static color and static density of each pixel in the video frame are generated. The results of the two modules are fused to obtain the reconstructed image corresponding to the video frame.

[0049] Specifically, in one embodiment of the present invention, a spatial sampling point sequence and a corresponding gaze direction sequence are input into a local dynamic neural radiation field to generate a reconstructed image corresponding to a video frame, including the following process:

[0050] First, the sequence of spatial sampling points and the corresponding sequence of gaze directions are input into the static module of the local dynamic neural radiation field to obtain the static color and static density of each spatial sampling point. The static module does not rely on time information, thereby achieving modeling of static objects in the video. The static module may include a multi-layer perceptron, which encodes the input information to obtain the static color and static density of the spatial sampling point. The static color represents the appearance characteristics of the point at the current viewing angle, and the static density represents the contribution of the point to light penetration in three-dimensional space.

[0051] While obtaining the static color and static density, the sequence of spatial sampling points and the corresponding sequence of sight directions are also input into the dynamic module of the local dynamic neural radiation field, and the dynamic color, dynamic density and dynamic mask of each spatial sampling point are obtained according to the time mark of the corresponding video frame. The dynamic module relies on time information to achieve modeling of moving objects in the scene. The dynamic module can be a multi-layer perceptron or deformation field containing time coding, etc., as long as it can process the input information according to time and generate the dynamic color, dynamic density and dynamic mask of each spatial sampling point, there is no specific limitation. Among them, the dynamic color represents the appearance characteristics of the spatial sampling point at the current time and viewing angle, the dynamic density represents the contribution of the spatial sampling point to the penetration of light in three-dimensional space at the current time, and the dynamic mask is used to represent the weight of the spatial sampling point belonging to the dynamic area at the current time.

[0052] Then, for each spatial sampling point, the dynamic color, dynamic density, static color, static density, and dynamic mask corresponding to the spatial sampling point are weightedly fused to obtain the fused color and fused density corresponding to the spatial sampling point. Finally, the fused color and fused density of each spatial sampling point are cumulatively weighted according to their depth order within the corresponding ray to generate a reconstructed image corresponding to the video frame. Specifically, this is achieved by simulating the color contribution and occlusion effects of light as it passes through spatial sampling points in three-dimensional space, sequentially superimposing the influence of each sampling point until the ray reaches its endpoint or the color converges. At this point, the pixel value corresponding to the ray can be reconstructed. Once all pixel values are reconstructed, the resulting reconstructed image is obtained.

[0053] S133. Calculate the difference between each reconstructed image and the corresponding video frame, and adjust the corresponding camera posture and parameters of the local dynamic neural radiation field based on the difference.

[0054] In the present invention, in order to effectively process the complex motion content in dynamic scenes, the local dynamic neural radiation field is divided into a static module and a dynamic module, and a dynamic mask and a multi-source supervisory signal are introduced to jointly optimize the image reconstruction quality and the accuracy of camera pose estimation. Specifically, in one embodiment of the present invention, the dynamic module of the local dynamic neural radiation field also generates a dynamic mask corresponding to the video frame, and the difference between the video frame and the reconstructed image generated by the local dynamic neural radiation field is calculated in the following manner: based on the reconstructed image, the video frame and the dynamic mask, the static photometric difference and the dynamic photometric difference are calculated accordingly; based on the dynamic mask and the corresponding real mask obtained in advance, the mask difference is calculated; the static photometric difference, the dynamic photometric difference and the mask difference are weighted to obtain the difference.

[0055] like Figure 2 As shown in the figure, the selected video frame sequence and the corresponding camera posture are input into the local dynamic neural radiation field. After each pixel is back-projected to generate rays, it enters the static module and the dynamic module respectively. For each spatial sampling point of each input video frame, the following processing is performed: the static module generates the static color c of the spatial sampling point s and static density σ s , the dynamic module generates the dynamic color c of the spatial sampling point d and static density σ d and a dynamic mask m d In order to fuse the output results of the two modules, the dynamic mask m is used d The density and color of the two modules are weighted fused to obtain the fusion density σ of the spatial sampling point f and fusion color c f , as shown in formula (1) (2):

[0056] σ f =σ s ·(1-m d )+σ d ·m d (1)

[0057] c f =c s ·(1-m d )+c d ·m d (2)

[0058] After all spatial sampling points of the current video frame are fused with density and fused color by the above method, the reconstructed image of the video frame is obtained by volume rendering. In order to achieve differential training of dynamic modules and static modules, the present invention adopts multiple loss functions to update parameters. Specifically, the static photometric difference and dynamic photometric difference of the current video frame are calculated respectively by formula (3) and (4):

[0059]

[0060] in, is the static photometric difference of the current video frame, C s is the static image obtained by the static module of the current video frame, C gt is the current video frame (as a supervisory signal), M gt is the real mask of the current video frame obtained in advance, which is used to eliminate dynamic areas and supervise the static module only on static areas. is the dynamic luminosity difference of the current video frame, C d is the dynamic image obtained by the dynamic module of the current video frame. In addition, the present invention also calculates the mask difference between the dynamic mask of the current video frame and the real mask as shown in formula (5):

[0061]

[0062] in, is the mask difference, M d The dynamic mask of the current video frame obtained by the dynamic module. As a mask supervision signal, the weight can be gradually attenuated during the training process to prevent overfitting of the prior mask. The above differences are weighted and summed to obtain the total difference between the current video frame and the corresponding reconstructed image. It should be noted that, if Figure 2 To ensure the robustness of camera pose estimation, in the dynamic module, the gradient is truncated in the part indicated by the red cross in the figure, that is, the dynamic mask backpropagation is not allowed to affect the update process of the camera pose, and only the static module updates the camera pose.

[0063] Specifically, if Figure 2 As shown, the generation process of this scheme is as follows: In the progressive optimization scheme, the image on the left represents the selection of different video frames in sequence according to time sequence, and each video frame has a camera pose (such as the yellow triangle). For the current radiation field, video frames and their initial camera poses are gradually added in chronological order, and the optimization of the corresponding local dynamic neural radiation field and camera pose, as well as the addition of new video frames and corresponding camera poses are cyclically performed in each round of training. If the current viewing angle range exceeds the camera pose range corresponding to the radiation field, the radiation field parameters after the last update are retained, and the corresponding camera pose range is updated using the last adjusted camera pose. Then a new radiation field is added to repeat the above process. For each selected local dynamic neural radiation field, after the pixel points are back-projected to generate rays, the line of sight direction d, spatial sampling point x and time identifier t are input into the dynamic module of the radiation field, and the corresponding dynamic color c is generated through the multi-layer deformation perceptron and the non-shared multi-layer perceptron. d , dynamic density σ d And the predicted mask m d , thus obtaining the dynamic image and prediction mask corresponding to the entire video frame. Input the realization direction d and spatial sampling point x into the static module of the radiation field to generate the corresponding static color c s and static density σ s , thus obtaining the static image corresponding to the entire video frame. Perform weighted fusion of color and density to obtain the corresponding fusion density σ f and fusion color c f , and volume rendering is used to generate a complete reconstructed image. The static image generated by the static module, the dynamic image generated by the dynamic module, and the fused reconstructed image are compared with the real image, and mask restrictions are applied using the real mask, which serves as a supervision basis for the radiation field training phase. Furthermore, the distribution curves of static and dynamic density values along ray distances show that different depth positions contribute differently to image rendering.

[0064] S134. Calculate the camera pose range and the offset between the camera pose adjusted based on the difference, and adjust the camera pose range according to the adjusted camera pose and / or the offset to complete the training of the local dynamic neural radiation field.

[0065] It should be noted that the present invention does not divide the radiation field according to a fixed number of frames or a fixed angle, but adopts a dynamic division strategy based on the posture range. For example, assuming that the camera posture range of the currently trained local dynamic neural radiation field A is [0°, 35°], when processing frames with video frame postures of 0°, 15°, 30°, and 31° in sequence, the camera postures obtained after adjustment of these frames are all within the preset camera posture range of the local dynamic neural radiation field A (i.e., [0°, 35°]), so these frames all belong to A and participate in its modeling. When A processes the next video frame with a camera posture of 38°, if the camera posture of the frame after adjustment exceeds the camera posture range of A, a new local dynamic neural radiation field B will be initialized, and the new local dynamic neural radiation field B will be trained starting from this frame or several frames before this frame.

[0066] Furthermore, when the video frame input to the current local dynamic neural radiation field does not include the last video frame of the video frame sequence, step S34 includes the following process:

[0067] The offset between the adjusted camera pose and the corresponding camera pose range is calculated, and it is determined whether the offset is greater than a preset offset threshold. If the offset is greater than the preset offset threshold, the corresponding camera pose range is adjusted based on the adjusted camera pose to complete the training of the current local dynamic neural radiation field. Conversely, if the offset is less than or equal to the offset threshold, several video frames are reselected from the video frame sequence and the current local dynamic neural radiation field is trained again. Specifically, if the current input video frame does not include the last video frame of the entire video frame sequence, it means that the current stage is in the intermediate stage. At this time, the offset between the camera pose updated in the current training round and the camera pose range set by the current local radiation field is calculated. The offset is used to measure whether the current viewing angle is still within the valid range of the original modeling of the radiation field. It can be used as the offset by Euclidean distance or by determining whether the two values are consistent. If the offset is greater than the offset threshold, it means that the current adjusted camera pose has significantly deviated from the applicable range of the local radiation field. At this time, the corresponding camera pose range is updated based on the last adjusted camera poses, and the training of the current local dynamic neural radiation field is completed. For example, the range of the camera pose after the last adjustment can be used as the updated camera pose range. Of course, other adjustment methods can also be used, which will not be detailed here. Conversely, if the offset is less than or equal to the offset threshold, it means that the current camera trajectory can still be represented by the existing radiation field. At this time, several new frames are selected from the video frame sequence and added to the previously selected video frames. These video frames are input into the current radiation field together and the next round of optimization is performed until the offset termination condition is reached or the entire video frame sequence is traversed. This mechanism ensures the dynamic controllability of the model range and the stability of the modeling accuracy during training.

[0068] In addition, when the video frame input to the local dynamic neural radiation field includes the last video frame of the video frame sequence, the offset between the camera pose range and the camera pose adjusted based on the difference is calculated, and the camera pose range is adjusted according to the adjusted camera pose and / or the offset to complete the training of the local dynamic neural radiation field, including: adjusting the camera pose range based on the adjusted camera pose to complete the training of the current local dynamic radiation field. When the video frame input to the radiation field includes the last video frame, it is not possible to determine whether to continue training the radiation field based on the offset. Instead, after updating the parameters of the radiation field and the camera pose, the updated camera pose is directly used to adjust the corresponding camera pose range.

[0069] S14. Generate an image reconstruction model by parallelizing the trained multiple local dynamic neural radiation fields.

[0070] After training multiple local dynamic neural radiation fields through the above process, these independent radiation fields are integrated in parallel to construct a complete video image reconstruction model, which is used to support high-quality synthesis and reconstruction of scene images at any time and perspective.

[0071] like Figure 3 As shown, the present invention also provides an image reconstruction method, comprising:

[0072] S31, selecting any time marker from the video frame sequence to be processed, and presetting a corresponding camera posture for the time marker;

[0073] S32, determining a target video frame corresponding to the time identifier from the video frame sequence;

[0074] S33. Determine an adjusted camera posture range that matches the camera posture, and determine a corresponding local dynamic neural radiation field in an image reconstruction model based on the adjusted camera posture; wherein the image reconstruction model is obtained by any of the above-mentioned image reconstruction model generation methods;

[0075] S34. Input the target video frame and camera posture into the local dynamic neural radiation field corresponding to the image reconstruction model to generate a reconstructed image under the current time mark and camera posture.

[0076] Specifically, in the actual application stage, the user can select any time point from the video frame sequence to be processed as the target reconstruction moment, and preset a camera posture for the time mark to indicate the reconstruction effect of the scene from the perspective. According to the selected time mark, the corresponding target video frame is located from the video frame sequence as a reference for the image content at that moment. According to the input camera posture, the corresponding camera posture range is matched, and the local dynamic neural radiation field containing the posture range is searched in the image reconstruction model. Among them, the image reconstruction model is pre-generated by training multiple local dynamic neural radiation fields, and each local dynamic neural radiation field covers a certain camera motion trajectory and viewing angle range. The target video frame and the preset camera posture are input into the matching local dynamic radiation field together to generate a reconstructed image at that time point and observation angle, thereby realizing three-dimensional reproduction of the video content and synthesis of a new perspective.

[0077] like Figure 4 As shown, the image reconstruction model training system 400 includes: a data acquisition module 410, a radiation field determination module 420, a radiation field training module 430, and a model generation module 440. The data acquisition module 410 is used to obtain a video frame sequence to be processed and a corresponding camera pose sequence to be adjusted; wherein the video frame sequence is derived from a dynamic scene, and each video frame has a corresponding time stamp. The radiation field determination module 420 is used to select a corresponding local dynamic neural radiation field based on the camera pose sequence to be adjusted; wherein each local dynamic neural radiation field is preset with a corresponding camera pose range. The radiation field training module 430 is used to train the selected local dynamic neural radiation field based on the video frame sequence and the camera pose sequence to be adjusted, adjust the corresponding camera pose and local dynamic neural radiation field parameters based on the difference between the video frame and the reconstructed image generated by the local dynamic neural radiation field, and adjust the corresponding camera pose range based on the adjusted camera pose and / or offset to complete the training of the local dynamic neural radiation field; wherein the offset is the deviation between the camera pose range and the adjusted camera pose. The model generation module 440 is used to generate an image reconstruction model by parallelizing the multiple local dynamic neural radiation fields that have been trained.

[0078] For the specific definition of the image reconstruction model training system, please refer to the definition of the image reconstruction model training method above, which will not be repeated here. The various modules in the above-mentioned image reconstruction model training system can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware format, or can be stored in the memory of the computer device in software format to facilitate the processor to call the operations corresponding to the above modules.

[0079] It should be noted that, in order to highlight the innovative part of the present invention, this embodiment does not introduce modules that are not closely related to solving the technical problems raised by the present invention, but this does not mean that there are no other modules in this embodiment.

[0080] like Figure 5 As shown, the electronic device 5 may include a memory 51, a processor 52 and a bus, and may also include a computer program stored in the memory 51 and executable on the processor 52, such as an image reconstruction model generation program or an image reconstruction program.

[0081] The memory 51 includes at least one type of readable storage medium, including flash memory, mobile hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a mobile hard disk of the electronic device 5. In other embodiments, the memory 51 may also be an external storage device of the electronic device 5, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 5. Furthermore, the memory 51 may include both an internal storage unit of the electronic device 5 and an external storage device. The memory 51 can be used not only to store application software installed in the electronic device 5 and various types of data, such as the code for generating an image reconstruction model or the code for image reconstruction, but also to temporarily store data that has been output or is to be output.

[0082] In some embodiments, the processor 52 may be comprised of an integrated circuit, such as a single packaged integrated circuit or a plurality of packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 52 is the control core (Control Unit) of the electronic device 5, connecting the various components of the entire electronic device 5 using various interfaces and circuits. It executes programs or modules stored in the memory 51 (such as an image reconstruction model generation program or image reconstruction program), and calls data stored in the memory 51 to perform various functions of the electronic device 5 and process data.

[0083] The processor 52 executes the operating system and various installed application programs of the electronic device 5. The processor 52 executes the application programs to implement the above-mentioned method for generating an image reconstruction model or the steps in the image reconstruction method.

[0084] Exemplarily, the computer program can be divided into one or more modules, one or more of which are stored in the memory 51 and executed by the processor 52 to complete the image reconstruction model training method or image reconstruction method of the present application. One or more modules can be a series of computer program instruction segments that can perform specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 5. For example, the computer program can be divided into a data acquisition module 410, a radiation field determination module 420, a radiation field training module 430, and a model generation module 440.

[0085] The above-mentioned integrated unit implemented in the form of a software functional module can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, computer device, or network device, etc.) or a processor to execute the image reconstruction model generation method or part of the image reconstruction method of each embodiment of the present application.

[0086] In summary, the present invention discloses a training and image reconstruction method, system, device and medium for an image reconstruction model, which processes a video frame sequence and a corresponding camera posture sequence to be adjusted in a dynamic scene, and combines the gradual training and optimization of multiple local dynamic neural radiation fields to achieve high-quality modeling and image reconstruction of complex three-dimensional dynamic scenes. Among them, during the radiation field training process, the camera posture is adjusted according to the difference between the reconstructed image and the original video frame, thereby significantly improving the accuracy and robustness of the posture estimation. Furthermore, by setting the camera posture range of the local dynamic neural radiation field and dynamically updating the camera posture range in combination with the offset, each local dynamic neural radiation field can be more accurately adapted to the current viewing angle, greatly improving the overall rendering quality, so that the final reconstructed image can be more accurate. Therefore, the present invention effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.

[0087] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.

Claims

1. A method for generating an image reconstruction model, characterized in that: The generation method comprises: Obtaining a video frame sequence to be processed and a corresponding camera pose sequence to be adjusted; wherein the video frame sequence is derived from a dynamic scene, and each video frame has a corresponding time stamp; Selecting a corresponding local dynamic neural radiation field according to the camera posture sequence to be adjusted; wherein each local dynamic neural radiation field is preset with a corresponding camera posture range; The selected local dynamic neural radiation field is trained according to the video frame sequence and the camera pose sequence to be adjusted, and the corresponding camera pose and parameters of the local dynamic neural radiation field are adjusted according to the difference between the video frame and the reconstructed image generated by the local dynamic neural radiation field. The corresponding camera pose range is adjusted according to the adjusted camera pose and / or offset to complete the training of the local dynamic neural radiation field; wherein the offset is the deviation between the camera pose range and the adjusted camera pose; Multiple trained local dynamic neural radiation fields are parallelized to generate an image reconstruction model.

2. The method for generating an image reconstruction model according to claim 1, wherein: For each selected local dynamic neural radiation field, the selected local dynamic neural radiation field is trained according to the video frame sequence and the camera pose sequence to be adjusted, the corresponding camera pose and parameters of the local dynamic neural radiation field are adjusted according to the difference between the video frame and the reconstructed image generated by the local dynamic neural radiation field, and the corresponding camera pose range is adjusted according to the adjusted camera pose and / or offset to complete the training of the local dynamic neural radiation field, including: Selecting a plurality of video frames in sequence from the video frame sequence; Input each selected video frame and the corresponding camera pose into the current local dynamic neural radiation field to generate the corresponding reconstructed image; Calculate the difference between each reconstructed image and the corresponding video frame, and adjust the corresponding camera pose and local dynamic neural radiation field parameters based on the difference; The camera pose range and the offset between the camera pose adjusted based on the difference are calculated, and the camera pose range is adjusted according to the adjusted camera pose and / or the offset to complete the training of the local dynamic neural radiation field.

3. The method for generating an image reconstruction model according to claim 1, wherein: When the video frame input to the current local dynamic neural radiation field does not include the last video frame of the video frame sequence, calculating the offset between the camera pose range and the camera pose adjusted based on the difference, and adjusting the camera pose range according to the adjusted camera pose and / or the offset to complete the training of the local dynamic neural radiation field, including: Calculate the offset between the adjusted camera pose and the corresponding camera pose range, and determine whether the offset is greater than the preset offset threshold: If so, adjust the corresponding camera pose range based on the adjusted camera pose to complete the training of the current local dynamic neural radiation field; Otherwise, a number of video frames are reselected from the video frame sequence to train the current local dynamic neural radiation field again.

4. The method for generating an image reconstruction model according to claim 1, wherein: When the video frame input to the local dynamic neural radiation field includes the last video frame of the video frame sequence, the camera posture range and the offset between the camera posture adjusted based on the difference are calculated, and the camera posture range is adjusted according to the adjusted camera posture and / or the offset to complete the training of the local dynamic neural radiation field, including: adjusting the camera posture range based on the adjusted camera posture to complete the training of the current local dynamic radiation field.

5. The method for generating an image reconstruction model according to claim 1, wherein: For each video frame, the video frame and the corresponding camera pose are input into the corresponding local dynamic neural radiance field to generate the corresponding reconstructed image. The process includes: Based on the camera posture corresponding to the video frame and the preset camera internal parameters, the video frame is back-projected to obtain the ray sequence corresponding to the video frame; For each ray of the ray sequence: uniformly sampling the ray according to a preset depth range to obtain a plurality of spatial sampling points along the ray direction; Arranging the spatial sampling points of the ray sequence in sequence according to a sampling order to obtain a spatial sampling point sequence; Determine the sight direction sequence corresponding to the spatial sampling point sequence based on the camera posture; The spatial sampling point sequence, the corresponding sight direction sequence and the time stamp corresponding to the video frame are input into the local dynamic neural radiation field to generate a reconstructed image corresponding to the video frame.

6. The method for generating an image reconstruction model according to claim 5, wherein: The step of inputting the spatial sampling point sequence and the corresponding sight direction sequence into the local dynamic neural radiation field to generate a reconstructed image corresponding to the video frame includes: Input the spatial sampling point sequence and the corresponding sight direction sequence into the static module of the local dynamic neural radiation field to obtain the static color and static density of each spatial sampling point; The spatial sampling point sequence and the corresponding gaze direction sequence are input into the dynamic module of the local dynamic neural radiation field. According to the time stamp of the corresponding video frame, the dynamic color, dynamic density and dynamic mask of each spatial sampling point are obtained. For each spatial sampling point: perform weighted fusion on the dynamic color, dynamic density, static color, static density and dynamic mask corresponding to the spatial sampling point to obtain the fused color and fused density corresponding to the spatial sampling point; The fusion color and fusion density of each spatial sampling point are accumulated and weighted according to their depth order in the corresponding ray to generate a reconstructed image corresponding to the video frame.

7. An image reconstruction method, characterized in that: The method comprises: Select any time stamp from the video frame sequence to be processed and preset the corresponding camera pose for the time stamp; Determining a target video frame corresponding to the time identifier from the video frame sequence; Determining an adjusted camera posture range that matches the camera posture, and determining a corresponding local dynamic neural radiation field in an image reconstruction model based on the camera posture; wherein the image reconstruction model is obtained by the image reconstruction model generation method according to any one of claims 1 to 6; The target video frame and camera posture are input into the local dynamic neural radiation field corresponding to the image reconstruction model to generate a reconstructed image under the current time mark and camera posture.

8. A training system for an image reconstruction model, characterized in that: The system comprises: A data acquisition module, configured to acquire a sequence of video frames to be processed and a corresponding sequence of camera poses to be adjusted; wherein the sequence of video frames is derived from a dynamic scene, and each video frame has a corresponding time stamp; A radiation field determination module is used to select a corresponding local dynamic neural radiation field according to the camera posture sequence to be adjusted; wherein each local dynamic neural radiation field is preset with a corresponding camera posture range; A radiation field training module is used to train the selected local dynamic neural radiation field based on the video frame sequence and the camera posture sequence to be adjusted, adjust the corresponding camera posture and parameters of the local dynamic neural radiation field according to the difference between the video frame and the reconstructed image generated by the local dynamic neural radiation field, and adjust the corresponding camera posture range according to the adjusted camera posture and / or offset to complete the training of the local dynamic neural radiation field; wherein the offset is the deviation between the camera posture range and the adjusted camera posture; The model generation module is used to generate an image reconstruction model by parallelizing multiple trained local dynamic neural radiation fields.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the method for generating an image reconstruction model as described in any one of claims 1 to 6 or the image reconstruction method as described in claim 7.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the method for generating an image reconstruction model according to any one of claims 1 to 6 or the image reconstruction method according to claim 7.

Citation Information

Cited By

  • Video driving vector data three-dimensional and real-time drawing method and system based on digital twinborn scene

    CN121458907A

  • Video-driven vector data 3Dization and real-time rendering method and system based on digital twin scene

    CN121458907B