Method for constructing and editing three-dimensional light field content
By acquiring multi-view images to construct a three-dimensional spatial coordinate system, and employing a three-dimensional reconstruction and two-dimensional image diffusion editing network, combined with staggered rendering and light field coding, the problem of viewpoint consistency in three-dimensional light field content editing is solved, achieving efficient and controllable three-dimensional light field content generation and improving user experience.
Patent Information
- Application Number
- CN202511118719.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies struggle to flexibly and efficiently edit 3D light field content while maintaining continuity between viewpoints and structural consistency.
Multi-view images are captured by a camera to construct a unified three-dimensional spatial coordinate system. An initial scene model is generated using a three-dimensional reconstruction algorithm. A two-dimensional image diffusion editing network is used for editing. Combined with a spatial attention module and a geometric constraint module, an image sequence with consistent viewpoints is generated. Finally, three-dimensional light field content is generated through shear rendering and light field encoding.
It enables efficient and controllable 3D light field content editing, improves the visual experience and display effect, lowers the technical threshold, and enhances the sense of spatial depth and realism.
Smart Images

Figure CN120915930A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of naked-eye three-dimensional display, in particular to a method for constructing and editing three-dimensional light field content. BACKGROUND
[0002] With the continuous development of naked-eye three-dimensional display technology, three-dimensional light field display devices have attracted widespread attention in recent years because they can present three-dimensional visual effects with real occlusion relationships and dynamic changes in viewing angles without the need to wear any auxiliary devices. Such devices achieve natural reproduction of real three-dimensional scenes by synchronously displaying parallax information at multiple viewing angles, significantly enhancing the immersive viewing experience of users. In three-dimensional light field display, the texture quality, viewing angle consistency, and depth information of the presented three-dimensional content largely determine the final visual effects and user experience. Currently, three-dimensional modeling and reconstruction techniques for real people or static scenes are relatively mature, and can effectively generate basic three-dimensional data for display. However, it is still a great challenge to flexibly and high-quality edit three-dimensional light field content while maintaining inter-view continuity and structural consistency. SUMMARY
[0003] In view of the above deficiencies in the prior art, the present application provides a method for constructing and editing three-dimensional light field content, which solves the problem of being difficult to flexibly and high-quality edit three-dimensional light field content while maintaining inter-view continuity and structural consistency in the prior art.
[0004] To achieve the above-mentioned application purposes, the technical scheme adopted by the present application is as follows: A method for constructing and editing three-dimensional light field content is provided, which includes the following steps: Acquiring the internal and external parameter information corresponding to each frame of image by capturing multi-view images of a real scene through a camera, and then constructing a unified three-dimensional space coordinate system; Processing the multi-view images using a three-dimensional reconstruction algorithm to generate an initial three-dimensional scene model; Adjustably editing the content of each frame of multi-view image using a two-dimensional image diffusion editing network to obtain edited images; wherein the edited images maintain structural consistency and visual continuity among multiple views; Updating the initial three-dimensional scene model through the edited images to obtain an updated three-dimensional scene model; Rendering the updated three-dimensional scene model through a skew-cut rendering method to generate an image sequence with multi-view consistency, denoted as a generated image sequence; Encoding the generated image sequence into three-dimensional light field content using a light field encoding algorithm.
[0005] Furthermore, specific methods for obtaining the intrinsic and extrinsic parameter information corresponding to each frame of the image, and then constructing a unified three-dimensional spatial coordinate system, include: The camera calibration algorithm is used to estimate the intrinsic parameter matrix, extrinsic parameter rotation matrix, and translation matrix of the camera corresponding to each frame image to obtain the spatial pose parameters of the camera. Based on the feature point matching results between multi-view images, the spatial relative relationship between cameras is recovered through the spatial pose parameters of the cameras to construct a unified three-dimensional coordinate system. The camera calibration algorithm includes motion reconstruction structure algorithm and multi-view geometry algorithm.
[0006] Furthermore, the multi-view images are sorted before acquiring the 3D scene model and obtaining the edited images. Specific methods include: Select the world coordinate system x The camera with the largest axis coordinate value is selected as the reference camera; the line-of-sight direction vectors of the remaining cameras in the world coordinate system are calculated, and they are initially sorted in ascending order based on the angle between this vector and the line-of-sight direction vector of the reference camera; when the angles are the same or less than a threshold, the cameras are further sorted based on their position in the world coordinate system. x shaft and y The images are sorted in ascending order of their planar distance relative to the reference camera along the axial direction to obtain a spatially continuous multi-view image arrangement.
[0007] Furthermore, specific methods for processing multi-view images using 3D reconstruction algorithms to generate initial 3D scene models include: A 3D scene model reflecting the geometric and appearance features of a real scene is reconstructed using a 3D reconstruction algorithm based on multi-view images and their corresponding camera pose information.
[0008] Furthermore, the two-dimensional image diffusion editing network is a network that further embeds a spatial attention module and a geometric constraint module into an image editing network framework based on a diffusion model; wherein: The diffusion model includes an encoding module, a noise-adding module, a noise-reducing module, and a decoding module connected in sequence. The encoding module is used to encode each frame of multi-view images into latent spatial features; The noise-adding module is used to add noise to the latent spatial features generated by the encoding module, simulate the forward diffusion process, and obtain the latent spatial features of the multi-view sequence image after adding noise. The denoising module uses a conditional diffusion denoising network built on the U-Net network structure to denoise and edit the latent spatial features of multi-view sequence images after adding noise, based on user-provided guiding data, and generate denoised and edited latent spatial features. The decoding module is configured to convert the denoised latent image features generated by the denoising module back to the pixel space to obtain a multi-view edited image sequence that is consistent with the user's editing intention, has consistent structure, and has continuous view angles, i.e., an edited image. In the U-Net network structure of the denoising module, at least one of the self-attention modules of the Transform layers is replaced by a spatial attention module, and each spatial attention module is followed by a geometric constraint module. The spatial attention module takes the feature representation of the current view output by the structure at the front end and the feature representations of the remaining views as query vectors and key vectors, generates enhanced feature map information considering other views, to realize spatial information fusion across views and improve the perception of the current view to the context information of other views in the denoising process; wherein the enhanced feature map information considering other views is the current view feature map. The geometric constraint module is configured to calculate the geometric correspondence between the current view feature map and the adjacent view feature map based on the known relative pose between the current view camera and the adjacent view camera and in combination with the camera intrinsic parameters, extract corresponding features from the adjacent view based on the geometric correspondence, and update the feature representation of the current view in this region by weighted fusion based on the extracted features, to obtain the current view enhanced feature map, so as to realize the consistency of the enhanced features and guide the features at corresponding positions in different views to maintain alignment and consistency in geometric structure.
[0009] Further, the expression of the enhanced feature map information considering other views is:
[0010] wherein represents the t frame edited image; is the query vector of the t frame edited image; is the key vector of the t frame edited image; is the key vector of the frame image; is the key vector of the frame image; , and are the value vector of the frame image, the value vector of the frame image, and the value vector of the frame image, respectively; is the feature dimension; is a softmax function.
[0011] Further, the expression of updating the feature representation of the region in the current view based on the extracted features through weighted fusion is as follows:
[0012] wherein is the feature representation of the edited image in the i-th frame at pixel position t is the feature representation of the edited image in the i-th frame at pixel position is the updated feature representation of the edited image in the i-th frame at pixel position is the feature representation of the i-th frame image; is the feature representation of the i-th frame image; is the geometric transformation function that maps the pixel position of the edited image in the i-th frame to the corresponding position in the i-th frame image based on the geometric constraint; t is the geometric transformation function that maps the pixel position of the edited image in the i-th frame to the corresponding position in the i-th frame image based on the geometric constraint; is the geometric transformation function that maps the pixel position is the weighted fusion operation; is the pose of the edited image in the i-th frame; t is the pose of the i-th frame image; is the pose of the i-th frame image; is the neighboring frame set of the edited image in the i-th frame. t Further, the specific method of updating the initial three-dimensional scene model through the edited image includes: using the edited image as a constraint condition in the three-dimensional scene model updating process, combining the camera pose parameters corresponding to the edited image, rendering the current three-dimensional scene model and generating a reconstructed image under the corresponding view, constructing a loss function by comparing the similarity of the rendered reconstructed image and the edited image, and performing reverse updating and optimization on the representation parameters that control the geometric structure and appearance information in the current three-dimensional scene model according to the loss function to obtain an updated three-dimensional scene model.
[0013] Further, the specific method of rendering the updated three-dimensional scene model through the anamorphic rendering method includes: arranging a group of virtual cameras along a uniform horizontal axis direction at a preset interval to form a camera array; introducing a horizontal offset based on the original main viewing direction of each virtual camera to realize different view observations of the three-dimensional scene content during the rendering process of the updated three-dimensional scene model, and generate an image sequence with horizontal parallax transition and consistency.
[0014] Further, the specific method of encoding the generated image sequence into three-dimensional light field content using a light field encoding algorithm includes:
[0015] According to the optical parameters and the geometric structure of the light field display, a mapping relationship between a viewpoint and a display pixel is established; by analyzing the pixels of each viewpoint image in the generated image sequence, the projection position of each pixel on the light field display is calculated, that is, the three-dimensional light field content is obtained.
[0016] The beneficial effects of the present application are: 1、The present application can efficiently convert a real static scene into three-dimensional light field content with complete structure and continuous perspective through multi-view image acquisition, three-dimensional reconstruction and rendering processing.
[0017] 2、The present application realizes controllable and high-fidelity editing of light field content through a two-dimensional image diffusion editing network integrated with a spatial attention mechanism and geometric consistency constraint, effectively reduces the technical threshold of three-dimensional content modification, and improves user interaction experience.
[0018] 3、The present application generates an image sequence with continuous parallax through the cross-cut rendering, so that the generated content better adapts to the light field display device, has stronger spatial depth and reality, and significantly improves the visual experience and display effect of three-dimensional light field content. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a flowchart of the present method; Figure 2 is a structural diagram of the two-dimensional image diffusion editing network. DETAILED DESCRIPTION
[0020] The specific embodiments of the present application will be described below to facilitate understanding by those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, any changes within the spirit and scope of the present application defined and determined by the appended claims are obvious, and all applications utilizing the concept of the present application are within the scope of protection.
[0021] Example 1: As shown in the figure, the method for constructing and editing three-dimensional light field content comprises the following steps: Figure 1 S1, acquiring multi-view images of a real scene by a camera, obtaining the internal and external parameter information corresponding to each frame of image, and then constructing a unified three-dimensional coordinate system; S2, processing the multi-view images by a three-dimensional reconstruction algorithm to generate an initial three-dimensional scene model; S3, constructing a three-dimensional light field content by a three-dimensional light field rendering algorithm based on the three-dimensional scene model; S3, using a two-dimensional image diffusion editing network with a fusion space attention mechanism and geometric constraints, performing adjustable content editing on each frame of multi-view images to obtain edited images; wherein the edited images maintain structural consistency and visual continuity between multi-views; S4, updating the initial three-dimensional scene model through the edited images to obtain an updated three-dimensional scene model; S5, rendering the updated three-dimensional scene model through a skew cut rendering method to generate an image sequence with multi-view consistency, denoted as a generated image sequence; S6, encoding the generated image sequence into three-dimensional light field content using a light field encoding algorithm.
[0022] In step S1, the specific method of acquiring multi-view images of a real scene through a camera is: acquiring multi-view image data of a real scene in a non-motion state through a camera.
[0023] The camera can be a camera array or a movable single camera system. During acquisition, the camera should cover the typical visible range of the light field display in the horizontal or vertical direction to achieve good spatial continuity and depth restoration capability.
[0024] The real scene in a non-motion state means that the target scene should be in a static state and the scene lighting conditions should be uniform and stable during acquisition, avoiding brightness and color differences introduced by shadows, reflections or changes in lighting, thereby avoiding reducing the reconstruction quality of the three-dimensional model.
[0025] In step S1, the specific method of obtaining the internal and external parameter information corresponding to each frame of image and constructing a unified three-dimensional coordinate system includes: Using a camera calibration algorithm to estimate the internal parameter matrix, external rotation matrix and displacement matrix of the camera corresponding to each frame of image to obtain the spatial pose parameters of the camera; based on the feature point matching results between multi-view images, the spatial relative relationship between cameras is recovered through the spatial pose parameters of the camera to construct a unified three-dimensional coordinate system; wherein the camera calibration algorithm includes a structure from motion algorithm (Structure from Motion, SfM) and a multi-view geometry algorithm (Multi-view Geometry Calibration).
[0026] The internal parameter matrix and the external rotation matrix are used to describe the spatial position and orientation of the camera relative to the world coordinate system. The camera internal parameters include focal length, principal point coordinates, pixel scale factor, etc., which are used to reflect the internal projection relationship in the camera imaging process.
[0027] In this embodiment, in order to obtain an image sequence with high spatial continuity and provide a better input arrangement for subsequent 3D reconstruction and editing operations, the multi-view images are sorted before acquiring the 3D scene model and obtaining the edited images. The specific method includes: Select the world coordinate system x The camera with the largest axis coordinate value is selected as the reference camera; the line-of-sight direction vectors of the remaining cameras in the world coordinate system are calculated, and they are initially sorted in ascending order based on the angle between this vector and the line-of-sight direction vector of the reference camera; when the angles are the same or less than a threshold, the cameras are further sorted based on their position in the world coordinate system. x shaft and y The images are sorted in ascending order of their planar distance relative to the reference camera along the axial direction to obtain a spatially continuous multi-view image arrangement.
[0028] If reference camera Its position in the boundary coordinate system is For each frame of image, the corresponding camera According to the rotation matrix in its extrinsic parameters Calculate the unit vector of the camera's line of sight (i.e., optical axis) in the world coordinate system:
[0029] Similarly, the line of sight of the reference camera can be calculated as follows: Calculate the line-of-sight angle between each camera and the reference camera. As a preliminary sorting criterion, it is defined as:
[0030] in for The transpose of .
[0031] In step S2, the specific method for processing multi-view images using a 3D reconstruction algorithm to generate an initial 3D scene model includes: A 3D reconstruction algorithm is employed based on multi-view images and their corresponding camera pose information to reconstruct a 3D scene model that reflects the geometric and appearance features of the real scene. The 3D scene model is represented using either a dense point cloud representation or a Gaussian distributed radiant field representation, with the Gaussian distributed radiant field representation being preferred. This representation supports fast forward projection image synthesis, improving the efficiency of subsequent 3D light field rendering. Furthermore, each Gaussian point serves as an independent adjustable unit, facilitating adjustments to local regions while maintaining structural consistency.
[0032] In step S3, as Figure 2As shown, the two-dimensional image diffusion editing network is a network in which a spatial attention module and a geometric constraint module are further embedded in an image editing network framework based on a diffusion model; wherein: The diffusion model comprises an encoding module, a noise adding module, a denoising module and a decoding module connected in sequence. The encoding module is configured to encode each frame of multi-view images into latent space features. The noise adding module is configured to add noise to the latent space features generated by the encoding module to simulate a forward diffusion process, thereby obtaining multi-view sequence image latent space features with added noise. The denoising module is a conditional diffusion denoising network based on a U-Net network structure, configured to denoise the multi-view sequence image latent space features with added noise in combination with the guide data provided by the user, thereby generating denoised latent space features. The decoding module is configured to convert the denoised latent image features generated by the denoising module back to the pixel space, thereby obtaining a multi-view edited image sequence that meets the user's editing intention, has consistent structure and continuous view angles, i.e., an edited image. After the improvement, the down-sampling layer in the U-Net network structure of the denoising module comprises a convolution layer, a spatial attention module, a geometric constraint module, a cross-attention module and a feedforward network connected in sequence. The convolution layer is configured to perform convolution processing on the input features to obtain corresponding feature representations. The spatial attention module takes the feature representation of the current view output by the structure at the front end and the feature representations of the remaining views as query vectors and key vectors, generates enhanced feature map information considering other views, to realize spatial information fusion across views and improve the perception ability of the current view to the context information of other views in the denoising process; wherein the enhanced feature map information considering other views is the current view feature map. The geometric constraint module is configured to calculate the geometric correspondence between the current view feature map and the adjacent view feature map based on the known relative pose between the current view camera and the adjacent view camera and in combination with the camera intrinsic parameters, extract corresponding features from the adjacent view based on the geometric correspondence, update the feature representation of the current view in this region through weighted fusion based on the extracted features, and obtain the current view enhanced feature map, to realize the consistency of the enhanced features and guide the features at corresponding positions in different views to maintain geometric structure alignment and consistency.
[0033] After the geometric constraint module completes the alignment of the current viewpoint feature map, the resulting enhanced feature map is input into the cross-attention module for conditional fusion with the user-provided guidance data. This cross-attention module uses the feature vectors at each position in the current viewpoint enhanced feature map as queries and the feature vectors in the guidance data as key / value pairs. Through an attention weighting mechanism, it selectively fuses the guidance information, effectively injecting the user's editing intent into the current viewpoint enhanced feature map. The feature map after cross-attention processing is further input into subsequent modules, including residual blocks and upsampling modules, ultimately outputting the predicted noise corresponding to the current time step. Combining this predicted noise with the noise image at the current time step, the denoised feature map in the latent space is gradually recovered using an anti-diffusion update formula. Finally, after completing all anti-diffusion steps, the denoised latent image features are obtained.
[0034] The expression for generating enhanced feature map information that takes into account other perspectives is:
[0035] in Indicates the first t The frame-edited image is computed using a spatial attention mechanism to obtain enhanced feature map information that takes into account other viewpoints; For the first t The query vector of the frame being edited; For the first t The key vector of the frame being edited; For the first The key vector of a frame image; For the first The key vector of a frame image; , and The first The value vector of the frame image, the first The value vector of the frame image and the first The value vector of the frame image; For feature dimensions; This refers to the softmax function.
[0036] The expression for updating the feature representation of the region from the current viewpoint using a weighted fusion method based on the extracted features is:
[0037] in For the first t The frame is edited at the pixel position of the image. The generated feature representation, i.e., the current edited image at the pixel location. The updated feature representation; Indicates the first a feature representation of the frame image; representing a current frame based on geometric constraints t a pixel position of the edited image of the frame mapped to a corresponding position in the frame a geometric transformation function mapping a pixel position of the edited image of the frame representing a weighted fusion operation; a pose of the edited image of the frame t a pose of the edited image of the frame a pose of the frame image; a pose of the frame image; a pose of the edited image of the frame t a neighboring frame set of the edited image of the frame.
[0038] On the basis of knowing the internal and external parameter matrices of the corresponding cameras of the adjacent views, the fundamental matrix between the adjacent image pairs can be calculated , and the fundamental matrix calculation function is:
[0039] wherein: , , is the camera internal parameter, rotation matrix and translation vector corresponding to the frame image, i , , is the camera internal parameter, rotation matrix and translation vector corresponding to the frame image. Based on this, the geometric correspondence relationship between images is derived, that is, for a pixel point in the image j , the corresponding point in the image frame i should satisfy: j
[0040] After completing the spatial attention mechanism for the association processing of image features, further based on the above geometric constraint formula, the geometric corresponding points of the key positions in the current frame image in the adjacent frames are determined. By extracting the feature information of these corresponding positions, the feature representation of the current frame is updated by weighting, so as to guide the image editing process of the current frame. Thus, the structural consistency and spatial continuity of the view editing result are enhanced.
[0041] Through the spatial attention mechanism and the feature guiding strategy based on the camera geometric mapping relationship, the spatial structural consistency and the inter-view continuity of the multi-view image in the editing process are ensured.
[0042] In step S4, the specific method of updating the initial three-dimensional scene model through the edited image includes: The edited image is used as a constraint in the 3D scene model update process. Combined with the camera pose parameters corresponding to the edited image, the current 3D scene model is rendered and a reconstructed image from the corresponding viewpoint is generated. By comparing the similarity between the reconstructed image and the edited image, a loss function is constructed. Based on the loss function, the representation parameters controlling the geometric structure and appearance information in the current 3D scene model are updated and optimized in reverse. This allows the updated 3D scene model to accurately reflect the changes introduced by the image editing, resulting in the updated 3D scene model.
[0043] The loss function can be Structural Similarity Index (SSIM), L1 norm, or Perceptual Loss, etc. The loss function used in this embodiment is:
[0044] in Representation of a 3D scene model This represents the 3D scene model that has been optimized to achieve the best rendering result. Indicates the first Multi-view images after frame editing Indicates the first Camera parameters corresponding to the frame image Indicates the use of the current 3D scene model From the perspective The reconstructed image generated below, Similarity loss functions in the image domain include, but are not limited to, LPIPS, SSIM, etc. This represents the total number of view frames.
[0045] By reducing the inconsistency between the edited and rendered images, the representation parameters in the 3D scene model are iteratively adjusted so that the updated model can restore the content and structural features of the edited image as much as possible during re-rendering, thereby realizing the update operation of the 3D scene content.
[0046] In step S5, the specific method for rendering the updated 3D scene model using the staggered rendering method includes: A set of virtual cameras is arranged at preset intervals along a unified horizontal axis to form a camera array. A lateral offset is introduced based on the original main view direction of each virtual camera, thereby enabling different perspectives of the 3D scene content during the rendering process of the updated 3D scene model. This generates an image sequence with lateral parallax transitions and consistency, exhibiting continuous parallax variation characteristics. The generated image sequence, after being processed by a light field encoding algorithm, is used to present a stronger sense of spatial depth and realistic motion parallax effects on a light field display device.
[0047] In step S6, the specific method for encoding the generated image sequence into three-dimensional light field content using a light field coding algorithm includes: Based on the optical parameters and geometry of the light field display, a mapping relationship between viewpoints and display pixels is established. By analyzing the pixels of each viewpoint image in the generated image sequence, the projection position of each pixel on the light field display is calculated, thus obtaining the three-dimensional light field content. This method ensures that during viewing, viewers can receive image information from the corresponding perspective when observing from different angles, thereby achieving a realistic three-dimensional visual experience.
[0048] Example 2: This embodiment is a further extension of Embodiment 1. In this embodiment, the static scene to be captured is arranged in an indoor environment, ensuring that the scene remains stationary during shooting and that the ambient lighting conditions are stable and uniform. A regular camera is used to slowly move around the scene with a controllable trajectory, the movement path being a semi-circular arc with the scene center as the focal point, covering a horizontal viewing angle of no less than 60 degrees. After shooting, 80 image frames are cropped from the continuous video at fixed time intervals.
[0049] The acquired multi-view images were calibrated using the open-source tool COLMAP. The intrinsic parameter matrix of the camera corresponding to each frame was estimated, along with extrinsic parameters such as rotation matrix and translation vector in a unified 3D world coordinate system. After calibration, the camera parameters were exported in a standardized format and saved as a .txt file.
[0050] Based on the above calibration results, select the world coordinate system. x The camera with the largest x-axis coordinate value is used as the reference camera, and its position and imaging direction are extracted. Based on the extrinsic information of the other cameras, the viewing axis direction of the corresponding camera for each frame is extracted and compared with the viewing axis of the reference camera. The angle between the viewing axes is calculated as the first criterion for sorting. If the viewing axis angle between two cameras is less than 1 degree, their imaging directions are considered to be approximately the same, and the relative positions of the reference cameras in the x-axis and y-axis directions are then used.
[0051] After sorting the above multi-view images, the 3D Gaussian Splatting (3DGS) algorithm is used for reconstruction. The input is the sorted image sequence and its corresponding camera pose information. After initializing the Gaussian point distribution, the spatial position, color, opacity, orientation and scale of each Gaussian point are optimized by differentiable rendering. The loss function adopts L1 loss and structural similarity index (SSIM).
[0052] An adjustable image content editing is performed on the sorted multi-view image sequence by using a two-dimensional image diffusion editing network with a fusion space attention mechanism and geometric constraints. The two-dimensional image diffusion editing network is constructed based on an InstructPix2Pix architecture. In the diffusion denoising process, the model introduces a cross-frame space attention mechanism to perform similarity calculation on the local features between the current frame and its adjacent multiple frames. At the same time, the fundamental matrix between the image pairs is calculated by using the camera extrinsic parameters of each frame image to obtain the epipolar constraint region, and the features at the geometrically corresponding positions are preferentially referenced during feature interaction. The input image resolution of the network is 512*512, the diffusion step number is set to 25 steps, the text prompt is encoded by CLIP to guide the denoising process, and finally an image sequence that meets the editing intention and is continuous and coordinated in the perspective is generated.
[0053] According to the camera pose corresponding to each frame image, the three-dimensional scene model is rendered to generate images under the corresponding perspective, and is compared with the edited images. By calculating the difference between the L1 loss and the perceptual loss, a joint loss function is constructed to optimize the parameters controlling the geometry and appearance in the Gaussian model in the reverse direction. The Adam optimizer is used in the optimization process, and the learning rate is set to 0.001.
[0054] The camera zero plane is set as the average value of the three-dimensional Gaussian point cloud in the z-axis direction, the rendering perspective range is set to 60 degrees, 80 virtual cameras are arranged at equal intervals along the horizontal direction, and the virtual cameras introduce lateral translation on the basis of the original main viewing direction to generate an image sequence with continuous parallax.
[0055] The 80 parallax images generated by the anamorphic rendering are encoded into content suitable for three-dimensional light field display by using a light field synthesis encoding algorithm. In the encoding process, the mapping relationship between the viewpoint and the screen pixel is established according to the optical parameters of the used light field display, that is, the three-dimensional light field content construction and editing are completed.
[0056] In summary, the application can efficiently convert a real static scene into a three-dimensional light field content with complete structure and continuous perspective through multi-view image acquisition, three-dimensional reconstruction and rendering processing. By using the two-dimensional image diffusion editing network with the fusion space attention mechanism and geometric consistency constraint, controllable and high-fidelity editing of the light field content is realized, the technical threshold of three-dimensional content modification is effectively reduced, and the user interaction experience is improved. The image sequence with continuous parallax is generated by anamorphic rendering, so that the generated content is better adapted to the light field display device, has stronger spatial depth and reality, and significantly improves the visual experience and display effect of the three-dimensional light field content.
Claims
1. A method of three-dimensional light field content construction and editing, characterized by, The method comprises the following steps: acquiring the internal parameter and external parameter information corresponding to each frame of image by capturing multi-view images of a real scene through a camera, and then constructing a unified three-dimensional space coordinate system; processing the multi-view images by using a three-dimensional reconstruction algorithm to generate an initial three-dimensional scene model; using a two-dimensional image diffusion editing network to perform adjustable content editing on each frame of multi-view image to obtain edited images; the edited images maintain structural consistency and visual continuity among the multi-view images; updating the initial three-dimensional scene model through the edited images to obtain an updated three-dimensional scene model; rendering the updated three-dimensional scene model through a cross-cut rendering method to generate an image sequence with multi-view consistency, denoted as a generated image sequence; encoding the generated image sequence into three-dimensional light field content by using a light field encoding algorithm.
2. The method of claim 1, wherein, The specific method for acquiring the internal parameter and external parameter information corresponding to each frame of image and then constructing a unified three-dimensional space coordinate system comprises: estimating the internal parameter matrix, external parameter rotation matrix and displacement matrix of the camera corresponding to each frame of image by using a camera calibration algorithm to obtain the spatial pose parameters of the camera; based on the feature point matching results between the multi-view images, the spatial relative relationship between the cameras is recovered through the spatial pose parameters of the cameras to construct a unified three-dimensional coordinate system; wherein the camera calibration algorithm comprises a motion recovery structure algorithm and a multi-view geometry algorithm.
3. The method of claim 1, wherein, The specific method for sorting the multi-view images before acquiring the three-dimensional scene model and obtaining the edited images comprises: Select the world coordinate system x The camera with the largest axis coordinate value is selected as the reference camera; the line-of-sight direction vectors of the remaining cameras in the world coordinate system are calculated, and they are initially sorted in ascending order based on the angle between this vector and the line-of-sight direction vector of the reference camera; when the angles are the same or less than a threshold, the cameras are further sorted based on their position in the world coordinate system. x shaft and y The images are sorted in ascending order of their planar distance relative to the reference camera along the axial direction to obtain a spatially continuous multi-view image arrangement.
4. The method of claim 1, wherein, The specific method for processing the multi-view images by using a three-dimensional reconstruction algorithm to generate an initial three-dimensional scene model comprises: reconstructing a three-dimensional scene model reflecting the geometric and appearance features of the real scene based on the multi-view images and the corresponding camera pose information by using a three-dimensional reconstruction algorithm.
5. The method of claim 1, wherein, The two-dimensional image diffusion editing network is a network further embedding a spatial attention module and a geometric constraint module in a diffusion model-based image editing network framework; wherein: the diffusion model comprises an encoding module, a noise adding module, a noise removing module and a decoding module connected in sequence; the encoding module is used for encoding each frame of multi-view image into latent space features; the noise adding module is used for adding noise to the latent space features generated by the encoding module to simulate a forward diffusion process and obtain latent space features of the multi-view sequence image after adding noise; the noise removing module adopts a conditional diffusion noise removing network constructed based on a U-Net network structure and is used for removing noise from the latent space features of the multi-view sequence image after adding noise in combination with guide data provided by a user to generate de-noised latent space features; the decoding module is used for converting the de-noised latent image features generated by the noise removing module back to a pixel space to obtain a multi-view edited image sequence in accordance with the user's editing intention, structural consistency and visual continuity, i.e. to obtain edited images; wherein the self-attention module of at least one Transform layer in the U-Net network structure of the noise removing module is replaced by a spatial attention module, and each spatial attention module is introduced with a geometric constraint module. The spatial attention module takes the feature representation of the current view and the feature representation of the remaining views as query vectors and key vectors, generates enhanced feature map information considering other views, to realize spatial information fusion across views and improve the perception ability of the current view to the context information on other views in the denoising process; wherein the enhanced feature map information considering other views is the current view feature map; The geometric constraint module is used to calculate the geometric correspondence between the current view feature map and the adjacent view feature map based on the known relative pose between the current view camera and the adjacent view camera, combined with the camera intrinsic parameters, and extract corresponding features from the adjacent view according to the geometric correspondence, update the feature representation of the current view in this region through weighted fusion based on the extracted features, and obtain the current view enhanced feature map, to realize the consistency of the enhanced features and guide the features at corresponding positions in different views to maintain geometric structure alignment and consistency.
6. The method of claim 5, wherein, The expression for generating enhanced feature map information considering other views is: wherein represents the t enhanced feature map information considering other view angles calculated by the spatial attention mechanism of the edited image of the frame; t query vector of the edited image of the frame; t key vector of the edited image of the frame; key vector of the image of the frame; key vector of the image of the , and are the value vector of the image of the frame, the value vector of the image of the frame, and the value vector of the image of the frame, respectively; is the feature dimension; is a softmax function.
7. The method of claim 5, wherein, The expression for updating the feature representation of the current view in this region through weighted fusion based on the extracted features is: in For the first t The frame is edited at the pixel position of the image. The generated feature representation, i.e., the current edited image at the pixel location. The updated feature representation; Indicates the first Feature representation of a frame image; This indicates that the current number is determined based on geometric constraints. t Pixel position of the frame being edited Mapped to the The geometric transformation function at the corresponding position in the frame image; This indicates a weighted fusion operation; For the first t The pose of the frame being edited; For the first The pose of the frame image; For the first t The set of neighboring frames of the image being edited.
8. The method of claim 1, wherein, The specific method for updating the initial three-dimensional scene model through the edited image includes: The edited image is used as a constraint condition in the three-dimensional scene model updating process, and the corresponding camera pose parameters of the edited image are combined to render the current three-dimensional scene model and generate a reconstructed image under the corresponding view, a loss function is constructed by comparing the similarity of the rendered reconstructed image and the edited image, and the representation parameters controlling the geometric structure and appearance information in the current three-dimensional scene model are updated and optimized in reverse according to the loss function, to obtain an updated three-dimensional scene model.
9. The method of claim 1, wherein, The specific method for rendering the updated three-dimensional scene model through the skew cut rendering method includes: A set of virtual cameras are arranged along a uniform horizontal axis direction at a preset interval to form a camera array; a horizontal offset is introduced based on the original main viewing direction of each virtual camera, so that different view observations of the three-dimensional scene content are realized during the rendering process of the updated three-dimensional scene model, and an image sequence with horizontal parallax transition and consistency is generated.
10. The method of claim 1, wherein, The specific method for encoding the generated image sequence into three-dimensional light field content using a light field encoding algorithm includes: According to the optical parameters and geometric structure of the light field display, the mapping relationship between the view point and the display pixel is established; by analyzing the pixels of each view image in the generated image sequence, the projection position of each pixel on the light field display is calculated, i.e. the three-dimensional light field content is obtained.
Citation Information
Cited By
Multi-view three-dimensional virtual-real fusion rendering method and related equipment
CN121600228A
Text-driven three-dimensional Gaussian scene editing method without training
CN122090022A