Dynamic three-dimensional scene reconstruction and rendering system, method and related equipment

By using image encoder, scene generator and scene decoder in the dynamic three-dimensional scene reconstruction and rendering system, a three-plane structure representation aligned with the camera's viewing angle is constructed and time-encoded, which solves the problems of weak generalization ability and strong dependence on prior information in the prior art, and achieves high-quality dynamic three-dimensional scene modeling and rendering.

CN120198601AInactive Publication Date: 2025-06-24SHANG HAI JIE YUE XING CHEN ZHI NENG KE JI YOU XIAN GONG SI

Patent Information

Application Number
CN202510677253.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing dynamic three-dimensional modeling methods can achieve high-quality reconstruction in specific environments, but the overall generalization ability is weak, strong dependence on prior information, and it is difficult to adapt to the rapidly changing scenario needs.

Method used

A dynamic three-dimensional scene reconstruction and rendering system is adopted to extract visual features of monocular video sequences through an image encoder. The scene generator constructs a three-planar structure representation that is aligned with the camera's viewing angle based on the extracted features and performs time encoding. The scene decoder generates a synthetic image based on the volume rendering mechanism.

Benefits of technology

It improves the system's generalization ability in unknown or rapidly changing scenarios, reduces the dependence on multi-view input and prior scene information, and realizes high-quality dynamic three-dimensional scene modeling and rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198601A_ABST
    Figure CN120198601A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic three-dimensional scene reconstruction and rendering system and method and related equipment, and the method comprises the steps: receiving a monocular video sequence through an image encoder, and carrying out the visual feature extraction of each frame of image; the scene generator constructs a three-plane structure expression aligned with the view angle of the current camera based on the multi-frame visual features extracted by the image encoder, performs time coding on the multi-frame visual features and maps the multi-frame visual features to the three-plane structure expression to generate a three-dimensional scene feature expression; and the scene decoder decodes the three-dimensional scene feature representation and generates a composite image under a corresponding view angle based on a volume rendering mechanism, and the composite image can display geometric details of the scene under different view angles and maintain geometric consistency under different view angles. Modeling and rendering of the dynamic three-dimensional scene are achieved through the monocular video sequence, multi-view input or prior scene information does not need to be depended on, and the adaptive capacity and generalization performance of the system in a rapid change environment are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a dynamic three-dimensional scene reconstruction and rendering system, method, and related devices. Background Art

[0002] With the rapid development of technologies such as augmented reality (AR), virtual reality (VR), and autonomous driving, the dynamic modeling and real-time rendering of three-dimensional scenes have become one of the key technologies in intelligent perception systems. As a neural rendering technology that has emerged in recent years, Neural Radiance Fields (NeRF) can perform implicit modeling of the geometry and appearance of a scene by using multi-layer perceptrons (MLPs), enabling high-quality image synthesis from arbitrary viewpoints, greatly promoting the development of three-dimensional reconstruction technologies, and being widely used in dynamic scene reconstruction.

[0003] On this basis, related research has begun to extend NeRF to the field of dynamic scene modeling. Typical methods include D-NeRF, Nerfies, DynIBaR, and MonoNeRF, etc. These methods usually map the observation viewpoints to a normalized scene space through a volume deformation field, and then model the dynamic changes in the scene. However, such methods generally rely on multi-view image inputs or require customized optimization for specific scenes during model training, limiting their applicability in single-camera inputs or rapidly changing environments.

[0004] To reduce the dependence on multiple views, subsequent research has attempted to achieve dynamic scene reconstruction based on monocular videos, but a large amount of prior information is still inevitably introduced in actual implementation. For example, methods such as DynIBaR and MonoNeRF usually rely on two-dimensional masks (masks) to distinguish dynamic and static scenes in three-dimensional space, or use optical flow techniques to achieve three-dimensional point tracking between frames. The above priors are difficult to accurately obtain in open scenes, resulting in significantly insufficient generalization ability of the model in unseen environments.

[0005] In summary, existing dynamic three-dimensional modeling methods can achieve high-quality reconstruction in specific environments, but overall, there are still problems such as weak generalization ability, strong dependence on prior information, and difficulty in adapting to the scene requirements of rapid changes. Therefore, there is an urgent need for a more general, data-driven, and dynamic modeling solution that does not require specific scene optimization to meet the real-time modeling and rendering requirements in variable scenes. Summary of the Invention

[0006] Aiming at the deficiencies of the existing technology, this application provides a dynamic three-dimensional scene reconstruction and rendering system, method, and related devices, aiming to improve the flexibility and robustness in the process of dynamic three-dimensional scene modeling, thereby enhancing the generalization ability of the system in unknown or rapidly changing scenes.

[0007] To achieve the above objectives and other advantages, the present application adopts the following technical solutions: In a first aspect, the present application provides a dynamic three-dimensional scene reconstruction and rendering system, including: An image encoder that receives a monocular video sequence and extracts visual features from each frame of the image; A scene generator that constructs a three-plane structure representation aligned with the current camera view based on the multi-frame visual features extracted by the image encoder, and maps the multi-frame visual features to the three-plane structure representation after temporal encoding to generate a three-dimensional scene feature representation; A scene decoder that decodes the three-dimensional scene feature representation and generates a synthesized image corresponding to the view based on a volume rendering mechanism. The synthesized image can display the geometric details of the scene from different views and maintain geometric consistency across different views.

[0008] According to a dynamic three-dimensional scene reconstruction and rendering system provided by the present application, the image encoder includes: A backbone network for performing multi-layer convolution operations on each frame of the image in the monocular video sequence to extract a deep visual feature map; A self-attention layer arranged after the backbone network for globally modeling the deep visual feature map to capture long-range dependencies.

[0009] According to a dynamic three-dimensional scene reconstruction and rendering system provided by the present application, after receiving the deep visual feature map, the self-attention layer performs position encoding processing on the deep visual feature map using a two-dimensional sine position encoding mechanism to embed spatial position information, so that each feature point carries spatial position information when participating in attention calculation, to further enhance the model's ability to model long-range dependencies.

[0010] According to a dynamic three-dimensional scene reconstruction and rendering system provided by the present application, the scene generator includes: A three-plane feature representation module for constructing three mutually perpendicular feature planes parallel to the XY, YZ, and ZX planes respectively. The feature planes take the camera center as the coordinate origin and dynamically adjust the direction and position according to the camera parameters of the input video, and store the geometric consistency features of images from different views as the three-plane structure representation aligned with the camera view; A temporal feature modeling module for performing temporal encoding processing on the multi-frame visual features, fusing the inter-frame dynamic change information, and then projecting the fused temporal features to the three-plane structure representation to construct the three-dimensional scene feature representation containing temporal dynamic information.

[0011] A dynamic three-dimensional scene reconstruction and rendering system provided by the present application, the scene generator further includes: An axial feature enhancement module for aggregating feature information on the same axis along the X, Y, or Z axis direction in each feature plane of the three-plane structure, and strengthening the information transmission along the axis through convolution operations to enhance the spatial continuity of the three-dimensional scene feature representation; A plane feature fusion module for establishing cross-plane feature interaction relationships between the three feature planes, and reallocating information between the three feature planes through a cross-attention mechanism to enhance the spatial consistency and overall expression ability of the three-dimensional scene feature representation.

[0012] A dynamic three-dimensional scene reconstruction and rendering system provided by the present application, the scene decoder includes: An upsampler for upsampling the low-resolution feature map of the three-dimensional scene constructed by the three-dimensional scene feature representation to generate a high-resolution three-plane feature map; A three-dimensional feature encoding module for extracting corresponding joint features from the high-resolution three-plane feature map for any position point in three-dimensional space, and performing sine position encoding in combination with the spatial coordinates of the position point to construct an input vector; A neural radiance field decoder for receiving the input vector and calculating the color value and density value of the position point in three-dimensional space through a multi-layer perceptron; A volume rendering module for calculating color and density information for multiple spatial sampling points on each ray based on the ray tracing path of the target view, performing weighted integration on the spatial sampling points through a volume rendering formula to obtain the color output of the ray, and arranging and combining the color outputs of each ray according to the corresponding image pixel positions to generate a synthesized image corresponding to the view, and the synthesized image represents the rendering result of the current three-dimensional scene from the target view.

[0013] In a second aspect, the present application provides a reconstruction and rendering method applicable to any one of the above systems, the method includes: Receiving a monocular video sequence and performing visual feature extraction on each frame of image; Based on the extracted multi-frame visual features, constructing a three-plane structure representation aligned with the current camera view, and mapping the multi-frame visual features to the three-plane structure representation after time encoding to generate a three-dimensional scene feature representation; Decoding the three-dimensional scene feature representation and generating a synthesized image corresponding to the view based on a volume rendering mechanism, and the synthesized image can display the geometric details of the scene from different views and maintain geometric consistency.

[0014] In a third aspect, the present application provides an electronic device, which includes: One or more processors; and a memory storing computer program instructions, which when executed cause the processors to execute the dynamic three-dimensional scene reconstruction and rendering method as described above.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, on which computer programs and / or instructions are stored, and when the computer programs and / or instructions are executed by a processor, the dynamic three-dimensional scene reconstruction and rendering method as described above is implemented.

[0016] In a fifth aspect, the present application provides a computer program product, including computer programs and / or instructions, which when executed by a processor implement the dynamic three-dimensional scene reconstruction and rendering method as described in any one of the above.

[0017] A dynamic three-dimensional scene reconstruction and rendering system, method and related device provided by the present application receive a monocular video sequence through an image encoder and extract visual features from each frame of the image; a scene generator constructs a three-plane structure representation aligned with the current camera view based on the multi-frame visual features extracted by the image encoder, and maps the multi-frame visual features to the three-plane structure representation after temporal encoding to generate a three-dimensional scene feature representation; a scene decoder decodes the three-dimensional scene feature representation and generates a synthesized image corresponding to the view based on a volume rendering mechanism. The synthesized image can display the geometric details of the scene from different views and maintain geometric consistency from different views. The present application realizes the modeling and rendering of a dynamic three-dimensional scene through a monocular video sequence, without relying on multi-view input or prior scene information, effectively improving the adaptability and generalization performance of the system in a rapidly changing environment. Using a three-plane structure to represent the three-dimensional space and combining a temporal encoding mechanism can accurately capture the spatio-temporal dynamic features in the scene and maintain the consistency of geometric information. By using a volume rendering mechanism to synthesize images from any view, clear and structurally consistent scene details can be presented at different viewing angles, thus significantly improving the rendering quality and modeling efficiency, and being applicable to dynamic vision task scenarios such as augmented reality, virtual reality and autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other implementation manners can be obtained according to these drawings without creative efforts.

[0019] Figure 1It is a schematic diagram of the logical structure of a dynamic three-dimensional scene reconstruction and rendering system provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a dynamic three-dimensional scene reconstruction and rendering method provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0020] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, is described in detail as follows.

[0021] It should be noted that those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict. Unless otherwise defined, the technical terms or scientific terms involved in the present application should be the general meanings understood by those with ordinary skills in the technical field to which the present application belongs. The terms "a", "one", "kind", "the" and other similar words involved in the present application do not indicate a quantity limitation and can represent a single or plural number. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; the terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0022] For the convenience of understanding the embodiments of the present application, the following are the explanations of the key terms / technical abbreviations in the present application: The deep convolutional neural network is a type of artificial neural network structure widely used in image recognition, computer vision, and pattern recognition tasks. Its core structure consists of multiple convolutional layers, activation functions, pooling layers, and fully connected layers, and can automatically extract hierarchical spatial features from the original image. By introducing multiple stacked convolutional layers, the network can learn complex representation forms from low-level edge information to high-level semantic information. This network structure has the characteristics of parameter sharing and local receptive fields, can effectively reduce the number of model parameters and improve the training efficiency, and is especially suitable for the processing and modeling of large-scale image data.

[0023] Neural Radiance Field (NeRF) is a neural network-based 3D scene representation method. By training a multi-layer perceptron model, it maps any 3D position and viewing direction in space to color and density values, and combines a volume rendering mechanism to generate images from new viewpoints. This technology enables high-fidelity reconstruction and realistic rendering of complex scenes.

[0024] Multi-Layer Perceptron (MLP) is a feedforward artificial neural network consisting of an input layer, one or more hidden layers, and an output layer. Each layer performs a non-linear transformation on the output of the previous layer through an activation function, used to approximate any complex functional relationship. It is used as a neural radiance field decoder in this application to predict the color and density information of spatial position points.

[0025] Sine position encoding is an encoding method for representing the position information of elements in space or sequences, commonly used in neural network models without position perception capabilities, such as the Transformer architecture. This method maps the input position coordinates to sine and cosine function values at different frequencies to form a continuous and generalizable position information representation, enabling the neural network to perceive position differences. Different from learnable embedding vectors, sine position encoding has the characteristics of being fixed, resolvable, and highly extrapolable, suitable for processing variable-length inputs or continuous space modeling tasks, and is widely used in fields such as image processing, natural language processing, and 3D reconstruction.

[0026] Ray tracing is a commonly used image generation technique that models based on the propagation path of light. By casting rays from the viewpoint along the pixel direction and interacting with the scene in 3D space for calculation, images are synthesized. This method can simulate physical processes such as light reflection, refraction, and shadows, and is widely used in fields such as realistic image rendering, 3D reconstruction, and virtual reality. In volume model-based representations, ray tracing is usually used to determine the spatial sampling path, providing input for subsequent image synthesis.

[0027] Volume rendering is an image generation method for displaying continuous media or dense field data in 3D space. This method samples multiple points along the ray direction in 3D volume data, estimates the attributes (such as color, density, transparency) of each point, and synthesizes the color output of the ray through weighted integration. Volume rendering can express 3D content with translucency, complex structures, and gradual changes, and is widely used in fields such as medical imaging, scientific visualization, and neural rendering.

[0028] A monocular camera refers to a camera device with a single imaging perspective. Different from a stereo or multi-camera system, it can only capture video frames from a single perspective. In this application, the system performs dynamic 3D scene modeling based on a continuous image sequence captured by a monocular camera, which has the advantages of simple hardware, low cost, and easy deployment.

[0029] Referring to Figure 1 As shown, an embodiment of this application provides a dynamic 3D scene reconstruction and rendering system, including: An image encoder that receives a monocular video sequence and extracts visual features from each frame of the image.

[0030] As an example, the image encoder is used to perform visual feature extraction on each frame of the input monocular video sequence. This encoder adopts a deep convolutional neural network structure, specifically including a backbone network and a self-attention layer.

[0031] In this embodiment, the image encoder specifically includes: A backbone network that performs multi-layer convolution operations on each frame of the monocular video sequence to extract deep visual feature maps; A self-attention layer, arranged after the backbone network, which is used to perform global modeling on the deep visual feature maps to capture long-range dependency relationships.

[0032] As an example, the image encoder is used to perform visual feature extraction on each frame of the input monocular video sequence. This encoder adopts a deep convolutional neural network structure, specifically including a backbone network and a self-attention layer.

[0033] The backbone network is a residual network structure similar to ResNet, which is used to perform layer-by-layer semantic feature extraction on the input image and improve the stability of deep information transmission and the training convergence speed through a residual connection mechanism. This backbone network includes three downsampling layers and nine residual modules. Each residual module contains two or three convolutional layers and is equipped with batch normalization and non-linear activation functions.

[0034] Exemplarily, for a single-frame image with an input size of 256×256×3, the initial convolutional layer includes a 7×7 convolutional layer (stride 2) and a max pooling layer, which are used to initially extract edge and texture features. The output feature map size is reduced from the input 256×256×3 to 128×128×64. The first downsampling stage uses a max pooling operation (3×3, stride 2) for the first spatial downsampling, obtaining a feature map size of 64×64×64, and then connects to 3 residual modules. Each residual module contains two consecutive 3×3 convolutional layers (with the number of channels unchanged), and uses an identity mapping for cross-layer connection. The second downsampling stage uses a convolutional layer with a stride of 2 to achieve spatial downsampling, obtaining a feature map size of 32×32×128, and also connects to 3 residual modules. The number of channels is expanded to 128, and a residual connection structure is also adopted. The third downsampling stage performs spatial downsampling again, reducing the feature map size to 16×16×256, connecting to 3 residual modules with the number of channels being 256, and continuing to use residual connection. Finally, a deep visual feature map with a size of 16×16×256 is obtained, providing rich context information for the subsequent self-attention mechanism.

[0035] In this embodiment, after receiving the deep visual feature map, the self-attention layer uses a two-dimensional sine position encoding mechanism to perform position encoding processing on the deep visual feature map to embed spatial position information, so that each feature point carries spatial position information when participating in attention calculation, thereby further enhancing the model's ability to model long-range dependence relationships.

[0036] To enhance the model's ability to model long-range dependence relationships in images, this embodiment introduces a self-attention layer after the backbone network. Based on the two-dimensional sine position encoding mechanism, this self-attention layer enables each feature point to carry explicit position information in the attention mechanism, thereby effectively capturing the global context information between different regions in the image and breaking through the modeling limitation of traditional convolutional neural networks that only have local receptive fields.

[0037] Specifically, to capture the global context relationship between different spatial positions in the image and retain spatial position information. Before calculating the attention weights, the self-attention layer performs two-dimensional sine position encoding on the two-dimensional feature map output by the backbone network. For each spatial position (x, y) in the feature map, a set of sine and cosine functions with fixed frequencies are used to encode it, constructing a position embedding vector with the same dimension as the number of feature channels, and fusing it with the original feature, so that each feature point can carry explicit spatial position information when participating in attention calculation.

[0038] The channel dimension of this attention layer can be set to 576, which is aligned with the encoded spatial features to ensure sufficient expressive power. The attention mechanism calculates the similarity between any two spatial positions and uses this as a weight to perform weighted aggregation on the features of all positions, finally outputting an image feature map with global perception ability.

[0039] In this way, the image features processed by the attention layer in combination with the two-dimensional sine position encoding technology not only retain the local semantic information extracted by the backbone network but also have the ability to model the relationships between remote regions, providing a more complete visual representation basis for subsequent dynamic 3D scene generation and rendering.

[0040] The scene generator constructs a three-plane structure representation aligned with the current camera view based on the multi-frame visual features extracted by the image encoder, and maps the multi-frame visual features to the three-plane structure representation after temporal encoding to generate a 3D scene feature representation.

[0041] In this embodiment, the scene generator specifically includes: The three-plane feature representation module is used to construct three mutually perpendicular feature planes parallel to the XY, YZ, and ZX planes respectively. The feature planes take the camera center as the coordinate origin and dynamically adjust the direction and position according to the camera parameters of the input video, and store the geometric consistency features of images from different perspectives as the three-plane structure representation aligned with the camera view.

[0042] As an example, the three-plane feature representation module is used to construct an ego-centric triplane structure representation. The triplane structure is composed of three mutually orthogonal two-dimensional feature planes, parallel to the XY, YZ, and ZX planes respectively. Each feature plane is used to store the image feature projection information from the corresponding direction. The triplane structure only stores information on three two-dimensional planes instead of storing features at each position in the entire three-dimensional space, replacing the traditional three-dimensional voxel grid or volumetric dense modeling, greatly reducing the memory occupancy and computational overhead, thus constituting a compressed but geometrically consistent three-dimensional expression method. That is, without sacrificing the accuracy of the spatial structure, a lighter way is used to represent the 3D scene.

[0043] To achieve dynamic alignment between the three-plane structure and the current frame view. First, the camera center corresponding to the current frame image is used as the reference origin in the three-dimensional space. This center can be obtained by parsing the extrinsic parameters of the camera in the input video frame (i.e., the pose of the camera in the world coordinate system). Subsequently, the system constructs a local coordinate system centered on the current camera based on the intrinsic and extrinsic parameters of this frame image, and defines three mutually orthogonal two-dimensional feature planes in this coordinate system, which are parallel to the XY, YZ, and ZX planes respectively, for storing the feature projection information of the scene in the three main viewing directions.

[0044] To adapt to the continuous change of the camera pose in the monocular video sequence, this three-plane structure is dynamically adjusted according to the camera parameters of each frame during processing, ensuring that its spatial orientation is aligned with the current viewing direction, so as to ensure the projection consistency of image features from different frames and different angles in the three-dimensional space.

[0045] It should be noted that although the three-plane structure representation does not use traditional three-dimensional data structures, through the combination of three two-dimensional planes, it can capture scene features in all directions and has spatial integrity and expressive ability.

[0046] The temporal feature modeling module is used to perform temporal encoding processing on multi-frame visual features, fuse the dynamic change information between frames, and then project the fused temporal features into the three-plane structure representation to construct a three-dimensional scene feature representation containing temporal dynamic information.

[0047] As an example, assuming the current target time is t, the system selects several frames of images before and after this time (for example ) as the input sequence. The image encoder extracts the corresponding visual feature vector F t+i for each frame image. Subsequently, the system assigns a corresponding time encoding vector E t+i to each frame. This time encoding can form a fixed position representation by embedding the time stamp through sine and cosine functions.

[0048] Concatenate the visual feature vector F t+i and the time encoding vector E t+i to form a feature representation with time information :

[0049] The system introduces a cross-attention mechanism to achieve cross-frame feature fusion. Specifically, for the current target time t, use as the query (Query), and the features of the remaining times as the key (Key) and value (Value). Through the calculation of attention weights, the system assigns weights to the features of each frame according to the time distance and visual similarity. For example, if t = 10, then I9 and I 11The weight of may be higher than that of I7 or I 13 , indicating that it is closer to the current target frame in terms of content and time.

[0050] According to the weights calculated above, the features of each neighboring frame are weighted and summed to obtain the finally fused temporal features; these temporal features are projected into the three-plane structure representation aligned with the current perspective and used as part of the final three-dimensional scene features.

[0051] Therefore, by introducing the self-attention mechanism based on spatial position encoding and the 4D scene generator with time modeling capabilities, the present invention can comprehensively capture the change features in different perspectives and time periods of dynamic videos, significantly enhancing the model's understanding and adaptation capabilities for complex dynamic scenes. In particular, the temporal feature modeling module in the scene generator integrates the cross-attention mechanism, time encoding, and visual similarity analysis, enabling the effective flow and collaborative modeling of information between frames, thereby strengthening the temporal continuity and dynamic consistency. This mechanism not only improves the modeling robustness of the model under complex change conditions such as fast motion and occlusion but also enables the dynamic information of each frame to be accurately mapped into the three-dimensional space, effectively supporting high-quality 4D dynamic scene reconstruction. This makes the system have stronger generality and generalization capabilities.

[0052] In this embodiment, the scene generator further includes: The axial feature enhancement module is used to aggregate the feature information on the same axis along the X, Y, or Z axis direction in each feature plane of the three-plane structure, and strengthen the information transmission along the axis through convolution operations to improve the spatial continuity of the three-dimensional scene feature representation.

[0053] Specifically, the three-plane structure consists of three two-dimensional feature planes parallel to the XY, YZ, and ZX planes respectively, and each feature plane can be regarded as a set of scene features projected in the corresponding axial direction. In actual modeling, if only local features are relied on, it may lead to insufficient transmission of structural information in the spatially continuous direction (such as the X, Y, or Z axis), thus affecting the coherence of the overall three-dimensional expression.

[0054] To solve the above problems, the axial feature enhancement module extracts the features on the same axis along the main axis direction (for example, along the X axis in the XY plane and along the Y axis in the YZ plane) within each feature plane and performs axial aggregation processing on these features through one-dimensional or two-dimensional convolution operations in deep learning.

[0055] For example, for the feature map in the XY plane, each row of features can be extracted with a fixed window in the horizontal direction (X axis), and weighted convolution can be applied for fusion to obtain new features with better contextual expression capabilities. Similarly, the YZ plane and ZX plane are processed in the Y and Z axis directions respectively to form a complete axial enhancement path. This axial aggregation processing method can be understood as establishing a "linear contextual connection" in the feature dimension, thereby achieving stronger global modeling capabilities in space.

[0056] The plane feature fusion module is used to establish cross-plane feature interaction relationships between the three feature planes and redistribute the information between the three feature planes through the cross-attention mechanism to enhance the spatial consistency and overall expression ability of the three-dimensional scene feature representation.

[0057] Although each plane can independently encode information in a specific direction, if there is a lack of an effective information sharing mechanism among the three, it may lead to inconsistencies in modeling results in different directions, affecting the final three-dimensional structure reconstruction effect.

[0058] To solve the above problems, the plane feature fusion module establishes a feature interaction mechanism between the three feature planes by constructing a cross-plane attention path. Specifically, first, the three feature planes are mapped to the embedding space of the same dimension through a linear transformation with a uniform number of channels to ensure comparability between them. Subsequently, the system uses one of the feature planes as a query item (Query) and the other two as keys (Key) and values ​​(Value) to input into the cross-attention calculation unit. By calculating the correlation scores between the planes, a weighted fusion representation is formed. Through cross-attention, the XY plane will query and fuse the contextual features of the same spatial position in the YZ and ZX planes according to its position, and then correct its own expression. The above process is performed alternately between the three planes, so that each plane can fuse the structural information provided by the other two planes. In this way, each plane is no longer modeled independently, but exchanges information and fuses perspectives with other planes through the attention mechanism to redistribute the information between the three feature planes, and finally forms a spatially coordinated, consistent, and complementary three-dimensional structural expression. In subsequent decoding or volume rendering, a geometrically continuous, structurally accurate, and unbroken three-dimensional scene representation can be generated.

[0059] Scene decoder,The scene decoder decodes the 3D scene feature representation and generates a synthetic image under the corresponding perspective based on the volume rendering mechanism. The synthetic image can show the geometric details of the scene under different perspectives and maintain the geometric consistency under different perspectives.

[0060] In this embodiment, the scene decoder includes: An upsampler for upsampling a low-resolution feature map of a three-dimensional scene constructed from a three-dimensional scene feature representation to generate a high-resolution tri-planar feature map.

[0061] As an example, the upsampler uses a transposed convolution technique to boost a low-resolution tri-planar feature map (e.g., 64×64) in the three-dimensional scene feature representation to a higher resolution (e.g., 256×256), enhancing feature details and spatial continuity and providing a more accurate input basis for subsequent high-quality rendering.

[0062] A three-dimensional feature encoding module for extracting corresponding joint features from a high-resolution tri-planar feature map for any position point in three-dimensional space and performing sine position encoding in combination with the spatial coordinates of the position point to construct an input vector.

[0063] As an example, after upsampling, the system can perform feature query and encoding processing for any spatial position point (x, y, z) in the high-resolution tri-planar feature map. Specifically, first, for the target position point, the corresponding projection positions are respectively found in three planes (XY, YZ, ZX), and the corresponding two-dimensional feature vectors are extracted from each plane; the features in the three directions are fused (such as concatenation or weighted averaging) to form a high-dimensional joint feature representation of this point; at the same time, the spatial coordinates (x, y, z) of this point are subjected to sine position encoding for embedding position information; finally, the feature vector and the position encoding are combined to construct an input vector for use by the subsequent decoder.

[0064] A neural radiance field decoder for receiving the input vector and calculating the color value and density value of the position point in three-dimensional space through a multi-layer perceptron.

[0065] As an example, a neural radiance field (NeRF) decoder for receiving the feature representation of any position point in three-dimensional space and outputting the color value and density value of this position point under the current view.

[0066] Specifically, the decoder structure is a two-layer multi-layer perceptron, whose input is the input vector constructed for the spatial point, and this vector consists of two parts: the spatial feature part and the position encoding part. The spatial feature part is the fused feature extracted from the high-resolution tri-planar structure, reflecting the context semantics of this position point. The position encoding part is a high-dimensional vector formed by encoding the three-dimensional coordinates (x, y, z) of this point through sine and cosine functions for embedding position correlation. After concatenating the above two parts, the input vector v is obtained and input into the multi-layer perceptron MLP for non-linear mapping:

[0067] where, is the output RGB color value; , representing the density value (or transparency), which is used for weighted calculation in volume rendering.

[0068] During the training process, the decoder automatically learns the mapping relationship from high-dimensional space features to colors and densities by optimizing the pixel loss between the rendered image and the real image, thereby realizing the realistic modeling of any spatial point in the scene.

[0069] The volume rendering module is used to calculate the color and density information for multiple spatial sampling points on each ray based on the ray tracing path of the target view, perform weighted integration on the spatial sampling points through the volume rendering formula to obtain the color output of the ray, and arrange and combine the color outputs of each ray according to the corresponding image pixel positions to generate a synthetic image corresponding to the view, and the synthetic image represents the rendering result of the current 3D scene from the target view.

[0070] In this embodiment, to complete the image rendering from the target view, the volume rendering module discretely samples each ray projected from the camera based on the ray tracing mechanism, and calculates the color and density information of each sampling point through the neural radiance field decoder, thereby completing the synthetic output of pixel values. Its processing flow includes the following steps: 1. Spatial sampling: For each ray (for example, a view direction projected from the camera to the scene), samples are taken at fixed intervals along its path to obtain a series of spatial point positions along the ray direction , and each point represents a three-dimensional coordinate.

[0071] 2. Feature extraction and position encoding: For each sampling point, extract the corresponding features of this point on the XY, YZ, and ZX planes from the high-resolution three-plane structure representation; after fusing the features in three directions, combine the spatial coordinates of this point, and embed its position information through the sine position encoding function to form the input vector v i of the neural radiance field decoder.

[0072] 3. Color and density prediction: Input each input vector v i into the neural radiance field decoder to output the RGB color value c i and density value σ i of the corresponding point.

[0073] 4. Volume rendering calculation: According to the color and density information of all sampling points, use the volume rendering formula to calculate the final color output of this ray:

[0074] Among them, a i is the opacity of the i-th sampling point, which is calculated by the calculation formula and is calculated as is the density value of the sampling point i, is the distance between adjacent sampling points, is the weight term of its forward visibility, which is calculated by the exponential decay method according to the density value.

[0075] 5. Image synthesis: By traversing all pixels in the image, the color output results of all rays are arranged and combined according to their pixel positions in the image, and finally the complete synthesized image under the current view is obtained.

[0076] The scene decoder combines upsampling technology and an efficient neural radiance field decoder, enabling the system to have good real-time performance in the inference stage, capable of achieving efficient real-time rendering capabilities on resource-constrained terminal devices, and capable of quickly generating high-quality, multi-view consistent 3D synthesized images.

[0077] Compared with the methods in the prior art that usually rely on a large amount of manually labeled data, mask information or need to tune parameters for each specific scene, the dynamic 3D scene reconstruction and rendering system proposed in this application has the significant advantage of not requiring scene-specific optimization. Without the need for additional prior information or training adjustment for a specific scene, it can directly model any dynamic video, greatly simplifying the model training process and deployment process, and improving the practicality and generality of the system.

[0078] Referring to Figure 2 shown, an embodiment of this application provides a reconstruction and rendering method applicable to any one of the above-mentioned systems, and this method includes: Step S1: Receive a monocular video sequence and perform visual feature extraction on each frame of the image; Step S2: Based on the extracted multi-frame visual features, construct a three-plane structure representation aligned with the current camera view, and map the multi-frame visual features to the three-plane structure representation after time encoding to generate a 3D scene feature representation; Step S3: Decode the 3D scene feature representation and generate a synthesized image corresponding to the view based on the volume rendering mechanism. The synthesized image can display the geometric details of the scene from different views and maintain geometric consistency.

[0079] Specifically, the reconstruction and rendering process of the dynamic 3D scene is completed by the cooperation of multiple modules, specifically including an image encoder, a scene generator, and a scene decoder. The specific execution process is as follows: First, the input monocular video sequence is received by the image encoder, and visual feature extraction operations are performed on each frame image in the monocular video sequence to obtain visual feature representations of multiple frames. Inside the image encoder, there are a backbone convolutional network and a self-attention layer, which are used to model the local and global context information of the image respectively.

[0080] Subsequently, the extracted visual features of multiple frames are input into the scene generator. The tri-plane feature representation module in the scene generator is responsible for constructing a tri-plane structure representation aligned with the current camera view. The tri-plane consists of three feature planes parallel to the XY, YZ, and ZX planes respectively, and with the spatial position of the camera of the current frame as the reference origin, the position and orientation of the tri-plane are dynamically adjusted through pose parameters.

[0081] Next, the visual features of multiple frames are further input into the temporal feature modeling module. This module combines the timestamp information of video frames through time encoding and introduces a cross-attention mechanism to fuse the dynamic change information between frames, thereby realizing the dynamic modeling of continuous changes in the time dimension. The fused temporal features are finally mapped into the tri-plane structure to form a complete three-dimensional scene feature representation.

[0082] After generating the three-dimensional feature representation, it enters the scene decoder stage. Among them, the upsampler first performs deconvolution processing on the low-resolution tri-plane feature map to generate a high-resolution tri-plane feature map to enhance the spatial detail expression ability.

[0083] Subsequently, the three-dimensional feature encoding module extracts the joint features corresponding to any sampling position point in the three-dimensional space from the above high-resolution tri-planes, and performs sine position encoding in combination with the spatial coordinates of this point to construct a high-dimensional input vector.

[0084] This input vector is fed into the neural radiance field decoder, which is a neural network composed of two layers of multi-layer perceptrons, and is used to predict the RGB color value and density value of this spatial position point.

[0085] After completing the color and density prediction of the sampling points, the volume rendering module performs multi-point spatial sampling along the light direction corresponding to each pixel based on the current target view, and calculates the color and density information of each sampling point. Subsequently, the volume rendering formula is used to perform weighted sum and integration on the sampling points on the light ray to calculate the final color output of the light ray.

[0086] Finally, the color output results of each light ray are arranged and stitched according to the corresponding image pixel positions to generate a synthesized image under the current view. This image can not only present the geometric details of the three-dimensional scene but also maintain the structural consistency between multiple views.

[0087] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0088] In summary, a dynamic three-dimensional scene reconstruction and rendering system and method provided by the present application receives a monocular video sequence through an image encoder, and extracts visual features from each frame of image; a scene generator constructs a three-plane structure representation aligned with the current camera view based on the multi-frame visual features extracted by the image encoder, and maps the multi-frame visual features to the three-plane structure representation after temporal encoding to generate a three-dimensional scene feature representation; a scene decoder decodes the three-dimensional scene feature representation and generates a synthesized image corresponding to the view angle based on a volume rendering mechanism. The synthesized image can display the geometric details of the scene from different view angles and maintain geometric consistency from different view angles. The present application realizes the modeling and rendering of a dynamic three-dimensional scene through a monocular video sequence, without relying on multi-view input or prior scene information, effectively improving the adaptability and generalization performance of the system in a rapidly changing environment. Using a three-plane structure to represent the three-dimensional space and combining a temporal encoding mechanism can accurately capture the spatio-temporal dynamic features in the scene and maintain the consistency of geometric information. By using a volume rendering mechanism to synthesize images at any view angle, clear and structurally consistent scene details can be presented at different viewing angles, significantly improving the rendering quality and modeling efficiency, and being applicable to dynamic vision task scenarios such as augmented reality, virtual reality, and autonomous driving.

[0089] In addition, some embodiments of the present application also provide an electronic device. The electronic device can be various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and so on. The electronic device can also be various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices.

[0090] The electronic device includes: one or more processors; and a memory storing computer program instructions, which when executed cause the processors to execute a dynamic three-dimensional scene reconstruction and rendering method provided by any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. As Figure 3As shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting the various components, including a high-speed interface and a low-speed interface. The various components are interconnected using different buses and can be mounted on a common motherboard or otherwise mounted as required. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory for graphical information to be displayed on an external input / output device (such as a display device coupled to the interface) to display a GUI. In some other embodiments, multiple processors and / or multiple buses can be used in conjunction with multiple memories and multiple memories if needed. Similarly, multiple electronic devices can be connected, with each device providing part of the necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Among them, the components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0091] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected by a bus or other means. Figure 3 Taking connection by bus as an example.

[0092] The input device 1103 can receive input digital or character information and generate key signal inputs related to the user settings and function controls of the electronic device, such as input devices like a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 1104 may include a display device, an auxiliary lighting device (such as an LED), and a haptic feedback device (such as a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0093] To provide interaction with a user, the electronic device may be a computer. The computer has: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or an LCD monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0094] In an embodiment of the present application, a computer program / instructions is stored on a computer-readable medium. When the computer program / instructions are executed by a processor, a method for dynamic three-dimensional scene reconstruction and rendering provided by any one or more of the above embodiments is implemented. The computer-readable medium may be included in the electronic device described in the above embodiment; or it may exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.

[0095] The memory 1102 may be used as a non-transitory computer-readable storage medium for storing non-transitory software programs, non-transitory computer-executable programs, and modules. By running the non-transitory software programs, instructions, and modules stored in the memory 1102, the processor 1101 executes various functional applications and data processing of the server to implement the program instructions / modules corresponding to the method provided by any one or more of the above embodiments in the embodiments of the present application.

[0096] The memory 1102 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely set relative to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0097] It should be noted that more specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0098] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical discs (CD-ROMs), digital versatile discs (DVDs) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0099] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0100] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of this application can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.

[0101] The computer program product provided by the embodiments of the present application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they wholly or partly generate the processes or functions described in the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive, SSD).

[0102] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0103] As described above, the foregoing are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily make changes or substitutions within the technical scope disclosed by the present application, and all such changes or substitutions should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A dynamic three-dimensional scene reconstruction and rendering system, characterized in that Comprising: An image encoder that receives a monocular video sequence and extracts visual features from each frame image; A scene generator that, based on the multi-frame visual features extracted by the image encoder, constructs a three-plane structure representation aligned with the current camera view, and maps the multi-frame visual features to the three-plane structure representation after temporal encoding to generate a three-dimensional scene feature representation; A scene decoder that decodes the three-dimensional scene feature representation and generates a synthesized image corresponding to the view based on a volume rendering mechanism, and the synthesized image can display the geometric details of the scene from different views and maintain geometric consistency across different views.

2. The dynamic three-dimensional scene reconstruction and rendering system according to claim 1, characterized in that The image encoder includes: A backbone network for performing multi-layer convolutional operations on each frame image in the monocular video sequence to extract a deep visual feature map; A self-attention layer arranged after the backbone network for globally modeling the deep visual feature map to capture long-range dependencies.

3. The dynamic three-dimensional scene reconstruction and rendering system according to claim 2, wherein After receiving the deep visual feature map, the self-attention layer performs position encoding on the deep visual feature map using a two-dimensional sine position encoding mechanism to embed spatial position information, so that each feature point carries spatial position information when participating in the attention calculation, to further enhance the model's ability to model long-range dependencies.

4. The dynamic three-dimensional scene reconstruction and rendering system according to claim 1, wherein The scene generator includes: A three-plane feature representation module for constructing three mutually perpendicular feature planes parallel to the XY, YZ, and ZX planes respectively, with the camera center as the coordinate origin, dynamically adjusting the direction and position according to the camera parameters of the input video, and storing the geometric consistency features of images from different views as the three-plane structure representation aligned with the camera view; A temporal feature modeling module for performing temporal encoding on the multi-frame visual features, fusing the inter-frame dynamic change information, and then projecting the fused temporal features to the three-plane structure representation to construct the three-dimensional scene feature representation containing temporal dynamic information.

5. The dynamic three-dimensional scene reconstruction and rendering system according to claim 4, characterized in that, The scene generator further includes: An axial feature enhancement module for aggregating the feature information on the same axis along the X, Y, or Z axis direction in each feature plane of the three-plane structure, and strengthening the information transmission along the axis through convolutional operations to improve the spatial continuity of the three-dimensional scene feature representation; A plane feature fusion module for establishing cross-plane feature interaction relationships between the three feature planes, and redistributing the information between the three feature planes through a cross-attention mechanism to enhance the spatial consistency and overall expression ability of the three-dimensional scene feature representation.

6. The dynamic three-dimensional scene reconstruction and rendering system according to claim 1, characterized in that The scene decoder includes: An upsampler for upsampling the low-resolution feature map of the three-dimensional scene constructed by the three-dimensional scene feature representation to generate a high-resolution three-plane feature map; A three-dimensional feature encoding module for extracting the corresponding joint features from the high-resolution three-plane feature map for any position point in the three-dimensional space, and performing sine position encoding in combination with the spatial coordinates of the position point to construct an input vector; A neural radiance field decoder for receiving the input vector and calculating the color value and density value of the position point in three-dimensional space through a multi-layer perceptron; A volume rendering module for calculating color and density information for multiple spatial sampling points on each ray based on the ray tracing path of the target view, performing weighted integration on the spatial sampling points through a volume rendering formula to obtain the color output of the ray, and arranging and combining the color outputs of each ray according to the corresponding image pixel positions to generate a synthesized image for the corresponding view, where the synthesized image represents the rendering result of the current three-dimensional scene from the target view.

7. A reconstruction and rendering method applicable to the system according to any one of claims 1-6, characterized in that, The method includes: Receiving a monocular video sequence and performing visual feature extraction on each frame image; Based on the extracted multi-frame visual features, constructing a three-plane structure representation aligned with the current camera view, and mapping the multi-frame visual features to the three-plane structure representation after temporal encoding to generate a three-dimensional scene feature representation; Decoding the three-dimensional scene feature representation and generating a synthesized image for the corresponding view based on a volume rendering mechanism, where the synthesized image can display the geometric details of the scene from different views and maintain geometric consistency.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; and a memory storing computer program instructions that, when executed, cause the processors to execute the dynamic three-dimensional scene reconstruction and rendering method according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program and / or instructions stored thereon, characterized in that, The computer program and / or instructions, when executed by the processor, implement the dynamic three-dimensional scene reconstruction and rendering method according to any one of claims 1-7.

10. A computer program product, comprising a computer program and / or instructions, characterized in that, The computer program and / or instructions, when executed by the processor, implement the dynamic three-dimensional scene reconstruction and rendering method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Three-dimensional scene reconstruction method based on neural radiation field

    CN116342788A

  • Neural radiation field NeRF-based three-dimensional complex scene refined reconstruction method and device

    CN118429526A

  • Neural rendering method and system based on global semantic information and local geometric perception

    CN118587340A

Cited By

  • Aluminum alloy pipe target forging forming method and system based on image processing

    CN120655867A

  • An image processing-based method and system for forging forming of an aluminum alloy tube target

    CN120655867B

  • Vehicle-mounted three-dimensional scene rendering method and device, storage medium and program product

    CN121661216A

  • A vehicle-mounted three-dimensional scene rendering method, device, storage medium and program product

    CN121661216B