Multi-view driven three-dimensional generation method

By constructing a multi-view driven three-dimensional generation model, using the multi-view Transformer network and correction flow model to generate high-quality three-dimensional models from multi-view images, the problems of high data acquisition cost and difficulty in controlling three-dimensional structure features in the existing technology are solved, and efficient and accurate three-dimensional reconstruction is achieved.

CN120259547APending Publication Date: 2025-07-04HUAZHONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510391887.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing three-dimensional generation technology is expensive to obtain data under complex lighting, occlusion and extreme weather conditions and has limited coverage dimensions. It is difficult for traditional methods to accurately control the three-dimensional structural characteristics of the target object, resulting in insufficient anti-interference ability and generalization performance of the model in real scenes.

Method used

A multi-view driver three-dimensional generation model is built, and an encoder, a multi-view Transformer network and a decoder are used to generate a three-dimensional Gaussian point cloud from a multi-view image through the correction flow model and a multi-view Transformer network. The relationship between images is captured using the multi-head self-attention and spatial attention modules. The decoder converts the implicit representation into an explicit structure.

Benefits of technology

It improves the accuracy and detailed performance of three-dimensional reconstruction, enhances the adaptability and robustness of the model, reduces the consumption of computing resources, and has a wide range of applications. It is suitable for three-dimensional modeling and data simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259547A_ABST
    Figure CN120259547A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view driven three-dimensional generation method, and belongs to the field of computer vision and three-dimensional reconstruction. According to the method, a multi-view image of an object is coded to generate corresponding feature representation, and support is provided for subsequent correction flow guide three-dimensional model optimization; a correction flow model is adopted to gradually convert a three-dimensional sample under prior distribution into implicit three-dimensional representation of an object under the condition of multi-view images of the object, a multi-view Transform network is adopted as a speed prediction model of the correction flow model, and the correction flow model starts from a Gaussian distribution sample and is guided according to a current state and a speed field, so that the speed of the object is predicted. The method comprises the following steps: continuously correcting a sample path, enabling a sample to be stably close to target distribution along a predetermined track, achieving efficient remodeling of an implicit expression of an object, and finally decoding the implicit expression into a three-dimensional Gaussian point cloud through a decoder, thereby realizing three-dimensional reconstruction of the object. The method has the advantages of high speed and high reconstruction quality, and is suitable for the fields of three-dimensional modeling, data simulation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and 3D reconstruction, and more specifically, relates to a multi-view driven 3D generation method. Background Art

[0002] In recent years, intelligent detection and recognition technologies based on deep learning have been widely applied in fields such as optical imaging reconnaissance, border control monitoring, and autonomous driving. However, existing methods face two major bottlenecks in actual deployment: firstly, the performance of the object detection model highly depends on large-scale and multi-scene training data. Especially under complex lighting, occlusion, and extreme weather conditions, the cost of obtaining real data is high and the coverage dimension is limited; secondly, traditional data augmentation means are difficult to accurately control the three-dimensional structural features of the target object, resulting in insufficient anti-interference ability and generalization performance of the model in real scenarios.

[0003] To address the problem of data scarcity, 3D reconstruction technology based on physical simulation provides a potential solution: by constructing a 3D model and generating multi-view 2D images based on a rendering pipeline, training data containing complex environmental conditions can be batch synthesized, reducing the data acquisition cost. However, traditional 3D reconstruction technology relies on precise multi-view acquisition devices and complex point cloud registration processes, and is prone to geometric structure distortion problems in real scenarios such as sparse viewpoints or occlusion scenarios, resulting in feature mismatches in detection models trained based on such data.

[0004] To break through the above limitations, 3D generation technology has become a research hotspot in recent years. 3D generation refers to the process of generating a 3D model from 2D images or other input data through computer technology. It can solve the problem of how to infer a complete 3D structure from limited visual information, thereby reconstructing a high-quality 3D model in the absence of direct 3D data.

[0005] Patent CN113112616A discloses an editable 3D hairstyle generation method, which adopts an improved Hair-GAN framework and can convert 2D information into a 3D directional voxel field, improving the generation efficiency. However, it mainly targets a specific scenario of hairstyles and is difficult to be directly applied to other types of 3D generation tasks. Patent CN118469827A relates to a geometric perception-based 3D synthesis method at all angles, which realizes image synthesis through a 3D generative adversarial network. However, in practical applications, this method has a strong dependence on the camera pose, and the camera pose information may be difficult to obtain in some scenarios. Summary of the Invention

[0006] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a multi-view driven 3D generation method, which realizes the efficient and accurate reconstruction of a high-quality 3D model from multiple images with different perspectives, and has a wide range of applications and does not depend on the camera pose.

[0007] To achieve the above object, according to the first aspect of the present invention, there is provided a method for constructing a multi-view driven 3D generation model, including:

[0008] Construct a multi-view driven 3D generation model and train it using a data set;

[0009] Wherein, the multi-view driven 3D generation model includes an encoder, a multi-view Transformer network and a decoder connected in sequence; the multi-view Transformer network includes several network layers and output heads, and each network layer includes a multi-head self-attention module, a multi-head spatial attention module and a feed-forward neural network connected in sequence; the data set is multi-view images of an object, the labeled images are part of the data set, and the sample images are other parts of the data set;

[0010] The training process includes:

[0011] S1, the encoder obtains the feature representation of the multi-view sample images The multi-head self-attention module performs self-attention calculation on the 3D sample Z at the current time t t to obtain the feature matrix O self , and the multi-head spatial attention module performs spatial attention calculation on and O self to obtain the fused feature matrix O space , and the feed-forward neural network performs a non-linear transformation on O space to predict the velocity feature O of the correction flow model ffn , and O ff is linearly transformed through the output head to be converted into the transformation velocity S of Z t ; wherein, the initial value of t is 0, and Z0 is a 3D sample randomly sampled from the prior distribution; t ;

[0012] S2, update t to t + Δt, and update Z t to Z t + S t Δt. If the current time reaches the preset time, output the updated Z t and enter S3, otherwise return to S1;

[0013] S3, the decoder converts the Z t output by S2 into a 3D Gaussian point cloud;

[0014] S4. Convert the three-dimensional Gaussian point cloud into a new perspective image, and update the learnable parameters of the multi-view Transformer network and the decoder with the goal of minimizing the loss between the new perspective image and the label image.

[0015] According to the second aspect of the present invention, there is provided a multi-view driven three-dimensional generation method, including:

[0016] Input the multi-view images of the target object into the multi-view driven three-dimensional model constructed by the method described in the first aspect to obtain the three-dimensional Gaussian point cloud of the target object.

[0017] According to the third aspect of the present invention, there is provided an electronic device, including: a computer-readable storage medium and a processor;

[0018] The computer-readable storage medium is used to store executable instructions;

[0019] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the method described in the first aspect or the second aspect.

[0020] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to execute the method described in the first aspect or the second aspect.

[0021] According to the fifth aspect of the present invention, there is provided a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by the processor, the method described in the first aspect or the second aspect is implemented.

[0022] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0023] 1. The multi-view driven 3D generation model provided by the present invention uses a correction flow model to improve the calculation speed. The correction flow model assumes that the connection path between samples is a straight-line path and learns an ordinary differential equation (ODE) model to achieve fast generation. The straight-line path can significantly reduce the number of ODE solution iterations and improve the operating efficiency of the model. The multi-view Transformer network has a powerful fitting ability, can effectively process complex multi-view information, capture complex relationships and features between images, thereby improving the accuracy and detail performance of 3D reconstruction, making the reconstruction result more realistic and accurate. The multi-view Transformer network uses a spatial attention mechanism to enable any number of image inputs, enabling the model to flexibly process different amounts of multi-view information, enhancing the adaptability and robustness of the model, and providing a wider applicable range for 3D reconstruction. Using implicit representation is more efficient, has strong compression ability, reduces the consumption of computing resources, and is convenient for subsequent expansion into various display 3D structures, providing multiple possibilities for downstream tasks.

[0024] 2. As a further preferred solution, the multi-head self-attention module provided by the present invention further includes a time encoding module for embedding time information into the multi-view Transformer network for calculation, which can improve the accuracy of the model's speed prediction without significantly increasing the amount of calculation.

[0025] 3. As a further preferred solution, the method provided by the present invention dynamically adjusts the sampling time interval Δt, which can improve the generation quality and speed up the generation speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic structural diagram of the multi-view driven 3D generation model provided by the embodiment of the present invention;

[0027] Figure 2 It is a schematic flow diagram of the multi-view driven 3D generation method provided by the embodiment of the present invention;

[0028] Figure 3 It is an example diagram of the multi-view Transformer network provided by the embodiment of the present invention;

[0029] Figure 4 It is a schematic diagram of the correction flow model inference provided by the embodiment of the present invention;

[0030] Figure 5 It is a schematic training flow diagram of the multi-view driven 3D generation model provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0032] Existing three-dimensional generation methods have some deficiencies. For example, some methods only rely on image information from a single perspective and are difficult to accurately capture the three-dimensional structure of an object; there are also some methods that may have problems such as loss of details or distortion of the structure in the generated three-dimensional models when dealing with complex scenes. In addition, some methods have low computational efficiency and are difficult to meet the requirements of real-time applications.

[0033] Based on this, an embodiment of the present invention provides a method for constructing a multi-view driven three-dimensional generation model, including:

[0034] Constructing a multi-view driven three-dimensional generation model and training the multi-view driven three-dimensional generation model using a data set;

[0035] Wherein, the multi-view driven three-dimensional generation model includes an encoder, a multi-view Transformer network, and a decoder connected in sequence; the multi-view Transformer network includes several network layers and output heads, and each network layer includes a multi-head self-attention module, a multi-head spatial attention module, and a feed-forward neural network connected in sequence; the data set is multi-view images of an object, the labeled images are part of the data set, and the sample images are other parts of the data set.

[0036] Preferably, the multi-head self-attention module further includes a temporal encoding module for encoding the current time t as t t before the multi-head self-attention module performs self-attention calculation on Z

[0037] and adding t e to Z e t

[0038] Preferably, the multi-view driven three-dimensional generation model further includes a preprocessing module for performing image normalization processing on the input image.

[0039] Specifically, the multi-view driven three-dimensional generation model includes:

[0040] A preprocessing module for preprocessing the input images of an object from multiple different perspectives, including image normalization operations, to improve the accuracy of subsequent encoding.

[0041] ​​An image encoding module for encoding the preprocessed image to generate a feature representation related to the image content.

[0042] The image encoding module can adopt any existing image encoding module, for example, DINO image encoding module, DINOV2 image encoding module, CLIP image encoding module, etc.

[0043] The multi-view driven 3D generation model provided by the present invention uses a correction flow model to gradually transform 3D samples (i.e., the initial 3D samples of the object: randomly initialized 3D voxels) under the prior distribution (i.e., the noise distribution) into an implicit 3D representation of the object under the condition guidance of images from multiple perspectives of the object, and uses a multi-view Transformer network as the speed prediction model of the correction flow model.

[0044] By randomly sampling initial sample points from a Gaussian distribution as the starting state of the 3D generation process, and based on the ordinary differential equation (ODE) constructed by the correction flow model, the initial sample points are gradually iteratively updated. In each iteration step, according to the state of the current sample point and the input multi-view reference images, the transformation speed of the sample point is predicted. The speed is determined by the feature representation generated by the multi-view Transformer network, guiding the sample point to approach the target distribution along the optimal path. After multiple rounds of iterative updates, the sample points gradually change from the Gaussian distribution state to a state close to the target distribution, and the sample points are decoded into 3D Gaussian point clouds through a decoder.

[0045] The multi-view Transformer network consists of several network layers and output heads, and each network layer includes a multi-head self-attention module, a multi-head spatial attention module, and a feed-forward neural network;

[0046] Among them, the multi-head self-attention module is used to perform interactive calculations on the 3D samples themselves to capture the global dependencies between different positions in the 3D samples, and the multi-head spatial attention module is used to perform interactive calculations between the 3D samples and different images to achieve effective fusion and correction of multi-view information, obtaining a fused feature matrix; the feed-forward neural network is used to predict the speed features of the correction flow model according to the fused feature matrix, and the speed features are converted into the transformation speed S of Z t by the output head t ;

[0047] A decoder for decoding the implicit 3D representation generated by the correction flow model into an explicit 3D structure, thereby realizing the 3D reconstruction of the object.

[0048] The decoder is implemented using common networks. For example, it can be implemented based on a Transformer network or based on a 3D convolutional network.

[0049] Training is carried out using the open-source 3D dataset Objaverse. First, a large number of multi-view images are rendered for each 3D object. During training, a part of these images is selected as conditional input, and another part is selected as supervision. The multi-view driven 3D generation model first generates a 3D object based on the conditional images, then uses differentiable rendering to generate images from different viewpoints, and finally uses these images for supervision to obtain a trained multi-view driven 3D generation model.

[0050] Specifically, after constructing the multi-view driven 3D generation model, a dataset needs to be constructed to train the model. The sample images in the dataset are multiple 3D object images from different viewpoints, and these images are generated by rendering using Blender software. Blender uses the Cycles rendering engine to ensure high-quality and realistic rendering effects. During the rendering process, a uniform lighting setting is used to avoid image quality differences caused by uneven lighting. The cameras are distributed on a sphere with a fixed radius, and the sphere radius is set according to the size of the object and the required viewing angle range to ensure that all details of the object can be comprehensively captured. The field of view (FOV) of the cameras ranges from 45° to 60° to simulate different viewing distances and viewing angle effects. The resolution of the rendered output images is 512×512 pixels to balance the requirements of image quality and computing resources.

[0051] In the process of constructing a high-quality training dataset, first, an image quality assessment model is used to quantitatively score each rendered image to obtain its quality index. Through statistical analysis of the quality scores of all rendered images, the average score is calculated. Based on this average score, combined with the data distribution characteristics and the model training requirements, a threshold is set according to experience. This threshold is usually slightly higher than the average score to screen out image data with better quality. This strategy can significantly reduce the interference of low-quality images on model training, enabling the model to extract key features and patterns from high-quality data, thereby improving the training effect and generation quality of the model.

[0052] The process of the training includes:

[0053] S1, the encoder obtains the feature representation of the multi-view sample images The multi-head self-attention module performs self-attention calculation on the 3D sample Z at the current time t t to obtain the feature matrix O of Z t , and the multi-head spatial attention module performs spatial attention calculation on self and O to fuse the features of the 3D sample with the features of the sample images, obtaining the fused feature matrix O self , and the feed-forward neural network performs a non-linear transformation on O space to predict the velocity feature O of the correction flow model space ​ffn , O ffn is linearly transformed through the output header to Z t with a transformation speed S t ;

[0054] S2, update t to t + Δt, and update Z t to Z t + S t Δt. If the current time reaches the preset time, output the updated Z t and enter S3; otherwise, return to S1;

[0055] S3, the decoder converts the Z output by S2 t into a three-dimensional Gaussian point cloud;

[0056] S4, convert the three-dimensional Gaussian point cloud into a new perspective image, and update the learnable parameters of the multi-view Transformer network and the decoder with the goal of minimizing the loss between the new perspective image and the label image.

[0057] The following further introduces each module in combination with the training process of the multi-view driven 3D generation model. The training process of the multi-view driven 3D generation model includes the following steps:

[0058] (1) Before the multi-view image data is input into the multi-view driven 3D generation model, it needs to be preprocessed to standardize the data and adapt to the input requirements of the network model. The preprocessing operations include image normalization to adjust the image pixels to a fixed range for convenient subsequent calculations of the network.

[0059] (2) Taking the image encoder of the CLIP (Contrastive Language-Image Pretraining) model in the image encoding module as an example, convert the preprocessed image into a high-dimensional vector representation.

[0060] Usually, a Vision Transformer (ViT) is used for feature extraction. Taking a single image as an example, I is divided into N x patches, i = 1, 2,..., N x , where H and W are the height and width of the image, P is the size of each patch, is the number of patches, and C is the number of channels of the image.

[0061] Each is linearly projected into an embedding vector where is the embedding vector of the i-th patch, is the linear projection matrix, and D xis the dimension of the embedding vector.

[0062] Add the positional encoding to the embedding vector:

[0063]

[0064] Process the initial embedding vector X0 through the L-layer Transformer encoder:

[0065] X l = TransformerLayer(X l-1 ), l = 1, 2, …, L

[0066] where is the output of the l-th layer of the Transformer, and TransformerLayer represents a single-layer Transformer encoder.

[0067] The output of the last layer of the Transformer serves as the high-dimensional feature of the image.

[0068] When inputting images from F viewpoints , perform the same above operations on each image to obtain a multi-view image feature set

[0069] Use Ψ to describe the image encoding process:

[0070]

[0071] (3) To reduce the computational amount, adopt the implicit representation of the three-dimensional model, and use the rectified flow model to transform the three-dimensional sample Z0 randomly sampled from the prior noise distribution π0 to the true distribution π1 of the implicit representation of the target three-dimensional model under the guidance of the multi-view image feature to obtain the implicit representation Z1 of the target three-dimensional model.

[0072] Usually, π0 is defined as a simple prior noise distribution. Generally, the standard normal distribution (Gaussian noise) is adopted, and a complex distribution such as the gamma distribution can also be used; while π1 is the true distribution of the implicit representation of the target three-dimensional model, and this distribution is unknown.

[0073] The rectified flow model assumes that there is a linear transformation path between the sample Z0 of the prior noise distribution and the sample Z1 of the true distribution of the implicit representation of the target three-dimensional model:

[0074] Z t = (1 - t)Z0 + tZ1

[0075] where t ∈ [0, 1] is the time parameter, and Z tis a three-dimensional sample at time t.

[0076] The above linear transformation can be described by an ordinary differential equation (ODE):

[0077]

[0078] The above equation describes the velocity of Z t changing with time t. However, in practical applications, Z1 is unknown, resulting in an unpredictable velocity.

[0079] Model the velocity of Z t changing with time t through a learnable neural network:

[0080]

[0081] where is the velocity field neural network, which describes the law of Z t changing with time t given the current sample Z , the current time t, and the multi-view image features t . θ is the learnable parameter of the neural network.

[0082] Therefore, randomly sample a three-dimensional sample Z0 from the prior noise distribution π0, and obtain the implicit representation Z1 of the target three-dimensional model by solving the ordinary differential equation in the time period from t = 0 to 1.

[0083] To reduce the computational complexity, it is preferable to use a numerical integration method to solve such as the Euler method to obtain the update formula of Z t :

[0084]

[0085] Z t+Δt = Z t + S t Δt

[0086] Starting from t = 0, iterate step by step to t = 1 to obtain Z1.

[0087] Considering the computational accuracy, the Runge-Kutta method can also be used to solve the ODE, and the update formula of Z t is:

[0088]

[0089] Z t+Δt = Z t + S t Δt

[0090] Preferably, in order to improve the generation quality and speed up the generation speed, the sampling time interval Δt is dynamically adjusted.

[0091] A larger Δt is used in areas where the velocity field changes slowly, while a smaller Δt is used in areas where the velocity field changes drastically, as defined below:

[0092]

[0093] where α is a regulating factor, is the gradient of the velocity field, and θ represents all learnable parameters of the multi-view Transformer network, including:

[0094] (4) Using the multi-view Transformer network as the velocity field neural network for the correction flow model The multi-view Transformer network is an improvement on the traditional Transformer network, which cannot effectively process any amount of image information. It consists of multiple layers of network layers and output heads with the same structure. Each network layer includes: a temporal encoding module, a multi-head self-attention module, a multi-head spatial attention module and a feedforward neural network module.

[0095] 1) Time coding module

[0096] As a preferred solution, the time encoding module is used to embed time information into the network for calculation, which can improve the accuracy of the model's speed prediction without significantly increasing the amount of calculation. Its input includes: batch 3D samples and time information t∈[0,1], where B is the batch size of 3D samples, N z is the sequence length of the three-dimensional sample, D z is the characteristic dimension of the three-dimensional sample.

[0097] The time factor t is usually converted into a time embedding vector by encoding

[0098] Time embedding vector t e t can be generated by sine and cosine functions e =[t0,t1,…,t 2i ,t 2i+1 ,…]:

[0099]

[0100] Where i is the dimension index, t 2i and t 2i+1 are the even and odd index values ​​of the temporal embedding vector respectively. Temporal embeddings can also be generated by a learnable embedding layer:

[0101] t e = E t [round(t·T)]

[0102] where is a learnable embedding matrix, [·] represents the indexing operation, round represents the rounding operation, and N e is the length of the embedding matrix.

[0103] The embedding vector t e is added to the input data Z t as the Z after time information embedding t :

[0104] Z t = Z t + t e

[0105] 2) Multi-Head Self-Attention Module

[0106] The multi-head self-attention module is used to perform interaction calculations on the three-dimensional sample itself to capture the global dependencies between different positions in the three-dimensional sample. Its input is the three-dimensional sample Z after time embedding t .

[0107] The input Z t is linearly transformed to obtain the query (Query), key (Key), and value (Value) variables:

[0108]

[0109] where is the learnable weight matrix, are the query, key, and value variables, and D is the feature dimension after linear transformation.

[0110] The multi-head mechanism splits Q self , K self , V self into H parts to reduce the computational complexity and improve the computational efficiency. That is where h = 1, 2,..., H, and D / H is the feature dimension of each part.

[0111] Calculate the attention weights of each part respectively:

[0112]

[0113] where T represents the matrix transpose, is the scaling factor to avoid the Softmax gradient vanishing due to the excessive dot product.

[0114] Calculate the output variables of each part respectively:

[0115]

[0116] Concatenate the output variables of all parts and perform a linear transformation to obtain the output:

[0117]

[0118] where is the output of the self-attention model, is the learnable weight matrix.

[0119] 3) Multi-head spatial attention module

[0120] The multi-head spatial attention module is used to perform interactive calculations between 3D samples and multi-view information, achieving effective fusion and correction of multi-view information. Its inputs include: multi-view image features and the output O self of the multi-head self-attention module, where B is the batch size, F is the number of views, N x is the length of each image feature sequence, and D x is the feature dimension of the image.

[0121] Perform a linear transformation on the input O self to obtain the query variable, and perform a linear transformation on to obtain the key and value variables:

[0122]

[0123] where is the learnable weight matrix, are the query, key, and value variables.

[0124] The multi-head mechanism splits Q space , K space , V space into H parts to reduce the computational load and improve computational efficiency. That is where h = 1, 2,..., H, and D / H is the feature dimension of each part.

[0125] Calculate the attention weights of each part respectively:

[0126]

[0127] Calculate the output variables of each part respectively:

[0128]

[0129] Concatenate the output variables of all parts and perform a linear transformation to obtain the output corresponding to each image under the multi-view image:

[0130]

[0131] where is the output corresponding to each image and needs to be adaptively fused, is a learnable weight matrix.

[0132] During the sample conversion process of the correction flow model, the three-dimensional samples gradually migrate from the π0 noise distribution to the target distribution π1, and the spatial attention module focuses on different regions of the view at different stages. In the early stage (t≈0), noise dominates, and the model mainly uses global features at this time to avoid the interference of noise on detailed information and ensure the stability of the overall structure of the generated data. Therefore, the variance of the attention weights is small. As the generation process progresses, the global structure of the three-dimensional data gradually recovers, but the local details are still imperfect. At this time, the model turns to using local features to recover high-frequency details, and the attention is concentrated in the local area. Therefore, the variance of the attention weights is large.

[0133] Based on the above temporal evolution characteristics, by estimating the difference between the variance and the expectation of the attention of the three-dimensional samples to each view, the credibility of each view is evaluated, and then the multi-view information fusion is adaptively adjusted.

[0134] Construct the overall attention weight matrix from the attention weights of all heads:

[0135]

[0136] Define the variance of the attention distribution of the three-dimensional samples to each view as:

[0137]

[0138] Construct a credibility index based on the above temporal evolution characteristics

[0139]

[0140] where is the theoretical maximum value of the variance.

[0141] Use a normalized exponential function with an adjustable temperature coefficient to convert the credibility index of each view into a normalized weight

[0142]

[0143] where τ is the temperature coefficient, which is used to control the sharpness or balance degree of the weight distribution.

[0144] Calculate the average of all heads of the normalized weights to obtain the average normalized weight

[0145]

[0146] The weighted fusion of the outputs corresponding to each image is performed using the average normalized weights to obtain the final output

[0147]

[0148] 4) Feed-forward neural network

[0149] The feed-forward neural network is used to perform a non-linear transformation on the output of the multi-head spatial attention module and can be implemented using any existing neural network structure. Its input O space is the output of the multi-head spatial attention module. As an example, the feed-forward neural network is expressed as:

[0150] O ffn = ReLU(O space W1 + b1)W2 + b2

[0151] where the output and are learnable weight matrices, is a learnable bias term, and D ff is the hidden layer dimension of the feed-forward neural network.

[0152] 5) Output head

[0153] The output head is used to transform the output O ffn of the last layer of the feed-forward neural network into a velocity term. It can be implemented using any existing network, such as a linear transformation network:

[0154] S t = O ffn W head + b head

[0155] where is the predicted velocity, is a learnable parameter. (6) The decoder network φ is used to decode the implicit representation Z1 of the target 3D model obtained by transforming the corrected flow model distribution into an explicit 3D model G.

[0156]

[0157] where are the learnable parameters of the decoder network.

[0158] Compared with other 3D model representations, the 3D Gaussian point cloud has the characteristics of being easy to optimize, having high rendering quality, and fast rendering speed. The 3D Gaussian point cloud is used as the output 3D model representation of the decoder.

[0159] 3D Gaussian point cloud contains K Gaussian points, where K = m·L, indicating the upsampling multiple.

[0160] Each Gaussian point G k has the following parameters: the center position indicating the position of the Gaussian in 3D space, μ = (μ x , μ y , μ z ); the scaling coefficient indicating the size of the Gaussian in 3D space, s = (s x , s y , s z ); the quaternion indicating the rotation direction of the Gaussian in 3D space, q = (q w , q x , s y , s z ); the opacity α ∈ [0, 1], indicating the visibility of the Gaussian; the spherical harmonic function coefficients used to model the three primary colors of the Gaussian under different viewing perspectives, y = (y r , y g , y b ), N y is the number of spherical harmonic coefficients corresponding to each primary color;

[0161] (7) To train the model, we use differentiable rendering to obtain the new perspective image of the 3D Gaussian point cloud.

[0162] Calculate the 3D covariance matrix of each 3D Gaussian point using the scaling transformation matrix and the rotation transformation matrix to calculate, and its mathematical representation is:

[0163] ∑ = RSS T R

[0164]

[0165] Given the pose transformation matrix of the camera and the camera focal length and center point offset project the 3D Gaussian onto the 2D image plane to obtain the 2D center and the 2D covariance matrix

[0166]

[0167] [R c |t c =Π

[0168] ∑′=JR c ∑R c T J T

[0169] where is the depth distance from the camera to each 3D Gaussian, is the Jacobian matrix of the projection transformation, is the rotation matrix of the camera, is the translation of the camera.

[0170] For each pixel perform volume rendering integration on each Gaussian point in the front-to-back order to obtain the pixel color c(p):

[0171]

[0172] where is the opacity of each 2D Gaussian for this pixel, is the color of each 2D Gaussian for this pixel, calculated by spherical harmonic functions:

[0173]

[0174] where are the coefficients of the spherical harmonic function of order l and degree m, is the spherical harmonic function of order l and degree m, is the camera viewing angle, calculated by the following method:

[0175]

[0176] Obtain the corresponding image by performing the above rendering on all pixels

[0177]

[0178] where represents the above rendering function, is the rendered image, are the internal and external parameters of the camera.

[0179] (9) Train the multi-view Transformer network and decoder model based on continuous flow matching and differentiable rendering:

[0180]

[0181] wherein is a weight function of time t, which can assign higher weights at the middle time steps (such as t = 0.5); t is the current time, Z0 is the initial three-dimensional sample, and I gt is the real label image, are the internal and external camera parameters of the real label image, and {I ref} is the input multi-view image; U(0,1) is a uniform distribution from 0 to 1, is a Gaussian distribution, is the distribution of the internal and external camera parameters of the real label image in the training set, p(I ref ) is the distribution of the input multi-view image in the training set, and p(I gt ) is the distribution of the real label image in the training set.

[0182] Through the gradient descent method, the model parameters can be iteratively updated, that is, the multi-view Transformer network θ and the learnable parameters of the decoder .

[0183] That is to say, after the decoder decodes the implicit representation into a three-dimensional Gaussian point cloud, the differentiable rendering technology is used to generate a new view image, and the loss function is used to compare the generated image with the supervised image to optimize the model parameters

[0184] An embodiment of the present invention provides a multi-view driven three-dimensional generation method, including:

[0185] Inputting the multi-view images of the target object into a multi-view driven three-dimensional model constructed by using the construction method of the multi-view driven three-dimensional generation model described in any one of the above embodiments to obtain the three-dimensional Gaussian point cloud of the target object.

[0186] An embodiment of the present invention provides an electronic device, including: a computer-readable storage medium and a processor;

[0187] The computer-readable storage medium is used to store executable instructions;

[0188] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the construction method of the multi-view driven three-dimensional generation model or the multi-view driven three-dimensional generation method described in any one of the above embodiments..

[0189] An embodiment of the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions for causing a processor to execute the construction method of the multi-view driven three-dimensional generation model or the multi-view driven three-dimensional generation method described in any one of the above embodiments.

[0190] An embodiment of the present invention provides a computer program product, including a computer program or instruction, which, when executed by a processor, implements the method for constructing a multi-view driven three-dimensional generation model or the multi-view driven three-dimensional generation method as described in any one of the above embodiments.

[0191] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a multi-view driven 3D generation model, characterized in that Including: Construct a multi-view driven 3D generation model and train it using a dataset. Among them, the multi-view driven 3D generation model includes an encoder, a multi-view Transformer network, and a decoder connected in sequence; the multi-view Transformer network includes several network layers and output heads, and each network layer includes a multi-head self-attention module, a multi-head spatial attention module, and a feed-forward neural network connected in sequence; the dataset is multi-view images of an object, the labeled images are part of the dataset, and the sample images are other parts of the dataset. The process of the training includes: S1. The encoder obtains the feature representation of the multi-view sample images The multi-head self-attention module performs self-attention calculation on the three-dimensional sample Z at the current moment t t to obtain the feature matrix O self , and the multi-head spatial attention module performs spatial attention calculation on and O self to obtain the fused feature matrix O space . The feed-forward neural network performs a non-linear transformation on O space to predict the velocity feature O of the correction flow model ffn . O ff is linearly transformed through the output head to be converted into the transformation velocity S of Z t ; where the initial value of t is 0, and Z0 is a three-dimensional sample randomly sampled from the prior distribution t ; S2, update t to t + Δt, Z t Update to Z t +S t Δt. If the current time reaches the preset time, output the updated Z t and enter S3, otherwise return to S1; S3, the decoder converts the Z output by S2 t into a three-dimensional Gaussian point cloud; S4. Convert the 3D Gaussian point cloud into a new view image, and update the learnable parameters of the multi-view Transformer network and the decoder with the goal of minimizing the loss between the new view image and the labeled image.

2. The method according to claim 1, characterized in that, Among them, O space,h = A space,h V space,h , the b, h, f, n z th elements are the b, h, f, n z th elements, τ is the temperature coefficient, N x is the length of each image feature sequence, v is an element of size B×H×F×N z , the b, h, f, n z th elements, is A space the b, h, fxn x , n z th elements, B is the batch size, H is the number of heads in the multi-head mechanism, F is the number of sample images, N z is the length of the three-dimensional sample sequence, N x is the length of each image feature sequence, b = 1, 2,..., B, h = 1, 2,..., H, f = 1, 2,..., F, n z = 1, 2,..., N z , n x = 1, 2,..., N x , A space = [A space,1 , A space,2 ,... A space,H , D is the feature dimension of Q space , Q space,h , K space,h , V space,h are respectively the h-th heads of Q space , K space , V space , are all learnable parameters of the multi-view Transformer network.

3. The method according to claim 1, wherein The multi-head self-attention module further includes a temporal encoding module, which is used to encode the current time t as t t before the multi-head self-attention module performs self-attention calculation on Z e and then add it to Z t .

4. The method according to claim 3, wherein Generate the time encoding vector \(t\) of the current moment \(t\) through sine and cosine functions e : \(t\) e = \([t_0, t_1, \ldots, t\) 2i , \(t\) 2i+1 , \(\ldots\), ], where \(i\) is the dimension index, \(t\) 2i and \(t\) 2i+1 are the even and odd index values of the time encoding vector respectively, and \(D\) z is the characteristic dimension of \(Z\) t ; Alternatively, generate the time encoding vector t at the current time t through a learnable embedding matrix e : t e = E t [round(t · N e )], where E t is a learnable embedding matrix, [·] represents the indexing operation, round represents the rounding operation, and N e is the length of the embedding matrix.

5. The method according to claim 1, characterized in that The multi-view driven 3D generation model further includes a preprocessing module for performing image normalization processing on the input image.

6. The method according to claim 1, characterized in that, Determine Δt using a dynamic adjustment method: where α is a regulation factor, is the gradient of the velocity field, and θ represents all learnable parameters of the multi-view Transformer network.

7. A multi-view driven 3D generation method, characterized in that Including: Input the multi-view images of the target object into the multi-view driven 3D model constructed by the method described in any one of claims 1-6 to obtain the 3D Gaussian point cloud of the target object.

8. An electronic device, characterized in that, Including: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the construction method of the multi-view driven 3D generation model described in any one of claims 1-6 or the multi-view driven 3D generation method described in claim 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to execute the construction method of the multi-view driven 3D generation model described in any one of claims 1-6 or the multi-view driven 3D generation method described in claim 7.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, the construction method of the multi-view driven 3D generation model described in any one of claims 1-6 or the multi-view driven 3D generation method described in claim 7 is implemented.

Citation Information

Patent Citations

  • Three-dimensional generation method based on improved editability

    CN113112616A

  • Three-dimensional synthesis of geometric perception at all angles

    CN118469827A

Cited By

  • Correction-flow-driven multi-mode unmanned aerial vehicle image enhancement method

    CN122415357A

  • A method for correcting flow-driven multimodal drone image enhancement

    CN122415357B