A three-dimensional model generation method based on multi-granularity visual feature guidance

By using a multi-granularity visual feature-guided method, the problems of insufficient local detail recovery and viewpoint consistency in single-image 3D generation are solved, achieving high-precision and stable 3D model generation and improving the detail quality and consistency of the generated results.

CN122454046APending Publication Date: 2026-07-24ZHONGKE FIFTH CENTURY (HANGZHOU) INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGKE FIFTH CENTURY (HANGZHOU) INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-04-29
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing single-image 3D generation schemes struggle to provide high-supervised density local detail restoration for areas such as edges, sharp corners, and thin walls when reconstructing high-quality 3D models. Furthermore, the generated results lack sufficient consistency with the perspective of the input image, and the transfer of 2D visual cues to 3D geometric representation is inadequate, affecting reconstruction accuracy and stability.

Method used

A method based on multi-granularity visual features is adopted, which extracts global and local visual features through geometric perception sampling, viewpoint consistency alignment, cross-modal feature fusion and latent generation mechanism, establishes an effective mapping between two-dimensional images and three-dimensional models, and improves the fine-grained geometric structure recovery capability and the directional consistency of the generated results.

Benefits of technology

It improves the ability to restore details in areas such as edges and sharp corners, enhances the consistency of the generated results with the input image from a different perspective, improves the overall accuracy and stability of the 3D model, and balances the overall topological rationality of the object with the fidelity of local contour details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454046A_ABST
    Figure CN122454046A_ABST
Patent Text Reader

Abstract

A three-dimensional model generation method based on multi-granularity visual feature guidance, including inputting image and three-dimensional shape data, geometric perception sampling, view consistency alignment, cross-modal joint coding to obtain latent representation, multi-granularity visual feature guided latent space generation, implicit decoding and surface extraction output three-dimensional grid, to improve the recovery ability of high geometric change areas such as edges, sharp corners and thin structures, make the generated three-dimensional model consistent or basically consistent with the input image in the observation direction, meanwhile, utilize the global semantic information and local detail information in the two-dimensional image, establish effective mapping and interaction between two-dimensional visual features and three-dimensional latent representation, improve the overall precision, detail quality and stability of single image driven three-dimensional model generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, computer vision, 3D reconstruction, and 3D content generation, and in particular to a 3D model generation method based on multi-granular visual features guided by a single 2D image, combined with geometric perception sampling, viewpoint consistency constraints, cross-modal feature fusion, and latent generation mechanisms. Background Technology

[0002] With the development of applications such as digital content generation, virtual reality, augmented reality, industrial design, digital twins, and robot simulation, how to quickly and accurately generate a 3D model of a target object from an input 2D image has become an important research direction in the field of 3D generation.

[0003] Existing single-image 3D generation schemes typically rely on neural implicit representations, generative models, or encoder-decoder networks to recover the 3D geometry of an object from a single input 2D image. Compared to multi-view reconstruction schemes or reconstruction schemes that rely on depth sensors, single-image 3D generation schemes have significant advantages in terms of input conditions, deployment costs, and ease of use.

[0004] However, existing technologies still have at least the following problems when recovering high-quality 3D models from 2D images: Training samples are typically sampled randomly or uniformly, which makes it difficult to provide higher supervision density for areas with drastic geometric changes such as edges, sharp corners, thin walls, and thin rod-like structures, resulting in insufficient recovery of local details.

[0005] The 3D shape data used in the training phase is usually in a standardized pose, while the target object corresponding to the input image is often in an arbitrary viewing angle. Existing solutions do not make sufficient use of the viewing angle conditions, which can easily lead to inconsistencies between the generated results and the input image in terms of orientation.

[0006] Existing solutions typically only utilize global semantic information of images or only local visual features, making it difficult to balance the overall topological rationality of the object with the fidelity of local contour details.

[0007] The lack of an effective interaction mechanism between two-dimensional image features and three-dimensional latent representations leads to insufficient transmission of two-dimensional visual cues to three-dimensional geometric representations, thus affecting the reconstruction accuracy and stability of the final three-dimensional model.

[0008] Therefore, it is necessary to propose a new 3D model generation scheme to improve the ability to restore fine-grained geometric structures, enhance the consistency of the generated results with the input image in orientation, and achieve the coordinated guidance of global semantic information and local detail information on the 3D generation process. Summary of the Invention

[0009] In view of the technical problems existing in the prior art, the purpose of this invention is to provide a three-dimensional model generation method based on multi-granularity visual features, so as to improve the recovery ability of high geometric change areas such as edges, sharp corners, and thin structures, make the generated three-dimensional model consistent or basically consistent with the input image in the viewing direction, and at the same time utilize global semantic information and local detail information in two-dimensional images to establish an effective mapping and interaction relationship between two-dimensional visual features and three-dimensional potential representation, thereby improving the overall accuracy, detail quality and stability of single-image driven three-dimensional model generation.

[0010] To achieve the above objectives, this invention provides a method for generating 3D models based on multi-granularity visual features, the method comprising: Obtain the input image of the object to be processed, and obtain the viewpoint parameters or pose parameters corresponding to the input image; Acquire 3D sample data for training or reconstruction, and construct geometrically sensed sampling results based on geometric change information; Perform viewpoint consistency alignment on the 3D sample data according to the viewpoint parameters; The two-dimensional image information, namely the image feature information obtained by the two-dimensional visual feature extraction module from the input image, and the three-dimensional shape information, namely the shape feature information obtained by sampling, alignment and three-dimensional encoding of the three-dimensional sample data, are jointly encoded to obtain a three-dimensional latent representation. Global visual features of the input image are extracted, including category semantics, overall topology and coarse geometric features, which can be modulated by the encoding results of viewpoint parameters or pose parameters, and local visual features, including contour boundaries, local texture changes, detailed structural information and pixel-level or patch-level appearance cues, to form multi-granularity visual conditions. The multi-granularity visual conditions are obtained by fusing the conditional encoding of global visual features, local visual features and viewpoint parameters or pose parameters. Guided by the multi-granularity visual conditions, latent space generation or denoising recovery is performed on the three-dimensional latent representation. The latent space generation or denoising recovery includes applying noise perturbation to the three-dimensional latent representation in the latent space and using a conditional generation network to perform prediction or iterative recovery under the constraints of the multi-granularity visual conditions to obtain the target latent representation. Based on the target latent representation, a three-dimensional implicit field of the target object is reconstructed. The three-dimensional implicit field is a continuous geometric field or other three-dimensional shape representation that takes the coordinates of the spatial query point as input and outputs the occupancy probability, symbolic distance value, probability field value or density value. The output can be a point cloud, voxels, three-plane features, or a 3D mesh model based on the described 3D shape representation. When using a 3D implicit field, the output can be a 3D mesh, point cloud, voxels, or three-plane features through isosurface extraction, mesh reconstruction, point sampling, or voxelization.

[0011] Preferably, according to the method, the characteristic is that, Step S1: Obtain input data: A single 2D image of the object to be reconstructed is acquired. Optionally, camera pose parameters, viewpoint parameters, extrinsic parameter matrices, or pose parameters estimated from the image are also acquired corresponding to the 2D image.

[0012] Step S2: Obtain 3D sample data: Acquire training data of a 3D model, mesh, point cloud, or implicit field corresponding to the target object. The training data may come from at least one of the following: a publicly available 3D model dataset, a 3D asset library obtained by manual modeling, 3D data obtained by scanning and reconstruction, or 3D samples generated by simulation. Calculate or extract at least one of the following: normal vector, curvature information, boundary information, and normal change rate information corresponding to surface points.

[0013] Step S3: Perform geometry-aware sampling. Based on at least one of the normal vector, curvature information, boundary information, and normal change rate information, the sampling weight of different regions is determined so that regions with larger geometric changes have a higher sampling probability than regions with smaller geometric changes, thereby obtaining a sampled 3D training point set or supervision point set.

[0014] Step S4: Perform viewpoint consistency alignment: Based on the viewpoint parameters corresponding to the input image, coordinate transformation is applied to the 3D training point set, supervision point set, or both simultaneously, so that the 3D sample data and the input image establish a correspondence in the viewing direction. Here, the alignment refers to the operation of performing coordinate transformation on the 3D training point set or supervision point set, and the correspondence refers to the mapping relationship between the viewing direction of the input image and the coordinate direction of the 3D sample after the alignment.

[0015] Step S5: Construct a unified latent representation: By using a cross-modal coding module to jointly encode the aligned 3D sample information and the input image information, a 3D latent representation for characterizing the shape of the target object is obtained.

[0016] Step S6: Extract multi-granularity visual features: Extract global visual features from the input image to characterize the overall structure, category semantics, and coarse geometry of the object, as well as local visual features to characterize contour boundaries, texture variations, and local appearance details.

[0017] Step S7: Execution conditions are generated. Using the global and local visual features as conditional information, the three-dimensional latent representation is generated, restored, or progressively denoised in the latent space. The generation, restoration, or progressive denoising process includes starting from the state of perturbation latent variables, random noise, or intermediate latent variables, predicting through one or more generation steps under multi-granularity visual constraints, or progressively removing noise components to obtain the target latent representation.

[0018] Step S8: Perform 3D decoding and reconstruction. The target latent representation is input into the decoding module to predict the occupancy probability, symbolic distance value, probability field value, or other three-dimensional implicit geometric parameters of the query point in order to recover the three-dimensional shape representation of the target object.

[0019] Step S9: Output the 3D model: Based on the three-dimensional shape representation, a three-dimensional mesh model of the target object is obtained through isosurface extraction, mesh restoration, or other surface reconstruction methods; or point clouds, voxel results, or other intermediate results that can be used for subsequent three-dimensional processing are directly output.

[0020] Preferably, in one embodiment, for each surface point in the three-dimensional surface point set, the surface point refers to a point located on the surface of the three-dimensional model or obtained by sampling from the surface of the three-dimensional model, having three-dimensional coordinates and being associated with normal vectors or local geometric properties, its corresponding normal vector is obtained, and the surface point is non-uniformly sampled according to the distribution of the normal vector.

[0021] Furthermore, the normal vectors can be mapped to the angle space, and the angle space can be divided into multiple partitions to statistically analyze the distribution of normal vectors in each partition. For partitions with denser normal vector distribution, the sampling ratio is reduced; for partitions with sparser normal vector distribution, the sampling ratio is increased, thereby reducing the problem of oversampling in flat areas and undersampling in complex areas.

[0022] Preferably, in another embodiment, a sampling weight function can be constructed based on local curvature, the included angle between adjacent normals, the density of boundary points, the local surface change rate, or any combination thereof, and the surface points can be weighted according to the sampling weight function to improve the supervision density of the geometric change region.

[0023] By using the aforementioned geometrically perceptual sampling method, the training process can pay more attention to fine-grained geometric regions such as edges, sharp corners, and thin structures, thereby improving the local detail recovery capability of the final 3D model.

[0024] Preferably, in one embodiment, the viewpoint parameters corresponding to the target object are obtained from the metadata, camera parameters, or pose estimation results of the input image, and rotation transformation, coordinate transformation, or pose mapping is performed on the three-dimensional training samples according to the viewpoint parameters.

[0025] For example, when the viewpoint parameter includes azimuth, a rotation transformation can be applied synchronously to the training point set and the supervision point set around a predetermined coordinate axis so that the direction of the three-dimensional sample data is consistent with or corresponds to the observation direction of the input image.

[0026] Preferably, in other embodiments, the correspondence between the input image orientation and the output 3D model orientation can also be established through coordinate transformation driven by camera extrinsic parameters, pose label supervision, orientation consistency loss constraint or other equivalent methods.

[0027] By introducing a viewpoint consistency alignment mechanism, the orientation deviation between the generated 3D model and the input image can be reduced, thereby enhancing the interpretability and consistency between the input image and the output 3D model.

[0028] Preferably, in one embodiment, the 3D sample data after geometric perception sampling and viewpoint consistency alignment is input into a 3D encoder to obtain 3D hidden features; at the same time, the input image is input into a 2D visual feature extraction module to obtain image features.

[0029] Furthermore, by querying features, attention interaction structures, or other cross-modal fusion structures, two-dimensional image features can participate in the update process of three-dimensional hidden features, thereby injecting local structural information, contour information, or texture boundary information in the two-dimensional image into the three-dimensional latent representation.

[0030] In a preferred embodiment, an asymmetric cross-modal attention structure is employed, using 3D hidden features as the query end and 2D image features as the key and value ends, to achieve guided fusion from 2D visual information to 3D shape representation. After subsequent feature refinement and mapping, the final 3D latent variables are obtained.

[0031] It should be noted that the above-mentioned cross-modal joint coding method is not limited to attention mechanism, but can also adopt gating fusion, conditional mapping, feature concatenation, modulation network or other equivalent structures that can achieve joint modeling of two-dimensional features and three-dimensional features.

[0032] Preferably, in one embodiment, the multi-granularity visual features include at least: Global visual features are used to characterize the category semantics, overall topology, and coarse geometric shape of the target object; Local visual features are used to characterize the outline boundaries, local texture variations, detailed structural information, or pixel-level appearance cues of a target object.

[0033] Furthermore, the pose information, camera parameters, or pose parameters corresponding to the input image can be encoded into a conditional vector, and the global visual features can be modulated using this conditional vector to make the global visual features viewpoint-dependent.

[0034] In a preferred embodiment, the global visual features modulated by pose conditions are jointly fused with the local visual features and used as conditional inputs in the latent space generation process, thereby maintaining both the overall structural rationality of the object and the consistency of local geometric details.

[0035] Preferably, in one embodiment, the latent space generation module applies noise perturbation to the three-dimensional latent variables, and then gradually recovers the target latent variables under the guidance of multi-granularity visual feature conditions, or outputs the target latent variables through an equivalent condition generation method.

[0036] The latent space generation module can be implemented based on Transformer structure, U-Net class structure, state space structure or other conditional generation network, and the present invention does not limit it in this regard.

[0037] During the decoding phase, the target latent variables are input into the decoder, and the corresponding geometric parameters are predicted based on the coordinates of the spatial query point to recover the three-dimensional implicit field, occupancy field, signed distance field, or other three-dimensional shape representation of the target object.

[0038] In a preferred embodiment, the decoding stage further incorporates the interaction between image features and 3D features to further improve the quality of local detail reconstruction and the consistency of the input image.

[0039] Preferably, in one embodiment, the training loss of the joint encoding module includes reconstruction loss and distribution constraint loss.

[0040] The reconstruction loss can include occupancy field prediction loss, symbolic distance regression loss, or other implicit field reconstruction loss; the distribution constraint loss can be used to constrain the stability of the underlying representation distribution.

[0041] The training loss of the latent space generation module may include noise prediction loss, consistency constraint loss, or other loss terms used to improve generation quality and convergence stability.

[0042] It should be understood that the specific form of the loss function, the number of sampling steps, the number of network layers, the feature dimension, the type of optimizer, and the training hyperparameters can all be adjusted according to the actual application scenario, and do not constitute a limitation on the scope of protection of this invention.

[0043] Compared with the prior art, the present invention has at least the following beneficial effects: By using a geometry-aware sampling mechanism, the supervision density of edges, sharp corners, and thin-structure regions can be increased, thereby enhancing the ability to express fine-grained geometry.

[0044] By using a viewpoint consistency alignment mechanism, the directional deviation between the generated model and the input image can be reduced, thereby improving the controllability and visual consistency of the results.

[0045] By employing a cross-modal joint coding mechanism, local structural information from two-dimensional images can be effectively injected into the three-dimensional latent representation, thereby improving the accuracy of three-dimensional geometric reconstruction.

[0046] By coordinating global and local visual features, the overall structural rationality of the object can be balanced with the fidelity of local details.

[0047] This invention is applicable to single-image driven 3D content generation scenarios, has good scalability, and is compatible with different types of visual backbone networks, generative networks, 3D implicit representations, and surface restoration algorithms. Attached Figure Description

[0048] Figure 1 This is a schematic diagram illustrating the overall process of a three-dimensional model generation method based on multi-granularity visual features according to a specific embodiment of the present invention. Figure 2 This is a geometric perception sampling diagram illustrating a specific embodiment of the three-dimensional model generation method based on multi-granularity visual features according to the present invention. Figure 3 This is a schematic diagram illustrating the viewpoint consistency alignment of a three-dimensional model generation method based on multi-granularity visual features according to a specific embodiment of the present invention. Figure 4 This is a schematic diagram illustrating the cross-modal joint coding of a three-dimensional model generation method based on multi-granularity visual features according to a specific embodiment of the present invention; Figure 5 This is a schematic diagram illustrating a multi-granularity visual feature guidance method for generating 3D models based on multi-granularity visual features according to a specific embodiment of the present invention. Figure 6 This diagram illustrates the latent space generation and 3D decoding reconstruction of a 3D model generation method based on multi-granularity visual features, according to a specific embodiment of the present invention. Detailed Implementation

[0049] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. Those skilled in the art will understand that this description is exemplary, and the scope of protection of the present invention is not limited to the following specific embodiments.

[0050] Figure 1 This is a schematic diagram of the overall process of the method of the present invention; Figure 2 This is a schematic diagram of geometric sensing sampling; Figure 3 This is a diagram illustrating the alignment for consistent viewing angles. Figure 4 This is a schematic diagram of cross-modal joint coding; Figure 5 A schematic diagram illustrating multi-granularity visual feature guidance; Figure 6 This is a schematic diagram of potential space generation and 3D decoding reconstruction.

[0051] like Figure 1 As shown, this invention generally includes five aspects: input image and 3D shape data, geometrically perceptual sampling + viewpoint consistency alignment, cross-modal joint encoding to obtain latent representation, multi-granularity visual feature-guided latent space generation, implicit decoding and surface extraction to output a 3D mesh. More specifically, it includes the following eight steps: 1. Obtain the input image of the object to be processed, and obtain the viewpoint parameters or pose parameters corresponding to the input image; 2. Acquire 3D sample data for training or reconstruction, and construct geometrically perceptual sampling results based on geometric change information; 3. Perform viewpoint consistency alignment on the 3D sample data according to the aforementioned viewpoint parameters; 4. Jointly encode the two-dimensional image information and the three-dimensional shape information to obtain a three-dimensional latent representation; 5. Extract global and local visual features from the input image to form multi-granularity visual conditions; 6. Guided by the multi-granularity visual conditions, perform latent space generation or denoising recovery on the three-dimensional latent representation to obtain the target latent representation; 7. Reconstruct the three-dimensional implicit field or other three-dimensional shape representation of the target object based on the target latent representation; 8. Output point cloud, voxel, three-plane feature or three-dimensional mesh model based on the three-dimensional shape representation.

[0052] Figure 2 This diagram illustrates geometrically perceptual sampling. The left image shows a schematic of a surface point set with normal vectors. The middle image shows the normal vectors from the left image mapped to angle space and discretized into B. The diagram on the right shows the balance point selection based on partition-based balanced sampling.

[0053] In one implementation, for each surface point in the three-dimensional surface point set, its corresponding normal vector is obtained, and the surface points are non-uniformly sampled according to the normal vector distribution.

[0054] Furthermore, the normal vectors can be mapped to the angle space, and the angle space can be divided into multiple partitions to statistically analyze the distribution of normal vectors in each partition. For partitions with denser normal vector distribution, the sampling ratio is reduced; for partitions with sparser normal vector distribution, the sampling ratio is increased, thereby reducing the problem of oversampling in flat areas and undersampling in complex areas.

[0055] In another embodiment, a sampling weight function can be constructed based on local curvature, the included angle between adjacent normals, the density of boundary points, the local surface change rate, or any combination thereof, and the surface points can be weighted according to the sampling weight function to improve the supervision density of the geometrically changing region.

[0056] By using the aforementioned geometrically perceptual sampling method, the training process can pay more attention to fine-grained geometric regions such as edges, sharp corners, and thin structures, thereby improving the local detail recovery capability of the final 3D model.

[0057] Figure 3 The diagram illustrates viewpoint consistency alignment. First, an input image with azimuth angle Θ and a standardized pose training point cloud are acquired (left image). Then, the point cloud and the supervision point are rotated synchronously using the same rotation matrix R(Θ), keeping the vertical axis direction unchanged, to obtain orientation-aligned training samples that allow the output 3D model to follow the image direction (right image).

[0058] In one implementation, the viewpoint parameters corresponding to the target object are obtained from the metadata, camera parameters, or pose estimation results of the input image, and rotation transformation, coordinate transformation, or pose mapping is performed on the 3D training samples based on the viewpoint parameters.

[0059] For example, when the viewpoint parameter includes azimuth, a rotation transformation can be applied synchronously to the training point set and the supervision point set around a predetermined coordinate axis so that the direction of the three-dimensional sample data is consistent with or corresponds to the observation direction of the input image.

[0060] In other implementations, the correspondence between the input image orientation and the output 3D model orientation can be established through coordinate transformation driven by camera extrinsic parameters, pose label supervision, orientation consistency loss constraints, or other equivalent methods.

[0061] By introducing a viewpoint consistency alignment mechanism, the orientation deviation between the generated 3D model and the input image can be reduced, thereby enhancing the interpretability and consistency between the input image and the output 3D model.

[0062] Figure 4 A schematic diagram of cross-modal joint coding is shown. As illustrated, local image features are injected into the 3D hidden state, and these local image features participate in feature updates as keys and values.

[0063] In one implementation, the 3D sample data, after geometric perception sampling and viewpoint consistency alignment, is input into a 3D encoder to obtain 3D hidden features; simultaneously, the input image is input into a 2D visual feature extraction module to obtain image features.

[0064] Furthermore, by querying features, attention interaction structures, or other cross-modal fusion structures, two-dimensional image features can participate in the update process of three-dimensional hidden features, thereby injecting local structural information, contour information, or texture boundary information in the two-dimensional image into the three-dimensional latent representation.

[0065] In a preferred embodiment, an asymmetric cross-modal attention structure is employed, using 3D hidden features as the query end and 2D image features as the key and value ends, to achieve guided fusion from 2D visual information to a 3D shape representation. After subsequent feature refinement and mapping, the final 3D latent variable is obtained. Here, the 3D feature token is a serialized feature unit obtained by mapping 3D points, voxels, implicit query points, or 3D hidden features; the cross-modal attention refers to attentional interaction using 3D feature tokens or 3D hidden states as queries and 2D image features as keys and values; the 3D hidden state is an intermediate feature representation output by the 3D encoder or the previous feature update layer, used to form a 3D latent representation after fusing image information.

[0066] It should be noted that the above-mentioned cross-modal joint coding method is not limited to attention mechanism, but can also adopt gating fusion, conditional mapping, feature concatenation, modulation network or other equivalent structures that can achieve joint modeling of two-dimensional features and three-dimensional features.

[0067] Figure 5 This diagram illustrates multi-granularity visual feature guidance.

[0068] In one implementation, local pixel-level or patch-level features simultaneously participate in conditional guidance, and the multi-granularity visual features include at least: Global visual features (i.e.) Figure 5 Global semantic features in the model are used to characterize the category semantics, overall topology, and coarse geometric shape of the target object. Local visual features (i.e.) Figure 5 Local pixel features (in the image) are used to characterize the outline boundary, local texture changes, detailed structural information, or pixel-level appearance cues of a target object.

[0069] Furthermore, the pose information, camera parameters, or position parameters corresponding to the input image can be encoded into a conditional vector, and this conditional vector can be used to modulate the global visual features so that the global visual features have viewpoint relevance (i.e., Figure 5 Pose embedding / modulation in (the process).

[0070] In a preferred embodiment, the pose-conditionally modulated global visual features and local visual features are jointly fused and used as conditional inputs in the latent space generation process (i.e., Figure 5 The concatenated input (potentially generating the main trunk) thus maintains both the overall structural rationality of the object and the consistency of local geometric details.

[0071] Figure 6 A schematic diagram of latent space generation and 3D decoding is shown.

[0072] In one implementation, the latent space generation module applies noise perturbation to the three-dimensional latent variables, and then gradually recovers the target latent variables under the guidance of multi-granularity visual feature conditions, or outputs the target latent variables through an equivalent condition generation method.

[0073] The latent space generation module can be implemented based on Transformer structure, U-Net class structure, state space structure or other conditional generation network, and the present invention does not limit it in this regard.

[0074] During the decoding phase, the target latent variable z is input into the decoder, and the corresponding geometric parameters (occupancy value / SDF / implicit geometric value) are predicted based on the coordinates of the spatial query point to recover the three-dimensional implicit field, occupancy field, signed distance field or other three-dimensional shape representation of the target object.

[0075] In a preferred embodiment, the decoding stage further incorporates the interaction between image features and 3D features to further improve the quality of local detail reconstruction and the consistency of the input image.

[0076] The loss function and training method of the joint coding module in the cross-modal joint coding scheme are explained in detail below.

[0077] In one implementation, the training loss of the joint encoding module includes a reconstruction loss and a distribution constraint loss. The reconstruction loss may include occupancy field prediction loss, symbolic distance regression loss, or other implicit field reconstruction losses; the distribution constraint loss can be used to constrain the stability of the underlying representation distribution.

[0078] The training loss of the latent space generation module may include noise prediction loss, consistency constraint loss, or other loss terms used to improve generation quality and convergence stability.

[0079] It should be understood that the specific form of the loss function, the number of sampling steps, the number of network layers, the feature dimension, the type of optimizer, and the training hyperparameters can all be adjusted according to the actual application scenario, and do not constitute a limitation on the scope of protection of this invention.

[0080] The following example illustrates the generation of an end-to-end 3D model based on a single RGB image.

[0081] First, input an RGB image of the object to be processed, and obtain the corresponding camera azimuth, pitch, and radius parameters, or obtain the corresponding camera extrinsic parameter matrix.

[0082] Surface points and their normal vectors are sampled from the 3D training mesh, and near-surface points and spatial points are further obtained as implicit field supervision samples.

[0083] Geometric sensing sampling is performed on surface points, and the sampling ratio of complex regions is increased based on normal vector distribution, curvature information, or local rate of change information to obtain a set of sampled training points.

[0084] Based on the viewpoint parameters corresponding to the input image, a synchronous coordinate transformation is performed on the sampled training point set and supervision point set to establish a correspondence with the direction of the input image.

[0085] The aligned 3D point set is input into the 3D encoder, the input image is input into the 2D visual feature extraction module, and the 3D latent variables are obtained through cross-modal joint encoding.

[0086] Global and local visual features of the input image are extracted and combined with pose information to form multi-granular visual conditions.

[0087] Under the guidance of the multi-granularity visual conditions, latent space generation is performed on the three-dimensional latent variables to obtain the target latent variables.

[0088] The target latent variables are input into the decoder to predict the geometric parameters of the query point and recover the three-dimensional implicit field of the target object.

[0089] Surface extraction is performed based on the recovered 3D implicit field, and a 3D mesh model of the target object is output.

[0090] In the optional parameter settings, the latent variable dimension, the number of latent space generation steps, the number of network layers, the image resolution, and the number of sampling points can all be configured according to task requirements. The above parameters are only used to illustrate the feasibility of this invention and do not constitute a limitation on the scope of protection of this invention.

[0091] In other implementations, the method output may also be a point cloud, voxel field, three-plane feature, or other intermediate result that can characterize a three-dimensional shape, and then converted into a mesh model through post-processing.

[0092] In summary, the present invention has been described in detail through specific embodiments. However, those skilled in the art will understand that this description is exemplary, and the present invention can be modified and altered in various ways. Such modifications and alterations should all fall within the protection scope of the present invention without departing from its spirit and intent. For example, the present invention can also have the following alternatives: the visual feature extraction module can employ a visual Transformer, convolutional neural network, hierarchical visual encoder, or other pre-trained visual model; the latent space generation module can employ a Transformer, U-Net-like network, state-space model, or other conditional generation network; pose information can be injected using Euler angles, quaternions, camera extrinsic matrix, position encoding, conditional vector concatenation, or other equivalent forms; the three-dimensional geometric representation can employ occupancy field, signed distance field, probability field, three-plane features, or other representations that can recover three-dimensional shapes; the cross-modal interaction method can employ cross-attention, gated fusion, feature concatenation, conditional convolution, or other equivalent methods; the surface recovery method can employ Marching Cubes, Dual Contouring, ball tracking, or other isosurface extraction methods, and so on. The scope of protection of the present invention is defined by the appended claims.

Claims

1. A method for generating 3D models based on multi-granularity visual features, the method comprising: Obtain the input image of the object to be processed, and obtain the viewpoint parameters or pose parameters corresponding to the input image; Acquire 3D sample data for training or reconstruction, and construct geometrically sensed sampling results based on geometric change information; Perform viewpoint consistency alignment on the 3D sample data according to the viewpoint parameters; Two-dimensional image information and three-dimensional shape information are jointly encoded to obtain a three-dimensional latent representation; Extract global and local visual features from the input image to form multi-granular visual conditions; Guided by the multi-granularity visual conditions, latent space generation or denoising recovery is performed on the three-dimensional latent representation to obtain the target latent representation; Based on the target latent representation, reconstruct the three-dimensional implicit field or other three-dimensional shape representation of the target object; The output is a point cloud, voxel, three-plane feature, or three-dimensional mesh model based on the three-dimensional shape representation.

2. The method according to claim 1, characterized in that, Obtain input data: Obtain a single 2D image of the object to be reconstructed; Acquire 3D sample data: Acquire 3D model, mesh, point cloud or implicit field training data corresponding to the target object, and calculate or extract at least one of the following: normal vector, curvature information, boundary information, normal change rate information corresponding to surface points; Perform geometric perception sampling: Based on at least one of the normal vector, curvature information, boundary information, and normal change rate information, determine the sampling weight of different regions so that regions with larger geometric changes have a higher sampling probability than regions with smaller geometric changes, thereby obtaining a sampled 3D training point set or supervision point set; Perform viewpoint consistency alignment: Based on the viewpoint parameters corresponding to the input image, apply coordinate transformation to the 3D training point set, supervision point set, or both simultaneously, so that the 3D sample data and the input image establish a correspondence in the viewing direction; Constructing a unified latent representation: The aligned 3D sample information and the input image information are jointly encoded using a cross-modal coding module to obtain a 3D latent representation for characterizing the shape of the target object; Extracting multi-granular visual features: Extracting global visual features from the input image to characterize the overall structure, category semantics and coarse geometric shape of the object, as well as local visual features to characterize the contour boundary, texture variation and local appearance details; Execution condition generation: Using the global visual features and local visual features as conditional information, the three-dimensional latent representation is generated, restored, or progressively denoised in the latent space to obtain the target latent representation; Perform 3D decoding and reconstruction: Input the latent representation of the target into the decoding module to predict the occupancy probability, symbolic distance value, probability field value or other 3D implicit geometric parameters of the query point in order to recover the 3D shape representation of the target object; Output 3D model: Based on the 3D shape representation, obtain the 3D mesh model of the target object through isosurface extraction, mesh restoration or other surface reconstruction methods; or directly output point cloud, voxel results or other intermediate results that can be used for subsequent 3D processing.

3. The method according to claim 2, characterized in that, For each surface point in the three-dimensional surface point set, its corresponding normal vector is obtained, and the surface points are sampled non-uniformly according to the distribution of the normal vector.

4. The method according to claim 3, characterized in that, Normal vectors can be mapped to angle space, and the angle space can be divided into multiple partitions to statistically analyze the distribution of normal vectors in each partition.

5. The method according to claim 3 or 4, characterized in that, A sampling weight function is constructed based on local curvature, the included angle between adjacent normals, boundary point density, local surface change rate, or any combination thereof, and surface points are sampled in a weighted manner according to the sampling weight function to improve the supervision density of the geometrically changing region.

6. The method according to any one of claims 1-4, characterized in that, Obtain the viewpoint parameters corresponding to the target object from the metadata, camera parameters, or pose estimation results of the input image, and perform rotation transformation, coordinate transformation, or pose mapping on the 3D training samples based on the viewpoint parameters.

7. The method according to any one of claims 1-4, characterized in that, By using camera extrinsic-driven coordinate transformation, pose label supervision, orientation consistency loss constraints, or other equivalent methods, the correspondence between the orientation of the input image and the orientation of the output 3D model is established.

8. The method according to any one of claims 1-4, characterized in that, The 3D sample data, after geometric perception sampling and viewpoint consistency alignment, is input into the 3D encoder to obtain 3D hidden features; at the same time, the input image is input into the 2D visual feature extraction module to obtain image features.

9. The method according to any one of claims 1-4, characterized in that, By querying features, attention interaction structures, or other cross-modal fusion structures, two-dimensional image features can participate in the update process of three-dimensional hidden features, thereby injecting local structural information, contour information, or texture boundary information in two-dimensional images into the three-dimensional latent representation.

10. The method according to claim 9, characterized in that, An asymmetric cross-modal attention structure is adopted, using 3D hidden features as the query end and 2D image features as the key and value ends, to achieve guided fusion from 2D visual information to 3D shape representation; after subsequent feature refinement and mapping, the final 3D latent variables are obtained.

11. The method according to claim 1, characterized in that, The multi-granularity visual features include at least: Global visual features are used to characterize the category semantics, overall topology, and coarse geometric shape of the target object; Local visual features are used to characterize the outline boundaries, local texture variations, detailed structural information, or pixel-level appearance cues of a target object.

12. The method according to claim 11, characterized in that, The global visual features modulated by the pose conditions are jointly fused with the local visual features and used as conditional inputs in the latent space generation process, thereby maintaining the rationality of the overall structure of the object and the consistency of local geometric details at the same time.

13. The method according to claim 1, characterized in that, The training loss for joint encoding includes reconstruction loss and distribution constraint loss; The reconstruction loss can include occupancy field prediction loss, symbolic distance regression loss, or other implicit field reconstruction losses; the distribution constraint loss can be used to constrain the stability of the underlying representation distribution. The training loss of the latent space generation module may include noise prediction loss, consistency constraint loss, or other loss terms used to improve generation quality and convergence stability.