Single-view image three-dimensional modeling method and system based on comparative learning
Through contrast learning and super-resolution image enhancement technology, the problems of texture inconsistency and structural incoherence in single-view 3D modeling are solved, and high-precision and stable 3D modeling effects are achieved, which is suitable for virtual human modeling, 3D asset generation and immersive content production.
Patent Information
- Application Number
- CN202510630274.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-05
AI Technical Summary
Existing single-view 3D modeling methods suffer from inconsistent textures, incoherent structure reconstruction, and insufficient semantic distinction capabilities, resulting in poor 3D reconstruction results. In particular, it is difficult to generate high-precision and stable 3D models in the absence of multi-view supervision.
A contrastive learning-based method is adopted to generate multi-view candidate images by pre-training a two-dimensional diffusion model, construct an image triplet structure, establish a perceptual contrast loss function and a quantity-aware triplet loss, and combine the super-resolution image enhancement module and the neural radiation field reconstruction module to optimize the three-dimensional modeling process.
It improves the texture consistency, structural coherence and stability of 3D modeling, and generates high-precision, high-fidelity 3D models suitable for virtual human modeling, 3D asset generation and immersive content production.
Smart Images

Figure CN120599170A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and three-dimensional modeling, and more specifically, to a single-view image three-dimensional modeling method and system based on contrast learning. Background Art
[0002] With the rapid development of computer graphics and computer vision technologies, the generation and reconstruction of three-dimensional content are becoming increasingly important in a variety of practical application scenarios, including virtual reality, augmented reality, film and television special effects, game development, medical imaging, and digital museums. Single-view 3D reconstruction, in particular, is a key area of current research. Due to its low input acquisition cost, ease of operation, and strong applicability, it has become a key technical path for automated 3D modeling. The goal of this task is to reconstruct a realistic, structurally sound, and textured 3D model from a single static RGB image, as far as possible, in the absence of multi-view observation information. This provides reliable data support for downstream 3D asset generation, digital human construction, and spatial understanding and interaction.
[0003] Traditional 3D reconstruction methods mostly rely on multi-view geometric information, such as structured light, stereo vision, depth sensor input, or multi-frame camera capture data. These methods typically construct spatial point clouds or volume meshes based on geometric projection consistency, feature matching, and triangulation. However, in the context of single-view input, due to the lack of depth information and spatial constraints, 3D reasoning based solely on the pixel intensity, edge cues, and semantic structure of the 2D image inevitably faces technical difficulties such as structural ambiguity, depth ambiguity, and texture distortion. Therefore, in recent years, researchers have gradually introduced deep neural networks into the task of single-view reconstruction, hoping to improve reconstruction performance by learning prior knowledge to model the relationship between structure and texture in 3D space.
[0004] In particular, the introduction of neural implicit modeling technologies, represented by Neural Radiance Fields (NeRF), provides a high-precision representation method based on continuous volume functions for 3D reconstruction. NeRF can simulate the distribution of color and density in a high-resolution space and restore high-quality images through volume rendering integrals. Under multi-view supervision, it demonstrates rendering fidelity far exceeding that of traditional methods. However, in single-view modeling scenarios, due to the lack of complete viewpoint information as training supervision, the performance of the NeRF model will degrade significantly, making it difficult to restore the accurate volume density distribution. The rendered images often suffer from severe geometric deformation and texture degradation.
[0005] To overcome the lack of supervision in single-view NeRF reconstruction, researchers have introduced generative models as an indirect source of supervision. Pre-trained generative networks, particularly text-image diffusion models like Stable Diffusion and DeepFloyd IF, have garnered widespread attention. Using techniques like Score Distillation Sampling (SDS), researchers have attempted to guide diffusion models to generate multiple images with varying viewpoints from a single input image. These images serve as pseudo-supervision signals to indirectly drive the NeRF model's learning of the spatial volume function, thereby enabling data-driven 3D reconstruction without the need for actual multi-view supervision.
[0006] Although SDS technology has significantly expanded the input modalities and supervision forms of 3D modeling and improved the flexibility of single-view reconstruction to a certain extent, it still relies on the quality and consistency of the image generated by the diffusion model. However, the current mainstream diffusion model has a poor performance in multi-view reconstruction. Figure 1 There are generally serious defects in consistency. First, the perspective images generated by the diffusion model often have problems such as missing texture details, blurred edges, and object deformation, making it difficult to form a truly continuous perspective change trajectory, seriously affecting the spatial continuity and texture integrity of the 3D model. Second, the lack of a clear semantic distinction mechanism between generated images makes it impossible to effectively distinguish and utilize high-quality and low-quality images during training. This causes a decline in the quality of the training loss signal and oscillations in the optimization path, reducing the stability and convergence efficiency of the modeling process.
[0007] Furthermore, most existing SDS frameworks employ per-image optimization or per-frame loss minimization strategies, failing to establish a comparative relationship or semantic ranking mechanism between candidate images. Lacking discriminative supervision objectives, they are unable to fully utilize potential negative sample information between images to enhance the model's discriminative capabilities. Furthermore, due to the resolution limitations of the diffusion model, the generated images generally suffer from insufficient clarity and missing details. These low-quality images easily interfere with the modeling of 3D geometric structures during training, ultimately weakening the model's ability to recover textures and reducing the accuracy of spatial shape fitting.
[0008] Therefore, current technology still urgently needs a new single-view 3D modeling method that can integrate contrast learning mechanism and high-resolution image enhancement means, effectively guide the diffusion model to output multi-view images with structural consistency and texture discernibility, and combine perceptual contrast loss function, quantity-aware triplet optimization strategy, super-resolution enhancement module and neural rendering network joint optimization to achieve high-precision, high-fidelity and high-stability single-view 3D reconstruction effect, so as to break through the technical bottlenecks of existing methods in geometric coherence, texture consistency and training stability. Summary of the Invention
[0009] In response to the technical problems existing in the prior art, the present invention provides a single-view image 3D modeling method and system based on contrastive learning. The present invention effectively solves the problems of inconsistent multi-view textures, incoherent structural reconstruction, and insufficient semantic differentiation ability in existing diffusion generation methods. It has the advantages of simple structure, good convergence stability, and high 3D modeling fidelity. It is suitable for virtual human modeling, 3D asset generation, and immersive content production scenarios.
[0010] According to a first aspect of the present invention, a single-view image 3D modeling method based on contrastive learning is provided, comprising: Obtaining a single-view image as the initial input for the 3D modeling process; Feed a single-view image into a pre-trained 2D image diffusion model to generate candidate images from multiple perspectives. Based on the obtained multiple candidate images, positive and negative samples are constructed to form an image triplet structure; Establish a perceptual contrast loss function based on image triplets, calculate the perceptual loss between images, and perform contrast optimization at the perceptual feature level; Establish a quantity-aware triplet loss weight mechanism to dynamically adjust the contribution ratio of positive and negative samples in training; The super-resolution image enhancement module is embedded in the candidate image optimization process to perform pixel-level enhancement on the generated candidate images of multiple perspectives. The overall optimization objectives of perceptual contrast loss and quantity-aware triplet loss are jointly trained with the neural radiation field reconstruction module to generate three-dimensional meshes or volume rendering results, and output the reconstructed three-dimensional representation results.
[0011] On the basis of the above technical solution, the present invention can also make the following improvements.
[0012] Optionally, the single-view image is a static image in three-channel RGB format; starting from a single input image, a pre-trained two-dimensional diffusion model is combined with score distillation sampling to generate multi-view candidate images.
[0013] Optionally, constructing positive and negative samples based on the obtained multiple candidate images includes: The candidate images are combined in pairs, and the geometric structure consistency and texture detail similarity between the images are evaluated based on the predefined structural similarity measurement function and texture distribution index; Images whose structure and texture are highly similar to the input image are selected as positive samples, and images whose geometric structure or texture distribution deviates significantly are selected as negative samples to form an image triplet structure.
[0014] Optionally, establishing a perceptual contrast loss function based on image triplets, calculating inter-image perceptual loss, and performing contrast optimization at the perceptual feature level includes: Construct a contrast loss module and use a pre-trained feature extraction network to perform forward propagation on the input image, positive sample image, and negative sample image to extract the multi-scale semantic feature tensor of the intermediate layer respectively; Normalize the feature tensor and project it into a feature space of the same dimension to construct a comparable embedding representation; The semantic difference between the input image and the positive and negative samples is measured using Euclidean distance, cosine similarity, or a perceptual metric based on feature difference, and a perceptual contrast loss function based on image triplets is established. The distance between the input image and the positive sample is minimized, while the distance between the input image and the negative sample is maximized. By performing gradient backpropagation on the loss function, the diffusion model is guided to gradually improve the consistency with the positive sample structure during the multi-view image generation process, and suppress the generation of blurred images similar to negative samples.
[0015] Optionally, establishing a quantity-aware triplet loss weight mechanism to dynamically adjust the contribution ratio of positive and negative samples in training includes: Dynamically count the number of positive and negative samples and their quality scores in each training batch; Assign a confidence-based dynamic weight coefficient to each triplet loss term to increase the training weight of high-quality positive samples and suppress the influence of low-quality negative samples; The weight mechanism is integrated into the overall contrast loss function to achieve adaptive control of the sample distribution of the global training process.
[0016] Optionally, embedding the super-resolution image enhancement module into the candidate image optimization process to perform pixel-level enhancement on the generated candidate images of multiple perspectives includes: Use any image super-resolution model with detail restoration capability to enlarge the candidate image; While keeping the structure unchanged, it enhances image edge details, texture clarity and resolution, and improves the quality of perceptual loss signals and training stability.
[0017] Optionally, the super-resolution image enhancement module adopts a super-resolution model structure based on a deep neural network, including an image enhancement network with detail restoration capabilities; the super-resolution image enhancement module is deployed between the diffusion model output image and the perceptual loss calculation module, with input being multiple low-resolution candidate images and output being an enhanced high-resolution image set.
[0018] Optionally, the neural radiation field reconstruction module is a neural radiation field or other volume rendering network, and the joint training of the overall optimization objectives of perceptual contrast loss and quantity-aware triplet loss with the neural radiation field reconstruction module includes: Receive the input image and its generated multi-view image set, and extract dense view information as conditional input; A mapping function between color field and density field is established in volume rendering space, and consistent image reconstruction output is synthesized through volume rendering integral. The volume rendering network parameters are jointly optimized based on the three-dimensional spatial consistency loss and the above-mentioned contrast loss to obtain high-quality three-dimensional model output with geometric coherence and texture consistency.
[0019] According to a second aspect of the present invention, there is provided a single-view image 3D modeling system based on contrastive learning, comprising: An input module, used to obtain a single-view image as an initial input for the 3D modeling process; Diffusion model generation module, used to generate multiple candidate images with different perspectives based on a single input RGB image; A perceptual contrastive learning module that constructs positive and negative sample pairs from candidate images generated by diffusion and performs feature-level contrastive training; Super-resolution image enhancement module, used to restore the diffusion-generated image at a high-quality detail level; A neural rendering and reconstruction module is used to fuse the perceptual loss signal with the multi-view image to output a 3D structure; Output module, used to output the final 3D modeling results.
[0020] According to a third aspect of the present invention, an electronic device is provided, comprising a memory and a processor, wherein the processor is configured to implement the steps of a single-view image three-dimensional modeling method based on contrastive learning when executing a computer program stored in the memory.
[0021] Technical effects and advantages of the present invention: This invention provides a single-view 3D modeling method and system based on contrastive learning. Starting from a single input image, this method utilizes a pre-trained 2D diffusion model combined with a score distillation sampling technique to generate multi-view candidate images. Next, candidate image pairs are classified into positive and negative samples based on texture fidelity and structural coherence. A perceptual contrastive learning mechanism is constructed, incorporating a VGG network or CLIP feature extractor to calculate a perceptual loss between images, thereby optimizing texture consistency during 3D content generation. Furthermore, a quantity-aware triplet loss function is proposed to dynamically adjust the contribution ratio of positive and negative samples during training, improving the stability and discriminative power of contrastive learning. A super-resolution reconstruction module is then embedded to enhance image detail and improve the quality of contrast gradient signals. Finally, the perceptual contrast loss is combined with a neural radiance field (NeRF) or volume rendering module to generate high-quality 3D meshes or volume renderings. This invention effectively addresses the problems of multi-view texture inconsistency, incoherent structural reconstruction, and insufficient semantic discrimination capability in existing diffusion generation methods. It offers advantages such as simple structure, good convergence stability, and high 3D modeling fidelity. It is suitable for scenarios such as virtual human modeling, 3D asset generation, and immersive content production. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 An overall flow chart of a single-view image 3D modeling method based on contrastive learning provided by an embodiment of the present invention; Figure 2 A schematic diagram of the modules of a single-view image 3D modeling system based on contrastive learning provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0024] It should be noted that, in terms of practical application, the single-view image 3D modeling method and system based on contrastive learning described in this embodiment of the invention are applicable to, but not limited to, the following fields: Virtual human modeling: Input a single head photo to generate a rotatable 3D head model, suitable for digital avatar and character change systems.
[0025] 3D asset generation: used for product modeling in e-commerce, metaverse, and game development, automatically generating rotatable models from product images.
[0026] Medical image inference: Using auxiliary models to input X-rays or single-frame CT images, a three-dimensional representation of organ structure is constructed to serve disease localization and preoperative simulation.
[0027] Industrial reverse modeling: Convert a single image of a device component into a measurable 3D structure for digital twinning and defect detection.
[0028] It is understandable that based on the defects in the background technology, the embodiment of the present invention proposes a single-view image 3D modeling method based on contrast learning, specifically as follows: Figure 1 As shown, the following steps are included: S1. Obtain an input single-view image as the initial input for the 3D modeling process; In this embodiment, the image is a static image in three-channel RGB format. This implementation method supports two typical input forms: one is a single static two-dimensional image in three-channel RGB input format with variable resolution and dimensions of 512×512; the other is an optional image prior, such as an image segmentation mask, a semantic label, or a view control encoding vector, which is used to guide the diffusion model's change strategy in spatial structure or posture during the generation process.
[0029] It should be noted that the single-view image 3D modeling method based on contrastive learning proposed in the embodiment of the present invention can be deployed in a standard GPU server or a local training workstation. During the implementation process, the input image format is a single three-channel RGB image, supporting mainstream formats such as JPG, PNG, WebP, and the recommended initial resolution is 512×512 or 768×768. After receiving the image, the system uniformly adjusts the size, center crops, and normalizes it to the [-1, 1] interval, and converts it into a tensor structure for network input. If the image contains a complex background or a multi-object scene, the image segmentation module can be used in conjunction to extract the main target area to enhance modeling stability and attention concentration.
[0030] S2, feed the single-view image into a pre-trained 2D image diffusion model to generate candidate images from multiple perspectives; It should be noted that this method generates multi-view candidate images from a single input image using a pre-trained 2D diffusion model combined with score distillation sampling. The 2D image diffusion model is guided by the Score Distillation Sampling mechanism. These candidate images exhibit perspective variability in spatial pose while maintaining semantic consistency with the input image in terms of content representation, enabling the construction of multi-view observation information for subsequent 3D modeling.
[0031] The embodiment of the present invention adopts the diffusion model-based score distillation sampling (SDS) mechanism as the core technical path for multi-view image generation. The goal of this module is to automatically generate multiple candidate images that are semantically consistent with the image but have different perspectives under the condition of providing only one input image, providing pseudo multi-view observation data for subsequent three-dimensional modeling. The generated image must meet three requirements at the same time: (1) consistency with the original image at the semantic level, (2) reasonable perspective perturbation at the spatial structure level, and (3) high resolution and detail fidelity at the texture quality level. To this end, this module design includes two major substructures: a guided generation network based on SDS, and a perspective perturbation mechanism based on posture control.
[0032] The score distillation sampling mechanism is essentially a diffusion sampling method guided by score matching. Its core principle is to use a pre-trained 2D text-image diffusion model (such as Stable Diffusion) as a pseudo-supervisory model. In successive diffusion iterations, it optimizes the generated samples using the image embedding gradient, thereby forming an image output path with conditional consistency and guided diversity. Unlike traditional stepwise diffusion sampling, the SDS mechanism compares the latent space difference between the image predicted by the diffusion model and the conditional target, guiding the diffusion process directly in the gradient space, making the generated image more consistent with the input semantics.
[0033] At the implementation level, the diffusion model accepts a single RGB input image as a conditional input, which is first mapped into a fixed-length latent vector through the CLIP image encoder. , serving as the target representation for the entire diffusion guidance process. The diffusion sampling process uses the random Gaussian noise image x_T as the initial state and gradually performs reverse denoising through the conditional diffusion network parameterized by the U-Net structure. The update process for each step t is:
[0034] Where η is the learning rate or step size coefficient, Represents the gradient direction relative to the image embedding target in the current diffusion state.
[0035] The embodiment of the present invention introduces an image-to-image guidance form in this process, that is, by encoding the input image instead of text description, the SDS mechanism can perform guidance generation without language labels, which is more in line with the structural fidelity requirements of the 3D reconstruction task. The guidance target is defined in the CLIP space as the original image embedding vector , and the gradient direction between the embedding vector h_t of the current image sampling result can be derived by calculating the feature difference (such as cosine similarity or Euclidean distance). The potential representation of the final output image of each sampling is fed into the image decoder to generate the final visualization image as one of the candidate multi-view samples.
[0036] This module generates K candidate images per sampling, where K is adjustable and typically ranges from 8 to 16. The output image resolution can be set to 256×256, 512×512, or 768×768, depending on the video memory budget. All candidate images are stored in a temporary buffer and fed into the positive and negative sample selection module in the next stage.
[0037] To further enhance the spatial structure diversity of the generated images and control their viewpoint offset, this embodiment of the present invention introduces a pose perturbation mechanism. This perturbation injects the camera pose information implicit in the conditional input during the diffusion process. Because the original Stable Diffusion model does not explicitly model the camera viewpoint, this method uses two indirect means to achieve this goal.
[0038] First, add a pseudo-view encoding vector before embedding the original image , which can be generated by learning the posture perturbation prior or sampling from a preset distribution using random perturbation. The perturbation vector is embedded in the original image. Add together to form the final guidance goal This strategy can be viewed as imposing a “perspective offset” on the original image guidance target in CLIP space, guiding the generated result to produce moderate changes in the spatial structure.
[0039] Secondly, during the image decoding phase, affine transformation, perspective remapping, or occlusion modeling is performed on the final image to simulate changes in camera shooting angles. This strategy achieves consistency with the sampling directions during NeRF training through image post-processing, providing more accurate perspective information support for the subsequent spatial consistency loss function.
[0040] Furthermore, to ensure that the geometric structure of the generated image does not deviate significantly from the original image, the amplitude of the view perturbation is limited to a structural consistency tolerance. Even with the perturbation, the generated image still needs to pass the perceptual feature screening of subsequent modules. The perturbation range can be controlled by the parameter θ, which is generally set to ±20° horizontal view change and ±10° elevation and pitch change.
[0041] Through the above mechanism, the multi-view image generation module designed in the embodiment of the present invention not only ensures the semantic consistency and quality controllability of the generated images, but also provides sufficiently rich spatial variations for positive and negative sample screening and feature enhancement in the subsequent comparative learning stage, constituting the first processing link of the entire three-dimensional modeling system.
[0042] S3. constructing positive and negative samples based on the obtained multiple candidate images to form an image triplet structure; To ensure that images entering the 3D modeling phase possess clear structure, accurate texture, and a reasonable perspective, the present invention employs an image quality assessment mechanism after the diffusion model generates candidate images. Based on this, positive and negative sample pairs are constructed for subsequent perceptual contrastive learning optimization. The core concept of this mechanism is to automatically determine the quality of candidate images based on perceptual feature differences without manual labeling, and to construct structured triples using the input image as an anchor. This process involves two main substeps: embedding feature extraction and similarity assessment, and positive and negative sample partitioning and triplet construction.
[0043] The constructing of positive and negative samples based on the obtained multiple candidate images includes: The candidate images are combined in pairs, and the geometric structure consistency and texture detail similarity between the images are evaluated based on the predefined structural similarity measurement function and texture distribution index; Images whose structure and texture are highly similar to the input image are selected as positive samples, and images whose geometric structure or texture distribution deviates significantly are selected as negative samples, forming an image triplet structure, namely (input image, positive sample image, negative sample image).
[0044] Specifically, this embodiment extracts embedded feature representations of the candidate images in the structural and texture dimensions based on the multiple candidate view images generated by the score distillation sampling mechanism in step S2; Calculate the structural consistency score and texture similarity index between all candidate images and the input image in the feature space; Sort candidate images according to the scoring indicators and set thresholds to distinguish between positive and negative sample images; The top N candidate images with the highest scores are selected as positive sample images, and the bottom M images with the lowest scores are selected as negative sample images to form a set of training triples. Among them, the number of positive and negative samples N and M can be dynamically adjusted according to the current training stage, and when the sample is insufficient, multiple negative samples are allowed to share the same positive sample to construct multiple training pairs; The sample pair construction process is performed automatically without manual labeling intervention, supporting seamless integration of the end-to-end modeling process and comparative supervision objectives.
[0045] In addition, embodiments of the present invention use a pre-trained image encoder as a perceptual feature extractor, which maps the input image and the generated candidate images into a unified semantic embedding space, used to measure the structural and texture differences between images. This embedding space can be defined by a VGG network, a CLIP visual encoder, or a multi-layer feature network based on convolution or Transformer.
[0046] In the specific implementation, each image First, it is uniformly scaled to a fixed resolution (such as 224×224 or 256×256) and normalized to the standard input domain of [0,1] or [-1,1] as the input tensor of the encoder. The encoder performs forward propagation to extract intermediate feature representations of the image at multiple semantic scales. Taking VGG-16 as an example, the outputs of its conv3_3, conv4_3, and conv5_3 layers respectively extract edge structure, texture block, and overall semantic information to form a feature tensor set. ,in .
[0047] The extracted tensor can be converted into a vector representation of uniform dimension through global average pooling, or fused through the convolutional attention module to form the final embedding vector
[0048] , where d is usually 512 or 768. In order to improve the discriminative ability of feature consistency measurement, all embedding vectors are L2 normalized before construction, that is, .
[0049] In the CLIP encoding path, image I passes through the visual Transformer encoder to obtain the projected embedding e_clip, whose structure contains a multi-layer Attention mechanism and residual connections, which can preserve spatial structure and semantic aggregation features. Its output is frozen after pre-training and does not participate in the current task fine-tuning.
[0050] After feature extraction is completed, the system embeds the vector of each candidate image of the input image e_anchor Perform similarity calculations. Common metrics include: (a) Cosine similarity:
[0051] (b) Euclidean distance:
[0052] (c) Perceptual loss (LPIPS): uses the weighted superposition of perceptual distances from multiple layers in the pre-trained network.
[0053] This embodiment of the present invention uses (a) cosine similarity as the primary metric and uses this metric to rank all candidate images against the anchor image, with higher rankings indicating closer similarity. An upper limit P_top for positive sample selection and a lower limit N_bottom for negative sample selection are set. In each round of construction, the first P_top images are selected as positive candidate samples, and the last N_bottom images are selected as negative candidate samples.
[0054] The system randomly samples a positive sample e_pos from the positive sample set and a negative sample e_neg from the negative sample set. Together with the positive sample e_anchor, these samples construct an image triplet (e_anchor, e_pos, e_neg). To increase training diversity and contrast, each anchor image can construct M sets of different positive and negative sample triplets, thereby increasing the number of training samples and gradient coverage.
[0055] In addition, to ensure the quality of triplet samples, the embodiment of the present invention introduces a triplet screening threshold, namely: Sim(e_anchor, e_pos)>τ_pos Sim(e_anchor, e_neg)<τ_neg Among them, τ_pos is usually set to 0.8 and τ_neg is set to 0.3. When some candidate images cannot meet the above conditions, the system can automatically discard the corresponding triples, or execute the image enhancement resampling mechanism to regenerate new candidate images to supplement the training sample set.
[0056] Through the above mechanism, the embodiment of the present invention automatically extracts the best and worst structural samples from the diffusion-generated image without manual labeling, constructs a perceptual comparison training set with discriminative and generalizable capabilities, and provides high-quality supervision signals for subsequent loss functions.
[0057] S4. Establish a perceptual contrast loss function based on image triplets, calculate the perceptual loss between images, and perform contrast optimization at the perceptual feature level; In order to improve the system's ability to automatically distinguish the quality of image structures in an unsupervised scenario, the embodiment of the present invention designs a perceptual contrast loss mechanism and introduces a quantity-aware triplet weight adjustment method, so that during the training process, the model's ability to aggregate high-quality candidate images can be enhanced and the response to low-quality samples can be suppressed, thereby significantly improving the performance of the generated images in terms of geometric consistency and texture fidelity.
[0058] The establishment of the perceptual contrast loss function based on the image triplet includes: Construct a contrast loss module and use a pre-trained feature extraction network to perform forward propagation on the input image, positive sample image, and negative sample image to extract the multi-scale semantic feature tensor of the intermediate layer respectively; Normalizing the feature tensor and projecting it into a feature space of the same dimension to construct a comparable embedding representation; Use Euclidean distance, cosine similarity, or perceptual metrics based on feature difference to measure the semantic difference between the input image and positive and negative samples; Establish a perceptual contrast loss function based on image triplets, where the distance between the input image and the positive sample is minimized, and the distance between the input image and the negative sample is maximized. By performing gradient backpropagation through the contrast loss function, the diffusion model is guided to gradually improve the consistency with the positive sample structure during the multi-view image generation process, and suppress the generation of blurred images similar to negative samples.
[0059] The contrast loss function is expressed as follows:
[0060]
[0061] Where a is the reference sample, p is the positive sample, and n is the negative sample.
[0062] In this embodiment, a VGG network or CLIP feature extractor is introduced to calculate the perceptual loss between images. The resulting candidate images are then compared with the input image, with those above a certain threshold being considered positive samples, and those below the threshold being considered negative samples. This optimizes texture consistency during 3D content generation. The pre-trained feature extraction network is preferably a VGG-16, CLIP, ResNet, or other perceptual network trained on large-scale vision tasks. Its output features are rich in texture, edge, and semantic information.
[0063] Furthermore, the construction of the perceptual contrast loss function specifically includes: Based on the constructed image triplet (e_anchor, e_pos, e_neg), the standard triplet loss function L_triplet is used as the main optimization objective, which is defined as follows:
[0064] in, represents the distance metric function between normalized feature vectors (such as Euclidean distance or cosine distance), and m represents the margin threshold, which controls the minimum distinction between positive and negative samples. The typical value is m = 0.2. The goal of this loss term is to minimize the distance between the anchor point and the positive sample, while maximizing the distance between the anchor point and the negative sample.
[0065] To enhance the ability to express structural details in the feature space, the triplet loss can also be integrated with multi-layer embedding distances in actual implementation. For example, the distances of the conv3_3 and conv5_3 layers in the VGG network are calculated separately and weighted summed to form a hierarchical perceptual contrast loss function:
[0066] in, and is the perceptual weight of each layer, and D represents the pixel-level perceptual distance between feature tensors (such as LPIPS or feature Euclidean difference). This strategy enables the model to learn the true differences between images at both the structure and texture levels.
[0067] S5. Establish a quantity-aware triplet loss weight mechanism to dynamically adjust the contribution ratio of positive and negative samples in training; In this embodiment, considering the gradient deviation problem that may be caused by uneven sample quality and distribution during training, the embodiment of the present invention designs a quantity-aware triple loss weighting mechanism. This mechanism dynamically assigns weights based on the structural score differences of each triple:
[0068] Where Sim_anchor_pos and Sim_anchor_neg are the similarity scores between the anchor and the positive and negative samples in the triplet, and β is the temperature control parameter. The weight w is introduced into the total loss to form a weighted loss term:
[0069] This mechanism automatically enhances the training intensity of high-value sample pairs at the tail of the distribution and suppresses the training perturbations of low-confidence sample pairs, thereby improving the stability and convergence speed of the overall contrastive learning.
[0070] Ultimately, all loss terms will be jointly trained with the 3D consistency loss in the neural rendering module (such as image reconstruction loss and forward light consistency) to form a unified multi-objective optimization framework.
[0071] The embodiment of the present invention introduces a quantity-aware triplet loss weight mechanism to solve the gradient bias problem caused by the imbalance of the positive and negative sample ratios. The dynamic adjustment of the contribution ratio of positive and negative samples in training includes: Dynamically count the number of positive and negative samples and their quality scores in each training batch; Assign a confidence-based dynamic weight coefficient to each triplet loss term to increase the training weight of high-quality positive samples and suppress the influence of low-quality negative samples; This weight mechanism is integrated into the overall contrast loss function to achieve adaptive control of sample distribution in the global training process.
[0072] S6. Embed a super-resolution image enhancement module into the candidate image optimization process to perform pixel-level enhancement on the generated candidate images of multiple perspectives; Because pre-trained diffusion models are often limited by video memory and sampling strategies during the generation process, the generated images often suffer from blurred textures, missing details, and limited resolution in the raw output stage. These defects not only affect the ability of contrastive learning to express feature differences, but may also cause the accumulation of ray casting errors in subsequent NeRF training. In this embodiment of the present invention, a super-resolution image enhancement module is introduced between the diffusion output and perceptual contrast modules to improve the spatial detail quality and feature accuracy of the candidate images.
[0073] This module receives a low-resolution image sequence generated by the diffusion model. , input it into the super-resolution network, and output the image set with enhanced resolution This paper recommends using an image enhancement model based on deep residual blocks and attention mechanisms, such as ESRGAN or SwinIR. Its structure includes a residual dense block (RRDB), an upsampling module (PixelShuffle / Deconv), and a feature fusion path.
[0074] The input image is first extracted through shallow convolution to extract low-level semantic features, and then through multi-layer RRDB to extract high-frequency structural details. The output stage is nonlinearly reconstructed to generate the original image. Figure 1 The enhancement module is deployed as an independent processing subsystem, whose parameters can be frozen or jointly optimized with the backbone path during the training phase. The specific strategy is determined by system resources and task complexity.
[0075] The enhanced image retains spatial alignment with the original input image, allowing it to be seamlessly fed into the subsequent VGG / CLIP embedder for perceptual loss calculation. Furthermore, because the super-resolution module enhances image edge and texture contrast, its output features exhibit greater discriminability in the contrast embedding space, helping to construct clearer triplet structures and optimize gradient paths.
[0076] The super-resolution image enhancement module includes: using any image super-resolution model with detail restoration capabilities (such as ESRGAN, Real-ESRGAN, etc.) to enlarge the candidate image; while maintaining the structure unchanged, it enhances the image edge details, texture clarity and resolution, thereby improving the quality of the perceptual loss signal and training stability.
[0077] The specific structure and functions of the super-resolution image enhancement module include: the module is independently deployed between the diffusion model output image and the perceptual loss calculation module, its input is multiple low-resolution candidate images, and its output is an enhanced high-resolution image set; The super-resolution image enhancement module preferably adopts a super-resolution model structure based on a deep neural network, including but not limited to any image enhancement network with detail restoration capabilities such as ESRGAN, SwinIR, EDSR, RCAN, etc.; the enhancement network can be introduced in the form of frozen parameters during the training phase, or it can participate in gradient updates in the joint optimization framework.
[0078] The high-resolution image output by the super-resolution image enhancement module retains the spatial arrangement information and structural features of the original image, and has stronger expressive ability in terms of texture edges and detail levels; the introduction of image enhancement can not only improve the gradient distribution quality of the perceptual contrast loss function, but also enhance the training signal amplitude of the triplet loss, thereby improving the overall three-dimensional modeling accuracy.
[0079] S7, jointly training the overall optimization objectives of perceptual contrast loss and quantity-aware triplet loss with the neural radiation field reconstruction module to generate a 3D mesh or volume rendering result; The neural radiation field reconstruction module is a neural radiation field or other volume rendering network. In order to achieve the final three-dimensional modeling and image consistency reconstruction, the embodiment of the present invention adopts the neural radiation field modeling method (Neural Radiance Fields, NeRF) as the core three-dimensional representation module. The NeRF structure is to transform each coordinate point in the three-dimensional space into a and viewing direction Mapping to color values and bulk density , and then the continuous sample points on the entire ray are integrated to synthesize a two-dimensional image through the volume rendering formula.
[0080] The neural network consists of a multilayer perceptron (MLP) with the structure f_θ(x, d) → (σ, r), where θ is a learnable parameter. The input 3D position x is first converted to a high-dimensional embedding using the positional encoding γ(x), which is then concatenated with the viewing direction before being fed into the network. The output value σ controls the transparency attenuation, and r controls the pixel color.
[0081] This module samples the camera ray direction under the perspective of the input candidate image and samples N points in three-dimensional space for each ray. , get its color and density value, and then use the volume rendering formula:
[0082] Where T_i represents the forward cumulative transparency, δ_i is the distance between each point, and σ_i and r_i are the neural network predictions. This formula outputs a rendered image for comparison with the input image to perform pixel-level or perceptual reconstruction loss calculations.
[0083] To enhance the realism and multi-view consistency of three-dimensional structures, an embodiment of the present invention jointly trains the neural rendering module with the aforementioned perceptual contrast module, introduces spatial consistency loss, depth regularization term and reconstruction constraint term into the loss function, and forms a multi-view constraint space through perspective perturbation, so that the NeRF network can learn a consistent and continuous volume density field.
[0084] The joint training of the overall optimization goal of perceptual contrast loss and quantity-aware triplet loss with the neural radiation field reconstruction module includes: Receive the input image and its generated multi-view image set, and extract dense view information as conditional input; A mapping function between color field and density field is established in volume rendering space, and consistent image reconstruction output is synthesized through volume rendering integral. The volume rendering network parameters are jointly optimized based on the three-dimensional spatial consistency loss and the above-mentioned contrast loss to obtain high-quality three-dimensional model output with geometric coherence and texture consistency.
[0085] It should be noted that the embodiment of the present invention adopts an end-to-end unified training framework to integrate multi-view image generation, image enhancement, contrast loss learning and neural rendering optimization into an integrated system. During the training process, the forward path sequentially calls the diffusion module → posture perturbation → super-resolution enhancement → embedding coding → positive and negative sample construction → contrast loss calculation → rendering module optimization to form a closed-loop structure. In the back propagation path, the contrast loss gradient can act on the output direction of the diffusion model, guiding it to gradually converge to the structural optimal image; and the rendering module loss affects the volume sampling and color prediction network at the same time, thereby achieving the perception-rendering joint modeling goal.
[0086] The system supports a batch-based triplet training scheduling mechanism, dynamically sampling anchor images and positive and negative samples within each batch to update the parameters θ_diffusion, θ_nerf, and θ_superres. When hardware conditions permit, NeRF and diffusion modules can be updated asynchronously in parallel to accelerate training efficiency. After training is complete, the model can be independently deployed, automatically completing high-quality 3D modeling with just a single RGB image as input.
[0087] S8. Output the reconstructed three-dimensional representation result.
[0088] The results include 3D visualization assets in the form of 3D mesh models, volume density fields or image sequences derived from the neural rendering model, which can be used for digital human modeling, 3D asset production and virtual reality systems.
[0089] In the final stage of this invention, the 3D content, which has undergone perceptual optimization and rendering reconstruction, is exported into a standard format that can be used by third-party programs. Considering the diversity of downstream applications, the following output content types are provided for developers or platforms to choose from: 3D implicit model file based on neural radiation field representation, saved in .pth or .pkl format, supports subsequent loading and rendering; Multi-view RGB image sequences synthesized based on volume rendering can be used to generate 3D turntable animations or construct texture datasets; Explicit 3D mesh files exported by voxel extraction algorithms, including standard formats such as .obj and .ply, can be loaded in 3D modeling platforms such as Blender and Unity; 3D density or color volume tensor, saved in .npz or .npy format, suitable for applications such as scientific visualization, volume data processing, and medical image modeling.
[0090] The contrast-enhanced 3D modeling method proposed in this embodiment has been tested on multiple public datasets (such as Shapenet, DTU, and RealEstate10K). Compared with existing typical methods (such as DreamFusion, Magic3D, and SparseNeRF), it achieves significant improvements in FID score, LPIPS indicator, PSNR reconstruction accuracy, and structural consistency score. The system is highly robust to input images, supports different resolutions and complex backgrounds, does not rely on depth maps or multi-frame data, and can also be deployed on mobile devices.
[0091] In summary, the single-view image 3D modeling method based on contrastive learning proposed in the embodiment of the present invention achieves a modeling effect with structural coherence, rich texture, and consistent 3D without multi-view supervision, and has broad industrial application potential.
[0092] In addition, an embodiment of the present invention provides a single-view image 3D modeling system based on contrastive learning, which is used to automatically construct a high-quality 3D model from a single input image. Figure 2 As shown, the system includes: An input module, used to obtain a single-view image as an initial input for the 3D modeling process; A diffusion model generation module is used to generate multiple candidate images with varying viewing angles based on a single input RGB image. The diffusion model generation module includes: A pre-trained 2D image diffusion model, which uses the Score Distillation Sampling mechanism to guide image generation. After inputting the single-view image, the model can sample multiple image samples with semantic consistency and perspective differences under specific noise guidance conditions; A candidate image sampling control unit controls the number of generated images, perspective difference, initial random noise distribution and number of denoising steps, and performs structural alignment and preprocessing on the resulting images.
[0093] Specifically, the pre-trained two-dimensional diffusion model in the diffusion model generation module adopts StableDiffusion, DeepFloyd IF, or any diffusion-type architecture with controllable text / image conditional input, and its specific structure includes: A U-Net type backbone neural network structure is used to conditionally denoise the image representation in the noise space and restore the image content; a conditional encoder that receives an input image and encodes it into a latent semantic guidance vector, which is used to control the direction of generated content during the diffusion process so that the generated multi-view images maintain semantic consistency; A scheduler and sampler module, which combines the diffusion step setting with the denoising sampling rule (such as DDIM and PLMS), performs multiple samplings under a specific guidance intensity to generate a set of candidate images with different perspectives but coherent structures; This module supports the input of external view control signals and enhances the geometric diversity and modeling space coverage of generated images by introducing camera pose encoding or implicit pose perturbations.
[0094] A perceptual contrast learning module is used to construct positive and negative sample pairs from the candidate images generated by diffusion and perform feature-level contrast training. The perceptual contrast learning module includes: An image embedding feature extraction submodule that uses VGG, CLIP, or other pre-trained image networks to extract embedded representations of the image at multiple intermediate layers, which are used for perceptual feature alignment and distance calculation; A triplet sample generation submodule automatically selects positive and negative sample images to form image triplets by calculating the texture similarity and structural consistency metrics between all candidate images and the input image; A contrastive loss construction and update unit embeds the above features into the input perception loss function and the triplet loss function to form a supervision signal with the goal of structural fidelity and semantic consistency; A sample quantity-aware regulation submodule calculates dynamic weights based on the ratio of positive and negative samples and the training confidence, and integrates them into the contrastive loss optimization path to achieve gradient stabilization and adaptive regulation of sample contribution.
[0095] The image embedding feature extraction submodule in the perceptual contrast learning module specifically includes: A convolutional neural network based on the VGG-16 or ResNet-50 architecture that outputs intermediate semantic features at multiple levels for contrastive loss calculation; A normalization processor that performs channel normalization and spatial average pooling on the extracted feature tensor to ensure a stable measurement basis for feature similarity between positive and negative samples; A projection transformation submodule maps each feature tensor to a feature embedding space of uniform dimension, making the features extracted under different network structures comparable; An optional feature selection mechanism performs weighted fusion of multi-scale semantic features to increase the dominant weight of structural edges and texture details in the perceptual loss, thereby improving the discriminative sensitivity of the triplet loss and the effectiveness of the training signal.
[0096] A super-resolution image enhancement module is used to restore the diffusion-generated image at a high-quality detail level. The super-resolution image enhancement module includes: A super-resolution neural network submodule that performs super-resolution upscaling of images based on ESRGAN, SwinIR, or other deep models with detail enhancement capabilities, restoring edge structures and texture details of candidate images; An image enhancement interpolation controller that sets the input image scaling factor, target resolution, and output size to ensure that the enhanced image remains semantically and spatially aligned with the input image; A neural rendering and reconstruction module is used to fuse the perceptual loss signal with the multi-view image to output a 3D structure. The neural rendering and reconstruction module includes: A neural radiation field modeling submodule that constructs color distribution functions and density distribution functions in space and reverse maps the projection of multi-view images into a three-dimensional coordinate space; A volume rendering integrator performs color fusion on continuous view sampling results based on the volume rendering integration method, and outputs a volume image or mesh representation with 3D structural consistency and texture visualization consistency; A joint loss optimization scheduling module is used to coordinate the parameter update paths of the neural rendering module and the perceptual contrast loss module to achieve multi-objective collaborative optimization of the overall training process.
[0097] The neural radiation field modeling submodule in the neural rendering and reconstruction module includes: A multi-layer perceptron (MLP) structure is used to encode the 3D coordinate position and camera direction and map them into continuous function outputs of color and volume density; A position encoder that performs frequency-upscale transformation on the input 3D coordinates to improve the model's ability to express high-frequency geometric structures; A camera ray manager that receives the view and direction information of each pixel in the candidate image and back-projects it to generate a set of spatial rays for sampling; A layered volume renderer that performs ray integration on sample points based on neural volume rendering theory to synthesize new perspective images; A spatial consistency constraint module compares the synthesized image with the actual candidate image and Figure 1 The consistency loss function and the perceptual contrast loss function jointly optimize the parameters in this module, so that the output three-dimensional model achieves high fidelity in both spatial structure and texture rendering.
[0098] The output module is used to output the final 3D modeling results. The results include: a 3D mesh generated by the neural rendering model, a volume field density representation, or a view-consistent image set in the form of an image sequence. The output module can also be exported to a standard 3D model format for subsequent editing, simulation, or visualization.
[0099] In summary, the system has the following advantages: It has low input dependency and only requires a single image to drive 3D modeling, improving acquisition efficiency.
[0100] The contrast loss mechanism is used to finely screen generated samples, significantly improving texture fidelity.
[0101] The super-resolution enhancement module improves the detail blurring problem in diffusion generation and enhances the quality of contrast features.
[0102] The rendering module and the perception module are jointly trained to optimize spatial consistency and realism.
[0103] The module structure is highly decoupled and has good scalability and replaceability. It can be used for text-driven modeling, graphic and text hybrid modeling, video 3D modeling and many other tasks.
[0104] In another embodiment of the present invention, an electronic device may include a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may invoke logic instructions in the memory to execute a single-view image 3D modeling method based on contrastive learning.
[0105] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0106] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
[0107] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A single-view image 3D modeling method based on contrastive learning, characterized in that: The following steps are involved: Obtaining a single-view image as the initial input for the 3D modeling process; Feed a single-view image into a pre-trained 2D image diffusion model to generate candidate images from multiple perspectives. Based on the obtained multiple candidate images, positive and negative samples are constructed to form an image triplet structure; Establish a perceptual contrast loss function based on image triplets, calculate the perceptual loss between images, and perform contrast optimization at the perceptual feature level; Establish a quantity-aware triplet loss weight mechanism to dynamically adjust the contribution ratio of positive and negative samples in training; The super-resolution image enhancement module is embedded in the candidate image optimization process to perform pixel-level enhancement on the generated candidate images of multiple perspectives. The overall optimization objectives of perceptual contrast loss and quantity-aware triplet loss are jointly trained with the neural radiation field reconstruction module to generate three-dimensional meshes or volume rendering results, and output the reconstructed three-dimensional representation results.
2. The single-view image 3D modeling method based on contrastive learning according to claim 1, characterized in that: The single-view image is a static image in three-channel RGB format. Starting from a single input image, a pre-trained two-dimensional diffusion model combined with score distillation sampling is used to generate multi-view candidate images.
3. The single-view image 3D modeling method based on contrastive learning according to claim 1, characterized in that: The constructing of positive and negative samples based on the obtained multiple candidate images includes: The candidate images are combined in pairs, and the geometric structure consistency and texture detail similarity between the images are evaluated based on the predefined structural similarity measurement function and texture distribution index; Images whose structure and texture are highly similar to the input image are selected as positive samples, and images whose geometric structure or texture distribution deviates significantly are selected as negative samples to form an image triplet structure.
4. The single-view image 3D modeling method based on contrastive learning according to claim 1, characterized in that: The process of establishing a perceptual contrast loss function based on image triplets, calculating inter-image perceptual loss, and performing contrast optimization at the perceptual feature level includes: A pre-trained feature extraction network is used to perform forward propagation on the input image, positive sample image, and negative sample image to extract the multi-scale semantic feature tensor of the intermediate layer respectively; Normalize the feature tensor and project it into a feature space of the same dimension to construct a comparable embedding representation; The semantic difference between the input image and the positive and negative samples is measured using Euclidean distance, cosine similarity, or a perceptual metric based on feature difference, and a perceptual contrast loss function based on image triplets is established. The distance between the input image and the positive sample is minimized, while the distance between the input image and the negative sample is maximized. By performing gradient backpropagation on the loss function, the diffusion model is guided to gradually improve the consistency with the positive sample structure during the multi-view image generation process, and suppress the generation of blurred images similar to negative samples.
5. The single-view image 3D modeling method based on contrastive learning according to claim 1, characterized in that: The establishment of a quantity-aware triplet loss weight mechanism to dynamically adjust the contribution ratio of positive and negative samples in training includes: Dynamically count the number of positive and negative samples and their quality scores in each training batch; Assign a confidence-based dynamic weight coefficient to each triplet loss term to increase the training weight of high-quality positive samples and suppress the influence of low-quality negative samples; The weight mechanism is integrated into the overall contrast loss function to achieve adaptive control of sample distribution in the global training process.
6. The single-view image 3D modeling method based on contrastive learning according to claim 1, characterized in that: The method of embedding the super-resolution image enhancement module into the candidate image optimization process and performing pixel-level enhancement on the generated candidate images of multiple perspectives includes: An image super-resolution model with detail restoration capability is used to enlarge the candidate image. While keeping the structure unchanged, it enhances image edge details, texture clarity and resolution, and improves the quality of perceptual loss signals and training stability.
7. The single-view image 3D modeling method based on contrastive learning according to claim 6, characterized in that: The super-resolution image enhancement module adopts a super-resolution model structure based on a deep neural network, including an image enhancement network with detail restoration capabilities; the super-resolution image enhancement module is deployed between the diffusion model output image and the perceptual loss calculation module, with input as multiple low-resolution candidate images and output as an enhanced high-resolution image set.
8. The single-view image 3D modeling method based on contrastive learning according to claim 1, characterized in that: The neural radiation field reconstruction module is a neural radiation field or volume rendering network, and the overall optimization objectives of perceptual contrast loss and quantity-perceptual triplet loss are jointly trained with the neural radiation field reconstruction module, including: Receive the input image and its generated multi-view image set, and extract dense view information as conditional input; A mapping function between color field and density field is established in volume rendering space, and consistent image reconstruction output is synthesized through volume rendering integral. The volume rendering network parameters are jointly optimized based on the three-dimensional spatial consistency loss and the above-mentioned contrast loss to obtain high-quality three-dimensional model output with geometric coherence and texture consistency.
9. A single-view image 3D modeling system based on contrastive learning, used to implement the single-view image 3D modeling method based on contrastive learning according to any one of claims 1 to 8, characterized in that: The system comprises: An input module, used to obtain a single-view image as an initial input for the 3D modeling process; Diffusion model generation module, used to generate multiple candidate images with different perspectives based on a single input RGB image; A perceptual contrastive learning module that constructs positive and negative sample pairs from candidate images generated by diffusion and performs feature-level contrastive training; Super-resolution image enhancement module, used to restore the diffusion-generated image at a high-quality detail level; Neural rendering and reconstruction module, which is used to fuse the perceptual loss signal with the multi-view image to output a 3D structure; Output module, used to output the final 3D modeling results.
10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the processor is configured to implement a single-view image three-dimensional modeling method based on contrastive learning as described in any one of claims 1 to 8 when executing a computer program stored in the memory.
Citation Information
Cited By
Performance index-based three-dimensional automobile model generation network, generation method and system
CN121052146A
Performance-based 3D car model generation network, generation method and system
CN121052146B
Three-dimensional texture super-resolution method and system based on multi-view fusion
CN122048672A
Small-sample target remote sensing image generation method based on structure perception and detail enhancement
CN122133737A
Few-shot Target Remote Sensing Image Generation Method Based on Structure Awareness and Detail Enhancement
CN122133737B