Image super-resolution method for three-dimensional reconstruction of automobile parts
Through the cross-view angle residual synergistic enhancement and multi-view angle consistency optimization module, combined with the image diffusion model and neural radiation field method, the three-dimensional reconstruction inconsistency problem of low-resolution automotive parts images is solved, high-resolution detail recovery and structural consistency are achieved, and the three-dimensional reconstruction accuracy and stability are improved.
Patent Information
- Application Number
- CN202510846906.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Traditional three-dimensional reconstruction technology is difficult to keep the image details consistent with the structure in low-resolution or blurred texture images of automobile parts, which affects the accuracy and stability of three-dimensional reconstruction, especially for complex parts such as engine housing and gears.
The cross-view angle residual synergistic enhancement mechanism and multi-view angle consistency optimization module are adopted, combined with the image diffusion model and neural radiation field method, and image super-resolution reconstruction is carried out through the multi-view angle consistency optimization module and the NeRF module to improve image consistency and clarity.
It significantly improves the three-dimensional reconstruction accuracy and stability of low-resolution automotive component images, achieves high-resolution details recovery and structural consistency, and supports subsequent detection and modeling tasks.
Smart Images

Figure CN120355579A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image enhancement, and particularly relates to an image super-resolution method for three-dimensional reconstruction of automotive parts. Background Art
[0002] With the development of the automotive industry towards intelligent manufacturing and high-precision inspection, higher requirements are put forward for the three-dimensional reconstruction and surface defect recognition of automotive parts; high-quality three-dimensional reconstruction results not only help to accurately restore the morphology of parts, but also provide key support for subsequent defect analysis; traditional three-dimensional reconstruction techniques often rely on multi-view images as input, and their reconstruction accuracy is limited by the resolution and consistency of the original images; in industrial environments, due to factors such as limited resolution of imaging devices, limited acquisition angles, and strong reflection on the surface of parts, the acquired images have problems such as insufficient clarity, inconsistent cross-view structures, and blurred textures, seriously affecting the accuracy and stability of three-dimensional reconstruction. Especially for parts with rich details or complex curved surfaces such as engine casings, gears, fastening connectors, etc., traditional methods are difficult to restore their complete spatial structure and surface texture.
[0003] In recent years, image super-resolution and neural rendering technologies have developed rapidly, providing new solutions for industrial vision inspection; image diffusion models perform excellently in detail restoration, while neural radiance field methods can construct three-dimensionally consistent implicit representations from multi-angle images, generating structurally continuous image views in three-dimensional space and showing good structural consistency; however, traditional neural radiance field methods usually rely on high-quality input images and are difficult to maintain image details and structural consistency in low-resolution or texture-blurred scenarios, limiting their applicability in automotive part inspection; to solve the above problems, the present invention proposes an image enhancement method for automotive parts that combines multi-view consistent diffusion super-resolution and three-dimensional modeling, which combines the detail generation ability of the diffusion model and the structural modeling ability of the neural radiance field method, and realizes multi-view image consistency super-resolution for three-dimensional reconstruction of automotive parts by introducing a cross-view residual collaborative enhancement mechanism and a multi-view consistency optimization module. Summary of the Invention
[0004] The present invention provides an image super-resolution method for three-dimensional reconstruction of automotive parts, aiming to improve the consistency and clarity of multi-angle automotive part images and three-dimensional modeling under the condition of low-resolution picture input by introducing a cross-view residual collaborative enhancement mechanism and a multi-view consistency optimization module, including the following steps.
[0005] S1. Use an industrial camera group installed on a bracket array to capture multi-angle low-resolution images of automotive parts; at the same time, calibrate each camera with the help of a calibration board to obtain the camera pose information corresponding to each image.
[0006] S2. Construct a cross-view collaborative enhancement module CVCEM, including linear projection, calculating attention correlation scores, residual fusion, and enhancing features. This module focuses on using adjacent views to enhance the latent variable representation of each view.
[0007] S3. Construct a multi-view consistency module MVCM, which is divided into four parts: sparse loss calculation, diffusion loss calculation, MultiViewUNet network, and joint optimization and parameter update of the total loss. Sparse loss calculation is used for residual sparsification to reduce redundant noise. Diffusion loss calculation constructs a reconstruction loss through the forward diffusion and reverse noise prediction processes. The MultiViewUNet network is used for noise prediction. Joint optimization and parameter update of the total loss is used for jointly optimizing the parameters of the entire module. This module focuses on updating the residuals while ensuring the consistency of the enhanced latent vector representations under multiple views.
[0008] S4. Construct a multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF, which is jointly composed of a multi-view consistent diffusion super-resolution model MVCDSR based on CVCEM and MVCM and a NeRF reconstruction module.
[0009] Preferably, in S1, first install multiple industrial cameras on a preset automotive parts shooting platform to form a camera bracket array at fixed angles. The industrial cameras are arranged around the automotive parts to be detected, covering multiple views of its typical structural surfaces. In the image acquisition stage, by synchronously triggering the camera group, a set of low-resolution images of the automotive parts at different angles is obtained. To ensure geometric consistency in the subsequent image enhancement and 3D modeling processes, use a calibration board for camera calibration operations, and use a multi-camera calibration tool to solve the internal and external parameters of each camera to obtain a set of camera pose parameters corresponding to the images. .
[0010] Preferably, in S2, when constructing the cross-view collaborative enhancement module CVCEM, the specific process includes: the input is the set of residuals corresponding to all views and a set of latent variables , where each corresponds to the latent vector representation obtained by encoding the low-resolution image of the th view through a variational autoencoder. For each input latent feature , this module updates the features through a cross-view residual enhancement mechanism. The specific process is as follows: First, and its neighbor view set are respectively transformed through learnable linear mapping matrices and to obtain query vectors And the key vector , calculate the cross-view attention weights based on the dot product similarity between the query vector and each neighbor key vector , Subsequently, according to the above attention weights, calculate the total residual of the current view. The specific process is as follows: , where is the residual of view , is the residual of the current view , is the set self-residual fusion coefficient. Finally, update the latent vector representation of the current view . The update process is , is a set of latent vector representations after enhancing the features and is the output of this module
[0011] Preferably, in S2, a strategy for enhancing the latent features of images based on cross-view residual guidance and attention fusion is proposed to improve the structural consistency and information integrity of multi-angle images of automotive parts in the latent representation stage; the representations of the query vector and the key vector are realized by constructing a learnable linear mapping matrix, and dynamic weights are generated under the dot product attention framework to realize the selective fusion of the residual information of each neighbor view; on this basis, the latent representation of each view is superimposed and updated with the fused residual information to obtain a more complete structure and richer features enhanced latent representation; while maintaining the semantic features of the original view, this module dynamically integrates the supplementary information from adjacent views, which can effectively alleviate the problem of image information loss caused by shooting angle deviation, reflection interference or low-resolution acquisition, and provides a stable and continuous expression basis for subsequent consistency optimization and image super-resolution reconstruction
[0012] Preferably, in S3, the specific process of constructing the multi-view consistency module MVCM is as follows: S31. The multi-view consistency module MVCM is divided into four parts: sparsity loss calculation, diffusion loss calculation, constructing the MultiViewUNet network, and joint optimization and parameter update of the total loss. For the sparsity loss calculation part, the input is a set of residuals of views , and the specific process is as follows: for each residual variable , calculate the residual sparsity loss corresponding to this view as , where , , correspond to the number of channels, height, and width of the residual features respectively; S32. For the part of constructing the MultiViewUNet network, which is used to predict the noise , the input of this network is the current view latent vector after adding noise , the set of potential vectors of adjacent viewpoints and the set of poses of adjacent viewpoints , the specific process is as follows: First, and are respectively flattened to obtain pixel sequence representations and , where , represents the height and width of the feature map, represents the number of channels, represents the flattening operation; Secondly, the pose is broadcast to each pixel to form a pose feature matrix , where represents the embedding encoding of the pose , represents the broadcast operation, is a unit vector with a length of ; Then the neighbor-added noise feature and the pose embedding are added element-wise to obtain the fused key vector , and the fused key vectors of all neighbor viewpoints are stacked along the pixel dimension to form a complete multi-view key set , where, is the number of neighbor viewpoints; Based on the cross-attention mechanism, the dot product relationship between the query vector of the current viewpoint and the neighbor-fused key vector is used to obtain the fused conditional feature as the final conditional input of the UNet, where function represents the normalization operation, ; Finally, the noisy latent vector of the current viewpoint, the time step , and the fused conditional feature are input into the UNet backbone network to generate the final noise prediction output ; S33. For the diffusion loss calculation part, the input is a set of enhanced latent vectors and a set of corresponding pose information . For each , forward diffusion is performed through to obtain the noisy vector and , is the cumulative noise attenuation coefficient corresponding to the th time step, which is generated by the diffusion scheduler, and the calculation formula is , where is the preset noise addition ratio per step in the diffusion schedule, represents standard Gaussian noise; Use the MultiViewUNet network to predict the noise ; Based on the noise prediction result , calculate the forward latent vector , and the calculation method is , and finally define the diffusion reconstruction loss as ; S34. For the total loss joint optimization and parameter update part, the input is a set of diffusion reconstruction losses and a set of sparsity losses . First, calculate the total loss of the current view as , where is a preset hyperparameter. Second, calculate the total loss of all views as . Third, update the residual based on the gradient descent of the total loss of the current view. The specific process is: , where is the preset learning rate, represents the gradient of the residual variable of the current view. Fourth, update the learnable linear mapping matrices and based on the gradient descent of the total loss of all views. The specific process is: , , where and are the gradients of the linear mapping matrices and under the overall loss.
[0013] Preferably, in S3, a multi-view consistency module MVCM is constructed. This module focuses on guiding the consistent expression between different views during the latent feature enhancement stage and is composed of four parts: sparsity loss calculation, diffusion reconstruction loss construction, MultiViewUNet network design, and total loss joint optimization. First, a sparse regularization term is constructed in the sparsity loss calculation part to encourage the network to generate sparse and clearly structured residual information and reduce redundant interference. Second, the MultiViewUNet network constructed based on the noisy latent vector introduces adjacent view images and pose information and combines the cross-attention mechanism for feature fusion to achieve high-precision prediction of noise. Third, through the forward diffusion and reverse reconstruction processes, a diffusion reconstruction loss in the latent space is constructed to effectively evaluate the consistency performance of the image during the progressive denoising process. Finally, the sparsity loss and the diffusion loss are jointly optimized, and the residual variable and the learnable attention mapping parameters are updated in a gradient descent manner to achieve the enhancement of consistency and the adaptive adjustment of residuals between multiple views. This module significantly improves the stability and coordination of the latent expression under multiple views and provides a high-quality feature basis for structural alignment in the subsequent process.
[0014] Preferably, in S4, a multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF is constructed. The specific process is: S41. The multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF includes an MVCDSR module and a NeRF reconstruction module. For the MVCDSR module, the input is a set of low-resolution images and their corresponding camera pose parameters . After , the latent vector representation of the image is obtained , representing the encoding process of the variational autoencoder; the latent vector representation is input into the cross-view collaborative enhancement module CVCEM module to obtain enhanced features ; subsequently, the enhanced features are input into the multi-view consistency module MVCM module to update the residuals and optimize the latent representation; the CVCEM and MVCM are executed Z times in a joint iterative manner to output the finally enhanced latent vector , and then through the variational autoencoder, after decoding, where represents the decoding process, is a set of output high-resolution multi-view images; S42. For the NeRF reconstruction module, the high-resolution multi-view images and their corresponding camera pose parameters are input into the neural radiance field model for three-dimensional reconstruction, and the reconstructed output image is used as the input for the next round of the MVCDSR module. The MVCDSR module and the NeRF reconstruction module are executed M times in a joint iterative manner, and finally an image representation with high-resolution details and a three-dimensional consistent structure and its three-dimensional model are output.
[0015] Preferably, in S4, by inputting multi-angle high-resolution images and their camera poses into the neural radiance field model, three-dimensional modeling and consistent view reconstruction of the surface of automotive parts are realized, and this is used as the input for a new round of image super-resolution enhancement to continuously optimize the image expression and geometric consistency, and finally an image representation with high-resolution details and a three-dimensional consistent structure and its three-dimensional model are output.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: It combines the advantages of the image diffusion model in detail enhancement and the capabilities of the neural radiance field in three-dimensional structure modeling. Aiming at the pain points of low-resolution and multi-angle images in terms of structural inconsistency and detail blurring, through the cross-view residual collaborative enhancement and multi-view consistency optimization strategies, it realizes the fine reconstruction and structural alignment of the latent expression of the image; at the same time, it introduces the NeRF module as a three-dimensional geometric constraint to achieve the joint iterative optimization of image enhancement and three-dimensional modeling; this method has significant advantages such as strong view consistency, high detail recovery accuracy, and stable and natural enhancement process, and can significantly improve the usability and accuracy in the three-dimensional reconstruction of low-quality automotive part images and subsequent detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 FIG. is a flowchart of an image super-resolution method for three-dimensional reconstruction of automotive parts provided by the present invention.
[0018] Figure 2 FIG. is a structural diagram of the cross-view collaborative enhancement module CVCEM provided by the present invention.
[0019] Figure 3 FIG. is a structural diagram of the multi-view consistency module MVCM provided by the present invention.
[0020] Figure 4 FIG. is a structural diagram of the MultiViewUNet network provided by the present invention.
[0021] Figure 5 FIG. is a structural diagram of the MVCDSR module provided by the present invention.
[0022] Figure 6 FIG. is a structural diagram of the multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] Please refer to Figures 1 to 6 , the present invention provides an image super-resolution method for three-dimensional reconstruction of automotive parts, aiming to generate high-resolution images and three-dimensional representations with clear details and structural consistency based on multi-view low-resolution images and camera pose information, so as to improve the detection accuracy of automotive parts and the ability to restore the spatial structure.
[0025] S1. Use an industrial camera group installed on a bracket array to capture multi-angle low-resolution images of automotive parts; meanwhile, calibrate each camera with the help of a calibration board to obtain the camera pose information corresponding to each image.
[0026] Further, in S1, select 8 industrial cameras to form a circular bracket array and install it around the detection station. The brackets are distributed in a hemispherical shape. Each camera is arranged around the part at a horizontal angle of 45° at equal intervals and is divided into two layers in the vertical direction; before shooting, place a planar calibration board with precise marking points at multiple spatial positions of the detection station in sequence, and use each camera to collect images of the calibration board. Calculate the internal parameter matrix and external parameter matrix of each camera through industrial calibration software to obtain the pose information corresponding to each image; in actual acquisition, fix the target automotive part on the platform, and through the multi-camera synchronous shooting method, collect the low-resolution images corresponding to each perspective. Each automotive part obtains 8 images and their corresponding 8 groups of camera pose information.
[0027] S2. Construct a cross-view collaborative enhancement module CVCEM, including linear projection, calculating attention correlation scores, residual fusion, and enhancing features. This module focuses on using adjacent views to enhance the latent variable expression of each view.
[0028] Further, in S2, construct a cross-view collaborative enhancement module CVCEM. The specific process includes: the input is the set of residuals corresponding to all views and a set of latent variables , each corresponding to the latent vector expression obtained by encoding the low-resolution image of the th view through a variational autoencoder. In this embodiment, the number of views is 8, the initial value of is 0. For each input latent feature , this module updates the features through a cross-view residual enhancement mechanism. The specific process is as follows: First, is transformed with its neighbor view set respectively through learnable linear mapping matrices and to obtain a query vector and a key vector . In this embodiment, the neighbor view is a set composed of the numbers of 4 views that are spatially adjacent to the view ; based on the dot product similarity between the query vector and each neighbor key vector, calculate the cross-view attention weight . Subsequently, according to the above attention weight, calculate the total residual of the current view. The specific process is: , where For the perspective is the residual, and is the residual of the current perspective, is the set self - residual fusion coefficient. In this embodiment, is set to 0.3. Finally, update the latent vector representation of the current perspective The update process is , is a set of latent vector representations after enhanced features and is the output of this module.
[0029] S3. Construct a multi - view consistency module MVCM, which is divided into four parts: sparsity loss calculation, diffusion loss calculation, MultiViewUNet network, and joint optimization and parameter update of the total loss. The sparsity loss calculation is used for residual sparsification to reduce redundant noise. The diffusion loss calculation constructs a reconstruction loss through the forward diffusion and reverse noise prediction processes. The MultiViewUNet network is used for noise prediction. The joint optimization and parameter update of the total loss is used to jointly optimize the parameters of the entire module. This module focuses on updating the residual while ensuring the consistency of the enhanced latent vector representations under multiple views.
[0030] Furthermore, in S3, construct a multi - view consistency module MVCM, and the specific process includes.
[0031] S31. The multi - view consistency module MVCM is divided into four parts: sparsity loss calculation, diffusion loss calculation, constructing the MultiViewUNet network, and joint optimization and parameter update of the total loss. For the sparsity loss calculation part, the input is a set of residuals of perspectives , and the specific process is: for each residual variable , calculate the residual sparsity loss corresponding to this perspective as , where , , correspond to the number of channels, height, and width of the residual features respectively. In this embodiment, set the number of channels , and the feature map size .
[0032] S32. For the part of constructing the MultiViewUNet network, which is used to predict noise , the input of this network is the current perspective latent vector after adding noise , the set of adjacent perspective latent vectors and the set of adjacent perspective poses . The specific process is: first, flatten and respectively to obtain pixel sequence representations and , where , represents the height and width of the feature map, represents the number of channels, represents the flattening operation; secondly, the pose is broadcast to each pixel to form a pose feature matrix , where represents the pose embedding encoding of, represents the broadcast operation, is a unit vector of length ; then the neighbor-added noise feature and the pose embedding are added element-wise to obtain the fused key vector , and the fused key vectors from all neighbor perspectives are stacked along the pixel dimension to form a complete multi-view key set , where, is the number of neighbor perspectives. In this embodiment, ; based on the cross-attention mechanism, using the dot product relationship between the query vector of the current perspective and the neighbor-fused key vector , the fused conditional feature is obtained as the final conditional input to the UNet, where function represents the normalization operation, ; finally, the noisy latent vector of the current perspective, the time step , and the fused conditional feature are input into the UNet backbone network to generate the final noise prediction output . In this embodiment, the total number of steps in the diffusion process is set to , and in each training iteration, a random time step is sampled to control the current diffusion strength.
[0033] S33. For the diffusion loss calculation part, the input is a set of enhanced latent vectors and a set of corresponding pose information . For each , through forward diffusion is performed to obtain the noisy vector and , is the cumulative noise decay coefficient corresponding to the th time step, generated by the diffusion scheduler, and the calculation formula is , where is the preset noise addition ratio per step in the diffusion schedule, represents standard Gaussian noise. In this embodiment, increases linearly with ; Predict noise using the MultiViewUNet network ; Based on the noise prediction result , calculate the forward latent vector , and the calculation method is , and finally define the diffusion reconstruction loss as .
[0034] S34. For the total loss joint optimization and parameter update part, the input is a set of diffusion reconstruction losses and a set of sparsity losses . First, calculate the total loss of the current view as , where is a preset hyperparameter. In this embodiment, Second, calculate the total loss of all views as . Third, update the residual based on the gradient descent of the total loss of the current view. The specific process is: , where is the preset learning rate. In this embodiment, , represents the gradient of the residual variable of the current view. Fourth, update the learnable linear mapping matrices and based on the gradient descent of the total loss of all views. The specific process is: , , where and are the gradients of the linear mapping matrices and under the overall loss.
[0035] S4. Construct a multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF, which is jointly composed of a multi-view consistent diffusion super-resolution model MVCDSR based on CVCEM and MVCM and a NeRF reconstruction module.
[0036] Furthermore, in S4, constructing the multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF, the specific process includes.
[0037] S41. The multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF includes an MVCDSR module and a NeRF reconstruction module. For the MVCDSR module, the input is a set of low-resolution images and their corresponding camera pose parameters . After , obtain the latent vector expression of the picture , represents the variational autoencoder encoding process; the latent vector expression Input into the cross-view collaborative enhancement module CVCEM module to obtain enhanced features ; Subsequently, input the enhanced features into the multi-view consistency module MVCM module to update the residuals and optimize the latent representation; Subsequently, input the enhanced features into the multi-view consistency module MVCM module to update the residuals and optimize the latent representation; The CVCEM and MVCM are executed Z times in a joint iterative manner to output the final enhanced latent vector , and then pass through the variational autoencoder decoding, where represents the decoding process, is a set of output images with a resolution of , in this embodiment, .
[0038] S42. For the NeRF reconstruction module, input the high-resolution image and its corresponding camera pose parameters into the neural radiance field model for 3D reconstruction. The neural radiance field model is trained for 5000 steps using dense sampling and ray integration techniques, sampling 1024 rays per step, and obtaining the reconstructed output image as the input for the next round of the MVCDSR module. The MVCDSR module and the NeRF reconstruction module are executed M times in a joint iterative manner, and finally output an image representation with high-resolution details and a 3D consistent structure and its 3D model. In this embodiment, .
[0039] Furthermore, the image super-resolution model for 3D reconstruction of automotive parts is coded using the Pycharm application and the Python language, and the Pytorch framework is used. The input resolution of the model is 8-view automotive part images with a resolution of 128×128×3.
[0040] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. An image super-resolution method for three-dimensional reconstruction of automotive parts, characterized in that, It includes the following steps: S1. Use an industrial camera group installed on a bracket array to capture multi-angle low-resolution images of automotive parts; meanwhile, calibrate each camera with the help of a calibration board to obtain the camera pose information corresponding to each image; S2. Construct a cross-view collaborative enhancement module CVCEM, including linear projection, calculating attention correlation scores, residual fusion, and enhancing features. This module focuses on using adjacent views to enhance the expression of latent variables for each view; S3. Construct a multi-view consistency module MVCM, which is divided into four parts: sparse loss calculation, diffusion loss calculation, MultiViewUNet network, and joint optimization and parameter update of the total loss. Sparse loss calculation is used for residual sparsification to reduce redundant noise. Diffusion loss calculation constructs a reconstruction loss through the forward diffusion and reverse noise prediction processes. The MultiViewUNet network is used for noise prediction. Joint optimization and parameter update of the total loss are used to jointly optimize the parameters of the entire module. This module focuses on updating the residuals while ensuring the consistency of the enhanced latent vector expressions under multiple views; S4. Construct a multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF, which is jointly composed of a multi-view consistent diffusion super-resolution model MVCDSR based on CVCEM and MVCM and a NeRF reconstruction module.
2. The image super-resolution method for three-dimensional reconstruction of automotive parts according to claim 1, in step S2, a cross-view collaborative enhancement module CVCEM is constructed, which is characterized in that: The input is the set of residuals corresponding to all perspectives and a set of latent variables , each corresponding to the latent vector representation obtained by encoding the low-resolution image of the th perspective through a variational autoencoder. For each input latent feature , this module updates the features through a cross-perspective residual enhancement mechanism. The specific process is as follows: First, is combined with its neighbor perspective set , and they are respectively transformed through learnable linear mapping matrices and to obtain a query vector and a key vector . Based on the dot product similarity between the query vector and each neighbor key vector, the cross-perspective attention weight is calculated. Subsequently, according to the above attention weight, the total residual of the current perspective is calculated. The specific process is as follows: , where is the residual of perspective , is the residual of the current perspective , is the set self-residual fusion coefficient. Finally, the latent vector representation of the current perspective is updated. The update process is , is the set of latent vector representations after enhanced features and is the output of this module.
3. For an image super-resolution method for 3D reconstruction of automotive parts according to claim 1, in step S3, when constructing the multi-view consistency module MVCM, it is characterized in that: S31. The multi-view consistency module MVCM is divided into four parts: sparsity loss calculation, diffusion loss calculation, constructing the MultiViewUNet network, and joint optimization and parameter update of the total loss. For the sparsity loss calculation part, the input is the residuals of a set of views. , and the specific process is as follows: For each residual variable , calculate the residual sparsity loss corresponding to this view as , where , , correspond to the number of channels, height, and width of the residual features respectively; S32. For the part of constructing the MultiViewUNet network for predicting noise , the input of this network is the current view latent vector after adding noise , the set of adjacent view latent vectors and the set of adjacent view poses . The specific process is as follows: First, flatten and respectively to obtain pixel sequence representations and , where , represents the height and width of the feature map, represents the number of channels, represents the flattening operation; Second, broadcast the pose to each pixel to form a pose feature matrix , where represents the embedding encoding of the pose , represents the broadcast operation, is a unit vector with a length of ; Then, add the neighbor noisy features and the pose embeddings element-wise to obtain the fused key vector . Stack the fused key vectors of all neighbor views along the pixel dimension to form a complete multi-view key set , where is the number of neighbor views; Based on the cross-attention mechanism, use the dot product relationship between the query vector of the current view and the neighbor fused key vector to obtain the fused conditional feature as the final conditional input of the UNet, where function represents the normalization operation, ; Finally, input the noisy latent vector of the current view, the time step , and the fused conditional feature into the UNet backbone network to generate the final noise prediction output ; S33. For the diffusion loss calculation part, the input is a set of enhanced latent vectors and a set of corresponding pose information . For each , forward diffusion is performed through to obtain the noisy vectors and . is the cumulative noise attenuation coefficient corresponding to the -th time step, generated by the diffusion scheduler, and the calculation formula is , where is the preset noise addition ratio per step in the diffusion schedule, represents standard Gaussian noise; the MultiViewUNet network is used to predict the noise . Based on the noise prediction result , the forward latent vector is calculated, and the calculation method is . Finally, the diffusion reconstruction loss is defined as . S34. For the total loss joint optimization and parameter update part, the input is a set of diffusion reconstruction losses and a set of sparsity losses . First, calculate the total loss of the current view as , where is a preset hyperparameter. Second, calculate the total loss of all views as . Third, update the residual by gradient descent based on the total loss of the current view. The specific process is as follows: , where is the preset learning rate, represents the gradient of the residual variable of the current view. Fourth, update the learnable linear mapping matrices and by gradient descent based on the total loss of all views. The specific process is as follows: , , where and are the gradients of the linear mapping matrices and under the overall loss.
4. For an image super-resolution method for 3D reconstruction of automotive parts according to claim 1, in step S4, when constructing the multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF, it is characterized in that: S41. The multi-view consistent diffusion super-resolution reconstruction module MVCDSR-NeRF includes the MVCDSR module and the NeRF reconstruction module. For the MVCDSR module, the input is a set of low-resolution images and their corresponding camera pose parameters , and after , the latent vector representation of the image is obtained , representing the variational autoencoder encoding process; Input the latent vector representation into the Cross-View Collaborative Enhancement Module (CVCEM) to obtain enhanced features ; subsequently, input the enhanced features into the Multi-View Consistency Module (MVCM) to update the residuals and optimize the latent representation; the CVCEM and MVCM are executed Z times in a joint iterative manner to output the finally enhanced latent vector , and then pass it through the variational autoencoder for decoding, where represents the decoding process and is a set of output high-resolution multi-view images; S42. For the NeRF reconstruction module, input the high-resolution multi-view images and their corresponding camera pose parameters into the neural radiance field model for 3D reconstruction, and obtain the reconstructed output image as the input for the next round of the MVCDSR module. The MVCDSR module and the NeRF reconstruction module are executed M times in a joint iterative manner, and finally output an image representation with high-resolution details and a 3D consistent structure and its 3D model.
Citation Information
Patent Citations
Space super-resolution reconstruction method for parallax-guided light field image
CN116823602A
Multistage light field super-resolution network training method, system and product
CN117078514A
Image super-resolution reconstruction method based on deep learning
CN119809933A
Super-resolution image reconstruction method, system and device based on dynamic frequency domain adaptive coding and contrast constraint optimization, and medium
CN119863364A
Label-free adaptive CT super-resolution reconstruction method, system and device based on generative network
US20240169610A1