Model training method based on blurred image, three-dimensional image reconstruction method and device

By introducing a Gaussian modulation network to simulate motion blur, a 3D reconstruction model is trained, which solves the problems of stability and accuracy of 3D reconstruction under fuzzy input conditions. It achieves efficient and robust 3D reconstruction under single fuzzy image conditions and is applicable to scenarios such as Building Information Modeling (BIM).

CN122454028APending Publication Date: 2026-07-24QINGDAO WANJINXIANG MUNICIPAL ENGINEERING CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO WANJINXIANG MUNICIPAL ENGINEERING CO LTD
Filing Date
2026-03-20
Publication Date
2026-07-24

Smart Images

  • Figure CN122454028A_ABST
    Figure CN122454028A_ABST
Patent Text Reader

Abstract

The application provides a model training method based on a blurred image, a three-dimensional image reconstruction method and equipment. The model training method comprises the following steps: constructing a three-dimensional reconstruction initial model; constructing a Gaussian regulation network for training a Gaussian generator network; training the three-dimensional reconstruction initial model by using a blurred image and the Gaussian regulation network, and obtaining a three-dimensional reconstruction model. The model training method, the three-dimensional image reconstruction method and the electronic equipment can effectively improve the processing capacity of the three-dimensional reconstruction model for the blurred image, realize efficient three-dimensional reconstruction based on a single blurred image, have high processing efficiency, and have the characteristics of high efficiency and robustness, and are suitable for three-dimensional reconstruction scenes such as building information modeling (BIM).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of model training technology, and in particular to a model training method based on blurred images, a three-dimensional image reconstruction method and device. Background Technology

[0002] Existing 3D reconstruction methods mostly rely on multiple viewpoints or clear inputs, and cannot achieve stable and reliable 3D representations under fuzzy input conditions. The reconstruction results are inaccurate and inefficient, and are prone to problems such as structural loss, boundary misalignment and texture distortion, making it difficult to meet the requirements of robustness and efficiency in practical applications. Summary of the Invention

[0003] In view of this, the purpose of this application is to propose a model training method, a three-dimensional image reconstruction method and device based on blurred images.

[0004] To achieve the above objectives, this application provides a model training method based on blurred images, comprising:

[0005] Construct an initial model for 3D reconstruction, wherein the initial model for 3D reconstruction includes a Gaussian generator network, which is used to generate 3D Gaussian parameters based on the input image; Construct a Gaussian modulation network for training a Gaussian generator network; The initial 3D reconstruction model is trained using the blurred image and the Gaussian modulation network to obtain the 3D reconstruction model.

[0006] Optionally, the step of training the initial 3D reconstruction model using the blurred image and the Gaussian modulation network to obtain the 3D reconstruction model includes: The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on the blurred image; The image is reconstructed using the Gaussian modulation network based on the corresponding three-dimensional Gaussian parameters to obtain the corresponding reconstructed image. A first loss function is constructed to calculate the loss value between the blurred image and the corresponding reconstructed image. The network parameters of the Gaussian generator network are iteratively updated by minimizing the first loss function until the loss value between the blurred image and the corresponding reconstructed image meets a first preset condition, or the number of iterations reaches a preset first threshold. The training of the initial model of the three-dimensional reconstruction is completed, and the three-dimensional reconstruction model is obtained.

[0007] Optionally, the step of using the Gaussian generator network to generate corresponding three-dimensional Gaussian parameters based on the blurred image includes... The Gaussian generator network is used to extract high-dimensional features corresponding to the blurred image; The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on the corresponding high-dimensional features.

[0008] Optionally, the step of using the Gaussian modulation network to reconstruct the image based on the corresponding three-dimensional Gaussian parameters to obtain the corresponding reconstructed image includes: The Gaussian tuning network is used to perform offset processing based on the corresponding three-dimensional Gaussian parameters to obtain the offset Gaussian parameters. The image is reconstructed using the Gaussian modulation network based on the offset Gaussian parameters to obtain the reconstructed image.

[0009] Optionally, the step of using the Gaussian modulation network to reconstruct the image based on the offset Gaussian parameters to obtain the reconstructed image includes: The reconstructed image is obtained by calculating the following formula using the Gaussian modulation network: ; ; in, To reconstruct the image, This refers to the number of discrete samples during the offset processing. , This represents the total number of three-dimensional Gaussians. , For the instantaneous sharp image of the i-th sample, For the j-th 3D Gaussian parameter of the i-th group, after offset, Let j be the mean value of the j-th three-dimensional Gaussian parameter after the sampling of the i-th group. For the i-th group, the rotation matrix is ​​obtained after offsetting the j-th 3D Gaussian parameter. The scaling matrix after offsetting the j-th 3D Gaussian parameter of the i-th sample group. The spherical harmonic function is the j-th three-dimensional Gaussian parameter offset for the i-th sample group.

[0010] Optionally, the first loss function includes: ; ; in, For the first loss function, The L2 loss is the difference between the blurred image and the corresponding reconstructed image. To reconstruct the image, For blurred images, For blurred images With the corresponding reconstructed image Structural similarity index between them As the first coefficient, The height of the blurred image, , The width of the blurred image, .

[0011] Optionally, the blurred image includes a uniform motion blurred image and a non-uniform motion blurred image, and the blurred image is generated by the following method: A first original image is acquired, and a linear motion blur processing is performed on the first original image to obtain the uniform motion blurred image; A second original image is acquired, and the second original image is blurred based on a random trajectory generation method to obtain the non-uniform motion blurred image.

[0012] Optionally, constructing the initial model for 3D reconstruction includes: Obtain the original 3D reconstruction model, which includes the Gaussian generator network and the differentiable Gaussian grating; The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on a clear image. The three-dimensional image is obtained by using the differentiable Gaussian grating to perform three-dimensional reconstruction based on the corresponding three-dimensional Gaussian parameters. A second loss function is constructed to calculate the loss value between the clear image and the corresponding 3D image. The network parameters of the Gaussian generator network are iteratively updated by minimizing the second loss function until the loss value between the clear image and the corresponding 3D image meets the second preset condition, or the number of iterations reaches the preset second threshold. The training of the original 3D reconstruction model is completed, and the initial 3D reconstruction model is obtained.

[0013] Based on the same inventive concept, this disclosure also provides a three-dimensional image reconstruction method, including: A three-dimensional reconstruction model is used to generate corresponding three-dimensional Gaussian parameters based on the input image, wherein the input image includes a blurred image; The three-dimensional reconstruction model is used to perform three-dimensional reconstruction based on the corresponding three-dimensional Gaussian parameters to obtain the three-dimensional image corresponding to the input image. The three-dimensional reconstruction model is trained based on a model training method based on blurred images as described in any of the preceding claims.

[0014] Based on the same inventive concept, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0015] Based on the same inventive concept, this disclosure also provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the method described above.

[0016] As can be seen from the above, the model training method, 3D image reconstruction method, and device based on blurred images provided in this application introduce a Gaussian modulation network during the training process of the 3D reconstruction model. This network simulates motion blur caused by pixel mixing, transforming blur disturbances into learnable and backpropagable supervised constraints. Based on the difference loss between the output of the Gaussian modulation network and the input blurred image, the training process effectively drives the Gaussian generator to form a more robust feature representation capability under blurred input conditions and establishes a stable mapping relationship from the blurred image to the deblurred 3D Gaussian parameters. This allows it to output reliable 3D Gaussian parameters even when relying solely on a single blurred image. The resulting 3D reconstruction model, with its Gaussian generator network, can extract effective 3D Gaussian features based on blurred images, thus providing reliable data for subsequent 3D reconstruction and rendering. This improves the efficiency, stability, robustness, and reconstruction accuracy of the 3D reconstruction model, making it widely applicable to 3D reconstruction scenarios such as Building Information Modeling (BIM). Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a model training method based on blurred images according to an embodiment of this application; Figure 2 This is a schematic diagram of a three-dimensional image reconstruction method according to an embodiment of this application; Figure 3 This is a schematic diagram of the results of Experiment 2 in the embodiments of this application; Figure 4 This is a schematic diagram of a model training device based on a blurred image according to an embodiment of this application; Figure 5 This is a schematic diagram of a three-dimensional image reconstruction apparatus according to an embodiment of this application; Figure 6 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0020] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0021] 3D reconstruction technology refers to the process of recovering the three-dimensional structure of an object from a two-dimensional image. This technology is widely used in virtual reality, augmented reality, and autonomous driving, providing fundamental support for building realistic and interactive 3D environments and achieving high-quality immersive experiences. 3D Gaussian Splatting (3D GS) is an efficient technique for representing and rapidly rendering 3D scenes. This technique represents points in 3D space as Gaussian distribution functions and projects them onto a 2D image plane for rapid rasterization rendering, thereby achieving efficient 3D reconstruction. However, this technique is highly dependent on high-quality multi-view input, and its performance is closely related to the quality of the input point cloud.

[0022] In practical applications, due to constraints such as acquisition time, viewpoint switching costs, and dynamic scene changes, it is often difficult to obtain clear and aligned image sequences from multiple perspectives, and only a single image can often be acquired. For example, in scenarios such as rapid drone inspection, vehicle-mounted on-the-go data acquisition, mobile snapshot modeling, or robot online perception, the equipment needs to complete imaging and processing in a short time. Factors such as platform jitter, target movement, and long exposure in low light can easily introduce motion blur, often resulting in only a single image, or even a single blurred image. Existing models have at least the following problems when processing this data: First, single-image input itself lacks multi-view geometric constraints, leading to insufficient depth and structural information; second, motion blur causes pixel mixing and loss of detail, creating a degradation deviation between the image and the actual imaging process, further weakening the reliability of feature extraction and 3D parameter regression, thus causing unstable reconstruction phenomena such as structural loss, boundary misalignment, and texture drift. Therefore, it is of great significance to develop a robust and efficient 3D reconstruction method for complex conditions with single-image fuzzy input. This method can still provide usable 3D geometric and appearance information even when data acquisition is limited and the scene changes rapidly, providing basic support for subsequent tasks such as inspection defect identification, measurement and evaluation, path planning and on-site decision-making.

[0023] In view of this, this application proposes a model training method based on blurred images, which can effectively improve the ability of 3D reconstruction models to process blurred images, realize efficient 3D reconstruction based on a single blurred image, and has the characteristics of high processing efficiency and robustness.

[0024] like Figure 1 As shown, the method includes: S101. Construct a three-dimensional reconstruction initial model, wherein the three-dimensional reconstruction initial model includes a Gaussian generator network, which is used to generate three-dimensional Gaussian parameters based on the input image. Specifically, a three-dimensional reconstruction initial model is constructed based on three-dimensional Gaussian sputtering technology. The three-dimensional reconstruction initial model includes a Gaussian Generator Network (GGN) and may also include a differentiable Gaussian Rasterizer (GR). The differentiable Gaussian Rasterizer is used to perform three-dimensional reconstruction based on the three-dimensional Gaussian parameters of the Gaussian Generator Network and outputs a three-dimensional image corresponding to the input image.

[0025] Furthermore, the Gaussian generator network adopts an encoder-decoder architecture. The encoder employs a layer-by-layer downsampling strategy to extract deep features at different resolutions of the input image. The decoder, based on the high-dimensional features extracted by the encoder, uses UNetBlock to progressively upsample and fuses them with the corresponding resolution features stored in the encoder stage through skip connections to fully utilize multi-scale information. Finally, it outputs the three-dimensional Gaussian parameters of the input image.

[0026] UNetBlock is a collective term for the basic feature transformation modules in the U-Net structure. It is typically reused at both the encoder and decoder ends to perform feature extraction and progressive reconstruction at different scales. On the encoder side, UNetBlock mainly consists of stacked operations such as convolution and normalization, used to extract deep features at various resolution scales and form multi-scale representations through layer-by-layer downsampling. On the decoder side, UNetBlock performs upsampling on high-dimensional features and introduces features of the same scale from the encoder stage for fusion (e.g., concatenation or addition by channel) through skip connections. Then, convolution, normalization, and activation operations are used to further refine the fused features. With the help of UNetBlock and skip connections at the decoder end, the network can preserve and supplement detailed information while restoring spatial resolution, thus making fuller use of multi-scale features and improving the quality of parameter regression and reconstruction results.

[0027] S102. Construct a Gaussian modulation network for training the Gaussian generator network; Specifically, the Gaussian Regulator Network (GRN) employs a hierarchical attention mechanism, with each layer containing 64 hidden units and using GELU as the activation function. The hidden layers consist of eight fully connected layers, with each layer alternately stacked with GELU activation to progressively transform features and enhance representation. To improve the expressive power of 3D coordinates in the Gaussian parameter space, the GRN also introduces an embedding function (with an input dimension of 3) and uses six different frequencies to embed the 3D coordinates, thereby enhancing the perception and modeling capabilities of spatial location features.

[0028] S103. The initial 3D reconstruction model is trained using the blurred image and the Gaussian modulation network to obtain the 3D reconstruction model.

[0029] Specifically, after inputting the blurred image into the initial 3D reconstruction model, the Gaussian generator network of the initial 3D reconstruction model encodes the features of the blurred image and regresses the corresponding 3D Gaussian parameters. Based on these 3D Gaussian parameters, the Gaussian modulation network performs image reconstruction; for example, after modulating the Gaussian parameters, a reprojected image is obtained through differentiable rendering. Then, the image reconstruction result is aligned and compared with the input blurred image to construct the difference loss between the two. Using this loss as the optimization objective, the parameters of the Gaussian generator network are updated and iteratively trained through backpropagation. As the training process progresses, the Gaussian generator network gradually learns effective feature representations and parameterized mapping relationships for blurred images, enabling the final 3D reconstruction model to output more reliable 3D Gaussian parameters based on the blurred image. This provides a more stable and usable parameter input for the subsequent differentiable Gaussian raster, improving the robustness and reconstruction accuracy of the 3D reconstruction process.

[0030] Existing 3D Gaussian reconstruction methods typically use sharp images as training input and supervision signals, lacking a dedicated training mechanism for blurred images. Furthermore, in conventional training processes, after the Gaussian generator network outputs 3D Gaussian parameters, the rendering result is directly obtained by a differentiable Gaussian rasterizer, and the backpropagation gradient of the rendering error is propagated back to the Gaussian generator network to update the parameters. However, the differentiable Gaussian rasterizer cannot characterize blurred imaging processes such as pixel blending and motion trajectory integration. Its forward rendering model and the exposure time integration observation model corresponding to motion-blurred images are mismatched. When the input image has motion blur, a deviation occurs between the rendering error and the actual imaging mechanism. The gradient obtained from backpropagation is more prone to instability or lacks effective directionality, making it difficult to form effective parameter updates for blur degradation. Consequently, the model struggles to obtain usable Gaussian parameter representations under blurred input conditions, making it difficult to stably adapt and effectively apply to new perspective rendering and 3D reconstruction tasks for blurred images.

[0031] In this application, based on steps S101-S103, a Gaussian modulation network is introduced during the training process of the 3D reconstruction model. This network simulates motion blur caused by pixel mixing, transforming blur perturbations into learnable and backpropagable supervised constraints. Based on the difference loss between the output of the Gaussian modulation network and the input blurred image, the training process effectively drives the Gaussian generator to develop more robust feature representation capabilities under blurred input conditions and establishes a stable mapping relationship from the blurred image to the deblurred 3D Gaussian parameters. This allows it to output reliable 3D Gaussian parameters even when relying solely on a single blurred image. The resulting 3D reconstruction model, with its Gaussian generator network, can extract effective 3D Gaussian features based on the blurred image, providing reliable data for subsequent 3D reconstruction and rendering. This improves the efficiency, stability, robustness, and reconstruction accuracy of the 3D reconstruction model, making it widely applicable to 3D reconstruction scenarios such as Building Information Modeling (BIM).

[0032] In some embodiments, training the initial 3D reconstruction model using the blurred image and the Gaussian modulation network to obtain the 3D reconstruction model includes: S21. Use the Gaussian generator network to generate corresponding three-dimensional Gaussian parameters based on the blurred image; Specifically, after the blurred image is input into the Gaussian generator network, the Gaussian generator network performs feature encoding and multi-scale feature extraction on the image, and establishes a parameterized mapping relationship from two-dimensional observation to three-dimensional representation based on the extracted features. Then, it performs three-dimensional Gaussian parameter regression on the scene corresponding to the input image and outputs three-dimensional Gaussian parameters that match the blurred image.

[0033] S22. Using the Gaussian modulation network based on the corresponding three-dimensional Gaussian parameters, the image is reconstructed to obtain the corresponding reconstructed image; Specifically, in real-world scenarios, image blurring is typically caused by two types of degradation processes: camera defocusing and camera motion. Motion blur can be represented as: ; in, This represents the blurred image. Representative from the sports field The unknown fuzzy kernel of the decision, Represents noise. Represents convolution operation. To achieve a clear image. In convolution operations, blurring stems from the weighted stacking of adjacent pixels, meaning the value of a single pixel is influenced by the varying intensities of its neighboring pixels, leading to image degradation.

[0034] Therefore, during image reconstruction based on the corresponding 3D Gaussian parameters in S22, the Gaussian modulation network modulates and reconstructs these 3D Gaussian parameters to address motion blur. By applying operations such as position offset, scale / covariance adjustment, and weight allocation to the Gaussian parameters corresponding to each pixel, it simulates the trajectory changes and pixel mixing effects caused by camera or object motion during exposure at the pixel level, causing the radiation contribution of the same spatial point to diffuse and superimpose on the imaging plane. Subsequently, the modulated Gaussian parameters are used for image synthesis to obtain a reconstructed image consistent with the motion blur imaging mechanism. This reconstructed image can be regarded as a differentiable simulation result of the input blurred image generation process.

[0035] S23. Construct a first loss function to calculate the loss value between the blurred image and the corresponding reconstructed image. Iteratively update the network parameters of the Gaussian generator network by minimizing the first loss function until the loss value between the blurred image and the corresponding reconstructed image meets the first preset condition, or the number of iterations reaches the preset first threshold, thereby completing the training of the initial model of the three-dimensional reconstruction and obtaining the three-dimensional reconstruction model.

[0036] Specifically, in each iteration, the Gaussian generator network generates and outputs corresponding 3D Gaussian parameters based on the input blurred image. These parameters are then used by the Gaussian modulation network to reconstruct the image. Subsequently, the loss value of the first loss function is calculated, and the gradient of the loss value relative to the Gaussian generator network parameters is obtained through backpropagation. A gradient descent-type optimization algorithm is then used to update the Gaussian generator network parameters. This process is repeated until a preset training termination condition is met: when the loss value between the blurred image and the reconstructed image meets the first preset condition (e.g., loss convergence or the decrease is below a threshold), or when the number of iterations reaches a preset first-stage threshold, training stops, and the updated network parameters are output. This completes the training of the initial 3D reconstruction model, resulting in the 3D reconstruction model. The preset first-stage threshold can be set to 50,000, 60,000, 40,000 iterations, or other values; no specific restrictions apply.

[0037] In this embodiment, based on steps S21-S23, a first loss function is constructed to explicitly quantify the difference between the blurred image and the reconstructed image. The Gaussian generator network is iteratively updated with minimizing this difference as the optimization objective. Due to the image reconstruction process, the Gaussian modulation network modulates and renders the three-dimensional Gaussian parameters, which can introduce the characterization of the motion blur pixel mixing effect during the optimization process. This allows the gradient of backpropagation to more effectively constrain the regression direction and update magnitude of the Gaussian parameters, thereby promoting stable convergence of the training process. The final three-dimensional reconstruction model can output a more consistent and reliable three-dimensional Gaussian parameter representation under blurred image input conditions, and achieves lower reprojection error and better structural consistency in subsequent rendering and reconstruction stages, thereby improving the robustness, stability, and reconstruction accuracy of new perspective rendering and three-dimensional reconstruction.

[0038] In some embodiments, the step of generating corresponding three-dimensional Gaussian parameters based on a blurred image using the Gaussian generator network includes: The Gaussian generator network is used to extract high-dimensional features corresponding to the blurred image; The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on the corresponding high-dimensional features.

[0039] Specifically, the process of generating three-dimensional Gaussian parameters will be described through a more detailed example.

[0040] The Gaussian generator network employs an encoder-decoder architecture. The blurred image is an RGB channel image; for the input blurred image... The UNetBlock in the encoder of the Gaussian generator network extracts features from the blurred image at different resolution scales. Each UNetBlock consists of a Group Normalization (GN) layer and... The network consists of stacked convolutional layers. During computation, feature maps corresponding to each resolution are stored in skip connections (SC) to preserve key information and enhance gradient propagation efficiency. At the 16×16 feature scale of encoding and decoding, the Gaussian generator network introduces a multi-head self-attention (MHSA) mechanism to capture global dependencies, thereby improving feature representation capabilities. Finally, the encoder of the Gaussian generator network outputs the high-dimensional features corresponding to the blurred image.

[0041] The decoder of the Gaussian generator network receives the encoder's output. Based on the high-dimensional features output by the encoder, it uses UNetBlock to progressively upsample the data and fuses it with the corresponding resolution features stored in the encoder stage through skip connections to fully utilize multi-scale information. The Gaussian generator network generates a Gaussian parameter set for each pixel in the blurred image. Its dimensions are Gaussian parameter set It can be expressed by the following formula: ; in, and These represent the height and width of the input image (i.e., the blurred image), respectively. The spatial distribution of the Gaussian parameters is aligned with the input image. The number of pixels in the input image (i.e., the blurred image).

[0042] Finally, the Gaussian generator network outputs the Gaussian parameter set corresponding to each pixel of the blurred image, thus obtaining the three-dimensional Gaussian parameters corresponding to the blurred image. Specifically, each pixel... The generated Gaussian parameter set It can be expressed as follows: ; in, For pixels The scaling matrix, For pixels The rotation matrix (represented in quaternion form). For pixels The mean (i.e., 3D position). For pixels Opacity For pixels The spherical harmonic function is used to characterize viewpoint-dependent appearance information.

[0043] In some embodiments, the step of using the Gaussian modulation network to reconstruct the image based on the corresponding three-dimensional Gaussian parameters to obtain the corresponding reconstructed image includes: S31. Using the Gaussian modulation network, perform offset processing based on the corresponding three-dimensional Gaussian parameters to obtain the offset Gaussian parameters. Specifically, the Gaussian control network simulates the displacement during camera exposure by generating multiple sets of offset three-dimensional Gaussian functions and performing a weighted average. Its mathematical expression is as follows: ; in, This is represented as the number of discrete samples simulating the camera displacement process. That is, to generate additional Group of three-dimensional Gaussian parameters; The network function representing a Gaussian-controlled network; Represents positional encoding; , This represents the total number of three-dimensional Gaussians. For parameter offset; The j-th 3D Gaussian rotation matrix, Let j be the scaling matrix of the j-th 3D Gaussian. Let be the mean of the j-th three-dimensional Gaussian. Let be the spherical harmonic function of the j-th three-dimensional Gaussian; For the i-th group of samples, the position The offset; For the i-th group of samples, the rotation The offset; For the i-th group of samples, the scale The offset; For the i-th group of samples, the appearance The offset.

[0044] The offset parameter set is represented by the following formula: ; ; ; ; ; in, Let be the mean value of the j-th 3D Gaussian parameter after offset in the i-th sampling group. Let be the rotation matrix after offsetting the j-th 3D Gaussian parameter in the i-th sample group. Let be the scaling matrix after offsetting the j-th 3D Gaussian parameter in the i-th sample group. Let be the spherical harmonic function after offsetting the j-th three-dimensional Gaussian parameter in the i-th sampling group.

[0045] S32. The image is reconstructed using the Gaussian modulation network based on the offset Gaussian parameters to obtain the reconstructed image.

[0046] Specifically, the Gaussian tuning network generates the result after offset processing. The three-dimensional Gaussian parameters are rasterized to generate... Zhang instantly clear image, then... The image is synthesized by linear averaging of the instantaneous clear images. Reconstructing the image The expression is: ; ; in This refers to the number of discrete samples during the offset processing. , This represents the total number of three-dimensional Gaussians. , This is a snapshot of the i-th sampled image rendered by a rasterizer using a Gaussian-controlled network. For the j-th 3D Gaussian parameter of the i-th group, after offset, Let j be the mean value of the j-th three-dimensional Gaussian parameter after the sampling of the i-th group. For the i-th group, the rotation matrix is ​​obtained after offsetting the j-th 3D Gaussian parameter. The scaling matrix after offsetting the j-th 3D Gaussian parameter of the i-th sample group. The spherical harmonic function is obtained by offsetting the j-th 3D Gaussian parameter of the i-th sample. Then, the reconstructed image is calculated. The error between the blurred image and the real input image can be optimized by using the backpropagation mechanism to optimize the network parameters of the Gaussian generator network.

[0047] In some embodiments, the first loss function includes: ; ; in, For the first loss function, The L2 loss is the difference between the blurred image and the corresponding reconstructed image. To reconstruct the image, For blurred images, For blurred images With the corresponding reconstructed image Structural similarity index between them As the first coefficient, The height of the blurred image, , The width of the blurred image, .

[0048] In this embodiment, the first loss function is constructed using the L2 loss function and the Structural Similarity Index Measure (SSIM). The L2 loss function constrains the pixel-level differences between the reconstructed image and the blurred image, while the SSIM constrains the structural differences between the reconstructed image and the blurred image. Then, weighted coefficients (i.e., the first coefficient) are used. The first loss function balances the contributions of L2 loss and SSIM loss, ensuring that pixel intensity consistency and local structure consistency are considered simultaneously during optimization. During training, L2 loss provides stable numerical error constraints, promoting alignment of low-frequency information such as overall brightness and color; SSIM loss strengthens the preservation of structural information such as edges, textures, and local contrast, reducing over-smoothing or detail loss caused by relying solely on pixel errors. Weighting coefficients balance the contributions of the two types of losses, forming a more reasonable error feedback and parameter update direction. This allows the Gaussian generator network to more fully learn the effective features in the blurred image and regress a more reliable 3D Gaussian parameter representation, thereby improving the consistency and stability of the reconstructed result with the input blurred image in terms of visual structure, providing higher-quality parameter input for subsequent new perspective rendering and 3D reconstruction.

[0049] Datasets available for training 3D reconstruction models include ShapeNet, Objaverse-lvis, and the Google Scans dataset. ShapeNet is a large-scale, high-quality dataset covering various common object categories such as airplanes, cars, and tables, and provides relatively complete parameter information. Objaverse-lvis contains a large number of object renderings, and the Google Scans dataset has thousands of real-world household item models. However, existing datasets lack large-scale, publicly available datasets that simultaneously provide realistic motion-blurred multi-view images and their corresponding camera parameters. To simulate motion blur in real-world scenes, this application further performs blurring processing on existing datasets to construct the blurred multi-view data required for training.

[0050] In some embodiments, the blurred image includes a uniform motion blurred image and a non-uniform motion blurred image, and the blurred image is generated by the following method: A first original image is acquired, and a linear motion blur processing is performed on the first original image to obtain the uniform motion blurred image; A second original image is acquired, and the second original image is blurred based on a random trajectory generation method to obtain the non-uniform motion blurred image.

[0051] Specifically, uniform motion blur processing and non-uniform motion blur processing are performed on the obtained original image data (i.e., the first original image and the second original image) to construct the blurred image data required for training.

[0052] Images from two categories, cars and chairs, were selected for blurring. Each category contained 251 RGB images from different viewpoints with a resolution of 128×128. The dataset also provides corresponding camera intrinsic and extrinsic parameter matrices to describe the camera's pose and imaging parameters in 3D space.

[0053] Based on the image degradation mathematical model, linear motion blur processing (i.e., uniform motion blur processing) is applied to chair category images (i.e., the first original image) to generate images with linear motion blur features, resulting in uniformly motion blurred images. The specific calculation formula is as follows: ; in, This is the frequency domain representation of the first image. For the frequency domain representation of a uniform motion-blurred image, the transfer function Defined as: ; in, The frequency components in the horizontal direction of the frequency domain coordinates. The frequency components in the vertical direction of the frequency domain coordinates. For exposure time, parameters and These represent the camera exposure time. Within, the displacement distance of the image in the horizontal and vertical directions, The phase term represents the frequency domain phase shift caused by motion, where... It is the imaginary unit.

[0054] Through the above calculations, motion blur is modeled as a convolution process between the image and the motion kernel, which can effectively simulate the blurring effect caused by the linear displacement of the camera during exposure, so as to quickly obtain a uniform motion-blurred image.

[0055] For non-uniform motion blur scenes, a random trajectory generation method is used to blur images of car categories (i.e., the second original image) to obtain non-uniform motion blur images. The random trajectory generation method constructs a complex-valued trajectory vector that follows two-dimensional random motion in a continuous domain, and combines this with a Markov process to simulate the object's trajectory. The blur kernel is obtained by sub-pixel interpolation of the trajectory vector. The trajectory generation process comprehensively considers factors such as the previous moment's position and velocity, inertia, Gaussian sway, and impulse perturbation, thus generating non-linear motion blur images with higher realism and complexity.

[0056] In some embodiments, constructing the initial model for 3D reconstruction includes: Obtain the original 3D reconstruction model, which includes the Gaussian generator network and the differentiable Gaussian grating; The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on a clear image. The three-dimensional image is obtained by using the differentiable Gaussian grating to perform three-dimensional reconstruction based on the corresponding three-dimensional Gaussian parameters. A second loss function is constructed to calculate the loss value between the clear image and the corresponding 3D image. The network parameters of the Gaussian generator network are iteratively updated by minimizing the second loss function until the loss value between the clear image and the corresponding 3D image meets the second preset condition, or the number of iterations reaches the preset second threshold. The training of the original 3D reconstruction model is completed, and the initial 3D reconstruction model is obtained.

[0057] Specifically, a two-stage training strategy is employed to obtain the final 3D reconstruction model. In the first stage, the original 3D reconstruction model is trained using clear images, enabling the Gaussian generator network to establish stable feature extraction and 3D Gaussian parameter regression capabilities under conditions free from blurring interference. This yields an initial 3D reconstruction model that accurately represents the geometry and appearance of a clear scene. Optionally, the Adam optimizer can be used, with a learning rate set to... The number of iterative training rounds (i.e., the preset threshold for the second round) can be set to 800,000, with no specific limitation. In the second stage, a Gaussian modulation network is introduced based on the initial 3D reconstruction model, and the model is further optimized using blurred images. This allows the model to consider the pixel mixing effect caused by motion blur during the optimization process, thereby enhancing the Gaussian generator network's adaptability to blurred inputs. This enables it to regress a more reliable deblurred 3D Gaussian parameter representation from blurred images, ultimately resulting in a more robust 3D reconstruction model for blurred images. Optionally, the learning rate of the Gaussian modulation network in the second stage can be set to... The number of iterative training rounds (i.e., the preset threshold for the first round) can be set to 50,000 rounds, with no specific limit. The second preset condition can be set to loss convergence or the rate of decrease falling below a threshold, etc., with no specific limit.

[0058] Based on the same inventive concept, this application also provides a three-dimensional image reconstruction method.

[0059] like Figure 2 As shown, the three-dimensional image reconstruction method includes: S201. Generate corresponding three-dimensional Gaussian parameters based on the input image using a three-dimensional reconstruction model, wherein the input image includes a blurred image; S202. Using the three-dimensional reconstruction model, perform three-dimensional reconstruction based on the corresponding three-dimensional Gaussian parameters to obtain the three-dimensional image corresponding to the input image; The three-dimensional reconstruction model is trained based on a model training method based on blurred images as described in any of the foregoing embodiments.

[0060] In this application, the 3D reconstruction model includes a Gaussian generator network and a differentiable Gaussian raster. The Gaussian generator network generates corresponding 3D Gaussian parameters based on the input image, and then the differentiable Gaussian raster performs 3D reconstruction based on the corresponding 3D Gaussian parameters to obtain the 3D image corresponding to the input image. In the above process, since the Gaussian generator network is trained and optimized under the constraints of the blurred image and the Gaussian modulation network, it has explicitly adapted to pixel mixing degradation caused by motion blur during parameter regression. Therefore, it can output more stable and reliable 3D Gaussian parameters under blurred input conditions, thus providing more consistent and usable parameter input for the differentiable Gaussian raster. Therefore, even if the input image is a blurred image with motion blur interference, a clear 3D image can still be reconstructed, improving the robustness and stability of the 3D reconstruction process, reducing reconstruction artifacts and structural deviations, thereby improving the rendering effect of new perspectives and increasing the accuracy of 3D reconstruction.

[0061] To further illustrate the technical effects of this application, the 3D reconstruction model trained by this application was compared with other models. The specific experiment and comparison process is as follows.

[0062] Multiple comparison groups were constructed based on the SI (Splatter Image) model, Restormer model, PixelNeRF joint model, and OpenLRM. The construction of each comparison group is shown in the table below: Table 1. Detailed Explanation of the Comparison Group

[0063] Evaluation metrics include Peak Signal to Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), Learned Perceptual Image Patch Similarity (LPIPS), and Frames Per Second (FPS).

[0064] PSNR measures the quality difference between two images based on Mean Square Error (MSE). SSIM evaluates image similarity from three dimensions: brightness, contrast, and structural information, based on the characteristics of the human visual system. LPIPS calculates the perceptual distance between images using a deep learning model, which is more in line with human visual perception characteristics. In the experiment, a pre-trained VGG deep network was used to extract image features to calculate the LPIPS score. To evaluate the inference speed and actual running efficiency of different 3D reconstruction methods, frame rate per second (FPS) was introduced as an evaluation metric, which measures the number of viewpoint images that the model can render per second. Higher PSNR and SSIM values ​​indicate better reconstruction quality, while lower LPIPS values ​​indicate higher perceptual similarity between the generated image and the target image.

[0065] Since this application is the first to address blurred single-view... Figure 3 For 3D image reconstruction, there is currently no directly comparable benchmark model. Therefore, in the comparison group, an indirect evaluation strategy of "deblurring first, then 3D reconstruction" was adopted. That is, the Restormer model was used to deblur the input image during the deblurring stage, and then 3D reconstruction was performed based on the deblurred image. The Restormer model is based on an efficient Transformer architecture, which can accurately capture the complex dependencies between distant pixels in the image, and exhibits excellent performance in image deblurring tasks.

[0066] Based on the above control group, three sets of experiments were conducted: (1) Experiment 1 Nonlinear motion blurring was performed on car category images from the Shapenet dataset to obtain non-uniform motion-blurred images. These non-uniform motion-blurred images were then input into comparison groups 1-3 and the 3D reconstruction model trained in this application. The unblurred original images were input into comparison group 5. The final test results are shown in the table below. Table 2 Comparison of various indicators between the control groups in Experiment 1 and this application

[0067] (2) Experiment 2 Linear motion blurring was performed on chair category images from the Shapenet dataset to obtain corresponding uniformly blurred images. These uniformly blurred images were then input into comparison groups 1-3 and the 3D reconstruction model trained in this application. The unblurred original images were input into comparison group 5. The final test results are shown in the table below: Table 3 Comparison of various indicators between the control groups in Experiment 2 and this application

[0068] In this experiment, the 3D reconstruction results of comparison groups 2, 3, and 5 with those of this application on chair category images are attached. Figure 3 As shown in the figure. The 3D reconstruction results of comparison group 1 are too blurry and are therefore not shown in the attached figures.

[0069] From the appendix Figure 3 It can be seen that, with a blurred chair image as input, the 3D reconstruction results of comparison groups 2 and 3 are significantly inferior to the 3D reconstruction model trained in this application in terms of structural integrity and detail clarity. Since the leg structure is difficult to discern in the blurred input chair image, comparison groups 2 and 3 failed to effectively reconstruct this part of the structure; however, the 3D reconstruction model trained in this application can accurately restore the complete shape of the chair legs, and its reconstruction results are basically consistent with the 3D image generated by SI (i.e., comparison group 5) under clear input conditions in terms of detail representation. This demonstrates that the model training method proposed in this application can significantly improve the quality of 3D reconstruction under blurred single-view conditions, and has obvious advantages in structural restoration ability and detail representation.

[0070] (3) Experiment 3 The Google Scan dataset was used, and the images in the dataset were subjected to nonlinear motion blur to obtain corresponding non-uniformly blurred images. These non-uniformly blurred images were then input into comparison groups 1-2, 4, and the 3D reconstruction model trained in this application. The unblurred original images were input into comparison group 5. The final test results are shown in the table below: Table 4. Comparison of various indicators between the control groups in Experiment 3 and this application.

[0071] The results of Experiments 1-3 (Tables 2-4) show that under clear input conditions, the SI model (comparison group 5) performs well on all datasets. However, under blurred input conditions (comparison group 1), its reconstruction performance significantly decreases due to a lack of robustness to blurred input. The test results of Comparison Groups 2 and 3 show that even with Restormer model preprocessing for image deblurring, the reconstruction metrics of the SI or PixelNeRF models still cannot be restored to the level of clear input. PixelNeRF, using known views as auxiliary input for neural rendering, exhibits good 3D reconstruction performance under sparse input conditions. However, PixelNeRF is extremely sensitive to the quality of the input image; even after deblurring, it is difficult to recover high-quality 3D information in real datasets. Further verification on the Google Scan dataset, as shown in Table 4, shows that under blurred input, SI also fails to achieve effective 3D reconstruction. The open-source, triplane-based large-scale reconstructor OpenLRM, even after deblurring using Restormer, still shows unsatisfactory reconstruction results.

[0072] In contrast, the 3D reconstruction model trained in this application, due to the effective training of the Gaussian generator network using a Gaussian modulation network during the training process, can directly optimize the generation of Gaussian parameters, enabling it to adapt to blurred inputs and generate high-quality new perspective images. As shown in Tables 2-4, the 3D reconstruction model of this application can achieve efficient and high-precision 3D reconstruction of blurred images. Comparison with group 5 in Tables 2-4 shows that the 3D reconstruction processing capability of this application for blurred images is close to that of the SI model for clear images, with processing results almost identical to those for clear images. This demonstrates that the training method of this application effectively improves the 3D reconstruction model's ability to reconstruct blurred images.

[0073] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.

[0074] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0075] Based on the same inventive concept, corresponding to the model training method described in any of the above embodiments, this application also provides a model training device based on blurred images.

[0076] refer to Figure 4 The model training device includes: The model building module 301 is used to build an initial model for three-dimensional reconstruction, wherein the initial model for three-dimensional reconstruction includes a Gaussian generator network, which is used to generate three-dimensional Gaussian parameters based on the input image; Network building module 302 is used to build a Gaussian modulation network for training the Gaussian generator network; Training module 303 is used to train the initial 3D reconstruction model using the blurred image and the Gaussian modulation network to obtain the 3D reconstruction model.

[0077] In some embodiments, the training module 303 is further configured to: The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on the blurred image; The image is reconstructed using the Gaussian modulation network based on the corresponding three-dimensional Gaussian parameters to obtain the corresponding reconstructed image. A first loss function is constructed to calculate the loss value between the blurred image and the corresponding reconstructed image. The network parameters of the Gaussian generator network are iteratively updated by minimizing the first loss function until the loss value between the blurred image and the corresponding reconstructed image meets a first preset condition, or the number of iterations reaches a preset first threshold. The training of the initial model of the three-dimensional reconstruction is completed, and the three-dimensional reconstruction model is obtained.

[0078] In some embodiments, the step of generating corresponding three-dimensional Gaussian parameters based on a blurred image using the Gaussian generator network includes: The Gaussian generator network is used to extract high-dimensional features corresponding to the blurred image; The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on the corresponding high-dimensional features.

[0079] In some embodiments, the step of using the Gaussian modulation network to reconstruct the image based on the corresponding three-dimensional Gaussian parameters to obtain the corresponding reconstructed image includes: The Gaussian tuning network is used to perform offset processing based on the corresponding three-dimensional Gaussian parameters to obtain the offset Gaussian parameters. The image is reconstructed using the Gaussian modulation network based on the offset Gaussian parameters to obtain the reconstructed image.

[0080] In some embodiments, the step of using the Gaussian modulation network to reconstruct the image based on the offset Gaussian parameters to obtain the reconstructed image includes: The reconstructed image is obtained by calculating the following formula using the Gaussian modulation network: ; ; in, To reconstruct the image, This refers to the number of discrete samples during the offset processing. , This represents the total number of three-dimensional Gaussians. , For the instantaneous sharp image of the i-th sample, For the j-th 3D Gaussian parameter of the i-th group, after offset, Let j be the mean value of the j-th three-dimensional Gaussian parameter after the sampling of the i-th group. For the i-th group, the rotation matrix is ​​obtained after offsetting the j-th 3D Gaussian parameter. The scaling matrix after offsetting the j-th 3D Gaussian parameter of the i-th sample group. The spherical harmonic function is the j-th three-dimensional Gaussian parameter offset for the i-th sample group.

[0081] In some embodiments, the first loss function includes: ; ; in, For the first loss function, The L2 loss is the difference between the blurred image and the corresponding reconstructed image. To reconstruct the image, For blurred images, For blurred images With the corresponding reconstructed image Structural similarity index between them As the first coefficient, The height of the blurred image, , The width of the blurred image, .

[0082] In some embodiments, the blurred image includes a uniform motion blurred image and a non-uniform motion blurred image. The model training device further includes an image blurring module, which generates the blurred image by means of the following methods: A first original image is acquired, and a linear motion blur processing is performed on the first original image to obtain the uniform motion blurred image; A second original image is acquired, and the second original image is blurred based on a random trajectory generation method to obtain the non-uniform motion blurred image.

[0083] In some embodiments, the model building module 301 is further configured to: Obtain the original 3D reconstruction model, which includes the Gaussian generator network and the differentiable Gaussian grating; The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on a clear image. The three-dimensional image is obtained by using the differentiable Gaussian grating to perform three-dimensional reconstruction based on the corresponding three-dimensional Gaussian parameters. A second loss function is constructed to calculate the loss value between the clear image and the corresponding 3D image. The network parameters of the Gaussian generator network are iteratively updated by minimizing the second loss function until the loss value between the clear image and the corresponding 3D image meets the second preset condition, or the number of iterations reaches the preset second threshold. The training of the original 3D reconstruction model is completed, and the initial 3D reconstruction model is obtained.

[0084] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.

[0085] The apparatus of the above embodiments is used to implement a model training method based on a blurred image in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0086] Based on the same inventive concept, corresponding to the three-dimensional image reconstruction method described in any of the above embodiments, this application also provides a three-dimensional image reconstruction apparatus.

[0087] refer to Figure 5 The three-dimensional image reconstruction device includes: The feature extraction module 401 is used to generate corresponding three-dimensional Gaussian parameters based on the input image using a three-dimensional reconstruction model, wherein the input image includes a blurred image; The 3D reconstruction module 402 is used to perform 3D reconstruction based on the corresponding 3D Gaussian parameters using the 3D reconstruction model to obtain the 3D image corresponding to the input image. The three-dimensional reconstruction model is trained based on a model training method based on blurred images as described in any of the foregoing embodiments.

[0088] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.

[0089] The apparatus of the above embodiments is used to implement a corresponding three-dimensional image reconstruction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0090] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements a model training method based on a fuzzy image and / or a three-dimensional image reconstruction method as described in any of the above embodiments.

[0091] Figure 6 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0092] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0093] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0094] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0095] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0096] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0097] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0098] The electronic devices described above are used to implement a model training method based on a blurred image and / or a three-dimensional image reconstruction method in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0099] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute a model training method based on a blurred image and / or a three-dimensional image reconstruction method as described in any of the above embodiments.

[0100] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0101] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute a model training method based on a blurred image and / or a three-dimensional image reconstruction method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0102] Based on the same concept, corresponding to any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to perform the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0103] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0104] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0105] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0106] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0107] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0108] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0109] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0110] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the claims of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

Claims

1. A model training method based on blurred images, characterized in that, include: Construct an initial model for 3D reconstruction, wherein the initial model for 3D reconstruction includes a Gaussian generator network, which is used to generate 3D Gaussian parameters based on the input image; Construct a Gaussian modulation network for training a Gaussian generator network; The initial 3D reconstruction model is trained using the blurred image and the Gaussian modulation network to obtain the 3D reconstruction model.

2. The model training method based on blurred images according to claim 1, characterized in that, The process of training the initial 3D reconstruction model using the blurred image and the Gaussian modulation network to obtain the 3D reconstruction model includes: The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on the blurred image; The image is reconstructed using the Gaussian modulation network based on the corresponding three-dimensional Gaussian parameters to obtain the corresponding reconstructed image. A first loss function is constructed to calculate the loss value between the blurred image and the corresponding reconstructed image. The network parameters of the Gaussian generator network are iteratively updated by minimizing the first loss function until the loss value between the blurred image and the corresponding reconstructed image meets a first preset condition, or the number of iterations reaches a preset first threshold. The training of the initial model of the three-dimensional reconstruction is completed, and the three-dimensional reconstruction model is obtained.

3. The model training method based on blurred images according to claim 2, characterized in that, The process of generating corresponding 3D Gaussian parameters based on a blurred image using the Gaussian generator network includes... The Gaussian generator network is used to extract high-dimensional features corresponding to the blurred image; The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on the corresponding high-dimensional features.

4. The model training method based on blurred images according to claim 2, characterized in that, The step of using the Gaussian modulation network to reconstruct the image based on the corresponding three-dimensional Gaussian parameters to obtain the corresponding reconstructed image includes: The Gaussian tuning network is used to perform offset processing based on the corresponding three-dimensional Gaussian parameters to obtain the offset Gaussian parameters. The image is reconstructed using the Gaussian modulation network based on the offset Gaussian parameters to obtain the reconstructed image.

5. The model training method based on blurred images according to claim 4, characterized in that, The step of using the Gaussian modulation network to reconstruct the image based on the offset Gaussian parameters to obtain the reconstructed image includes: The reconstructed image is obtained by calculating the following formula using the Gaussian modulation network: ; ; in, To reconstruct the image, This refers to the number of discrete samples during the offset processing. , This represents the total number of three-dimensional Gaussians. , For the instantaneous sharp image of the i-th sample, For the j-th 3D Gaussian parameter of the i-th group, after offset, Let j be the mean value of the j-th three-dimensional Gaussian parameter after the sampling of the i-th group. For the i-th group, the rotation matrix is ​​obtained after offsetting the j-th 3D Gaussian parameter. The scaling matrix after offsetting the j-th 3D Gaussian parameter of the i-th sample group. The spherical harmonic function is the j-th three-dimensional Gaussian parameter offset for the i-th sample group.

6. The model training method based on blurred images according to claim 2, characterized in that, The first loss function includes: ; ; in, For the first loss function, The L2 loss is the difference between the blurred image and the corresponding reconstructed image. To reconstruct the image, For blurred images, For blurred images With the corresponding reconstructed image Structural similarity index between them As the first coefficient, The height of the blurred image, , The width of the blurred image, .

7. The model training method based on blurred images according to claim 1, characterized in that, The blurred image includes uniform motion blurred images and non-uniform motion blurred images, and the blurred image is generated by the following methods: A first original image is acquired, and a linear motion blur processing is performed on the first original image to obtain the uniform motion blurred image; A second original image is acquired, and the second original image is blurred based on a random trajectory generation method to obtain the non-uniform motion blurred image.

8. The model training method based on blurred images according to claim 1, characterized in that, The construction of the initial 3D reconstruction model includes: Obtain the original 3D reconstruction model, which includes the Gaussian generator network and the differentiable Gaussian grating; The Gaussian generator network is used to generate corresponding three-dimensional Gaussian parameters based on a clear image. The three-dimensional image is obtained by using the differentiable Gaussian grating to perform three-dimensional reconstruction based on the corresponding three-dimensional Gaussian parameters. A second loss function is constructed to calculate the loss value between the clear image and the corresponding 3D image. The network parameters of the Gaussian generator network are iteratively updated by minimizing the second loss function until the loss value between the clear image and the corresponding 3D image meets the second preset condition, or the number of iterations reaches the preset second threshold. The training of the original 3D reconstruction model is completed, and the initial 3D reconstruction model is obtained.

9. A three-dimensional image reconstruction method, characterized in that, include: A three-dimensional reconstruction model is used to generate corresponding three-dimensional Gaussian parameters based on the input image, wherein the input image includes a blurred image; The three-dimensional reconstruction model is used to perform three-dimensional reconstruction based on the corresponding three-dimensional Gaussian parameters to obtain the three-dimensional image corresponding to the input image. The three-dimensional reconstruction model is trained based on a model training method based on blurred images as described in any one of claims 1 to 8.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 9.