Image generation method and system, image generation model training method and system, equipment and medium

By combining local and global feature extraction networks with a cross-attention mechanism, three-dimensional Gaussian splash parameters are generated, which solves the problems of holes, blurring artifacts and modeling difficulties in point cloud rendering and achieves high-quality image reconstruction effects.

CN120655827APending Publication Date: 2025-09-16SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510754770.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies have holes and blurring artifacts in point cloud rendering, difficulty in modeling surface texture mutations, high complexity in radiation field reconstruction, and the problem of balancing rendering speed and quality. In addition, neural radiation field rendering lacks end-to-end forward prediction capabilities, which limits its practicality.

Method used

By acquiring sparse point cloud data, using local feature extraction network and global feature extraction network, combined with the cross-attention mechanism, the local and global features of the sparse point cloud are extracted, and the fusion network of the image generation model is used to generate three-dimensional Gaussian splash parameters for rendering to generate high-quality reconstructed images.

Benefits of technology

It improves the accuracy and realism of image reconstruction, can truly restore the geometric shape and appearance texture of the target object, and solves the problem of inaccurate image reconstruction in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655827A_ABST
    Figure CN120655827A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an image generation method and system, an image generation model training method and system, equipment and a medium. The method comprises the following steps: acquiring sparse point cloud data of a target object and camera parameters of the target object at a preset observation angle; determining difference features between a plurality of points in the sparse point cloud data and neighborhood points of the points, and performing aggregation processing to obtain local features of the sparse point cloud data; processing the sparse point cloud data based on a cross attention mechanism to obtain global features of the sparse point cloud data; fusing the local features and the global features of the sparse point cloud data to obtain comprehensive features of the sparse point cloud data; obtaining a three-dimensional Gaussian splashing parameter of the sparse point cloud data; and rendering the three-dimensional Gaussian splash parameters of the sparse point cloud data based on the camera parameters, and generating a reconstructed image of the target object under the camera parameters. According to the invention, the accuracy of image reconstruction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method, system, device and medium for generating an image and training an image generation model. Background Art

[0002] The core challenge currently faced by point cloud rendering stems from its discrete nature as a sparse sampling of the object surface: irregular point distribution is prone to holes and blurring artifacts when directly rendered, and the modeling difficulties brought about by sudden changes in surface texture (such as text and geometric edges) further exacerbate the complexity of reconstructing the radiation field (color, material and other attributes). Traditional graphics methods usually rely on rasterization projection or disk / sphere-based splattering technology to alleviate this problem, but are limited by parameter estimation errors and insufficient adaptive modeling capabilities, and their effectiveness is limited. In recent years, neural decoding methods have attempted to project point clouds into two-dimensional feature maps and reconstruct images through neural networks. However, due to the inherited sparsity of point clouds and the lack of real three-dimensional space modeling capabilities, they often result in missing details.

[0003] With the development of neural representation technology, it has continuously evolved from global latent coding, voxel grids to three-plane representations, improving local modeling capabilities. However, it still faces problems such as information loss caused by voxelization or limited three-plane expression capabilities. The neural radiance field rendering method has achieved a breakthrough in visual quality through volume rendering. However, it is still difficult to get rid of the high difficulty of continuous field modeling, the high cost of relying on scene-by-scene optimization, and the trade-off between rendering speed and image quality. As an emerging solution, 3D Gaussian splattering parameterizes the scene with a set of discrete Gaussian kernels and combines it with differentiable rasterization technology to achieve efficient real-time rendering. However, due to its reliance on scene-by-scene optimization and lack of end-to-end forward prediction capabilities, it still limits its practicality in general tasks. Therefore, it is necessary to provide an image generation and image reconstruction method, system, device and medium. Summary of the Invention

[0004] In view of the above shortcomings of the prior art, the purpose of the present invention is to provide a method, system, device and medium for generating an image and training an image generation model, so as to improve the problem of low accuracy of image reconstruction in the prior art.

[0005] To achieve the above-mentioned and other related purposes, the present invention provides an image generation method, comprising: obtaining sparse point cloud data of a target object and camera parameters of the target object at a preset observation angle; inputting the sparse point cloud data into a local feature extraction network of an image generation model, determining differential features between multiple points in the sparse point cloud data and their neighborhood points and performing aggregation processing to obtain local features of the sparse point cloud data; inputting the sparse point cloud data into a global feature extraction network of the image generation model, processing the sparse point cloud data based on a cross-attention mechanism to obtain global features of the sparse point cloud data; inputting the local features and global features of the sparse point cloud data into a fusion network of the image generation model for fusion to obtain comprehensive features of the sparse point cloud data; inputting the comprehensive features of the sparse point cloud data into a prediction network of the image generation model to obtain three-dimensional Gaussian splash parameters of the sparse point cloud data; rendering the three-dimensional Gaussian splash parameters of the sparse point cloud data based on the camera parameters to generate a reconstructed image of the target object under the camera parameters.

[0006] In one embodiment of the present invention, obtaining sparse point cloud data of a target object includes: obtaining initial sparse point cloud data of the target object; and performing denoising on the initial sparse point cloud data to obtain final sparse point cloud data.

[0007] In one embodiment of the present invention, sparse point cloud data is input into a local feature extraction network of an image generation model, and differential features between multiple points in the sparse point cloud data and their neighborhood points are determined and aggregated to obtain local features of the sparse point cloud data, including: selecting multiple points from the sparse point cloud data as target points, and for each selected target point: based on the K-nearest neighbor method, selecting a preset number of neighborhood points closest to the target point from the sparse point cloud data, and calculating the position difference and color difference between each neighborhood point and the target point to obtain differential features between each neighborhood point and the target point; inputting each differential feature into a multi-layer perceptron of the local feature extraction network to obtain corresponding neighborhood embedding features; inputting all neighborhood embedding features into the pooling layer of the local feature extraction network for maximum pooling to obtain local features of the target point; and counting the local features of each target point to obtain local features of the sparse point cloud data.

[0008] In one embodiment of the present invention, sparse point cloud data is input into a global feature extraction network of an image generation model, and the sparse point cloud data is processed based on a cross-attention mechanism to obtain global features of the sparse point cloud data, including: selecting multiple points from the sparse point cloud data as target points, and for each selected target point: performing Fourier position encoding on the position of the target point, and concatenating the encoding result with the color of the target point, and then inputting the result into a linear mapping module of the global feature extraction network to obtain initial features of the target point; inputting the initial features of the target point into a cross-attention module of the global feature extraction network, and interactively processing the initial features and a preset first number of query features based on the cross-attention mechanism to obtain a global feature set of the target point; and statistically analyzing the global feature sets of each target point to obtain global features of the sparse point cloud data.

[0009] In one embodiment of the present invention, the initial features of the target point are input into the cross-attention module of the global feature extraction network, and the cross-attention between the initial features and a preset first number of query features is calculated to obtain a global feature set of the target point, including: inputting the initial features of the target point into the cross-attention layer of the cross-attention module, interactively processing the initial features and the preset first number of query features based on the cross-attention mechanism, and obtaining the first number of initial global features of the target point; inputting the first number of initial global features into the encoding layer of the cross-attention module, capturing the dependency between each initial global feature, and updating the corresponding initial global features accordingly, obtaining the final first number of global features of the target point, and constructing a global feature set; wherein the encoding layer includes a cascaded Transformer and U-Net structure.

[0010] In one embodiment of the present invention, the linear mapping module is a lightweight multi-layer perceptron.

[0011] In one embodiment of the present invention, a training method for an image generation model is also provided, the training method comprising: obtaining sparse point cloud data of a target object, camera parameters of the target object at a preset observation angle, and a real image of the target object under the camera parameters; inputting the sparse point cloud data into a local feature extraction network of the image generation model, determining differential features between multiple points in the sparse point cloud data and their neighborhood points and performing aggregation processing to obtain local features of the sparse point cloud data; inputting the sparse point cloud data into a global feature extraction network of the image generation model, processing the sparse point cloud data based on a cross-attention mechanism, and obtaining The global features of sparse point cloud data; the local features and global features of the sparse point cloud data are input into the fusion network of the image generation model for fusion to obtain the comprehensive features of the sparse point cloud data; the comprehensive features of the sparse point cloud data are input into the prediction network of the image generation model to obtain the three-dimensional Gaussian splash parameters of the sparse point cloud data; the three-dimensional Gaussian splash parameters of the sparse point cloud data are rendered based on the camera parameters to generate a reconstructed image of the target object under the camera parameters; the image reconstruction loss and perceptual loss between the reconstructed image and the real image are calculated, and the parameters of the image generation model are updated accordingly to obtain a trained image generation model.

[0012] In one embodiment of the present invention, an image generation system is also provided, which includes: a data acquisition module for acquiring sparse point cloud data of a target object and camera parameters of the target object at a preset observation angle; a local feature extraction module for inputting the sparse point cloud data into a local feature extraction network of an image generation model, determining differential features between multiple points in the sparse point cloud data and their neighborhood points, and performing aggregation processing to obtain local features of the sparse point cloud data; a global feature extraction module for inputting the sparse point cloud data into a global feature extraction network of the image generation model, processing the sparse point cloud data based on a cross-attention mechanism, and obtaining global features of the sparse point cloud data; a fusion module for inputting the local features and global features of the sparse point cloud data into a fusion network of the image generation model for fusion to obtain comprehensive features of the sparse point cloud data; a parameter prediction module for inputting the comprehensive features of the sparse point cloud data into a prediction network of the image generation model to obtain three-dimensional Gaussian splash parameters of the sparse point cloud data; and a reconstruction module for rendering the three-dimensional Gaussian splash parameters of the sparse point cloud data based on the camera parameters to generate a reconstructed image of the target object under the camera parameters.

[0013] In one embodiment of the present invention, an electronic device is also provided, including: one or more processors; a storage device for storing one or more programs, which, when the one or more programs are executed by one or more processors, enables the electronic device to implement any of the above-mentioned image generation methods or image generation model training methods.

[0014] In one embodiment of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a computer processor, the computer executes any of the above-mentioned image generation methods or image generation model training methods.

[0015] As described above, the present invention provides an image generation and image generation model training method, system, device and medium, which have the following beneficial effects: the sparse point cloud data of the target object is input into the local feature extraction network, the differential features between the points in the sparse point cloud and their neighborhood points are extracted, and local features with spatial structure perception capabilities are formed through aggregation operations, thereby improving the recognition ability of local geometric shapes and texture changes. In addition, through the global feature extraction network, the cross-attention mechanism can be used to extract global features from the sparse point cloud data, aiming to improve the image generation model's ability to model the overall spatial semantics. The comprehensive features obtained by fusing local features and global features have both local fine-grained details and global consistency features, so that under specified camera parameters, the reconstructed image rendered based on the comprehensive features can more realistically restore the geometric shape and appearance texture of the target object at that perspective, greatly improving the accuracy and realism of image reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic flow chart of an image generation method provided by an embodiment of the present invention;

[0017] Figure 2 Shown is a schematic diagram of comprehensive feature extraction provided by an embodiment of the present invention;

[0018] Figure 3 A schematic flow chart showing a method for training an image generation model according to an embodiment of the present invention;

[0019] Figure 4 A schematic diagram showing the effect of reconstructing an image according to the solution provided in an embodiment of the present invention;

[0020] Figure 5 Shown is a structural block diagram of an image generation system provided by an embodiment of the present invention;

[0021] Figure 6 Shown is a structural schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0023] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0024] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0025] The present invention provides an image generation method, which inputs sparse point cloud data of a target object into a local feature extraction network, extracts differential features between points in the sparse point cloud and their neighborhood points, and forms local features with spatial structure perception capabilities through aggregation operations, thereby improving the recognition capability of local geometric forms and texture changes. In addition, through a global feature extraction network, global features in sparse point cloud data can be extracted using a cross-attention mechanism, aiming to improve the image generation model's ability to model overall spatial semantics. The comprehensive features obtained by fusing local features with global features have both local fine-grained details and global consistency features, so that under specified camera parameters, the reconstructed image rendered based on the comprehensive features can more realistically restore the geometric shape and appearance texture of the target object at that perspective, greatly improving the accuracy and realism of image reconstruction.

[0026] like Figure 1 As shown, the image generation method includes the following steps:

[0027] S11. Acquire sparse point cloud data of the target object and camera parameters of the target object at a preset observation angle.

[0028] Sparse point cloud data of the target object can be obtained through depth cameras, LiDAR devices, etc., where the sparse point cloud data includes multiple points, each of which includes the three-dimensional spatial coordinates and color of the point. In addition, the camera parameters of the target object at a preset observation angle must be obtained so that the three-dimensional points can be projected onto the image plane to obtain an image at that observation angle. The camera parameters can include camera intrinsic parameters such as focal length and principal point position, as well as camera extrinsic parameters such as the camera's attitude and position.

[0029] Furthermore, considering that the collected sparse point cloud data may have measurement errors or uneven spatial distribution, resulting in poor accuracy in subsequent data processing, in order to improve the above situation and enhance the quality of point cloud data, optionally, the sparse point cloud data is obtained in the following way: first, the initial sparse point cloud data of the target object is obtained, and then the initial sparse point cloud data is denoised to obtain the final sparse point cloud data.

[0030] S12. Input the sparse point cloud data into the local feature extraction network of the image generation model, determine the differential features between multiple points in the sparse point cloud data and their neighborhood points, and perform aggregation processing to obtain local features of the sparse point cloud data; wherein the differential features include position differences and color differences.

[0031] The sparse point cloud data is fed into a local feature extraction network. Multiple points are selected from the sparse point cloud. For each selected point, a K-nearest neighbor algorithm is used to find several spatially adjacent points. The differential features between the point and each of these neighboring points are then calculated. These differential features include positional and color differences between the point and its neighbors. All of these differential features are fed into a multi-layer perceptron, where maximum pooling is performed to extract the local features of the point. The local features of all selected points are aggregated to form the local features of the sparse point cloud data.

[0032] Specifically, step S12 includes the following process:

[0033] Select multiple points from the sparse point cloud data as target points, and perform the following processing for each selected target point:

[0034] First, based on the K-nearest neighbor method, a preset number of neighborhood points closest to the target point are selected from the sparse point cloud data, and the position difference and color difference between each neighborhood point and the target point are calculated respectively to obtain the differential features between each neighborhood point and the target point.

[0035] Based on the K-nearest neighbor algorithm, the Euclidean distance between the current target point and each remaining point in the sparse point cloud data is calculated, and the K points closest to it are selected as neighborhood points. These neighborhood points together constitute a local receptive region centered on the target point, which can be used to describe the distribution of relevant features in the local space around the target point. For each selected neighborhood point, the difference between the neighborhood point and the target point in the three-dimensional coordinate space is calculated to obtain the position difference, and the difference between the neighborhood point and the target point in the RGB color space is calculated to obtain the color difference. The position difference and color difference are spliced ​​together to obtain the differential features between the neighborhood point and the target point.

[0036] After obtaining the differential features, each differential feature is input into the multi-layer perceptron of the local feature extraction network to obtain the corresponding neighborhood embedding features.

[0037] Specifically, the following process is performed on the differential features between the target point and each neighboring point: the differential features are input into a multi-layer perceptron for feature transformation. Through multi-layer feature mapping, the corresponding neighborhood embedding features between the target point and the current neighboring point are obtained. The neighborhood embedding features are used to represent the comprehensive differences in spatial position and color information of the neighboring point relative to the target point.

[0038] All neighborhood embedding features are input into the pooling layer of the local feature extraction network for maximum pooling to obtain the local features of the target point.

[0039] All neighborhood embedding features of the current target point are input into the pooling layer together, and all neighborhood embedding features are aggregated through the maximum pooling operation, so that the eigenvalues ​​with the strongest response in each dimension can be extracted to obtain the local features of the current target point.

[0040] S13. Input the sparse point cloud data into the global feature extraction network of the image generation model, process the sparse point cloud data based on the cross-attention mechanism, and obtain the global features of the sparse point cloud data.

[0041] Step S13 includes the following process:

[0042] Select multiple points from the sparse point cloud data as target points, and perform the following process for each selected target point:

[0043] First, the position of the target point is Fourier encoded, and the encoding result is spliced ​​with the color of the target point, and then input into the linear mapping module of the global feature extraction network to obtain the initial features of the target point.

[0044] Each point in the sparse point cloud data includes two attribute parameters: three-dimensional spatial position and color information. For the target point, the three-dimensional spatial position of the target point in the entire sparse point cloud data can be Fourier encoded, thereby mapping the original position coordinates of the target point to a high-dimensional periodic vector. The encoded result is spliced ​​with the color information of the target point to obtain a composite feature containing spatial and color information. The composite feature of the target point obtained after splicing is input into the linear mapping module, and the initial feature of the current target point is obtained through feature mapping, which can be expressed as q(P0)=MLP q (f o ), q(P0) is the initial feature of the target point P0, MLP q (f o ) is the linear mapping module MLP q For the composite feature f o The linear mapping module may be a multi-layer perceptron or a plurality of linear layers. Preferably, in order to increase the processing speed, the linear mapping module is a lightweight multi-layer perceptron.

[0045] The initial features of the target point are then input into the cross-attention module of the global feature extraction network. Based on the cross-attention mechanism, the initial features and the preset first number of query features are interactively processed to obtain the global feature set of the target point.

[0046] Specifically, for each query feature z i Perform the following process: respectively query feature z i The first and second multi-layer perceptrons input to the cross attention module obtain the key vector and value vector, as shown in formula (1) (2):

[0047] k(z i )=MLP k (z i ) (1)

[0048] v(z i )=MLP v (z i ) (2)

[0049] Among them, k(z i ) is the key vector of the i-th query feature, v(z i ) is the value vector of the i-th query feature, MLP k (z i ) is the first multi-layer perceptron processing the i-th query feature, MLP v (z i ) is the second multi-layer perceptron that processes the i-th query feature. And the i-th query feature z is calculated using formula (3) iThe attention weight α for the initial feature q(p0) of the target point p0 i :

[0050]

[0051] Where d is the dimension of the query feature, and L is the number of query features (i.e., the first number). After the calculation is completed, the global feature F of the target point p0 is generated by weighted summation. GLS (P o ), as shown in formula (4):

[0052]

[0053] Furthermore, there are multiple attention heads in the cross-attention module, and each attention head is processed in the above manner to obtain a global feature. The global features obtained by all attention heads are summarized to serve as the global feature set of the target point.

[0054] In order to improve the expressive power of global features and enable sparse point cloud data to obtain structurally complete and semantically rich global context support, the Transformer and U-Net structures can be optionally used to optimize the global features. Specifically, the global feature set of the target point is obtained by the following method:

[0055] First, the initial features of the target point are input into the cross-attention layer of the cross-attention module. Based on the cross-attention mechanism, the initial features and the preset first number of query features are interactively processed to obtain multiple initial global features of the target point.

[0056] The cross attention layer includes a preset second number of attention heads, each of which obtains an initial global feature using the above formula (4). After all attention heads are processed according to the above process, the second number of initial global features of the target point are obtained.

[0057] Then, multiple initial global features are input into the encoding layer of the cross-attention module to capture the dependencies between the initial global features, and the corresponding initial global features are updated accordingly to obtain the final multiple global features of the target point and construct a global feature set; among them, the encoding layer includes a cascaded Transformer and U-Net structure.

[0058] The second number of initial global features is input into the encoding layer. The Transformer's multi-head self-attention mechanism captures the dependencies between the initial global features, resulting in a second number of intermediate global features. These second number of intermediate global features are then passed through a U-Net structure using hierarchical encoding and decoding to further enhance the structural semantic information of these global features at different scales. This results in the generation of a second number of final global features, ultimately forming the global feature set for the target point.

[0059] The global feature set of each target point is counted to obtain the global features of the sparse point cloud data.

[0060] It is understandable that in the present invention, several points can be selected from the sparse point cloud data as target points based on preset selection rules to speed up the processing speed of the entire solution. In order to improve the accuracy of subsequent image generation, each point in the sparse point cloud data can be used as a target point. The specific method is not limited here.

[0061] S14. Input the local features and global features of the sparse point cloud data into the fusion network of the image generation model for fusion to obtain the comprehensive features of the sparse point cloud data.

[0062] For each target point in the sparse point cloud obtained above, the following process is performed: the global and local features of the target point are fused by concatenation or element-by-element weighting to obtain the comprehensive features of the target point. After all target points have been processed according to the above process, the comprehensive features of each target point are summarized to obtain the comprehensive features of the sparse point cloud data.

[0063] like Figure 2 As shown, for the target point P o , whose kth neighborhood point includes position x k and color c k Two attributes, by calculating the differential features between the target point and each neighboring point, and using maximum pooling to obtain the local features of the target point. Through a set of learnable query features Each target point is subjected to cross-attention interaction with each query feature to obtain the global feature of the target point. The comprehensive feature of the target point is obtained by fusing the local features and the global features.

[0064] S15. Input the comprehensive features of the sparse point cloud data into the prediction network of the image generation model to obtain the three-dimensional Gaussian splash parameters of the sparse point cloud data.

[0065] The comprehensive features of each point in the sparse point cloud data are input into the prediction network of the image generation model. This prediction network consists of multiple parallel output branches, each of which can be a multi-layer perceptron. Each branch predicts a type of three-dimensional Gaussian splash parameter for the point. The three-dimensional Gaussian splash parameters include the point's center position, color, opacity, scaling vector, and rotation quaternion. These parameters can be used to represent the corresponding point in the sparse point cloud data as a three-dimensional Gaussian kernel with directionality and range perception. The center position is used to define the position offset of the Gaussian kernel in space, the scaling vector is used to control the size of the Gaussian shape, and the rotation quaternion is used to define the Gaussian direction.

[0066] S16. Rendering the three-dimensional Gaussian splash parameters of the sparse point cloud data based on the camera parameters to generate a reconstructed image of the target object under the camera parameters.

[0067] The following process is performed for the three-dimensional Gaussian kernel of each point in the sparse point cloud data: the center position of the three-dimensional Gaussian kernel is projected onto the image plane using the camera parameters, and the Gaussian influence area of ​​the Gaussian kernel in the two-dimensional image is determined using the scaling vector and rotation quaternion of the three-dimensional Gaussian kernel to obtain the Gaussian projection value of the Gaussian kernel. After all points in the sparse point cloud data are processed as described above, the image space is initialized to a blank pixel grid. The projection area of ​​each Gaussian point may cover multiple pixel positions in the image. For each pixel, all Gaussian projections covering the pixel are collected, and the final color value of the pixel is calculated by α blending based on the color and transparency, as shown in formula (5):

[0068]

[0069] Among them, C is the color value of the pixel in the final output image, c n is the color of the nth Gaussian point, α' n , α′ j are the blending weights of the nth and jth Gaussian points at the current pixel, respectively. These weights are equal to the product of the opacity of the Gaussian point and the corresponding two-dimensional Gaussian projection value. N is the total number of Gaussian points that affect the pixel, sorted by depth. This processing is repeated for all pixels, gradually generating a complete reconstructed image. This approach can generate reconstructed images for any camera parameters.

[0070] See Figure 3 The present invention also provides a training method for an image generation model, the training method comprising the following steps:

[0071] S31 , obtaining sparse point cloud data of the target object, camera parameters of the target object at a preset observation angle, and a real image of the target object under the camera parameters.

[0072] S32. Input the sparse point cloud data into the local feature extraction network of the image generation model, determine the differential features between multiple points in the sparse point cloud data and their neighborhood points, and perform aggregation processing to obtain local features of the sparse point cloud data.

[0073] S33. Input the sparse point cloud data into the global feature extraction network of the image generation model, process the sparse point cloud data based on the cross-attention mechanism, and obtain the global features of the sparse point cloud data.

[0074] S34. Input the local features and global features of the sparse point cloud data into the fusion network of the image generation model for fusion to obtain the comprehensive features of the sparse point cloud data.

[0075] S35. Input the comprehensive features of the sparse point cloud data into the prediction network of the image generation model to obtain the three-dimensional Gaussian splash parameters of the sparse point cloud data.

[0076] S36. Rendering the three-dimensional Gaussian splash parameters of the sparse point cloud data based on the camera parameters to generate a reconstructed image of the target object under the camera parameters.

[0077] S37. Calculate the image reconstruction loss and the perceptual loss between the reconstructed image and the real image, and update the parameters of the image generation model accordingly to obtain a trained image generation model.

[0078] The generation process of the reconstructed image is consistent with the above and will not be described in detail here. After obtaining the reconstructed image, the image reconstruction loss and the perceptual loss can be calculated based on the difference between the reconstructed image and the real image acquired under the camera parameters. Among them, the image reconstruction loss may include a weighted combination of the L1 loss and the structural similarity (D-SSIM, Differentiable Structural Similarity Index) loss between the predicted reconstructed image and the real image. The perceptual loss is to extract features from the reconstructed image and the real image respectively through a pre-trained perceptual network to obtain reconstructed features and real features, and calculate the difference between the reconstructed features and the real features in the semantic space, which is used as the perceptual loss. Among them, the perceptual network includes but is not limited to VGG or AlexNet, etc., as long as it can extract multi-layer semantic features from the image. The image reconstruction loss and the perceptual loss are weighted to obtain the total loss, and the parameters of the image generation model are updated according to the above total loss to obtain a trained image generation model.

[0079] like Figure 4As shown in the figure, the left side is the sparse point cloud data of the input target object, such as a horse, a child in yellow, or a triceratops. As can be seen from the figure, the point cloud is very sparse, the color is incomplete, the details are blurred, and only the general shape and texture distribution are retained. The middle is the reconstructed image after rendering using this method. It can be seen from the figure that this scheme can realize the reconstructed image of any perspective, for example, it can present an image with continuously changing multiple perspectives, and the reconstructed image has a high sense of reality. Therefore, this scheme can successfully realize the mapping from sparse point cloud to high-quality image. The right side is the three-dimensional mesh model obtained by the image reconstruction algorithm. It can be seen that this scheme can also be used to obtain the three-dimensional mesh model corresponding to the target object from the sparse point cloud data, and the mesh result has a complete shape and good details, which can be used for subsequent three-dimensional reconstruction tasks.

[0080] like Figure 5 As shown, the image generation system 500 includes: a data acquisition module 510, a local feature extraction module 520, a global feature extraction module 530, a fusion module 540, a parameter prediction module 550, and a reconstruction module 560. The data acquisition module 510 is used to obtain sparse point cloud data of the target object and the camera parameters of the target object at a preset observation angle. The local feature extraction module 520 is used to input the sparse point cloud data into the local feature extraction network of the image generation model, determine the differential features between multiple points in the sparse point cloud data and their neighboring points, and perform aggregation processing to obtain local features of the sparse point cloud data. The global feature extraction module 530 is used to input the sparse point cloud data into the global feature extraction network of the image generation model, process the sparse point cloud data based on the cross-attention mechanism, and obtain global features of the sparse point cloud data. The fusion module 540 is used to input the local features and global features of the sparse point cloud data into the fusion network of the image generation model for fusion, to obtain comprehensive features of the sparse point cloud data. The parameter prediction module 550 is used to input the comprehensive features of the sparse point cloud data into the prediction network of the image generation model to obtain the 3D Gaussian splatter parameters of the sparse point cloud data. The reconstruction module 560 is used to render the 3D Gaussian splatter parameters of the sparse point cloud data based on the camera parameters to generate a reconstructed image of the target object under the camera parameters.

[0081] The specific definition of the image generation system can be found in the definition of the image generation method above and will not be repeated here. Each module in the above-mentioned image generation system can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware format, or can be stored in the memory of the computer device in software format to facilitate the processor to call the corresponding operations of each of the above modules.

[0082] It should be noted that, in order to highlight the innovative part of the present invention, this embodiment does not introduce modules that are not closely related to solving the technical problems raised by the present invention, but this does not mean that there are no other modules in this embodiment.

[0083] like Figure 6 As shown, the electronic device 6 may include a memory 61 , a processor 62 and a bus, and may also include a computer program stored in the memory 61 and executable on the processor 62 , such as an image generation program.

[0084] Among them, the memory 61 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 61 can be an internal storage unit of the electronic device 6, such as a mobile hard disk of the electronic device 6. In other embodiments, the memory 61 can also be an external storage device of the electronic device 6, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. equipped on the electronic device 6. Furthermore, the memory 61 can also include both an internal storage unit of the electronic device 6 and an external storage device. The memory 61 can not only be used to store application software installed in the electronic device 6 and various types of data, such as code for generating images, but can also be used to temporarily store data that has been output or is to be output.

[0085] In some embodiments, the processor 62 may be comprised of an integrated circuit, such as a single packaged integrated circuit or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 62 is the control core (Control Unit) of the electronic device 6, connecting the various components of the entire electronic device 6 using various interfaces and circuits. It executes or runs programs or modules (such as image generation programs) stored in the memory 61 and accesses data stored in the memory 61 to perform various functions of the electronic device 6 and process data.

[0086] The processor 62 executes the operating system and various installed application programs of the electronic device 6. The processor 62 executes the application programs to implement the steps in the above-mentioned image generation method.

[0087] Exemplarily, the computer program may be divided into one or more modules, one or more of which are stored in the memory 61 and executed by the processor 62 to complete the present application. One or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 6. For example, the computer program may be divided into a data acquisition module 510, a local feature extraction module 520, a global feature extraction module 530, a fusion module 540, a parameter prediction module 550, and a reconstruction module 560.

[0088] The above-mentioned integrated unit implemented in the form of a software functional module can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, computer device, or network device, etc.) or a processor to perform part of the functions of the image generation method of each embodiment of the present application.

[0089] In summary, the present invention discloses a method, system, device and medium for generating an image and training an image generation model, which inputs the sparse point cloud data of the target object into a local feature extraction network, extracts the differential features between the points in the sparse point cloud and their neighborhood points, and forms local features with spatial structure perception capabilities through aggregation operations, thereby improving the recognition ability of local geometric shapes and texture changes. In addition, through the global feature extraction network, the cross-attention mechanism can be used to extract global features from the sparse point cloud data, aiming to improve the image generation model's ability to model the overall spatial semantics. The comprehensive features obtained by fusing local features and global features have both local fine-grained details and global consistency features, so that under specified camera parameters, the reconstructed image rendered based on the comprehensive features can more realistically restore the geometric shape and appearance texture of the target object at that perspective, greatly improving the accuracy and realism of image reconstruction. Therefore, the present invention effectively overcomes the various shortcomings of the prior art and has high industrial utilization value.

[0090] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.

Claims

1. An image generation method, characterized in that: The generation method comprises: Acquire sparse point cloud data of a target object and camera parameters of the target object at a preset observation angle; Inputting the sparse point cloud data into a local feature extraction network of an image generation model, determining differential features between a plurality of points in the sparse point cloud data and their neighboring points and performing aggregation processing to obtain local features of the sparse point cloud data; Inputting the sparse point cloud data into the global feature extraction network of the image generation model, processing the sparse point cloud data based on the cross attention mechanism, and obtaining the global features of the sparse point cloud data; Inputting the local features and global features of the sparse point cloud data into the fusion network of the image generation model for fusion to obtain comprehensive features of the sparse point cloud data; Inputting the comprehensive features of the sparse point cloud data into the prediction network of the image generation model to obtain three-dimensional Gaussian splatter parameters of the sparse point cloud data; The three-dimensional Gaussian splash parameters of the sparse point cloud data are rendered based on the camera parameters to generate a reconstructed image of the target object under the camera parameters.

2. The image generation method according to claim 1, wherein: The step of obtaining sparse point cloud data of the target object includes: Acquiring initial sparse point cloud data of the target object; The initial sparse point cloud data is subjected to denoising processing to obtain final sparse point cloud data.

3. The image generation method according to claim 1, wherein: Inputting the sparse point cloud data into a local feature extraction network of an image generation model, determining differential features between a plurality of points in the sparse point cloud data and their neighboring points and performing aggregation processing to obtain local features of the sparse point cloud data, includes: Select multiple points from the sparse point cloud data as target points, and for each selected target point: Based on the K-nearest neighbor method, a preset number of neighborhood points closest to the target point are selected from the sparse point cloud data, and the position difference and color difference between each neighborhood point and the target point are calculated to obtain the differential features between each neighborhood point and the target point; Input each differential feature into the multi-layer perceptron of the local feature extraction network to obtain the corresponding neighborhood embedding feature; Input all neighborhood embedded features into the pooling layer of the local feature extraction network for maximum pooling to obtain the local features of the target point; The local features of each target point are counted to obtain the local features of the sparse point cloud data.

4. The image generation method according to claim 1, wherein: Inputting the sparse point cloud data into the global feature extraction network of the image generation model, processing the sparse point cloud data based on the cross attention mechanism to obtain the global features of the sparse point cloud data, includes: Select multiple points from the sparse point cloud data as target points, and for each selected target point: Performing Fourier position encoding on the position of the target point, and concatenating the encoding result with the color of the target point, inputting the result into the linear mapping module of the global feature extraction network to obtain the initial features of the target point; Inputting the initial features of the target point into the cross attention module of the global feature extraction network, interactively processing the initial features and a preset first number of query features based on the cross attention mechanism to obtain a global feature set of the target point; The global feature set of each target point is counted to obtain the global feature of the sparse point cloud data.

5. The image generation method according to claim 4, characterized in that The initial features of the target point are input into the cross attention module of the global feature extraction network, and the cross attention between the initial features and a preset first number of query features is calculated to obtain a global feature set of the target point, including: Inputting the initial features of the target point into the cross attention layer of the cross attention module, interactively processing the initial features and a preset first number of query features based on the cross attention mechanism, and obtaining a first number of initial global features of the target point; The first number of initial global features are input into the encoding layer of the cross-attention module to capture the dependencies between the initial global features, and the corresponding initial global features are updated accordingly to obtain the final first number of global features of the target point and construct a global feature set; wherein the encoding layer includes a cascaded Transformer and U-Net structure.

6. The image generation method according to claim 4, characterized in that The linear mapping module is a lightweight multi-layer perceptron.

7. A training method for an image generation model, characterized in that: The training method comprises: Acquire sparse point cloud data of a target object, camera parameters of the target object at a preset observation angle, and a real image of the target object under the camera parameters; Inputting the sparse point cloud data into a local feature extraction network of an image generation model, determining differential features between a plurality of points in the sparse point cloud data and their neighboring points and performing aggregation processing to obtain local features of the sparse point cloud data; Inputting the sparse point cloud data into the global feature extraction network of the image generation model, processing the sparse point cloud data based on the cross attention mechanism, and obtaining the global features of the sparse point cloud data; Inputting the local features and global features of the sparse point cloud data into the fusion network of the image generation model for fusion to obtain comprehensive features of the sparse point cloud data; Inputting the comprehensive features of the sparse point cloud data into the prediction network of the image generation model to obtain three-dimensional Gaussian splatter parameters of the sparse point cloud data; Rendering the three-dimensional Gaussian splash parameters of the sparse point cloud data based on the camera parameters to generate a reconstructed image of the target object under the camera parameters; The image reconstruction loss and the perceptual loss between the reconstructed image and the real image are calculated, and the parameters of the image generation model are updated accordingly to obtain a trained image generation model.

8. An image generation system, characterized in that: The system comprises: A data acquisition module, configured to acquire sparse point cloud data of a target object and camera parameters of the target object at a preset observation angle; a local feature extraction module, configured to input the sparse point cloud data into a local feature extraction network of an image generation model, determine differential features between a plurality of points in the sparse point cloud data and their neighboring points, and perform aggregation processing to obtain local features of the sparse point cloud data; a global feature extraction module, configured to input the sparse point cloud data into a global feature extraction network of the image generation model, process the sparse point cloud data based on a cross-attention mechanism, and obtain global features of the sparse point cloud data; A fusion module, configured to input the local features and global features of the sparse point cloud data into a fusion network of the image generation model for fusion, thereby obtaining comprehensive features of the sparse point cloud data; a parameter prediction module, configured to input the comprehensive features of the sparse point cloud data into a prediction network of the image generation model to obtain three-dimensional Gaussian splatter parameters of the sparse point cloud data; A reconstruction module is used to render the three-dimensional Gaussian splash parameters of the sparse point cloud data based on the camera parameters to generate a reconstructed image of the target object under the camera parameters.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the image generation method according to any one of claims 1 to 6 or the training method of the image generation model according to claim 7.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the image generation method according to any one of claims 1 to 6 or the image generation model training method according to claim 7.