Three-dimensional image generation model training method and device, equipment and medium

By processing the original image generation model from a sparse perspective, extracting semantic features and performing regularized training in comparison learning, the poor performance of the three-dimensional image generation model caused by limited acquisition equipment is solved, and efficient three-dimensional image generation and accuracy improvement are achieved.

CN120279170APending Publication Date: 2025-07-08PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279100.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, due to the limited number of acquisition devices, it is difficult to obtain high-quality sample data, resulting in poor performance of the three-dimensional image generation model obtained by training, which in turn affects its target three-dimensional image accuracy during actual application.

Method used

By obtaining the original image from the sparse perspective, generating a frame image, and performing motion structure recovery processing to obtain the initial Gaussian point set, performing Gaussian optimization processing to obtain the coarse Gaussian point set, extracting semantic features from the sparse perspective and the new perspective, performing comparison learning regularization processing, and training the initial three-dimensional image generation model based on the optimization loss value and the comparison loss value until the preset conditions are reached.

Benefits of technology

The performance of the three-dimensional image generation model is improved, the model's ability to distinguish similar and dissimilar samples is enhanced, the computing resource consumption is reduced, and the accuracy of the target three-dimensional image during actual application is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279170A_ABST
    Figure CN120279170A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a training method and device for a three-dimensional image generation model, equipment and a medium, and belongs to the technical field of three-dimensional modeling. The method comprises the following steps: generating an adjacent frame image corresponding to an original image under a sparse view angle; performing motion structure recovery processing on the original image and the adjacent frame image to obtain an initial Gaussian point set, performing Gaussian optimization processing on the initial Gaussian point set to obtain a coarse Gaussian point set, and determining an optimization loss value based on the coarse Gaussian point set; performing comparative learning regularization processing based on a first semantic feature under a sparse view angle extracted from the coarse Gaussian point set and a second semantic feature under a new view angle, and determining a comparative loss value; and training the initial three-dimensional image generation model based on the optimization loss value and the comparison loss value to obtain a target three-dimensional image generation model. According to the invention, the performance of the three-dimensional image generation model obtained through training can be improved, and the accuracy of the target three-dimensional image output by the model in actual application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of 3D modeling technology, and particularly to a training method, device, equipment and medium for a 3D image generation model. Background Art

[0002] A 3D image generation model is a type of model based on artificial intelligence and deep learning technology. This model can automatically generate a 3D geometric model or a 3D rendered image from input data (such as 2D images, point clouds, text descriptions, etc.) through 3D reconstruction. The generated target 3D image digitally restores the 3D structure and appearance features of the corresponding object or environment in the real world.

[0003] In related technologies, sample data is obtained based on images collected by multiple acquisition devices, and a 3D image generation model is trained based on these sample data. However, in actual situations, the number of acquisition devices is very limited, so it is often difficult to obtain high-quality sample data, resulting in poor performance of the trained 3D image generation model, and further causing low accuracy of the target 3D image output by the model in actual use. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a training method, device, equipment and medium for a 3D image generation model, aiming to improve the performance of the trained 3D image generation model, and further improve the accuracy of the target 3D image output by the model in actual use.

[0005] To achieve the above object, a first aspect of the embodiments of the present application proposes a training method for a 3D image generation model, including:

[0006] Obtain the original image under a sparse view, and input the original image into the initial 3D image generation model to generate the adjacent frame image corresponding to the original image;

[0007] Perform motion structure recovery processing on the original image and the adjacent frame image to obtain an initial Gaussian point set, perform Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, and determine an optimization loss value based on the rough Gaussian point set;

[0008] Extract the first semantic feature under the sparse view and the second semantic feature under a new view from the rough Gaussian point set, where the new view is different from the sparse view;

[0009] Perform contrast learning regularization processing based on the first semantic feature and the second semantic feature to determine a contrast loss value;

[0010] Train the initial 3D image generation model based on the optimization loss value and the contrast loss value until a preset training condition is reached, and obtain the trained target 3D image generation model.

[0011] In some embodiments, the initial Gaussian point set includes a plurality of Gaussian basis elements;

[0012] Performing Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, including:

[0013] For any Gaussian basis element, count the number of neighboring Gaussian basis elements within a spherical region centered on the Gaussian basis element and with a preset value as the radius;

[0014] If the number of neighboring Gaussian basis elements is less than a preset basis element threshold, remove the Gaussian basis element from the initial Gaussian point set to obtain a rough Gaussian point set.

[0015] In some embodiments, determining an optimization loss value based on the rough Gaussian point set, including:

[0016] Performing parsing processing on the adjacent frame image to obtain an extracted feature point set and an extracted association graph set;

[0017] Performing semantic rendering processing on the rough Gaussian point set to obtain a semantic feature point set and a semantic association graph set of the adjacent frame image;

[0018] Performing convolution processing on the semantic feature point set and the semantic association graph set respectively to obtain a first convolution result and a second convolution result;

[0019] Determining a first feature loss value based on the first convolution result and the extracted feature point set, and determining a second feature loss value based on the second convolution result and the extracted association graph set;

[0020] Superposing the first feature loss value and the second feature loss value to obtain an optimization loss value.

[0021] In some embodiments, performing contrast learning regularization processing based on the first semantic feature and the second semantic feature to determine a contrast loss value, including:

[0022] Distinguishing a first positive semantic feature and a first negative semantic feature from the first semantic feature, and performing contrast learning regularization processing from a sparse perspective based on the first positive semantic feature and the first negative semantic feature to obtain a first contrast loss value;

[0023] Distinguishing a second positive semantic feature and a second negative semantic feature from the second semantic feature, and performing contrast learning regularization processing from a new perspective based on the second positive semantic feature and the second negative semantic feature to obtain a second contrast loss value;

[0024] Superposing the first contrast loss value and the second contrast loss value to obtain a contrast loss value.

[0025] In some embodiments, after obtaining the second contrast loss value, further including:

[0026] Perform color rendering processing on the set of coarse Gaussian points to determine the color features from a sparse perspective;

[0027] Distinguish positive color features and negative color features from the color features, and perform contrast learning regularization processing from a sparse perspective based on the positive color features and negative color features to obtain a color contrast loss value;

[0028] Update the first contrast loss value using the color contrast loss value, and superimpose the updated first contrast loss value and the second contrast loss value to obtain a contrast loss value.

[0029] In some embodiments, the training conditions include a first training condition and a second training condition;

[0030] Train the initial three-dimensional image generation model based on the optimization loss value and the contrast loss value until a preset training condition is reached to obtain a trained target three-dimensional image generation model, including:

[0031] Determine an initial loss value based on the initial set of Gaussian points, where the initial loss value includes an initial color loss value and an initial semantic loss value;

[0032] Perform first-stage joint training on the initial three-dimensional image generation model based on the initial loss value and the optimization loss value until a preset first training condition is reached to obtain an intermediate reconstruction model;

[0033] Perform second-stage joint training on the initial three-dimensional image generation model based on the initial loss value and the contrast loss value until a preset second training condition is reached to obtain a trained target three-dimensional image generation model.

[0034] In some embodiments, after obtaining the trained target three-dimensional image generation model, it further includes:

[0035] Obtain a target image from a sparse perspective;

[0036] Input the target image into the trained target three-dimensional image generation model to output a corresponding target three-dimensional image of the target image.

[0037] To achieve the above object, a second aspect of the embodiments of the present application proposes a training device for a three-dimensional image generation model, including:

[0038] An acquisition module, which acquires an original image from a sparse perspective and inputs the original image into the initial three-dimensional image generation model to generate a neighboring frame image corresponding to the original image;

[0039] An optimized loss value determination module is used to perform motion structure recovery processing on the original image and the adjacent frame image to obtain an initial Gaussian point set, perform Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, and determine the optimized loss value based on the rough Gaussian point set;

[0040] A semantic feature extraction module is used to respectively extract a first semantic feature from the rough Gaussian point set from a sparse perspective and a second semantic feature from a new perspective, where the new perspective is different from the sparse perspective;

[0041] A contrast loss value determination module is used to perform contrast learning regularization processing based on the first semantic feature and the second semantic feature to determine the contrast loss value;

[0042] A target training module is used to train the initial three-dimensional image generation model based on the optimized loss value and the contrast loss value until a preset training condition is reached, and obtain the trained target three-dimensional image generation model.

[0043] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the training method of the three-dimensional image generation model in the first aspect above.

[0044] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the training method of the three-dimensional image generation model in the first aspect above.

[0045] Training method, device, equipment and medium for three-dimensional image generation model proposed in this application. It obtains original images from sparse viewpoints and inputs the original images into an initial three-dimensional image generation model to generate corresponding adjacent frame images of the original images; performs motion structure recovery processing on the original images and adjacent frame images to obtain an initial Gaussian point set, and performs Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, determines an optimization loss value based on the rough Gaussian point set. The adjacent frame images enrich the information for training the initial three-dimensional image generation model, laying a foundation for generating a high-performance target three-dimensional image generation model in the future; then, respectively extracts a first semantic feature from the sparse viewpoint and a second semantic feature from a new viewpoint from the rough Gaussian point set, where the new viewpoint is different from the sparse viewpoint; performs contrast learning regularization processing based on the first semantic feature and the second semantic feature to determine a contrast loss value; trains the initial three-dimensional image generation model based on the optimization loss value and the contrast loss value until a preset training condition is reached to obtain a trained target three-dimensional image generation model. The optimization loss value is used to measure the quality of the obtained rough Gaussian point set after optimization, ensuring that the generated rough Gaussian point set is as close as possible to the actual target object or scene; while the contrast loss value is used to measure the ability of the initial three-dimensional image generation model to distinguish between similar and dissimilar samples, thereby improving the performance of the finally trained target three-dimensional image generation model. At the same time, the training of the initial three-dimensional image generation model can be completed based on a small number of original images, reducing the consumption of computer resources while improving the training effect, and further improving the accuracy of the target three-dimensional image output by the finally generated model in actual application. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 FIG. is a schematic diagram of an application scenario of a training device for a three-dimensional image generation model provided by an embodiment of the present application;

[0047] Figure 2 FIG. is an alternative flowchart of a training method for a three-dimensional image generation model provided by an embodiment of the present application;

[0048] Figure 3 FIG. is an alternative schematic diagram of original image acquisition of a training method for a three-dimensional image generation model provided by an embodiment of the present application;

[0049] Figure 4 FIG. is an alternative schematic diagram of a sphere region of a training method for a three-dimensional image generation model provided by an embodiment of the present application;

[0050] Figure 5 FIG. is an alternative schematic diagram of model training and application of a training method for a three-dimensional image generation model provided by an embodiment of the present application;

[0051] Figure 6It is an optional module schematic diagram of the training device for the three-dimensional image generation model provided by the embodiments of the present application;

[0052] Figure 7 It is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0053] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.

[0054] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart in the flowchart. Terms such as "first" and "second" in the specification, claims and the above drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application, and are not intended to limit this application.

[0056] First, several nouns involved in the present application are analyzed:

[0057] Artificial Intelligence (AI) is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.

[0058] Three-dimensional reconstruction refers to the process of generating a three-dimensional geometric model of an object or scene through computer technology from two-dimensional images, point cloud data or information collected by other sensors. Its goal is to digitize the three-dimensional space information in the real world for visualization, analysis and interaction in the computer.

[0059] A three-dimensional image generation model is a type of model based on artificial intelligence and deep learning technologies. This model can automatically generate a three-dimensional geometric model or a three-dimensional rendered image from input data (such as two-dimensional images, point clouds, text descriptions, etc.) through three-dimensional reconstruction. The generated target three-dimensional image digitally restores the three-dimensional structure and appearance features of the corresponding object or scene in the real world.

[0060] In related technologies, sample data is obtained based on images collected by multiple acquisition devices, and a three-dimensional image generation model is trained based on these sample data. However, in actual situations, the number of acquisition devices is very limited, so it is often difficult to obtain high-quality sample data, resulting in poor performance of the trained three-dimensional image generation model, and further causing low accuracy of the target three-dimensional image output by this model during actual use.

[0061] Based on this, the embodiments of the present application provide a training method, device, equipment, and medium for a three-dimensional image generation model, aiming to improve the performance of the trained three-dimensional image generation model, and further improve the accuracy of the target three-dimensional image output by this model during actual use.

[0062] Exemplarily, as Figure 1 shown, Figure 1 is a schematic diagram of the application scenario of the training device for the three-dimensional image generation model provided by the embodiments of the present application. In an optional application scenario, the client 11 is communicatively connected to the server 12, and the training device for the three-dimensional image generation model proposed by the embodiments of the present application is deployed in the server 12: The client 11 sends the original image under a sparse view to the server 12. After the training device for the three-dimensional image generation model set in the server 12 obtains the original image under the sparse view, it inputs the original image into the initial three-dimensional image generation model to generate a neighboring frame image corresponding to the original image; then, perform motion structure recovery processing on the original image and the neighboring frame image to obtain an initial Gaussian point set, and perform Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, and determine an optimization loss value based on the rough Gaussian point set; then, respectively extract the first semantic feature under the sparse view and the second semantic feature under the new view from the rough Gaussian point set, where the new view is different from the sparse view; after that, perform contrast learning regularization processing based on the first semantic feature and the second semantic feature to determine a contrast loss value; finally, train the initial three-dimensional image generation model based on the optimization loss value and the contrast loss value until the preset training conditions are reached, and obtain the trained target three-dimensional image generation model. During the actual use process, the server 12 performs three-dimensional scene reconstruction processing on the target image based on the trained target three-dimensional image generation model to efficiently generate the target three-dimensional image, where the target three-dimensional image highly accurately restores the target object or scene represented by the target image.

[0063] It should be noted that in the embodiments of the present application, when it comes to information related to user characteristics such as user basic information or user identity, user permission or consent will be obtained first. Moreover, the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of users, separate permission or separate consent of the users will be obtained first. After clearly obtaining the separate permission or separate consent of the users, the necessary data for the normal operation of the embodiments of the present application will be obtained. For example, before obtaining the original image from a sparse perspective, the consent of the relevant regulatory personnel to which the original image belongs will be obtained first. Otherwise, the original image that cannot be applied to the embodiments of the present application will be obtained. In addition, other relevant data obtained by the training device of the three-dimensional image generation model in the embodiments of the present application are all authorized data after obtaining the consent of relevant personnel, which will not be elaborated here one by one.

[0064] In the embodiments of the present application, a description will be made from the perspective of the training device of the three-dimensional image generation model, and the training device of the three-dimensional image generation model can be integrated in a computer device, such as a server. As Figure 2 shown, Figure 2 is an optional flowchart of the training method of the three-dimensional image generation model provided by the embodiments of the present application. Figure 2 The method in Figure 2 may include but is not limited to the following steps 101 to 105. When the training device of the three-dimensional image generation model executes the training method of the three-dimensional image generation model, the specific process is as follows. It should be noted first that the order of steps 101 to 105 in

[0065] is not specifically limited in this embodiment, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.

[0065] Step 101, obtain the original image from a sparse perspective, and input the original image into the initial three-dimensional image generation model to generate the corresponding adjacent frame image of the original image.

[0066] The following will describe step 101 in detail.

[0067] Among them, the original image refers to the unprocessed image data obtained by an acquisition device through image acquisition of a target object or scene, which retains all the description information during acquisition. The sparse view refers to a limited number of views, which corresponds to the dense view; if the dense view represents a continuous and rich number of views, then the sparse view represents a discrete and limited number of views. The process of obtaining the original image under the sparse view means that when acquiring the original image, the target object or scene is photographed from a limited number of angles (views), rather than densely acquiring images from all directions or multiple angles. It can be understood that since the description information of different views provided by the original image acquired under the sparse view is relatively few, it is usually not sufficient to completely reconstruct the corresponding target three-dimensional image of the target object or scene, or the accuracy of the reconstructed target three-dimensional image is poor.

[0068] Furthermore, the acquisition device can be an ordinary camera, a depth camera, a multi-spectral / hyperspectral camera, a radar sensor, a scanning device, etc. Of course, the acquisition device can also be other types of devices. Here, only examples are given, and it does not mean that the specific types of acquisition devices are limited in the embodiments of the present application.

[0069] Exemplarily, as Figure 3 shown, Figure 3 is an optional schematic diagram of original image acquisition for the training method of the three-dimensional image generation model provided by the embodiments of the present application. Figure 3 It represents that the acquisition devices A and B are used to perform image acquisition and processing on the target object. Figure 3 The target object in

[0070] is a table. Due to the very limited view of the image acquisition and processing, that is, only the original image a under the view A can be acquired through the acquisition device A, and the original image b under the view B can be acquired through the acquisition device B. In traditional three-dimensional reconstruction methods, the information available for three-dimensional reconstruction provided by the original images a and b acquired under the sparse view is relatively few. Furthermore, the initial three-dimensional image generation model includes a motion generation sub-model, and the motion generation sub-model is used to generate adjacent-frame images according to the input original images

[0071] Further, the motion generation sub-model can be a motion control model (MotionCtrl), a diffusion model (MotionDiffusion), a generative adversarial network-based model (GAN-based Models), etc. Of course, the motion generation sub-model can also be other types of models. Here, only examples are given, and it does not mean that the embodiments of the present application limit the specific type of the motion generation sub-model.

[0072] Further, in the traditional model training process, the collected original images usually come from sparse viewpoints. For example, when the acquisition time and acquisition equipment are limited, only a limited number of acquisition devices can be used to obtain the original images from a limited number of viewpoints; or, in actual situations, the specified acquisition points that can be collected are fixed and limited, and the viewpoint information provided by the original images collected from a small number of viewpoints is less. In this case, the embodiments of the present application incorporate a motion generation sub-model when constructing the initial three-dimensional image generation model. The purpose is to generate adjacent-frame images from other viewpoints based on the original images obtained from the current limited viewpoints. The adjacent-frame images enrich the information for training the initial three-dimensional image generation model, laying a foundation for generating a high-performance target three-dimensional image generation model later.

[0073] Step 102: Perform motion structure recovery processing on the original image and the adjacent-frame image to obtain an initial Gaussian point set, perform Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, and determine an optimization loss value based on the rough Gaussian point set.

[0074] The following is a detailed description of step 102.

[0075] Among them, structure-from-motion (SfM) is a computer vision technology that estimates the pose (position and orientation) of the camera and the three-dimensional structure of the scene by analyzing the feature point matching relationships in multiple images. Further, the three-dimensional point data obtained through structure-from-motion processing is the initial Gaussian point set. The initial Gaussian point set contains the initial geometric information of the target object or scene to be three-dimensionally reconstructed, which provides a basis for subsequent Gaussian optimization and three-dimensional reconstruction.

[0076] Among them, the rough Gaussian point set refers to a higher-quality three-dimensional point data obtained by performing Gaussian optimization processing on the initial Gaussian point set during the three-dimensional reconstruction process. Compared with the initial Gaussian point set, the rough Gaussian point set has a higher density, more accurate position information, and fewer outliers (abnormal points).

[0077] Among them, the optimization loss value is a quantitative index that measures the difference between the result obtained by the current three-dimensional reconstruction and the ideal target. The optimization loss value can guide subsequent model training and improve the training quality.

[0078] In some embodiments, Gaussian optimization is performed on the initial Gaussian point set to obtain a rough Gaussian point set, including the following steps:

[0079] (102.a.1) For any Gaussian basis element, count the number of neighboring Gaussian basis elements within a spherical region centered at the Gaussian basis element with a preset value as the radius;

[0080] (102.a.2) If the number of neighboring Gaussian basis elements is less than a preset basis element threshold, remove the Gaussian basis element from the initial Gaussian point set to obtain a rough Gaussian point set.

[0081] The following will describe steps (102.a.1) to (102.a.2) in detail.

[0082] In some embodiments, in the case of obtaining an initial Gaussian point set representing initial geometric information, 3D Gaussian Splatting (3DGS) is performed on the initial Gaussian point set to obtain a rough Gaussian point set. Among them, 3DGS realizes efficient and high-quality 3D reconstruction and rendering by modeling the surface of the target object or scene as a set of Gaussian distributions.

[0083] Among them, the initial Gaussian point set includes multiple basic units (Gaussian basis elements), and each Gaussian basis element contains information such as position (3D coordinates), scale (uncertainty range), and direction; the spherical region refers to a 3D spherical space defined with a certain Gaussian basis element as the center and a preset value as the radius; neighboring Gaussian basis elements are other Gaussian basis elements located within the spherical region except the central Gaussian basis element, and the number of other Gaussian basis elements obtained by statistics is the number of neighboring Gaussian basis elements.

[0084] Furthermore, the outlier Gaussian basis elements in the initial Gaussian point set may appear at positions far from the center of the set. Too many outlier Gaussian basis elements can easily complicate the subsequent 3D reconstruction process and hinder scene convergence. Based on this, in the process of performing 3DGS processing in the embodiments of the present application, discrete Gaussian basis element groups are removed by comparing the number of neighboring Gaussian basis elements with a preset basis element threshold, achieving a balance between removing outliers and retaining basic scene elements, effectively reducing the spread of outliers while maintaining the integrity of the scene representation.

[0085] Among them, the primitive threshold is used to determine whether to retain a certain Gaussian primitive. Specifically, if the number of neighboring Gaussian primitives of a certain Gaussian primitive is less than this threshold, it is considered that the spherical region where this Gaussian primitive is located is a sparse region, and this Gaussian primitive is an outlier caused by noise or incorrect matching. Removing this isolated Gaussian primitive can make the distribution of the three-dimensional point cloud data more uniform, while reducing unnecessary calculations to reduce the consumption of computing power resources and improving the quality of the target three-dimensional image generation model obtained by training. It should be noted that the radius of the spherical region and the primitive threshold can both be set according to the actual situation, and the embodiments of the present application do not limit this.

[0086] Exemplarily, as Figure 4 shown, Figure 4 FIG. is an optional schematic diagram of a spherical region of the training method of the three-dimensional image generation model provided by the embodiment of the present application. Among them, the radius of the spherical region G is R, and it is centered on the Gaussian primitive g0, and it is assumed that the primitive threshold is 5; in this example, in addition to the Gaussian primitive g0 in the spherical region G, there are also neighboring Gaussian primitives including the Gaussian primitive g1, the Gaussian primitive g2, and the Gaussian primitive g3, and the number of neighboring Gaussian primitives is 3. Since 3 (the number of neighboring Gaussian primitives) < 5 (the primitive threshold), the Gaussian primitive g0 is removed from the corresponding initial Gaussian point set; the Gaussian optimization operation is repeatedly performed on each Gaussian primitive to obtain a coarsely Gaussian point set after optimization processing.

[0087] In some embodiments, determining the optimization loss value based on the coarsely Gaussian point set includes the following steps:

[0088] (102.b.1) Parse the adjacent frame image to obtain an extracted feature point set and an extracted association graph set;

[0089] (102.b.2) Perform semantic rendering processing on the coarsely Gaussian point set to obtain a semantic feature point set and a semantic association graph set of the adjacent frame image;

[0090] (102.b.3) Perform convolution processing on the semantic feature point set and the semantic association graph set respectively to obtain a first convolution result and a second convolution result;

[0091] (102.b.4) Determine a first feature loss value based on the first convolution result and the extracted feature point set, and determine a second feature loss value based on the second convolution result and the extracted association graph set;

[0092] (102.b.5) Superimpose the first feature loss value and the second feature loss value to obtain the optimization loss value.

[0093] The following will describe steps (102.b.1) to (102.b.5) in detail.

[0094] In some embodiments, by parsing the adjacent frame images the obtained set of extracted feature points and the set of extracted association graphs provide additional semantic constraints. Specifically, the optimized loss value is obtained by the following formula <1>

[0095]

[0096] where and respectively perform semantic rendering processing on the set of rough Gaussian points to obtain the set of semantic feature points and the set of semantic association graphs of the adjacent frame images; ω f and ω s are both 1×1 convolutional networks. Using ω f to perform convolutional processing on the set of semantic feature points to obtain the first convolutional result Using ω s to perform convolutional processing on the set of semantic association graphs to obtain the second convolutional result

[0097] Furthermore, using the loss function, based on the first convolutional result and the set of extracted feature points to determine the first feature loss value; using the loss function, based on the second convolutional result and the set of association graphs to determine the second feature loss value. Where the loss function is a function related to cosine similarity calculation, which is used to measure the directional difference between two vectors; the loss function is used to measure the difference between the predicted probability distribution and the true label distribution.

[0098] Furthermore, it includes an image parsing sub-model, and the input image is parsed by the image parsing sub-model. Among them, the image parsing sub-model can be a multi-modal model (Contrastive Language-Image Pre-training, CLIP), a deep feature extraction model, a pre-trained visual model, etc. Of course, the specific type of the image parsing sub-model can be set according to the actual situation. The embodiments of the present application only make an example illustration and do not represent a limitation of the embodiments of the present application in this regard.

[0099] Further, in addition to parsing and processing the adjacent-frame image, the image parsing sub-model can also parse and process the original image in a manner similar to steps (102.b.1) to (102.b.5) to obtain the optimized loss value of the original image; then, the total parsing loss value is obtained by superimposing the optimized loss value of the original image and the optimized loss value of the adjacent-frame image; the initial three-dimensional image generation model is trained according to the total parsing loss value to more comprehensively evaluate the model training performance and improve the training accuracy of the model.

[0100] Step 103: Respectively extract the first semantic feature from the sparse perspective and the second semantic feature from the new perspective from the set of coarse Gaussian points.

[0101] The following provides a detailed description of step 103.

[0102] Among them, the new perspective refers to a new angle different from the sparse perspective, and the new perspective can be an unobserved perspective or a virtual perspective generated by interpolation. The first semantic feature is the feature extracted from the set of coarse Gaussian points that reflects the scene semantic information from the sparse perspective; the second semantic feature is the feature extracted from the set of coarse Gaussian points that reflects the scene semantic information from the new perspective. Among them, the new perspective is different from the sparse perspective, and the new perspective is also different from the perspective corresponding to the adjacent-frame image. Therefore, the second semantic feature is actually a pseudo-feature predicted from the new perspective. The pseudo-feature can introduce additional effective constraints, thereby enhancing the semantic feature learning from the sparse perspective and improving the training effect of the initial three-dimensional image generation model.

[0103] Further, the initial three-dimensional image generation model also includes a feature extraction sub-model, which is used to respectively extract the first semantic feature and the second semantic feature from the coarse Gaussian set. Among them, the feature extraction sub-model can be a convolutional neural network model, a graph neural network model, or other models. The specific type of the feature extraction sub-model can be set according to the actual situation, and the embodiments of the present application do not limit this.

[0104] Step 104: Perform contrastive learning regularization processing based on the first semantic feature and the second semantic feature to determine the contrast loss value.

[0105] The following provides a detailed description of step 104.

[0106] In some embodiments, inspired by the principle that similar semantic representations should be pulled closer and dissimilar semantic representations should be pushed farther apart, the embodiments of the present application introduce semantic-guided contrastive learning regularization processing to obtain a contrastive loss value for measuring the difference index of the feature distributions between positive sample pairs and negative sample pairs. Contrastive learning regularization processing refers to optimizing the model by introducing a regularization term during the contrastive learning process to improve the generalization ability of the model. Among them, regularization is a technique for reducing model overfitting. By adding a penalty term to the loss function, the complexity of the model is restricted, so that the model can also exhibit good generalization performance on unseen data.

[0107] In some embodiments, performing contrastive learning regularization processing based on the first semantic feature and the second semantic feature to determine the contrastive loss value includes the following steps:

[0108] (104.a.1) Distinguish the first positive semantic feature and the first negative semantic feature from the first semantic feature, and perform contrastive learning regularization processing from a sparse perspective based on the first positive semantic feature and the first negative semantic feature to obtain the first contrastive loss value;

[0109] (104.a.2) Distinguish the second positive semantic feature and the second negative semantic feature from the second semantic feature, and perform contrastive learning regularization processing from a new perspective based on the second positive semantic feature and the second negative semantic feature to obtain the second contrastive loss value;

[0110] (104.a.3) Superimpose the first contrastive loss value and the second contrastive loss value to obtain the contrastive loss value.

[0111] The following provides a detailed description of steps (104.a.1) to (104.a.3).

[0112] In some embodiments, U z is a sample set containing multiple first semantic features z. Select sample i from the multiple first semantic features z as the reference sample T z (i) ≡ U z {i}; then, obtain the semantic information of the first semantic feature through language rendering processing Then, determine the reference sample T z (i) corresponding semantic label Furthermore, define the first positive semantic sample set P z (i) includes all first positive semantic features having the same semantic label as the reference sample i, and define the first negative semantic sample set T z (i)\P z (i) includes all features other than the first positive semantic features (first negative semantic features).

[0113] Further, the first contrast loss value is calculated under the sparse perspective through the following formula <2>

[0114]

[0115] wherein, · represents the inner product operation, τ represents the temperature parameter that can be set, and e represents the natural constant. The definitions of the same parameters in the embodiments of the present application are the same, and the repeated parts will not be elaborated.

[0116] Further, For a sample set including multiple second semantic features z + select sample i as the reference sample from the multiple second semantic features z + Then, the semantic information of the second semantic feature is obtained through semantic rendering processing Next, determine the reference sample Then, determine the corresponding semantic label of the reference sample corresponding semantic label Further, define the second positive semantic sample set including all second positive semantic features having the same semantic label as the reference sample i + and define the second negative semantic sample set including all features (second negative semantic features) except the second positive semantic features.

[0117] Further, the second contrast loss value is calculated under the new perspective through the following formula <3>

[0118]

[0119] Further, obtain the contrast loss value

[0120] In some embodiments, after obtaining the second contrast loss value, the following steps are further included:

[0121] (104.b.1) Perform color rendering processing on the set of coarse Gaussian points to determine the color features under the sparse perspective;

[0122] (104.b.2) Distinguish positive color features and negative color features from the color features, and perform contrast learning regularization processing under the sparse perspective based on the positive color features and negative color features to obtain the color contrast loss value;

[0123] (104.b.3) Update the first contrast loss value using the color contrast loss value, and superimpose the updated first contrast loss value and the second contrast loss value to obtain the contrast loss value.

[0124] The following describes steps (104.b.1) to (104.b.3) in detail.

[0125] In some embodiments, considering that the color representations of the same target object or scene are more similar than those of different target objects or scenes, during the contrastive learning regularization process, soft color contrastive learning regularization is applied to enhance color representation learning. Specifically, U x is a sample set containing multiple color features x, and sample i is selected as the reference sample T from the multiple color features x x (i) ≡ U x {i}; then, the color information S of the color features is obtained through color rendering processing; then, the semantic label q corresponding to the reference sample T x (i) is determined as q = argmax(σ(ψ(S))); further, the positive color sample set P x (i) includes all positive color features with the same semantic label as the reference sample i, and the negative color sample set T x (i)\P x (i) includes all features (negative color features) except the positive color features.

[0126] Further, the color contrast loss value L under the sparse perspective is calculated by the following formula <4>[ CCL :

[0127]

[0128] Further, the contrastive learning regularization under the sparse perspective is updated using the color contrast loss value to obtain the first contrast loss value, and the updated first contrast loss value is Further, as shown in the following formula <5>, the contrast loss value is obtained by superimposing the updated first contrast loss value and the second contrast loss value

[0129]

[0130] Step 105: Train the initial three-dimensional image generation model based on the optimization loss value and the contrast loss value until the preset training conditions are met, and obtain the trained target three-dimensional image generation model.

[0131] The following provides a detailed description of step 105.

[0132] In some embodiments, the optimization loss value is used to measure the quality of the optimized set of coarse Gaussian points, ensuring that the generated set of coarse Gaussian points is as close as possible to the actual target object or scene; while the contrast loss value is used to measure the ability of the initial three-dimensional image generation model to distinguish between similar and dissimilar samples. By training the initial three-dimensional image generation model using various types of loss values, the model can not only generate high-quality Gaussian basis elements, but also maintain good consistency of each Gaussian basis element from multiple perspectives, thereby improving the performance of the finally trained target three-dimensional image generation model; at the same time, the training of the initial three-dimensional image generation model can be completed based on a small number of original images, reducing the consumption of computer resources while improving the training effect, and thus improving the accuracy of the target three-dimensional image output by the finally generated model in actual use.

[0133] Among them, the training conditions can be that the training time meets a preset time threshold, or the model parameters meet a preset parameter threshold. The model parameters can include the learning rate, batch size, number of iterations, etc., and can be specifically set according to the actual situation, and the embodiments of the present application do not limit this.

[0134] In some embodiments, training the initial three-dimensional image generation model based on the optimization loss value and the contrast loss value until a preset training condition is reached to obtain the trained target three-dimensional image generation model includes the following steps:

[0135] (105.a.1) Determine the initial loss value based on the initial set of Gaussian points, where the initial loss value includes an initial color loss value and an initial semantic loss value;

[0136] (105.a.2) Perform the first-stage joint training on the initial three-dimensional image generation model based on the initial loss value and the optimization loss value until a preset first training condition is reached to obtain an intermediate reconstruction model;

[0137] (105.a.3) Perform the second-stage joint training on the initial three-dimensional image generation model based on the initial loss value and the contrast loss value until a preset second training condition is reached to obtain the trained target three-dimensional image generation model.

[0138] The following details steps (105.a.1) to (105.a.3).

[0139] Among them, the initial color loss value is used to measure the color difference degree between the generated initial Gaussian point set and the initial Gaussian point set in the real state. The initial color loss value can be determined by means such as the mean square error function, the perceptual loss function, and the color consistency loss function. The initial semantic loss value is used to measure the semantic difference degree between the generated initial Gaussian point set and the initial Gaussian point set in the real state. The initial semantic loss value can be determined by means such as the cross-entropy loss function and the similarity metric function.

[0140] Further, after obtaining the initial color loss value and the initial semantic loss value the first loss value is obtained by the following formula <6>

[0141]

[0142] Further, in the first stage of training, based on the initial loss value and the optimized loss value the initial three-dimensional image generation model is jointly trained. The purpose is to let the initial three-dimensional image generation model learn how to generate a high-quality initial Gaussian point set that meets the expected standards, so as to obtain an intermediate reconstruction model.

[0143] At this time, the model training is not over yet, and the training will enter the second stage. Specifically, the second loss value is obtained by the following formula <7>

[0144]

[0145] Further, in the second stage of training, based on the initial loss value and the contrast loss value the initial three-dimensional image generation model is jointly trained. The purpose is to let the intermediate reconstruction model learn how to perform further Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set. When the training in the second stage ends, the target three-dimensional image generation model is obtained.

[0146] It should be noted that the training conditions include the first training condition and the second training condition. The first training condition and the second training condition are one or more of the training conditions. The first training condition and the second training condition can be the same or different, and can be specifically set according to the actual situation.

[0147] In some embodiments, after obtaining the trained target three-dimensional image generation model, the following steps are further included:

[0148] (105.b.1) Obtain the target image under the sparse view;

[0149] (105.b.2) Input the target image into the trained target 3D image generation model to output the corresponding target 3D image of the target image.

[0150] The following provides a detailed description of steps (105.b.1) to (105.b.2).

[0151] In some embodiments, the target image obtained from a sparse view is input into the trained target 3D image generation model to obtain the target 3D image. Among them, the target 3D image is essentially a 3D model, and its prediction restores the 3D structure and appearance features of the target object or scene represented by the target image, providing intuitive and three-dimensional visual effect information with high accuracy. In practical applications, the target 3D image can be applied in fields such as architectural design and film and television entertainment.

[0152] Among them, the target image is an image obtained based on a sparse view during the actual application process. It should be noted that the perspective information provided by the target image for 3D reconstruction is very limited. However, the target 3D image generation model trained in the embodiments of the present application can still restore the target 3D image with high accuracy based on the target image from a sparse view. In addition, the sparse view for obtaining the target image and the sparse view for obtaining the original image can be the same or different. The acquisition method of the target image is similar to that of the original image, which will not be elaborated here.

[0153] To enable readers to further understand the beneficial effects of the embodiments of the present application, the following provides a complete example. The specific implementation methods of each step in this example can be found in the relevant detailed explanations in the specification.

[0154] As Figure 5 shown, Figure 5 is an optional model training and application schematic diagram of the 3D image generation model training method provided by the embodiments of the present application. Among them, Figure 5 Push in

[0155] (1) The training stage includes two modules, namely the coarse Gaussian point set generation training module and the contrastive learning regularization module.

[0156] (1.1) In the coarse Gaussian point set generation training module:

[0157] ① Obtain the original image from a sparse view;

[0158] ② Input the original image into the initial 3D image generation model to generate the corresponding adjacent frame image of the original image;

[0159] ③Perform motion structure recovery processing on the original image and the adjacent frame image to obtain an initial Gaussian point set;

[0160] ④Perform Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set;

[0161] ⑤Determine the optimization loss value based on the rough Gaussian point set

[0162] (1.2) Further, after the training module is completed based on the rough Gaussian point set, in the contrast learning regularization module:

[0163] ⑥Perform initialization processing on the rough Gaussian point set;

[0164] ⑦Perform Gaussian optimization processing on the initialized rough Gaussian point set. Then, the Differentiable Parallel Rasterizer renders the results after Gaussian optimization from sparse viewpoints and new viewpoints respectively; among them, the Differentiable Parallel Rasterizer is a part set in the initial 3D image generation model, which is used to convert the 3D model into a 2D image. It should be noted that the rendering from sparse viewpoints and new viewpoints is carried out simultaneously;

[0165] ⑧Calculate the first contrast loss value under the sparse viewpoint and the color contrast loss value and calculate the second contrast loss value under the new viewpoint Based on the first contrast loss value the color contrast loss value and the second contrast loss value Calculate the contrast loss value Then, use the optimization loss value and the contrast loss value Train the initial 3D image generation model until the preset training conditions are met to obtain the trained target 3D image generation model.

[0166] (2) In the application stage:

[0167] ①Obtain the target image;

[0168] ②Input the target image into the trained target 3D image generation model and output the corresponding target 3D image of the target image.

[0169] As Figure 6 shown, Figure 6 is an optional module schematic diagram of the training device of the 3D image generation model provided by the embodiment of the present application. The training device of the 3D image generation model includes the following modules 201 to module 204:

[0170] An acquisition module 201, which acquires the original image from a sparse perspective and inputs the original image into an initial three-dimensional image generation model to generate a neighboring frame image corresponding to the original image;

[0171] An optimized loss value determination module 202, which is used to perform motion structure restoration processing on the original image and the neighboring frame image to obtain an initial Gaussian point set, perform Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, and determine an optimized loss value based on the rough Gaussian point set;

[0172] A semantic feature extraction module 203, which is used to respectively extract a first semantic feature from the sparse perspective and a second semantic feature from a new perspective from the rough Gaussian point set, where the new perspective is different from the sparse perspective;

[0173] A contrast loss value determination module 204, which is used to perform contrast learning regularization processing based on the first semantic feature and the second semantic feature to determine a contrast loss value;

[0174] A target training module 205, which is used to train the initial three-dimensional image generation model based on the optimized loss value and the contrast loss value until a preset training condition is reached, and obtain a trained target three-dimensional image generation model.

[0175] The training method, device, equipment and medium of the three-dimensional image generation model proposed in this application obtain the original images from sparse viewpoints and input the original images into the initial three-dimensional image generation model to generate the corresponding adjacent frame images of the original images. The original images and the adjacent frame images are processed for motion structure recovery to obtain an initial Gaussian point set, and the initial Gaussian point set is processed by Gaussian optimization to obtain a rough Gaussian point set. The optimization loss value is determined based on the rough Gaussian point set. The adjacent frame images enrich the information for training the initial three-dimensional image generation model, laying a foundation for generating a high-performance target three-dimensional image generation model in the future. Then, the first semantic feature from the sparse viewpoint and the second semantic feature from a new viewpoint are respectively extracted from the rough Gaussian point set, where the new viewpoint is different from the sparse viewpoint. Based on the first semantic feature and the second semantic feature, contrastive learning regularization processing is performed to determine the contrast loss value. The initial three-dimensional image generation model is trained based on the optimization loss value and the contrast loss value until the preset training conditions are met, and the trained target three-dimensional image generation model is obtained. The optimization loss value is used to measure the quality of the obtained rough Gaussian point set after optimization, ensuring that the generated rough Gaussian point set is as close as possible to the actual target object or scene. The contrast loss value is used to measure the ability of the initial three-dimensional image generation model to distinguish between similar and dissimilar samples, thereby improving the performance of the finally trained target three-dimensional image generation model. At the same time, the initial three-dimensional image generation model can be trained based on a small number of original images, reducing the consumption of computer resources while improving the training effect, and further improving the accuracy of the target three-dimensional image output by the finally generated model in actual use.

[0176] The specific implementation manner of the training device of the three-dimensional image generation model is basically the same as the specific embodiment of the training method of the three-dimensional image generation model described above, and will not be elaborated here.

[0177] The embodiment of this application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the training method of the three-dimensional image generation model described above. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0178] As Figure 7 shown, Figure 7 is the hardware structure schematic diagram of the electronic device provided by the embodiment of this application. The electronic device includes:

[0179] The processor 301 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0180] The memory 302 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 302 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 302 and are called by the processor 301 to execute the training method of the three-dimensional image generation model in the embodiments of the present application;

[0181] The input / output interface 303 is used to implement information input and output;

[0182] The communication interface 304 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0183] The bus 305 transmits information between the various components of the device (such as the processor 301, the memory 302, the input / output interface 303, and the communication interface 304);

[0184] Among them, the processor 301, the memory 302, the input / output interface 303, and the communication interface 304 achieve communication connections with each other inside the device through the bus 305.

[0185] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the training method of the above-mentioned three-dimensional image generation model is implemented.

[0186] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0187] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0188] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0191] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above figures are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0192] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression below refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0193] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0194] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0196] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0197] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.

Claims

1. A training method for a three-dimensional image generation model, characterized in that Including: Obtain the original image from a sparse perspective and input the original image into an initial three-dimensional image generation model to generate a neighboring frame image corresponding to the original image; Perform motion structure recovery processing on the original image and the neighboring frame image to obtain an initial Gaussian point set, perform Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, and determine an optimization loss value based on the rough Gaussian point set; Extract a first semantic feature in the sparse perspective and a second semantic feature in a new perspective from the rough Gaussian point set, where the new perspective is different from the sparse perspective; Perform contrast learning regularization processing based on the first semantic feature and the second semantic feature to determine a contrast loss value; Train the initial three-dimensional image generation model based on the optimization loss value and the contrast loss value until a preset training condition is reached to obtain a trained target three-dimensional image generation model.

2. The training method of the three-dimensional image generation model according to claim 1, characterized in that The initial Gaussian point set includes a plurality of Gaussian basis elements; The performing Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set includes: For any one of the Gaussian basis elements, count the number of neighboring Gaussian basis elements within a spherical region centered on the Gaussian basis element with a preset value as the radius; If the number of neighboring Gaussian basis elements is less than a preset basis element threshold, remove the Gaussian basis element from the initial Gaussian point set to obtain the rough Gaussian point set.

3. The training method of the three-dimensional image generation model according to claim 1, wherein The determining the optimization loss value based on the rough Gaussian point set includes: Perform parsing processing on the neighboring frame image to obtain an extracted feature point set and an extracted association graph set; Perform semantic rendering processing on the rough Gaussian point set to obtain a semantic feature point set and a semantic association graph set of the neighboring frame image; Perform convolution processing on the semantic feature point set and the semantic association graph set respectively to obtain a first convolution result and a second convolution result; Determine a first feature loss value based on the first convolution result and the extracted feature point set, and determine a second feature loss value based on the second convolution result and the extracted association graph set; Superimpose the first feature loss value and the second feature loss value to obtain the optimization loss value.

4. The training method of the three-dimensional image generation model according to claim 1, wherein The performing contrast learning regularization processing based on the first semantic feature and the second semantic feature to determine a contrast loss value includes: Distinguish a first positive semantic feature and a first negative semantic feature from the first semantic feature, and perform contrast learning regularization processing in the sparse perspective based on the first positive semantic feature and the first negative semantic feature to obtain a first contrast loss value; Distinguish a second positive semantic feature and a second negative semantic feature from the second semantic feature, and perform contrast learning regularization processing in the new perspective based on the second positive semantic feature and the second negative semantic feature to obtain a second contrast loss value; Superimpose the first contrast loss value and the second contrast loss value to obtain the contrast loss value.

5. The training method of the three-dimensional image generation model according to claim 4, wherein After obtaining the second contrast loss value, it further includes: Perform color rendering processing on the rough Gaussian point set to determine the color feature in the sparse perspective; Distinguish the positive color feature and the negative color feature from the color features, and perform contrastive learning regularization processing in the sparse perspective based on the positive color feature and the negative color feature to obtain a color contrast loss value; Update the first contrast loss value using the color contrast loss value, and superimpose the updated first contrast loss value and the second contrast loss value to obtain the contrast loss value.

6. The training method of the three-dimensional image generation model according to claim 1, wherein The training conditions include a first training condition and a second training condition; Training the initial three-dimensional image generation model based on the optimization loss value and the contrast loss value until a preset training condition is reached to obtain a trained target three-dimensional image generation model, including: Determine an initial loss value based on the initial Gaussian point set, where the initial loss value includes an initial color loss value and an initial semantic loss value; Perform a first-stage joint training on the initial three-dimensional image generation model based on the initial loss value and the optimization loss value until the preset first training condition is reached to obtain an intermediate reconstruction model; Perform a second-stage joint training on the initial three-dimensional image generation model based on the initial loss value and the contrast loss value until the preset second training condition is reached to obtain the trained target three-dimensional image generation model.

7. The training method of the three-dimensional image generation model according to claim 1, characterized in that After obtaining the trained target three-dimensional image generation model, it further includes: Obtain a target image in the sparse perspective; Input the target image into the trained target three-dimensional image generation model to output the corresponding target three-dimensional image of the target image.

8. A training device for a three-dimensional image generation model, characterized in that, Including: An acquisition module that acquires an original image in the sparse perspective and inputs the original image into the initial three-dimensional image generation model to generate a neighboring frame image corresponding to the original image; An optimization loss value determination module for performing motion structure restoration processing on the original image and the neighboring frame image to obtain an initial Gaussian point set, performing Gaussian optimization processing on the initial Gaussian point set to obtain a rough Gaussian point set, and determining an optimization loss value based on the rough Gaussian point set; A semantic feature extraction module for respectively extracting a first semantic feature in the sparse perspective and a second semantic feature in a new perspective from the rough Gaussian point set, where the new perspective is different from the sparse perspective; A contrast loss value determination module for performing contrastive learning regularization processing based on the first semantic feature and the second semantic feature to determine a contrast loss value; A target training module for training the initial three-dimensional image generation model based on the optimization loss value and the contrast loss value until a preset training condition is reached to obtain a trained target three-dimensional image generation model.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the training method of the three-dimensional image generation model according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method of the three-dimensional image generation model according to any one of claims 1 to 7.