Model training method, image generation method and related equipment
By generating predicted images based on collected viewpoint information and dynamically adjusting the gradient scaling factor during model training, the problem of slow convergence speed in traditional model training is solved, thus improving model training efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MOORE THREADS TECH CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-15
AI Technical Summary
Traditional model training methods result in slow model convergence, leading to low training efficiency.
By generating a predicted image based on the acquisition viewpoint information corresponding to the target image in the current round, and dynamically adjusting the gradient scaling factor using the gradient information and visibility count of the target Gaussian sphere, the model convergence speed is improved.
It improves the efficiency of model training, avoids the problem of slowed overall convergence caused by gradient information in traditional model training methods, and achieves faster model convergence.
Smart Images

Figure CN122049564A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a model training method, an image generation method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the continuous development of artificial intelligence technology, neural network models are widely used in various business scenarios to solve users' actual business needs; for example, in image generation scenarios, image processing models can be used to generate the images required by users.
[0003] Image processing models of related technologies need to be trained before application. Since the traditional model training method has a slow convergence speed, the model training efficiency is low. Therefore, how to improve the model convergence speed and training efficiency has become an urgent technical problem to be solved. Summary of the Invention
[0004] This disclosure provides a model training method, an image generation method, a model training device, an image generation device, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] In a first aspect, this disclosure provides a model training method, which includes: generating a predicted image using an image processing model of the current round based on the acquisition viewpoint information corresponding to the target image of the current round, wherein the predicted image is generated based on a target Gaussian sphere in the image processing model of the current round, the image processing model includes multiple Gaussian spheres, the target Gaussian sphere is at least one Gaussian sphere corresponding to the acquisition viewpoint information, and the current round is any round in a multi-round model training; determining a gradient scaling factor for the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the predicted image in the multi-round model training; and training the target Gaussian sphere based on the gradient information and the gradient scaling factor to obtain the image processing model of the next round.
[0006] Secondly, this disclosure provides a model training device, which includes: an image generation module, a gradient determination module, a factor determination module, and a model training module.
[0007] The image generation module is configured to generate a predicted image based on the acquisition viewpoint information corresponding to the target image of the current round and using the image processing model of the current round. The predicted image is generated based on the target Gaussian sphere in the image processing model of the current round. The image processing model includes multiple Gaussian spheres, and the target Gaussian sphere is at least one Gaussian sphere corresponding to the acquisition viewpoint information. The current round is any round in the multi-round model training.
[0008] The gradient determination module is configured to determine the gradient information of the target Gaussian sphere based on the predicted image and the target image.
[0009] The factor determination module is configured to determine the gradient scaling factor of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training.
[0010] The model training module is configured to train the target Gaussian sphere based on the gradient information and the gradient scaling factor to obtain the image processing model for the next round.
[0011] Thirdly, this disclosure provides an image generation method, which includes: generating a predicted image using an image processing model based on target acquisition viewpoint information, wherein the image processing model is determined according to the above-mentioned model training method.
[0012] Fourthly, this disclosure provides an image generation apparatus, which includes an image generation module.
[0013] The image generation module is configured to generate a predicted image based on the target acquisition viewpoint information and using an image processing model, wherein the image processing model is determined according to the model training method described above.
[0014] Fifthly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the model training method or image generation method described above.
[0015] In a sixth aspect, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described model training method or image generation method.
[0016] In a seventh aspect, this disclosure provides a computer program product that includes computer-readable code or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the model training method or image generation method described above.
[0017] The model training method provided in this disclosure can generate a predicted image using the target Gaussian sphere in the image processing model of the current round, based on the acquisition viewpoint information corresponding to the target image of the current round. Then, the gradient information of the target Gaussian sphere is determined based on the predicted image and the target image. In order to improve the model convergence speed, this method generates a dynamically adjusted gradient scaling factor based on the number of times the target Gaussian sphere is visible, and trains the target Gaussian sphere together with the gradient scaling factor and gradient information. This allows the gradient information to be dynamically adjusted during the model training process, thereby improving the model convergence speed and avoiding the problem of overall slowdown in convergence caused by gradient information in traditional model training methods, thus improving the model training efficiency.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0020] Figure 1 This is a flowchart of a model training method provided in an embodiment of the present disclosure.
[0021] Figure 2 This is a schematic diagram illustrating the application of a model training method provided in an embodiment of this disclosure.
[0022] Figure 3 This is a flowchart of an image generation method provided in an embodiment of the present disclosure.
[0023] Figure 4 This is a block diagram of a model training device provided in an embodiment of the present disclosure.
[0024] Figure 5 This is a block diagram of an image generation apparatus provided in an embodiment of the present disclosure.
[0025] Figure 6 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0026] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0027] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0028] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0029] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0030] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0031] The model training method or image generation method according to embodiments of this disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. Alternatively, the method can be executed by a server.
[0032] The technical terms used in the embodiments of this disclosure are explained below.
[0033] 3DGS model (3D Gaussian Splatting): is a 3D scene representation and rendering technique. It uses tens of thousands of learnable 3D Gaussian spheres to represent the scene and, in conjunction with a rasterization renderer, achieves high-quality real-time (or real-time-level) new view compositing.
[0034] Visibility probability: The probability that a Gaussian sphere in a 3DGS model will be "seen" or "observed" under given observation conditions (such as camera viewpoint, occlusion, lighting, sampling strategy, etc.).
[0035] Adam optimizer: It is an adaptive learning rate optimization algorithm that dynamically adjusts the learning rate independently for each parameter (i.e., the Gaussian ball).
[0036] Adaptive Learning Rate: This refers to an adaptive learning rate, a mechanism that automatically adjusts the learning rate based on the performance of each parameter during training.
[0037] Element-wise Regret Bound: This is an analytical method that defines total regret separately for each dimension of the parameter vector.
[0038] Epoch: The process by which the model completes one forward and back propagation cycle on the entire training set.
[0039] Iteration: The process of processing a batch of data; the model updates its parameters once after processing each batch of data (1 Batch), and this process is called 1 Iteration.
[0040] PSNR (Peak Signal-to-Noise Ratio): An objective metric used to quantify the difference between a reconstructed image and the original image. Its core idea is to measure the degree of distortion by calculating the mean square error (MSE), and to express it in decibels (dB) as the ratio of the maximum signal power to the noise power.
[0041] With the continuous development of artificial intelligence technology, neural network models are widely used in various business scenarios to solve users' actual business needs. For example, in image generation scenarios, image processing models can be used to generate the images required by users. Image processing models using related technologies need to be trained before application. However, traditional model training methods have slow convergence speeds, resulting in low training efficiency.
[0042] Taking 3D Gaussian Splatting (3DGS model) as an example, the 3DGS model is an explicit radiative field representation method that reconstructs a 3D scene by optimizing the attribute parameters (position, covariance, color, opacity, etc.) of thousands of 3D Gaussian spheres in the scene. The training process for the 3DGS model can be implemented using stochastic gradient descent (SGD) or the Adam optimizer. In each iteration, an image (or a patch of an image) is randomly sampled from the training set, and then the 3DGS model renders the image from that viewpoint and calculates the loss, which is then used for backpropagation to update the parameters.
[0043] In a standard 3DGS training process, parameter updates can depend on their visibility in the current viewpoint. However, differences in visibility can lead to difficulties in overall model convergence. A specific implementation method involves randomly and uniformly sampling training images during the model training loop. For each iteration, only the Gaussian spheres retained after frustum culling and occlusion culling participate in gradient calculation and parameter updates.
[0044] The drawback of this implementation method is that, due to the complexity of the geometry of 3D scenes (such as occlusion and edge regions), different Gaussian spheres exhibit significant differences in their "parameter visibility probability" (i.e., the probability of being observed) within the training set. Specific drawbacks include:
[0045] 1. Uneven convergence speed: Gaussian spheres located in the center of the scene and without occlusion have a high visibility probability and converge quickly; while Gaussian spheres located at the edge or behind occlusions (such as under a table or in a corner of a wall) have an extremely low visibility probability, resulting in a geometric decrease in the number of training iterations and extremely slow convergence.
[0046] 2. Optimizer Failure: The Adam optimizer relies on gradient momentum estimation. For parameters with extremely low visibility probabilities, gradient updates are sparse, leading to inaccurate momentum estimation (decaying to 0), which causes the adaptive learning rate mechanism to fail.
[0047] 3. Poor final rendering quality: Due to the above reasons, low-visibility areas in the scene cannot converge within a limited number of iterations, resulting in blurring, artifacts, or missing geometric structures.
[0048] Based on the above, it can be seen that in the standard 3DGS training process, the uniform sampling strategy does not consider the visibility imbalance at the parameter level, leading to underfitting in low-visibility regions. Furthermore, simply increasing the global learning rate causes oscillations in high-visibility regions; simply increasing the learning rate of low-visibility probability parameters causes the L-smoothness assumption to fail, resulting in training divergence.
[0049] Based on this, this disclosure provides a model training method, an image generation method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product, as detailed in the following embodiments.
[0050] Figure 1 A flowchart illustrating a model training method provided in an embodiment of this disclosure. (Refer to...) Figure 1 The method includes steps S11 to S14.
[0051] Step S11: Based on the acquisition viewpoint information corresponding to the target image in the current round, generate a predicted image using the image processing model of the current round. The predicted image is generated based on the target Gaussian sphere in the image processing model of the current round. The image processing model includes multiple Gaussian spheres, and the target Gaussian sphere is at least one Gaussian sphere corresponding to the acquisition viewpoint information. The current round is any round in the multi-round model training.
[0052] The target image can be understood as an image used to train the image processing model, and the target image can be a sample label; the target image can be an image obtained by capturing images of the target object based on the acquisition viewpoint information; there can be multiple target images, which can be understood as multiple images obtained by capturing images of the target object based on different acquisition viewpoint information. The target object can be a three-dimensional object or a living organism. This disclosure does not specifically limit the target object. For example, the target object can be an apple, a pet, etc.; multiple target images can be four apple images or four pet images captured (shot) from four perspectives (front, back, left, and right).
[0053] The acquisition perspective information can be understood as the information needed to acquire the target image, including but not limited to: the camera's position and orientation in the world. The position can be a description of the camera's coordinates in 3D space; this position can be represented by a three-dimensional vector or three-dimensional coordinates. The orientation can be represented by rotation matrices, quaternions, Euler angles, etc. The world can be the real world or a virtual world (e.g., a game world, a virtual reality scene, etc.). In some embodiments, the acquisition perspective information may also include camera intrinsic data, including but not limited to: focal length, principal point (i.e., the image center point), image size (i.e., the number of pixels), etc. For example, the acquisition perspective information can be the camera pose.
[0054] In some embodiments, multi-round model training can be understood as model training for multiple iterations of an image generation model, while a single multi-round model training can be model training for one epoch of an image generation model. That is to say, in the embodiments provided in this disclosure, multiple epochs of model training can be performed on the image generation model, and a single epoch of model training includes model training for multiple iterations (i.e., the multi-round model training).
[0055] In a multi-round model training process, the target images and acquisition viewpoint information corresponding to each round of model training are different; for example, in a multi-round model training process of 1000 rounds (Iteration), the corresponding multiple target images can be 1000 images. In some embodiments, the multiple target images can be acquired under the same or different acquisition viewpoint information.
[0056] In some embodiments, the image processing model can be understood as a model used to generate a predicted image based on the input acquisition viewpoint information. The image processing model can be a 3D Gaussian model, a 3D reconstruction model, etc., and this disclosure does not impose any specific limitations on it; for example, the image processing model can be a 3DGS model.
[0057] In an image processing model, multiple Gaussian spheres can be used to generate a predicted image; these Gaussian spheres can be referred to as Gaussian spheres or Gaussian ellipsoids. The target Gaussian sphere can be understood as the Gaussian sphere used to generate the predicted image corresponding to the acquired viewpoint information.
[0058] A predicted image can be understood as an image generated by an image processing model, which can be an image corresponding to the target image.
[0059] In some embodiments, generating a predicted image using the image processing model of the current wheel based on the acquisition viewpoint information corresponding to the target image of the current wheel includes: determining the target Gaussian sphere from multiple Gaussian spheres of the image processing model of the current wheel using the acquisition viewpoint information; and performing image rendering using the target Gaussian sphere to generate the predicted image.
[0060] Specifically, the model training method provided in this disclosure will input the acquisition viewpoint information corresponding to the target image of the current round into the image processing model of the current round during the current round of model training. The image processing model of the current round can be the image processing model that needs to be trained in the current round.
[0061] Using the acquired viewpoint information, a target Gaussian sphere corresponding to the acquired viewpoint information can be determined from multiple Gaussian spheres in the current round of the image processing model. Furthermore, the image processing model can be used to render the target Gaussian sphere to obtain a predicted image corresponding to the acquired viewpoint information.
[0062] As can be seen from the above embodiments, this disclosure can utilize the acquisition perspective information to generate a predicted image through the target Gaussian sphere in the image processing model. This enables the image processing model to optimize parameters based on different acquisition perspectives during model training, thereby improving the image generation model's ability to generate images from multiple perspectives. Furthermore, it allows the predicted image generated during model inference to be closer to the real image in terms of structure and visual effects.
[0063] In some embodiments, the predicted image is a three-dimensional model image; determining the target Gaussian sphere from multiple Gaussian spheres of the current round's image processing model using the acquisition viewpoint information includes: adjusting the parameters of multiple Gaussian spheres of the current round's image processing model using the acquisition viewpoint information to obtain adjusted multiple Gaussian spheres, wherein the adjusted multiple Gaussian spheres are used to construct a three-dimensional model corresponding to the target image; determining multiple target Gaussian spheres visible under the acquisition viewpoint information from the adjusted multiple Gaussian spheres; and generating the predicted image by rendering the image using the target Gaussian spheres includes: rendering the image based on the multiple target Gaussian spheres to obtain a three-dimensional model image of the three-dimensional model visible under the acquisition viewpoint information.
[0064] The target image can be understood as an image used for 3D reconstruction. The image processing model can use multiple Gaussian spheres to construct a 3D model corresponding to the target object in the target image and output a predicted image of the 3D model under a specific acquisition viewpoint. For example, the target object can be an apple. The image processing model can use multiple Gaussian spheres to generate a 3D model corresponding to the apple. When the acquisition viewpoint information of the current round is input into the image processing model, the image processing model can output an image of the 3D model visible under that acquisition viewpoint information (i.e., the predicted image).
[0065] The 3D model image can be understood as an image generated from the visible portion of the 3D model (i.e., the target Gaussian sphere) under the current acquisition viewpoint information. For example, if the 3D model is an apple model and the acquisition viewpoint information is a top-down view, the 3D model image can be understood as an image generated from the visible portion of the apple model under the top-down viewpoint observation. In other words, the predicted image can be a 3D model image of the 3D model constructed by the image processing model under the current acquisition viewpoint information.
[0066] Adjusting the parameters of multiple Gaussian spheres includes adjusting their distribution positions, covariance, and / or display parameters (such as color and transparency).
[0067] Taking the application of the model training method provided in this disclosure in a 3D reconstruction scene as an example, the model training method is explained. The image processing model in this model training method can be a 3DGS model, and the viewpoint information is the camera pose. Based on this, the model training method of this disclosure uses the 3DGS model to generate the predicted image as follows.
[0068] 1. Initialize the 3D Gaussian sphere set.
[0069] Initialize thousands of 3D Gaussian spheres; each Gaussian sphere contains learnable parameters, including but not limited to: position (3D coordinates), covariance (controlling shape and orientation), color (usually represented by spherical harmonics), transparency (alpha value), etc.
[0070] 2. Obtain the camera pose corresponding to the target image of the current round and input it into the 3DGS model.
[0071] 3. Render the predicted image from the current viewpoint using the 3DGS model.
[0072] The specific execution method is as follows: the 3DGS model performs adjustment operations such as frustum culling (only retaining Gaussian spheres within the field of view) and depth sorting on all Gaussian spheres according to the camera pose, and adjusts the display parameters of the Gaussian spheres within the field of view to obtain multiple adjusted Gaussian spheres.
[0073] The Gaussian sphere within the field of view is projected onto a 2D image plane (rasterization); then, alpha blending is used to synthesize and render the final RGB image (3D model image). It should be noted that, due to the issue of Gaussian spheres within the field of view being occluded by other Gaussian spheres, this RGB image is generated from the Gaussian sphere visible from the camera's pose (the target Gaussian sphere).
[0074] As can be seen from the above embodiments, the model training method provided in this disclosure can utilize a 3DGS model to reconstruct a renderable 3D scene representation (i.e., a 3D model) from multi-view images, and can generate corresponding 2D images for any acquisition viewpoint. The 3D model obtained using the 3DGS model can be a set of precisely arranged 3D Gaussian spheres, which can accurately reconstruct all acquisition viewpoints and generate corresponding 2D images; thus, it achieves the ability to reconstruct the 3D scene corresponding to the original multi-view images with high precision in a continuous and explicit geometric and appearance expression.
[0075] Step S12: Determine the gradient information of the target Gaussian sphere based on the predicted image and the target image.
[0076] Specifically, after obtaining the predicted image, this disclosure calculates a loss value using the predicted image and the target image (the actual acquired image), and calculates the gradient information of the target Gaussian sphere based on the loss value.
[0077] Step S13: Determine the gradient scaling factor of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training.
[0078] The visible number can be understood as the number of times the target Gaussian sphere is rendered to the prediction image or generated during multiple rounds of model adjustment.
[0079] The gradient scaling factor can be understood as a parameter used to scale, update, or adjust the gradient information of the target Gaussian sphere. This gradient scaling factor can be a numerical value or a weight.
[0080] Specifically, the model training method provided in this disclosure can determine the number of times a target Gaussian sphere is rendered into the prediction image during multiple rounds of model training, and process the number of times of visibility according to the scaling factor calculation module to obtain the gradient scaling factor of the target Gaussian sphere; the scaling factor calculation module can be a sub-model or agent in the model training process, used to process the input number of times of visibility to obtain the gradient scaling factor of the target Gaussian sphere.
[0081] In some embodiments, determining the gradient scaling factor of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training includes: determining the visibility probability of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training; and determining the gradient scaling factor of the target Gaussian sphere based on the visibility probability.
[0082] The visibility probability is used to represent the probability that the target Gaussian sphere is rendered to the prediction image or generated as a prediction image during multiple rounds of model adjustment. The visibility probability can be a probability value, and each target Gaussian sphere has a corresponding visibility probability. For example, the visibility probability can be the parameter visibility probability of the Gaussian sphere.
[0083] Specifically, this disclosure can calculate the visibility probability of a target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during multiple rounds of model training.
[0084] After obtaining the visibility probability of the target Gaussian sphere, the gradient scaling factor of the target Gaussian sphere can be calculated based on the visibility probability and preset calculation parameters. The preset calculation parameters can be weights or smoothing terms, used to limit the range of the gradient scaling factor and avoid unstable problems such as the gradient scaling factor being too large, too small, or fluctuating too much.
[0085] For example, in backpropagation to calculate the original gradient After (i.e., gradient information), but before the optimizer updates the parameters, the gradient needs to be corrected. Modifying the gradient requires calculating the corresponding gradient scaling factor (also known as the visibility compensation factor). This disclosure calculates the gradient scaling factor in the following ways:
[0086] Calculate the first Visibility compensation factor for a Gaussian sphere . It is inversely proportional to its visibility probability and is defined as the gradient scaling factor. ,in, This is a smoothing term. In other words, the gradient scaling factor of the target Gaussian sphere is inversely proportional to the visibility probability.
[0087] In some embodiments provided in this disclosure, the gradient scaling factor is designed to counteract the slow convergence speed caused by the gradient of each target Gaussian sphere. The specific justification is as follows.
[0088] This disclosure may assume that the gradients between different parameters (i.e., Gaussian spheres) are decoupled (i.e., the Hessian matrix is a diagonal matrix); for example, most techniques for Adaptive Learning Rate in the "Occam's Razor" principle implicitly or explicitly assume, when analyzing convergence, that the gradients between different parameters are independent of each other, or that the focus is on the Element-wise Regret Bound.
[0089] Before describing the derivation process, the definitions of the symbols in the following derivation formulas will be explained, as follows.
[0090] Let the scene parameter set be , of which The parameters of a Gaussian Primitive are .
[0091] Define the sample space as the image index in the training set. .
[0092] The cumulative number of times a point is visible is divided by the number of sample labels to obtain the visibility probability of the visible Gaussian sphere.
[0093] Visibility event : refers to random events, representing parameters Visible in the currently sampled training view (not culled or occluded by the view frustum).
[0094] Visibility Probability This refers to the fact that in the rendering pipeline, primitives may be discarded after the "Visibility Test (rendering operation)". Primitive parameter The probability that is visible.
[0095] Geometric gradient random variable : refers to when When it occurs, the gradient is calculated by backpropagation. That is to say, due to factors such as viewing angle differences and rasterization noise, the gradient disclosed in this invention is a random variable.
[0096] : refers to the loss function.
[0097] Random mask matrix ,in , This is the current iteration round. The parameter update rule is: Here, `diag()` is the "diagonal" operation, which can transform a vector into a diagonal matrix or extract diagonal elements from a matrix. `Bernoulli()` is the Bernoulli distribution function, which randomly outputs 1 or 0 based on a given probability. : Represents the gradient (i.e., gradient information); : Indicates the learning rate.
[0098] The model training method disclosed herein assumes that the loss function is... -Strongly convex and - Smooth; the discussion is restricted to the vicinity of the optimal solution, and the loss function is considered to be strongly abrupt and smooth, which can be understood as a standard regularity condition used to ensure that the function has a good shape.
[0099] Depend on -Strong convexity yields:
[0100] Where T is the time step; y and x are any two parameter vectors in the function input space (such as two different sets of parameters of the model), used to describe the "curvature" of the function over the entire domain.
[0101] In the case of gradient decoupling between parameters, it can be further derived that:
[0102] Among them, in the above formula The vector that indicates the current position is the local optimum.
[0103] Depend on - Smoothing results in:
[0104] Based on the above formula, in order to prove that setting independent visibility compensation for the gradient can offset the convergence effect, this disclosure provides the following proof process.
[0105] This disclosure defines the first The step error is This allows us to examine the expected change of the square of the error modulus after one iteration.
[0106] For the One parameter component (assuming local decoupling between parameters): .
[0107] Regarding the first Each parameter component, subtract the optimal solution from both sides. Square the result and take the conditional expectation. :
[0108]
[0109] remember Substitute At the same time, due to ,have The above equation can be simplified to:
[0110]
[0111] At this time, due to strong convexity Smoothing with L get:
[0112]
[0113]
[0114] Based on the above proof, the convergence factor of this disclosure is: To ensure convergence, it is usually taken that... Ignoring higher-order small quantities, the convergence speed is mainly determined by the linear terms. .when When the value is small, the convergence of this parameter will be significantly slower, reducing the overall number of iterations required for convergence. Therefore, it is necessary to set independent visibility compensation for the gradient of each parameter. This offsets the aforementioned convergence effect.
[0115] In some embodiments provided in this disclosure, determining the visibility probability of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training includes: obtaining the number of times the target Gaussian sphere is visible, wherein the number of times the target Gaussian sphere is rendered into the prediction image in the current round and the historical rounds, and the historical rounds are the rounds that have been completed in the multi-round model training; and calculating the visibility probability of the target Gaussian sphere based on the number of times it is visible and the number of target images corresponding to the multi-round model training.
[0116] Specifically, this method can obtain the number of times the target Gaussian sphere is visible from storage units (such as databases, memory, etc.), and obtain the visibility probability of the target Gaussian sphere by calculating the number of visibilitys by dividing the number of visibilitys by the number of target images corresponding to multiple rounds of model training. For example, the visibility probability of the visible Gaussian sphere is obtained by dividing the cumulative number of visibilitys of the visible Gaussian sphere by the number of sample labels.
[0117] As can be seen from the above embodiments, this method uses the objective quantitative information of visibility probability to identify which Gaussian spheres are frequently visible and contribute stably to the reconstruction quality during the model training process, and which are sparsely appearing or easily occluded edge Gaussian spheres. Subsequently, the model can be adjusted according to the visibility probability, thereby improving training efficiency and model convergence stability.
[0118] In some embodiments, obtaining the visibility count of the target Gaussian sphere includes: obtaining the index mask of the target Gaussian sphere in the current round; determining the target count position corresponding to the target Gaussian sphere from the global counting vector according to the index mask, wherein the global counting vector contains multiple count positions, and one count position is used to record the cumulative visibility count of a Gaussian sphere in the multi-round model training; updating the historical visibility count recorded in the target count position to obtain the visibility count corresponding to the current round, wherein the historical visibility count is the number of times the target Gaussian sphere has been rendered into the prediction image in the historical round.
[0119] The index mask can be understood as the aforementioned random mask matrix, used to identify the target Gaussian sphere that is visible during the current rendering round.
[0120] A global counting vector can be understood as a vector used to record the number of times a global sphere is visible. The value corresponding to the count bit of the global counting vector can be 1, 10, 50, etc., which is used to indicate that the number of times the Gaussian sphere corresponding to a count bit is visible is 1, 10, or 50.
[0121] Following the previous example, the model training method in this disclosure proposes a hybrid training optimization strategy, which mainly includes three core modules: a global visibility statistics module, an adaptive gradient scaling module, and a variance-constrained importance sampling module.
[0122] For this global visibility statistics module, a global visibility probability statistician can be constructed. During training, this global visibility probability statistician maintains a global counting vector of length N, which is equal to the total number of Gaussian spheres. .
[0123] After each forward propagation rasterization stage, obtain the visibility mask of the target Gaussian sphere currently being rendered.
[0124] The index mask is used to determine the count bit corresponding to the Gaussian sphere currently being rendered from the global count vector. Furthermore, for all visible Gaussian spheres... Update its cumulative visibility count in the counter: In other words, after the rasterization stage, this disclosure updates the Gaussian sphere by incrementing by 1. The cumulative number of times it can be viewed.
[0125] As can be seen from the above embodiments, this disclosure accurately records the cumulative number of times the target counter is visible during multiple rounds of model training by using a global counting vector. Subsequently, it is possible to quickly and accurately obtain the cumulative number of times visible, which improves the calculation efficiency of the gradient scaling factor and thus improves the model convergence speed.
[0126] Step S14: Train the target Gaussian sphere according to the gradient information and the gradient scaling factor to obtain the image processing model for the next round.
[0127] The next round of image processing model can be understood as the image processing model that needs to be trained in the next round of multi-round model training.
[0128] In some embodiments, if the current round is the last round, the image processing model for the next round is determined to be the image processing model after the multi-round model training; thus, it is determined that the current multi-round model training has been completed.
[0129] For example, if the current round is not the last round, the trained 3DGS model is determined as the 3DGS model for the next round; if the current round is the last round, the trained 3DGS model is determined as the 3DGS model after multiple rounds of model adjustment.
[0130] As can be seen from the above embodiments, this disclosure can achieve high training efficiency while ensuring reconstruction quality. After training, the 3DGS model can be directly rendered quickly through rasterization, significantly reducing the computational overhead of generating images from new perspectives.
[0131] In some embodiments, training the target Gaussian sphere according to the gradient information and the gradient scaling factor to obtain the image processing model for the next round includes: updating the gradient information using the gradient scaling factor to obtain updated gradient information of the target Gaussian sphere; and adjusting the parameters of the target Gaussian sphere according to the updated gradient information to obtain the image processing model for the next round.
[0132] Specifically, after obtaining the gradient scaling factor, this disclosure allows updating the gradient information using the gradient scaling factor to obtain the updated gradient information of the target Gaussian sphere. The gradient scaling factor can be used to update the gradient information in the following ways: multiplying the gradient scaling factor and the gradient information to obtain the updated gradient information; adding the gradient scaling factor and the gradient information to obtain the updated gradient information; or dividing the gradient scaling factor and the gradient information to obtain the updated gradient information.
[0133] After obtaining the updated gradient information, the optimizer (such as Adam) can use this updated gradient information to adjust the parameters of the target Gaussian sphere to obtain the image processing model for the next round. For example, the optimizer (such as Adam) can adjust the parameters (position, covariance, color, opacity) of the target Gaussian sphere participating in the rendering based on the updated gradient information to obtain the 3DGS model for the next round.
[0134] As can be seen from the above embodiments, updating gradient information by using a gradient scaling factor can improve the model convergence speed and avoid the overall slowdown in convergence caused by gradient information in traditional model training methods.
[0135] In some embodiments, updating the gradient information using the gradient scaling factor to obtain the updated gradient information of the target Gaussian sphere includes: when the gradient scaling factor is less than or equal to a preset scaling threshold, multiplying the gradient scaling factor of the target Gaussian sphere and the gradient information to obtain the updated gradient information of the target Gaussian sphere; and when the gradient scaling factor is greater than the preset scaling threshold, multiplying the preset scaling threshold and the gradient information of the target Gaussian sphere to obtain the updated gradient information of the target Gaussian sphere.
[0136] The preset scaling threshold can be set according to the actual application scenario. This preset scaling threshold can be understood as the maximum variance tolerance threshold. For example, the preset scaling threshold could be... (For example =10). In the embodiments provided in this disclosure, the method can update the gradient information using a gradient scaling factor and a preset scaling threshold to obtain updated gradient information. It should be noted that the preset scaling threshold in this disclosure can prevent the update step size from exceeding the L smoothness trust region, thereby improving the stability of training.
[0137] Following the previous example, after introducing visibility compensation in this disclosure, the update rule becomes: when visible, .in It is a gradient random variable under visible conditions, satisfying The variance is (Due to inconsistent perspectives).
[0138] The smoothness assumption implies a key constraint: the quadratic approximation of the loss function is only valid within a certain radius. If the single-step update... If the radius is too large, a second-order approximation failure may occur, and the loss may increase instead of decrease. This safety radius (trust region) is typically related to... Proportional. After introducing visibility compensation, the expected value of the module length update in a single step is:
[0139]
[0140] for The parameters, with the step size magnified by 100 times, cannot be guaranteed. The descent condition still holds. Therefore, after introducing visibility compensation, the L-smoothing property remains valid. When extremely small parameters fail, convex optimization theory no longer guarantees a decrease in loss, and the model enters a chaotic state (divergence). Therefore, visibility compensation is not applicable. To handle these extremely small parameters, this disclosure requires the introduction of a new method: importance sampling.
[0141] Based on this, the present disclosure can perform operations to calculate the gradient scaling factor and variance determination, and calculate the original gradient during backpropagation. Then, but before the optimizer updates the parameters, the gradient is corrected.
[0142] Before making any corrections, this disclosure requires the calculation of the gradient scaling factor. And perform variance determination.
[0143] The variance determination method is as follows: a preset maximum variance tolerance threshold is set. (For example =10), this threshold represents the maximum multiplier that can be achieved simply by increasing the step size without exceeding the L-smoothing trust domain.
[0144] Then, this disclosure corrects the gradient by implementing a hybrid compensation strategy, as follows.
[0145] according to and The Gaussian sphere is divided into two categories: "usually visible" and "usually invisible".
[0146] exist In this case, the Gaussian sphere can be divided into normally visible (case A). In this case, the compensation strategy adopted is to use visibility compensation only, specifically by directly applying the gradient of the Gaussian sphere. (i.e., gradient information) multiplied by To obtain updated gradient information .
[0147] The principle behind this operation is that the gradient remains within the trust domain after scaling, and directly amplifying the gradient is equivalent to simulating multiple training iterations with zero computational cost.
[0148] exist In this case, the Gaussian sphere can be divided into cases that are usually invisible (case B). In this case, the compensation strategy adopted is: compensation factor truncation + importance sampling. The specific operations include operation 1 (compensation factor truncation) and operation 2 (updating the sampler).
[0149] Operation 1 includes: to prevent the update step size from exceeding the L-smoothing trust region, truncating the scaling factor of visibility compensation to... In other words, using Multiply To obtain updated gradient information .
[0150] Operation 2 includes: increasing the sampling probability of the target image (i.e., patch) corresponding to the "usually invisible" parameter in the next epoch of training. times.
[0151] Based on the principle of importance sampling, through this operation, the gradient of the target image corresponding to the "usually invisible" parameter should be scaled in the next epoch of training. It falls back into the L-smooth trust domain, transforming into case A.
[0152] It should be noted that the above operations of calculating the gradient scaling factor and determining the variance, as well as the operation of executing the hybrid compensation strategy, are implemented based on the adaptive gradient scaling module and the variance-constrained importance sampling module.
[0153] As can be seen from the above embodiments, this disclosure achieves an automatic balance between the convergence speed of high and low visibility probability parameters under the premise of controllable mathematical variance, thus accelerating training in blind spots while ensuring the overall model training stability. The model training method provided by this disclosure can rapidly improve early convergence speed, with particularly outstanding results in time-sensitive reconstruction scenarios (time-sensitive scenarios require early termination of training, reducing the number of iterations from the default 30,000 rounds to 10,000 rounds). With 10,000 iterations, the PSNR can be improved by 0.5 to 1.0. Alternatively, while maintaining the same PSNR, the number of iterations can be halved (5,000 iterations).
[0154] In some embodiments, the method further includes: determining the target image of the current round and the acquisition viewpoint information corresponding to the target image of the current round from the sample set using a sampler corresponding to the image processing model, based on the sampling probability of the target image; wherein, the sampling probability is the probability that the target image is selected by the sampler in multiple rounds of model training; the sampling probability is obtained by adjusting the sampling probability of the target image in the previous multi-round model training by using the preset scaling threshold and the visibility probability in the previous multi-round model training, when the gradient scaling factor in the previous multi-round model training is greater than a preset scaling threshold.
[0155] Specifically, the sampling probability of the target image in the previous multi-round model training is obtained by adjusting the preset scaling threshold and visibility probability from the previous multi-round model training. This can be understood as follows: in the previous multi-round model training, if the gradient scaling factor is greater than the preset scaling threshold, the sampling probability of the target image in the previous multi-round model training can be amplified or reduced according to the preset scaling threshold and visibility probability from the previous multi-round model training, thereby obtaining the sampling probability corresponding to the current multi-round model training.
[0156] Following the previous example, the model training method provided in this disclosure offers a compensation strategy for importance sampling. In this case, the Gaussian sphere can be divided into normally invisible categories, and for the target image corresponding to the "normally invisible" parameter, its sampling probability can be increased in the next epoch of training. times.
[0157] Based on this, during this multi-round model training process, the sampler can adjust its parameters based on the previous multi-round model training process. magnified With a sampling probability of 1:1, select the target image and corresponding acquisition viewpoint information for the current round from the sample set (i.e., the training set); then, perform the model training operation for the current round based on the target image and corresponding acquisition viewpoint information for the current round.
[0158] As can be seen from the above embodiments, this disclosure utilizes importance sampling as a backup strategy when the visibility probability is extremely low, thereby achieving uniform and rapid convergence across the entire scenario.
[0159] The model training method provided in this disclosure can generate a predicted image using the target Gaussian sphere in the image processing model of the current round, based on the acquisition viewpoint information corresponding to the target image of the current round. Then, the gradient information of the target Gaussian sphere is determined based on the predicted image and the target image. In order to improve the model convergence speed, this method generates a dynamically adjusted gradient scaling factor based on the number of times the target Gaussian sphere is visible, and trains the target Gaussian sphere together with the gradient scaling factor and gradient information. This allows the gradient information to be dynamically adjusted during the model training process, thereby improving the model convergence speed and avoiding the problem of overall slowdown in convergence caused by gradient information in traditional model training methods, thus improving the model training efficiency.
[0160] The model training method provided in this disclosure is illustrated by taking its application in a 3D reconstruction scenario as an example. Figure 2 This is a schematic diagram illustrating the application of a model training method provided in an embodiment of this disclosure, based on... Figure 2 As can be seen, the model training method disclosed herein provides a visibility compensation training process based on variance constraints, specifically including steps S21 to S213.
[0161] Step S21: Start iterative training.
[0162] Step S22: Weighted data sampling (image k).
[0163] Specifically, the model training method disclosed herein provides a compensation strategy for importance sampling, in In this case, the Gaussian sphere can be divided into normally invisible categories, and for the images corresponding to the "normally invisible" parameter, the sampling probability can be increased in the next epoch of training. times.
[0164] Based on this, during this multi-round model training process, the sampler can adjust its parameters based on the previous multi-round model training process. magnified With a sampling probability of times, the image k (i.e., the target image) and the corresponding camera pose (i.e., the acquisition viewpoint information) for the current round are selected from the sample set (i.e., the training set).
[0165] Step S23: Forward rendering (rasterization).
[0166] Specifically, this disclosure involves inputting the camera pose into a 3DGS model.
[0167] The 3DGS model can perform adjustments on all Gaussian spheres within the 3DGS model based on the camera pose, such as frustum culling (retaining only Gaussian spheres within the field of view) and depth sorting. Furthermore, it can adjust the display parameters of the Gaussian spheres within the field of view to obtain multiple adjusted Gaussian spheres.
[0168] The Gaussian sphere within the field of view is projected onto a 2D image plane (rasterization); and then synthesized and rendered into a final RGB image (3D model image) using alpha blending.
[0169] It should be noted that, since the Gaussian sphere in the field of view may be occluded by other Gaussian spheres, this RGB image is generated from the Gaussian sphere (target Gaussian sphere) that is visible from the camera pose.
[0170] Step S24: Global visibility statistics.
[0171] Specifically, during training, the global visibility probability statistician maintains a global counting vector of the same length as the total number of Gaussian spheres N. .
[0172] After each forward propagation rasterization phase, obtain the visibility mask of the Gaussian sphere currently being rendered (the target Gaussian sphere).
[0173] The index mask is used to determine the count bit corresponding to the Gaussian sphere currently being rendered from the global count vector. Furthermore, for all visible Gaussian spheres... Update its cumulative visibility count in the counter: In other words, after the rasterization stage, this disclosure updates the Gaussian sphere by incrementing by 1. The cumulative number of times it can be viewed.
[0174] Then, the global visibility probability counter uses the global counting vector to count all globally visible Gaussian spheres. The cumulative number of times the data is viewed is recorded and stored in the database.
[0175] Step S25: Loss calculation.
[0176] Specifically, the loss value is calculated using the 3D model image and image k (the actual captured image).
[0177] Step S26: Backpropagation (original gradient g).
[0178] Specifically, a visible Gaussian sphere is calculated based on the loss value. gradient information .
[0179] Step S27: Calculate the compensation coefficient.
[0180] Modification operations on the gradient require calculation of the corresponding gradient scaling factor. This method calculates the gradient scaling factor in the following ways:
[0181] Calculate the first Visibility compensation factor for a Gaussian sphere . It is inversely proportional to its visibility probability and is defined as the gradient scaling factor. ,in, This is a smoothing term.
[0182] Set a preset maximum variance tolerance threshold. (For example =10), this threshold represents the maximum multiplier that can be achieved simply by increasing the step size without exceeding the L-smoothing trust domain.
[0183] Step S28: Variance check: S> If not, proceed to step S29; if yes, proceed to step S210.
[0184] Specifically, S equals Furthermore, in the variance check: S> If no, it indicates low variance, and step S29 is executed. If yes, it indicates high variance, and step S210 is executed.
[0185] Step S29: Gradient scaling.
[0186] Specifically, the compensation strategy adopted is to use visibility compensation only, which involves directly applying the gradient of the Gaussian sphere. (i.e., gradient information) multiplied by To obtain updated gradient information .
[0187] Step S210: Gradient truncation.
[0188] The compensation strategy adopted is: compensation factor truncation + importance sampling. The specific operations include operation 1 (compensation factor truncation) and operation 2 (updating the sampler).
[0189] Operation 1 includes: to prevent the update step size from exceeding the L-smoothing trust region, truncating the scaling factor of visibility compensation to... In other words, using Multiply To obtain updated gradient information .
[0190] Operation 2 includes: increasing the sampling probability of the target image (i.e., patch) corresponding to the "usually invisible" parameter in the next epoch of training. times.
[0191] The sampling probability can be understood as the dataset weight. After obtaining the amplified dataset weight, the importance weight can be updated (indicating insufficient feedback), and the updated importance weight (i.e., sampling probability) can be stored in the database.
[0192] Step S211: Arithmetic unit steps (updated) ).
[0193] Specifically, this disclosure obtains updated gradient information. Then, the gradient information is updated. Parameters of the Gaussian sphere Adjustments are made to complete the model training for this iteration, and then the arithmetic unit steps (i.e., the number of iterations) are updated.
[0194] Step S212: Has the maximum number of iterations been reached?
[0195] Specifically, this disclosure determines whether the number of steps has reached the maximum number of iterations. If so, step S213 is executed; otherwise, step S22 is executed.
[0196] Step S213: End training.
[0197] Based on the above steps, it can be seen that the model training method in this disclosure is a 3D Gaussian model training optimization method based on parameter visibility compensation; this disclosure relates to the fields of computer vision, 3D reconstruction and neural rendering technology, specifically to an optimization method for the training convergence speed and rendering quality of 3D Gaussian Splatting (3DGS) models.
[0198] This disclosure addresses the overall convergence slowdown problem caused by parameter visibility imbalance and proposes an effective solution. The disclosure provides a "visibility compensation under L-smoothness constraint" training framework, which dynamically adjusts the gradient scaling factor by statistically analyzing the cumulative visibility count of each Gaussian sphere, and utilizes importance sampling as a fallback strategy when the visibility probability is extremely low. This achieves uniform and rapid convergence across the entire scene, avoiding the overall convergence slowdown problem caused by parameter visibility imbalance.
[0199] Figure 3 A flowchart illustrating an image generation method provided in this disclosure. (Refer to...) Figure 3 The method includes step S31.
[0200] Step S31: Based on the target acquisition perspective information, generate a predicted image using an image processing model, wherein the image processing model is determined according to the model training method described above.
[0201] In this context, the target acquisition viewpoint information can be understood as the acquisition viewpoint information required when generating a predicted image using the image processing model during the application of the image processing model. This predicted image can be an image visible under the target acquisition viewpoint information, such as a 3D model image of the visible portion of a 3D object constructed by the image processing model under the target acquisition viewpoint information.
[0202] In some instances, generating a predicted image using an image processing model based on target acquisition viewpoint information includes: determining a target Gaussian sphere from a plurality of Gaussian spheres in the image processing model using the target acquisition viewpoint information; and generating the predicted image by rendering the image using the target Gaussian sphere, wherein the image processing model includes a plurality of Gaussian spheres, and the target Gaussian sphere is at least one Gaussian sphere corresponding to the target acquisition viewpoint information.
[0203] The image generation method provided in this disclosure utilizes the image processing model to determine the target Gaussian sphere among multiple Gaussian spheres of the image processing model that corresponds to the target acquisition viewpoint information, and generates a predicted image based on the target Gaussian sphere, thereby achieving fast rendering and significantly reducing the computational overhead of generating new viewpoint images.
[0204] The above is an illustrative scheme of an image generation method according to this embodiment. It should be noted that the technical solution of this image generation method and the technical solution of the model training method described above belong to the same concept. For details not described in detail in the technical solution of the image generation method, please refer to the description of the technical solution of the model training method described above.
[0205] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0206] In addition, this disclosure also provides a model training device, an image generation device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any of the model training devices or image generation methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding section of the method and will not be repeated here.
[0207] Figure 4 This is a block diagram of a model training device provided in an embodiment of the present disclosure.
[0208] Reference Figure 4 This disclosure provides a model training device, which includes: an image generation module 401, a gradient determination module 402, a factor determination module 403, and a model training module 404.
[0209] The image generation module 401 is configured to generate a predicted image based on the acquisition viewpoint information corresponding to the target image of the current round and using the image processing model of the current round. The predicted image is generated based on the target Gaussian sphere in the image processing model of the current round. The image processing model includes multiple Gaussian spheres, and the target Gaussian sphere is at least one Gaussian sphere corresponding to the acquisition viewpoint information. The current round is any round in the multi-round model training.
[0210] The gradient determination module 402 is configured to determine the gradient information of the target Gaussian sphere based on the predicted image and the target image.
[0211] The factor determination module 403 is configured to determine the gradient scaling factor of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training.
[0212] The model training module 404 is configured to train the target Gaussian sphere based on the gradient information and the gradient scaling factor to obtain the image processing model for the next round.
[0213] In some embodiments, the model training module 404 is further configured to: update the gradient information using the gradient scaling factor to obtain updated gradient information of the target Gaussian sphere; and adjust the parameters of the target Gaussian sphere according to the updated gradient information to obtain the image processing model for the next round.
[0214] In some embodiments, the model training module 404 is further configured to: when the gradient scaling factor is less than or equal to a preset scaling threshold, multiply the gradient scaling factor of the target Gaussian sphere with the gradient information to obtain the updated gradient information of the target Gaussian sphere; and when the gradient scaling factor is greater than the preset scaling threshold, multiply the preset scaling threshold with the gradient information of the target Gaussian sphere to obtain the updated gradient information of the target Gaussian sphere.
[0215] In some embodiments, the factor determination module 403 is further configured to: determine the visibility probability of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training; and determine the gradient scaling factor of the target Gaussian sphere based on the visibility probability.
[0216] In some embodiments, the factor determination module 403 is further configured to: obtain the number of times the target Gaussian sphere is visible, wherein the number of times the target Gaussian sphere is rendered into the prediction image in the current round and the historical round, and the historical round is the round that has been completed in the multi-round model training; and calculate the visibility probability of the target Gaussian sphere based on the number of times it is visible and the number of target images corresponding to the multi-round model training.
[0217] In some embodiments, the factor determination module 403 is further configured to: obtain the index mask of the target Gaussian sphere in the current round; determine the target count position corresponding to the target Gaussian sphere from the global counting vector according to the index mask, wherein the global counting vector contains multiple count positions, and one count position is used to record the cumulative number of times a Gaussian sphere is visible in the multi-round model training; update the historical visibility count recorded in the target count position to obtain the visibility count corresponding to the current round, wherein the historical visibility count is the number of times the target Gaussian sphere is rendered into the prediction image in the historical round.
[0218] In some embodiments, the model training apparatus further includes a sampling module configured to: determine the target image of the current round and the acquisition viewpoint information corresponding to the target image of the current round from the sample set using a sampler corresponding to the image processing model, based on the sampling probability of the target image; wherein, the sampling probability is the probability that the target image is selected by the sampler in multiple rounds of model training; the sampling probability is obtained by adjusting the sampling probability of the target image in the previous multi-round model training by using the preset scaling threshold and the visibility probability in the previous multi-round model training, when the gradient scaling factor in the previous multi-round model training is greater than a preset scaling threshold.
[0219] In some embodiments, the image generation module 401 is further configured to: determine the target Gaussian sphere from multiple Gaussian spheres of the image processing model of the current round using the acquired viewpoint information; and perform image rendering using the target Gaussian sphere to generate the predicted image.
[0220] In some embodiments, the predicted image is a three-dimensional model image; the image generation module 401 is further configured to: adjust the parameters of a plurality of Gaussian spheres of the image processing model of the current round using the acquisition viewpoint information to obtain a plurality of adjusted Gaussian spheres, wherein the plurality of adjusted Gaussian spheres are used to construct a three-dimensional model corresponding to the target image; determine a plurality of target Gaussian spheres visible under the acquisition viewpoint information from the plurality of adjusted Gaussian spheres; and perform image rendering based on the plurality of target Gaussian spheres to obtain a three-dimensional model image of the three-dimensional model visible under the acquisition viewpoint information.
[0221] In some embodiments, the model training apparatus further includes: if the current round is the last round, determining the image processing model for the next round as the image processing model after the multi-round model training.
[0222] The model training apparatus provided in this embodiment can generate a predicted image using the target Gaussian sphere in the image processing model of the current round, based on the acquisition viewpoint information corresponding to the target image of the current round. Then, it determines the gradient information of the target Gaussian sphere based on the predicted image and the target image. In order to improve the model convergence speed, this method generates a dynamically adjusted gradient scaling factor based on the number of times the target Gaussian sphere is visible, and trains the target Gaussian sphere together with the gradient scaling factor and gradient information. This allows the gradient information to be dynamically adjusted during the model training process, thereby improving the model convergence speed and avoiding the problem of overall slowdown in convergence caused by gradient information in traditional model training methods, thus improving the model training efficiency.
[0223] The above is an illustrative scheme of a model training device according to this embodiment. It should be noted that the technical solution of this model training device and the technical solution of the model training method described above belong to the same concept. For details not described in detail in the technical solution of the model training device, please refer to the description of the technical solution of the model training method described above.
[0224] Figure 5 This is a block diagram of an image generation apparatus provided in an embodiment of the present disclosure.
[0225] Reference Figure 5 This disclosure provides an image generation apparatus, which includes an image generation module 501. The image generation module is configured to generate a predicted image using an image processing model based on target acquisition viewpoint information, wherein the image processing model is determined according to the model training method described above.
[0226] In some embodiments, the image generation module 501 is further configured to: determine a target Gaussian sphere from a plurality of Gaussian spheres in the image processing model using the target acquisition viewpoint information; and generate the predicted image by rendering the image through the target Gaussian sphere, wherein the image processing model includes a plurality of Gaussian spheres, and the target Gaussian sphere is at least one Gaussian sphere corresponding to the target acquisition viewpoint information.
[0227] The image generation apparatus provided in this embodiment utilizes the image processing model to determine the target Gaussian sphere among multiple Gaussian spheres of the image processing model that corresponds to the target acquisition viewpoint information, and generates a predicted image based on the target Gaussian sphere, thereby achieving fast rendering and significantly reducing the computational overhead of generating new viewpoint images.
[0228] The above is an illustrative scheme of an image generation apparatus according to this embodiment. It should be noted that the technical solution of this image generation apparatus and the technical solution of the image generation method described above belong to the same concept. For details not described in detail in the technical solution of the image generation apparatus, please refer to the description of the technical solution of the image generation method described above.
[0229] Figure 6 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0230] Reference Figure 6 This disclosure provides an electronic device, which includes: at least one processor 601; at least one memory 602; and one or more I / O interfaces 603 connected between the processor 601 and the memory 602; wherein the memory 602 stores one or more computer programs that can be executed by the at least one processor 601, and the one or more computer programs are executed by the at least one processor 601 to enable the at least one processor 601 to perform the above-described model training method or image generation method.
[0231] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the model training method or image generation method described above. The computer-readable storage medium may be volatile or non-volatile.
[0232] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described model training method or image generation method.
[0233] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0234] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0235] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0236] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0237] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0238] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0239] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0240] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0241] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0242] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A model training method, characterized in that, include: Based on the acquisition viewpoint information corresponding to the target image in the current round, a predicted image is generated using the image processing model of the current round. The predicted image is generated based on the target Gaussian sphere in the image processing model of the current round. The image processing model includes multiple Gaussian spheres, and the target Gaussian sphere is at least one Gaussian sphere corresponding to the acquisition viewpoint information. The current round is any round in a multi-round model training. Based on the predicted image and the target image, determine the gradient information of the target Gaussian sphere; The gradient scaling factor of the target Gaussian sphere is determined based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training. The target Gaussian sphere is trained based on the gradient information and the gradient scaling factor to obtain the image processing model for the next round.
2. The method according to claim 1, characterized in that, The step of training the target Gaussian sphere based on the gradient information and the gradient scaling factor to obtain the image processing model for the next round includes: The gradient information is updated using the gradient scaling factor to obtain the updated gradient information of the target Gaussian sphere; The parameters of the target Gaussian sphere are adjusted based on the updated gradient information to obtain the image processing model for the next round.
3. The method according to claim 2, characterized in that, The step of updating the gradient information using the gradient scaling factor to obtain the updated gradient information of the target Gaussian sphere includes: If the gradient scaling factor is less than or equal to a preset scaling threshold, the gradient scaling factor of the target Gaussian sphere and the gradient information are multiplied together to obtain the updated gradient information of the target Gaussian sphere. If the gradient scaling factor is greater than the preset scaling threshold, the preset scaling threshold is multiplied by the gradient information of the target Gaussian sphere to obtain the updated gradient information of the target Gaussian sphere.
4. The method according to any one of claims 1 to 3, characterized in that, The step of determining the gradient scaling factor of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training includes: The visibility probability of the target Gaussian sphere is determined based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training. Based on the visibility probability, the gradient scaling factor of the target Gaussian sphere is determined.
5. The method according to claim 4, characterized in that, Determining the visibility probability of the target Gaussian sphere based on the number of times the target Gaussian sphere is rendered into the prediction image during the multi-round model training includes: Obtain the number of times the target Gaussian sphere is visible, wherein the number of times the target Gaussian sphere is rendered into the prediction image in the current round and the historical round, and the historical round is the round that has been completed in the multi-round model training; The visibility probability of the target Gaussian sphere is calculated based on the number of visibilitys and the number of target images corresponding to the multi-round model training.
6. The method according to claim 5, characterized in that, The process of obtaining the number of times the target Gaussian sphere is visible includes: Get the index mask of the target Gaussian ball in the current round; Based on the index mask, the target count bit corresponding to the target Gaussian sphere is determined from the global count vector, wherein the global count vector contains multiple count bits, and one count bit is used to record the cumulative number of times a Gaussian sphere is visible in the multi-round model training; The historical visibility count recorded in the target counter is updated to obtain the visibility count corresponding to the current round, wherein the historical visibility count is the number of times the target Gaussian sphere is rendered into the prediction image in the historical round.
7. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Based on the sampling probability of the target image, the sampler corresponding to the image processing model is used to determine the target image of the current round and the acquisition viewpoint information corresponding to the target image of the current round from the sample set; Wherein, the sampling probability is the probability that the target image is selected by the sampler during multiple rounds of model training; The sampling probability is obtained by adjusting the sampling probability of the target image in the previous multi-round model training, using the preset scaling threshold and the visibility probability, when the gradient scaling factor in the previous multi-round model training is greater than a preset scaling threshold.
8. The method according to claim 1, characterized in that, The step of generating a predicted image based on the acquisition viewpoint information corresponding to the target image of the current round and using the image processing model of the current round includes: Using the acquired perspective information, the target Gaussian sphere is determined from multiple Gaussian spheres in the current round's image processing model; The predicted image is generated by rendering the image using the target Gaussian sphere.
9. The method according to claim 8, characterized in that, The predicted image is a three-dimensional model image; The step of determining the target Gaussian sphere from multiple Gaussian spheres in the current round of image processing model using the acquired viewpoint information includes: Using the acquired perspective information, the parameters of multiple Gaussian spheres of the current round image processing model are adjusted to obtain multiple adjusted Gaussian spheres, wherein the multiple adjusted Gaussian spheres are used to construct a three-dimensional model corresponding to the target image; From the adjusted multiple Gaussian spheres, identify multiple target Gaussian spheres that are visible under the acquired viewing angle information; The step of rendering the image using the target Gaussian sphere to generate the predicted image includes: Image rendering is performed on the multiple target Gaussian spheres to obtain a 3D model image of the 3D model visible under the acquired viewpoint information.
10. The method according to any one of claims 1 to 3, characterized in that, The method further includes: If the current round is the last round, the image processing model for the next round is determined to be the image processing model after the multi-round model training.
11. An image generation method, characterized in that, include: Based on the target acquisition perspective information, a predicted image is generated using an image processing model, wherein the image processing model is determined according to the model training method of any one of claims 1-10.
12. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-11.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1-11.
14. A computer program product, characterized in that, Includes computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the method as described in any one of claims 1-11.