Method, device and equipment for training image generation model

CN122820906APending Publication Date: 2026-09-25BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611305406.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,自然语言在描述精确的空间几何关系方面存在天然的模糊性

Benefits of technology

[0009]通过上述技术方案,获取包括相机视点的样本位姿参数、样本描述文本和参考样本图像的训练样本;对样本位姿参数进行编码处理得到样本视点嵌入向量;确定样本位姿参数对应的视点令牌的目标数量并生成目标数量的样本视点令牌;根据样本位姿参数构建用于编码背景场景透视几何关系的样本场景上下文令牌;将样本视点令牌和样本场景上下文令牌与样本描述文本进行融合得到样本提示信息;将样本提示信息输入至图像生成模型生成预测图像;基于预测图像和参考样本图像获得模型损失,基于模型损失调整图像生成模型的模型参数,得到训练后的图像生成模型。通过该方法,实现了令牌数量与视点信息量的精准匹配,并通过场景上下文令牌实现了前背景几何解耦,使训练得到的图像生成模型能够根据相机视点的位姿参数精确控制生成图像的视角,同时保证前景物体与背景场景的几何一致性,提升了图像生成的精度和视觉真实感。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820906A_ABST
    Figure CN122820906A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device and equipment for training an image generation model. The method comprises: obtaining a training sample, which comprises a sample pose parameter of a camera viewpoint, a sample description text and a reference sample image; encoding the sample pose parameter to obtain a sample viewpoint embedding vector; determining a target number of viewpoint tokens and generating a corresponding number of sample viewpoint tokens; constructing a sample scene context token for encoding background perspective geometric relationship according to the sample pose parameter; inputting the sample viewpoint token and the sample scene context token and the sample description text into the image generation model after fusing to obtain sample prompt information, generating a predicted image; obtaining a model loss based on the predicted image and the reference sample image and adjusting the model parameters. Through the method, the trained image generation model can accurately control the viewpoint of the generated image according to the camera pose parameter and ensure the consistency of the foreground and background geometry, solving the problems of ambiguous viewpoint control and geometric inconsistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically, to a training method, apparatus, and device for an image generation model. Background Technology

[0002] Text-to-image generation technology has made significant breakthroughs in recent years, enabling the generation of high-quality, highly realistic images simply by inputting natural language descriptions. However, natural language inherently possesses ambiguity in describing precise spatial geometric relationships. When a user desires to generate a vehicle viewed from a 30-degree top-down angle on the left, current mainstream text-to-image models often struggle to accurately understand and execute such viewpoint instructions, frequently generating images with incorrect poses, default frontal perspectives, or geometrically inconsistent results. Summary of the Invention

[0003] To address the aforementioned technical issues, this disclosure provides a training method, apparatus, and device for an image generation model, which can precisely control the viewing angle of the generated image based on the pose parameters of the camera viewpoint, thereby generating an image that meets the user's viewing angle requirements.

[0004] To achieve the above objectives, in a first aspect, this disclosure provides a method for training an image generation model, the method comprising: Acquire training samples, which include sample description text, sample pose parameters of the camera viewpoint, and reference sample images corresponding to the sample pose parameters and sample description text; The sample pose parameters are encoded to obtain the sample viewpoint embedding vector; Determine the target number of viewpoint tokens corresponding to the sample pose parameters, and generate the target number of sample viewpoint tokens based on the sample viewpoint embedding vector; A sample scene context token is constructed based on the sample pose parameters. The sample scene context token is used to encode the perspective geometry of the background scene. The sample viewpoint token and the sample scene context token are fused with the sample description text to obtain sample prompt information; The sample prompt information is input into the image generation model to generate a predicted image; The model loss is obtained based on the predicted image and the reference sample image. The model parameters of the image generation model are adjusted based on the model loss to obtain the trained image generation model.

[0005] Secondly, this disclosure provides a training apparatus for an image generation model, the apparatus comprising: The data acquisition module is used to acquire training samples, which include sample description text, sample pose parameters of the camera viewpoint, and reference sample images corresponding to the sample pose parameters and sample description text. The encoding module is used to encode the sample pose parameters to obtain the sample viewpoint embedding vector; The token generation module is used to determine the target number of viewpoint tokens corresponding to the sample pose parameters, and generate the target number of sample viewpoint tokens based on the sample viewpoint embedding vector. The token generation module is further configured to construct a sample scene context token based on the sample pose parameters, wherein the sample scene context token is used to encode the perspective geometric relationships of the background scene. The fusion module is used to fuse the sample viewpoint token and the sample scene context token with the sample description text to obtain sample prompt information; The image generation module is used to input the sample prompt information into the image generation model to generate a predicted image; The model training module is used to obtain the model loss based on the predicted image and the reference sample image, and to adjust the model parameters of the image generation model based on the model loss to obtain the trained image generation model.

[0006] Thirdly, this disclosure provides an electronic device, including: Processor; and A memory storing a computer program or readable instructions, which, when executed by the processor, implement the training method for the image generation model as described above.

[0007] Fourthly, this disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described training method for the image generation model.

[0008] Fifthly, this disclosure provides a computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the above-described training method for the image generation model.

[0009] The above technical solution acquires training samples including sample pose parameters of the camera viewpoint, sample descriptive text, and reference sample images; encodes the sample pose parameters to obtain sample viewpoint embedding vectors; determines the target number of viewpoint tokens corresponding to the sample pose parameters and generates the target number of sample viewpoint tokens; constructs sample scene context tokens to encode the perspective geometry of the background scene based on the sample pose parameters; fuses the sample viewpoint tokens and sample scene context tokens with the sample descriptive text to obtain sample prompt information; inputs the sample prompt information into the image generation model to generate a predicted image; obtains the model loss based on the predicted image and the reference sample image; adjusts the model parameters of the image generation model based on the model loss to obtain the trained image generation model. This method achieves precise matching between the number of tokens and the amount of viewpoint information, and decouples the foreground and background geometry through scene context tokens. This allows the trained image generation model to precisely control the viewpoint of the generated image based on the camera viewpoint pose parameters, while ensuring the geometric consistency between the foreground object and the background scene, thus improving the accuracy and visual realism of image generation.

[0010] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0011] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a training method for an image generation model according to an exemplary embodiment of the present disclosure.

[0012] Figure 2 This is another schematic flowchart illustrating a training method for an image generation model according to an exemplary embodiment of the present disclosure.

[0013] Figure 3 This is a block diagram of a training apparatus for an image generation model according to an exemplary embodiment of the present disclosure.

[0014] Figure 4 This is a schematic diagram of the structure of an electronic device for performing a training method for an image generation model, according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0015] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.

[0016] Figure 1This is a flowchart illustrating a training method for an image generation model according to an exemplary embodiment of this disclosure. The method can be executed by an electronic device, which can be a terminal device or a server; that is, the method can be executed by the terminal device or the server alone, or by the terminal device and the server working together.

[0017] Taking the server as the execution entity as an example, the specific implementation process of this method is as follows: Step S110: Obtain training samples.

[0018] The training samples include sample description text, sample pose parameters of the camera viewpoint, and reference sample images corresponding to the sample pose parameters and sample description text.

[0019] The sample text description is used to represent the semantic content information of the images in the training samples, including the category, attributes, and action state of foreground objects, as well as the type and spatial layout of the background scene. For example, the sample description text could be a white cat lying on a sofa with a potted green plant next to it.

[0020] Sample pose parameters are used to characterize the camera's position and orientation in three-dimensional space. In one possible implementation, the sample pose parameters include at least two of the following parameters: azimuth, elevation, distance, pitch, and yaw. The azimuth represents the camera's horizontal position relative to the object, the elevation represents the camera's vertical position relative to the object, the distance represents the distance between the camera and the object, the pitch represents the up-down orientation of the camera's optical axis (i.e., the camera's head orientation), and the yaw represents the camera's rotation about its optical axis (i.e., the camera's tilt).

[0021] In one possible implementation, the sample pose parameters include azimuth, elevation, distance, pitch, and yaw.

[0022] To ensure semantic consistency in the training data, an object-centered coordinate system can be used, stipulating that all objects face the positive direction of the coordinate system, ensuring that concepts such as "left side" and "back side" have consistent meanings for all objects. In one possible implementation, the azimuth angle can be encoded using both sine and cosine values ​​to eliminate periodic angle jumps. All parameters are uniformly normalized to a reasonable numerical range to ensure the stability of subsequent training of the image generation model.

[0023] In one possible implementation, training samples can be constructed by: selecting objects with clear semantic frontal orientation from a 3D asset library, covering categories such as animals, vehicles, people, and furniture; uniformly sampling multiple camera viewpoints for each 3D object, covering the range of 0-360 degrees azimuth and 0-45 degrees elevation, generating multiple images for each object; and rendering results with precisely known camera parameters to provide standard-answer level geometric supervision signals for model learning.

[0024] In another possible implementation, the training samples can also include realism enhancement data: high-quality objects are selected from a 3D asset library, rendered from multiple perspectives, and then image editing models are used to remove artificial artifacts from the 3D rendering, adding realistic appearance details and diverse backgrounds; unsatisfactory enhancement results are filtered out through quality screening, retaining high-quality images containing realistic background scenes, which are the core data for calculating the foreground-background geometric consistency loss. During training, samples can be taken in equal proportions from both rendered data and realism data.

[0025] Step S120: Encode the sample pose parameters to obtain the sample viewpoint embedding vector.

[0026] In one possible implementation, a multilayer perceptron (MLP) network can be used to map the sample pose parameters into embedding vectors of the same dimension as the text words in the sample description text, thus obtaining the sample viewpoint embedding vector.

[0027] Specifically, a lightweight 3-layer MLP network can be designed as the viewpoint base encoding network to map 6-dimensional geometric parameters (the sine and cosine values ​​of the azimuth angle are encoded separately, plus the elevation angle, distance, pitch angle, and yaw angle, for a total of 6 dimensions) to the same dimension as the text vocabulary, thus obtaining the sample viewpoint embedding vector.

[0028] It is worth mentioning that this MLP network has a very small number of parameters (only about a few million), which can be used as a network model in an image generation model. Compared with the backbone network used for image generation (billions of parameters), the number is negligible, and the training efficiency is extremely high.

[0029] Step S130: Determine the target number of viewpoint tokens corresponding to the sample pose parameters, and generate the target number of sample viewpoint tokens based on the sample viewpoint embedding vector.

[0030] In one possible implementation, determining the target quantity includes: inputting sample pose parameters into a multilayer perceptron network (which acts as a learnable complexity predictor), obtaining probability distributions corresponding to multiple token quantity categories through multilayer perceptron prediction, and determining the target quantity of sample viewpoint tokens corresponding to the sample pose parameters based on the probability distributions corresponding to multiple token quantity categories.

[0031] For example, the Learnable Complexity Predictor (LCP) is a standalone 2-layer MLP network. Its input is a sample viewpoint embedding vector (e.g., the previously extracted 6-dimensional geometric parameters), and its output is a probability distribution for three categories, corresponding to the probabilities of using 1, 2, and 3 viewpoint tokens, respectively. The number of tokens corresponding to the highest probability is determined as the target number.

[0032] In one possible implementation, generating a target number of sample viewpoint tokens based on the sample viewpoint embedding vector includes: When there is only one target, the sample viewpoint embedding vector is used as the unique sample viewpoint token.

[0033] When there are multiple targets, a hierarchical generation strategy is used to generate multiple sample viewpoint tokens based on the sample viewpoint embedding vectors. Specifically, when there are three targets, an attention mechanism is used to generate a first sample viewpoint token based on the sample viewpoint embedding vectors. This first sample viewpoint token encodes macroscopic spatial location information (such as the camera's azimuth and elevation angles). Based on the first sample viewpoint token, a second sample viewpoint token is generated through a cross-attention mechanism. This second sample viewpoint token encodes incremental geometric information relative to the first sample viewpoint token (such as rotation parameters like pitch and yaw angles). Based on the first and second sample viewpoint tokens, a third sample viewpoint token is generated through a cross-attention mechanism. This third sample viewpoint token encodes residual geometric information relative to the first and second sample viewpoint tokens (such as residual offset information between the viewpoint and the standard reference position).

[0034] It is worth mentioning that when there are two targets, the first sample viewpoint token and the second sample viewpoint token are generated using the above generation method.

[0035] In one possible implementation, generating a second sample viewpoint token based on a first sample viewpoint token using a cross-attention mechanism includes: using the first sample viewpoint token as a key-value pair and the sample viewpoint embedding vector as a query, and generating the second sample viewpoint token through the cross-attention mechanism. Generating a third sample viewpoint token based on the first and second sample viewpoint tokens using a cross-attention mechanism includes: using both the first and second sample viewpoint tokens as a key-value pair and the sample viewpoint embedding vector as a query, and generating the third sample viewpoint token through the cross-attention mechanism.

[0036] The hierarchical generation method described above ensures that each token is incrementally encoded based on the information expressed by the previous token, avoiding the duplication of encoding the same geometric dimensions by different tokens, while also guaranteeing the semantic hierarchy of the token sequence (from coarse-grained to fine-grained). All generated tokens are inserted into the text sequence in hierarchical order, and the attention mechanism of the image generation model can automatically aggregate geometric constraints at different levels to achieve precise viewpoint control.

[0037] Step S140: Construct a sample scene context token based on the sample pose parameters.

[0038] Among them, the sample scene context token is used to encode the perspective geometry of the background scene.

[0039] In one possible implementation, the construction of the Scene Context Token (SCT) includes: deriving key geometric attributes of the background scene based on sample pose parameters, the key geometric attributes including at least one of the horizon height position, vanishing perspective direction, and main light and shadow direction; and mapping the key geometric attributes to the scene context token through a multilayer perceptron network.

[0040] Specifically, the relative height of the horizon in the image can be estimated based on the elevation angle parameter (the higher the elevation angle, the lower the horizon), the direction of the vanishing point of perspective can be calculated based on the azimuth and distance parameters, and the main direction of the background light and shadow (i.e. the angle at which the object's shadow is projected on the background) can be calculated based on the combination of elevation and pitch angles. The above key geometric attributes are then mapped into scene context tokens of the same dimension as the viewpoint tokens using a lightweight MLP.

[0041] Step S150: Fuse the sample viewpoint token and sample scene context token with the sample description text to obtain sample prompt information.

[0042] In one possible implementation, the viewpoint token can be associated with descriptive words about foreground objects in the description text, and the scene context token can be associated with descriptive words about the background scene in the description text. The associated viewpoint token, scene context token, and description text can then be concatenated to obtain sample prompt information.

[0043] When scene context tokens and sample viewpoint tokens are inserted into the sample description text, the sample viewpoint token is adjacent to the object description words in the sample description text, and the scene context token is adjacent to the scene / background description words in the sample description text. Semantic separation has been achieved from the text position, and regional division of labor is realized in conjunction with the attention mechanism.

[0044] Step S160: Input the sample prompt information into the image generation model to generate the predicted image.

[0045] Optionally, the image generation model may include a viewpoint-based coding network (the aforementioned lightweight 3-layer MLP network), a learnable complexity predictor (a 2-layer MLP network), and a diffusion model (such as mainstream diffusion models like SD, SDXL, and Flux). Sample cue information is input as a condition to the diffusion model, and a predicted image is generated through a denoising process. The predicted image is a complete scene image containing both objects and the background, rather than a cropped image of a local area of ​​the object.

[0046] In the denoising process of the diffusion model, attention weights for sample viewpoint tokens and sample scene context tokens can be assigned through region-aware attention routing. Region-aware attention routing includes: predicting the confidence map of whether each spatial location in the current latent variable features belongs to the foreground or background region; based on the confidence map, assigning high weights to sample viewpoint tokens and low weights to sample scene context tokens for latent variable features in the foreground region, and assigning high weights to sample scene context tokens and low weights to sample viewpoint tokens for latent variable features in the background region.

[0047] Specifically, in the denoising process of the diffusion model, for the latent variable features at each spatial location, a lightweight semantic classification head (running in the latent space, with only 2 convolutional layers) predicts whether the location belongs to a foreground object region or a background scene region, outputting a foreground confidence map F_map. Based on F_map, the attention weights of the viewpoint token and scene context token are soft-weighted: latent variable features of the foreground region are given high weight by F_map to focus on the viewpoint token (camera geometric constraints), while being given low weight by (1-F_map) to focus on the scene context token; the opposite applies to the background region. This ensures that foreground objects are primarily controlled by the camera viewpoint, while the background scene is primarily constrained by perspective geometry. Region-aware routing is implemented with soft constraints, without breaking the end-to-end differentiability of the attention mechanism, and can directly participate in gradient backpropagation training without requiring additional segmentation and annotation data.

[0048] Step S170: Obtain the model loss based on the predicted image and the reference sample image, and adjust the model parameters of the image generation model based on the model loss to obtain the trained image generation model.

[0049] In one possible implementation, obtaining the model loss includes: calculating a diffusion loss based on the predicted image and a reference sample image; calculating a token orthogonal constraint loss based on the information redundancy between viewpoint tokens of the target number; calculating a foreground-background geometric consistency loss based on the geometric consistency between the foreground and background regions of the predicted image; and obtaining the model loss based on at least one of the diffusion loss, the token orthogonal constraint loss, and the foreground-background geometric consistency loss.

[0050] For example, the model loss can be expressed as L_total = L_diffusion + λ1·L_orth + λ2·L_geo, where L_diffusion is the difference diffusion loss, L_orth is the token orthogonality constraint loss, and L_geo is the foreground-background geometric consistency loss. λ1 = 0.1 (orthogonality constraint weight, ensuring it does not affect the main generation quality), λ2 = 0.05 (geometric consistency weight, only effective on realistic augmented data); a warm-up phase is set up in the early stage of training, where only L_diffusion is optimized in the first 1000 steps. After the viewpoint tokens are basically learned and stabilized, the auxiliary loss is introduced to avoid instability in the early training stage.

[0051] In one possible implementation, the foreground-background geometric consistency loss L_geo includes at least one of the shadow direction consistency loss L_shadow and the vanishing point alignment loss L_vp. Specifically, the shadow direction consistency loss can be obtained by: analytically calculating the theoretical shadow projection angle θ_shadow of the object on the background based on camera parameters (elevation angle, azimuth angle) and the default light source assumption (top light); detecting the actual shadow direction θ_actual in the background region of the generated image using image gradients; and obtaining the shadow direction consistency loss as L_shadow = ||θ_shadow - θ_actual|| using the following formula, thus prompting the model to generate physically correct shadows. The vanishing point alignment loss can be obtained by: calculating the theoretical vanishing point position p_theory where horizontal lines (such as grout lines in the ground or building edges) in the scene should converge based on the camera's pitch and yaw angles; detecting straight lines in the background region of the generated image using Hough transform and estimating the actual vanishing point position p_actual; and calculating the vanishing point alignment loss L_vp = ||p_theory - p_actual||² using the following formula.

[0052] In this case, the foreground-background geometric consistency loss is L_geo = α·L_shadow + β·L_vp. The background geometric consistency loss L_geo can constrain the geometric plausibility of the generated image's background. The geometric consistency loss is only calculated on realism-enhanced data (images with realistic backgrounds); pure 3D rendering data (transparent backgrounds) is not included in the optimization of this loss.

[0053] In one possible implementation, a token orthogonality constraint loss can be introduced into the model loss, defined as the sum of squares of the cosine similarities between any two viewpoint tokens. A smaller loss indicates better orthogonality of the information encoded by each token. This token orthogonality constraint loss is added to the total loss function (weight λ=0.1) as a regularization term to constrain different tokens to actively learn and encode different geometric dimensions, preventing information redundancy between tokens. The orthogonality constraint works synergistically with the hierarchical attention mechanism: hierarchical attention ensures incremental encoding in the generation structure, while the orthogonality constraint further strengthens the information differentiation of each token at the loss level.

[0054] In one possible implementation, the model parameters of the image generation model can be adjusted based on the model loss. Specifically, a hierarchical learning rate strategy can be used to update the parameters of the image generation model. Specifically, a first learning rate (e.g., 2×10⁻⁶) can be used for the learnable complexity predictor and the multilayer perceptron network used to encode pose parameters in the image generation model. -4 For diffusion models, a second learning rate (e.g., 2×10) is used. -5 The region-aware attention routing module uses a third learning rate (e.g., 1×10). -4 To update the model parameters, a hierarchical learning rate strategy can be adopted to ensure that each network can be updated stably.

[0055] In one possible implementation, during the training of the image generation model, in the early stages of training, such as the first preset iterations (e.g., the first 1000 iterations), only the difference-based diffusion loss is used to iteratively train the image generation model. After the preset number of iterations has been exceeded, token orthogonality constraint loss and foreground-background geometric consistency loss are introduced to participate in the model training. Furthermore, the sample images used during training are complete scene images containing both objects and background, rather than cropped images of local object regions, enabling the image generation model to learn the spatial relationships between objects and background from different viewpoints. Cross-object category generalization enhancement: The training data covers a variety of different objects, and the training samples are randomly shuffled across different categories to prevent overfitting of viewpoint information to the appearance of specific objects.

[0056] This application provides a training method for an image generation model. It encodes multi-parameter geometric features of the camera viewpoint into viewpoint embeddings, adaptively generates hierarchical viewpoint tokens using a learnable complexity predictor to match the information complexity of different viewpoints, and simultaneously constructs scene context tokens to encode background perspective geometry. This is combined with region-aware attention routing to decouple foreground projection geometry from background perspective geometry. Furthermore, it incorporates token orthogonality constraints and foreground-background geometric consistency loss for multi-objective joint training. This enables the trained image generation model to precisely control the viewpoint of the generated image based on camera pose parameters, while ensuring that foreground object deformation and background spatial perspective are coordinated and unified. This fundamentally solves the technical problems of ambiguous natural language viewpoint descriptions, limited expression with a fixed number of tokens, and inconsistencies between foreground and background geometry in existing text-to-image generation methods, significantly improving the viewpoint accuracy, geometric realism, and cross-class generalization ability of image generation.

[0057] Optionally, after training the image generation model through steps S110-S170 above, the trained image generation model can be used to perform the following steps to generate an image: The method further includes: Step S210: Obtain the pose parameters and description text of the camera viewpoint.

[0058] Step S220: Encode the pose parameters to obtain the viewpoint embedding vector.

[0059] Step S230: Determine the number of viewpoint tokens based on the pose parameters, and generate the corresponding number of viewpoint tokens based on the viewpoint embedding vector.

[0060] Step S240: Construct a scene context token based on the pose information.

[0061] Step S250: Combine the viewpoint token, scene context token, and description text to obtain the prompt information.

[0062] Step S260: Input the prompt information into the trained image generation model to generate the target image.

[0063] For details on the specific implementation principles of steps S210-S260, please refer to the description of steps S110-S160 in the foregoing embodiments, which will not be repeated here.

[0064] It is worth mentioning that during the inference phase, the learnable complexity predictor can directly take the argmax of its output to obtain the number of tokens, without the need for Gumbel-Softmax relaxation processing.

[0065] Through the above steps S210-S250, the pose parameters and descriptive text of the camera viewpoint are obtained. The pose parameters are encoded into viewpoint embeddings. The number of tokens is determined by a learnable complexity predictor and a corresponding number of viewpoint tokens are generated. At the same time, scene context tokens are constructed. The viewpoint tokens and scene context tokens are fused with the foreground and background related descriptions in the descriptive text, respectively, and then input into the trained image generation model to generate a target image corresponding to the camera viewpoint that meets the user's needs. This achieves high-efficiency inference speed while ensuring the accuracy of viewpoint control.

[0066] Optionally, the image generation model described above can be used to generate continuous images (i.e., videos). In this case, the training samples may include a sequence of sample pose parameters corresponding to multiple continuously changing camera viewpoints in the viewpoint trajectory, sample description text, and a sequence of reference sample images corresponding to multiple continuously changing camera viewpoints and sample description text. Step S150 may specifically encode the differences in sample pose parameters between adjacent viewpoints to obtain inter-frame relative change codes. The sample viewpoint tokens, inter-frame relative change codes, and sample description text of each camera viewpoint in the viewpoint trajectory are fused to obtain video sample prompt information. Step S160 may specifically input the video sample prompt information into the image generation model to generate a predicted image sequence, which includes a predicted image corresponding to each camera viewpoint.

[0067] Specifically, when encoding the differences in pose parameters between adjacent viewpoints to obtain inter-frame relative change codes, for the pose parameters of adjacent frames t and t-1 in the viewpoint trajectory, the differences between them in azimuth, elevation, distance, pitch, and yaw are calculated. This difference vector is input into a lightweight MLP network and mapped to an inter-frame relative change code of the same dimension as the viewpoint token. Further, when fusing the sample viewpoint tokens, inter-frame relative change codes, and sample description text for each camera viewpoint in the viewpoint trajectory to obtain video sample prompt information, specifically, for each viewpoint, its corresponding sample viewpoint token and inter-frame relative change code are concatenated along the sequence dimension, and then concatenated or fused with the embedding vector of the sample description text through cross-attention to form complete prompt information containing frame-by-frame viewpoint information and inter-frame change information. The video sample prompt information is input into an image generation model to generate a predicted image sequence, which includes predicted images corresponding to each camera viewpoint.

[0068] It's also worth mentioning that during the training phase, sampling consecutive multi-frame 3D rendering sequences (such as video clips rotating around an object) allows the model to learn the smooth transition patterns of image content as the viewpoint changes continuously. Temporal consistency loss can be introduced: the difference in viewpoint tokens between adjacent frames should be proportional to the difference in image features to prevent image jumps caused by abrupt viewpoint changes; the inter-frame change rate of scene context tokens should match the camera motion rate to prevent abrupt changes in background perspective. Contrastive learning ensures the consistency of the same object's identity across different viewpoints, maintaining stable appearance features such as color and texture even with significant viewpoint changes. Furthermore, the changes in pose parameters between adjacent frames in each sequence are kept relatively smooth (e.g., the azimuth step size is fixed at 5 degrees), serving as a temporal consistency supervision signal.

[0069] By acquiring sample pose parameter sequences, sample description text sequences, and reference sample image sequences corresponding to multiple continuously changing camera viewpoints in the viewpoint trajectory, the pose parameter differences between adjacent viewpoints are encoded to obtain inter-frame relative change codes. The viewpoint tokens of each viewpoint, the inter-frame relative change codes, and the sample description text are fused and input into the image generation model to generate predicted image sequences. A temporal consistency loss is introduced, which includes two constraints: the first constraint constrains the viewpoint token differences between adjacent frames to be positively correlated with the image feature differences of the corresponding frames; the second constraint constrains the inter-frame change rate of the scene context tokens to match the camera motion rate. This enables the trained model to generate video sequences with smooth viewpoint transitions and coherent background perspective relationships, fundamentally solving the technical problems of viewpoint jumps and background abrupt changes in existing methods during video generation.

[0070] Furthermore, after training the image generation model using S110-S170, the trained image generation model can be used to perform the following steps to generate continuous images (videos): obtaining descriptive text and camera viewpoint pose parameters, wherein the camera viewpoint pose parameters include a sequence of pose parameters corresponding to multiple continuously changing camera viewpoints; fusing viewpoint tokens, scene context tokens, and descriptive text, including: encoding the pose parameter differences between adjacent viewpoints to obtain inter-frame relative change codes; fusing the viewpoint token sequences and inter-frame relative change codes of each viewpoint in the viewpoint trajectory with the descriptive text to obtain video prompt information; inputting the video prompt information into the trained image generation model to generate a target image sequence, the target image sequence including multiple images, each image corresponding to a camera viewpoint.

[0071] It is worth mentioning that when the descriptive text includes multiple texts corresponding one-to-one with each viewpoint, the tokens of each text and its corresponding viewpoint are fused separately and then organized in frame order; when the descriptive text is a single global text, the global text is sequentially fused with the tokens of all viewpoints. Specifically, the viewpoint tokens of each viewpoint, the inter-frame relative change codes, and the embedding vectors of the descriptive text are organized into a sequence in frame order and used as the conditional input of the video diffusion model.

[0072] By acquiring pose parameter sequences and descriptive text containing multiple continuously changing camera viewpoints, the pose parameter differences between adjacent viewpoints are encoded to obtain inter-frame relative change codes. The viewpoint token sequences of each viewpoint, the inter-frame relative change codes, and the descriptive text are fused and input into the trained image generation model to generate target image sequences. When the descriptive text is a single global text, it is fused with the tokens of all viewpoints. When it is multiple texts that correspond one-to-one with each viewpoint, they are fused separately in frame order. Finally, a coherent video with each frame matching the corresponding camera viewpoint is generated, achieving efficient end-to-end video generation.

[0073] Please see Figure 3 This application embodiment also provides a training device 300 for an image generation model, comprising: The data acquisition module 310 is used to acquire training samples, which include sample pose parameters of the camera viewpoint, sample description text, and reference sample images corresponding to the pose parameters and description text. Encoding module 320 is used to encode the sample pose parameters to obtain the sample viewpoint embedding vector; The token generation module 330 is used to determine the target number of viewpoint tokens corresponding to the sample pose parameters, and generate the target number of sample viewpoint tokens based on the sample viewpoint embedding vector. The token generation module 330 is further configured to construct a sample scene context token based on the sample pose parameters, wherein the sample scene context token is used to encode the perspective geometric relationships of the background scene. The fusion module 340 is used to fuse the sample viewpoint token and the sample scene context token with the sample description text to obtain sample prompt information; Image generation module 350 is used to input the sample prompt information into the image generation model to generate a predicted image; The model training module 360 ​​is used to obtain the model loss based on the predicted image and the reference sample image, and adjust the model parameters of the image generation model based on the model loss to obtain the trained image generation model.

[0074] Optionally, the model training module 360 ​​is further configured to calculate a diffusion loss based on the difference between the predicted image and the target image; calculate a token orthogonal constraint loss based on the information redundancy between the viewpoint tokens of the target number; calculate a foreground-background geometric consistency loss based on the geometric consistency between the foreground and background regions of the predicted image; and obtain a model loss based on the diffusion loss, the token orthogonal constraint loss, and the foreground-background geometric consistency loss.

[0075] Optionally, the sample pose parameters include at least two of the following parameters: azimuth angle, elevation angle, distance, pitch angle, and yaw angle; the encoding module 320 is further configured to use a multilayer perceptron network to map the sample pose parameters into an embedding vector of the same dimension as the text words in the sample description text, thereby obtaining a sample viewpoint embedding vector.

[0076] Optionally, the token generation module 330 is further configured to input the sample pose parameters into the multilayer perceptron network, and obtain the probability distributions corresponding to multiple token quantity categories through the multilayer perceptron prediction; and determine the target number of sample viewpoint tokens corresponding to the sample pose parameters based on the probability distributions corresponding to multiple token quantity categories.

[0077] Optionally, the token generation module 330 is further configured to: when the number of targets is one, use the sample viewpoint embedding vector as a unique sample viewpoint token; when the number of targets is multiple, generate multiple sample viewpoint tokens based on the sample viewpoint embedding vector using a hierarchical generation strategy, wherein generating multiple sample viewpoint tokens using a hierarchical generation strategy includes: generating a first sample viewpoint token based on the sample viewpoint embedding vector using an attention mechanism, the first sample viewpoint token being used to encode macroscopic spatial location information; generating a second sample viewpoint token based on the first sample viewpoint token using a cross-attention mechanism, the second sample viewpoint token being used to encode incremental geometric information relative to the first sample viewpoint token; and generating a third sample viewpoint token based on the first sample viewpoint token and the second sample viewpoint token using a cross-attention mechanism, the third sample viewpoint token being used to encode residual geometric information relative to the first sample viewpoint token and the second sample viewpoint token.

[0078] Optionally, the token generation module 330 is further configured to use the first sample viewpoint token as a key-value pair, with the sample viewpoint embedding vector as a query, to generate the second sample viewpoint token through a cross-attention mechanism; and to use the first sample viewpoint token and the second sample viewpoint token together as a key-value pair, with the sample viewpoint embedding vector as a query, to generate the third sample viewpoint token through a cross-attention mechanism.

[0079] Optionally, the training samples include a sequence of sample pose parameters corresponding to multiple continuously changing camera viewpoints in the viewpoint trajectory, sample description text, and a sequence of reference sample images corresponding to multiple continuously changing camera viewpoints and sample description text. The fusion module 340 is also used to encode the differences in sample pose parameters between adjacent viewpoints to obtain inter-frame relative change codes; and to fuse the sample viewpoint tokens of each camera viewpoint in the viewpoint trajectory, the inter-frame relative change codes, and the sample description text to obtain video sample prompt information. The image generation module 350 is further configured to input the video sample prompt information into the image generation model to generate a predicted image sequence, wherein the predicted image sequence includes a predicted image corresponding to each camera viewpoint.

[0080] Optionally, the data acquisition module 310 is also used to acquire the pose parameters and descriptive text of the camera viewpoint; The encoding module 320 is also used to encode the pose parameters to obtain a viewpoint embedding vector; The token generation module 330 is also used to determine the number of viewpoint tokens based on the pose parameters, and generate a corresponding number of viewpoint tokens based on the viewpoint embedding vector; The token generation module 330 is also used to construct a scene context token based on the pose parameters, wherein the scene context token is used to encode the perspective geometry of the background scene; The fusion module 340 is also used to fuse the viewpoint token, the scene context token, and the description text to obtain a prompt message; The image generation module 350 is also used to input the prompt information into the trained image generation model to generate the target image.

[0081] Optionally, the pose parameters of the camera viewpoint include a sequence of pose parameters corresponding to multiple continuously changing camera viewpoints; the fusion module 340 is further used to encode the pose parameter differences between adjacent viewpoints to obtain inter-frame relative change codes; and to fuse the viewpoint token sequence of each viewpoint in the viewpoint trajectory, the inter-frame relative change codes, and the description text to obtain video prompt information. The image generation module 350 is also used to input the video prompt information into the trained image generation model to generate a target image sequence, wherein the target image sequence includes multiple images, and each image corresponds to a camera viewpoint.

[0082] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0083] Figure 4 This is a block diagram illustrating an electronic device 100 according to an exemplary embodiment. Figure 4 As shown, the electronic device 100 may include a processor 101 and a memory 102. The electronic device 100 may also include one or more of a multimedia component 103, an input / output (I / O) interface 104, and a communication component 105. Specifically, the electronic device may be a server or terminal, or other device with data processing capabilities, and may be able to run training methods for image generation models.

[0084] The processor 101 controls the overall operation of the electronic device 100 to complete all or part of the steps in the training method for the image generation model described above. The memory 102 stores various types of data to support the operation of the electronic device 100. This data may include, for example, instructions for any application or method operating on the electronic device 100, and application-related data. The memory 102 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 103 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 102 or transmitted via communication component 105. The audio component also includes at least one speaker for outputting audio signals. I / O interface 104 provides an interface between processor 101 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 105 is used for wired or wireless communication between the electronic device 100 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G, or a combination thereof, is not limited here. Therefore, the corresponding communication component 105 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0085] In an exemplary embodiment, the electronic device 100 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the training method of the image generation model described above.

[0086] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the image generation model training method described above. For example, the computer-readable storage medium may be the memory 102 including the program instructions described above, which may be executed by the processor 101 of the electronic device 100 to complete the image generation model training method described above.

[0087] In another exemplary embodiment, a computer program product is also provided, comprising a computer program executable by a programmable device, the computer program having a code portion for performing the training method of the image generation model described above when executed by the programmable device.

[0088] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0089] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0090] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

Claims

1. A training method for an image generation model, characterized in that, The method includes: Acquire training samples, which include sample description text, sample pose parameters of the camera viewpoint, and reference sample images corresponding to the sample pose parameters and sample description text; The sample pose parameters are encoded to obtain the sample viewpoint embedding vector; Determine the target number of viewpoint tokens corresponding to the sample pose parameters, and generate the target number of sample viewpoint tokens based on the sample viewpoint embedding vector; A sample scene context token is constructed based on the sample pose parameters. The sample scene context token is used to encode the perspective geometry of the background scene. The sample viewpoint token and the sample scene context token are fused with the sample description text to obtain sample prompt information; The sample prompt information is input into the image generation model to generate a predicted image; The model loss is obtained based on the predicted image and the reference sample image. The model parameters of the image generation model are adjusted based on the model loss to obtain the trained image generation model.

2. The method according to claim 1, characterized in that, The process of obtaining the model loss based on the predicted image and the reference sample image includes: Calculate the diffusion loss based on the predicted image and the reference sample image; Calculate the token orthogonal constraint loss based on the information redundancy among the target number of viewpoint tokens; Calculate the foreground-background geometric consistency loss based on the geometric consistency between the foreground and background regions of the predicted image; The model loss is obtained based on the diffusion loss, the token orthogonality constraint loss, and the foreground-background geometric consistency loss.

3. The method according to claim 1, characterized in that, The sample pose parameters include at least two of the following parameters: azimuth, elevation, range, pitch, and yaw. The process of encoding the sample pose parameters to obtain the sample viewpoint embedding vector includes: The sample pose parameters are mapped to embedding vectors of the same dimension as the text words in the sample description text using a multilayer perceptron network, thus obtaining the sample viewpoint embedding vector.

4. The method according to claim 1, characterized in that, Determining the target number of sample viewpoint tokens corresponding to the pose parameters includes: The sample pose parameters are input into a multilayer perceptron network, and the probability distributions corresponding to various token quantity categories are predicted by the multilayer perceptron. The target number of sample viewpoint tokens corresponding to the sample pose parameters is determined based on the probability distribution corresponding to various token quantity categories.

5. The method according to claim 1, characterized in that, The step of generating the target number of sample viewpoint tokens based on the sample viewpoint embedding vector includes: When the number of targets is one, the sample viewpoint embedding vector is used as the unique sample viewpoint token; When there are multiple targets, multiple sample viewpoint tokens are generated based on the sample viewpoint embedding vector using a hierarchical generation strategy. The generation of multiple sample viewpoint tokens using the hierarchical generation strategy includes: Based on the sample viewpoint embedding vector, an attention mechanism is used to generate a first sample viewpoint token, which is used to encode macroscopic spatial location information. Based on the first sample viewpoint token, a second sample viewpoint token is generated through a cross-attention mechanism. The second sample viewpoint token is used to encode incremental geometric information relative to the first sample viewpoint token. Based on the first sample viewpoint token and the second sample viewpoint token, a third sample viewpoint token is generated through a cross-attention mechanism. The third sample viewpoint token is used to encode residual geometric information relative to the first sample viewpoint token and the second sample viewpoint token.

6. The method according to claim 5, characterized in that, The step of generating a second sample viewpoint token based on the first sample viewpoint token through a cross-attention mechanism includes: The first sample viewpoint token is used as a key-value pair, and the sample viewpoint embedding vector is used as a query to generate the second sample viewpoint token through a cross-attention mechanism; The step of generating a third sample viewpoint token based on the first and second sample viewpoint tokens through a cross-attention mechanism includes: The first sample viewpoint token and the second sample viewpoint token are used together as a key-value pair, and the sample viewpoint embedding vector is used as the query to generate the third sample viewpoint token through a cross-attention mechanism.

7. The method according to claim 1, characterized in that, The training samples include a sequence of sample pose parameters corresponding to multiple continuously changing camera viewpoints in the viewpoint trajectory, sample description text, and a sequence of reference sample images corresponding to multiple continuously changing camera viewpoints and sample description texts. The step of fusing the sample viewpoint token and the sample scene context token with the sample description text to obtain sample prompt information includes: The differences in pose parameters between adjacent viewpoints are encoded to obtain the inter-frame relative change encoding. The sample viewpoint tokens of each camera viewpoint in the viewpoint trajectory, the inter-frame relative change code, and the sample description text are fused to obtain video sample prompt information; The step of inputting the sample prompt information into the image generation model to generate a predicted image includes: The video sample prompt information is input into the image generation model to generate a predicted image sequence, which includes a predicted image corresponding to each camera viewpoint.

8. The method according to claim 1, characterized in that, The method further includes: Obtain the pose parameters and descriptive text of the camera viewpoint; The pose parameters are encoded to obtain the viewpoint embedding vector; Based on the pose parameters, determine the number of viewpoint tokens, and generate a corresponding number of viewpoint tokens based on the viewpoint embedding vector; A scene context token is constructed based on the pose parameters, and the scene context token is used to encode the perspective geometry of the background scene; The viewpoint token, the scene context token, and the description text are merged to obtain the prompt information; The prompt information is input into the trained image generation model to generate the target image.

9. The method according to claim 8, characterized in that, The pose parameters of the camera viewpoint include a sequence of pose parameters corresponding to multiple continuously changing camera viewpoints in the viewpoint trajectory; By fusing the viewpoint token, the scene context token, and the descriptive text, the resulting prompt information includes: The differences in pose parameters between adjacent viewpoints are encoded to obtain the inter-frame relative change encoding. The viewpoint token sequence of each viewpoint in the viewpoint trajectory, the inter-frame relative change encoding, and the description text are fused to obtain video prompt information; The step of inputting the prompt information into the trained image generation model to generate the target image includes: The video prompt information is input into the trained image generation model to generate a target image sequence, which includes multiple images, each corresponding to a camera viewpoint.

10. A training device for an image generation model, characterized in that, The device includes: The data acquisition module is used to acquire training samples, which include sample description text, sample pose parameters of the camera viewpoint, and reference sample images corresponding to the sample pose parameters and sample description text. The encoding module is used to encode the sample pose parameters to obtain the sample viewpoint embedding vector; The token generation module is used to determine the target number of viewpoint tokens corresponding to the sample pose parameters, and generate the target number of sample viewpoint tokens based on the sample viewpoint embedding vector. The token generation module is further configured to construct a sample scene context token based on the sample pose parameters, wherein the sample scene context token is used to encode the perspective geometric relationships of the background scene. The fusion module is used to fuse the sample viewpoint token and the sample scene context token with the sample description text to obtain sample prompt information; The image generation module is used to input the sample prompt information into the image generation model to generate a predicted image; The model training module is used to obtain the model loss based on the predicted image and the reference sample image, and to adjust the model parameters of the image generation model based on the model loss to obtain the trained image generation model.

11. An electronic device, characterized in that, include: processor; as well as A memory storing a computer program or readable instructions, which, when executed by the processor, implement the method as described in any one of claims 1-9.

12. A computer-readable storage medium having a computer program or readable instructions stored thereon, characterized in that, When the computer program or readable instructions are executed, they implement the steps of the method according to any one of claims 1-9.