Image generation method and device, storage medium, electronic equipment and chip

By generating sequentially predicted image tokens through diagonal generation, the four-dimensional rotation position coding and direction embedding in the autoregressive Transformer model solves the accuracy problem caused by the distance between tokens in image generation, and improves the accuracy and quality of image generation.

CN120259462APending Publication Date: 2025-07-04BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308845.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the distance between image tokens is relatively long due to the order of row expansion during image generation, which affects the accuracy and quality of image generation.

Method used

The diagonal line is used to generate sequential prediction image tokens, which are used to generate from the upper left corner of the image, and alternately along the diagonal direction of the preset angle from the lower left corner to the upper right corner and from the upper right corner to the lower left corner. The four-dimensional rotation position coding and direction embedding in the autoregressive Transformer model are used to improve prediction accuracy.

Benefits of technology

It effectively shortens the distance between image tokens and improves the accuracy and quality of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259462A_ABST
    Figure CN120259462A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device, a storage medium, electronic equipment and a chip. The method comprises the following steps: firstly, acquiring a generation condition of an image and determining a category token corresponding to the generation condition; the image tokens are sequentially predicted according to the diagonal generation sequence based on the tokens of the category, and the diagonal generation sequence is that generation from the lower left corner to the upper right corner and generation from the upper right corner to the lower left corner are alternately carried out in the direction of the diagonal of the preset angle from the upper left corner of the image; and generating a complete image according to the predicted image tokens. The image tokens are sequentially predicted according to the diagonal generation sequence, so that the distance between the image tokens can be effectively shortened, the prediction accuracy can be improved when the next image token is predicted based on the previous image token, a complete image can be generated according to the accurately predicted image tokens, the image generation accuracy can be improved, and the image generation efficiency can be improved. And the image generation quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to an image generation method, apparatus, storage medium, electronic device, and chip. Background Art

[0002] Image generation refers to the process of automatically creating images through computer algorithms. This technology is widely used in multiple fields, such as art creation, design, entertainment, medical image analysis, etc. In recent years, with the development of artificial intelligence, especially deep learning technology, image generation has reached an unprecedented level.

[0003] Currently, the image to be generated can be regarded as multiple rows, and each image token in the current row is predicted in sequence from left to right, and then the next row is continued to predict each image token in sequence from left to right. After all the image tokens in all rows are predicted, a complete image is generated based on these image tokens.

[0004] However, each time a new row is started, since the distance between the last image token in the current row and the first image token in the next row is relatively far, the accuracy of predicting the first image token in the next row based on the last image token in the current row is poor, which in turn affects the accuracy of image generation and thus the quality of image generation. Summary of the Invention

[0005] The present disclosure provides an image generation method, apparatus, storage medium, electronic device, and chip, mainly aiming to improve the technical problem that the current method of generating image tokens in the order of row by row affects the accuracy of image generation and thus the quality of image generation.

[0006] According to the first aspect of the embodiments of the present disclosure, an image generation method is provided, including:

[0007] Obtain the generation conditions of the image;

[0008] Determine the category token corresponding to the generation conditions;

[0009] Predict each image token in sequence according to the diagonal generation order based on the category token, where the diagonal generation order starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the direction of the preset angle diagonal;

[0010] Generate an image based on the respective image tokens.

[0011] Optionally, the predicting each image token in sequence according to the diagonal generation order based on the category token includes:

[0012] Input the class token into an autoregressive Transformer model to sequentially predict each image token according to the diagonal generation order. Among them, in the autoregressive Transformer model, four-dimensional rotational position encoding is used to inject the information of the current position and the predicted position into the attention matrix simultaneously when predicting each image token, so as to realize the prediction of image tokens along the direction of the predicted position.

[0013] Optionally, the adaptive normalization layer in the autoregressive Transformer model is conditioned on the direction embedding, and the direction embedding is used to inject the direction information corresponding to the diagonal generation order into the adaptive normalization layer.

[0014] Optionally, the training process of the autoregressive Transformer model includes:

[0015] Encode the sample image into a discrete sequence of image tokens;

[0016] Rearrange the image token sequence according to the diagonal generation order;

[0017] Concatenate the sample class token corresponding to the sample image at the forefront of the rearranged image token sequence, and train the autoregressive Transformer model on the concatenated image token sequence in the way of predicting the next token task.

[0018] Optionally, the step of inputting the class token into the autoregressive Transformer model to sequentially predict each image token according to the diagonal generation order includes:

[0019] Input the class token into the trained autoregressive Transformer model, and sequentially predict each image token from the class token in the way of predicting the next token task according to the diagonal generation order.

[0020] Optionally, the step of encoding the sample image into a discrete sequence of image tokens includes:

[0021] Use the codebook in the image tokenizer to encode the sample image into a discrete sequence of image tokens.

[0022] Optionally, the step of generating an image based on the respective image tokens includes:

[0023] Restore the respective image tokens to the arrangement order of being unfolded row by row;

[0024] Convert the respective image tokens in the arrangement order into image pixels to obtain the generated image.

[0025] Optionally, converting each image token in the arrangement order into image pixels to obtain a generated image includes:

[0026] Using a decoder in the image tokenizer to convert each image token in the arrangement order into image pixels to obtain a generated image.

[0027] According to a second aspect of the embodiments of the present disclosure, there is provided an image generation device, including:

[0028] An acquisition module configured to acquire generation conditions of an image;

[0029] A determination module configured to determine class tokens corresponding to the generation conditions;

[0030] A prediction module configured to sequentially predict each image token based on the class tokens in a diagonal generation order, where the diagonal generation order starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the diagonal direction of a preset angle;

[0031] A generation module configured to generate an image based on the respective image tokens.

[0032] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the image generation method described in the first aspect is implemented.

[0033] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the image generation method described in the first aspect is implemented.

[0034] By means of the above technical solutions, the present disclosure provides an image generation method, device, storage medium, electronic device, and chip. Compared with the current related technologies, the present disclosure does not use the method of generating image tokens in a row-by-row expansion order. Specifically, first, class tokens corresponding to the generation conditions of the image are determined; then, based on these class tokens, each image token is sequentially predicted in a diagonal generation order, where the diagonal generation order starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the diagonal direction of a preset angle, so as to effectively shorten the distance between image tokens and improve the prediction accuracy when predicting the next image token based on the previous image token; furthermore, a complete image can be generated based on these accurately predicted image tokens. By applying the technical solutions of the present disclosure, the accuracy of image generation can be improved, and thus the quality of image generation can be improved.

[0035] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0037] Figure 1 A schematic diagram showing an example provided by an embodiment of the present disclosure;

[0038] Figure 2 A schematic flowchart showing a method for generating an image provided by an embodiment of the present disclosure;

[0039] Figure 3 A schematic diagram showing another example provided by an embodiment of the present disclosure;

[0040] Figure 4 A schematic flowchart showing another method for generating an image provided by an embodiment of the present disclosure;

[0041] Figure 5 A schematic structural diagram showing yet another example provided by an embodiment of the present disclosure;

[0042] Figure 6 A schematic structural diagram showing an image generation device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] Some embodiments of the present disclosure will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the present disclosure, except for operations that must be performed in a specific order. Additionally, descriptions of features known in the art may be omitted for increased clarity and conciseness. It should be noted that, without conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0044] The embodiments described in some embodiments of the present disclosure below do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0045] For image generation technology, as an example, the image to be generated can be regarded as multiple rows, such asFigure 1 As shown, using the autoregressive prediction algorithm, each image token in the current row is predicted in sequence from left to right. For example, starting from token 1, tokens 2, 3, 4, and 5 are predicted in sequence, then moving to the next row and predicting each image token in the next row in sequence from left to right, such as tokens 6, 7, 8, 9, and 10. After predicting all the image tokens in all rows, tokens 1 to 25 are obtained, and a complete image is generated based on these image tokens.

[0046] However, each time a new row is started, since the distance between the last image token in the current row and the first image token in the next row is relatively large, the accuracy of predicting the first image token in the next row based on the last image token in the current row is poor. For example, when predicting token 2, it can be based on token 1 and the relative position between token 2 and token 1; when predicting token 3, it can be based on tokens 1, 2 and the relative position between token 3 and token 2; and so on. When predicting token 6, it can be based on tokens 1 to 5 and the relative position between token 5 and token 6. However, the distance between token 5 and token 6 is relatively large, which will affect the prediction accuracy of token 6. If token 6 is predicted inaccurately, it will then affect the prediction of subsequent image tokens, and when predicting tokens 11, 16, and 21, the accuracy of the prediction results will be further reduced, thus affecting the quality of image generation.

[0047] To improve the above technical problem that the method of generating image tokens in the order of row-by-row expansion (as shown in the example) affects the accuracy of image generation and thus affects the quality of image generation, the embodiments of the present disclosure provide an image generation method. As shown, this method can be applied to an image generation device or equipment for execution, and the method includes the following steps. Figure 1 To improve the above technical problem that the method of generating image tokens in the order of row-by-row expansion (as shown in the example) affects the accuracy of image generation and thus affects the quality of image generation, the embodiments of the present disclosure provide an image generation method. As shown, this method can be applied to an image generation device or equipment for execution, and the method includes the following steps. Figure 2 As shown, this method can be applied to an image generation device or equipment for execution, and the method includes the following steps.

[0048] Step 101: Obtain the generation conditions of the image.

[0049] The generation conditions refer to the parameters or information used to guide or constrain image generation. These conditions can be specific numerical values, text descriptions, category labels, etc., in order to generate an image that meets specific requirements.

[0050] For example, the generation conditions may include: category labels, or text descriptions, or image attributes, etc. Among them, the category label can indicate the specific category to which the generated image belongs, such as categories like cats, dogs, etc.; the text description can be used to describe the content of the image to be generated; the image attributes can contain information about specific attributes of the image, such as color, shape, etc.

[0051] The generation conditions of the image can be input by the user or the system or input based on condition triggering, etc.

[0052] Step 102: Determine the class token corresponding to the generation condition of the image.

[0053] The class token can be an identifier used to represent a specific class, which can be embedded into the neural network in vector form and can be used as part of the model input to help the model distinguish features of different classes.

[0054] In some embodiments, according to the generation condition of the image, its corresponding class token can be obtained, such as mapping class information such as class labels or text descriptions and encoding to obtain the class token.

[0055] Step 103: Predict each image token in sequence according to the diagonal generation order based on the class token.

[0056] Among them, the diagonal generation order (or can be called the diagonal scanning order, etc.) starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the direction of the preset angle diagonal. The preset angle can be 45° angle, etc. For example, there are four generation directions in the diagonal generation order: to the right, down, to the upper right, and to the lower left. As Figure 3 shown, based on the class token, image token 1 is predicted, then based on the class token and image token 1, image token 2 is predicted downward, then based on the class token and image tokens 1 and 2, image token 3 is predicted to the upper right, then based on the class token and image tokens 1 to 3, image token 4 is predicted to the right, then based on the class token and image tokens 1 to 4, image token 5 is predicted to the lower left, and then based on the class token and image tokens 1 to 5, image token 6 is predicted to the lower left. And so on, alternating between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the 45° angle diagonal direction to obtain image tokens 1 to 25.

[0057] This method, compared with the method of generating image tokens in the order of row expansion as Figure 1 shown, does not have the situation where the actual distance between adjacent index image tokens at the end of each row is very far, can effectively shorten the distance between image tokens, and can improve the prediction accuracy when predicting the next image token based on the previous image token; furthermore, a complete image can be generated based on these accurately predicted image tokens.

[0058] Step 104: Generate an image based on each predicted image token.

[0059] Compared with the current related technologies, the embodiments of the present disclosure do not use the method of generating image tokens in a row-by-row expansion order. Specifically, first, the category token corresponding to the generation condition of the image is determined; then, based on this category token, each image token is predicted in turn according to the diagonal generation order, where the diagonal generation order starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the direction of the diagonal line at a preset angle. In this way, the distance between image tokens can be effectively shortened, and the prediction accuracy can be improved when predicting the next image token based on the previous image token; furthermore, a complete image can be generated based on these accurately predicted image tokens. By applying the technical solutions of the embodiments of the present disclosure, the accuracy of image generation can be improved, and thus the quality of image generation can be improved.

[0060] The image generation scheme provided by the embodiments of the present disclosure aims to improve the quality of image generation and can be applied in many aspects. For example, in terms of artistic creation: designers can use the image generation scheme provided by the embodiments of the present disclosure to create unique artworks and design patterns; in terms of film and television production: the image generation scheme provided by the embodiments of the present disclosure can be used for special effects production, scene reconstruction, virtual character generation, etc.; in terms of e-commerce: the image generation scheme provided by the embodiments of the present disclosure can be used to generate product display pictures to improve the user experience.

[0061] To further illustrate the specific implementation process of the method as Figure 2 shown, as an optional way, the embodiments of the present disclosure provide the specific method as Figure 4 shown, and this method includes:

[0062] Step 201, obtain the generation condition of the image and determine the category token corresponding to the generation condition of the image.

[0063] Step 202, input the category token into the autoregressive Transformer model to predict each image token in turn according to the diagonal generation order.

[0064] The autoregressive Transformer model used in the embodiments of the present disclosure can be used to predict each image token in turn according to the diagonal generation order based on the category token. The diagonal generation order can start from the upper left corner of the image and alternate between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the direction of the diagonal line at a preset angle. The preset angle can be 45° and so on.

[0065] In some embodiments, in order to better implement predicting each image token in turn according to the diagonal generation order, a four-dimensional rotation position encoding is adopted in the autoregressive Transformer model to inject the information of the current position and the predicted position into the attention matrix at each prediction of the image token, so as to realize predicting the image token along the direction of the predicted position. For example, asFigure 3 In the example shown, when predicting image token 2 based on category token and image token 1, four-dimensional rotation position encoding is used to simultaneously inject the position information of image token 1 and image token 2 into the attention matrix to achieve the prediction of image token 2 along the downward direction.

[0066] In some examples, the disclosed embodiments define four-dimensional coordinates for each image token index as follows:

[0067]

[0068] The four dimensions are the row and column of the image token with index n, and the row and column of the image token with index n+1. Then, the rotation matrix is ​​defined as follows:

[0069]

[0070] Among them, θ t represents the tth frequency, d head Represents the dimension of each attention head in the autoregressive Transformer model. The query and key in the autoregressive Transformer model are converted to complex form and then multiplied element-by-element with the rotation matrix to inject the position information of the current position and the predicted position into the attention matrix, i.e., four-dimensional rotation position encoding.

[0071] In some embodiments, the adaptive normalization layer in the autoregressive Transformer model is conditioned on a directional embedding, which is used to inject directional information corresponding to the diagonal generation order into the adaptive normalization layer. There are four generation directions in the diagonal generation order: right, downward, upper right, and lower left. The disclosed embodiment uses four learnable directional embeddings to represent the four generation directions, each index selects the corresponding directional embedding, and then the category embeddings are added to calculate the scale and shift parameters in the adaptive normalization layer.

[0072] The embodiment of the present disclosure improves the traditional autoregressive Transformer model to obtain the autoregressive Transformer model used in the embodiment of the present disclosure, such as Figure 5As shown, the Transformer model (Transformer Block) is the core part, which receives input tokens and performs a series of transformations, including the multi-head attention mechanism (Multi-head Attention) and the feed-forward neural network (Feed-Forward Network). Among them, the four-dimensional rotational position encoding (4D-RoPE) is used in combination with the multi-head attention mechanism (Multi-head Attention with 4D-RoPE), which can inject the information of the current position and the predicted position into the attention matrix simultaneously when predicting each image token, so as to realize the prediction of image tokens along the direction of the predicted position. After the multi-head attention mechanism, normalization techniques such as scaling (Scale), shifting (Shift), and root mean square normalization (RMSNorm) are applied. These operations help to stabilize the training process and improve the convergence speed and performance of the model. The output after normalization is fed into a feed-forward neural network to further extract features. The feed-forward network usually contains two fully connected layers with an activation function (such as ReLU) in the middle. Among them, the multi-layer perceptron (MLP) can be used for classification or regression tasks. For example, it can combine the class embedding and the direction embedding to generate the prediction result. The direction embedding can be used to inject the direction information corresponding to the diagonal generation order into the adaptive normalization layer.

[0073] In some embodiments, the training process of the autoregressive Transformer model may specifically include: first, collecting different sample images, and then encoding the sample images into a discrete sequence of image tokens; then rearranging the sequence of image tokens in the diagonal generation order, as Figure 5 shown; finally, splicing the sample class token corresponding to the sample image at the front of the rearranged sequence of image tokens, and training the autoregressive Transformer model on the spliced sequence of image tokens in the way of predicting the next token task. In this way, the autoregressive Transformer model applied in the embodiments of the present disclosure can be accurately trained.

[0074] In some examples, the above-mentioned encoding of the sample image into a sequence of discrete image tokens may specifically include: using the codebook in the image tokenizer to encode the sample image into a sequence of discrete image tokens. Compared with the need to re-learn the image representation during the training of the autoregressive model, the representation ability of the image tokenizer is not fully utilized. Embodiments of the present disclosure may use the codebook in the image tokenizer to replace the image embedding that needs to be re-learned and freeze these parameters. For example, when training an autoregressive Transformer model in a traditional way, it is necessary to re-learn the image representation, which affects the efficiency of model training and increases the difficulty of model training. However, embodiments of the present disclosure can use the codebook in the pre-trained image tokenizer to encode the sample image into a sequence of discrete image tokens without re-learning the image representation, effectively improving the efficiency of model training and reducing the difficulty of model training.

[0075] Embodiments of the present disclosure propose an autoregressive image generation model that can be used for class-conditional generation tasks. Embodiments of the present disclosure use a pre-trained image tokenizer to encode the image into a sequence of discrete tokens, rearrange the sequence of image tokens in the order of diagonal scanning, and then splice a class token at the forefront. Then, embodiments of the present disclosure use the task of predicting the next token to train an autoregressive Transformer model on the rearranged sequence of image tokens. The training task used in embodiments of the present disclosure is the task of predicting the next token, and the same method of predicting the next token is also used for autoregressive image generation during inference (or can be referred to as application or prediction, etc.). Embodiments of the present disclosure improve this training task by changing the arrangement order of unfolding by rows to the order of diagonal scanning, thereby training an autoregressive Transformer model that can make predictions according to the diagonal generation order.

[0076] Based on the trained autoregressive Transformer model, in some embodiments, step 202 may specifically include: inputting the class token into the trained autoregressive Transformer model, and starting from this class token in the way of using the task of predicting the next token, sequentially predicting each image token according to the diagonal generation order.

[0077] Step 203, restore each image token to the arrangement order of unfolding row by row.

[0078] For example, after predicting each image token in the manner as Figure 3 shown, restore these image tokens to the arrangement order of unfolding row by row, that is, in the form as Figure 1 shown.

[0079] Step 204, convert each image token in the arrangement order of unfolding row by row into image pixels to obtain the generated image.

[0080] In some embodiments, step 204 may specifically include: using a decoder in the image tokenizer to convert each image token arranged in the row-by-row unfolding order into image pixels to obtain the generated image.

[0081] The embodiments of the present disclosure can start from a separate class token in the way of predicting the next token, gradually generate a complete sequence of image tokens, then restore the image tokens arranged in the diagonal scan order to the row-by-row unfolding order, and finally use the decoder in the image tokenizer to convert the discrete sequence of image tokens into image pixels to obtain the generated image.

[0082] The image generation scheme provided by the embodiments of the present disclosure can be applied to image generation in computer vision. It should be noted that the scheme provided by the embodiments of the present disclosure can be extended to multi-modal foundation models, such as multi-modal large models that can generate text and images simultaneously. By applying the technical solution of the embodiments of the present disclosure, an image token sequence is generated in the diagonal scan order, which solves the problem that the distance between adjacent tokens is too far and improves the quality of the generated image. And the four-dimensional rotation position encoding injects the information of the current position and the predicted position into the attention matrix at the same time, and the direction embedding injects the direction information into the adaptive normalization layer. These two methods explicitly input the generation direction into the model, endowing the model with the ability to perceive the generation direction, reducing the difficulty of model training, and improving the quality of the generated image. In addition, using the codebook in the image tokenizer as the image embedding reduces the difficulty of model training.

[0083] Figure 6 is a block diagram of an image generation device shown according to some embodiments of the present disclosure. Referring to Figure 6 This device includes: an acquisition module 31, a determination module 32, a prediction module 33, and a generation module 34.

[0084] The acquisition module 31 is configured to acquire the generation conditions of an image;

[0085] The determination module 32 is configured to determine the class token corresponding to the generation conditions;

[0086] The prediction module 33 is configured to sequentially predict each image token based on the class token in the diagonal generation order, where the diagonal generation order starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the direction of the preset angle diagonal;

[0087] The generation module 34 is configured to generate an image based on each image token.

[0088] In some embodiments, the prediction module 33 is specifically configured to input the category token into an autoregressive Transformer model to sequentially predict each image token according to the diagonal generation order. In the autoregressive Transformer model, four-dimensional rotational position encoding is used to inject the information of the current position and the predicted position into the attention matrix simultaneously when predicting each image token, so as to predict the image token along the direction of the predicted position.

[0089] In some embodiments, the adaptive normalization layer in the autoregressive Transformer model is conditioned on the direction embedding, and the direction embedding is used to inject the direction information corresponding to the diagonal generation order into the adaptive normalization layer.

[0090] In some embodiments, the apparatus further includes a training module;

[0091] The training module is configured to encode the sample image into a discrete sequence of image tokens; rearrange the sequence of image tokens according to the diagonal generation order; concatenate the sample category token corresponding to the sample image at the forefront of the rearranged sequence of image tokens, and train the autoregressive Transformer model on the concatenated sequence of image tokens in the way of predicting the next token task.

[0092] In some embodiments, the prediction module 33 is further specifically configured to input the category token into the trained autoregressive Transformer model, and sequentially predict each image token starting from the category token in the way of predicting the next token task according to the diagonal generation order.

[0093] In some embodiments, the training module is specifically configured to use the codebook in the image tokenizer to encode the sample image into a discrete sequence of image tokens.

[0094] In some embodiments, the generation module 34 is specifically configured to restore the respective image tokens into a row-by-row expanded arrangement order; convert the respective image tokens in the arrangement order into image pixels to obtain the generated image.

[0095] In some embodiments, the generation module 34 is further specifically configured to use the decoder in the image tokenizer to convert the respective image tokens in the arrangement order into image pixels to obtain the generated image.

[0096] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0097] It should be noted that for other corresponding descriptions of each functional unit involved in the image generation device provided in the embodiments of the present disclosure, reference can be made to Figures 1 to 5 the corresponding description in

[0098] Based on the method as described above in Figures 1 to 5 shown, correspondingly, the embodiments of the present disclosure further provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as described above in Figures 1 to 5 shown is implemented.

[0099] Based on such an understanding, the technical solution of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of the present disclosure.

[0100] Based on the method as described above in Figures 1 to 5 shown, and Figure 6 the virtual device embodiment shown, in order to achieve the above object, the embodiments of the present disclosure further provide an electronic device, such as a high-definition TV, a computer, a projector, etc. with a liquid crystal display, a plasma display, a digital light processing display, etc., and the device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as described above in Figures 1 to 5 shown.

[0101] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc., and optionally the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.

[0102] Those skilled in the art can understand that the above-mentioned physical device structure provided by the embodiments of the present disclosure does not limit the physical device, and may include more or fewer components, or combine some components, or have different component arrangements.

[0103] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of an information processing program and other software and / or programs. The network communication module is used to implement communication between the components inside the storage medium, and communication with other hardware and software in the information processing physical device.

[0104] Based on the method as described above Figures 1 to 5 and the virtual device embodiments as Figure 6 shown, embodiments of the present disclosure further provide a chip, including one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal from a memory of an electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to execute the method as Figures 1 to 5 shown above.

[0105] Through the description of the above embodiments, those skilled in the art can clearly understand that the present disclosure can be implemented by means of software plus a necessary general hardware platform, or can be implemented by hardware. Compared with the current related technologies, the embodiments of the present disclosure do not use the method of generating image tokens in a row-by-row expansion order. Specifically, first, a category token corresponding to the generation condition of the image is determined; then, based on the category token, each image token is predicted in turn according to the diagonal generation order, where the diagonal generation order starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the direction of a preset angle diagonal, so as to effectively shorten the distance between image tokens and improve the prediction accuracy when predicting the next image token based on the previous image token; and then a complete image can be generated based on these accurately predicted image tokens. By applying the technical solution of the embodiments of the present disclosure, the accuracy of image generation can be improved, and thus the quality of image generation can be improved.

[0106] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0107] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. An image generation method, characterized in that, Including: Obtain the generation conditions of the image; Determine the class token corresponding to the generation conditions; Predict each image token in turn according to the diagonal generation order based on the class token, where the diagonal generation order starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the direction of the preset angle diagonal; Generate an image based on each image token.

2. The method according to claim 1, wherein The predicting each image token in turn according to the diagonal generation order based on the class token includes: Input the class token into an autoregressive Transformer model to predict each image token in turn according to the diagonal generation order, where a four-dimensional rotation position encoding is adopted in the autoregressive Transformer model to inject the information of the current position and the predicted position into the attention matrix at each prediction of the image token, so as to realize predicting the image token along the direction of the predicted position.

3. The method according to claim 2, wherein The adaptive normalization layer in the autoregressive Transformer model is conditioned on the direction embedding, and the direction embedding is used to inject the direction information corresponding to the diagonal generation order into the adaptive normalization layer.

4. The method according to claim 2, wherein The training process of the autoregressive Transformer model includes: Encode the sample image into a discrete sequence of image tokens; Rearrange the sequence of image tokens according to the diagonal generation order; Concatenate the sample class token corresponding to the sample image at the forefront of the rearranged sequence of image tokens, and train the autoregressive Transformer model on the concatenated sequence of image tokens in the way of predicting the next token task.

5. The method according to claim 4, characterized in that, The inputting the class token into an autoregressive Transformer model to predict each image token in turn according to the diagonal generation order includes: Input the class token into the trained autoregressive Transformer model, and predict each image token in turn from the class token in the way of predicting the next token task according to the diagonal generation order.

6. The method according to claim 4, wherein The encoding the sample image into a discrete sequence of image tokens includes: Use the codebook in the image tokenizer to encode the sample image into a discrete sequence of image tokens.

7. The method according to any one of claims 1 to 6, characterized in that, The generating an image based on each image token includes: Restore each image token in the arrangement order to a row-by-row expansion order; Convert each image token in the arrangement order into image pixels to obtain the generated image.

8. The method according to claim 7, characterized in that, The converting each image token in the arrangement order into image pixels to obtain the generated image includes: Use the decoder in the image tokenizer to convert each image token in the arrangement order into image pixels to obtain the generated image.

9. An image generation device, characterized in that, Including: An obtaining module configured to obtain the generation conditions of the image; A determining module configured to determine the class token corresponding to the generation conditions; A prediction module, configured to sequentially predict each image token based on the category token in a diagonal generation order, where the diagonal generation order starts from the upper left corner of the image and alternates between generating from the lower left corner to the upper right corner and from the upper right corner to the lower left corner along the direction of a preset angle diagonal; A generation module, configured to generate an image based on the respective image tokens.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 8.

11. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.