Visual task generation method based on field Token

CN120655756AInactive Publication Date: 2025-09-16BEIJING DIGITAL FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510746992.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing autoregressive image generation methods lack the effective integration of fine-grained domain knowledge in domain-specific image generation, resulting in the generated images having low constraints on the target domain's unique visual style, content elements, or structure, making it difficult to achieve accurate and flexible domain feature control.

Method used

A domain token-based visual task generation method is adopted. Through domain feature extraction and lemmatization, domain lemmas are used to guide image generation in the autoregressive prediction process. The self-attention and cross-attention mechanisms are combined to ensure that the generated images conform to the subtle features of the target domain.

Benefits of technology

It improves the fidelity and consistency of image generation in specific domains, enhances the fine controllability of the domain style, elements and structure of the generated content, and solves the problem that fine-grained domain knowledge is difficult to integrate into the autoregressive generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655756A_ABST
    Figure CN120655756A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence image generation, in particular to a visual task generation method based on domain Token. Comprising the steps of obtaining field feature information of a target field and generating or determining at least one field lexical element; initializing an initial representation of an image to be generated; in an autoregression generation process, according to an image lexical element sequence corresponding to a previously generated image part and the at least one field lexical element, utilizing an autoregression model to predict a next image lexical element of the current position; repeating the prediction step until an image lexical element sequence representing the complete image is generated; and generating a final pixel space image by decoding based on the image lexical element sequence representing the complete image. According to the method, by introducing and fusing the field lexical elements, the field specific attributes of the generated image can be more accurately controlled, the generation fidelity and consistency of the image in the specific field are remarkably improved, and the controllability of the field style, element and structure of the generated content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence image generation technology, and specifically to a visual task generation method based on domain tokens. Background Art

[0002] In recent years, with the development of deep learning technology, AI has made significant progress in the field of image generation. With the tremendous success of the Transformer architecture in natural language processing (NLP), its powerful sequence modeling capabilities have been applied to image generation tasks. However, existing methods still face numerous challenges in domain-specific image generation. While general-purpose image generation models can generate diverse images, they often struggle when it comes to generating highly realistic images that reflect the unique visual characteristics of specific domains (such as specific artistic styles, specific types of medical images, and satellite imagery of specific scenes). Current conditional generation methods typically rely on global conditional information, such as category labels or textual cues. While these global conditions can guide generation, they often struggle to control fine-grained domain characteristics, such as specific textures, subtle structures, and unique lighting atmospheres.

[0003] In order to achieve more refined content and style control, researchers have explored a variety of embedding or word-based methods. For example, by learning learnable identifiers to represent fine-grained categories, or by using multimodal embeddings to fuse control signals from different sources. Some work is also dedicated to developing unified visual word-mergers, such as UniTok and UniToken, which aims to bridge the different needs of visual understanding and generation tasks for word-merge representation, which indirectly involves how to make words more semantically rich, which may help to express domain information. However, most of these methods do not design an explicit and efficient mechanism specifically for autoregressive models to utilize specialized "domain words" to directly guide and constrain the autoregressive generation process of images, thereby ensuring that the generated content strictly follows the subtle characteristics of the target domain at the pixel or image block level.

[0004] In summary, despite significant progress in autoregressive image generation, there is still a lack of an autoregressive method that can effectively integrate fine-grained domain knowledge, achieve precise domain feature control, and generate high-fidelity images for domain-specific image generation. Existing methods still have room for improvement in terms of the specificity of domain information representation, the depth of integration with the autoregressive process, and the quality of generation. Summary of the Invention

[0005] In response to the problems existing in existing autoregressive image generation methods in generating images in specific domains, such as the low conformity of the generated images with the unique visual style, content elements or structural constraints of the target domain, the lack of precision and flexibility in controlling specific domain features, and the difficulty in effectively integrating fine-grained, highly specific domain knowledge into the autoregressive word-by-word generation process, the present invention provides a visual task generation method based on domain tokens.

[0006] A method for generating visual tasks based on domain tokens, comprising: S1. Domain feature extraction and lemma: obtaining domain feature information of at least one target domain, and generating or determining at least one domain lemma based on the domain feature information.

[0007] Preferably, the domain feature information may be derived from a set of representative sample images of the target domain, metadata describing the domain (such as text description, attribute tags, etc.), or a combination thereof.

[0008] Preferably, the generation or determination of the domain word-unit may include: processing the domain feature information through a pre-trained or jointly trained domain encoder network with the main model to obtain an embedding vector as the domain word-unit, or searching a predefined domain codebook for a codeword index that best matches the domain feature information as the domain word-unit.

[0009] S2. Image representation initialization: Initialize an initial representation of the image to be generated, for example, a special starting word sequence, a low-resolution image patch, or a potential representation composed of random noise.

[0010] S3. Domain-aware autoregressive prediction: Based on the image word-meta sequence corresponding to the previously generated image part and the at least one domain word-meta generated or determined in step S1, a pre-set or trained autoregressive model is used to predict the next image word-meta at the current position.

[0011] Preferably, the domain word-grams influence the prediction of the autoregressive model through one or more preset fusion mechanisms, wherein the fusion mechanisms include: splicing the domain word as a conditional input into the representation of the previously generated image word sequence, and feeding both into the main structure of the autoregressive model; Adjust the attention mechanism within the autoregressive model (e.g., the self-attention or cross-attention module in the Transformer) so that the generation of image tokens can perceive and utilize the domain information carried by the domain tokens. For example, domain tokens can be used as keys and values ​​in the cross-attention module, while the image token representation can be used as the query. The domain word is directly modulated by the autoregressive model to predict the output probability distribution of the next image word. For example, the generation probability of different candidate image words is adjusted through a gating mechanism or bias term to bias the generation of words that are more consistent with the domain characteristics.

[0012] S4 iterative generation and termination: repeatedly performing the domain-aware autoregressive prediction described in step S3, appending the newly generated image word-grams to the image word-gram sequence, and using them as the previously generated image portion for the next prediction, until a predetermined number of image word-grams are generated, or a specific termination word-gram is generated, thereby obtaining an image word-gram sequence representing the complete image.

[0013] S5. Image decoding and output: Based on the image word sequence representing the complete image obtained in step S4, it is decoded and synthesized into the final pixel space image through an image decoder (such as the decoder part of VQ-VAE or an independent generation network).

[0014] The present invention also includes a domain token-based visual task generation system for executing any of the above methods, comprising a domain token processing module, an image representation initialization module, an autoregressive prediction module, an iteration control module, and an image decoding module; The domain word-gram processing module is used to obtain domain feature information of at least one target domain, and generate or determine at least one domain word-gram based on the domain feature information; The image representation initialization module is used to initialize an initial representation of an image to be generated; The autoregressive prediction module is configured to predict the next image word at the current position using a predetermined autoregressive model based on a sequence of image word-grams corresponding to a previously generated image portion and the at least one domain word-gram, wherein the at least one domain word-gram influences the prediction of the autoregressive model through one or more preset fusion mechanisms; The iterative control module is used to repeatedly activate the autoregressive prediction module until an image word sequence representing a complete image is generated; The image decoding module is used to generate a final pixel space image by decoding through an image decoder based on the image word sequence representing the complete image.

[0015] Compared with the prior art, the advantages of the present invention are: Improving the fidelity and consistency of domain-specific image generation: By introducing domain tokens that directly represent the core features of the target domain, this method can more accurately guide the autoregressive model to generate images that conform to the visual specifications of the specific domain. These domain tokens carry richer and more detailed domain information than global category labels or general textual cues, making the generated images highly consistent with real-domain samples in terms of style, texture, structure, color, etc., thereby improving the fidelity and domain consistency of the generated images.

[0016] Enhanced control over the domain style, elements, and structure of generated content: The domain tokens of the present invention can be designed to represent different aspects of a domain. By manipulating or combining these domain tokens, users can exercise more detailed and direct control over the domain-related properties of generated images, enabling them to produce images that better suit specific creative intent or application requirements. For example, they can change the subject matter of generated content while maintaining the "Van Gogh style."

[0017] This effectively addresses the difficulty of integrating fine-grained domain knowledge into the autoregressive generation process: When processing images, traditional autoregressive models may not contain sufficient domain-specific information in their image tokens. This invention provides a direct path for incorporating fine-grained domain knowledge through specially designed domain tokens and their specific integration with the autoregressive prediction core (such as the attention mechanism or output layer). This enables the model to "perceive" and "follow" domain constraints at every decision-making step, effectively addressing the issues of insufficient domain knowledge transfer and limited impact in previous approaches. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of a method for generating visual tasks based on domain tokens proposed in the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0021] Example 1: Reference Figure 1One embodiment of the present invention provides a method for generating visual tasks based on domain tokens, including: S1. Domain feature extraction and lemmatization: First, the system receives information about the target domain, which can be a domain ID, a set of reference images, text descriptions, etc. Then, the domain lemma processing module processes this information and outputs one or more domain lemmas. These domain terms Encapsulates the key features of the target domain.

[0022] S2. Image representation initialization: Initialize an initial representation of the image to be generated. The initial representation can be a special "sequence start" word, or a latent vector initialized by random noise, or a very small image block.

[0023] S3. Domain-aware autoregressive prediction: at time step (Assuming that Start), the autoregressive prediction module performs the following operations: Receive end time step Generated image word sequence (for , the sequence is empty or contains only the initial representation).

[0024] At the same time, receive the determined domain word .

[0025] Utilizes its internal Transformer layer and one or more domain word fusion mechanisms to process the input image word sequence and domain word Its goal is to compute the word given the previous image and domain constraints Under the condition of The conditional probability distribution of .

[0026] Sample or select (e.g. greedily select the word with the highest probability) an image word from the probability distribution as the output of the current time step .

[0027] S4, iterative generation and termination: after generating a new image word After that, the system determines whether all word units representing the complete image have been generated by: Reaching a predetermined sequence length: If the image is represented as a fixed number of tokens (e.g., an N×N token grid), the generation process ends when the number of generated tokens reaches this predetermined value.

[0028] Generate a termination token: Similar to natural language generation, you can include a special "end of sequence" token in your vocabulary. When the model generates this "end of sequence" token, it indicates that image generation is complete.

[0029] If the complete image has not been generated (“no”), the newly generated word Append it to the generated sequence, and the process returns to step S3 to continue predicting the next word. If the complete image has been generated (judged "yes"), proceed to the next step.

[0030] Image decoding and output: When the complete image word sequence (where T is the final sequence length) Once generated, the sequence is fed into the image decoding module. This module converts the discrete token sequence back into a continuous pixel space, generating the final image. For example, if using the VQ-VAE framework, the decoder maps each token index to the corresponding vector in the codebook and then reconstructs the image through a series of operations such as transposed convolutions.

[0031] Example 2: Based on the domain token-based visual task generation method provided in Example 1, this example specifically describes the definition, representation, and fusion mechanism of domain tokens.

[0032] The fields of the present invention can be defined according to various criteria, for example: Artistic style: such as "Impressionist painting style", "Japanese Ukiyo-e style", "Cyberpunk art style", etc. Its characteristics may include specific brushstrokes, color application, composition preferences, light and shadow processing, etc.

[0033] Image categories and content: For example, "images of specific bird species," "CAD renderings of industrial parts," "city nightscape photos," etc., whose characteristics may include the object's morphological structure, typical background, and common element combinations.

[0034] Imaging conditions and sources: such as "cell images under a specific type of microscope", "remote sensing images taken by a specific satellite", "black-and-white photos from a specific historical period", etc., whose characteristics may include the image's noise pattern, resolution characteristics, specific artifacts or distortions, etc.

[0035] The present invention aims to capture the visual and semantic information that is specific to these domains and can distinguish between different domains.

[0036] In the present invention, domain terms can be extracted in the following ways: Learned Domain Embeddings: A dedicated embedding layer or a small neural network is trained for each predefined domain. The input can be a discrete ID representing the domain identity (e.g., represented by a one-hot encoding) or a set of global features extracted from representative images of that domain. The network learns to map the domain identity or domain features into a compact, fixed-dimensional vector space. The resulting embedding vector is the domain token for that domain.

[0037] Dynamic Domain Token Extraction: An additional encoder network (called domain encoder or style encoder) is trained that takes a reference image as input and outputs one or more domain tokens that characterize the domain of the reference image. The reference image is provided by the user during generation.

[0038] Fusion-based domain tokens: Domain information from different sources is fused to jointly generate domain tokens. For example, a textual hint describing the target domain can be converted into a text embedding through a text encoder, while visual feature embeddings are extracted from a reference image. These embeddings of different modalities are then fused into a unified domain token through a fusion module (such as an attention mechanism, a multi-layer perceptron, or a more complex fusion network).

[0039] It should be noted that the domain word-unit can be represented as a continuous vector or a discrete word-unit. A continuous vector is a real-valued vector of fixed dimension, whose dimension can be adjusted based on domain complexity and model capacity. Discrete word-units can be codeword indices selected from a codebook specifically designed for domain feature learning.

[0040] In the present invention, the generated domain word-units are integrated into the prediction process of autoregressive prediction, and the following fusion strategy can be adopted: As conditional input: The domain token (or its transformed representation) can be concatenated or added with the image token embedding (or the context embedding of the previous image token sequence) at each time step and then used together as input to the autoregressive model.

[0041] Adjusting the attention mechanism: In Transformer-based autoregressive models, the attention mechanism is the core of capturing dependencies. Domain tokens can be used to adjust the behavior of the attention module: Cross-Attention: A cross-attention module can be introduced into the Transformer decoder layer. In this module, the representation of the image portion currently being generated serves as the query, and the domain tokens serve as the key and value. This allows the model to "query" the domain token for each image token generated, obtaining domain information relevant to the current generation context.

[0042] Modulating Self-Attention: Domain tokens can be used to modulate the parameters of the self-attention module. For example, affine transformations can be used to adjust the attention weights or output values ​​calculated by the self-attention algorithm, thereby making the self-attention mechanism focus more on image features relevant to the target domain. Domain tokens can also be used through gating mechanisms to control the activation level or information flow of different attention heads.

[0043] Directly Influencing the Prediction Head: Domain tokens can directly affect the prediction head of the autoregressive model, the part responsible for generating the probability distribution of the next image token from the model's final hidden state. For example, domain tokens can be fused with the model's final hidden state, or used to generate a bias or scaling factor that directly adjusts the logit values ​​of each image token in the output vocabulary, thereby increasing the probability of generating image tokens consistent with the target domain.

[0044] Example 3: Based on the domain token-based visual task generation method provided in Example 1, this example specifically describes the training process of the autoregressive model, including: Dataset preparation: Collect a high-quality image dataset and determine the specific form of the image dataset based on how the domain tokens are generated: Datasets with domain labels: If domain tokens are generated by learning domain embeddings or require a supervised signal to learn to extract from reference images, an explicit domain label is associated with each image in the training dataset (for example, "landscape painting - impressionism", "medical imaging - lung CT - nodule positive", "animal portrait - cat - ragdoll cat", etc.).

[0045] Datasets without domain labels but groupable by domain: If domain tokens are extracted from reference images in a completely unsupervised manner, the training data is divided into different domain subsets to provide the model with samples in the same domain as reference during training.

[0046] Multimodal dataset: If fused domain tokens are used, such as combining text descriptions and reference images, the dataset needs to contain images and their paired domain description text and / or domain reference images.

[0047] Multi-model training: Image encoder, domain encoder and autoregressive model are trained based on the above dataset: Image encoder VQ-VAE pre-training: Input: Raw image data.

[0048] Goal: Encode the image into discrete tokens through vector quantization and reconstruct the image through the decoder.

[0049] Loss function: Reconstruction loss: measures the pixel difference between the generated image and the original image.

[0050] Commitment loss: ensures that the encoder output is close to the codebook vector.

[0051] Codebook loss: maintains the stability of codebook vector updates.

[0052] Domain encoder pre-training: Input: domain sample images + metadata.

[0053] Goal: To make domain-specific features represented by domain-words.

[0054] Loss function: contrast loss, which forces the domain word distance of samples in the same domain to be smaller than that of samples in different domains.

[0055] End-to-end joint training of autoregressive models: Input data organization: Image word sequence: obtained by discretization through VQ-VAE, with a length of 256 (16×16), and a starting word is added.

[0056] Domain word: Generated by the domain encoder, dimension 128, and used as conditional input throughout the autoregressive process.

[0057] Forward propagation process: Input the initial word and domain word into the autoregressive model.

[0058] The model processes the generated word sequence through the self-attention module and fuses the domain word information through the cross-attention module.

[0059] Predict the probability distribution of the next word, sample or greedily select the word, and compare it with the actual word sequence to calculate the loss.

[0060] Loss function design: Autoregressive loss: cross entropy loss, which measures the distribution difference between the predicted word and the true word.

[0061] Domain consistency loss: Introduce a domain classifier (such as ResNet-18) to determine the domain of the generated image.

[0062] The loss function is cross entropy, which ensures that the generated image conforms to the characteristics of the target domain.

[0063] Perceptual Loss: Use the pre-trained VGG network to extract high-level features of generated images and real images.

[0064] The loss function is the L2 distance of the feature vector to ensure style and structure consistency.

[0065] Total loss (in, is the autoregressive loss, is the domain consistency loss, is the perceptual loss).

[0066] Optimized configuration: Optimizer: AdamW, initial learning rate 1e-4, using cosine annealing decay strategy.

[0067] Batch size: 32, enable mixed precision training to accelerate convergence.

[0068] Training rounds: 100 rounds, with FID and domain classification accuracy evaluated on the validation set every 5 rounds.

[0069] Early stopping strategy: If the validation set loss does not decrease for 10 consecutive rounds, stop training.

[0070] Example 4: Based on Example 1, an embodiment of the present invention provides a domain token-based visual task generation system, including a domain word unit processing module, an image representation initialization module, an autoregressive prediction module, an iteration control module, and an image decoding module; The domain word-gram processing module is used to obtain domain feature information of at least one target domain, and generate or determine at least one domain word-gram based on the domain feature information; The image representation initialization module is used to initialize an initial representation of an image to be generated; The autoregressive prediction module is configured to predict the next image word at the current position using a predetermined autoregressive model based on a sequence of image word-grams corresponding to a previously generated image portion and the at least one domain word-gram, wherein the at least one domain word-gram influences the prediction of the autoregressive model through one or more preset fusion mechanisms; The iterative control module is used to repeatedly activate the autoregressive prediction module until an image word sequence representing a complete image is generated; The image decoding module is used to generate a final pixel space image by decoding through an image decoder based on the image word sequence representing the complete image.

[0071] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0072] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A visual task generation method based on domain tokens, characterized in that: include: S1. Domain feature extraction and lemma: obtaining domain feature information of at least one target domain, and generating or determining at least one domain lemma based on the domain feature information; S2. Image representation initialization: Initializing an initial representation of the image to be generated, wherein the initial representation includes but is not limited to a special starting word sequence, a low-resolution image patch, or a potential representation composed of random noise; S3, domain-aware autoregressive prediction: predicting the next image word at the current position using a pre-trained autoregressive model based on the image word sequence corresponding to the generated image portion and the at least one domain word generated or determined in step S1, wherein the domain word influences the prediction of the autoregressive model through multiple preset fusion mechanisms; S4, iterative generation and termination: repeatedly performing the domain-aware autoregressive prediction in step S3, appending the newly generated image tokens to the image token sequence and using them as the previously generated image portion for the next prediction, until a predetermined number of image tokens are generated, or a specific termination token is generated, thereby obtaining an image token sequence representing the complete image; S5, image decoding and output: Based on the image word sequence representing the complete image obtained in step S4, an image decoder is used to decode it and synthesize it into a final pixel space image; The feature information is derived from a set of representative sample images of the target domain, metadata describing the target domain, or a combination thereof.

2. A method for generating visual tasks based on domain tokens according to claim 1, characterized in that: The generating or determining at least one domain word-unit based on the domain feature information includes: processing the domain feature information through a domain encoder network to obtain an embedding vector as the domain word-unit.

3. A method for generating visual tasks based on domain tokens according to claim 1, characterized in that: The generating or determining at least one domain word-element based on the domain characteristic information further includes: searching a predefined domain codebook for a codeword index that best matches the domain characteristic information as the domain word-element.

4. A method for generating visual tasks based on domain tokens according to claim 1, characterized in that: The preset fusion mechanism includes: splicing the at least one domain word as a conditional input into the representation of the previously generated image word sequence, and feeding them together into the main structure of the predetermined autoregressive model.

5. The method for generating visual tasks based on domain tokens according to claim 1, characterized in that: The autoregressive model includes an attention mechanism, and the preset fusion mechanism includes: using the at least one domain word to adjust the attention mechanism so that the generation of image words can perceive and utilize the domain information carried by the domain word.

6. A method for generating visual tasks based on domain tokens according to claim 5, characterized in that: The attention mechanism is a cross-attention module, the domain word-gram serves as the key and / or value of the cross-attention module, and the representation of the generated image word-gram sequence serves as the query.

7. The method for generating visual tasks based on domain tokens according to claim 1, characterized in that: The preset fusion mechanism includes: using the at least one domain word to directly modulate the autoregressive model to predict the output probability distribution of the next image word.

8. The method for generating visual tasks based on domain tokens according to claim 1, characterized in that: The autoregressive model is an improved model based on the Transformer architecture or the PixelCNN architecture.

9. The method for generating visual tasks based on domain tokens according to claim 1, characterized in that: It also includes using a loss function including autoregressive loss and domain consistency loss or perceptual loss when training the autoregressive model.

10. A visual task generation system based on domain tokens, characterized by: It includes domain word processing module, image representation initialization module, autoregressive prediction module, iteration control module and image decoding module; The domain word-gram processing module is used to obtain domain feature information of at least one target domain, and generate or determine at least one domain word-gram based on the domain feature information; The image representation initialization module is used to initialize an initial representation of an image to be generated; The autoregressive prediction module is configured to predict the next image word at the current position using a predetermined autoregressive model based on a sequence of image word-grams corresponding to a previously generated image portion and the at least one domain word-gram, wherein the at least one domain word-gram influences the prediction of the autoregressive model through one or more preset fusion mechanisms; The iterative control module is used to repeatedly activate the autoregressive prediction module until an image word sequence representing a complete image is generated; The image decoding module is used to generate a final pixel space image by decoding through an image decoder based on the image word sequence representing the complete image.

Citation Information

Cited By

  • System and method for processing brain image

    CN120997405A

  • Systems and methods for processing brain images

    CN120997405B