Figure graph model, text graph method, text graph device and storage medium
By improving the text-based image model and combining local and global feature extraction branches, the problem of balancing image output quality and efficiency in traditional text-based image models is solved, resulting in high-resolution images with clear details and natural global structure.
Patent Information
- Application Number
- CN202511108005.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-12-09
AI Technical Summary
Traditional text-based graph models struggle to balance output quality and efficiency. Models based on distillation techniques produce poor output, while models based on the Transformer architecture have high computational demands and low output efficiency.
A text-based image model is designed, comprising a first latent image generator, a second latent image generator, and a latent image decoder connected in sequence. The second latent image generator employs a symmetrically configured image compression network and image expansion network, an intermediate connection network, and a skip connection network. It uses fast Fourier convolutional blocks and feature fusion blocks, and combines local and global feature extraction branches to improve the U-Net network structure.
It improves image output efficiency and quality, producing images with clear details and natural global structure, making it suitable for generating high-resolution images and reducing inconsistencies in local areas.
Smart Images

Figure CN121095370A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a text-to-image model, a text-to-image method, a text-to-image device, and a storage medium. Background Technology
[0002] With the continuous development of artificial intelligence technology, text-based graph models have emerged like mushrooms after rain. Among them, text-based graph models such as StableDiffusion (a model of latent diffusion) have attracted widespread attention.
[0003] However, traditional text-based image models still struggle to balance output efficiency and quality. Therefore, researching and developing a text-based image model that can achieve both is of great significance for promoting technological innovation and development in text-based image models. Summary of the Invention
[0004] In view of this, embodiments of this application provide a text-to-image model, method, device, and storage medium that can balance image output quality and output efficiency.
[0005] A first aspect of this application provides a text-based image model, including:
[0006] A first latent image generator, a second latent image generator, and a latent image decoder are connected in sequence;
[0007] The second latent image generator includes a symmetrically arranged image compression network and image expansion network, an intermediate connection network, and a skip connection network; wherein, the image compression network and the image expansion network are connected through the intermediate connection network; the skip connection network is located between the image compression network and the image expansion network;
[0008] The image compression network includes multiple downsampling networks connected in series, and each downsampling network includes a first basic residual module and a downsampling module connected in series; the image expansion network includes multiple upsampling networks connected in series, and each upsampling network includes a second basic residual module and an upsampling module connected in series.
[0009] Both the first basic residual module and the second basic residual module include at least one fast Fourier convolution block and a first feature fusion block;
[0010] The Fast Fourier Convolutional Block includes a first feature extraction unit, a second feature extraction unit, a local feature extraction branch connected to the first feature extraction unit, a global feature extraction branch connected to the second feature extraction unit, a feature concatenation unit connected to the local feature extraction branch and the global feature extraction branch, and a third feature extraction unit connected to the feature concatenation unit.
[0011] The structure of the latent image decoder is the same as that of the image extension network.
[0012] A second aspect of this application provides a text-based image method based on the text-based image model of the first aspect, including:
[0013] The text description is input into the first latent image generator of the text-generated image model for encoding and feature transformation to obtain the first latent image.
[0014] The first latent image is input into the second latent image generator of the Wensheng image model for latent image generation processing to obtain the second latent image.
[0015] The second latent image is input into the latent image decoder of the Wensheng image model for decoding processing to obtain the reconstructed image.
[0016] A third aspect of this application provides a text-to-image device configured to implement the text-to-image method of the second aspect.
[0017] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0018] Compared with the prior art, the beneficial effects of this application embodiment include at least the following: This application embodiment improves upon the traditional Stable Cascade (a text-to-image generation model developed by Stability AI) architecture, designing a novel text-to-image model. This model includes: a first latent image generator, a second latent image generator, and a latent image decoder connected sequentially; the second latent image generator includes a symmetrically arranged image compression network and image expansion network, an intermediate connection network, and a skip connection network; wherein the image compression network and image expansion network are connected through the intermediate connection network; the skip connection network is located between the image compression network and the image expansion network; the image compression network includes multiple downsampling networks connected in series, and each downsampling network includes a first basic residual module and a downsampling module connected in series. The image expansion network comprises multiple cascaded upsampling networks. Each upsampling network includes a cascaded second basic residual module and an upsampling module. Both the first and second basic residual modules include at least one Fast Fourier Convolutional (FFT) block and a first feature fusion block. The FFT block includes a first feature extraction unit, a second feature extraction unit, a local feature extraction branch connected to the first feature extraction unit, a global feature extraction branch connected to the second feature extraction unit, a feature concatenation unit connected to the local and global feature extraction branches, and a third feature extraction unit connected to the feature concatenation unit. The latent image decoder has the same structure as the image expansion network. The text-based image model provided in this application can balance image output quality and efficiency, has broad market application prospects, and is of great significance for promoting the technological innovation and development of text-based image models. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the structure of a text image model provided in an embodiment of this application;
[0021] Figure 2 This is a schematic diagram of the structure of the first basic residual module provided in an embodiment of this application;
[0022] Figure 3 This is a schematic diagram of the structure of a Fast Fourier Convolutional Block provided in an embodiment of this application;
[0023] Figure 4This is a schematic diagram of the structure of a fast Fourier transform unit provided in an embodiment of this application;
[0024] Figure 5 This is a schematic diagram of the structure of a text information capture unit provided in an embodiment of this application;
[0025] Figure 6 This is a schematic diagram of the structure of a downsampling module provided in an embodiment of this application;
[0026] Figure 7 This is a schematic flowchart of a text-to-image method provided in an embodiment of this application;
[0027] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0028] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0029] Among related technologies, the most widely used textural image models are the distillation-based diffusion model and the Transformer-based diffusion model. However, while the distillation-based diffusion model reduces the number of denoising steps and speeds up image generation, the resulting image quality is poor. The Transformer-based diffusion model still produces good images after distillation, but it requires significant computation, has low efficiency, and demands high-performance computing equipment. Therefore, current textural image models still struggle to balance image quality and efficiency.
[0030] In view of this, this application improves upon the traditional Stable Cascade architecture (a text-to-image generation model developed by StabilityAI) to design a text-to-image model. This model includes three stages connected in sequence: a first latent image generator, a second latent image generator, and a latent image decoder (similar to the three-stage architecture of Stage C, Stage B, and Stage A in the Stable Cascade architecture). Each stage can efficiently and focusedly complete a specific task, thereby improving the image output efficiency. Secondly, the structure of the first latent image generator follows the network structure of Stage C in the traditional Stable Cascade architecture, while the second latent image generator and the latent image decoder adopt the improved U-Net network structure uniquely designed in this application, which can improve the image output effect.
[0031] The following will describe in detail, with reference to the accompanying drawings, a text-to-image model and a text-to-image method according to embodiments of this application.
[0032] Figure 1 This is a schematic diagram of the structure of a text-based image model provided in one embodiment of this application. Please refer to [link / reference]. Figure 1 The textural image model in this embodiment includes a first latent image generator 101, a second latent image generator 102, and a latent image decoder 103 connected in sequence. The second latent image generator 102 includes a symmetrically arranged image compression network 1021 and image expansion network 1022, an intermediate connection network 1023, and a skip connection network 1024. The image compression network 1021 and image expansion network 1022 are connected through the intermediate connection network 1023; the skip connection network 1024 is located between the image compression network 1021 and the image expansion network 1022. The image compression network 1021 includes multiple downsampling networks connected in series (e.g., downsampling network 1, downsampling network 2, ..., downsampling network n, where n is an integer ≥ 4). Each downsampling network includes a first basic residual module and a downsampling module connected in series. The image expansion network 1022 includes multiple upsampling networks connected in series (e.g., upsampling network 1, upsampling network 2, ..., upsampling network n, where n is an integer ≥ 4). Each upsampling network includes a second basic residual module and an upsampling module connected in series. The latent image decoder 103 has the same structure as the image expansion network 1022. The intermediate connection network 1023 includes a first basic residual module.
[0033] The first and second basic residual modules described above can each include at least one Fast Fourier Convolution Block (FFCB) and a first feature fusion block (Add). The number of FFCBs in the first and second basic residual modules can be set to one or more; for example, the number of FFCBs in the first and second basic residual modules can be set to two.
[0034] Figure 2 This is a schematic diagram of the structure of the first basic residual module provided in an embodiment of this application. Please refer to... Figure 2 The first basic residual module in this embodiment includes two cascaded Fast Fourier Convolutional Blocks (FFCB) and a first feature fusion block (Add). The first feature fusion block (Add) is connected to the downsampling module in the downsampling network.
[0035] Figure 3 This is a schematic diagram of the structure of a Fast Fourier Convolutional Block provided in an embodiment of this application. Please refer to... Figure 3 The Fast Fourier Convolutional Block (FFCB) of this application includes: a first feature extraction unit (which may be a combination of a fully convolutional block with a kernel size of 1×1, group normalization, and activation function (such as ReLU activation function), denoted as Conv1+GN+AC, where Conv1 represents a fully convolutional block with a kernel size of 1×1, GN (Group Normalization) represents group normalization, and AC represents activation function), a second feature extraction unit (which may be Conv1+GN+AC), a local feature extraction branch connected to the first feature extraction unit (which may be Conv1+GN+AC), a global feature extraction branch connected to the second feature extraction unit (which may be Conv1+GN+AC), a feature concatenation unit (Concat) connected to the local feature extraction branch (Local) and the global feature extraction branch (Global), and a third feature extraction unit (which may be Conv1+GN+AC) connected to the feature concatenation unit (Concat).
[0036] The technical solution provided in this application is based on the traditional Stable Cascade architecture (a text-to-image generation model developed by Stability AI) and improves it to design a text-to-image model. This text-to-image model includes three stages connected in sequence: a first latent image generator, a second latent image generator, and a latent image decoder (similar to the three-stage architecture of Stage C, Stage B, and Stage A in the Stable Cascade architecture). Each stage can efficiently and focusedly complete a specific task, thereby improving the efficiency of generating images from text, that is, improving the output efficiency.
[0037] Secondly, the first latent image generator adopts the Stage C network structure in the traditional Stable Cascade architecture. The second latent image generator and latent image decoder adopt the improved U-Net network structure uniquely designed in this application. It can capture the detailed features of the image (such as edges and texture features) by improving the local feature extraction branch (Local) in the U-Net network structure to ensure that the details of the generated image are clear. It captures the global features of the image (such as the overall structure of the image, color distribution, common shapes and positional relationships of objects, etc.) and the long-range dependencies of the image by the global feature extraction branch (Global) to ensure that the global structure of the generated image is natural and conforms to human visual perception, reducing the problem of inconsistency in local areas. At the same time, the global feature extraction branch (Global) has good resolution robustness and is suitable for generating high-resolution images. It can avoid performance degradation due to resolution changes, thereby improving the quality of the generated image, that is, improving the output image quality.
[0038] In some embodiments, please continue reading Figure 3The aforementioned local feature extraction branch may include: a first fully convolutional unit (which may be a 3×3 kernel-sized fully convolutional unit, denoted as Conv3), a second fully convolutional unit (which may be Conv3), a first feature fusion unit (Add), and a first activation normalization unit (which may be a combination of group normalization and activation function (such as ReLU activation function), denoted as GN+AC, where GN (Group Normalization) represents group normalization and AC represents activation function); the first fully convolutional unit (which may be Conv3) and the second fully convolutional unit (which may be Conv3) are respectively connected to the first feature extraction unit (which may be Conv1+GN+AC); the first fully convolutional unit (which may be Conv3) is connected to the first feature fusion unit (Add), the first feature fusion unit (Add) is connected to the first activation normalization unit (which may be GN+AC), and the first activation normalization unit (which may be GN+AC) is connected to the feature concatenation unit (Concat). The weights of the first fully convolutional unit (which may be Conv3) and the second fully convolutional unit (which may be Conv3) are different.
[0039] The aforementioned Local feature extraction branch can extract local features of an image (such as edge features, texture features, and other details) through spatial domain convolution. It can operate directly in the pixel domain and is suitable for capturing high-frequency information of the image, thereby making the details of the generated image clearer.
[0040] In some embodiments, please continue reading Figure 3 The aforementioned global feature extraction branch (Global) may include: a text information capture unit (Context Unit), a fast Fourier transform unit (Fourier Unit), a second feature fusion unit (Add), and a second activation normalization unit (which may be GN+AC, where GN (Group Normalization) represents group normalization and AC represents an activation function (such as ReLU activation function, etc.)). The text information capture unit (Context Unit) is connected to both the second feature extraction unit (Add) and the first feature fusion unit (Add). The fast Fourier transform unit (Fourier Unit) is connected to both the second feature extraction unit (Conv1+GN+AC) and the second feature fusion unit (Add). The second fully convolutional unit (which may be Conv3) is connected to the second feature fusion unit (Add), the second feature fusion unit (Add) is connected to the second activation normalization unit (which may be GN+AC), and the second activation normalization unit (which may be GN+AC) is connected to the feature concatenation unit (Concat).
[0041] Figure 4This is a schematic diagram of the structure of a Fast Fourier Transform unit provided in one embodiment of this application. Please refer to... Figure 4 The Fast Fourier Transform (FFT) unit in this embodiment may include: a first feature extraction layer (which may be Conv1+GN+AC, where Conv1 represents a fully convolutional kernel of size 1×1, GN (Group Normalization) represents group normalization, and AC represents an activation function (such as ReLU activation function)) connected in sequence, a two-dimensional real fast Fourier transform layer (Real FFT2d), a second feature extraction layer (which may be Conv1+GN+AC), a two-dimensional real inverse fast Fourier transform layer (Inv Real FFT2d), a first feature fusion layer (Add), and a third feature extraction layer (which may be Conv1+GN+AC); the first feature extraction layer (which may be Conv1+GN+AC) is connected to the second feature extraction unit (which may be Conv1+GN+AC) and the first feature fusion layer (Add); the third feature extraction layer (which may be Conv1+GN+AC) is connected to the second feature fusion unit (Add).
[0042] The aforementioned Global feature extraction branch captures low-frequency information and global features of an image (such as the overall structure of the image, color distribution, common shapes of objects, and positional relationships) through frequency domain convolution. This results in a more natural and realistic global structure in the generated image, conforming to human visual perception. Furthermore, the Global feature extraction branch can handle images of different resolutions, making it suitable for generating high-resolution images and avoiding performance degradation caused by resolution changes. It can also capture long-range dependencies in the image, ensuring global consistency and reducing inconsistencies in local areas.
[0043] Figure 5 This is a schematic diagram of the structure of a text information capture unit provided in an embodiment of this application. Please refer to [link / reference]. Figure 5The text information capture unit (Context Unit) in this application embodiment may include: a fully convolutional layer (which may be Conv1, representing a fully convolutional layer with a kernel size of 1×1), a softmax layer, a second feature fusion layer, a third feature extraction layer (which may be Conv1+GN+AC, where Conv1 represents a fully convolutional layer with a kernel size of 1×1, GN (GroupNormalization) represents group normalization, and AC represents an activation function (such as the ReLU activation function)), a fourth feature extraction layer (which may be Conv3+GN+AC, where Conv1 represents a fully convolutional layer with a kernel size of 3×3, GN (GroupNormalization) represents group normalization, and AC represents an activation function (such as the ReLU activation function)) and a third feature fusion layer, respectively; the third feature fusion layer, the fully convolutional layer (which may be Conv1), and the second feature fusion layer are respectively connected to the second feature extraction unit (which may be Conv1+GN+AC); the third feature fusion layer is connected to the first feature fusion unit (Add).
[0044] The text information capture unit described above can be used for global modeling to capture contextual information.
[0045] Figure 6 This is a schematic diagram of the structure of a downsampling module provided in one embodiment of this application. Please refer to... Figure 6 The downsampling module in this application embodiment may include: a feature map discrete decomposition block, a second feature fusion block (Concat), and a fully convolutional block (which may be a fully convolutional block with a kernel size of 1×1, denoted as Conv1) connected in sequence; the feature map discrete decomposition block is connected to the first basic residual module.
[0046] The downsampling module in this embodiment can employ a Haar wavelet transform downsampling module. Compared to traditional pooling and strided convolution downsampling methods, the downsampling module in this embodiment uses Haar wavelet transform downsampling, which can reduce the spatial resolution of the feature map, retain more information, and improve the output image quality.
[0047] Figure 7 This is a schematic flowchart illustrating a text-to-image method provided in an embodiment of this application. Please refer to... Figure 7 The text-to-image generation method of this application embodiment can be executed by a server or a terminal device, and the text-to-image generation method includes the following steps:
[0048] Step S701: Input the text description content into the first latent image generator of the text-generated image model for encoding and feature transformation processing to obtain the first latent image;
[0049] Step S702: Input the first latent image into the second latent image generator of the Wensheng image model for latent image generation processing to obtain the second latent image;
[0050] In step S703, the second latent image is input into the latent image decoder of the Wensheng image model for decoding processing to obtain the reconstructed image.
[0051] The text description contains a specific description or requirement of the image the user wants to generate. These specific descriptions or requirements can be a detailed description of the image the user wants to generate, including but not limited to: the image's subject, scene, objects, colors, style, perspective, and other aspects. For example, the text description could be: "A seaside town at sunset, with many houses in the town, a few people strolling on the beach, and waves crashing on the sea; the overall style of the scene is a realistic oil painting."
[0052] As an example, combined with Figure 1 First, the server or terminal device can obtain the text description content input by the user (which can be touch screen input, voice input, etc.); then, the text description content is input into the first latent image generator 101 of the text-based graph model, and the first latent image generator 101 encodes and performs feature transformation processing on the text description content in a highly compressed latent space to obtain the first latent image; next, the first latent image is input into the second latent image generator 102 of the text-based graph model for a first decoding process to obtain the second latent image; then, the second latent image is input into the latent image decoder 103 of the text-based graph model for a second decoding process to obtain the reconstructed image, and the reconstructed image is output.
[0053] The text-based image generation method provided in this application can quickly generate a high-quality reconstructed image that matches the text description by inputting the text description content into the text-based image generation model provided in this application.
[0054] In some embodiments, the first latent image is input into the second latent image generator of the text image model for a decoding process to obtain the second latent image, including:
[0055] The first latent image is input into the image compression network of the second latent image generator for downsampling to obtain a downsampled feature map;
[0056] The downsampled feature map is input into the intermediate connection network of the second latent image generator for processing to obtain the intermediate feature map;
[0057] The intermediate feature map is input into the image extension network of the second latent image generator for upsampling to obtain the upsampled feature map;
[0058] The second latent image is obtained by concatenating and stitching the downsampled feature map and the upsampled feature map through the skip connection network of the second latent image generator.
[0059] As an example, combined with Figure 1 Assuming the image compression network 1021 includes four cascaded downsampling networks and the image expansion network 1022 includes four cascaded upsampling networks, and n is 4; then the process of obtaining the second latent image by decoding the first latent image through the second latent image generator 102 mainly includes the following steps:
[0060] First, the first latent image output by the first latent image generator 101 is input into the downsampling network 1 of the image compression network 1021 of the second latent image generator 102. Feature extraction is performed by the first basic residual module of the downsampling network 1 to obtain a first residual feature map. The first residual feature map is then downsampled by the downsampling module in the downsampling network 1 to obtain a first downsampled feature map. Next, the first downsampled feature map is input into the downsampling network 2. Feature extraction is performed by the first basic residual module of the downsampling network 2 to obtain a second residual feature map. The second residual feature map is then downsampled by the downsampling module in the downsampling network 2 to obtain a second downsampled feature map. Next, the second downsampled feature map is input into downsampling network 3, where it undergoes feature extraction via the first basic residual module to obtain a third residual feature map. This third residual feature map is then downsampled via the downsampling module in downsampling network 3 to obtain another third downsampled feature map. Next, the third downsampled feature map is input into downsampling network 4, where it undergoes feature extraction via the first basic residual module to obtain a fourth residual feature map. This fourth residual feature map is then downsampled via the downsampling module in downsampling network 4 to obtain another fourth downsampled feature map. Finally, the fourth downsampled feature map is input into intermediate connection network 1023 for feature extraction. The intermediate feature map is obtained. Then, the intermediate feature map is input into upsampling network 1, where it is upsampled by the upsampling module to obtain a first upsampled feature map. This first upsampled feature map is then concatenated (i.e., spliced) with the fourth residual feature map by skip connection network 1024 to obtain a first spliced feature map. The first spliced feature map is then processed by the second basic residual module of upsampling network 1 for feature extraction to obtain a fifth residual feature map. Next, the fifth residual feature map is input into upsampling network 2, where it is upsampled by the upsampling module to obtain a second upsampled feature map. This second upsampled feature map is then processed by skip connection network 1024 to obtain a second upsampled feature map. The sampled feature map and the third residual feature map are concatenated (i.e., spliced) to obtain the second spliced feature map. The second spliced feature map is then processed by the second basic residual module of the sampling network 2 for feature extraction to obtain the sixth residual feature map. Next, the sixth residual feature map is input into the upsampling network 3 and upsampled by the upsampling module of the upsampling network 3 to obtain the third upsampled feature map. The third upsampled feature map is then concatenated (i.e., spliced) with the second residual feature map by the skip connection network 1024 to obtain the third spliced feature map. The third spliced feature map is then processed by the second basic residual module of the sampling network 3 for feature extraction to obtain the seventh residual feature map.Next, the seventh residual feature map is input into the upsampling network 4. After upsampling by the upsampling module of the upsampling network 4, a fourth upsampling feature map is obtained. Then, the fourth upsampling feature map is concatenated (i.e., stitched) with the first residual feature map via the skip connection network 1024 to obtain a fourth stitched feature map. This fourth stitched feature map is then used for feature extraction by the second basic residual module of the upsampling network 4 to obtain the eighth residual feature map (i.e., the second latent image).
[0061] As an example, please refer to Figures 2-6 The process of extracting features from the first latent image through the downsampling network 1 in the image compression network 1021 of the second latent image generator 102 includes the following steps:
[0062] The first latent image is input into the first Fast Fourier Convolutional Block (FFCB) of the first basic residual module of the downsampling network 1 in the image compression network 1021 for convolution processing to obtain a first convolutional feature map. Then, the first convolutional feature map is input into the second Fast Fourier Convolutional Block (FFCB) of the first basic residual module of the downsampling network 1 for convolution processing to obtain a second convolutional feature map. Finally, the second convolutional feature map and the first latent image are input into the first feature fusion block (Add) of the first basic residual module of the downsampling network 1 for feature map addition to obtain a first fused feature map. The feature map addition operation refers to adding corresponding elements of two or more feature maps with the same shape.
[0063] The process of inputting the first latent image into the first Fast Fourier Convolutional Block (FFCB) of the first basic residual module for feature extraction to obtain the first convolutional feature map mainly includes the following steps:
[0064] The first potential image is input into the first feature extraction unit (which can be Conv1+GN+AC) of the Fast Fourier Convolutional Block (FFCB) for feature extraction to obtain the first image features; the first image features are input into the first fully convolutional unit (which can be Conv3) of the Local Feature Extraction Branch for feature extraction to obtain the first feature map; the first image features are input into the second fully convolutional unit (which can be Conv3) of the Local Feature Extraction Branch for feature extraction to obtain the second feature map.
[0065] The first potential image is input into the second feature extraction unit (which can be Conv1+GN+AC) of the Fast Fourier Convolutional Block (FFCB) for feature extraction to obtain the second image features; the second image features are input into the text information capture unit (Context Unit) of the global feature extraction branch (Global) of the Fast Fourier Convolutional Block (FFCB) for feature extraction to obtain the third feature map; the second image features are input into the Fast Fourier Transform unit (Fourier Unit) of the global feature extraction branch (Global) of the Fast Fourier Convolutional Block (FFCB) for feature extraction to obtain the fourth feature map.
[0066] The first feature map output by the first fully convolutional unit (which can be Conv3) and the third feature map output by the text information capture unit (Context Unit) are input into the first feature fusion unit (Add) of the local feature extraction branch (Local) for feature map addition to obtain the first fused feature map; the first fused feature map is then input into the first activation normalization unit (which can be GN+AC) of the local feature extraction branch (Local) for processing to obtain the local feature map.
[0067] The second feature map output by the second fully convolutional unit (which can be Conv3) and the fourth feature map output by the fast Fourier transform unit are input into the second feature fusion unit (Add) of the global feature extraction branch (Global) for feature map addition to obtain the second fused feature map; the second fused feature map is then input into the second activation normalization unit (GN+AC) of the global feature extraction branch (Global) for processing to obtain the global feature map.
[0068] The local and global feature maps are input into the feature concatenation unit (Concat) to concatenate the feature maps and obtain the third image feature; the third image feature is then input into the third feature extraction unit (which can be Conv1+GN+AC) to extract the feature and obtain the first convolutional feature map.
[0069] Similarly, the process of inputting the first convolutional feature map into the second Fast Fourier Convolutional Block (FFCB) of the first basic residual module of the downsampling network 1 for convolution processing to obtain the second convolutional feature map is basically the same as the process of inputting the first latent image into the first Fast Fourier Convolutional Block (FFCB) of the first basic residual module for feature extraction to obtain the first convolutional feature map, and will not be repeated here.
[0070] As an example, the process of inputting the second image features output by the second feature extraction unit (which can be Conv1+GN+AC) into the Fast Fourier Transform unit (Fourier Unit) for feature extraction to obtain the fourth feature map mainly includes the following steps:
[0071] Please see Figure 4 The second image features are input into the first feature extraction layer (which can be Conv1+GN+AC) of the Fast Fourier Transform Unit (FFT) for feature extraction, resulting in a first extracted feature map. The first extracted feature map is then sequentially input into a two-dimensional real FFT2d layer, a second feature extraction layer (which can be Conv1+GN+AC), and a two-dimensional real FFT2d inverse FFT layer (Inv Real FFT2d) for feature extraction, resulting in a second extracted feature map. The first and second extracted feature maps are then input into a first feature fusion layer (Add) for feature map addition, resulting in a third fused feature map. The third fused feature map is then input into a third feature extraction layer (which can be Conv1+GN+AC) for feature extraction, resulting in a fourth feature map.
[0072] The first extracted feature map is transformed to the frequency domain by using a two-dimensional real fast Fourier transform layer (Real FFT2d), and convolution is performed in the frequency domain. Then, it is transformed back to the spatial domain by using a two-dimensional real fast Fourier transform layer (Inv Real FFT2d). This can better capture the low-frequency information and global structural information of the image (such as the overall structure of the image, color distribution, etc.).
[0073] As an example, the process of extracting the second image features output by the second feature extraction unit (which can be Conv1+GN+AC) into the text information capture unit (Context Unit) to obtain the third feature map mainly includes the following steps:
[0074] Please see Figure 5The second image features are sequentially input into the fully convolutional layer (which can be Conv1) and the softmax layer of the text information capture unit (Context Unit) for processing to obtain the third extracted feature map. The third extracted feature map and the second image features are input into the second feature fusion layer of the text information capture unit (Context Unit) for feature map multiplication to obtain the fourth fused feature map. The fourth fused feature map is sequentially passed through the third feature extraction layer (which can be Conv1+GN+AC) and the fourth feature extraction layer (which can be Conv3+GN+AC) of the text information capture unit (Context Unit) for feature extraction to obtain the fourth extracted feature map. The fourth extracted feature map and the second image features are input into the third feature fusion layer of the text information capture unit (Context Unit) for feature map addition to obtain the third feature map.
[0075] As an example, the process of inputting the first residual feature map output by the first basic residual module of the downsampling network 1 into the downsampling module of the downsampling network 1 for downsampling to obtain the first downsampled feature map mainly includes the following steps:
[0076] Please see Figure 6 First, the first residual feature map is input into the feature map discrete decomposition block of the downsampling module of downsampling network 1 for two-dimensional discrete wavelet first-order decomposition, resulting in one low-frequency sub-map (LL sub-map) and three high-frequency sub-maps (LH sub-map, HL sub-map, and HH sub-map). This yields the downsampled lossless encoding of the first residual feature map. Next, the LL sub-map, LH sub-map, HL sub-map, and HH sub-map are sequentially input into the second feature fusion block (Concat) of the downsampling module of downsampling network 1 for feature map concatenation. Then, the feature maps are passed through a fully convolutional block (e.g., a 1×1 kernel full convolution) to learn frequency domain features, resulting in the first downsampled feature map.
[0077] Feature map concatenation is the operation of connecting and merging multiple feature maps along a specific dimension. In convolutional neural networks, feature maps are typically tensors with multiple dimensions, commonly including height (H), width (W), and number of channels (C). Generally, feature map concatenation is performed along the channel dimension, connecting the channels of different feature maps.
[0078] The LL submap is an approximate representation of the original image (here referring to the first residual feature map), which is obtained by convolving the horizontal and vertical directions with low-pass wavelet filters using wavelet coefficients.
[0079] The HL sub-image is a horizontal detail sub-image of the original image, used to highlight the singular characteristics of the image in the horizontal direction. It is obtained by convolving the image horizontally with a high-pass wavelet filter and then convolving it vertically with a low-pass wavelet filter.
[0080] The LH sub-image is a vertical detail sub-image of the original image, used to highlight the singular characteristics of the image in the vertical direction. It is obtained by convolving a low-pass wavelet filter in the horizontal direction and then convolving it in the vertical direction with a high-pass wavelet filter.
[0081] The HH subimage is a diagonal detail subimage of the original image, used to highlight the diagonal edge characteristics of the image. It is obtained by convolving the horizontal and vertical directions with high-pass wavelet filters and using wavelet coefficients.
[0082] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0083] In some embodiments, the training processes of the first latent image generator 101, the second latent image generator 102, and the latent image decoder 103 in the text-based graph model provided in this application are performed independently. The first latent image generator 101 can employ the Omnivision-968M model (a compact, less than one billion parameter (968M) multimodal model developed by Nexa AI) as the multimodal model for extracting textual information from the text description content. The second latent image generator 102 and the latent image decoder 103 can employ the Rectified Flow diffusion model. Compared to traditional text-based graph diffusion models based on distillation techniques, the Rectified Flow diffusion model offers advantages such as better output quality and fewer sampling steps.
[0084] As an example, the training process of the second latent image generator 102 and the latent image decoder 103 is as follows:
[0085] 1) Obtain training data pairs (X0, X1), where X0 is the initial noise latent sampled from the noise distribution π0, and X1 is the real image latent sampled from the target image distribution π1.
[0086] 2) Use the training data to train the initial Rectified Flow model until the preset convergence condition is met (such as a low or basically unchanged loss function value) to obtain the trained Rectified Flow diffusion model.
[0087] In the iterative training process of step 2) above, the training objective is to minimize the loss function L. An optimizer (such as Adam) can be used to update the model parameters during training.
[0088] The loss function L can be the mean squared error (MSE), and its mathematical expression is as follows:
[0089]
[0090] In the formula, X0 is the initial noise latent sampled from the noise distribution π0; X1 is the real image latent sampled from the target image distribution π1; v represents the target velocity field; X t X represents the midpoint between a point on the noise distribution π0 and a point on the latent distribution π1 of the target image. t = tX1 + (1-t)X0, where t∈[0,1] represents the time step; v(X t ,t) represents the predicted velocity field.
[0091] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0092] This application also provides a text-to-image device configured to implement the text-to-image method described above.
[0093] Figure 8 This is a schematic diagram of the electronic device 800 provided in an embodiment of this application. For example... Figure 8 As shown, the electronic device 800 of this embodiment includes: a processor 801, a memory 802, and a computer program 803 stored in the memory 802 and executable on the processor 801. When the processor 801 executes the computer program 803, it implements the steps in the various method embodiments described above. Alternatively, when the processor 801 executes the computer program 803, it implements the functions of each module / unit in the various device embodiments described above.
[0094] Electronic device 800 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 800 may include, but is not limited to, a processor 801 and a memory 802. Those skilled in the art will understand that... Figure 8 This is merely an example of electronic device 800 and does not constitute a limitation on electronic device 800. It may include more or fewer components than shown, or different components.
[0095] The processor 801 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0096] The memory 802 can be an internal storage unit of the electronic device 800, such as a hard disk or RAM of the electronic device 800. The memory 802 can also be an external storage device of the electronic device 800, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the electronic device 800. The memory 802 can also include both internal and external storage units of the electronic device 800. The memory 802 is used to store computer programs and other programs and data required by the electronic device.
[0097] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0098] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in a computer-readable medium can be appropriately added to or subtracted according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0099] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A text-based image model, characterized in that, include: A first latent image generator, a second latent image generator, and a latent image decoder are connected in sequence; The second latent image generator includes a symmetrically arranged image compression network and image expansion network, an intermediate connection network, and a skip connection network; wherein the image compression network and image expansion network are connected through the intermediate connection network; and the skip connection network is located between the image compression network and the image expansion network. The image compression network includes multiple downsampling networks connected in series, and each downsampling network includes a first basic residual module and a downsampling module connected in series; the image expansion network includes multiple upsampling networks connected in series, and each upsampling network includes a second basic residual module and an upsampling module connected in series. Both the first basic residual module and the second basic residual module include at least one fast Fourier convolution block and a first feature fusion block; The fast Fourier convolution block includes a first feature extraction unit, a second feature extraction unit, a local feature extraction branch connected to the first feature extraction unit, a global feature extraction branch connected to the second feature extraction unit, a feature splicing unit connected to the local feature extraction branch and the global feature extraction branch, and a third feature extraction unit connected to the feature splicing unit. The latent image decoder has the same structure as the image extension network.
2. The text-based image model according to claim 1, characterized in that, The local feature extraction branch includes a first fully convolutional unit, a second fully convolutional unit, a first feature fusion unit, and a first activation normalization unit; The first fully convolutional unit and the second fully convolutional unit are respectively connected to the first feature extraction unit; The first fully convolutional unit is connected to the first feature fusion unit, the first feature fusion unit is connected to the first activation normalization unit, and the first activation normalization unit is connected to the feature splicing unit.
3. The text-based image model according to claim 2, characterized in that, The global feature extraction branch includes a text information capture unit, a fast Fourier transform unit, a second feature fusion unit, and a second activation normalization unit; The text information capture unit is connected to the second feature extraction unit and the first feature fusion unit; The fast Fourier transform unit is connected to the second feature extraction unit and the second feature fusion unit; The second fully convolutional unit is connected to the second feature fusion unit, the second feature fusion unit is connected to the second activation normalization unit, and the second activation normalization unit is connected to the feature splicing unit.
4. The text-based image model according to claim 3, characterized in that, The Fast Fourier Transform (FFT) unit comprises a first feature extraction layer, a two-dimensional real-number FFT layer, a second feature extraction layer, a two-dimensional real-number inverse FFT layer, a first feature fusion layer, and a third feature extraction layer connected in sequence. The first feature extraction layer is connected to the second feature extraction unit and the first feature fusion layer; The third feature extraction layer is connected to the second feature fusion unit.
5. The text-based image model according to claim 3, characterized in that, The text information capture unit includes a fully convolutional layer, a softmax layer, a second feature fusion layer, a third feature extraction layer, a fourth feature extraction layer, and a third feature fusion layer connected in sequence. The third feature fusion layer, the fully convolutional layer, and the second feature fusion layer are respectively connected to the second feature extraction unit; The third feature fusion layer is connected to the first feature fusion unit.
6. The text-based image model according to claim 1, characterized in that, The downsampling module includes a feature map discrete decomposition block, a second feature fusion block, and a fully convolutional block connected in sequence. The feature map discrete decomposition block is connected to the first basic residual module.
7. A text-based image generation method based on the text-based image model as described in any one of claims 1 to 6, characterized in that, include: The text description is input into the first latent image generator of the text-generated image model for encoding and feature transformation to obtain the first latent image. The first latent image is input into the second latent image generator of the Wensheng image model for latent image generation processing to obtain the second latent image; The second latent image is input into the latent image decoder of the Wensheng image model for decoding processing to obtain the reconstructed image.
8. The method according to claim 7, characterized in that, The first latent image is input into the second latent image generator of the Wensheng image model for a decoding process to obtain the second latent image, including: The first latent image is input into the image compression network of the second latent image generator for downsampling to obtain a downsampled feature map; The downsampled feature map is input into the intermediate connection network of the second latent image generator for processing to obtain an intermediate feature map; The intermediate feature map is input into the image extension network of the second latent image generator for upsampling to obtain an upsampled feature map. The downsampled feature map and the upsampled feature map are concatenated and stitched together by the skip connection network of the second latent image generator to obtain the second latent image.
9. A text-to-image processing device, characterized in that, The texturing device is configured to implement the texturing method of claim 7 or 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in claim 7 or 8.