A conditional image generation method, apparatus, device and medium
Patent Information
- Application Number
- CN202610708289.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-18
AI Technical Summary
[0004]本发明的主要目的在于提供一种条件图像生成方法、装置、设备与介质,旨在解决现有技术中基于流模型框架进行条件图像生成时难以适应条件生成的动态需求和计算开销大的技术问题,提高条件图像生成的准确性和生成效率
[0009]Beneficial Effects: This invention discloses a conditional image generation method, apparatus, device, and medium. Compared to existing technologies, this invention receives input conditions for guiding image generation and encodes these conditions to obtain corresponding conditional vectors. Initial noise is randomly sampled as initial latent variables, and the initial latent variables and the conditional vectors are input into a pre-trained streaming model. The streaming model iteratively transforms the initial latent variables and conditional vectors for a preset number of time steps. Each time step performs dynamic sparse constraint attention feature fusion and transformation based on the current latent variables and conditional vectors. After completing the preset number of iterative transformations, the final output feature transformation result is used as the target image corresponding to the input conditions. This invention can be applied to business scenarios such as fintech and healthcare. By performing dynamic sparse constraint attention feature fusion and transformation on latent variables and conditional vectors during iterative transformation, image generation can dynamically adjust the attention mode and constrain the attention range based on different conditions. This saves computational overhead while adapting to different generation requirements, improving image generation quality and efficiency.
Smart Images

Figure CN122597543A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image generation technology and can be applied to business areas such as fintech and healthcare. In particular, it relates to a conditional image generation method, apparatus, device, and medium. Background Technology
[0002] Conditional image generation, a type of generative artificial intelligence, refers to generating images that meet specific conditions (such as text descriptions, category labels, and other images). It has significant applications in the financial and healthcare sectors. For example, in anti-fraud applications in finance, conditionally generated image data matching transaction scenarios (such as abnormal payment interfaces) can be used to train fraud detection models, improving their ability to identify forged transactions and false identities. Alternatively, in financial product marketing, personalized product marketing images can be efficiently generated based on product characteristics and promotional needs. In the medical education and training sector, conditional generation can produce relevant teaching and training images. It can also batch generate medical image samples based on specified conditions to provide controllable training samples for AI models in medical image recognition and target localization.
[0003] Currently, conditional generation based on flow models has made significant progress in the field of image synthesis, achieving controllable image synthesis by injecting demand information into the flow model framework. However, existing conditional image generation based on flow models still faces challenges in adapting to the dynamic requirements of conditional generation and incurring high computational costs, thus affecting the accuracy and efficiency of conditional image generation. Summary of the Invention
[0004] The main objective of this invention is to provide a conditional image generation method, apparatus, device, and medium, aiming to solve the technical problems of difficulty in adapting to the dynamic requirements of conditional image generation and high computational overhead when using a flow model framework in the prior art, thereby improving the accuracy and efficiency of conditional image generation.
[0005] The technical solution of the present invention is as follows: The first aspect of this invention provides a conditional image generation method, comprising: Receive input conditions for guiding image generation, and encode the input conditions to obtain a corresponding condition vector; Randomly sample initial noise as initial latent variables, and input the initial latent variables and the conditional vector into the pre-trained stream model; The initial latent variables and condition vectors are iteratively transformed by the flow model for a preset number of time steps. Each time step performs feature fusion and transformation based on the current latent variables and condition vectors with dynamic sparse constraint attention. After completing the iterative transformation for a preset number of time steps, the final output feature transformation result is used as the target image corresponding to the input conditions.
[0006] A second aspect of the present invention provides a conditional image generation apparatus, comprising: The receiving module is used to receive input conditions for guiding image generation, and to encode the input conditions to obtain a corresponding condition vector; The sampling input module is used to randomly sample initial noise as initial latent variables, and input the initial latent variables and the condition vector into the pre-trained stream model; The feature transformation module is used to perform a preset number of time steps of iterative transformation on the initial latent variables and condition vectors through the flow model. Each time step is based on the current latent variables and condition vectors to perform feature fusion and transformation with dynamic sparse constraint attention. The image output module is used to output the final feature transformation result as the target image corresponding to the input conditions after completing the iterative transformation of a preset number of time steps.
[0007] A third aspect of the present invention provides a computer device including at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the conditional image generation method described above.
[0008] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the conditional image generation method described above.
[0009] Beneficial Effects: This invention discloses a conditional image generation method, apparatus, device, and medium. Compared to existing technologies, this invention receives input conditions for guiding image generation and encodes these conditions to obtain corresponding conditional vectors. Initial noise is randomly sampled as initial latent variables, and the initial latent variables and the conditional vectors are input into a pre-trained streaming model. The streaming model iteratively transforms the initial latent variables and conditional vectors for a preset number of time steps. Each time step performs dynamic sparse constraint attention feature fusion and transformation based on the current latent variables and conditional vectors. After completing the preset number of iterative transformations, the final output feature transformation result is used as the target image corresponding to the input conditions. This invention can be applied to business scenarios such as fintech and healthcare. By performing dynamic sparse constraint attention feature fusion and transformation on latent variables and conditional vectors during iterative transformation, image generation can dynamically adjust the attention mode and constrain the attention range based on different conditions. This saves computational overhead while adapting to different generation requirements, improving image generation quality and efficiency. Attached Figure Description
[0010] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0011] Figure 1 A schematic diagram of an application environment for the conditional image generation method provided in an embodiment of the present invention; Figure 2 A flowchart of a conditional image generation method provided in an embodiment of the present invention; Figure 3 A schematic diagram of the functional modules of the conditional image generation device provided in an embodiment of the present invention; Figure 4 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments of the invention are described below in conjunction with the accompanying drawings.
[0013] The conditional image generation method provided in this embodiment of the invention can be applied to, for example... Figure 1In the application environment, it includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0014] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as financial clients, healthcare clients, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).
[0015] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0016] Server 105 can be a server providing various services, such as a backend server supporting the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. Server 105 can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"), such as high management difficulty and weak business scalability. Server 105 can also be a server for a distributed system or a server combined with blockchain.
[0017] It should be noted that the conditional image generation method provided in this application embodiment can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the conditional image generation apparatus provided in this embodiment can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Alternatively, the conditional image generation method provided in this embodiment can generally be executed by the server 105. Correspondingly, the conditional image generation apparatus provided in this embodiment can generally be disposed in the server 105.
[0018] It should be understood that the number of terminal devices, networks, and servers listed above is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be used.
[0019] like Figure 2 As shown, the conditional image generation method provided in this embodiment of the invention specifically includes the following steps: S201. Receive input conditions for guiding image generation, and encode the input conditions to obtain a corresponding condition vector.
[0020] In this embodiment, the input conditions are semantic information used to guide image generation. The specific type can be flexibly set according to the actual application scenario, such as text description, category labels, reference images, etc., thereby providing semantic guidance for the streaming model to generate the target image and ensuring that the generated image accurately matches the user's needs. Among them, the text description is the image description entered by the user in natural language (such as "a white Shiba Inu standing in the snow, with fluffy fur and gentle eyes"), the category label can be a preset image category (such as "cat", "car", "landscape"), and the reference image can be an image uploaded by the user or a pre-stored image in the database.
[0021] Upon receiving the input conditions, the input conditions are encoded, transforming unstructured or high-dimensional input conditions into low-dimensional, structured condition vectors that can be recognized and processed by the streaming model. Specifically, different types of input conditions are encoded using corresponding encoders. For example, if the input condition is a text description, a pre-trained text encoder (such as CLIP text encoder, BERT encoder, etc.) is used to encode the text description into a fixed-dimensional text condition vector; if the input condition is a category label, an embedding layer is used to encode the category label into a low-dimensional condition vector; if the input condition is a reference image feature, an image encoder is used to extract the global features of the reference image as the condition vector. Preferably, the encoded vectors can also be normalized to normalize the numerical range of the condition vectors to the [0,1] interval, avoiding problems such as gradient instability and feature fusion distortion during the training and inference of the streaming model due to excessively large or small vector values, thus ensuring the stability and consistency of the condition vectors.
[0022] For example, in a financial data visualization scenario, a user inputs the text description "a trend chart of the change in the non-performing loan ratio of a certain bank's personal loans from the first quarter to the fourth quarter of 2024, with the horizontal axis representing the quarter and the vertical axis representing the percentage of non-performing loans, using a blue gradient color scheme and smooth lines" through a terminal device. The CLIP text encoder encodes this input condition, performs word segmentation and word embedding, and then extracts semantic features from the text (such as "bank non-performing loan ratio", "quarterly trend", "blue gradient", etc.) through the multi-layer attention mechanism of the text encoder. Finally, a corresponding condition vector is output. This condition vector represents the core semantic information such as the financial data dimension and visualization requirements in the text description, providing semantic guidance for the subsequent generation of a financial trend chart that meets the requirements.
[0023] For example, in the scenario of generating materials for medical and health promotion, a health manager can input the text "A frontal diagram of a healthy adult lung, with clear markings of lung lobe regions, soft color scheme, and simple and easy-to-understand annotations." Similarly, the CLIP text encoder is called to encode the text, extract the core semantics of the text (such as "healthy lungs," "lung lobe region markings," and "soft color scheme"), and finally output a conditional vector representing the location, color scheme, and annotation requirements of the health promotion image, so as to guide the subsequent generation of lung diagrams that meet the needs of health science popularization.
[0024] S202. Randomly sample initial noise as initial latent variables, and input the initial latent variables and the condition vector into the pre-trained flow model.
[0025] In this embodiment, the initial latent variable is the initial input for image generation, specifically random noise following a preset probability distribution. Its dimension matches the input dimension of the flow model, thus providing initial features for the iterative transformation of the flow model. Through the reversible transformation of the flow model, the random noise is gradually transformed into image features with clear semantics. Preferably, sampling is performed from a standard Gaussian distribution to obtain the initial noise as the initial latent variable. That is, each element of the initial latent variable follows a standard Gaussian distribution. Using random noise with a Gaussian distribution ensures the randomness and diversity of the initial latent variable, avoiding the homogenization problem of the generated images.
[0026] The initial latent variables and the previously encoded conditional vectors are input into a pre-trained streaming model. This model is based on normalized streaming (a Glow-like model) and has been pre-trained using a large amount of training data. Its core structure includes a conditional fusion module, a dynamic sparse constraint attention module incorporating a routing network, and a low-rank projection module. The trained streaming model can precisely guide the transformation direction of the initial latent variables based on the input conditional vectors, thereby achieving iterative transformation of the initial latent variables and conditional vectors, gradually converting random noise into the target image.
[0027] For example, in different fields of guided image generation scenarios, the initial noise of the corresponding dimension can be sampled according to the scenario requirements. For example, in the financial data visualization scenario, the resolution requirement of the financial trend chart is relatively low. The initial latent variable with a dimension of 128×128×3 can be sampled from the standard Gaussian distribution. At the same time, the financial scenario condition vector is mapped to a 3-channel vector so that it is consistent with the channel dimension of the initial latent variable. Then, both are simultaneously input into the pre-trained streaming model for image generation processing.
[0028] For example, in health promotion scenarios, where health science illustrations require high resolution, initial latent variables with dimensions of 256×256×3 are sampled from a standard Gaussian distribution. Similarly, the dimensions of the medical and health scenario condition vector and the channel dimensions of the initial latent variables are adjusted to be consistent before being input into the flow model, and the iterative transformation process is started to achieve image generation.
[0029] S203. The initial latent variables and condition vectors are iteratively transformed by a preset number of time steps through the flow model. Each time step is based on the current latent variables and condition vectors to perform feature fusion and transformation of dynamic sparse constraint attention.
[0030] In this embodiment, the preset number of time steps refers to the number of iterations in the flow model. The specific value can be flexibly adjusted according to the resolution and complexity of the generated image; that is, the higher the resolution and the more complex the image, the more preset time steps are required. For example, to generate a 256×256 resolution image, the preset number of time steps is 60; to generate a 512×512 resolution image, the preset number of time steps is 100, etc. This embodiment does not impose any limitations on this. Each time step performs feature fusion and transformation based on the current latent variables and conditional vectors using dynamic sparse constraint attention, gradually optimizing the semantic expression of the latent variables to make them increasingly closer to the feature distribution of the target image.
[0031] By employing dynamic sparse constraint attention feature fusion and transformation at each time step, the computational complexity of attention can be reduced while achieving precise semantic alignment between the condition vector and the current latent variable. This avoids the problems of high complexity and inaccurate semantic fusion inherent in traditional static attention computation. Specifically, at each time step, the flow model first concatenates and fuses the latent variable and condition vector through the condition fusion module. Based on the condition vector, a portion of the expert network in the routing network is dynamically activated to perform sparse constraint attention computation on the concatenated vector. After dynamically constraining the attention range, a sparse attention map is obtained. Finally, a low-rank projection transformation is performed through the low-rank projection module to obtain the latent variable for the next time step, completing one iteration transformation. Through dynamic sparse constraint attention fusion at each time step, the focus area of attention can be dynamically adjusted according to the feature state of the current latent variable and condition vector, strengthening features related to the condition vector and weakening redundant features, ensuring the accuracy and efficiency of feature transformation.
[0032] For example, in the scenario of generating a financial trend chart, randomly sampled initial latent variables and financial scenario text condition vectors are input into a flow model for iterative transformation processing with a total of 80 time steps. In the first time step, the flow model standardizes and channels the initial latent variables, concatenates the processed latent variables with the condition vector, activates several of the most relevant expert networks (e.g., responsible for trend lines, coordinate axes, and color rendering respectively) through a routing network, calculates a sparse attention map, and then obtains the latent variables for the second time step through a low-rank projection transformation. From the second to the 80th time step, the above process is repeated. Each time step is based on the latent variables of the previous time step and a fixed condition vector, and performs feature fusion and transformation of dynamic sparse constraint attention, gradually transforming random noise into financial trend chart features containing elements such as quarterly coordinate axes, non-performing loan rate trend lines, and blue gradient color scheme.
[0033] In the scenario of generating health science popularization illustrations, the preset time steps can be adjusted to 100 to meet the image resolution requirements. The initial latent variables and the medical and health scenario text condition vector are input into the flow module. In each time step, the flow model standardizes and performs channel mixing on the initial latent variables and then concatenates them with the condition vector. The routing network selects and activates the corresponding expert network to achieve attention focus. For example, it focuses on the expert networks responsible for lung contour, lung lobe partition labeling, color matching adjustment, and text label clarity, respectively. After obtaining the sparse attention map, the latent variables for the next time step are obtained through low-rank projection transformation. The feature fusion and transformation of dynamic sparse constraint attention are performed iteratively step by step. After each time step, the illustration details of the latent variables are improved until the iterative transformation of all time steps is completed.
[0034] S204. After completing the iterative transformation of a preset number of time steps, the final output feature transformation result is used as the target image corresponding to the input conditions.
[0035] In this embodiment, after completing the iterative transformation of a preset number of time steps, the feature transformation result output by the flow model is a feature tensor with the same resolution as the target image. This feature tensor has completed the transformation from random noise to target image features, contains all the semantic information corresponding to the input conditions, and has clear image contours, details, and color distribution. Therefore, the final output feature transformation result is used as the target image corresponding to the input conditions. The generated target image can be stored in the server's database and simultaneously fed back to the terminal device for users to view, save, or further edit.
[0036] For example, in the scenario of generating a financial trend chart, after completing the iterative transformation of the total time step, the feature transformation result output by the flow model is a 128×128×3 RGB feature tensor. The numerical range of this tensor is mapped to [0,255], and after rounding, a color financial trend chart with a resolution of 128×128 is obtained. This financial trend chart clearly presents the changing trend of the non-performing loan ratio of a certain bank's personal loans in the four quarters of 2024. The coordinate axes, lines, and color scheme are highly matched with the input text description.
[0037] For example, in the scenario of generating a medical and health promotion poster, after completing the iterative transformation of a preset number of time steps, the feature transformation result output by the flow model is an RGB feature tensor of 256×256×3. After mapping and rounding, a healthy lung diagram with a resolution of 256×256 is obtained. This diagram shows the outline of a healthy adult lung and the labeling of lung lobe areas. The color scheme is soft and the labeling is simple to suit the usage scenario of health science popularization posters.
[0038] In the above embodiments, this invention discloses a conditional image generation method. The method receives input conditions to guide image generation and encodes these conditions to obtain a corresponding conditional vector. Initial noise is randomly sampled as an initial latent variable, and the initial latent variable and the conditional vector are input into a pre-trained streaming model. The streaming model iteratively transforms the initial latent variable and the conditional vector for a preset number of time steps. At each time step, dynamic sparse constraint attention feature fusion and transformation are performed based on the current latent variable and the conditional vector. After completing the preset number of iterative transformations, the final output feature transformation result is used as the target image corresponding to the input conditions. This invention can be applied to business scenarios such as fintech and healthcare. By performing dynamic sparse constraint attention feature fusion and transformation on the latent variable and conditional vector during iterative transformation, image generation can dynamically adjust the attention mode and constrain the attention range based on different conditions. This saves computational overhead while adapting to different generation requirements, improving image generation quality and efficiency.
[0039] In one embodiment, step S203 includes: Obtain the latent variables at the current time step, and perform preprocessing on the latent variables at the current time step to obtain the preprocessed latent variables; The preprocessed latent variables and the condition vector at the current time step are concatenated to obtain the concatenated vector at the current time step. Based on the concatenation vector at the current time step, a portion of the expert network of the routing network in the flow model is dynamically activated, and sparse constraint attention is calculated on the concatenation vector at the current time step to obtain a sparse attention map. Based on the sparse attention map and the spliced vector, a low-rank projection transformation is performed to obtain the latent variables for the next time step. Continue to perform feature fusion and transformation of the latent variables and the condition vector at the next time step using dynamic sparse constraint attention, and so on, until the iterative transformation of the preset number of time steps is completed.
[0040] In this embodiment, during the iterative feature fusion and transformation processing of dynamic sparse constrained attention, the latent variables at the current time step are first obtained. If the current time step is the first time step, the current latent variables are the initial latent variables randomly sampled in the first step; if the current time step is the second or subsequent time steps, the current latent variables are the feature transformation results of the previous time step. Preprocessing is then performed on the latent variables at the current time step. The specific preprocessing process can be set according to actual needs, typically including feature standardization and channel mixing. In some scenarios, noise removal processing can be further added to ensure that the preprocessed latent variables have good feature distribution and semantic expressive ability. Preprocessing standardizes and optimizes the features of the current latent variables, eliminating redundant information and noise in the latent variables, ensuring the accuracy of subsequent feature fusion and transformation, and avoiding problems such as unstable model training and distorted generated images caused by uneven distribution of latent variables.
[0041] The preprocessed latent variables and conditional vectors at the current time step are then concatenated to fuse them, forming a concatenated vector containing image feature carriers and semantic guidance information. This provides a unified input for subsequent sparse constraint attention calculations and expert network activation. Based on this concatenated vector, the routing network in the flow model is dynamically activated, activating a subset of expert networks. Each expert network focuses on processing a specific type of feature (such as contour features, color features, or detail features). Dynamically activating these expert networks allows the selection of the most suitable network for processing the current feature based on the feature state of the concatenated vector. Unselected expert networks remain dormant and do not participate in feature processing at the current time step, avoiding the computational waste caused by simultaneous activation of all expert networks and improving the model's computational efficiency.
[0042] Based on the routing network that has activated some expert networks, sparse constraint attention calculation is performed on the concatenated vector at the current time step. Specifically, a sparse mask is generated to constrain the global attention score. Only the attention weights related to the condition vector are retained to obtain the attention-constrained sparse attention map. This sparse attention map retains only the high-weight attention information related to the condition vector, thereby removing redundant attention weights, reducing the complexity of attention calculation, and ensuring the accuracy of feature fusion.
[0043] Finally, a low-rank projection transformation is performed based on the sparse attention map and the concatenated vector to compress the value vector of the concatenated vector, further reducing computational complexity. Simultaneously, the core feature information in the concatenated vector is preserved to obtain the latent variables for the next time step, completing the iterative transformation for the current time step. The latent variables for the next time step are used as the new latent variables for the current time step, and the above process is repeated. At each time step, based on the current latent variables and a fixed condition vector, preprocessing, concatenation, sparse constraint attention calculation, and low-rank projection transformation are performed to progressively optimize the feature representation of the latent variables. Iteration stops after a predetermined number of iterative transformations, ultimately obtaining the target image.
[0044] It is understandable that the expert network activated by the routing network at each time step during the iterative transformation process may be different because the latent variable feature states are different at each time step. The routing network will dynamically adjust the activated expert network according to the feature changes of the latent variables to ensure that the feature processing at each time step can match the current feature state and improve the accuracy of feature fusion and transformation.
[0045] In one embodiment, obtaining the latent variables at the current time step and performing preprocessing on the latent variables at the current time step to obtain preprocessed latent variables includes: Obtain the initial latent variable or the feature transformation result of the previous time step, and use it as the latent variable of the current time step; The latent variables at the current time step are subjected to feature standardization to obtain normalized intermediate variables; The normalized intermediate variables are subjected to channel mixing to obtain preprocessed latent variables.
[0046] In this embodiment, during the preprocessing of latent variables, the initial latent variables or the feature transformation result of the previous time step are obtained as the latent variables of the current time step. Specifically, when the current time step is the first time step, since no iterative transformation has been performed, the latent variables of the current time step are the initial latent variables obtained by random sampling in the first step; when the current time step is the second or subsequent time steps, the latent variables of the current time step are the feature transformation result of the previous time step, that is, the latent variables of the next time step obtained by the low-rank projection transformation of the previous time step. These latent variables have undergone feature optimization of the previous time step, contain some semantic information, and are the basis for the iterative transformation of the current time step.
[0047] Next, feature standardization is performed on the latent variables at the current time step to eliminate differences in their feature distribution. The numerical range of the latent variables is normalized to a preset interval to obtain normalized intermediate variables, making subsequent feature fusion and transformation more stable. Then, channel mixing is performed on the normalized intermediate variables, specifically through a 1×1 convolutional layer. A 1×1 convolutional layer can transform and fuse channel dimensions without changing the dimension of the feature map space. By inputting the normalized intermediate variables into the 1×1 convolutional layer for convolution, a channel-mixed feature tensor is obtained. This feature tensor is the preprocessed latent variable, thus fusing feature information from different channels, improving the feature representation ability of the latent variables, and avoiding the problem of inaccurate semantic fusion caused by channel feature separation.
[0048] In one embodiment, the step of dynamically activating a portion of the expert network of the routing network in the flow model based on the concatenated vector at the current time step, and performing sparse constrained attention calculation on the concatenated vector at the current time step to obtain a sparse attention map, includes: Dynamically activate a portion of the expert network of the routing network in the flow model based on the splicing vector at the current time step; Global attention is calculated between all feature pairs on the concatenated vector at the current time step to obtain the attention score matrix; Based on the currently dynamically activated routing network, a corresponding sparse mask is generated for the spliced vector; The attention score matrix and the sparse mask are multiplied element-wise to obtain a sparse attention graph with attention constraints.
[0049] In this embodiment, when performing sparse constrained attention calculation, a portion of the expert networks in the routing network of the flow model is dynamically activated based on the concatenation vector at the current time step. Specifically, the routing network adopts a hierarchical sparse MoE architecture, in which E expert networks are predefined { Each expert corresponds to a specific feature subspace, focusing on processing a particular type of feature (such as contour features, color features, and texture features). The routing weight is calculated as follows:
[0050] Where t is the current time step. For potential representation, Let [·,·] be the condition vector, and let [·,·] denote the concatenation operation. The learnable weight matrix for the routing network is obtained by inputting the concatenated vector at the current time step into the routing network and mapping the concatenated vector to E expert scores. Then, the activation score vector is subjected to Softmax normalization to obtain the activation probability of each expert network. Finally, the top-M experts are selected through sparse gating. That is, the expert networks with the highest activation probabilities (M) are selected for activation and participate in the subsequent sparse constraint attention calculation. The expert networks that are not selected do not participate in the feature processing of the current time step and are in a dormant state.
[0051] Global attention is calculated between feature pairs in the concatenated vector at the current time step. This global attention calculation uses a self-attention mechanism. Specifically, the concatenated vector at the current time step... The inputs are fed into three independent linear layers to obtain the query vector. Key vector Sum value vector ,Right now:
[0052] The attention score matrix is calculated using a scaled dot product:
[0053] in, ∈ d×dk , ∈ m×dk and These are learnable parameters, [·,·] represent concatenation operations, and d k It is the dimension of the key / query feature vector.
[0054] The attention score matrix is obtained by calculating the correlation between all feature pairs in the concatenated vector. This attention score matrix can reflect the correlation between each feature and all other features, providing a basis for subsequent sparsity constraints.
[0055] Based on the attention score, a sparse mask is generated for the concatenated vector according to the currently dynamically activated routing network. This sparse mask is used to constrain the attention score matrix, filter out redundant attention weights, and retain only the high-weight attention information related to the condition vector, thereby achieving sparsity of attention and reducing computational complexity.
[0056] Specifically, after activating a portion of the expert network, the routing network generates corresponding sparse activation values based on the feature processing preferences of the activated expert networks and the global features of the concatenated vectors.
[0057] in and These are routing network parameters, and 'e' is the hidden layer dimension.
[0058] For this sparse activation value Top-K filtering is performed to retain the K elements with the largest sparse activation values, where K is a preset sparsity threshold. The K largest elements are set to 1, and the remaining elements are set to 0, resulting in the final sparse mask. ,Right now The specific value of K can be flexibly adjusted according to computational efficiency and generation quality. The larger K is, the denser the attention, the higher the computational complexity, and the richer the details of the generated image; the smaller K is, the sparser the attention, the lower the computational complexity, and the simpler the details of the generated image may be.
[0059] The attention score matrix and the sparse mask are then multiplied element-wise to obtain the attention-constrained sparse attention graph:
[0060] Here, ⊙ represents element-wise multiplication. By multiplying elements-wise, redundant attention weights in the attention score matrix are filtered out. That is, after element-wise multiplication, the sparse attention map only retains the attention weights marked as 1 in the sparse mask. In other words, only the high-weight attention information related to the condition vector is retained, thus realizing the sparse constraint of attention. By obtaining the sparse attention map, the feature association information related to the condition vector can be accurately reflected, while significantly reducing the computational complexity and improving the model's computational efficiency.
[0061] In one embodiment, the step of performing a low-rank projection transformation based on the sparse attention map and the concatenated vector to obtain the latent variables for the next time step includes: The low-rank value matrix is obtained by performing a low-rank decomposition projection on the value vector of the concatenated vector. The sparse attention map is weighted and fused with the low-rank matrix to obtain the corresponding attention features; The spliced vector is subjected to multilayer perceptron transformation to obtain nonlinear compensation features; The attention features and nonlinear compensation features are summed to obtain the potential variables for the next time step.
[0062] In this embodiment, since the dimension of the value vector of the concatenated vector is usually high, directly performing weighted fusion with the sparse attention map would result in a large amount of computation. Therefore, the dimension of the concatenated vector is compressed by low-rank decomposition projection, which reduces the computational complexity while preserving the core feature information, and at the same time ensures the reversibility of the flow model.
[0063] Specifically, the value vector of the concatenated vector is first subjected to low-rank decomposition and projection to obtain a low-rank value matrix:
[0064] in ∈ r×d This is the value vector of the concatenated vectors. ∈ d×r It is a left orthogonal matrix. ∈ r×r为 diagonal matrix Let r be a right orthogonal matrix. d is the low-rank dimension. The value of the low-rank dimension r can be flexibly adjusted according to the quality of the generated image and the computational efficiency. The larger r is, the more complete the feature information is retained and the higher the quality of the generated image, but the higher the computational complexity. Conversely, the lower the computational complexity, the more likely the generated image will be distorted.
[0065] The sparse attention map is then weighted and fused with the low-rank matrix to obtain the corresponding attention features. By using weighted fusion, the attention weight information in the sparse attention map is fused with the feature information in the low-rank matrix to obtain the attention feature. This feature contains both the attention weights related to the conditional vector and the core features of the concatenated vector, which can accurately reflect the image features related to the conditional vector.
[0066] Furthermore, the spliced vectors are processed using a multilayer perceptron transform. The spliced vector is transformed by multilayer perceptron transformation to compensate for the nonlinear feature information lost in the low-rank decomposition projection process. This complements the attention feature, enhances the expressive power of the feature, and ensures that the latent variables in the next time step have rich feature details.
[0067] Finally, the attention features and nonlinear compensation features are summed to obtain the latent variables for the next time step:
[0068] The latent variables obtained at this time can combine the advantages of attention features and nonlinear compensation features to obtain latent variables for the next time step that have attention weight information, core feature information and nonlinear feature information, ensuring that the iterative transformation of the next time step can be based on more comprehensive and accurate features.
[0069] In one embodiment, the flow model is obtained through gradient-stabilized training via the following steps: Collect training images and corresponding training input conditions, and encode the training input conditions into a training condition vector; An initial flow model is constructed, and the training condition vector and the acquired image are subjected to dynamic sparse constraint attention feature fusion and transformation through the initial flow model to obtain the corresponding latent variables; The corresponding loss value is calculated based on the latent variables and the pre-constructed loss function; The initial flow model is updated with parameters using a gradient-stabilized sparse backpropagation mechanism combined with the loss value until the preset convergence condition is met, at which point the trained flow model is obtained.
[0070] In this embodiment, the flow model is pre-trained with gradient stabilization using a large number of training samples. First, training images and corresponding training input conditions are collected. The training images can include various types of images such as animals, plants, landscapes, and people. The training input conditions correspond one-to-one with the training images, and their types are consistent with the input conditions in the inference process, such as text descriptions, category labels, and reference images. At the same time, the training input conditions are encoded into training condition vectors. The specific encoding process is consistent with the encoding process of the input conditions in the inference process.
[0071] The initial flow model, built upon normalized flow, comprises a core structure including a conditional fusion module, a dynamic sparse constraint attention module incorporating a routing network, and a low-rank projection module. At this stage, the initial model's parameters are randomly initialized, lacking feature fusion and transformation capabilities. By inputting the training image and its corresponding training conditional vector into the initial flow model, the model performs a forward iterative transformation on the training image (the opposite of the inverse iterative transformation during inference). In the inference process, the target image is obtained through an inverse transformation starting from the initial noise; during training, the corresponding noise distribution, i.e., latent variables, is obtained through a forward iterative transformation starting from the training image. This forward iterative transformation process is the inverse of the inverse iterative transformation of the inference process, ensuring the reversibility of the flow model. Similarly, during training, feature fusion and transformation are performed at each time step using dynamic sparse constraint attention, gradually transforming the training image into latent variables. The distribution of these latent variables is related to the feature distribution of the training image. The training objective of the flow model is to make the distribution of the latent variables as close as possible to a standard Gaussian distribution, so that the trained model can transform the sampled noise conforming to the standard Gaussian distribution into a training image corresponding to the training input conditions through an inverse transformation.
[0072] The corresponding loss value is calculated based on the latent variables and the pre-built loss function. The loss function is the basis for updating the model parameters. Specifically, the pre-built loss function is a four-term joint loss function, including the flow model negative log-likelihood loss (L_flow), sparse regularization loss (L_sparse), coverage regularization loss (L_cover), and expert balance loss (L_balance). This allows for a comprehensive evaluation of the flow model's parameter performance. It improves the reliability of expert activation and sparse masking while assessing the difference between the latent variable distribution and the standard Gaussian distribution, as well as the difference between the model-generated image and the training image. This guides the gradual optimization of model parameters and improves the generation quality and stability of the model.
[0073] Specifically, the calculation methods for each loss function are as follows: Negative log-likelihood loss (L_flow):
[0074] in As a latent variable, p( L_flow is the probability density function of the standard Gaussian distribution. It is used to measure the difference between the distribution of the latent variables and the standard Gaussian distribution. The smaller L_flow is, the closer the distribution of the latent variables is to the standard Gaussian distribution, and the better the invertibility of the model.
[0075] Sparse regularized loss (L_sparse):
[0076] L_sparse is calculated using L1 regularization, where, Let t be the weight and t be the time step. L_sparse is the sparse activation value used to constrain the sparsity of the sparse attention map, avoiding an increase in computational complexity caused by an overly dense attention map. The smaller L_sparse is, the higher the sparsity of the sparse mask and the lower the computational complexity.
[0077] Coverage regularization loss (L_cover):
[0078] As weight, Let be the sparse mask matrix at time step t. The value of topK is used to prevent the sparse attention map from missing key feature regions related to the conditional vector. The smaller L_cover is, the more comprehensive the key feature regions covered by the sparse mask are.
[0079] Expert-balanced loss (L_balance):
[0080] in As weight, Let L_balance be the routing weight of the e-th hidden layer dimension at time step t, where T is the total number of time steps and E is the number of experts. L_balance is used to avoid the expert networks being activated by the routing network being too concentrated, ensuring that all expert networks can be fully utilized. The smaller L_balance is, the more balanced the activation probability of the expert networks is, thereby avoiding parameter degradation caused by some expert networks being dormant for a long time.
[0081] The final total loss is By calculating the total loss value, the current training effect of the model is comprehensively evaluated, providing a basis for subsequent parameter updates.
[0082] When adjusting model parameters, the discreteness (0 or 1) of the mask in the traditional backpropagation mechanism can cause gradient calculation jumps, leading to gradient vanishing or exploding problems, which affect the convergence stability of the model. Therefore, this embodiment uses a gradient-stabilized sparse backpropagation mechanism combined with the loss value to update the parameters of the initial flow model, thereby smoothing gradient changes and ensuring gradient stability. Specifically, the gradients of each parameter of the model are calculated using the gradient-stabilized sparse backpropagation mechanism based on the total loss value L_total. Based on the calculated gradients, optimization algorithms such as stochastic gradient descent and Adam are used to update the parameters of the initial flow model (including the parameters of the routing network, expert network, linear layer, attention layer, etc.). After the update is completed, the next round of training begins, and the above process is repeated until the preset convergence condition is met. At this point, the training stops, the current model parameters are saved, and the trained flow model is obtained. The trained flow model has stable feature fusion and transformation capabilities and can generate high-quality target images based on the input conditions.
[0083] In one embodiment, updating the parameters of the initial flow model using the gradient-stabilized sparse backpropagation mechanism combined with the loss value includes: The true gradient and the approximate gradient are calculated based on the loss value, and the approximate gradient is obtained by ignoring the discreteness of the sparse mask during the transformation process. Obtain the adaptive adjustment coefficient for the current training round, and then perform gradient stability correction based on the adaptive adjustment coefficient, the true gradient, and the approximate gradient to obtain the final surrogate gradient. The parameters of the initial flow model are updated based on the surrogate gradient, and the adaptive adjustment coefficients are updated before entering the next training round.
[0084] In this embodiment, when updating the parameters of the convection model, the true gradient is first calculated based on the loss value. and approximate gradient Specifically, the true gradient is obtained by differentiating the model parameters based on the total loss value L_total. During the differentiation process, the discreteness of the sparse mask is strictly preserved, that is, the value of the sparse mask is always 0 or 1. The approximate gradient is also based on differentiating the model parameters based on the total loss value L_total, but the discreteness of the sparse mask is ignored during the differentiation process. It is regarded as a continuous variable with a value range between [0,1]. Specifically, a continuously differentiable function (such as the sigmoid function) can be used to approximate the sparse mask to make the mask continuously differentiable.
[0085] Because the true gradient takes into account the discreteness of the sparse mask, it can accurately reflect the influence of the sparse mask on the gradient of the model parameters. However, since the sparse mask is a binary discrete variable, directly calculating the true gradient will result in gradient jumps and discontinuities, easily leading to gradient vanishing or exploding, thus affecting the model's convergence stability. The approximate gradient, on the other hand, ignores the discreteness of the sparse mask and treats it as a continuous variable for gradient calculation, resulting in smooth gradient values and avoiding gradient jumps. However, it will have some gradient bias and cannot completely and accurately reflect the characteristics of the true gradient. Therefore, it is necessary to combine and modify the characteristics of the true gradient and the approximate gradient to obtain a modified gradient that balances stability and accuracy.
[0086] Specifically, the adaptive adjustment coefficients for the current training epoch are first obtained. These coefficients are the initial adjustment coefficients during the first training epoch and are dynamically updated as training progresses in subsequent epochs. Based on these adaptive adjustment coefficients, the true gradient, and the approximate gradient, gradient stability correction is performed to obtain the final surrogate gradient.
[0087] Adaptive adjustment coefficient Adjust according to the following strategies:
[0088] in, η is the initial adjustment coefficient, η is the decay factor, and τ is the adjustment period.
[0089] The surrogate gradient obtained through gradient stability correction not only retains the accurate capture of the discreteness of the sparse mask by the real gradient, but also incorporates the smoothness of the approximate gradient, providing a reliable basis for the stable update of model parameters and effectively solving the problem of gradient instability.
[0090] The parameters of the initial flow model are updated based on the corrected surrogate gradient. After updating the adaptive adjustment coefficients according to the above strategy, the next training round is entered. The process of calculating the true gradient and approximate gradient, correcting the gradient, updating the parameters and adaptive adjustment coefficients is repeated. Each training round is based on the new model parameters to calculate a more accurate and stable gradient, gradually optimize the model parameters to reduce the total loss value, and finally obtain the trained flow model, thereby improving the stability and reliability of the flow model training.
[0091] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0092] Further reference Figure 3 As a response to the above Figure 2 The present invention provides an embodiment of a conditional image generation apparatus, which is similar to the method shown. Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0093] like Figure 3 As shown, the conditional image generation device 30 described in this embodiment includes: The receiving module 301 is used to receive input conditions for guiding image generation, and to encode the input conditions to obtain a corresponding condition vector; The sampling input module 302 is used to randomly sample initial noise as initial latent variables, and input the initial latent variables and the condition vector into the pre-trained stream model; The feature transformation module 303 is used to perform a preset number of time steps of iterative transformation on the initial latent variables and condition vectors through the flow model. Each time step is based on the current latent variables and condition vectors to perform feature fusion and transformation with dynamic sparse constraint attention. The image output module 304 is used to output the final feature transformation result as the target image corresponding to the input conditions after completing the iterative transformation of a preset number of time steps.
[0094] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the conditional image generation and execution process. For specific implementation methods of each module, please refer to the corresponding method embodiments above, which will not be repeated here.
[0095] In one embodiment, the feature transformation module 303 includes: The preprocessing unit is used to obtain the latent variables of the current time step, perform preprocessing on the latent variables of the current time step, and obtain the preprocessed latent variables. The splicing unit is used to splice the preprocessed latent variables and the condition vector at the current time step to obtain the spliced vector at the current time step. The sparse attention unit is used to dynamically activate a portion of the expert network of the routing network in the flow model based on the splicing vector at the current time step, and to perform sparse constraint attention calculation on the splicing vector at the current time step to obtain a sparse attention graph. The low-rank projection unit is used to perform a low-rank projection transformation based on the sparse attention map and the splicing vector to obtain the latent variables for the next time step. The iterative control unit is used to continue to perform feature fusion and transformation of the latent variables and the condition vector of the next time step with dynamic sparse constraint attention, and so on, until the iterative transformation of a preset number of time steps is completed.
[0096] In one embodiment, the preprocessing unit includes: The acquisition unit is used to acquire the initial latent variable or the feature transformation result of the previous time step, and use it as the latent variable of the current time step. The feature standardization unit is used to perform feature standardization on the latent variables of the current time step to obtain normalized intermediate variables. The channel mixing unit is used to perform channel mixing processing on the normalized intermediate variables to obtain preprocessed potential variables.
[0097] In one embodiment, the sparse attention unit includes: The dynamic activation unit is used to dynamically activate a portion of the expert network of the routing network in the flow model based on the splicing vector at the current time step. The attention score calculation unit is used to perform global attention calculation on all feature pairs of the concatenated vector at the current time step to obtain the attention score matrix; A sparse mask generation unit is used to generate a corresponding sparse mask for the concatenated vector based on the currently dynamically activated routing network. The sparse attention constraint unit is used to perform element-wise multiplication based on the attention score matrix and the sparse mask to obtain a sparse attention graph with attention constraints.
[0098] In one embodiment, the low-rank projection unit includes: A low-rank decomposition unit is used to perform low-rank decomposition projection on the value vector of the concatenated vector to obtain a low-rank value matrix. The weighted fusion unit is used to weightedly fuse the sparse attention map with the low-rank matrix to obtain the corresponding attention features; The compensation transformation unit is used to perform multilayer perceptron transformation processing on the spliced vector to obtain nonlinear compensation features; The feature summation unit is used to sum the attention features and the nonlinear compensation features to obtain the potential variables for the next time step.
[0099] In one embodiment, the device 30 further includes: The training data acquisition module is used to acquire training images and training input conditions corresponding to the training images, and to encode the training input conditions into a training condition vector. The construction and training module is used to construct an initial flow model, and to perform feature fusion and transformation of the training condition vector and the acquired image through dynamic sparse constraint attention using the initial flow model to obtain the corresponding latent variables. The loss calculation module is used to calculate the corresponding loss value based on the latent variables and the pre-built loss function; The parameter update module is used to update the parameters of the initial flow model using the gradient-stabilized sparse backpropagation mechanism combined with the loss value until the preset convergence condition is met to obtain the trained flow model.
[0100] In one embodiment, the parameter update module includes: The gradient calculation unit is used to calculate the true gradient and the approximate gradient based on the loss value, wherein the approximate gradient is obtained by ignoring the discreteness of the sparse mask during the transformation process. The gradient correction unit is used to obtain the adaptive adjustment coefficient of the current training round, and to obtain the final surrogate gradient after performing gradient stability correction based on the adaptive adjustment coefficient, the true gradient and the approximate gradient. The parameter update unit is used to update the parameters of the initial flow model according to the surrogate gradient, and then update the adaptive adjustment coefficients before entering the next training round.
[0101] In the above embodiments, the present invention discloses a conditional image generation device. By performing feature fusion and transformation of latent variables and conditional vectors with dynamic sparse constraint attention during iterative transformation, image generation can dynamically adjust the attention mode and constrain the attention range based on different conditions. This saves computational overhead while adapting to the generation needs of different conditions, and improves the quality and efficiency of image generation.
[0102] Specific limitations regarding the conditional image generation device can be found in the limitations of the conditional image generation method described above, and will not be repeated here. Each module in the aforementioned conditional image generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0103] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0104] Another embodiment of the present invention provides a computer device, such as... Figure 4 As shown, the computer device 40 includes: One or more processors 401 and memory 402, Figure 4 The following section uses a processor 401 as an example. The processor 401 and the memory 402 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.
[0105] The processor 401 is used to perform various control logics of the computer device 40. It can be any conventional processor, microprocessor, state machine, general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0106] The memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the conditional image generation method in the embodiments of the present invention. The processor 401 executes various functional applications and data processing of the computer device 40 by running the non-volatile software programs, instructions, and units stored in the memory 402, thereby implementing the conditional image generation method in the above method embodiments.
[0107] Another embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, perform the steps of the conditional image generation method in any of the above method embodiments.
[0108] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0109] Based on the above description of the embodiments, those skilled in the art will understand that the methods described in the embodiments can be implemented using software plus necessary general-purpose hardware platforms. Of course, they can also be implemented using hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0110] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The computer program can be stored in a non-volatile, computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a memory, magnetic disk, floppy disk, flash memory, optical storage, etc.
[0111] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0112] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A conditional image generation method, characterized in that, include: Receive input conditions for guiding image generation, and encode the input conditions to obtain a corresponding condition vector; Randomly sample initial noise as initial latent variables, and input the initial latent variables and the conditional vector into the pre-trained stream model; The initial latent variables and condition vectors are iteratively transformed by the flow model for a preset number of time steps. Each time step performs feature fusion and transformation based on the current latent variables and condition vectors using dynamic sparse constraint attention. After completing the iterative transformation for a preset number of time steps, the final output feature transformation result is used as the target image corresponding to the input conditions.
2. The conditional image generation method according to claim 1, characterized in that, The process involves iteratively transforming the initial latent variables and condition vectors using the flow model for a predetermined number of time steps. Each time step performs feature fusion and transformation based on the current latent variables and condition vectors using dynamic sparse constraint attention, including: Obtain the latent variables at the current time step, and perform preprocessing on the latent variables at the current time step to obtain the preprocessed latent variables; The preprocessed latent variables and the condition vector at the current time step are concatenated to obtain the concatenated vector at the current time step. Based on the concatenation vector at the current time step, a portion of the expert network of the routing network in the flow model is dynamically activated, and sparse constraint attention is calculated on the concatenation vector at the current time step to obtain a sparse attention map. Based on the sparse attention map and the spliced vector, a low-rank projection transformation is performed to obtain the latent variables for the next time step. Continue to perform feature fusion and transformation of the latent variables and the condition vector at the next time step using dynamic sparse constraint attention, and so on, until the iterative transformation of the preset number of time steps is completed.
3. The conditional image generation method according to claim 2, characterized in that, The process of obtaining the latent variables at the current time step and performing preprocessing on the latent variables at the current time step to obtain preprocessed latent variables includes: Obtain the initial latent variable or the feature transformation result of the previous time step, and use it as the latent variable of the current time step; The latent variables at the current time step are subjected to feature standardization to obtain normalized intermediate variables; The normalized intermediate variables are subjected to channel mixing to obtain preprocessed latent variables.
4. The conditional image generation method according to claim 2, characterized in that, The step of dynamically activating a portion of the expert network in the routing network of the flow model based on the concatenated vector at the current time step, and performing sparse constrained attention calculation on the concatenated vector at the current time step to obtain a sparse attention map includes: Dynamically activate a portion of the expert network of the routing network in the flow model based on the splicing vector at the current time step; Global attention is calculated between all feature pairs on the concatenated vector at the current time step to obtain the attention score matrix; Based on the currently dynamically activated routing network, a corresponding sparse mask is generated for the spliced vector; The attention score matrix and the sparse mask are multiplied element-wise to obtain a sparse attention graph with attention constraints.
5. The conditional image generation method according to claim 2, characterized in that, The step of performing a low-rank projection transformation based on the sparse attention map and the concatenated vector to obtain the latent variables for the next time step includes: The low-rank value matrix is obtained by performing a low-rank decomposition projection on the value vector of the concatenated vector. The sparse attention map is weighted and fused with the low-rank matrix to obtain the corresponding attention features; The spliced vector is subjected to multilayer perceptron transformation to obtain nonlinear compensation features; The attention features and nonlinear compensation features are summed to obtain the potential variables for the next time step.
6. The conditional image generation method according to claim 1, characterized in that, The flow model is obtained through gradient stabilization training via the following steps: Collect training images and corresponding training input conditions, and encode the training input conditions into a training condition vector; An initial flow model is constructed, and the training condition vector and the acquired image are subjected to dynamic sparse constraint attention feature fusion and transformation through the initial flow model to obtain the corresponding latent variables; The corresponding loss value is calculated based on the latent variables and the pre-constructed loss function; The initial flow model is updated with parameters using a gradient-stabilized sparse backpropagation mechanism combined with the loss value until the preset convergence condition is met, at which point the trained flow model is obtained.
7. The conditional image generation method according to claim 6, characterized in that, The step of updating the parameters of the initial flow model using the gradient-stabilized sparse backpropagation mechanism combined with the loss value includes: The true gradient and the approximate gradient are calculated based on the loss value, and the approximate gradient is obtained by ignoring the discreteness of the sparse mask during the transformation process. Obtain the adaptive adjustment coefficient for the current training round, and then perform gradient stability correction based on the adaptive adjustment coefficient, the true gradient, and the approximate gradient to obtain the final surrogate gradient. The parameters of the initial flow model are updated based on the surrogate gradient, and the adaptive adjustment coefficients are updated before proceeding to the next training round.
8. A conditional image generation apparatus, characterized in that, include: The receiving module is used to receive input conditions for guiding image generation, and to encode the input conditions to obtain a corresponding condition vector; The sampling input module is used to randomly sample initial noise as initial latent variables, and input the initial latent variables and the condition vector into the pre-trained stream model; The feature transformation module is used to perform a preset number of time steps of iterative transformation on the initial latent variables and condition vectors through the flow model. Each time step is based on the current latent variables and condition vectors to perform feature fusion and transformation with dynamic sparse constraint attention. The image output module is used to output the final feature transformation result as the target image corresponding to the input conditions after completing the iterative transformation of a preset number of time steps.
9. A computer device, characterized in that, Includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the conditional image generation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the conditional image generation method according to any one of claims 1-7.