A food image generation method based on a two-stream diffusion control model

By introducing a dual-flow diffusion control model and multiple mechanisms into the food image generation technology, the problem of global structure and local details is solved, and high-quality and delicate food image generation is achieved.

CN119723567BActive Publication Date: 2025-05-27SHANDONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510247700.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-27
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

Existing food image generation technology is difficult to balance the global structure and local details, resulting in the generated food image lacking a sense of reality and delicateness.

Method used

The food image generation method based on the dual-current diffusion control model is adopted, and the two-way information flow between the generation network and the control network is enhanced by introducing the dual-current interaction mechanism, gate mechanism, adaptive scaling mechanism and attention mechanism, and the two-way information flow between the generation network and the control network is realized, and feature transmission and fusion are achieved.

Benefits of technology

It significantly improves the quality and detail capture capabilities of food images, ensuring that the generated food images have a good global structure and can present local details in detail, enhancing the realism and visual expression of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723567B_ABST
    Figure CN119723567B_ABST
Patent Text Reader

Abstract

The present invention discloses a food image generation method based on a dual-stream diffusion control model, comprising the following steps: constructing a food image generation architecture based on a diffusion model; introducing an adaptive scaling mechanism and a global and local attention module to optimize the semantic information and detailed texture in the food image; a fusion network based on a gating mechanism to dynamically adjust the fusion ratio of the control signal and the generation signal and adaptively optimize the generation process; based on a dual-stream diffusion generation strategy, the generation model gradually denoises the noise image through the diffusion model to generate high-quality images. By adopting the above-mentioned food image generation method based on a dual-stream diffusion control model, the present invention significantly improves the quality and detail capture ability of the generated food images by introducing a dual-stream interaction mechanism, a gating fusion mechanism and an adaptive scaling mechanism; and effectively solves the problem of balancing the global structure and local details in food image generation by increasing the bidirectional information flow between the generation network and the control network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of food image generation, and in particular to a food image generation method based on a two-stream diffusion control model. Background Art

[0002] The technological development in the field of food image generation mainly relies on generative adversarial networks (GANs) and diffusion models. GAN-based models (such as CookGAN) perform well in the overall structure of the generated images, but have significant defects in the presentation of details (such as the surface texture and gloss reflection of food), especially in the generation of complex food images, where the lack of details is particularly prominent. In addition, although diffusion models can better restore the global structure when generating high-quality food images, they are still insufficient in capturing local details.

[0003] For example, when existing diffusion models process food images with complex details, they cannot accurately capture local details such as texture and gloss, resulting in the generated food images lacking a sense of reality and fineness. The guiding effect of control conditions (such as edge maps, depth maps, etc.) on the generated images is weak, and there is insufficient fine control of details during the generation process. Therefore, existing methods fail to fully utilize the differential features of multi-modal conditions, resulting in the inability to flexibly adjust the contributions between different conditions during the image generation process, thus limiting the generation quality.

[0004] Therefore, based on the above technical defects, the present invention proposes a food image generation method based on a two-stream diffusion control model to effectively solve the problem of balancing the global structure and local details in food image generation. Summary of the Invention

[0005] The object of the present invention is to provide a food image generation method based on a two-stream diffusion control model, which significantly improves the quality and detail capture ability of the generated food images by introducing a two-stream interaction mechanism, a gating mechanism, an adaptive scaling mechanism, and an attention mechanism; by adding bidirectional information flow between the generation network and the control network on the basis of the diffusion model and realizing feature transfer through zero convolutional layers, effectively solving the problem of balancing the global structure and local details in food image generation.

[0006] To achieve the above object, the present invention provides a food image generation method based on a two-stream diffusion control model, including the following steps:

[0007] Step S1, construct a food image generation architecture based on a two-stream diffusion control model;

[0008] Step S2, introduce an adaptive scaling mechanism and a global and local attention module to optimize the semantic information and detailed texture in the food image;

[0009] Step S3: Based on the fusion network with a gating mechanism, dynamically adjust the fusion ratio of the control signal and the generated signal, and adaptively optimize the generation process according to the characteristics of the food image and the control conditions;

[0010] Step S4: Based on the two-stream diffusion generation strategy, through the generation model of the food image generation architecture based on the two-stream diffusion control model, gradually denoise the noise image to generate a high-quality image that meets the text description and control conditions.

[0011] Preferably, in step S1, the food image generation architecture based on the two-stream diffusion control model includes a CLIP encoder, a variational autoencoder, a temporal encoder, a control network based on the gating mechanism, an image generation network, and a decoder;

[0012] The food image generation architecture based on the two-stream diffusion control model combines the image generation network and the control network to generate food images in an end-to-end manner.

[0013] Preferably, the input of the food image generation architecture based on the two-stream diffusion control model includes text, time step, conditional image, and Gaussian noise; the time step is processed by the temporal encoder, the text is processed by the CLIP encoder, and the generation process is further guided through the cross-attention mechanism;

[0014] The control network based on the gating mechanism takes the conditional image as the input, and extracts features related to the food shape and spatial information through the image annotation module;

[0015] The image generation network works in parallel with the control network, and their interaction is dynamically controlled by the gating mechanism and then fused, and this mechanism selectively determines the flow of features to the subsequent stages.

[0016] Preferably, in the food image generation architecture based on the two-stream diffusion control model, an adaptive scaling mechanism is introduced to adjust the feature allocation between the backbone network and the skip connections;

[0017] Global and local attention modules are introduced into each stage of the decoder to further focus on the semantic information and detailed texture in the food image.

[0018] Preferably, in step S2, an adaptive scaling mechanism is introduced to adjust the feature allocation between the backbone network and the skip connections, and the specific process is as follows:

[0019] Step S211: Input noise After being processed by the U-Net encoder, extract multi-scale features ; where, is the batch size, is the number of channels, and respectively represent the spatial dimensions of the feature map;

[0020] Step S212: Enhance the backbone features;

[0021] Backbone features After channel-level scaling processing, the channel-level scaling is adjusted through learnable parameters to enhance the global consistency of the backbone network, as follows:

[0022] ;

[0023] Among them, is the learnable channel-level scaling parameter of the backbone features; Acts on each channel, enabling the features to obtain different scaling ratios on different channels;

[0024] Step S213: Perform frequency-domain adjustment on the skip connection features to optimize the skip connection features;

[0025] Step S214: Concatenate the optimized backbone features and skip connection features in the channel dimension to obtain the fused features, as follows:

[0026] ;

[0027] Among them, represents the fused features; and respectively represent the optimized backbone features and skip connection features.

[0028] Preferably, in step S213, the skip connection features are directly passed from the encoder to the decoder for frequency-domain adjustment to optimize information transmission. The specific process is as follows:

[0029] (1) First, convert the skip connection features to the frequency domain to obtain the frequency-domain representation:

[0030] ;

[0031] Among them, represents the Fourier transform, which converts the features in the spatial domain to the frequency domain so that different frequency components can be processed separately;

[0032] (2) To distinguish high-frequency and low-frequency information, calculate the spectral energy distribution of the skip connection features , measure the contributions of different frequency components, as follows:

[0033] ;

[0034] Among them, represents the radius The total energy of frequency components within, representing the proportion of information in this frequency range;

[0035] To dynamically determine the demarcation point between low frequency and high frequency , set the cumulative energy ratio threshold , as follows:

[0036] ;

[0037] Among them, represents the proportion of energy controlling the low-frequency part;

[0038] (3) To dynamically scale the low-frequency and high-frequency features, introduce a frequency mask , as follows:

[0039] ;

[0040] Among them, is the learnable channel-level scaling parameter of the skip connection feature, and the skip connection feature after being adjusted by the frequency mask , as follows:

[0041] ;

[0042] (4) Convert the adjusted frequency-domain features back to the spatial domain through inverse Fourier transform, as follows:

[0043] .

[0044] Preferably, in step S2, add global and local attention modules at each stage of the decoder to further focus on the semantic information and detailed texture in the food image. The specific process is as follows:

[0045] Step S221, local attention divides the input fused feature into 3×3 fixed windows and calculates queries, keys, and values to achieve this, as follows:

[0046] ;

[0047] ;

[0048] ;

[0049] Among them, 、 and respectively represent the query vector, key vector, and value vector of the local feature; 、 and respectively represent the transformation matrix of the query vector of local features, the transformation matrix of the key vector, and the transformation matrix of the value vector; represents the local features obtained by dividing through a 3×3 fixed window;

[0050] The calculation method of local attention is as follows:

[0051] ;

[0052] where, represents the feature output weighted by the local attention mechanism; represents the key vector the transpose of, which is used to calculate the similarity between the query vector and the key vector of local features; represents the scaling factor, which is equal to the dimension of the query vector and the key vector;

[0053] Step S222, Global attention is achieved by performing global pooling on the input fused feature and calculating the query, key, and value, as follows:

[0054] ;

[0055] ;

[0056] ;

[0057] where, , and respectively represent the query vector, key vector, and value vector of global features; , and respectively represent the transformation matrix of the query vector of global features, the transformation matrix of the key vector, and the transformation matrix of the value vector; represents the feature obtained through global pooling;

[0058] The calculation method of global attention is as follows:

[0059] ;

[0060] where, represents the feature output weighted by the global attention mechanism; represents the key vector the transpose of, which is used to calculate the similarity between the query vector and the key vector of global features;

[0061] Step S223, Fuse the features obtained by the local and global attention mechanisms to obtain an optimized feature representation:

[0062] ;

[0063] Among them, represents the optimized feature; α is a learnable weight that balances the contributions of local and global features; the optimized feature is input into the decoder, and combined with the upsampling operation to restore image details and generate a denoised image.

[0064] Preferably, in step S3, based on the fusion network of the gating mechanism, the fusion ratio of the control signal and the generated signal is dynamically adjusted, and the generation process is adaptively optimized according to the characteristics of the food image and the control conditions. The specific process is as follows:

[0065] Step S31: Obtain the initial fused feature by adding the generated feature and the control feature , as follows:

[0066] ;

[0067] Step S32: The generated feature and the control feature interact through the multi-head attention mechanism; among them, the multi-head attention mechanism is used in each layer to enable the network to concurrently focus on different features in multiple subspaces;

[0068] The update formula of the generated feature is as follows:

[0069] ;

[0070] Among them, is the query from the fused feature ; and are the key and value of the generated feature;

[0071] The update formula of the control feature is as follows:

[0072] ;

[0073] Among them, is the query from the fused feature ; and are the key and value of the control feature;

[0074] Step S33: Calculate the gating coefficients and through the gating mechanism, and these coefficients dynamically adjust the importance of the generated feature and the control feature in each layer;

[0075] The gating coefficients determine the contributions of the generated feature and the control feature in different contexts, as follows:

[0076] ;

[0077] ;

[0078] Among them, represents the gating coefficient of the generated feature of the th layer; represents the weight matrix corresponding to the generated feature in the th layer; represents the generated feature of the th layer; represents the control feature of the th layer; represents the bias term of the generated feature of the th layer; represents the gating coefficient of the control feature of the th layer; represents the weight matrix corresponding to the control feature in the th layer; represents the bias term of the control feature of the th layer;

[0079] Step S34, the updated generated feature and the control feature are weighted and summed, and normalized by LayerNorm to obtain the finally updated feature, as follows:

[0080] ;

[0081] ;

[0082] Step S35, the finally updated features are concatenated to obtain the fused feature , as follows:

[0083] ;

[0084] The concatenated feature is further enhanced by the Softmax and tanh mechanisms to aggregate the features, as follows:

[0085] ;

[0086] Among them, represents the transpose of the learnable weight vector; represents the weight matrix, which is used to perform a linear transformation on the fused feature ; represents the bias term;

[0087] Based on the above process, the generated features and control features interact and update effectively between each layer, and finally a food image is generated.

[0088] Preferably, in step S4, based on the dual-stream diffusion generation strategy, through the generation model of the food image generation architecture based on the dual-stream diffusion control model, the noise image is gradually denoised, and the training objective function , is as follows:

[0089] ;

[0090] where, is the target image; is the image after noise processing; is the current time step; is the text condition; is the control condition; is the noise term; is the noise predicted by the model.

[0091] Therefore, the present invention adopts the above-mentioned method for generating a food image based on a dual-stream diffusion control model, and the beneficial effects are as follows:

[0092] (1) Fine control of the generation process: By introducing a dual-stream interaction mechanism and a gating mechanism, the present invention enables the features between the generation network and the control network to optimize each other, ensuring a fine balance between the generated features and control features of the image, and making the generated food image more realistic and delicate; compared with traditional methods, the method proposed in this aspect can capture the detailed features of food more accurately;

[0093] (2) Allocation of the contributions of the backbone network and skip connections: The adaptive addition and scaling mechanism enables the Unet architecture to dynamically adjust the weights of the backbone network features and skip connection features according to specific requirements during the generation process, and combines global and local attention mechanisms to further optimize the features in the network, enabling it to dynamically focus on different important regions during the generation process, ensuring the overall coordination and detailed authenticity of the food image;

[0094] (3) Meeting diverse food generation requirements: The present invention supports multi-modal input of text embedding and conditional images, and can guide the generation network to flexibly respond to different control requirements, enabling the appearance, shape, and details of the food image to be customized according to user-specified requirements; this multi-modal guidance not only improves the flexibility of generation but also ensures a high degree of consistency between the generated results and user input conditions;

[0095] (4)Excellent effects verified by experiments: Through experiments on multiple food image datasets, the method proposed in the present invention is significantly superior to existing diffusion models and generative adversarial network methods; in the generated food images, it performs excellently in capturing details, reproducing complex textures, and generating gloss, and the visual quality and control effect of the generated images have been significantly improved.

[0096] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0097] Figure 1 It is a schematic diagram of the dual-stream diffusion control generation framework of the present invention;

[0098] Figure 2 is Figure 1 The working process schematic diagram of the cross-attention downsampling module, downsampling module, cross-attention intermediate sampling module, global-local attention upsampling module, and cross-global-local attention upsampling module in

[0099] Figure 3 It is a schematic diagram of the global and local attention module of the present invention;

[0100] Figure 4 It is a schematic diagram of the fusion module based on the dual-stream gating mechanism of the present invention. Detailed Embodiments

[0101] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0102] As Figure 1 shown, a method for generating food images based on a dual-stream diffusion control model includes the following steps:

[0103] Step S1, construct a food image generation architecture based on a dual-stream diffusion control model;

[0104] Step S2, introduce an adaptive scaling mechanism and a global and local attention module to optimize the semantic information and detailed texture in food images;

[0105] Step S3, based on a fusion network with a gating mechanism, dynamically adjust the fusion ratio of the control signal and the generation signal, and adaptively optimize the generation process according to the characteristics and control conditions of the food images;

[0106] Step S4, based on a dual-stream diffusion generation strategy, gradually denoise the noise image through the generation model of the food image generation architecture based on the dual-stream diffusion control model to generate high-quality images that meet the text description and control conditions.

[0107] Embodiment

[0108] Step S1: Construct a food image generation architecture based on a two-stream diffusion control model.

[0109] As Figure 1 and Figure 2 shown, the food image generation architecture based on the two-stream diffusion control model includes a CLIP encoder, a variational autoencoder, a temporal encoder, a control network based on a gating mechanism, an image generation network, and a decoder.

[0110] The food image generation architecture based on the two-stream diffusion control model combines the image generation network with the control network to generate food images in an end-to-end manner, improving the quality and control accuracy of the generated images. Among them, the inputs include text, time steps, conditional images, and Gaussian noise. The time steps are processed by the temporal encoder, and the text is processed by the CLIP encoder. As an additional condition, the generation process is further guided through the cross-attention mechanism.

[0111] The control network based on the gating mechanism takes the conditional image as input and extracts features related to the food shape and spatial information through the image annotation module. The image generation network works in parallel with the control network, and their interaction is dynamically controlled by the gating mechanism and then fused. This mechanism selectively determines the flow of features to subsequent stages.

[0112] In addition, in the food image generation architecture based on the two-stream diffusion control model, an adaptive scaling mechanism is introduced to adjust the feature allocation between the backbone network and the skip connections. Global and local attention modules are introduced into each stage of the decoder to further focus on the semantic information and detailed texture in the food image. As the features propagate through each network layer, the predicted noise is obtained. After t time steps, denoising is completed to obtain the generated food picture.

[0113] Step S2: Introduce an adaptive scaling mechanism and global and local attention modules to optimize the semantic information and detailed texture in the food image.

[0114] Step S21: Introduce an adaptive scaling mechanism to adjust the feature allocation between the backbone network and the skip connections.

[0115] Step S211: Input noise After being processed by the U-Net encoder, multi-scale features are extracted ; where is the batch size, is the number of channels, and respectively represent the spatial dimensions of the feature maps.

[0116] The generation task requirements of food images are to capture both overall information and retain local detailed textures. Therefore, the feature allocation of the backbone network and skip connections needs to be optimized to ensure a balance between global consistency and detailed expression. To optimize this allocation, the backbone features are first enhanced, and then the skip connection features are adjusted in the frequency domain to optimize information transmission.

[0117] Step S212: Enhance the backbone features.

[0118] Backbone features Undergo channel-level scaling processing to enhance their global information expression ability. Channel-level scaling is adjusted through learnable parameters to enhance the global consistency of the backbone network and make it have a stronger contribution in the final feature fusion, as follows:

[0119] ;

[0120] Among them, is the learnable channel-level scaling parameter of the backbone features; acts on each channel, enabling the features to obtain different scaling ratios on different channels.

[0121] This adjustment can ensure that the backbone features have a stronger expression ability, thus playing a stable global control role in the skip connection and decoding stages.

[0122] Step S213: Adjust the skip connection features in the frequency domain to optimize the skip connection features.

[0123] Skip connection features Are directly transmitted from the encoder to the decoder, but the unprocessed skip connection features may carry too much high-frequency information, thus introducing noise and affecting the generation quality. Therefore, the skip connection features are adjusted in the frequency domain to optimize information transmission.

[0124] (1) First, convert the skip connection features to the frequency domain to obtain the frequency domain representation:

[0125] ;

[0126] Among them, Represents the Fourier transform, which converts the features in the spatial domain to the frequency domain so that different frequency components can be processed separately.

[0127] (2) To distinguish high-frequency and low-frequency information, calculate the spectral energy distribution of the skip connection features to measure the contributions of different frequency components, as follows:

[0128] ;

[0129] Among them, represents the total energy of frequency components within the radius , indicating the proportion of information in this frequency range.

[0130] To dynamically determine the demarcation point between low frequency and high frequency , a cumulative energy ratio threshold is set as follows:

[0131] ;

[0132] Among them, represents the proportion of the energy of the low-frequency part, ensuring that the model can adaptively adjust the demarcation point between low frequency and high frequency according to different input images to meet the detail requirements of different food images.

[0133] (3) To dynamically scale the low-frequency and high-frequency features, a frequency mask is introduced as follows:

[0134] ;

[0135] Among them, is the learnable channel-level scaling parameter of the skip connection feature. The skip connection feature after being adjusted by the frequency mask is as follows:

[0136] ;

[0137] (4) Subsequently, the adjusted frequency-domain features are converted back to the spatial domain through inverse Fourier transform as follows:

[0138] ;

[0139] This step ensures that the skip connection feature can be restored to the spatial domain after frequency-domain adjustment and carries optimized low-frequency and high-frequency information to enhance the quality of the generated image.

[0140] Step S214: Concatenate the backbone feature and the skip connection feature and in the channel dimension to obtain a fused feature as follows:

[0141] ;

[0142] This concatenation operation ensures that the global consistency of the backbone feature and the local detail information of the skip connection feature can be fully fused to improve the quality of the final generated image.

[0143] Step S22: Add global and local attention modules at each stage of the decoder to further focus on the semantic information and detailed texture in the food image.

[0144] Since the global features contain the semantic information of the whole food, and the local features contain the surface texture and details of the food, so at the decoder stage, global-local attention is used to further optimize the expression of global and local information, as Figure 3 shown.

[0145] Step S221: Local attention divides the input fused features into 3×3 fixed windows and calculates queries, keys, and values to achieve this, as follows:

[0146] ;

[0147] ;

[0148] ;

[0149] where , and represent the query vector, key vector, and value vector of the local features respectively; , and represent the transformation matrices of the query vector, key vector, and value vector of the local features respectively; represents the local features obtained by dividing through the 3×3 fixed window; The calculation method of local attention is as follows:

[0150] ;

[0151] where represents the feature output weighted by the local attention mechanism; represents the transpose of the key vector for calculating the similarity between the query vector and key vector of the local features; represents the scaling factor, equal to the dimension of the query vector and key vector;

[0152] Step S222: Global attention is achieved by performing global pooling on the input fused features and calculating queries, keys, and values, as follows:

[0153] ;

[0154] ;

[0155] ;

[0156] Among them, , and respectively represent the query vector, key vector, and value vector of the global feature; , and respectively represent the transformation matrix of the query vector of the global feature, the transformation matrix of the key vector, and the transformation matrix of the value vector; represents the feature obtained through global pooling;

[0157] The calculation method of global attention is as follows:

[0158] ;

[0159] Among them, represents the feature output weighted by the global attention mechanism; represents the transpose of the key vector , which is used to calculate the similarity between the query vector and the key vector of the global feature.

[0160] Step S223. Finally, fuse the features obtained by the local and global attention mechanisms to obtain an optimized feature representation:

[0161] ;

[0162] Among them, represents the optimized feature; α is a learnable weight that balances the contributions of local and global features.

[0163] Optimized feature is further input into the decoder, and the image details are gradually restored by combining the upsampling operation, and finally a denoised image is generated, ensuring that the generated food image can retain details and maintain global consistency.

[0164] Through the adaptive scaling mechanism, balance the contributions of the backbone network and the skip connection during the generation process, so that the generation network can process global information and local details simultaneously. This mechanism ensures that the generated food image pays more attention to detail processing on the basis of maintaining global consistency, especially in the local features such as the texture, gloss, and hierarchy of the food. In addition, by combining the global and local attention mechanisms, the network can focus on the global structure and local details respectively according to the importance of different regions, further enhancing the realism and visual expressiveness of the food image. Compared with the standard U-Net, this method improves the generation quality, making the processing details of the food image more realistic while maintaining the overall color and shape stable.

[0165] Step S3: Based on the fusion network with a gating mechanism, dynamically adjust the fusion ratio of the control signal and the generated signal, and adaptively optimize the generation process according to the characteristics of the food image and the control conditions.

[0166] In the food image generation task based on the dual-stream diffusion control model network, the interaction between the generated features and the control features is crucial. The generated features represent the feature representation of the existing food images, while the control features are the conditional information that guides the generation process. By interacting these two types of features, more meaningful controls can be added to the generation process, enhancing the diversity and accuracy of the generated images. The purpose of the dual-stream interaction is to utilize the relationship between the generated features and the control features to ensure that the generation process not only depends on the existing image features but also can be precisely guided by the control features, thereby generating food images that meet various requirements, specifically as Figure 4 shown.

[0167] Step S31: First, obtain the initial fused feature by adding the generated features and the control features , as follows:

[0168] ;

[0169] Step S32: Next, the generated features and the control features will interact through the multi-head attention mechanism.

[0170] At each layer, using the multi-head attention mechanism enables the network to concurrently focus on different features in multiple subspaces, thereby enhancing the richness of the interaction. The update formula for the generated features is as follows:

[0171] ;

[0172] where is the query from the fused feature ; and are the key and value of the generated features. Through this formula, the generated features can interact with the control features at a fine-grained level in each layer, enhancing the model's generation ability.

[0173] The control features also go through the multi-head attention calculation to achieve the update of the control features, as follows:

[0174] ;

[0175] where is the query from the fused feature ; and are the key and value of the control features.

[0176] The control features and generation features are updated at this layer, thus enabling cross-modal information exchange.

[0177] Step S33: Calculate the gating coefficients through a gating mechanism and , and these coefficients will dynamically adjust the importance of the generation features and control features in each layer.

[0178] The gating coefficients help the network determine the contributions of the generation features and control features in different contexts, as follows:

[0179] ;

[0180] ;

[0181] where represents the gating coefficient of the generation features of the -th layer; represents the weight matrix corresponding to the generation features in the -th layer; represents the generation features of the -th layer; represents the control features of the -th layer; represents the bias term of the generation features of the -th layer; represents the gating coefficient of the control features of the -th layer; represents the weight matrix corresponding to the control features in the -th layer; represents the bias term of the control features of the -th layer.

[0182] These gating coefficients adjust their contribution ratios according to the combination of the generation features and control features of the current layer. This step enables the generation process to flexibly weight each feature according to different conditions of the input at different levels, thereby improving the quality of image generation.

[0183] Step S34: The updated generation features and control features are weighted and summed, and then normalized through LayerNorm to obtain the final updated fused features, as follows:

[0184] ;

[0185] ;

[0186] The purpose of this step is to ensure that the generated features and control features are reasonably weighted in each layer, so that the final fused features can reflect the best combination of the two-modal information.

[0187] Step S35: Concatenate the updated features to obtain the fused features , as follows:

[0188] ;

[0189] Through concatenation, the representations of the generated features and control features can be uniformly processed, making the next feature update more effective. The concatenated features will continue to be updated, and the feature aggregation will be further strengthened through the Softmax and tanh mechanisms, as follows:

[0190] ;

[0191] where, represents the transpose of a learnable weight vector; represents the weight matrix, which is used to perform a linear transformation on the fused features ; represents the bias term, a learnable parameter.

[0192] This step effectively fuses the generated features and control features through weighted methods, further improving the quality of the information flow, enabling the network to generate more accurate images that meet the conditions.

[0193] Through these steps, the generated features and control features can effectively interact and update between each layer, finally generating food images with higher quality and diversity. When it is necessary to emphasize the detailed areas, the gating mechanism will enhance the role of the control signal to ensure more refined adjustment of details such as texture, while in other areas, the control signal will be appropriately weakened to avoid over-influencing the image generation and maintain the naturalness and realism of the generated results.

[0194] In addition, the gating mechanism further improves the quality of image generation by dynamically adjusting the feature flow between the generation network and the control network. Through the gating mechanism, the model can automatically and selectively transmit and enhance specific features according to the different requirements of each generation stage, avoiding the problem of information flow delay and ensuring the synchronous optimization of image details and structure.

[0195] Step S4: Based on the two-stream diffusion generation strategy, through the generation model of the food image generation architecture based on the two-stream diffusion control model, gradually denoise the noise image to generate high-quality images that meet the text description and control conditions.

[0196] During the generation process, the control network not only receives text conditions but also combines additional image inputs to further guide the details of the generated images. In this way, the generated images not only conform to the text description but also show more control in details, such as the appearance, color, and texture of the food.

[0197] Specific training objective function , as follows:

[0198] ;

[0199] Among them, is the target image; is the image after noise processing; is the current time step; is the text condition; is the control condition; is the noise term; is the noise predicted by the model.

[0200] During the training process, the model optimizes the generation process by minimizing the difference between the predicted noise and the actual noise, thereby generating clear and detailed food images.

[0201] Therefore, the present invention adopts the above-mentioned food image generation method based on a two-stream diffusion control model. By introducing a two-stream interaction mechanism, a gating mechanism, an adaptive scaling mechanism, and an attention mechanism, the quality of the generated food images and the ability to capture details are significantly improved; by adding bidirectional information flow between the generation network and the control network on the basis of the diffusion model and realizing feature transfer through zero convolutional layers, the problem of balancing the global structure and local details in food image generation is effectively solved.

[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A food image generation method based on a dual-flow diffusion control model, characterized in that: The following steps are involved: Step S1, constructing a food image generation architecture based on a dual-stream diffusion control model, including a CLIP encoder, a variational autoencoder, a temporal encoder, a control network based on a gating mechanism, an image generation network, and a decoder; The food image generation architecture based on the two-stream diffusion control model combines the image generation network with the control network to generate food images in an end-to-end manner; The input of the food image generation architecture based on the two-stream diffusion control model includes text, time steps, conditional images and Gaussian noise; the time steps are processed by the temporal encoder, and the text is processed by the CLIP encoder, and the generation process is further guided by the cross-attention mechanism; The control network based on the gating mechanism takes the conditional image as input and extracts features related to the shape and spatial information of food through the image annotation module; The image generation network works in parallel with the control network, and their interaction is dynamically controlled and fused by a gating mechanism that selectively determines the flow of features to subsequent stages; Step S2, introducing an adaptive scaling mechanism and global and local attention modules to optimize the semantic information and detail texture in food images; Step S3: Based on the fusion network of the gating mechanism, dynamically adjust the fusion ratio of the control signal and the generated signal, and adaptively optimize the generation process according to the characteristics of the food image and the control conditions; Step S4: Based on the dual-stream diffusion generation strategy, the noisy image is gradually denoised through the generation model of the food image generation architecture based on the dual-stream diffusion control model to generate a high-quality image that meets the text description and control conditions.

2. The method for generating food images based on a dual-flow diffusion control model according to claim 1, characterized in that: In the food image generation architecture based on the two-stream diffusion control model, an adaptive scaling mechanism is introduced to adjust the feature allocation of the backbone network and the skip connection; Global and local attention modules are introduced into each stage of the decoder to further focus on the semantic information and detailed textures in food images.

3. The method for generating food images based on a dual-flow diffusion control model according to claim 1, characterized in that: In step S2, an adaptive scaling mechanism is introduced to adjust the feature allocation of the backbone network and the skip connection. The specific process is as follows: Step S211: Input noise After being processed by the U-Net encoder, multi-scale features are extracted ;in, is the batch size, is the number of channels, and Respectively represent the spatial size of the feature map; Step S212, enhancing the main features; Main features After channel-level scaling, channel-level scaling is achieved through learnable parameters Adjustments are made to enhance the global consistency of the backbone network as follows: ; in, Learnable channel-level scaling parameters for backbone features; Acting on each channel, the features are scaled differently on different channels; Step S213: performing frequency domain adjustment on the skip connection feature to optimize the skip connection feature; Step S214: Concatenate the optimized backbone features and skip connection features in the channel dimension to obtain fusion features, as shown below: ; in, Indicates fusion features; and They represent the optimized backbone features and skip connection features respectively.

4. The method for generating food images based on a dual-flow diffusion control model according to claim 3, characterized in that: In step S213, the skip connection feature Directly passed from the encoder to the decoder, frequency domain adjustment is performed to optimize information transmission. The specific process is as follows: (1) First, the skip connection features are converted to the frequency domain to obtain the frequency domain representation: ; in, represents Fourier transform, which converts the features of spatial domain into frequency domain so that different frequency components can be processed separately; (2) In order to distinguish high-frequency and low-frequency information, the spectral energy distribution of the skip connection feature is calculated , which measures the contribution of different frequency components, as follows: ; in, Indicates the radius The total energy of the frequency components within represents the proportion of information within this frequency range; In order to dynamically determine the dividing point between low frequency and high frequency , set the cumulative energy ratio threshold , as shown below: ; in, Indicates the energy proportion of controlling the low-frequency part; (3) In order to dynamically scale low-frequency and high-frequency features, a frequency mask is introduced , as shown below: ; in, is the learnable channel-level scaling parameter of the skip connection feature, and the skip connection feature after frequency mask adjustment , as shown below: ; (4) The adjusted frequency domain features are converted back to the spatial domain through inverse Fourier transform as shown below: 。 5. The method for generating food images based on a dual-flow diffusion control model according to claim 1, characterized in that: In step S2, global and local attention modules are added to each stage of the decoder to further focus on the semantic information and detail texture in the food image. The specific process is as follows: Step S221: local attention takes the input fusion features It is implemented by dividing into 3×3 fixed windows and calculating the query, key and value as follows: ; ; ; in, , and The query vector, key vector and value vector represent the local features respectively; , and The transformation matrices of the query vector, the key vector, and the value vector respectively represent the local features; Represents the local features obtained by dividing the fixed window of 3×3; The local attention is calculated as follows: ; in, Represents the feature output after weighting by the local attention mechanism; Represents the key vector The transpose of is used to calculate the similarity between the query vector and the key vector of the local feature; represents the scaling factor, which is equal to the dimensions of the query vector and the key vector; Step S222: Global attention is obtained by fusion of input features. This is done by performing global pooling and calculating the query, key, and value as follows: ; ; ; in, , and Represent the query vector, key vector and value vector of global features respectively; , and The transformation matrices of the query vector, the key vector, and the value vector respectively represent the global features; Represents the features obtained after global pooling; The global attention is calculated as follows: ; in, Represents the feature output after weighting by the global attention mechanism; Represents the key vector The transpose of is used to calculate the similarity between the query vector and the key vector of the global feature; Step S223: Fuse the features obtained by the local and global attention mechanisms to obtain an optimized feature representation: ; in, represents the optimized features; α is a learnable weight that balances the contribution of local and global features; the optimized features The input is sent to the decoder and combined with the upsampling operation to restore the image details and generate a denoised image.

6. The method for generating food images based on a dual-flow diffusion control model according to claim 1, characterized in that: In step S3, based on the fusion network of the gating mechanism, the fusion ratio of the control signal to the generated signal is dynamically adjusted, and the generation process is adaptively optimized according to the characteristics of the food image and the control conditions. The specific process is as follows: Step S31, generate features by and control characteristics Add to get the initial fusion features , as shown below: ; Step S32: generating features and controlling features interact with each other through a multi-head attention mechanism; wherein the multi-head attention mechanism is used in each layer to allow the network to focus on different features in multiple subspaces in parallel; The update formula for generating features is as follows: ; in, It comes from the fusion feature Query; and are the keys and values ​​of the generated features; The update formula of the control feature is as follows: ; in, It comes from the fusion feature Query; and are the keys and values ​​of the control features; Step S33: Calculate the gating coefficient through the gating mechanism and ,These coefficients dynamically adjust the importance of generated features and control features in each layer; The gating coefficient determines the contribution of generated features and controlled features in different contexts as follows: ; ; in, Indicates The gating coefficients of the generated features of the layer; Indicates The weight matrix corresponding to the generated features in the layer; Indicates Generative characteristics of layers; Indicates Control characteristics of the layer; Indicates The bias term of the generated features of the layer; Indicates The gating coefficients of the control features of the layer; Indicates The weight matrix corresponding to the control features in the layer; Indicates The bias term of the control feature of the layer; Step S34: Updated generated features and control characteristics Perform weighted summation and normalize through LayerNorm to obtain the final updated features as shown below: ; ; Step S35: concatenate the updated features to obtain the fused features , as shown below: ; Features after stitching The aggregation of features is further enhanced through the Softmax and tanh mechanisms, as shown below: ; in, represents the transpose of the learnable weight vector; Represents a weight matrix, which is used to fusion features Perform linear transformation; represents the bias term; Based on the above process, the generated features and control features interact and update effectively at each layer, and finally generate food images.

7. The method for generating food images based on a dual-flow diffusion control model according to claim 1, characterized in that: In step S4, based on the two-stream diffusion generation strategy, the noisy image is gradually denoised through the generation model of the food image generation architecture based on the two-stream diffusion control model, and the objective function is trained. , as shown below: ; in, is the target image; is the image after noise processing; is the current time step; It is a text condition; is the control condition; is the noise term; is the noise in the model prediction.

Citation Information

Patent Citations

  • Fast double-flow denoising diffusion model and image fusion method based on soft attention control and without classifier guidance

    CN118864295A

  • Real world image super-resolution method based on stable diffusion

    CN118918009A