Training and optimizing method for e-commerce map generation model
Through multi-stage visual prior embedding and diffusion denoising training combined with reinforcement learning to optimize the e-commerce biographic model, the problem of insufficient image quality and details generated by traditional models in e-commerce scenarios is solved, and the high-quality image generation and advertising delivery effect is improved.
Patent Information
- Application Number
- CN202510475072.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional generative models are difficult to take into account the high-quality generation of products, models and backgrounds in the same model, and have weak personalized details control capabilities, and lack closed-loop feedback and continuous optimization mechanisms based on advertising delivery effects, resulting in the images generated in e-commerce scenarios that cannot meet the quality and detail requirements.
The multi-stage visual prior embedding module and diffusion denoising training method are adopted, combined with reinforcement learning to optimize the e-commerce biographical model, and image details are reconstructed through the multi-stage visual prior embedding module, and the model is optimized using the feed-off data, and the deep understanding and expansion of Chinese prompt words are supported.
The generated product images, model details and background environment are naturally integrated, which conform to the characteristics of e-commerce scenarios, improves generation quality and details, improves the click and conversion rate of advertising delivery, and reduces the burden of user input.
Smart Images

Figure CN120449976A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and image generation technology, and in particular relates to a training and optimization method for an e-commerce image generation model. Background Art
[0002] Simply put, e-commerce raw image models are AI-based, automatically generating image models suitable for e-commerce platforms. These image models are not only high-quality and attractive, but can also be customized to meet the specific needs of e-commerce platforms. They are widely used in scenarios such as product display, personalized recommendations, and advertising design, improving user experience and merchant efficiency. In e-commerce scenarios, high quality is required for model, background, and product images. The goal is to ensure that clothing looks natural, products are prominently positioned, and that human-computer and human-product interactions flow seamlessly.
[0003] The defects of traditional generative models: it is difficult to take into account the high-quality generation of multiple dimensions of products, models and backgrounds in the same model; the control ability of personalized details (posture, clothing, scene) is weak; the prior of single conditions (such as labels or text) is insufficient, which easily leads to missing details; there is a lack of closed-loop feedback and continuous optimization mechanism based on advertising effectiveness. Due to the above defects of traditional generative models, the images generated in e-commerce scenarios cannot meet the quality and detail requirements. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes a training and optimization method for an e-commerce raw image model to solve the problems existing in the above-mentioned prior art.
[0005] To achieve the above objectives, the present invention provides a method for training and optimizing an e-commerce raw image model, comprising:
[0006] Obtaining e-commerce image data, labeling the image data, and obtaining category priors;
[0007] Constructing an e-commerce raw image model and a visual prior embedding module, wherein the e-commerce raw image model is a backbone diffusion model, wherein the backbone diffusion model adopts a neural network model, and wherein the visual prior embedding module includes an encoder and a decoder;
[0008] The backbone diffusion model is trained by random noise injection and diffusion denoising;
[0009] Encoding the image data to obtain latent variables, and inputting the category priors combined with the latent variables into the backbone diffusion model to obtain a preliminary image;
[0010] The preliminary image is compressed and reconstructed through the visual prior embedding module to obtain visual prior information. The visual prior information and the category prior are spliced together and then input into the backbone diffusion model in combination with the input conditions to obtain the optimized e-commerce display image.
[0011] Optionally, the e-commerce image data includes image data of different commodities, models, backgrounds, and interactions of various scenes.
[0012] Optionally, the backbone diffusion model adopts a Transformer-based network structure.
[0013] Optionally, after obtaining the optimized e-commerce display image, the following is also included:
[0014] Based on the optimized e-commerce display image, the process of compression reconstruction, splicing and backbone diffusion model input is repeated until the maximum number of iteration stages is reached to obtain the final optimized e-commerce display image.
[0015] Optionally, the visual prior embedding module includes an 8-layer encoder and a 4-layer decoder, wherein the encoder is used to encode and compress the e-commerce image generated in the previous stage to obtain a compressed latent variable, and the decoder is used to decode and reconstruct the compressed latent variable to obtain visual prior information.
[0016] Optionally, after building the e-commerce raw image model and visual prior embedding module, the following steps are also included:
[0017] The e-commerce raw image model and visual prior embedding module are optimized through reinforcement learning.
[0018] Optionally, the process of optimizing the e-commerce raw image model and visual prior embedding module through reinforcement learning includes:
[0019] Acquire interactive data in real time, wherein the interactive data is advertising effect data, wherein the advertising effect data includes click-through rate, conversion rate and dwell time;
[0020] The corresponding reward function value is calculated through the interaction data, and the model parameters of the backbone diffusion model and the visual prior embedding module are updated according to the reward function value through the reinforcement learning algorithm.
[0021] Optionally, after building the e-commerce raw image model and visual prior embedding module, the following steps are also included:
[0022] Obtain Chinese prompt words, identify them through the semantic parsing model, and obtain Chinese description condition vectors;
[0023] The Chinese description condition vector is used as input data for the e-commerce image generation model.
[0024] Compared with the prior art, the present invention has the following advantages and technical effects:
[0025] 1) Significantly improved generation quality and details: With the help of multi-stage visual priors, the generated product images, model details, and background environment can be naturally integrated to better highlight the characteristics of e-commerce scenarios.
[0026] 2) Display logic that complies with e-commerce scenarios: The model has built-in prior knowledge of logic such as "product center position" and "interaction between people and products," making the automatically generated images more in line with e-commerce needs in terms of layout and visual focus.
[0027] 3) Continuous optimization to improve conversion rates: By collecting advertising results (CTR, CVR, etc.), we continuously iterate model parameters to ensure that the generated images have higher click-through and conversion rates in advertising.
[0028] 4) High comprehension of Chinese prompt words: Allow users to describe with minimal Chinese phrases, and the model can automatically expand and understand them, outputting image results that better meet the needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0030] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention. DETAILED DESCRIPTION
[0031] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0032] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0033] This paper aims to provide a method for generating and continuously optimizing e-commerce scene images based on multi-stage visual priors, with the following main objectives:
[0034] Improved generation quality: Utilizing multi-stage diffusion sampling and visual prior embedding mechanisms, we improve the controllability and realism of product, model, background, and their interaction details.
[0035] Scalability and real-time optimization: We continuously optimize the model through reinforcement learning based on advertising data (click-through rate, conversion rate, etc.), ensuring that the generated images have higher conversion rates and market adaptability in advertising.
[0036] Chinese scene friendliness: Develop in-depth understanding and expansion strategies for Chinese prompt words to minimize user input barriers and generate high-quality images that better match e-commerce needs.
[0037] like Figure 1 As shown in the figure, the present invention is based on the combination of multi-stage diffusion sampling (Diffusion on Diffusion) and the latent embedding module (LEM), further incorporating e-commerce scenario characteristics and data feedback process. The main steps are as follows:
[0038] 1. Data Collection and Preprocessing
[0039] 1) Diversified industry data: Collect massive amounts of high-quality image data from e-commerce platforms or industry partners, including different products, models, backgrounds, and various scene interactions; at the same time, annotate product categories, shooting angles, backgrounds, model parameters, etc.
[0040] 2) Data annotation and enhancement: Label products according to categories, person characteristics, and background environments to form type priors; use image enhancement (rotation, flipping, cropping, color adjustment) to expand data and improve the model's adaptability to diverse e-commerce scenarios.
[0041] 3) VAE encoding (optional): Encode the input image (such as the VAE mechanism in Stable Diffusion) to reduce the resolution space pressure and perform subsequent diffusion training in the latent space.
[0042] 2. Multi-stage visual prior diffusion model
[0043] 1) Backbone Diffusion Model:
[0044] Use Transformer or other neural network structures combined with diffusion denoising scheduling (such as Rectified Flow);
[0045] Classifier-Free Guidance or other conditional integration methods are introduced, using input image data or VAE-encoded input image information, i.e., encoded latent variables, as input data, and embedding conditions such as product category, model information, or time step into the backbone network to provide basic semantic priors.
[0046] 2) Visual Prior Embedding Module (LEM):
[0047] After each generation phase, the latent variables or feature maps of the output image are input into the module, and high-level visual prior information is obtained after "compression-reconstruction";
[0048] Only key structures and semantics are retained, and redundant texture details are removed to prevent the next stage of sampling from excessively "replicating" the output of the previous stage, leaving room for secondary generation and refinement.
[0049] 3) Multi-stage sampling strategy:
[0050] Stage 1: Only using category priors, time information, etc., a diffusion sampling is performed to generate a preliminary image;
[0051] Stage i (i≥2): Input the results generated in the previous stage into LEM to obtain a "visual prior"; combined with other conditions such as product category, diffuse sampling is performed again to continuously refine the interaction between the product texture and the model background;
[0052] Multiple stages of operation can be repeated until the desired image quality or e-commerce display requirements are achieved.
[0053] 3. Delivery data feedback and continuous optimization
[0054] 1) Advertising data collection: Obtain key performance indicators such as click-through rate (CTR), conversion rate (CVR), and viewing time from various advertising channels (e-commerce websites, social media, etc.);
[0055] 2) Data Analysis: Analyze the correlation between different image features (model pose, product center display, background color, etc.) and the advertising effect to identify the key factors that influence conversions / clicks.
[0056] 3) Reinforcement Learning Optimization: Using ad performance as a reward signal, reinforcement learning algorithms (such as policy gradient, DQN, or deep Q-learning) are used to update the parameters of the diffusion model and / or LEM, making the model more likely to generate images that perform better in the delivery process.
[0057] 4) Continuous iterative training: Regularly use updated data feedback to fine-tune or retrain the model to ensure that the model always has high advertising effectiveness and conversion rate under market dynamics.
[0058] 4. Understanding and expanding prompt words (for Chinese scenarios)
[0059] 1) Natural Language Parsing: This strengthens the multi-stage visual prior diffusion model's ability to understand minimalist Chinese prompts. Through the built-in semantic parsing module, it automatically expands user-entered phrases or keywords into more detailed descriptions.
[0060] 2) Multi-dimensional control interface: Based on e-commerce design requirements, users are allowed to simply describe key elements such as "model temperament, product display style, and background atmosphere" in Chinese. The system automatically generates model-friendly conditional vectors to guide diffusion generation;
[0061] Effect evaluation and fine-tuning: Score the match between the prompt word and the generated result. If there is any deviation, fine-tune it through small samples or additional supervision signals to improve the consistency of "prompt word-generated image".
[0062] While addressing the above-mentioned pain points of the prior art, the present invention makes full use of the following ideas:
[0063] Multi-stage visual prior: The image generated in the previous stage is compressed and extracted as the visual prior for the next stage.
[0064] Delivery data feedback (reinforcement learning, etc.): Continuously use advertising performance data such as click-through rate and conversion rate to optimize model parameters to achieve the delivery effect of "high conversion rate and high click-through rate".
[0065] Deep understanding of Chinese prompt words: Targeting the Chinese e-commerce environment, we improve prompt word parsing and keyword expansion capabilities to reduce user input burden.
[0066] As some embodiments, the above technical solutions are described in detail:
[0067] Multi-stage visual prior e-commerce image generation
[0068] Data preparation:
[0069] Collect image data from e-commerce platforms containing multiple categories of products (clothing, food, home furnishings, etc.), models, and backgrounds, with a resolution of, for example, 512×512;
[0070] The image data is labeled (product category, background environment, model attributes) to obtain label information, i.e., category prior, and then the image data is converted into a latent space representation through the VAE encoder, and the latent space representation is a latent variable.
[0071] Model building and training:
[0072] Backbone Diffusion Model: Based on the Transformer structure, it combines with diffusion denoising scheduling and incorporates category priors.
[0073] In the process of incorporating category priors, the Transformer structure is used as the main structure, and a conditional embedding layer, a category-conditional cross-attention module, and an adaptive normalization layer are added to the main structure. Specifically, the following contents are included:
[0074] Conditional Embedding Layer Design:
[0075] Add a category conditional embedding layer after the Transformer input embedding layer to map the category prior (product type, background environment, model attributes) into a vector through a learnable embedding matrix and the input image features (input image or encoded features) Concatenate (concat) As model input for subsequent data feature extraction and processing;
[0076] Cross-attention mechanism: A category-conditional cross-attention module is introduced in the middle layer of the Transformer. The cross-attention mechanism is added after the self-attention mechanism of the Transformer structure, allowing the model to first capture the relationship within the image patch and then interact with the image features and category conditions. The specific structure is self-attention mechanism → layernormal → cross-attention → layernormal → FFN → layernormal. In the cross-attention mechanism, the category prior is embedded as the Key (K)-Value (V) pair, and the image feature is used as the Query (Q). The attention weight is calculated:
[0077]
[0078] Where Q = fWq, K = cWk, V = cWv, Wq, Wk, Wv represent different weights, and dk represents the length of the word vector. In this way, the category semantics is injected into the feature space;
[0079] Adaptive Normalization Layer:
[0080] Introducing category conditions in the layer normalization operation of Transformer:
[0081]
[0082] Where γ(c) and β(c) are generated by category embedding through MLP, and the normalization parameters are dynamically adjusted.
[0083] Visual prior embedding module: 8-layer encoder + 4-layer decoder, used for "compression-reconstruction" of the e-commerce images generated in the previous stage. The e-commerce images are compressed and encoded by the encoder to generate latent variables, which are then reconstructed into visual prior information based on the latent variables.
[0084] The encoder structure consists of 8 layers of encoders. Each layer of encoder contains an input layer, a multi-head attention layer, a residual connection, a normalization layer and a feedforward network.
[0085] Among them, the multi-head self-attention layer adopts a 4-head attention mechanism to calculate cross-region feature correlations, and performs layer normalization after adding residual connections to the multi-head self-attention layer. The output f' of the layer normalization is: f'=LayerNorm(f+Attention(f)), where f represents the input feature of the multi-head self-attention layer, Attention represents the multi-head self-attention layer processing, and LayerNorm represents the layer normalization processing.
[0086] After layer normalization, the layer normalized output features are fed into a feedforward network, which consists of two fully connected layers, with the middle dimension expanded by 4 times and the activation function being GELU. Output dimension: Input latent variable z∈R 64*64*256 , the output of each layer maintains the same dimension.
[0087] The decoder structure consists of 4 layers of decoders. Each layer of decoder contains masked multi-head self-attention, encoder-decoder cross attention, residual connection + layer normalization and transposed convolution;
[0088] Among them, the masked multi-head self-attention masks the features of the future position of the input data and processes the masked information using multi-head self-attention to focus only on the features before the current position.
[0089] Encoder-decoder cross attention: The final output of the encoder, i.e. the latent variable, is used as the key-value, and the decoder uses the features obtained by masked multi-head self-attention processing as the query.
[0090] Add residual connections to the encoder-decoder cross attention and perform layer normalization. The residual connection + layer normalization process is the same as the encoder, f'=LayerNorm(f+Attention(f)), where f represents the input feature of the encoder-decoder cross attention, and Attention represents the encoder-decoder cross attention.
[0091] Output reconstruction: The latent variable is finally upsampled to the original resolution (e.g. 512×512) through transposed convolution.
[0092] Training process:
[0093] Perform random noise injection and diffusion denoising on the latent variables of the encoded input image;
[0094] Use only category priors with a certain probability (such as the cf.guidance method), or replace the results of the previous stage with learnable tokens in a multi-stage training logic;
[0095] Repeat the iteration until convergence.
[0096] Inference and Sampling:
[0097] Stage 1: Input conditions such as product category and time step, denoise the noise latent variables, and obtain a preliminary image;
[0098] Stage 2: The output of Stage 1 is passed through LEM to obtain visual prior information. This is then combined with the category prior and sampled again to obtain an optimized background-product-model image.
[0099] During the training process, noise injection and diffusion denoising are used for training.
[0100] The noise injection process includes the following:
[0101] Noise Scheduling: Adopting Linear Noise Scheduling T = 1000, where βt is the noise intensity, t represents the injection time step, and T represents the maximum injection time step.
[0102] Latent variable diffusion: gradually add noise to the latent variable z0 (the latent variable of the image data after the VAE encoder), and the noise latent variable z in the tth step t for:
[0103]
[0104] in Where s represents the time step before time step t, β s represents the noise intensity at step s, a t represents the noise retention coefficient, ∈ represents standard Gaussian noise, and N represents normal distribution.
[0105] Denoising training process
[0106] Input construction: Random sampling time step t ~ Uniform(1,T), generating noisy latent variable z t .
[0107] Noise prediction: Model fθ(z t ,t,c) Output prediction noise ε θ , the loss function is:
[0108]
[0109] Among them, E represents expectation, and the model weight is adjusted through the loss function.
[0110] Multi-stage training: only the category prior is used (the visual prior input is blocked) with a probability of 30%, and in the rest of the cases, multi-stage features are spliced together, i.e., visual prior;
[0111] During the multi-stage inference and sampling process:
[0112] Stage 1 initial generation:
[0113] Input category condition is category prior c, from pure noise z T Starting from ~N(0,I), denoising is iterated through the backbone diffusion model:
[0114]
[0115] Predict the current noise ε using the backbone diffusion model θ (z t ,t,c), and use the current denoised latent variable to generate a preliminary image I1, where η represents Gaussian noise.
[0116] Stage 2 visual prior enhancement:
[0117] LEM compression: Input I1 into the 8-layer encoder to obtain a low-dimensional latent variable
[0118] Prior splicing: It is concatenated with the category embedding c as the conditional input of Stage2, and the encoded latent variable is denoised again. The optimized image I2 is generated by combining the denoised latent variable through the backbone diffusion model.
[0119] If higher precision is required, multiple stages of iteration can be repeated until the desired quality is achieved.
[0120] Optimize data reinforcement learning:
[0121] Collect advertising results: Place the generated images on e-commerce ads or social media to collect data such as click-through rate, conversion rate, and dwell time.
[0122] Reward mechanism: Set a reward function based on the advertising effect, so that images with high conversion rates and high click-through rates correspond to higher rewards;
[0123] Parameter update: Use reinforcement learning algorithms to update some parameters of the diffusion model or LEM, iteratively improving the performance of the model in e-commerce delivery.
[0124] The reinforcement learning part includes the following:
[0125] Reinforcement Learning State-Action-Reward Definition System
[0126] Status: It consists of three parts:
[0127]
[0128] in: The latent variable output by the LEM module; c: product category embedding vector; m t ∈R5 :Real-time advertising indicators (CTR, CVR, dwell time, conversion amount, sharing rate);
[0129] Action:
[0130] a t =Δθ∈R d
[0131] Fine-tune the channel attention weight θ of the diffusion model UNet part, Δ represents the gradient, dimension d = 1.2M
[0132] Reward function design:
[0133] R = 0.4 sigmoid(CTR) + 0.5 log(CVR + 1e-6) + 0.1 tanh(stay time / 30)
[0134] 2. Strategy optimization implementation
[0135] Using the PPO algorithm, the loss function L CLIP Include:
[0136]
[0137] Advantage function Calculated by GAE, Et represents expectation, rt represents importance sampling ratio, clip represents clip function, and ε represents clipping threshold;
[0138] The policy network is the UNet part of the diffusion model. The value function network is a 3-layer MLP (512-256-1).
[0139] The policy network is used to form a subsequent predicted strategy, and based on the reward function, the reward value of the predicted strategy is predicted through the value function network to determine the direction in which the current state needs to be adjusted and adjust the current state.
[0140] Chinese minimalist prompt word expansion:
[0141] Prompt word parsing: Users only need to enter a description such as "red dress, summer outdoor scene", and the system will automatically expand it into a detailed description, such as "red summer dress, model showing in an outdoor scene, upper body close-up" and convert it into a conditional vector;
[0142] Multi-stage visual prior sampling (Diffusion on Diffusion)
[0143] The output of the previous stage is used as the visual prior for the next stage, and the texture details of products and models are refined across stages to achieve more realistic e-commerce display images.
[0144] The “compression-reconstruction” mechanism of the Visual Prior Embedding Module (LEM)
[0145] The Encoder-Decoder structure removes redundant textures from the image generated in the previous stage, retaining only the core semantic and structural information. This avoids simple reconstruction during subsampling and significantly improves the generated details and diversity.
[0146] Reinforcement learning optimization integrating e-commerce delivery feedback
[0147] Based on real delivery data such as advertising click-through rate and conversion rate, the model weights are updated through reinforcement learning to form a closed-loop continuous optimization to achieve higher delivery effects.
[0148] Understanding and expanding Chinese minimalist prompt words
[0149] Automatically expand and analyze Chinese prompt words in multiple dimensions to reduce the difficulty of user input and help generate results that better meet the personalized and aesthetic needs of e-commerce scenarios.
[0150] In summary, the present invention can significantly improve the generation quality and personalization capabilities of product images in e-commerce scenarios, and continuously iterate and optimize with the support of delivery data feedback, which is of great significance for meeting rapidly changing market demands and improving advertising conversion rates.
[0151] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A training and optimization method for an e-commerce raw image model, characterized in that: include: Obtaining e-commerce image data, labeling the image data, and obtaining category priors; Constructing an e-commerce raw image model and a visual prior embedding module, wherein the e-commerce raw image model is a backbone diffusion model, wherein the backbone diffusion model adopts a neural network model, and wherein the visual prior embedding module includes an encoder and a decoder; The backbone diffusion model is trained by random noise injection and diffusion denoising; Encoding the image data to obtain latent variables, and inputting the category priors combined with the latent variables into the backbone diffusion model to obtain a preliminary image; The preliminary image is compressed and reconstructed through the visual prior embedding module to obtain visual prior information. The visual prior information and the category prior are spliced together and then input into the backbone diffusion model in combination with the input conditions to obtain the optimized e-commerce display image.
2. The method according to claim 1, characterized in that The e-commerce image data includes image data of different products, models, backgrounds and various scene interactions.
3. The method according to claim 1, characterized in that The backbone diffusion model adopts a Transformer-based network structure.
4. The method according to claim 1, wherein After obtaining the optimized e-commerce display image, it also includes: Based on the optimized e-commerce display image, the process of compression reconstruction, splicing and backbone diffusion model input is repeated until the maximum number of iteration stages is reached to obtain the final optimized e-commerce display image.
5. The method according to claim 4, characterized in that The visual prior embedding module includes an 8-layer encoder and a 4-layer decoder, wherein the encoder is used to encode and compress the e-commerce image generated in the previous stage to obtain the compressed latent variable, and the decoder decodes and reconstructs the compressed latent variable to obtain visual prior information.
6. The method according to claim 1, characterized in that After building the e-commerce raw image model and visual prior embedding module, we also need to: The e-commerce raw image model and visual prior embedding module are optimized through reinforcement learning.
7. The method according to claim 6, characterized in that The process of optimizing the e-commerce raw image model and visual prior embedding module through reinforcement learning includes: Acquire interactive data in real time, wherein the interactive data is advertising effect data, wherein the advertising effect data includes click-through rate, conversion rate and dwell time; The corresponding reward function value is calculated through the interaction data, and the model parameters of the backbone diffusion model and the visual prior embedding module are updated according to the reward function value through the reinforcement learning algorithm.
8. The method according to claim 1, characterized in that After building the e-commerce raw image model and visual prior embedding module, we also need to: Obtain Chinese prompt words, identify the Chinese prompt words through a semantic parsing model, and obtain a Chinese description condition vector; and use the Chinese description condition vector as input data for the e-commerce image generation model.
Citation Information
Cited By
Abnormal image generation method and system, computer and storage medium
CN122244222A