Image controllable generation method and system for packaging field
By combining rolling active learning and multimodal expert knowledge labeling with a fully bidirectional interactive Transformer block, the problems of professional attribute filtering and structural consistency in image generation in the field of packaging design are solved, and efficient and accurate packaging image generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ADVANCED INST OF INFORMATION TECH (AIIT) PEKING UNIV
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to effectively address issues such as professional attribute filtering, accurate text annotation, and structural consistency generation of high-quality images in the packaging design field. This is particularly true in food and cosmetic packaging design, where general image generation models are inefficient and lack accuracy, and existing controllable generation solutions suffer from structural offsets and insufficient support for Chinese text.
A rolling active learning strategy combined with the RankNet model is used to select images with high aesthetic scores. Through multimodal expert knowledge labeling and a fully bidirectional interactive Transformer block, deep interaction between text, images and control charts is carried out to construct noisy latent variables to generate packaging design images.
It significantly reduces data annotation costs, improves the accuracy and efficiency of aesthetic preferences and annotation, enhances the spatial structure consistency and Chinese text fidelity of generated images, and ensures the high quality and professionalism of generated images.
Smart Images

Figure CN122066814B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an image generation method, and more specifically to a controllable image generation method and system for use in the packaging field. Background Technology
[0002] In recent years, AI-generated content technology has made significant progress in image generation, particularly in DiT (Diffusion Transformer) models based on diffusion models and their combination with the Transformer architecture. These technologies can generate high-quality images based on text prompts, demonstrating great potential in multiple application areas. However, when it comes to highly specialized packaging design fields, such as food and cosmetic packaging design, general image generation models are often difficult to apply directly. This is because packaging images not only require extremely high aesthetic quality but also a deep understanding of the professional attributes of packaging, such as specific box shapes, materials, and printing processes. Furthermore, absolute precision in packaging structure, such as creases and edges, and zero-error rendering of text on the packaging surface are essential. Therefore, accurately filtering out low-quality packaging images, accurately annotating the professional attributes and text of high-quality images without illusions, and using this as a basis to achieve highly consistent packaging image generation have become core elements for improving the quality of AIGC in the packaging field.
[0003] To meet the requirements of industrial-grade deployment, three core stages are typically required: cleaning and aesthetically screening high-quality professional data, accurate multimodal annotation of image-text pairs, and training of a controllable generative architecture based on a foundational model. First, in the packaging data cleaning and aesthetic scoring stage, existing technologies typically employ manual absolute value scoring or use general-purpose state-of-the-art aesthetic scoring models such as HPSv2 and Image Score. While manual scoring provides relatively accurate results, it is inefficient and highly subjective, with a single person only able to annotate around 250 images per day. While general-purpose aesthetic models are efficient, their lack of expertise specific to the packaging domain results in an accuracy rate of less than 50% for human preferences in the packaging field. Second, in the multimodal professional annotation stage, supervised image-text pair training requires detailed text descriptions. Current mainstream solutions directly use general-purpose large language models to infer annotations from the images. However, the packaging field involves extremely high levels of professional knowledge, such as distinguishing between "clasp boxes" and "airplane boxes," and between gravure printing and flexographic printing. Furthermore, packaging surfaces often have complex text layouts, so directly relying on large models for marking can lead to serious "illusions"—the model may fabricate non-existent information or make identification errors. Finally, in the controllable generation architecture stage of the DiT base, existing technologies typically employ controllable generation schemes such as EasyControl, OmniControl, and ControlNet, but these have shortcomings in feature interaction and structural position offset, limiting the quality and consistency of the generated images.
[0004] Currently, mainstream solutions fall into two main categories: data processing solutions based on general large models and controllable generation solutions based on diffusion models. In terms of data processing, general large models have a wide range of applications, but they are prone to "illusions" in vertical expertise and specific text positions, leading to feature mismatches, and their accuracy in aesthetic preferences in the packaging field is extremely low. If purely manual absolute value scoring and labeling are used, there are problems such as strong subjectivity, high cost, and extremely low classification efficiency. Regarding controllable generation, mainstream solutions can be compatible with multiple control modes, but they suffer from incomplete interaction between control flow and image flow, forced scaling leading to structural position shifts, and large extrinsic parameters resulting in slow and unstable training. More importantly, existing control strengths are often static, ignoring the diffusion network's principle of "determining structure in large time steps and refining details in small time steps," resulting in weak structural fidelity of the final generated image to the control map and extremely poor support for Chinese text. For example, EasyControl uses a KV-cache approach, resulting in only unidirectional attention computation between image features and control features; OmniControl simply concatenates control features and image features together and inputs them into the network; ControlNet, when migrated to the DiT architecture, exposed problems of slow and unstable training. None of these solutions effectively address the high cost of data annotation, the illusion of large models, and the issues of spatial structure consistency and Chinese text fidelity in the packaging domain.
[0005] Therefore, it is necessary to design a new method to significantly reduce data annotation costs and improve the accuracy and efficiency of aesthetic preferences, while addressing the shortcomings of existing controllable generation models in terms of structural consistency and Chinese text support, thereby enhancing the spatial structural consistency and Chinese text fidelity of the final generated image. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for controllable image generation in the field of packaging.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: a method for controllable image generation in the field of packaging, comprising:
[0008] Get the initial image; Visual models and resolution thresholds are used to filter packaging images with clear and complete subjects from the initial images, and images with high aesthetic scores are labeled by a rolling active learning strategy combined with RankNet to obtain the initial screening images; Based on the initial screening images, batch automatic labeling is performed, and the text content and its position information are fixed to obtain multimodal expert knowledge labeling results; Based on the multimodal expert knowledge labeling results, feature sequences are extracted by the encoder and noisy latent variables are constructed. The fully bidirectional interactive Transformer block performs deep interaction between text, images and control charts to generate packaging design images.
[0009] The further technical solution is as follows: The process of using a visual model and resolution threshold to filter packaging images with clear and complete subjects from the initial images, and using a rolling active learning strategy combined with RankNet to label images with high aesthetic scores, to obtain the initial screened images, includes: The initial images are initially screened using a resolution threshold and a visual backbone network, retaining images with the required length and width and clear subjects to obtain a basic candidate pool; Random images are randomly generated from the basic candidate pool. Each image is paired with several other images and labeled to form several initial comparison pairs. The confusion index of the remaining images in the basic candidate pool is predicted based on the RankNet model. Images with confusion indices that meet the requirements are selected to form new comparison pairs with images in the initial comparison pairs. This process is repeated several times to utilize the forced loop link mechanism to enable the RankNet model to stably evaluate the aesthetics of the images. Images with high aesthetic scores are selected using the RankNet model to obtain the initial screening images.
[0010] The further technical solution is as follows: The RankNet model uses DINOv3-7b as the underlying feature extractor, and connects a multilayer perceptron at the end of the network to map the image into a one-dimensional aesthetic scalar score, and uses a model based on cross-entropy to learn the comparison relationship.
[0011] The further technical solution is as follows: The step of performing batch automatic labeling based on the initial screening images, and fixing the text content and its position information to obtain multimodal expert knowledge labeling results includes: A structured tagging system was established based on manual verification and a proprietary database. A multi-label classification network is trained based on the structured labeling system, and the trained multi-label classification network is used to automatically annotate the preliminary screening images with professional features to obtain the annotation results. The annotation results are used to identify and record the text content and its location information on each packaging image to obtain the recording results; The annotation results and the recording results are injected into the multimodal large language model to guide the generation of text descriptions that strictly follow the control conditions, so as to obtain the multimodal expert knowledge labeling results.
[0012] The further technical solution is as follows: the structured label system includes professional knowledge labels in multiple dimensions, including style, box type, industry, material and printing process.
[0013] The further technical solution is as follows: the control conditions include following the conditions of supplementary lighting, color of screen elements and environmental background, without modifying the existing material, box shape and text position information; the multimodal expert knowledge labeling results include tags composed of discrete words, short prompts describing the core subject and key text, and long prompts containing complete natural language descriptions of light and shadow, environment and detail texture.
[0014] Its further technical solution is as follows: Based on the multimodal expert knowledge labeling results, feature sequences are extracted by an encoder and noisy latent variables are constructed. A fully bidirectional interactive Transformer block performs deep interaction between text, images, and control charts to generate packaging design images, including: Obtain the control chart and use the multimodal expert knowledge labeling results as input information; Extract the features of the input information and construct noisy latent variables related to the time step to obtain the extraction results; The extraction results are segmented into blocks and linearly projected to a unified hidden layer dimension. Temporal and spatially aware positional information is added to the initial image and the control map through time-step embedding and two-dimensional rotational position encoding to maintain frequency domain consistency and obtain the generated result. The generated results are processed by using a pure three-stream fully bidirectional interactive Transformer block combined with adaptive layer normalization and global joint attention mechanism to achieve deep feature fusion. Feature updates are optimized by sequence segmentation and attention residual gating. A feedforward neural network with shared MLP weights is used to enhance the feature mapping of the noisy image stream and the control image stream in the common visual space to obtain the mapping result. The LoRA mechanism, a time-aware control mechanism, is used to adjust the control strength of the model at different time steps. The loss function is used to optimize the predicted velocity field. A clear image is generated from the noise of the mapping result by solving ODE and decoding VAE to obtain the packaging design image.
[0015] The further technical solution is as follows: the extraction of the respective features of the input information and the construction of noisy latent variables related to the time step to obtain the extraction result includes: The multimodal expert knowledge labeling results are converted into text feature sequences, the initial image is encoded to obtain initial latent variables, and noisy latent variables are generated based on linear interpolation at time steps. The control chart is encoded as a latent variable to obtain the extraction results.
[0016] The further technical solution is as follows: The LoRA mechanism, which uses time-aware control, adjusts the control strength of the model at different time steps, and uses a loss function to optimize the predicted velocity field. A clear image is then generated from the noise of the mapping result through ODE solving and VAE decoding to obtain the packaging design image, including: When calculating the attention layer and the MLP projection matrix, a time-aware low-rank adaptation is introduced to adjust the model control strength at different time steps. The mean square error is used as the loss function to optimize the predicted velocity field to approximate the target velocity field. Starting from pure noise, the latent representation is obtained by inverse integration of the mapping result using the ODE solver, and then restored to the finished image by the VAE decoder to obtain the packaging design image.
[0017] The present invention also provides an image controllable generation system for the packaging field, comprising: The acquisition unit is used to acquire the initial image; The image screening unit is used to screen packaging images with clear and complete subjects from the initial images using a visual model and a resolution threshold, and to label images with high aesthetic scores using a rolling active learning strategy combined with RankNet to obtain the screening images. The labeling unit is used to perform batch automatic labeling based on the initial screening images and fix the text content and its position information to obtain multimodal expert knowledge labeling results; The generation unit is used to extract feature sequences and construct noisy latent variables based on the multimodal expert knowledge labeling results, and to generate packaging design images through deep interaction between text, images and control graphs by a fully bidirectional interactive Transformer block.
[0018] The advantages of this invention compared to existing technologies are as follows: This invention uses a rolling active learning strategy combined with the RankNet model to select high-aesthetic-score initial images, and performs batch automatic labeling based on these high-quality images, fixing the text content and its positional information, thereby constructing multimodal expert knowledge labeling results. Then, an encoder is used to extract feature sequences and construct noisy latent variables, utilizing a fully bidirectional interactive Transformer block to promote deep interaction between text, images, and control graphs, thereby generating packaging design images. This method significantly reduces data labeling costs, improves the accuracy and efficiency of aesthetic preferences, and simultaneously addresses the shortcomings of existing controllable generation models in terms of structural consistency and Chinese text support, enhancing the spatial structural consistency and Chinese text fidelity of the final generated images. By introducing a time-aware control flow and a forced control graph alignment mechanism, not only is the model's ability to control details enhanced, but the high quality and professionalism of the generated images are also ensured.
[0019] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating a method for controllable image generation in the packaging field provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of a pure three-stream generation architecture provided in an embodiment of the present invention; Figure 3 A schematic diagram of packaging design images provided for embodiments of the present invention; Figure 4 A schematic block diagram of an image controllable generation system for the packaging field provided in an embodiment of the present invention; Figure 5 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0024] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0025] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0026] Please see Figure 1 , Figure 1 This is a flowchart illustrating a controllable image generation method for the packaging field provided in this invention. This method is applied to a server. It utilizes a rolling active learning strategy combined with a RankNet model to screen high-aesthetic-score initial images, and employs a multimodal large language model to perform batch automatic labeling and construct a structured labeling system for these images, effectively reducing data annotation costs and improving the accuracy and efficiency of aesthetic preference labeling. Based on extracted feature sequences and noisy latent variables, a fully bidirectional interactive Transformer block is used to achieve deep interaction between text, images, and control graphs. Adaptive layer normalization and a global joint attention mechanism optimize feature updates, ensuring the spatial structural consistency of the final generated images. Simultaneously, the introduction of a time-aware control LoRA mechanism and a method for reverse-generating clear images using a VAE decoder not only enhances control accuracy at different time steps but also significantly improves Chinese text support capabilities, thus addressing the shortcomings of existing controllable generation models in terms of structural consistency and Chinese text fidelity. This method significantly improves the quality and professionalism of generated packaging design images, meeting the needs of a specific field.
[0027] Figure 1 This is a flowchart illustrating a method for controllable image generation in the packaging field provided by an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110 to S140.
[0028] S110, Obtain the initial image.
[0029] In this embodiment, the initial image refers to an unprocessed or unfiltered packaging image obtained from the original data source. These images may originate from various sources, including but not limited to the Internet, internal corporate databases, or actual product packaging photographs captured by specific devices. The quality and characteristics of the initial images may vary significantly, such as varying resolutions, different levels of sharpness of the included subjects, and varying background complexity.
[0030] First, a large number of packaging images are collected as initial images through various means. These images can cover different industries (such as food and cosmetics), packaging types (such as snap-on boxes and corrugated boxes), and materials (such as white cardboard and specialty paper) to ensure the diversity of the dataset. Next, all collected images undergo preliminary screening. Predefined resolution thresholds (e.g., at least 512 pixels in length and width) are used to exclude low-quality images that clearly do not meet the requirements. This step aims to quickly remove images that are obviously unsuitable for subsequent processing, such as those with too low resolution or blurry images. Then, a visual model (such as SigLIP2) is further used to construct a packaging quantity classification model and a packaging integrity model to evaluate whether the number of packages in each image meets the standards (e.g., 0-6+ classes) and whether the packaging is intact. The purpose of this stage is to identify and eliminate images containing cluttered elements or incomplete subjects, ensuring that only high-quality and clearly informative images proceed to the next round of processing. After the above two screening steps, the remaining images are included in the basic candidate pool. These images have high subject clarity and their content basically meets the requirements of packaging design. These images will serve as the foundational material for the next step of the active learning and aesthetic cleaning module, used to train the RankNet model to select images with higher aesthetic scores.
[0031] In summary, step S110 is primarily responsible for selecting a batch of high-quality, clearly defined candidate images from a massive amount of raw packaging images, providing reliable data support for subsequent, more refined data processing and model training. This process involves not only the application of technical methods but also manual review to ensure the accuracy and effectiveness of the selection results.
[0032] S120. Using a visual model and resolution threshold, select packaging images with clear and complete subjects from the initial images, and use a rolling active learning strategy combined with RankNet to label images with high aesthetic scores to obtain the initial screening images.
[0033] In this embodiment, the initial screening images refer to high-quality, aesthetically pleasing images selected from a large number of original packaging images after a series of screening and evaluation processes. These images not only have a clear subject and complete packaging structure, but also meet specific aesthetic standards, making them suitable for subsequent multimodal labeling and generative network training.
[0034] In one embodiment, step S120 described above may include steps S121 to S123.
[0035] S121. The initial images are initially screened using a resolution threshold and a visual backbone network, retaining images with the required length and width and clear main subject to obtain a basic candidate pool.
[0036] In this embodiment, the basic candidate pool refers to the set of packaging images that, after initial screening, meet minimum quality requirements (e.g., a resolution of at least 512x512 pixels) and have high subject sharpness. The specific operation is as follows: Resolution threshold: First, set a resolution threshold (e.g., at least 512 pixels in length and width) to quickly exclude samples with too low resolution or blurry images.
[0037] Visual backbone network: Next, two classification models are built using a pre-trained visual backbone network (such as SigLIP2): one is a classification model for identifying the number of packages (labeled as classes 0 to 6+), and the other is a model for determining the integrity of the packages (complete / partially incomplete). These two models are trained by manually sampling and labeling a portion of the dataset (e.g., 3000 incomplete and 3000 complete images each).
[0038] Screening criteria: Images with more than 5 images or those deemed incomplete are directly removed to ensure that the remaining images have clear subjects and explicit information, forming a basic candidate pool.
[0039] S122. Random images are randomly selected from the basic candidate pool. Each image is paired with several other images and labeled to form several initial comparison pairs. The confusion index of the remaining images in the basic candidate pool is predicted based on the RankNet model. Images with confusion indices that meet the requirements are selected to establish new comparison pairs with images in the initial comparison pairs. This process is repeated several times to utilize the forced loop link mechanism to enable the RankNet model to stably evaluate the aesthetics of the images.
[0040] In this embodiment, the RankNet model uses DINOv3-7b as the underlying feature extractor and connects a multilayer perceptron at the end of the network to map the image into a one-dimensional aesthetic scalar score. It uses a cross-entropy-based forced model to learn the comparison relationship.
[0041] In this stage, a rolling active learning strategy was primarily used to optimize the RankNet model, thereby improving the accuracy of image aesthetic scoring. The specific steps are as follows: Initial construction: 100 images are randomly selected from the basic candidate pool. Each image is paired with 5 other randomly selected images and manually labeled as "A is better than B" or "B is better than A", forming 500 initial comparison pairs.
[0042] Confusion filtering and loop formation: The RankNet model is trained using the comparison pairs from the current round. Then, the trained model is used to globally score and predict all remaining images in the basic candidate pool, calculating the "confusion index" of unlabeled images. The 100 new images with the highest confusion index are extracted, and each new image is then compared with 5 previously labeled images.
[0043] Iterative process: Repeat the above process for about 20 rounds. Since each new image is compared with the historical anchor image, a strongly connected loop link is logically formed, which greatly improves the information density, so that only about 10,000 comparison pairs need to be manually labeled to complete the model convergence.
[0044] S123. Use the RankNet model to filter out images with high aesthetic scores to obtain the initial screening images.
[0045] In the final stage, the pre-trained RankNet model is used to comprehensively score and rank the images in the entire basic candidate pool, eliminating low-scoring data and selecting approximately 40,000 high-aesthetic-score images as initial screening images. These images will then enter the next stage of processing, namely the multimodal expert knowledge labeling module, to further refine and prepare the training input for the pure three-stream temporal-aware generative architecture.
[0046] This series of steps not only effectively improves the quality and efficiency of image screening, but also provides a solid data foundation for subsequent deep learning and image generation.
[0047] In this embodiment, the entire step S120 is to enable the generative model to learn clean and aesthetically pleasing packaging features by performing multi-level purification from 400,000 original packaging data, eliminating samples that are messy, incomplete, and aesthetically poor.
[0048] First, to ensure image quality and integrity, a basic resolution threshold was set to filter images. Specifically, any image with a width and height less than 512 pixels was directly excluded because these low-resolution images could not provide sufficient detail, which was detrimental to subsequent aesthetic scoring and model training. Next, two key classification models were built: one for identifying the number of packages (labeled as classes 0 to 6+), and another for determining package integrity (complete / partially incomplete). These two models were trained on a manually labeled dataset containing 3000 incomplete images and 3000 complete images. By using SigLIP2 as the visual backbone network, images containing multiple packages or incomplete package structures were effectively identified and removed from the candidate pool. Finally, after this step, only images with clear subjects and only one or a few packages were retained, forming the basic candidate pool. This process ensured at least 98% accuracy, thus guaranteeing the image quality in the basic candidate pool.
[0049] To overcome the inefficiency of traditional absolute value scoring (processing only about 250 data points per day), this embodiment employs a pairwise comparison scoring method based on RankNet, significantly improving efficiency and enabling the processing of up to 3000 pairs of comparison data per day. This pairwise comparison-based method not only improves efficiency but also enhances the model's ability to evaluate the relative aesthetic standards of images. The following are the specific implementation steps of rolling active learning: One hundred images are randomly selected from the initial candidate pool. Each image is then paired with five randomly selected images from the pool, meaning each image undergoes a one-to-one aesthetic comparison with the other five images. Professionals then label these paired images based on their personal aesthetic preferences, determining whether "A is better than B" or "B is better than A." This process creates the initial 500 comparison pairs.
[0050] Using the aforementioned 500 comparison pairs as training data, a RankNet model is trained. Subsequently, the trained model is used to perform global score predictions on the remaining images in the entire base candidate pool. The "confusion index" of unlabeled images is calculated; this index reflects the degree of uncertainty in the distribution of image scores across all images. Generally, if an image's score is concentrated near the median, or its score differs very little from other image scores, the image is considered a sample with a high confusion index. The top 100 new images with the highest confusion index are selected, and each new image is forced to establish a new comparison pair with 5 previously labeled images. This ensures that each newly added image is tightly integrated with the existing evaluation system, forming a strongly connected loop.
[0051] The above process is repeated approximately 20 times. Each iteration adds new images to the existing comparison pairs, making the evaluation criteria of the entire system more stable and accurate. Since each new image is compared with historical anchor images, a strongly connected loop is logically formed: "A is compared with B, B is compared with C, and A and C are deduced." This method greatly improves information density, allowing the model to converge with only about 10,000 manually labeled comparison pairs, significantly reducing manual costs while improving model accuracy.
[0052] In the design of the RankNet aesthetic model, DINOv3-7b was chosen as the underlying feature extractor, which offers stronger cognitive capabilities compared to other models such as SigLIP2 and CLIP. Based on DINOv3-7b, a multilayer perceptron (MLP) is integrated to map images to a one-dimensional aesthetic scalar score. For any pair of images (A, B), if A is labeled as superior to B, then the target probability is... RankNet Loss uses the cross-entropy loss function, which is expressed as follows: , , This represents the aesthetic scalar scores predicted by the model for images A and B. In this context, the model maps the input image to an aesthetic score on a real number, where a higher score means the model considers the image more aesthetically pleasing. This refers to the difference between the aesthetic score of image A and the aesthetic score of image B. If this difference is positive, it means the model considers image A more aesthetically pleasing than image B; conversely, if the difference is negative, it means image B is considered more aesthetically pleasing than image A. In this way, the model continuously adjusts its parameters to widen the score difference between positive and negative samples, achieving gradient descent. The core function of this loss function is to guide the model to adjust its parameters so that for image pairs labeled "A is better than B," the model assigns a higher score to A rather than B, and optimizes the model by minimizing the value of the loss function, making it better conform to human aesthetic standards. This mechanism ensures that the model can accurately score images aesthetically in subsequent processing.
[0053] Ultimately, the fully trained RankNet model was able to comprehensively score and rank all 400,000 original images, selecting approximately 40,000 high-aesthetic-score packaged images as input for the next stage of processing. This method not only effectively improved the quality and efficiency of image selection but also provided solid data support for subsequent deep learning and image generation.
[0054] S130. Based on the initial screening images, perform batch automatic labeling and fix the text content and its position information to obtain multimodal expert knowledge labeling results.
[0055] In this embodiment, the multimodal expert knowledge labeling result refers to a comprehensive text description that includes tags composed of discrete words, short prompts describing the core subject and key text, and long prompts describing the lighting and environmental background.
[0056] In one embodiment, step S130 described above may include steps S131 to S134.
[0057] S131. Establish a structured labeling system based on manual verification and proprietary database.
[0058] In this embodiment, the structured labeling system includes professional knowledge labels in multiple dimensions, including style, box type, industry, material, and printing process.
[0059] Specifically, the structured labeling system includes professional knowledge labels across multiple dimensions, such as style (e.g., ink wash, realistic), box type (e.g., snap-bottom box, corrugated box), industry (e.g., food, cosmetics), material (e.g., white cardboard, specialty paper), and printing process (e.g., gravure, hot stamping). Through manual verification and the support of a professional database, it is ensured that each label accurately reflects the actual attributes of the packaging. This process lays the foundation for subsequent automated labeling.
[0060] S132. Train a multi-label classification network based on the structured label system, and use the trained multi-label classification network to automatically annotate the preliminary screening images with professional features to obtain the annotation results.
[0061] In this embodiment, the annotation result refers to the professional feature labels automatically assigned to each initially screened image by a trained multi-label classification network, covering information such as style, box type, industry, material and printing process.
[0062] A multi-label classification network is trained using a large-scale dataset incorporating the aforementioned structured labeling system. This network automatically identifies and assigns corresponding professional feature labels to each initially screened image. These labels not only describe the appearance attributes of the packaging but also contain information about its functionality and purpose. This step is the core of the entire process because it directly determines whether each image accurately reflects its intended professional characteristics.
[0063] S133. Identify and record the text content and its location information on each packaging image based on the annotation results to obtain the recording results.
[0064] In this embodiment, the recording result refers to all the text content on each packaging image and its precise location (bounding box) in the image, which is identified and recorded using OCR technology, to ensure that the spatial positioning of the text information is accurate.
[0065] Each packaging image is scanned using specialized OCR technology to extract all text content and their precise pixel-level bounding boxes within the image. This step is crucial because it ensures that even if the same text appears on different surfaces, it can be accurately located, thus preventing errors or illusions from occurring in subsequent processing of the large model.
[0066] S134. Inject the annotation results and the recording results into the multimodal large language model to guide the generation of text descriptions that strictly follow the control conditions, so as to obtain the multimodal expert knowledge labeling results.
[0067] In this embodiment, the control conditions include following the conditions of supplementing information lighting, color of image elements and environmental background, without modifying existing material, box shape and text position information; the multimodal expert knowledge labeling results include tags composed of discrete words, short prompts describing the core subject and key text, and long prompts containing complete natural language descriptions of light and shadow, environment and detail texture.
[0068] In this embodiment, the control conditions include not modifying the existing material, box shape, and text position information, but only supplementing details such as lighting, color of image elements, and environmental background.
[0069] Specifically, the final content generated by the multimodal expert knowledge labeling results includes: Labels composed of discrete words: These labels concisely describe the main features of the packaging.
[0070] Short keywords describing the core elements and key text: This section of text concisely and accurately summarizes the core elements and key text of the packaging.
[0071] Long prompts containing complete natural language descriptions of lighting effects, environment, and detailed textures: This descriptive method not only ensures the professional accuracy of the generated images but also enriches their expressiveness.
[0072] Through the above steps, a high-quality, hallucination-free, and feature-perfectly aligned "packaging image-control conditions-multimodal text" training set was constructed from the original image, providing a solid data foundation and technical support for subsequent image generation. This method significantly improves the efficiency and quality of packaging design while reducing labor costs and error rates.
[0073] In this embodiment, step S130 above, in order to enable the generated model to accurately understand the packaging structure and avoid the "illusion" of a large model, involves deep professional feature fusion and labeling of the 40,000 images selected above. This process not only ensures the quality and accuracy of the data, but also provides a solid foundation for the subsequent generation of high-quality packaging images.
[0074] Because professional tagging has a very high barrier to entry, requiring deep professional knowledge and experience, this module combines manual verification with its own database to pre-structure a multi-dimensional domain tagging system. Specifically, it includes: Style: such as ink painting, realism, etc.; Box types: such as snap-bottom boxes, corrugated boxes, etc.; Industries: Covering multiple sectors including food and cosmetics; Materials: White cardstock, specialty paper, and many other options are available; Printing processes: gravure printing, hot stamping, and other techniques.
[0075] Based on these dimensions, a high-precision database of 10,000 manually constructed "packaging images - professional knowledge" entries was used to train an image multi-label classification network. This step not only ensured the professionalism and comprehensiveness of the labeling system but also provided a reliable data foundation for automatic annotation. The network was then used to automatically label 40,000 images in batches, extracting the corresponding discrete professional labels. This method significantly improved the efficiency and accuracy of labeling while reducing labor costs.
[0076] To prevent "illusions" in text content from arising in the Large Language Model (LLM) during the labeling process—that is, the creation of non-existent text information—this embodiment employs dedicated OCR technology. This technology not only recognizes and outputs all text content on the packaging surface but also provides precise pixel-level bounding boxes (BOKs) for each segment of text. This approach uses the OCR output as the "sole basis for the true text," ensuring the authenticity and accuracy of the text content. It also solves the problem of spatial positioning when the same text appears multiple times on different surfaces, further enhancing the model's understanding ability.
[0077] Large-scale natural language fusion under constraints The professional discrete labels obtained in the previous steps, along with the text and its coordinate bounding box information acquired based on OCR, are forcibly injected into the system prompts of a multimodal large language model (such as Qwen-VL). The purpose of this is to ensure that the large model strictly adheres to the given information, without any modification or assumptions. The specific instructions are as follows: "Strictly follow the above material, box shape, and text position information; do not modify; only supplement lighting, image element colors, and environmental background based on the image." Based on this, the large model will output text descriptions in three formats: Tags: Composed of discrete words separated by commas, concisely and clearly summarizing the main features of the packaging; Short keywords: These only describe the core content and key words, facilitating quick understanding and application; Long prompts: These are complete natural language descriptions containing lighting effects, environmental details, and textures, enriching the expressiveness and realism of the images.
[0078] Using the above method, this embodiment successfully constructed a high-quality, hallucination-free, and feature-completely aligned "packaged image-control condition-multimodal text" training set of 40,000 texts, which was then used as the training input for the next module. This process yielded the following significant beneficial effects: Improve data quality: By using professional tags and OCR technology, we ensure the authenticity and accuracy of every piece of data, reducing errors and biases.
[0079] Enhanced Model Understanding: The adoption of a fully bidirectional attention interaction mechanism enables the model to understand and process multimodal information more deeply, improving the professionalism and consistency of the generated results.
[0080] Improved production efficiency: Automated processes significantly shorten the time from raw images to the final product, improving the efficiency of the entire workflow.
[0081] Reduced labor costs: By reducing reliance on manual verification and labeling, labor costs are lowered while work efficiency is improved.
[0082] In summary, this embodiment, through a series of innovative technologies and methods, not only effectively solves many problems in the prior art, but also demonstrates outstanding effects and potential in practical applications.
[0083] S140. Based on the multimodal expert knowledge labeling results, feature sequences are extracted by the encoder and noisy latent variables are constructed. The fully bidirectional interactive Transformer block performs deep interaction between text, images and control charts to generate packaging design images.
[0084] In one embodiment, step S140 described above may include steps S141 to S145.
[0085] S141. Obtain the control chart and use the multimodal expert knowledge labeling results as input information.
[0086] In this embodiment, this step mainly involves collecting necessary input data, including but not limited to: high-quality image-text pairs (tags, short prompts, long prompts) tagged with multimodal expert knowledge, and the original clear image of the target (…). ) and control charts representing the packaging structure ( This input information will serve as the basis for subsequent processing.
[0087] S142. Extract the respective features of the input information and construct noisy latent variables related to the time step to obtain the extraction results.
[0088] In this embodiment, the extraction result refers to converting the multimodal expert knowledge labeling result into a text feature sequence, encoding the initial image and control chart to obtain their respective latent variables, and generating noisy latent variables based on linear interpolation at time steps, thereby preparing a unified feature representation for subsequent processing.
[0089] Specifically, the multimodal expert knowledge labeling results are converted into text feature sequences, the initial image is encoded to obtain initial latent variables, and noisy latent variables are generated based on linear interpolation at time steps. The control chart is encoded as a latent variable to obtain the extraction results.
[0090] Using a pre-trained text encoder, natural language prompts are converted into text feature sequences. .
[0091] Initial clear image Initial latent variables are obtained by compressing to the latent space using a variational autoencoder (VAE). Meanwhile, the sampling standard Gaussian noise Based on linear optimal transport interpolation, noisy latent variables for the current time step are constructed within continuous time steps t∈[0,1]. .
[0092] Ensure control chart and They have the same resolution and obtain the control latent variables through the same VAE encoder. .
[0093] Specifically, the multimodal expert knowledge labeling results are input into a pre-trained text encoder to extract text feature sequences. : ,in, The length of the text sequence. For text feature dimensions.
[0094] Initial image The encoder input to the variational autoencoder (VAE) In the process, the initial latent variables are obtained by compressing them into the latent space. Meanwhile, the sampling standard Gaussian noise In continuous time steps (in Represents real data. (Representing pure noise), construct the noisy latent variable for the current time step based on the linear optimal transmission interpolation of flow matching. : ; The corresponding target velocity field (VectorField, i.e., time-dependent velocity field) The derivative of the derivative is constant: ; Control chart Forced retention and Using the same resolution and inputting the same VAE encoder, the control hidden variables are obtained. : .
[0095] S143. The extraction result is segmented into blocks and linearly projected to a unified hidden layer dimension. Temporal and spatially aware positional information is added to the initial image and the control map through time step embedding and two-dimensional rotational position encoding to maintain frequency domain consistency and obtain the generated result.
[0096] In this embodiment, the generated result refers to the gradual denoising and eventual generation of a target image that matches the control map by backpropagating the noisy latent variables through a pre-trained diffusion model.
[0097] First, the tensor from the previous step is flattened and linearly projected onto a unified hidden layer dimension D. Then, sinusoidal positional encoding is used to map the noisy time steps into temporal features. Furthermore, it innovatively constructs a zero-step time step tensor mapping as a control-specific time feature. Finally, to enable the one-dimensional sequence to regain two-dimensional spatial awareness, two-dimensional rotational position coding (MSRoPE) was applied to ensure consistency in the frequency domain.
[0098] In this embodiment, the tensor is flattened and linearly projected onto a unified hidden layer dimension. middle: ; ; (ensure ).
[0099] Noise addition time step Sine position encoding is mapped to time features. Innovatively, a zero-time-step tensor is forced to be constructed (i.e., (representing noise-free real data endpoints) is mapped to control-specific time features. Ensure that the control conditions are absolutely clear: ; .
[0100] To enable a one-dimensional sequence to regain two-dimensional spatial awareness, this embodiment adopts the implementation of qwenimage to generate two-dimensional absolutely aligned positional frequencies (Freqs) for the image and the control token. For feature dimensions... Divide it in half into a height dimension and a width dimension. Let the height index be... The width index is The fundamental rotational frequency is First, calculate the frequency matrices in the height and width directions: ; Concatenate to obtain the joint two-dimensional frequency vector of each token Before applying to the Query and Key matrices, discard all image scaling factors (i.e., forced scaling factors). ),ensure and Their frequency domains are absolutely consistent. Features The rotation formula after RoPE is: in, For element-wise multiplication, This is the conjugate flip of the feature in the complex field.
[0101] S144. The generated result is processed by using a pure three-stream fully bidirectional interactive Transformer block combined with adaptive layer normalization and global joint attention mechanism to achieve deep feature fusion. Feature updates are optimized by sequence segmentation and attention residual gating. A feedforward neural network with shared MLP weights is used to enhance the feature mapping of the noisy image stream and the control image stream in the common visual space to obtain the mapping result.
[0102] In this embodiment, the mapping result refers to projecting the input data into another space or dimension through a specific transformation or learning model to reveal a new representation or feature of the data.
[0103] This process includes the following key steps: Adaptive layer normalization modulation: Modulation is applied to the text, noisy image, and control graph separately.
[0104] Global Joint Attention and Position Encoding Injection: Generate query matrix Q, key matrix K, and value matrix V through independent linear projection, and concatenate the QKV of the text and joint image stream in the sequence dimension, and perform global scaling dot product attention operation.
[0105] Sequence segmentation and attention residual gating: The global attention output is segmented into three feature branches, and residual connections are performed by calculating the gating scalar at the corresponding time step.
[0106] Feedforward Neural Network with Shared Visual Space: The noisy image stream and the control image stream not only share the layer normalization weights, but also completely share the weights of the feedforward neural network, thus constraining them to be in the same visual manifold space.
[0107] Specifically, the computation flow for each TransformerBlock layer is as follows: (1) Adaptive Layer Normalization (AdaLN) Modulation: ; ; ; The modulated noisy map token and the control map token are directly concatenated along the sequence dimension, resulting in a length of [length missing]. : ; To achieve deep fusion of multimodal features, this network performs independent linear projections on the text and joint image streams, generating query, key, and value matrices. In this step, the previously calculated 2D rotational position frequency (MSRoPE) is explicitly applied to the Q and K matrices to inject 2D spatial geometric priors, while the V matrix remains unchanged. For text stream features: ; ; ;in, The projection weights are for the text stream. For the joint image stream (a concatenation of noisy and control images): ; ; ;in, The projection weights of the image stream.
[0108] After computation, before proceeding to the core attention operation, the QKV of the text and the QKV of the joint image are globally concatenated in the sequence dimension, thereby breaking the two-stream barrier and constructing a fully bidirectional attention space: ; ; ; Finally, the attention mask is set to all 1s, and the features of the text, noisy image, and control image are subjected to standard scaled dot-product attention in the global space to achieve mutual visibility and bidirectional deep understanding among the three paths: ; Based on the original length Output global attention It is divided into three independent feature branches: Then, the gate scalar is calculated at the corresponding time step to perform residual joins: ; ; ; Before entering the feedforward neural network, the three feature streams undergo a second layer normalization and modulation. The key innovation of this embodiment lies in the fact that the noisy image stream and the control image stream not only share the layer normalization weights (…). ), and also fully share the weights of the feedforward neural network ( This forces the feature maps of both to reside in the same visual manifold space. For text streams: ; ; For noisy image streams: ; ; For controlling the image flow (forcing the use of zero time step) Modulate and reuse ): ; .
[0109] Ultimately, the generated As input to the next layer of TransformerBlock Perform iterations.
[0110] S145. The control strength of the model is adjusted at different time steps by using the time-aware control LoRA mechanism, and the predicted velocity field is optimized using the loss function. A clear image is generated from the noise of the mapping result by solving ODE and decoding VAE to obtain the packaging design image.
[0111] Specifically, time-aware low-rank adaptation is introduced when calculating the attention layer and MLP projection matrix to adjust the model control strength at different time steps; mean square error is used as the loss function to optimize the predicted velocity field to approximate the target velocity field; starting from pure noise, the latent representation is obtained by inverse integration of the mapping result using the ODE solver, and then restored to the finished product image through the VAE decoder to obtain the packaging design image.
[0112] The final steps involve: Introducing time-aware low-rank adaptation: When calculating the attention layer and MLP projection matrix, time-step-aware LoRA is introduced to dynamically adjust the model control strength at different time steps.
[0113] Loss function optimization: The mean squared error is used as the loss function to optimize the predicted velocity field to approximate the target velocity field.
[0114] Reverse generation of sharp images: Using an ordinary differential equation (ODE) solver, a sharp latent representation is derived step-by-step from pure noise. And then, through the VAE decoder, it is restored to the finished image, completing the generation of the packaging design image, such as... Figure 3 As shown.
[0115] Through the detailed steps described above, this embodiment realizes a controllable generation method for packaging images based on active learning and time-aware control flow, which solves several challenges in the prior art, such as low efficiency of aesthetic marking, large model illusion problem, and insufficient understanding of control conditions, and significantly improves the quality and accuracy of generated packaging design images.
[0116] In this embodiment, in the computation attention layer and the projection matrix of MLP At this time, time-aware LoRA is introduced. The weight update formula is: ;in, For the frozen backbone weights, It is a low-rank trainable matrix. The innovation lies in controlling the intensity. For time step Functions: ;when (In the pure noise stage, the overall structure of the model needs to be determined.) Extremely large; when (At the stage of approaching realistic images, the model needs to be supplemented with details.) attenuation.
[0117] This embodiment employs flow matching as the objective function. After end-to-end forward propagation, the model outputs a predicted value for the velocity field at the current time step. The mean squared error (MSE) is used as the training optimization objective to approximate the target velocity field. Backpropagation updates the weights of LoRA and the perception layer: ; During the inference phase, the model transitions from pure noise. Departure (corresponding time step) Using an ordinary differential equation (ODE) solver, within the time interval... Inside, the velocity field predicted by a pure three-flow network. Solve by integration: ; A clear latent representation is derived iteratively. Ultimately, Input to VAE decoder In the middle, restore to the real pixel space: ; the generated finished product image It possesses both the aesthetic appeal of high-resolution packaging materials and spatial structure such as creases and edges that align with the input control sketches. Reached Absolute consistency.
[0118] In this embodiment, please refer to Figure 2 Step S140 employs an end-to-end controllable image generation architecture, specifically designed to convert text, noisy images, and control graphs into high-quality pixel-level final images. This architecture aims to address the problems encountered by existing DiT models in feature interaction and scaling processes. By introducing a pure three-stream architecture and a time-aware control LoRA method, it achieves more accurate image generation results.
[0119] Current hybrid two-stream image generation models suffer from insufficient feature interaction and are prone to shifts during scaling. To overcome these challenges, this embodiment reconstructs based on Qwen Image and proposes a novel framework aimed at improving the interaction efficiency between multimodal inputs and ensuring accurate maintenance of feature consistency and clarity across different time steps.
[0120] First, the model receives three types of raw inputs: multimodal expert knowledge labeling results, the initial target image, and a control chart representing the packaging structure. Text feature sequences are extracted by encoding the text; the raw image and control chart are encoded using a variational autoencoder (VAE) to obtain their respective latent variables, and the noisy latent variables and their velocity fields are calculated using the linear optimal transmission interpolation formula.
[0121] Serialization, zero-timestep embedding, and 2D rotational position encoding: This step includes block segmentation, projection, timestep embedding, and 2D rotational position encoding. Specifically, the all-zero tensor at time step t=0 is mapped to control-specific temporal features to ensure absolute clarity of control conditions. Simultaneously, the 2D rotational position encoding (MSRoPE) technique is employed to enable one-dimensional sequences to regain two-dimensional spatial awareness.
[0122] Pure three-stream fully bidirectional interactive Transformer blocks: Each Transformer Block implements functions such as adaptive layer normalization modulation, sequence concatenation, and global joint attention mechanisms. Through this design, the features of text, noisy images, and control images can interact in the global space, forming deep understanding and fusion.
[0123] Temporally Aware Low-Rank Adaptation (LoRA) Mechanism and Flow Matching Loss Function: A temporally aware low-rank adaptation (LoRA) mechanism is introduced to adjust the weight update strategy, and a flow matching loss function is used as the training optimization objective to ensure that the predicted velocity field is as close as possible to the real target velocity field. - ).
[0124] Final Decoding and Inference Generation Stage: In the inference stage, the model starts from pure noise, gradually deriving a clear latent representation using an ordinary differential equation solver, and then restores it to the real pixel space through a VAE decoder, generating a high-resolution, aesthetically pleasing representation that is consistent with the controlled grass. Figure 1 The final product image.
[0125] Through the innovative technologies and methods described above, this embodiment not only improves the quality of image generation but also enhances the ability to precisely control input conditions, making it suitable for various application scenarios that require high-quality image generation.
[0126] In this embodiment, to address the problem of professional data screening, a rolling active learning method based on RankNet is proposed. By calculating the confusion index, this method can extract samples from candidate graphs and perform a "1-to-5 forced loop comparison" with historical graphs, thereby effectively cleaning high-quality packaged data and improving the accuracy and efficiency of data labeling.
[0127] To address the potential illusion problem that may occur when large models are labeled in specific domains (such as packaging aesthetics), a multimodal annotation method combining specialized OCR technology and a packaging knowledge base is introduced. This method injects text coordinates (bounding boxes) and professional labels as hard constraints into a large language model (LLM) to generate accurate text-image pairs, ensuring the high quality and authenticity of the training data.
[0128] To enhance the interactivity between existing control models, a pure three-stream network architecture based on a pure two-stream foundation is proposed. This innovation enables fully bidirectional attention interaction between Image Token, Control Token, and Text Token, enhancing the mutual understanding between features of different modalities.
[0129] To address the issue of structural shifts that can easily occur during the control graph generation process, a 1:1 size alignment and zero timestep strategy is proposed. This method forces the control graph to not participate in the noise addition process, maintaining absolutely clear structural constraints at all times, thus solving the problem of easy shifts in the generated image.
[0130] Based on the generation rules of diffusion models, a time-aware Control LoRA method is proposed. By introducing a time-step-aware layer, the control strength is dynamically adjusted at different diffusion stages—strong control in the early stages to determine the structure, and weak control in the later stages to supplement details. This adaptive adjustment not only ensures high structural consistency but also shortens the training time required for the model.
[0131] Therefore, compared to general aesthetic models (such as HPSv2), the rolling active learning and "1-to-5 forced loop comparison" strategy of this embodiment significantly reduce the cost of packaging aesthetic labeling while improving accuracy. Experiments show that only 10,000 pieces of manual data are needed to achieve a 90% preference accuracy, and the labeling efficiency has increased from 250 pairs per day to 3,000 pairs. By introducing dedicated OCR position hard constraints and expert knowledge base tag injection, the errors that may occur when large models identify packaging professional attributes and text layout are effectively eliminated. This ensures that the three multimodal descriptions (tags, short prompts, and long prompts) are accurately aligned, improving data quality. Compared to existing DiT models, the proposed pure three-flow architecture significantly enhances the model's understanding and loyalty to control conditions (such as structural sketches). The fully bidirectional attention mechanism enables a deep mutual understanding between image, control chart, and text features, greatly improving the interaction effect of multimodal tokens. By employing the zero-timestep alignment and time-aware Control LoRA method, this approach not only solves the structural offset problem in existing solutions but also optimizes detail processing, ensuring extremely high structural consistency and high-quality Chinese character rendering. Furthermore, this method significantly reduces the time required for model training.
[0132] Specifically, considering the stringent requirements of packaging-related data and the weaknesses of existing control generation models such as weak interaction capabilities and susceptibility to structural shifts, this embodiment proposes a controllable packaging image generation method based on active learning data processing and time-aware control flow. First, an image feature classification model is used to filter out messy and incomplete low-quality samples. Next, a candidate pool is constructed, and a RankNet aesthetic model is trained using a rolling active learning strategy and a "1-to-5 forced loop comparison" method to select a high-quality packaging base library. This method ensures that even with a small amount of manually labeled data (e.g., 10,000 records), a preference accuracy of over 90% can be achieved in testing, significantly improving labeling efficiency. Then, the text and its precise location recognized by dedicated OCR technology are used as hard constraints and combined with a manually constructed domain-specific lexicon (labels), injected into the context of a large language model. This eliminates the illusion problem that may occur when large models label specific domains, while generating three formats of natural language descriptions containing rich details and free of illusions, thus completing the construction of a high-quality multimodal dataset. Finally, based on a pure dual-stream Transformer architecture, a three-stream fully bidirectional interactive network incorporating images, control graphs, and text was constructed. Furthermore, a time-aware Control LoRA module was introduced. This module dynamically adjusts the control strength based on the diffusion time step, ensuring the control graph's time step is 0 and its dimensions are 1:1 aligned (strong control in the early stages to determine the structure, and weak control in the later stages to refine details). This mechanism not only fundamentally solves the problems of insufficient interactive capabilities, easy structural shifts, and non-compliance with diffusion laws in existing controllable generative models, but also significantly reduces the required training time and substantially improves the spatial structural consistency of the final generated packaged images and the accuracy of the Chinese text.
[0133] In summary, the method in this embodiment effectively solves the problems of high data annotation costs and the illusion of large models in vertical domains, achieving extremely high aesthetic preference accuracy and annotation efficiency with a very small amount of data. At the same time, it optimizes the model architecture, solving several key pain points in the existing technology, which not only speeds up the training process but also improves the quality of the generated results.
[0134] The aforementioned controllable image generation method for the packaging field employs a rolling active learning strategy combined with the RankNet model to select high-aesthetic-score initial images. These high-quality images are then automatically labeled in batches, fixing the text content and its positional information to construct multimodal expert knowledge labeling results. Next, an encoder extracts feature sequences and constructs noisy latent variables. A fully bidirectional interactive Transformer block is used to promote deep interaction between text, images, and control graphs, thereby generating packaging design images. This method significantly reduces data annotation costs, improves the accuracy and efficiency of aesthetic preferences, and addresses the shortcomings of existing controllable generation models in terms of structural consistency and Chinese text support, enhancing the spatial structural consistency and Chinese text fidelity of the final generated images. By introducing a time-aware control flow and a forced control graph alignment mechanism, not only is the model's ability to control details enhanced, but the high quality and professionalism of the generated images are also ensured.
[0135] Figure 4 This is a schematic block diagram of an image controllable generation system 300 for the packaging field provided in an embodiment of the present invention. Figure 4 As shown, corresponding to the above-described image controllable generation method for the packaging field, the present invention also provides an image controllable generation system 300 for the packaging field. This image controllable generation system 300 for the packaging field includes a unit for executing the above-described image controllable generation method for the packaging field, and the system can be configured in a desktop computer, tablet computer, laptop computer, or other terminal. Specifically, please refer to... Figure 4 The image controllable generation system 300 for the packaging field includes an acquisition unit 301, an image screening unit 302, a marking unit 303, and a generation unit 304.
[0136] The acquisition unit 301 is used to acquire an initial image; the image screening unit 302 is used to screen packaging images with clear and complete subjects from the initial image using a visual model and resolution threshold, and to label images with high aesthetic scores using a rolling active learning strategy combined with RankNet to obtain the screening images; the labeling unit 303 is used to perform batch automatic labeling based on the screening images, and to fix the text content and its position information to obtain multimodal expert knowledge labeling results; the generation unit 304 is used to extract feature sequences and construct noisy latent variables based on the multimodal expert knowledge labeling results through an encoder, and to perform deep interaction between text, images and control graphs through a fully bidirectional interactive Transformer block to generate packaging design images.
[0137] In one embodiment, the image initial screening unit 302 includes: The initial screening subunit uses a resolution threshold and a visual backbone network to initially screen the initial images, retaining images with acceptable dimensions and clear subjects to obtain a basic candidate pool. The iterative processing subunit randomly selects images from the basic candidate pool, pairs each image with several other images, and labels them to form several initial comparison pairs. Based on the RankNet model, it predicts the confusion index of the remaining images in the basic candidate pool, selects images with acceptable confusion indices to establish new comparison pairs with images from the initial comparison pairs, and iterates several times to utilize a forced loop link mechanism to ensure the RankNet model stably evaluates image aesthetics. The image filtering subunit uses the RankNet model to filter out images with high aesthetic scores to obtain the initial screened images.
[0138] In one embodiment, the marking unit 303 includes: The system includes a sub-unit for establishing a structured labeling system based on manual verification and a proprietary database; a labeling sub-unit for training a multi-label classification network based on the structured labeling system, and using the trained multi-label classification network to automatically annotate the initially screened images with professional features to obtain labeling results; a recording sub-unit for recognizing and recording the text content and its location information on each packaging image based on the labeling results to obtain recording results; and a description generation sub-unit for injecting the labeling results and the recording results into a multimodal large language model to guide the generation of text descriptions that strictly follow control conditions to obtain multimodal expert knowledge labeling results.
[0139] In one embodiment, the generation unit 304 includes: The system comprises the following subunits: an input subunit for acquiring the control map and using the multimodal expert knowledge labeling results as input information; an extraction subunit for extracting the features of the input information and constructing noisy latent variables related to the time step; a location information generation subunit for segmenting the extraction results into blocks and linearly projecting them to a unified hidden layer dimension, adding time- and space-aware location information to the initial image and the control map through time step embedding and two-dimensional rotational position encoding to maintain frequency domain consistency and obtain the generated result; and a mapping subunit for applying a pure three-stream fully bidirectional interactive Tra map to the generated result. The nsformer block combines adaptive layer normalization and global joint attention mechanism to achieve deep feature fusion. It optimizes feature updates through sequence segmentation and attention residual gating, and uses a feedforward neural network with shared MLP weights to enhance the feature mapping of the noisy image stream and the control image stream in a common visual space to obtain the mapping result. The image generation subunit is used to adjust the control strength of the model at different time steps through the time-aware control LoRA mechanism, and uses a loss function to optimize the predicted velocity field. It reverse-generates a clear image from the noise of the mapping result through ODE solving and VAE decoding to obtain the packaging design image.
[0140] In one embodiment, the extraction subunit is used to convert the multimodal expert knowledge labeling result into a text feature sequence, encode the initial image to obtain initial latent variables, generate noisy latent variables based on time-step linear interpolation, and encode the control graph as latent variables to obtain the extraction result.
[0141] In one embodiment, the image generation subunit is used to introduce time-aware low-rank adaptation when calculating the attention layer and the MLP projection matrix to adjust the model control strength at different time steps; using mean square error as the loss function, the predicted velocity field is optimized to approximate the target velocity field; starting from pure noise, the latent representation is solved by inverse integration of the mapping result using the ODE solver; and the image is restored to the finished product image through the VAE decoder to obtain the packaging design image.
[0142] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned image controllable generation system 300 for the packaging field and its various units can be referred to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity, these details will not be repeated here.
[0143] The aforementioned image controllable generation system 300 for the packaging field can be implemented as a computer program, which can, for example... Figure 5 It runs on the computer device shown.
[0144] Please see Figure 5 , Figure 5This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0145] See Figure 5 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0146] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform an image controllable generation method for the packaging field.
[0147] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0148] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an image controllable generation method for the packaging field.
[0149] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0150] The processor 502 is used to run a computer program 5032 stored in a memory to implement all the steps of the image controllable generation method for the packaging field.
[0151] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0152] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0153] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform all the steps of the image controllable generation method for the packaging field.
[0154] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0155] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0156] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0157] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the system of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0158] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0159] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. Method for the controlled generation of images for the packaging field, characterized by, include: Get the initial image; The process involves using a visual model and resolution threshold to filter packaging images with clear and complete subjects from the initial images. A rolling active learning strategy combined with RankNet is used to label images with high aesthetic scores to obtain initial screening images. This includes using a resolution threshold and a visual backbone network to initially screen the initial images, retaining images with acceptable dimensions and clear subjects to obtain a basic candidate pool. Images are randomly selected from the basic candidate pool, and each image is paired with several other images, labeled to form several initial comparison pairs. The confusion index of the remaining images in the basic candidate pool is predicted based on the RankNet model. Images with a confusion index meeting the requirements are selected to form new comparison pairs with images in the initial comparison pairs. This process is iterated several times to utilize a forced loop link mechanism to ensure the RankNet model stably evaluates image aesthetics. Finally, the RankNet model is used to filter images with high aesthetic scores to obtain the initial screening images. Based on the initial screening images, batch automatic labeling is performed, and the text content and its position information are fixed to obtain multimodal expert knowledge labeling results; Based on the multimodal expert knowledge labeling results, feature sequences are extracted by an encoder and noisy latent variables are constructed. A fully bidirectional interactive Transformer block performs deep interaction between text, images, and control charts to generate packaging design images, including: obtaining control charts and using the multimodal expert knowledge labeling results as input information; The features of the input information are extracted and noisy latent variables related to the time step are constructed to obtain the extraction result. The extraction result is block-segmented and linearly projected to a unified hidden layer dimension. Temporal and spatially aware positional information is added to the initial image and the control map through time step embedding and two-dimensional rotational position encoding to maintain frequency domain consistency, thereby obtaining the generated result. The generated result is processed by using a pure three-stream fully bidirectional interactive Transformer block combined with adaptive layer normalization and global joint attention mechanism to achieve deep feature fusion. Feature updates are optimized through sequence segmentation and attention residual gating. A feedforward neural network with shared MLP weights is used to enhance the feature mapping of the noisy image stream and the control image stream in the common visual space to obtain the mapping result. The control strength of the model is adjusted at different time steps through a temporally aware control LoRA mechanism, and the predicted velocity field is optimized using a loss function. A clear image is generated from the noise of the mapping result through ODE solving and VAE decoding to obtain the packaging design image.
2. The method for the controlled generation of images for the packaging field according to claim 1, characterized in that, The RankNet model uses DINOv3-7b as the underlying feature extractor and connects a multilayer perceptron at the end of the network to map the image into a one-dimensional aesthetic scalar score. It uses a cross-entropy-based forced model to learn the comparison relationship.
3. The method for controllable generation of images for the packaging field according to claim 1, characterized in that, The step of performing batch automatic labeling based on the initial screening images, and fixing the text content and its position information to obtain multimodal expert knowledge labeling results includes: A structured tagging system was established based on manual verification and a proprietary database. A multi-label classification network is trained based on the structured labeling system, and the trained multi-label classification network is used to automatically annotate the preliminary screening images with professional features to obtain the annotation results. The annotation results are used to identify and record the text content and its location information on each packaging image to obtain the recording results; The annotation results and the recording results are injected into the multimodal large language model to guide the generation of text descriptions that strictly follow the control conditions, so as to obtain the multimodal expert knowledge labeling results.
4. The image controllable generation method for the packaging field according to claim 3, characterized in that, The structured labeling system includes professional knowledge labels across multiple dimensions, such as style, box type, industry, material, and printing process.
5. The image controllable generation method for the packaging field according to claim 3, characterized in that, The control conditions include following the conditions of supplementing information lighting, color of image elements and environmental background, without modifying existing material, box shape and text position information; the multimodal expert knowledge labeling results include labels composed of discrete words, short prompts describing the core subject and key text, and long prompts containing complete natural language descriptions of light and shadow, environment and detail texture.
6. The image controllable generation method for the packaging field according to claim 1, characterized in that, The step of extracting the respective features of the input information and constructing noisy latent variables related to the time step to obtain the extraction result includes: The multimodal expert knowledge labeling results are converted into text feature sequences, the initial image is encoded to obtain initial latent variables, and noisy latent variables are generated based on linear interpolation at time steps. The control chart is encoded as a latent variable to obtain the extraction results.
7. The image controllable generation method for the packaging field according to claim 1, characterized in that, The LoRA mechanism, which adjusts the control strength of the model at different time steps through time-aware control, and uses a loss function to optimize the predicted velocity field, generates a clear image from the noise of the mapping result through ODE solving and VAE decoding to obtain the packaging design image, including: When calculating the attention layer and the MLP projection matrix, a time-aware low-rank adaptation is introduced to adjust the model control strength at different time steps. The mean square error is used as the loss function to optimize the predicted velocity field to approximate the target velocity field. Starting from pure noise, the latent representation is obtained by inverse integration of the mapping result using the ODE solver, and then restored to the finished image by the VAE decoder to obtain the packaging design image.
8. An image controllable generation system for use in the packaging field, characterized in that, The system uses the image controllable generation method for the packaging field as described in any one of claims 1 to 7, including: The acquisition unit is used to acquire the initial image; The image screening unit is used to screen packaging images with clear and complete subjects from the initial images using a visual model and a resolution threshold, and to label images with high aesthetic scores using a rolling active learning strategy combined with RankNet to obtain the screening images. The labeling unit is used to perform batch automatic labeling based on the initial screening images and fix the text content and its position information to obtain multimodal expert knowledge labeling results; The generation unit is used to extract feature sequences and construct noisy latent variables based on the multimodal expert knowledge labeling results, and to generate packaging design images through deep interaction between text, images and control graphs by a fully bidirectional interactive Transformer block.