Product image generation method and system based on semantic consistency correction
Through image reverse semantic vector parsing and structured label evaluation, the problems of weak semantic control ability and low iteration efficiency in product image generation in existing technologies are solved, and efficient and personalized automatic generation of product appearance designs is achieved.
Patent Information
- Application Number
- CN202510776215.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing product image generation methods in industrial design have the problems of inaccurate semantic parsing, uncontrollable generation results, lack of consistency judgment and feedback mechanism, and lack of personalized accumulation mechanism, resulting in low iteration efficiency and insufficient personalization.
By introducing the image reverse semantic vector parsing mechanism, the consistency between image content and user intention is evaluated item by item, structured tags are used for local optimization and regeneration, and personalized accumulation is combined with user historical preferences to achieve automatic evaluation and correction of multi-dimensional consistency of image generation.
The accuracy of image generation results and the intelligence of system responses are significantly improved, achieving efficient and consistent matching of images and user intentions and personalized control.
Smart Images

Figure CN120635240A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of product design, and in particular to a product image generation and system based on semantic consistency correction. Background Art
[0002] Currently, product design relies primarily on professional designers modeling and drawing based on user needs. This process is not only time-consuming but also inefficient in human-computer interaction, making it difficult to meet the actual needs of rapid iteration and personalized design.
[0003] In recent years, with the development of diffusion models and text-to-image generation technologies such as Stable Diffusion, image synthesis has gradually entered the practical application stage, providing new possibilities for product appearance design. However, existing image generation methods still have many shortcomings, limiting their practical value in industrial design scenarios.
[0004] Specifically, current approaches suffer from the following common problems: First, imprecise semantic parsing: Most generative models only support a fuzzy understanding of natural language, making it difficult to extract structural design elements from it, resulting in significant deviations between generated images and user intent. Second, uncontrollable generated results: Existing methods have limited control over fine-grained attributes such as color, style, material, and function. In particular, during the modification and regeneration process, the system lacks a clear judgment mechanism, often requiring users to conduct multiple rounds of trial and error and manual comparison. Third, there is a lack of consistency judgment and feedback mechanisms: Current models typically generate images based on a single text input and lack a means to assess the consistency of the resulting image with the original intent. This means the system cannot automatically determine whether the results need optimization or identify the source of deviation. Fourth, there is a lack of a personalized accumulation mechanism, making it impossible to fully utilize user historical preference information to improve the consistency and personalization of generated results.
[0005] Therefore, there is an urgent need for an image generation method that can extract structured semantic information from natural language and realize automatic evaluation and correction of multi-dimensional consistency between images and semantics, so as to achieve an automated product appearance generation process with consistent images and texts, precise control and efficient iteration. Summary of the Invention
[0006] This paper proposes a method for automatically generating product appearance images based on semantic consistency vector comparison and local label updating. This method aims to address the weak semantic control, inefficient interactive iteration, and poor style consistency issues inherent in existing natural language-based product image generation methods. By introducing an image reverse semantic vector parsing mechanism, the system evaluates the consistency between image content and user intent at the structured label level, driving automated local optimization and image regeneration, significantly improving the accuracy of image representation and the intelligence of system responses.
[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions: A method for generating product images based on semantic consistency correction includes the following steps: Step 1: Based on the user's natural language description of the needs, extract structured tags that represent product design elements; for items with unclear tags, complete them through probabilistic sampling based on historical preference data; Step 2: Encode the structured label into a semantic style vector and generate an initial image based on an image generation model; Step 3: Perform reverse semantic parsing on the generated initial image to construct an image style vector, and calculate the semantic consistency score between the image style vector and the semantic style vector of the corresponding structured label item one by one; Step 4: Based on the preset semantic consistency threshold, identify the structured label items that do not meet the standard and trigger the correction strategy for the corresponding elements; Step 5: Generate a new image based on the modified features through iterative updates until the semantic consistency threshold requirement or termination condition is met; Step 6: Record the image generation results accepted by the user and dynamically update the user's personalized feature representation for personalized guidance of subsequent generation tasks.
[0008] Furthermore, the structured tags in step 1 include but are not limited to: product category, usage scenario, style, color, material, structure, size and function.
[0009] Furthermore, in step 1, the completion based on historical preference data through probability sampling includes: calculating the label according to the historical usage frequency and smoothing factor of the candidate label The sampling probability of for:
[0010] in, α >0 is the smoothing factor, is the historical usage frequency of the i-th tag, is the historical usage frequency of the jth tag, is the total number of candidate labels.
[0011] Furthermore, the semantic style vector encoding method in step 2 is: each label is independently embedded into a vector and then concatenated to generate the overall style vector :
[0012] in, Indicates the embedding of product category label content. Indicates that the product style tag content is embedded. Indicates that product function label content is embedded; Generate the initial image I0:
[0013] in, Represents the image generation step.
[0014] Furthermore, the reverse semantic parsing in step 3 is implemented through the BLIP-2 model.
[0015] Furthermore, the semantic consistency score of the image style vector and the semantic style vector of the corresponding structured label item is for:
[0016] in, and Represents text and image in i The vector representation of the labels is Represents cosine similarity.
[0017] Furthermore, the semantic consistency threshold preset in step 4 for:
[0018] in, and are the mean and standard deviation of the similarity of the corresponding elements in the training data.
[0019] Furthermore, the correction strategy for the elements in step 4 includes at least one of the following: Generate new feature descriptions based on language model reasoning; Select high-frequency tags from user history records for replacement.
[0020] Furthermore, the dynamic update of the user personalized feature representation in step 6 is achieved through a time series smoothing formula:
[0021] in, represents the updated user personalized features of the t+1th step, λ represents the smoothing coefficient, represents the user personalized features of the tth step before the update, is the final style vector accepted by the user.
[0022] In another aspect, the present invention provides a product image generation system based on semantic consistency correction, comprising: Natural language parsing module: This module extracts structured tags representing product design elements based on the user's natural language description of needs. For items with unclear tags, probabilistic sampling is used to complete them based on historical preference data. An initial image generation module is used to encode the structured label into a semantic style vector and generate an initial image based on an image generation model; Consistency assessment module: It is used to perform reverse semantic analysis on the generated initial image, construct the image style vector, and calculate the semantic consistency score between the image style vector and the semantic style vector of the corresponding structured label item one by one; Dynamic correction module: This module is used to identify structured label items that do not meet the requirements based on the preset semantic consistency threshold and trigger the correction strategy for the corresponding elements; Iterative optimization module: It is used to iteratively update and generate a new image based on the corrected features until the semantic consistency threshold requirement or termination condition is met; User preference learning module: records the image generation results accepted by the user, dynamically updates the user's personalized feature representation, and is used for personalized guidance of subsequent generation tasks.
[0023] Compared with the prior art, the present invention has the following beneficial effects: The present invention introduces image reverse style vector parsing and label-level similarity calculation mechanisms to construct a closed-loop consistency evaluation model from user natural language input to image semantic structure output. Compared with the method of using only text description similarity for overall judgment, the present invention can identify semantic deviation dimensions in a fine-grained manner, and locally trigger structured label correction and image regeneration when the similarity is insufficient, greatly improving the fit of the image generation results to the user's design intent. In addition, the present invention uses training data to set the label-level similarity threshold, realizing an adjustable automatic optimization control process with high adaptability. By generating personalized style vectors through multiple rounds of user behavior modeling, long-term style evolution and consistency migration can also be achieved, significantly improving the personalization capabilities and industrial application value of the image generation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 It is an overall flow chart of the method of the present invention.
[0026] Figure 2This is a flowchart of label extraction, completion, and Prompt construction in an embodiment of the present invention.
[0027] Figure 3 2 is a schematic diagram of similarity calculation according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below. Example 1 like Figure 1 As shown, a product image generation method based on semantic consistency correction includes the following steps: Step 1: Natural Language Parsing and Structured Tag Extraction. The user enters product design requirements in natural language. A pre-trained large-scale language model (such as the GPT series) is used to semantically understand the input text and automatically identify and extract structured tag fields related to product design. Extracted content includes, but is not limited to, product category, usage scenario, style, color, material, structure, size, function, etc. For tags not described by the user, the system completes the tags based on probability weighting from the historical preference knowledge base.
[0029] Let the candidate label set be , and its historical usage frequency is , then the corresponding sampling probability is:
[0030] in, >0 is a smoothing factor used to avoid zero probability of label sampling. The empirical value is = 1. Step 2: Style Vector Construction and Image Generation. The structured label set is processed by the semantic encoding module and converted into a style vector. This vector serves as the latent variable input in image generation models (such as Stable Diffusion) to control the generated style. Simultaneously, the system automatically generates a text description based on a predefined prompt template. This text description, along with the style vector, is input into the image generation module to generate the initial product image.
[0031] In this embodiment, the system encodes the above structured labels into a style vector v text ∈R dAs the control input for image generation, each dimension of the style vector corresponds to a label item. Each label is vectorized using the embedding method and then concatenated into the overall style vector:
[0032] Then, the system constructs the prompt text and inputs the image generation model and style vector to drive it together to generate the initial image I0:
[0033] Step 3: Image style vector decomposition and tag-level similarity calculation. The system uses the image semantic parsing module BLIP-2 to analyze the generated image, extract the tag-level sum of the image, and convert it into the image style vector. The system calculates cosine similarity between the image style vector and the user-required style vector along each tag dimension. The calculated similarity score is used to determine whether the image expression for each tag meets the user's intent.
[0034] In this embodiment, the image I0 is parsed into a set of structured semantic labels, and the image style vector v is further constructed. img ∈R d The dimension of this vector is consistent with the user semantic vector, which represents the semantic expression of the image in various design elements.
[0035] Then, the system calculates the similarity between the image and the user semantics for each label dimension. Let the label dimension set be , then the similarity of the corresponding label items is:
[0036] in, and denote the vector representations of text and image on the i-th design label respectively.
[0037] Step 4: Local Threshold Determination and Structured Label Correction. The system sets a consistency threshold for each label dimension. If the consistency score for a label dimension does not meet the threshold, the label is deemed inconsistent with the user's semantics in the image representation. The system automatically marks the dimension as "needs optimization" and triggers a local update of the structured label. Update methods include re-inference and historical label replacement.
[0038] For each label dimension L i , if its similarity score <θ i , then the image expression of this dimension is judged to be not up to standard, and the system will automatically trigger the correction of the label field and image regeneration. i It can be set according to the training data:
[0039] in, 、 are the mean and standard deviation of the similarity of the image-text samples on the i-th label dimension in the training set. If the empirical data is: style dimension =0.88, =0.05, then Label dimensions that meet this condition will trigger the system to update the corresponding items in the structured labels. If necessary, new label items can be generated based on user feedback or language model suggestions.
[0040] Step 5: Prompt reconstruction and local regeneration. Based on the corrected label field, the system reconstructs the prompt text and calls the image generation module again to perform image regeneration based on the updated style vector. Generate image:
[0041] The process is automatically iterated until the similarities of all label dimensions meet their respective thresholds. ≥θ i Or reach the maximum number of iterations N max .
[0042] If multiple dimensions fail to meet the requirements simultaneously, they can be combined and corrected before being reconstructed all at once. If only individual labels fail to meet the requirements, a local enhancement strategy can be implemented. This process can be repeated until the similarity scores of all label dimensions are no less than their corresponding thresholds, or the maximum number of iterations is reached.
[0043] Step 6: Dynamic evolution of the user style model. During multiple user generation operations, the system records the final accepted label combination and style vector and inputs it into the style evolution module for user preference modeling.
[0044] In this embodiment, the system records the label combination and image style vector finally accepted by the user in multiple rounds of interaction, models the user style, and forms a personalized style vector representation v user , used for initialization input and preference guidance of future generation tasks.
[0045] The user style vector can be updated through the following temporal smoothing method:
[0046] in, is the final style vector accepted by the user.
[0047] like Figure 2 As shown, this embodiment demonstrates specific user needs and explains the tag extraction, completion and Prompt construction processes.
[0048] Regarding user needs: "I want to design a Bluetooth speaker suitable for office use. The overall appearance should be simple and modern. It should preferably be made of metal, not too large, and support touch operation." By analyzing natural language input using a large language model, the following design tags can be extracted: Product Category: Speakers Usage scenario: Office Style: Modern and simple Color: (to be supplemented) Material: Metal Structure: Hide Button Size: Small Function: Bluetooth, touch The following element tags are stored in the design knowledge base: (speaker, office, modern and simple, (to be supplemented), metal, hidden button, small, Bluetooth + touch) For missing product design elements, labels are selected from the design knowledge base based on sampling probability. Assuming that the color matching candidate labels are pure white (f1=3), metallic gray (f2=5) and dark blue (f3=2), then according to the completion formula Calculation: pure white p1 = 4 / 14, metallic gray p2 = 6 / 14, dark blue p3 = 3 / 14.
[0049] Based on this sampling, the system selects "Metallic Gray" as the color tag. It then combines the tags extracted from user requirements with user preference tags sampled from the design knowledge base to form an overall style vector. The overall style necklace is translated into English and entered into the image generation model Prompt, resulting in the following result: "A modern minimalist-style speaker designed for office, made of metal, featuring hidden buttons and Bluetooth +touch control. It has a metallic gray color scheme and a small size. High-quality product rendering, studio lighting, clean background." like Figure 3 As shown, this embodiment shows a similarity calculation mechanism.
[0050] For the currently generated image, perform reverse semantic generation on it to obtain the reverse semantic style vector v img, and perform cosine similarity calculation on each label dimension with the demand vector used when generating the image. The similarity calculation formula is: .
[0051] Example 2 A specific embodiment of the present invention provides a product image generation system based on semantic consistency correction, comprising: Natural language parsing module: This module extracts structured tags representing product design elements based on the user's natural language description of needs. For items with unclear tags, probabilistic sampling is used to complete them based on historical preference data. An initial image generation module is used to encode the structured label into a semantic style vector and generate an initial image based on an image generation model; Consistency assessment module: It is used to perform reverse semantic analysis on the generated initial image, construct the image style vector, and calculate the semantic consistency score between the image style vector and the semantic style vector of the corresponding structured label item one by one; Dynamic correction module: This module is used to identify structured label items that do not meet the requirements based on the preset semantic consistency threshold and trigger the correction strategy for the corresponding elements; Iterative optimization module: It is used to iteratively update and generate a new image based on the corrected features until the semantic consistency threshold requirement or termination condition is met; User preference learning module: records the image generation results accepted by the user, dynamically updates the user's personalized feature representation, and is used for personalized guidance of subsequent generation tasks.
[0052] The above is only a preferred specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed in this application should be covered by the scope of protection of the present application.
[0053] It should be understood that parts not elaborated in detail in this specification belong to the prior art.
[0054] It should be understood that the above description of the preferred embodiment is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.
Claims
1. A product image generation method based on semantic consistency correction, characterized in that: The following steps are involved: Step 1: Based on the user's natural language description of the needs, extract structured tags that represent product design elements; for items with unclear tags, complete them through probabilistic sampling based on historical preference data; Step 2: Encode the structured label into a semantic style vector and generate an initial image based on an image generation model; Step 3: Perform reverse semantic parsing on the generated initial image to construct an image style vector, and calculate the semantic consistency score between the image style vector and the semantic style vector of the corresponding structured label item one by one; Step 4: Based on the preset semantic consistency threshold, identify the structured label items that do not meet the standard and trigger the correction strategy for the corresponding elements; Step 5: Generate a new image based on the modified features through iterative updates until the semantic consistency threshold requirement or termination condition is met; Step 6: Record the image generation results accepted by the user and dynamically update the user's personalized feature representation for personalized guidance of subsequent generation tasks.
2. The method for generating product images based on semantic consistency correction according to claim 1, characterized in that: The structured tags in step 1 include but are not limited to: product category, usage scenario, style, color, material, structure, size and function.
3. The method for generating product images based on semantic consistency correction according to claim 1, characterized in that: In step 1, the completion based on historical preference data through probability sampling includes: calculating the label according to the historical usage frequency and smoothing factor of the candidate label The sampling probability of for: in, α >0 is the smoothing factor, For the i The historical usage frequency of tags, For the j The historical usage frequency of tags, is the total number of candidate labels.
4. The method for generating product images based on semantic consistency correction according to claim 1, characterized in that: The semantic style vector encoding method described in step 2 is: each label is embedded into a vector independently and then concatenated to generate the overall style vector : in, Indicates the embedding of product category label content. Indicates that the product style tag content is embedded. Indicates that product function label content is embedded; Generate the initial image I0: in, Represents the image generation step.
5. The method for generating product images based on semantic consistency correction according to claim 1, characterized in that: The reverse semantic parsing in step 3 is implemented through the BLIP-2 model.
6. The method for generating product images based on semantic consistency correction according to claim 1, characterized in that: Semantic consistency score between the image style vector and the semantic style vector of the corresponding structured label item for: in, and Represents text and image in i The vector representation of the labels is Represents cosine similarity.
7. The method for generating product images based on semantic consistency correction according to claim 1, characterized in that: The semantic consistency threshold preset in step 4 for: in, and are the mean and standard deviation of the similarity of the corresponding elements in the training data.
8. The method for generating product images based on semantic consistency correction according to claim 1, characterized in that: The correction strategy for the elements described in step 4 includes at least one of the following: Generate new feature descriptions based on language model reasoning; Select high-frequency tags from user history records for replacement.
9. The method for generating product images based on semantic consistency correction according to claim 1, characterized in that: The dynamic update of the user personalized feature representation described in step 6 is achieved through the time series smoothing formula: in, represents the updated user personalized features of the t+1th step, λ represents the smoothing coefficient, represents the user personalized features of the tth step before the update, is the final style vector accepted by the user.
10. A product image generation system based on semantic consistency correction, characterized in that: include: Natural language parsing module: This module is used to extract structured tags representing product design elements based on the user's natural language description of the requirements; For items with unspecified labels, they are completed through probabilistic sampling based on historical preference data; An initial image generation module is used to encode the structured label into a semantic style vector and generate an initial image based on an image generation model; Consistency assessment module: It is used to perform reverse semantic analysis on the generated initial image, construct the image style vector, and calculate the semantic consistency score between the image style vector and the semantic style vector of the corresponding structured label item one by one; Dynamic correction module: This module is used to identify structured label items that do not meet the requirements based on the preset semantic consistency threshold and trigger the correction strategy for the corresponding elements; Iterative optimization module: It is used to iteratively update and generate a new image based on the corrected features until the semantic consistency threshold requirement or termination condition is met; User preference learning module: records the image generation results accepted by the user, dynamically updates the user's personalized feature representation, and provides personalized guidance for subsequent generation tasks; The product image generation system based on semantic consistency correction is used to execute the steps of the product image generation method based on semantic consistency correction according to any one of claims 1 to 9.