Image generation method, apparatus and electronic device

By extracting the object features of the target object, updating and generating product images that match the style and preferences based on published product images, the problem of AI-generated images not being able to be smoothly integrated into commercial shooting environments is solved, realizing an intelligent closed loop and improving generation efficiency and conversion rate.

CN122492875APending Publication Date: 2026-07-31ZHEJIANG TMALL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG TMALL TECH CO LTD
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing AI-generated product images cannot be smoothly integrated into commercial photography environments and lack specificity, resulting in low click-through rates. Traditional manual processes are time-consuming and costly.

Method used

By extracting the object features of the target object, updating the product image based on the published product image, and combining the pre-trained indicator evaluation model and visual template, a product image that conforms to the style and preferences of the target object is generated, thus achieving an intelligent closed loop.

Benefits of technology

The generated product images match the style and preferences of the target audience, eliminating reliance on manual processing, improving generation efficiency and quality, and increasing click-through rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492875A_ABST
    Figure CN122492875A_ABST
Patent Text Reader

Abstract

This application provides an image generation method, apparatus, and electronic device, relating to the field of image processing technology. The image generation method includes: updating a first product image displayed in a first product image based on object features of a target object and first product features corresponding to a first product image, to obtain a second product image; obtaining a target secondary product image corresponding to a secondary product paired with the second product based on object features and second product features corresponding to the second product image, and generating a target pairing image based on the second product image and the target secondary product image; determining a target visual template corresponding to the second product image based on the second product features and the object features; and generating a to-be-published image of the target object based on the target pairing image and the target visual template. This application can break down the barriers between product image design and publishing, requiring no manual intervention, and greatly improving the quality and efficiency of product image generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image generation method, apparatus and electronic device. Background Technology

[0002] The core growth driver of e-commerce platforms lies in the continuous and frequent supply of high-quality, high-conversion original new products from a massive number of merchants. Traditional product launch processes typically involve a lengthy series of steps: manual trend research, product selection / design, factory prototyping and trial and error, product photography, post-production retouching, and manual cross-platform listing. This process can take weeks or even months, incurring high trial-and-error costs and heavy reliance on manual labor. While AI-powered image generation tools are used to address this, existing tools often operate in a fragmented manner, resulting in product images that cannot be seamlessly integrated into commercial photography environments. Furthermore, AI-generated images are often generic and lack specificity, leading to high-quality but low-conversion rates. Therefore, a smart and efficient solution for generating product images is urgently needed. Summary of the Invention

[0003] This application provides an image generation method, apparatus, and electronic device to alleviate or solve one or more technical problems existing in the prior art.

[0004] In a first aspect, embodiments of this application provide an image generation method, including: Based on the object features of the target object and the first product features corresponding to the first product image, the first product displayed in the first product image is updated to obtain the second product image; the object features are determined based on the published product images of the target object. Based on the object features and the second product features corresponding to the second product image, obtain the target secondary product image corresponding to the secondary product that is paired with the second product, and generate a target pairing image based on the second product image and the target secondary product image; Based on the second product characteristics and the object characteristics, a target visual template corresponding to the second product image is determined; the target visual template includes a target human body model and / or a target display scene for displaying the second product. Based on the target image and the target visual template, generate a publishable image of the target object.

[0005] Optionally, updating the first product displayed in the first product image based on the object features of the target object and the first product features corresponding to the first product image to obtain the second product image includes: The first product is updated based on the object features and the first product features to obtain multiple candidate product images; The first indicator of each candidate product image is evaluated using a pre-trained indicator evaluation model to obtain a first evaluation result. Based on the first evaluation result, the second product image is determined from the plurality of candidate product images.

[0006] Optionally, updating the first product based on the object features and the first product features to obtain multiple candidate product images includes: Based on the published product images, obtain material information that matches the characteristics of the object; the material information includes at least one of the following: product type, product style, product visual characteristics, product-related accessory characteristics, and visual template characteristics; The first product feature is matched with the material information to obtain candidate material information that matches the first product feature; Based on the candidate material information, target elements matching the features of the first product are generated, and the first product is updated using the target elements to obtain the candidate product image.

[0007] Optionally, the step of evaluating a first indicator of each candidate product image using a pre-trained indicator evaluation model to obtain a first evaluation result includes: Based on the image metrics of the published product images, determine the metric data of the target element corresponding to each candidate product image; The indicator data, the object features, and the multiple candidate product images are input into the indicator evaluation model to obtain a first indicator value corresponding to each candidate product image; the first evaluation result includes the first indicator value.

[0008] Optionally, obtaining the target by-product image corresponding to the by-product paired with the second product based on the object features and the second product features corresponding to the second product image includes: Acquire multiple candidate by-product images that match the characteristics of the second product; Calculate the multimodal cross-fit degree among the by-product features, the second product features, and the object features corresponding to each candidate by-product image; Based on the multimodal cross-fit, at least one target candidate off-product is determined from the plurality of candidate off-product images.

[0009] Optionally, generating a target matching image based on the second product image and the target by-product image includes: Based on the target by-product image and the second product image, at least one candidate matching image is generated; The second indicator of the candidate matching image is evaluated using the indicator evaluation model to obtain a second evaluation result; Based on the second evaluation result, the target matching image is determined from the at least one candidate matching image.

[0010] Optionally, determining the target visual template corresponding to the second product image based on the second product features and the object features includes: Obtain multiple candidate visual templates that match the features of the object; Based on the template feature matching relationship corresponding to the target object, the target visual template that matches the second product feature is determined from the plurality of candidate visual templates; the template feature matching relationship is used to represent the matching relationship between the product feature of the target object and the template feature.

[0011] Optionally, the method further includes: A first visual template corresponding to the published product image is determined; the first visual template is used to represent at least one of the following: a first human body model and a first display scene corresponding to the published product of the target object; Extract the product features of the published product corresponding to the published product image, and the template features of the first visual template; the template features include at least one of the following: the model features of the first human body model, and the scene features of the first display scene; A matching relationship is constructed between the product image and the template features, which serves as the template feature matching relationship corresponding to the target object.

[0012] Optionally, generating the image to be published of the target object based on the second product image, the target by-product image, and the target visual template includes: Based on the target matching image and the target visual template, generate a first image of the second product; Determine the display features of the first image; the display features include at least one of the following: human model posture features, display scene features, product local features, and by-product local features; Based on the display features, specified elements in the first image are modified to obtain at least one second image; the specified elements include at least one of the following: the pose of the human model, the display scene, the partial style of the product, and the partial style of the derivative product; the image to be published includes the first image and the second image.

[0013] Optionally, after generating the image to be published of the target object based on the target matching image and the target visual template, the method further includes: The first image quality detection model, which is pre-trained, is used to detect inferior materials in the image to be published, and the image detection results are obtained. If the image detection result meets the preset image requirements, the image quality score of the image to be published is determined by a pre-trained second image quality detection model. The preset image requirements include at least one of the following: the confidence that the image to be published does not contain inferior elements is higher than or equal to a first confidence threshold, and the confidence that the image to be published is not a low-quality image is higher than or equal to a second confidence threshold. If the image quality score is higher than or equal to a preset score threshold, the image to be published is published.

[0014] Optionally, before updating the first product displayed in the first product image based on the object features of the target object and the first product features corresponding to the first product image to obtain the second product image, the method further includes: Based on the image metrics of the published product images, images that meet preset index conditions are obtained from the published product images and used as feature reference images of the target object; the preset index conditions include at least one of the following: image exposure rate is greater than or equal to a preset exposure rate threshold, image click rate is greater than or equal to a preset click rate threshold, and the conversion rate of the published product is greater than or equal to a preset conversion rate threshold. Determine the product image preference information of the target object; the product image preference data is used to represent the target object's preference information for product elements contained in the image to be published; Based on the product image preference information and the feature reference image, generate image labels with multiple specified dimensions; The multiple image labels are combined to generate the object features.

[0015] Secondly, embodiments of this application provide an image generation apparatus, comprising: The update module is used to update the first product displayed in the first product image according to the object features of the target object and the first product features corresponding to the first product image, so as to obtain the second product image; the object features are determined based on the published product images of the target object; The acquisition module is used to extract the second product features corresponding to the second product image and acquire a target by-product image that matches the second product features; the by-product displayed in the target by-product image is associated with the second product displayed in the second product image; The determining module is used to determine a target visual template corresponding to the second product image based on the second product features and the object features; the target visual template includes a target human body model and / or a target display scene for displaying the second product; The generation module is used to generate a publishable image of the target object based on the second product image, the target by-product image, and the target visual template.

[0016] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the methods of embodiments of this application when executing the computer program.

[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method of any one of the embodiments of this application.

[0018] Fifthly, embodiments of this application provide a computer program product, including a computer program, which, when executed by a processor, implements any of the methods described in the embodiments of this application.

[0019] According to the technical solution of this application embodiment, the first product displayed in the first product image is updated based on the object features of the target object and the first product features corresponding to the first product image to obtain a second product image; a target secondary product image corresponding to the secondary product paired with the second product is obtained based on the object features and the second product features corresponding to the second product image, and a target pairing image is generated based on the second product image and the target secondary product image; a target visual template corresponding to the second product image is determined based on the second product features and object features, the target visual template including a target human body model and / or a target display scene for displaying the second product; and a target object's image to be published is generated based on the target pairing image and the target visual template. Since the object features are determined based on the published product images of the target object, they can reflect the style and preferences of the target object's product images to a certain extent. Therefore, this application uses the object features of the target object as a data bus throughout, so that the final generated image to be published conforms to the style and preferences of the target object, ensuring the visual consistency between the image to be published and the published product images of the target object. In addition, it has achieved an intelligent and complete closed loop of "extracting object features, guiding product image updates based on object features, guiding visual template selection based on object features, and generating images to be published", breaking down the barriers between product image design and publishing, eliminating the need for manual intervention, and greatly improving the quality and efficiency of generating product images.

[0020] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0021] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments according to this application and should not be construed as limiting the scope of this application.

[0022] Figure 1 A flowchart of the image generation method provided in an embodiment of this application is shown; Figure 2 Flowcharts of image generation methods provided in other embodiments of this application are shown; Figure 3 A block diagram of an image generation apparatus provided in an embodiment of this application is shown; Figure 4 A block diagram of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0023] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the concept or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0024] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application.

[0025] The following terms will be used in the following text: High-dimensional caption extraction: refers to using a multimodal large language model to perform deep reverse semantic analysis on image materials, and accurately extract structured image parameters, such as "color = khaki, collar type = V-neck, sleeve length = long sleeve".

[0026] Store DNA: A structured feature vector. By performing multimodal analysis on images and text of historically high-converting products from a specific store, such as extracting tags covering 500 dimensions including design, styling, and lighting, a feature model representing the store's unique visual style and audience preferences is formed.

[0027] Matching Engine: A deep learning model that integrates retrieval and multi-level cross-ranking. This engine extracts features from the main product (i.e., modified image), candidate secondary product features, and "store gene" feature vectors, calculates the multimodal cross-fit score of the three, and thus accurately filters and outputs secondary product template schemes with calculable scores that conform to the specific store style from a massive pool of data.

[0028] CTR (Click-Through Rate) quantitative prediction model: A deep neural network trained on massive historical exposure and click conversion data of e-commerce, used to quantitatively evaluate the expected click probability of AI-generated product images for the target audience.

[0029] De-defect model: A computer vision risk control model trained to detect common defects in AI-generated images (such as limb deformities, physical clipping, and light and shadow paradoxes), used to replace manual methods to achieve fully automated quality interception.

[0030] This application aims to provide an image generation method that can be applied to the scenario of launching new products on e-commerce platforms. The products can be any online items with product update needs, such as clothing and daily necessities. The image generation method provided by this application can provide e-commerce platforms with an intelligent production and distribution pipeline for product elements. Relying on "store DNA" and dual review (including conversion prediction and visual interception), it achieves a high-conversion closed loop from finding best-selling products to distributing them. It not only overcomes the pain points of relying on manual labor and deviating from the store's style and tone, but also breaks down the contextual fragmentation of independent AI generation tools, solving the efficiency bottleneck of low-quality AI-generated images (such as illusions and flaws) requiring manual review.

[0031] It should be noted that the application scenarios or examples provided in the embodiments of this application are for ease of understanding, and the embodiments of this application do not specifically limit the application of the technical solutions. In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0032] The technical solution of this application and how it solves the aforementioned technical problems are described in detail below with specific embodiments. The listed specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0033] Figure 1A flowchart of the image generation method provided in an embodiment of this application is shown, such as... Figure 1 As shown, the method may include steps S101, S102, S103 and S104.

[0034] Step S101: Update the first product displayed in the first product image according to the object features of the target object and the first product features corresponding to the first product image to obtain the second product image; the object features are determined based on the published product images of the target object.

[0035] Optionally, the method for determining object features includes the following steps: First, based on the image metrics of published product images, images meeting preset metric conditions are obtained from the published product images of the target object and used as feature reference images of the target object; second, the product image preference information of the target object is determined, whereby the product image preference data represents the target object's preference information for product elements contained in the image to be published; third, based on the product image preference information and the feature reference image, multiple image tags of specified dimensions are generated; and the multiple image tags are combined to generate the object features of the target object. The preset metric conditions include at least one of the following: image exposure rate greater than or equal to a preset exposure rate threshold, image click-through rate greater than or equal to a preset click-through rate threshold, and conversion rate of published products greater than or equal to a preset conversion rate threshold.

[0036] Optionally, when generating image labels of a specified dimension based on product image preference information and feature reference images, the anchor point information of the target object can be determined by combining the product image preference information and guidance data of the target object. Then, multiple image labels of a specified dimension can be generated based on the anchor point information and feature reference images. In the clothing image generation scenario, guidance data includes the store's sales guidance data and / or R&D guidance data. Sales guidance data includes non-price-driven transaction and sales data, while R&D guidance data includes effective new product launch data. Product image preference data may include at least one of the following: clothing elements that the store expects to retain, future strategies, categories, styles, elements, and visual features to avoid.

[0037] Optionally, the feature reference image is used as positive sample material, and images that do not meet the preset index conditions are used as negative sample material. Multimodal parsing is performed on the positive sample material, negative sample material, and anchor point information to obtain multiple labels for the target object, such as user positioning labels, style proposition labels, design element labels, matching formula labels, and shooting proposition labels. Then, based on the label distribution differences between the positive and negative sample materials and the label weights corresponding to the anchor point information, object features of the target object are generated. Furthermore, the object features can be updated based on user selection feedback, image index feedback, and image posting conversion feedback.

[0038] The target object's characteristics can include image tags across multiple dimensions, such as at least one of the following: user positioning characteristics, style proposition characteristics, design element characteristics (e.g., fabric, design points), matching formula characteristics, and shooting proposition characteristics (e.g., lighting and atmosphere). These dimensions can be configured uniformly by the e-commerce platform or customized by the user. Specifically, style proposition represents the target object's aesthetic style and desired expression; design elements represent commonly used fabrics, patterns, colors, designs, accessories, and techniques; matching formulas represent common main and secondary item combinations, wearing methods, accessories, and styling highlights; and shooting propositions represent background, model, lighting, composition, camera position, and shooting type.

[0039] In apparel stores, the first product is clothing, and the image of the clothing displayed on the e-commerce platform is the first product image. After updating the clothing image, the updated clothing becomes the second product, and the image used to display the second product is the second product image. The first product characteristics are used to represent the types of elements possessed by the first product, and may include at least one of the following: sleeve length, fabric, collar type, color, etc.

[0040] Step S102: Based on the object features and the second product features corresponding to the second product image, obtain the target product image corresponding to the product paired with the second product.

[0041] In this context, the second product can be understood as the main product, while the secondary products are other categories that complement the main product. For example, in a clothing image generation scenario, the main product is the updated top, and the secondary products are bottoms, accessories, etc., that can be paired with the updated top.

[0042] Step S103: Based on the characteristics of the second product and the characteristics of the object, determine the target visual template corresponding to the image of the second product; the target visual template includes a target human body model and / or a target display scene for displaying the second product.

[0043] The second product feature is used to indicate the type of elements that the second product has, such as including at least one of the following: sleeve length, fabric, collar type, color, etc.

[0044] Optionally, a target visual template can be selected from a plurality of pre-configured visual templates. Each visual template may include a human model and / or a display scene for showcasing the product.

[0045] In the clothing image generation scenario, the human body model can be a model. The human body model can include at least one of the following information: style, body type, gender, facial features, etc. The display scenario can include at least one of the following information: outdoor shooting scenario, studio shooting scenario, lighting, composition, etc.

[0046] Step S104: Generate the target object's image to be published based on the target matching image and the target visual template.

[0047] Optionally, the target image and target visual template can be combined to obtain the image to be published.

[0048] In some embodiments, the images to be published include multiple images, which may include the same target human body model and target display scene, and the specified elements of the target human body model in the multiple images to be published are different. The specified elements may include at least one of the following: the pose of the human body model, the display scene, the product's partial style, and the accessory's partial style. When generating multiple images to be published, firstly, a first image of the second product is generated based on the target matching image and the target visual template; then, the display features of the first image are determined, which may include at least one of the following: human body model pose features, display scene features, product partial features, and accessory partial features. Afterwards, the specified elements in the first image are modified according to the display features of the first image to obtain at least one second image. The images to be published include the first image and the second image.

[0049] In the clothing image generation scenario, the first image can serve as the main image for clothing display, and the second image can serve as a set of images associated with the main image. A set of images can include one or more. Differences between the set and the main image can include: different human poses, different shot sizes, different back views, different details of clothing and / or accessories, different fabrics, different display scenarios, etc. Since the second product image is generated based on the store's genetic characteristics (i.e., object characteristics) and the first product characteristics, and the store's genetic characteristics represent the store's style and preferences, and both the target matching image and the target visual template are generated from the second product characteristics and the store's genetic characteristics, the main image and the set of images used to display clothing share the same or matching main product structure constraints, matching constraints, and store visual constraints. Furthermore, the final generated images to be published can conform to the store's style, style preferences, secondary product matching preferences, and product display visual preferences, etc.

[0050] According to the technical solution of this application embodiment, the first product displayed in the first product image is updated based on the object features of the target object and the first product features corresponding to the first product image to obtain a second product image; a target secondary product image corresponding to the secondary product paired with the second product is obtained based on the object features and the second product features corresponding to the second product image, and a target pairing image is generated based on the second product image and the target secondary product image; a target visual template corresponding to the second product image is determined based on the second product features and object features, the target visual template including a target human body model and / or a target display scene for displaying the second product; and a target object's image to be published is generated based on the target pairing image and the target visual template. Since the object features are determined based on the published product images of the target object, they can reflect the style and preferences of the target object's product images to a certain extent. Therefore, this application uses the object features of the target object as a data bus throughout, so that the final generated image to be published conforms to the style and preferences of the target object, ensuring the visual consistency between the image to be published and the published product images of the target object. In addition, it has achieved an intelligent and complete closed loop of "extracting object features, guiding product image updates based on object features, guiding visual template selection based on object features, and generating images to be published", breaking down the barriers between product image design and publishing, eliminating the need for manual intervention, and greatly improving the quality and efficiency of generating product images.

[0051] In some embodiments, updating the first product displayed in the first product image to obtain a second product image based on the object characteristics of the target object and the first product characteristics corresponding to the first product image can be performed as follows: Steps A1, A2, and A3. Step A1: Update the first product based on the object features of the target object and the features of the first product to obtain multiple candidate product images.

[0052] In step A1, firstly, based on the published product image, material information matching the object characteristics of the target object is obtained. This material information represents at least one of the following: product type, product style, product visual characteristics, associated accessory characteristics, and visual template characteristics. Then, the first product characteristic is matched with the material information to obtain candidate material information matching the first product characteristic. Next, target elements matching the first product characteristic are generated based on the candidate material information, and the first product is updated using these target elements to obtain the candidate product image.

[0053] The target object's object features include multiple material tags. When acquiring material information that matches the object features, high-weight material tags can be extracted from the target object's object features, including product type, product style (such as pattern, silhouette, collar type, sleeve type, garment length, fabric, accessories, etc.), product visual features (such as color scheme, pattern, craftsmanship, design point tags, etc.), accessory features (such as trouser length, skirt length, accessory style, etc.), and visual template features (such as mannequin style, scene shooting style, etc.). After matching the first product features with the material information, target elements that can be transferred to the first product image can be identified. These target elements constrain the updating of the first product to the object features of the target object and the first product features. In the clothing image generation scenario, these constraints can include main product structure preservation, variable design areas, replaceable design elements, color constraints, fabric constraints, craftsmanship constraints, and negative constraints.

[0054] In some embodiments, a material library for the target object can be pre-configured. Multiple material images matching the object characteristics of the target object are acquired and stored in the material library. When updating the first product, elements from the material images in the target object's material library can be selected, such as design point materials (e.g., specific prints, sleeve types, etc.), and the selected elements are then used to update the first product.

[0055] For example, in a clothing image generation scenario, each source image includes clothing that matches the object features of the target object, matching clothes and / or accessories, a mannequin, and a display scene. When updating the clothing to be updated, the clothing features of the clothing are first matched with the source images to obtain candidate source images that match the clothing features. Specifically, if the style of the clothing shown in the source image matches the style of the clothing to be updated, the source image is considered to match the clothing features. Optionally, the similarity between the clothing features and the product features (such as tops, bottoms, and accessories) shown in the source images is calculated. If the similarity reaches a preset similarity threshold, the source image is considered to match the clothing features. After obtaining the candidate source images, design point materials (i.e., target elements), such as specific prints or sleeve styles, are extracted from the candidate source images. The extracted design point materials are then used to modify the clothing to be updated, resulting in the updated clothing.

[0056] When multiple target elements are extracted, one element or a combination of multiple elements can be selected from these elements to update the first product, resulting in multiple candidate product images. For example, using the design point material "puff sleeves" to modify the top to be updated yields one updated top style; then using the design point material "specific print" to modify the top to be updated yields another updated top style. Each top style corresponds to one candidate product image.

[0057] Step A2: The first indicator of each candidate product image is evaluated using a pre-trained indicator evaluation model to obtain the first evaluation result.

[0058] When performing step A2, firstly, based on the image metrics of the published product images of the target object, the metric data of the target element corresponding to each candidate product image is determined; then, the metric data of the target element, the object features of the target object, and multiple candidate product images are input into the metric evaluation model to obtain the first metric value corresponding to each candidate product image. The first evaluation result includes the first metric value.

[0059] Image metrics for published product images can include at least one of the following: click-through rate (CTR), conversion rate, exposure rate, favorites rate, add-to-cart rate, and user feedback. When determining audience preferences based on image metrics, published images that meet preset criteria can be selected. These preset criteria can include at least one of the following: CTR reaching a preset CTR threshold, conversion rate reaching a preset conversion rate threshold, and product conversion rate reaching a preset conversion rate threshold. Published product images that meet these preset criteria can, to some extent, indicate a target audience's best-selling products.

[0060] Based on the selected published images, product features of the products displayed in those images are extracted. These product features represent the types of elements possessed by a best-selling product, and may include at least one of the following: sleeve length, fabric type, collar type, color, etc. Based on the types of elements possessed by the best-selling product, the target audience's preference information can be determined. For example, the audience preference information for a clothing store might include: pure cotton fabric, minimalist style, and light colors.

[0061] The first metric may include at least one of the following: click-through rate, conversion rate, exposure rate, collection rate, add-to-cart rate, and feedback data.

[0062] Step A3: Based on the first evaluation result, determine the second product image from multiple candidate product images.

[0063] Optionally, based on the first indicator value corresponding to each candidate product image, at least one of the following candidate product images is selected as the second product image: the first indicator value is higher than a preset indicator value threshold, or the first indicator value is among the top N values. N is a positive integer.

[0064] The indicator evaluation model can be a CTR quantitative prediction model, and the training method of the CTR quantitative prediction model can include the following steps: First, acquire the published product images and their corresponding image data; the image data includes metric data and product-related data. Metric data may include at least one of the following: exposure rate, click-through rate, favorite rate, add-to-cart rate, conversion rate, etc. Product-related data may include at least one of the following: product type, corresponding text description information, product attribute information, target object attribute information, etc.

[0065] Secondly, the published product images and their corresponding image data are combined into training data pairs. Based on these training data pairs, the CTR quantification prediction model is iteratively trained in a supervised manner. The supervision labels for the CTR quantification prediction model can be generated based on feedback data from the published product images. Feedback data may include at least one of the following: whether there was a click after exposure, click-through rate (CTR), CTR rank percentile, conversion rate, etc.

[0066] During training, the CTR quantification prediction model learns the mapping relationship between image features (including product features and indicator features) corresponding to the image data of published product images, object features of the target object, and feedback data. Then, when using the CTR quantification prediction model for indicator evaluation, the candidate product image, the image data corresponding to the target element of the candidate product image, and the object features of the target object can be input into the trained CTR quantification prediction model to output the first indicator value of the candidate product image. Here, the image data corresponding to the target element refers to the image data of the published product image corresponding to the target element.

[0067] In this embodiment, candidate product images are generated based on the target object's object type, ensuring that these images match the target object's style and preferences. This guarantees visual consistency between the final generated product image and the target object's published product images. Furthermore, a pre-trained metric evaluation model is used to evaluate the candidate product images, and a second product image is determined based on the evaluation results. This ensures that the final second product image not only meets the target object's audience preferences but also achieves high metric values, significantly improving the quality and conversion rate of the generated product images.

[0068] In some embodiments, obtaining the target byproduct image corresponding to the byproduct paired with the second product based on the object features and the second product features corresponding to the second product image can be performed as follows: Steps B1, B2, and B3. Step B1: Obtain multiple candidate byproduct images that match the characteristics of the second product.

[0069] During step B1, multiple candidate images of derivative products can be obtained from a resource library containing images that match the object features of the target object. The second product features are then matched against the images in the resource library to select the image that matches the second product features as the candidate derivative product image.

[0070] For example, in a clothing image generation scenario, each source image includes clothing that matches the object features of the target object, clothes and / or accessories that match the clothing, a mannequin, and a display scene. When matching the second product feature with source images in the source library, the similarity between the second product feature and the product features of the products displayed in the source image (such as tops, bottoms, and accessories) can be calculated. If the similarity reaches a preset similarity threshold, the source image is determined to match the second product feature, and the matched source image is identified as a candidate derivative image.

[0071] Step B2: Calculate the multimodal cross-fit degree between the by-product features, second product features, and object features corresponding to each candidate by-product image.

[0072] Step B3: Based on the multimodal cross-fit, determine at least one target product image from multiple candidate product images.

[0073] Among them, the characteristics of the candidate by-product images refer to the features of the by-products shown in the candidate by-product images. The candidate by-product images show by-products that match the product and visual templates. Optionally, the candidate by-product images may also show other products that are different from the product to be updated.

[0074] Multimodal cross-fit can be characterized as a multimodal cross-fit score, which indicates the degree of fit between the by-product features, the second product features, and the object features. In step B3, candidate by-product images with multimodal cross-fit scores higher than or equal to a preset fit threshold can be selected, and these selected candidate by-product images are determined as the target by-product images.

[0075] After determining the target byproduct image, a target matching image is generated based on the second product image and the target byproduct image. Specifically, this can be performed as follows: First, based on the target product image and the second product image, at least one candidate matching image is generated.

[0076] Secondly, the second indicator of each candidate combination image is evaluated using an indicator evaluation model to obtain the second evaluation result.

[0077] Next, based on the second evaluation results, a target outfit image is determined from at least one candidate outfit image. Each candidate outfit image corresponds to one candidate outfit scheme, which may include at least one of the following information: accessory type, accessory color, accessory style, accessory information, category matching relationship, etc. Taking the clothing image generation scenario as an example, the candidate outfit scheme may include the style and type of the clothing, the relationship between garment length / skirt length / trouser length, the amount of skin exposed on the upper and lower body, how it is worn, outfit highlights (such as layering highlights), and matching proportion information, etc.

[0078] The second metric can include at least one of the following: click-through rate, conversion rate, exposure rate, favorites rate, add-to-cart rate, and user feedback. The metric evaluation model is pre-trained based on multiple sample product images and their corresponding second metrics. Candidate combination images are obtained by combining candidate secondary product images and the second product image. For example, if the second product image contains an updated top, and the candidate secondary product image contains bottoms and other tops paired with the updated top, then when generating the combination image, the top in the candidate secondary product image can be replaced with the top in the second product image to obtain the candidate combination image.

[0079] Optionally, the indicator evaluation model for evaluating the second indicator of the candidate paired images can be the same model as the indicator evaluation model for evaluating the first indicator of the candidate product images. The indicator evaluation model can be a CTR quantization prediction model, the training method of which has been detailed in the above embodiments and will not be repeated here.

[0080] In this embodiment, a matching engine extracts features of the main product (i.e., the modified image), candidate secondary product features, and object features. The multimodal cross-fit score of these three features is calculated, allowing for precise selection from a massive material library and output of secondary product and visual template schemes that have calculable scores and conform to specific object characteristics. This not only ensures that the generated secondary product images match the style and preferences of the target object, maintaining visual consistency, but also ensures that the overall generated matching image has high index values, greatly improving the quality and conversion rate of the generated product images.

[0081] In some embodiments, determining the target visual template corresponding to the second product image based on the second product features and object features can be performed by the following steps C1 and C2: Step C1: Obtain multiple candidate visual templates that match the object's features.

[0082] When performing step C1, the target object's preference features for visual templates (hereinafter referred to as visual preference features) can be determined first based on the object's characteristics, and then multiple candidate visual templates can be determined based on the visual preference features.

[0083] Visual preference features may include human model preference features and / or display scene preference features. In the clothing image generation scene, human model preference features may include at least one of the following: head-to-body ratio features, face wrapping features, face occlusion features, posture features, etc. Display scene preference features may include at least one of the following: background features, tonal features, lighting quality features, color saturation features, commercial photography type features, shot size features, composition features, camera position features, etc.

[0084] Step C2: Based on the template feature matching relationship corresponding to the target object, determine the target visual template that matches the second product feature from multiple candidate visual templates.

[0085] The template feature matching relationship is used to represent the matching relationship between the product features of the target object and the template features. The process of constructing the template feature matching relationship includes the following steps: First, determine the first visual template corresponding to the published product image; the first visual template is used to represent at least one of the following: the first human model and the first display scene corresponding to the published product of the target object.

[0086] Secondly, extract the product features corresponding to the published product images, as well as the template features of the first visual template. Template features include at least one of the following: model features of the first human model and scene features of the first display scene. Model features may include at least one of the following: style, body type, gender, facial features, facial occlusion, posture, etc. Scene features may include at least one of the following: outdoor shooting scene, studio shooting scene, lighting, composition, landscape shooting, camera position, light quality, etc. Product features are used to represent the types of elements possessed by the product. For example, when the product is clothing, product features may include at least one of the following: sleeve length, fabric, collar type, color, etc.

[0087] Next, construct the matching relationship between the product features and template features of the published products, which will serve as the template feature matching relationship for the target object.

[0088] After the target visual template is determined, it is used to decide on the visual generation control parameters, including at least one of the following: human model control parameters, scene control parameters, lighting control parameters, composition control parameters, posture control parameters, shot size control parameters, and negative constraint parameters.

[0089] Furthermore, when determining the target visual template, user needs information can also be considered. User needs information may include at least one of the following: the style, body shape, and gender of the human model, the type of shooting location, lighting effects, etc. After obtaining multiple candidate visual templates that match the object characteristics, the target visual template that meets the user needs information can be selected from the multiple candidate visual templates; or, the target visual template can be determined from the multiple candidate visual templates by combining the second product characteristics with the user needs information.

[0090] In this embodiment, by using the product features corresponding to the updated product image (i.e. the second product image) and the object features of the target object as query conditions, the most suitable visual template is automatically matched, so that the target visual template matched for the second product image simultaneously conforms to the style and preferences of the target object and the product features of the currently updated new product, thereby improving the quality and conversion rate of the final generated product image.

[0091] In some embodiments, after generating the image to be published, a quality check is performed on the image before publication. If the quality check passes, the image is then published. Optionally, a pre-trained first image quality detection model is used to detect poor-quality material in the image to be published, obtaining an image detection result; if the image detection result meets preset image requirements, a pre-trained second image quality detection model is used to further determine the image quality score of the image to be published; if the image quality score is higher than or equal to a preset score threshold, the image is published.

[0092] The first image quality detection model may include at least one of the following: a poor-quality material recognition model, a low-quality AI image recognition model, or an image quality scoring model.

[0093] The substandard material identification model is used to identify whether there are substandard materials in the image to be published, such as missing subjects, subject occlusion, abnormal composition, insufficient clarity, incomplete clothing display, abnormal background interference, etc. The low-quality AI image identification model is used to identify whether the image to be published is a low-quality image, such as whether there are limb deformities, abnormal hands, abnormal faces, clothing clipping, disordered textures, contradictory lighting and shadows, abnormal perspective, etc. The image quality scoring model is used to score the quality of the image to be published. A higher score indicates a lower quality image. The quality score is determined based on at least one of the following: image clarity, subject integrity, clothing consistency, reasonable composition, and image realism. Preset image requirements include at least one of the following: the confidence level that the image to be published does not contain substandard elements is higher than or equal to a first confidence threshold; the confidence level that the image to be published is not a low-quality image is higher than or equal to a second confidence threshold.

[0094] The second image quality detection model may include at least one of the following: a click-through rate (CTR) quantification and prediction scoring model, an object feature matching scoring model, and an image aesthetic scoring model. The CTR quantification and prediction scoring model is used to predict the CTR of the image to be published. The object feature matching scoring model is used to score the degree of matching between the materials displayed in the image to be published (including main products, secondary products, and visual templates) and the object features of the target object. The image aesthetic scoring model is used to score the aesthetic quality of the materials displayed in the image to be published. When there are multiple scoring dimensions, the scores from these multiple dimensions can be weighted and calculated to obtain the final image quality score.

[0095] Optionally, after determining that the image detection result meets the preset image requirements and / or the image quality score is higher than or equal to the preset score threshold, the image to be published can be output to the target object's publishing pool. The publishing pool stores at least one image of the target object to be published, and each image to be published can be configured to be published at a specified time.

[0096] If the image detection result does not meet the preset image requirements and / or the image quality score is lower than the preset score threshold, the corresponding problem diagnosis information can be generated and fed back, so that the system can update the relevant parameters of image generation according to the problem diagnosis information and regenerate the image to be published.

[0097] In this embodiment, by automatically identifying image defects in the image to be published, the system can automatically review and redraw defects such as clipping and multiple fingers that may occasionally occur in AI-generated images, eliminating the need for manual review and achieving a truly fully automated product image generation effect.

[0098] In some embodiments, after generating the image to be published, feature parsing is performed on the image to extract structured product features, and these extracted features are then populated into the new product publication form of the target object. Optionally, structured product features, such as SKU (StockKeeping Unit) information and specifications, are automatically populated into the new product publication form using field mapping rules. Users can confirm the information in the publication form; after receiving confirmation from the user, the system calls the target object's interactive interface to complete the image publication.

[0099] For example, feature analysis is performed on the image to be published, and the structured product features are extracted as follows: color = khaki, collar type = round neck, sleeve type = long sleeve.

[0100] In this embodiment, by extracting the structured product features of the image to be published and automatically filling the structured product features into the new product release form for product release, a truly fully automated distribution and product placement effect is achieved.

[0101] Figure 2 Flowcharts of image generation methods provided in other embodiments of this application are shown, such as... Figure 2 As shown, the method may include steps S201 to S209.

[0102] Step S201: Based on the image metrics of the published product images of the target object, obtain images that meet the preset index conditions from the published product images and use them as feature reference images of the target object.

[0103] The preset indicator conditions include at least one of the following: image exposure rate is greater than or equal to preset exposure rate threshold, image click rate is greater than or equal to preset click rate threshold, and conversion rate of published products is greater than or equal to preset conversion rate threshold.

[0104] Step S202: Determine the product image preference information and guidance data of the target object; generate multiple image labels of specified dimensions based on the product image preference information, guidance data and feature reference images; and generate the object features of the target object based on the multiple image labels.

[0105] The process begins by determining anchor point information for the target object based on product image preference information and guidance data. Then, based on the anchor point information and feature reference images, multiple image labels with specified dimensions are generated. Guidance data includes the store's sales guidance data and / or R&D guidance data. Sales guidance data includes non-price-driven transaction and sales data, while R&D guidance data includes data on effective new product launches. Product image preference data may include at least one of the following: clothing elements the store intends to retain, future strategies, categories, styles, elements, and visual features to avoid.

[0106] In the clothing image generation scenario, the target object is clothing stores, and the published product images are the clothing images already listed in the store. Published product images that meet preset index conditions can represent clothing types with high audience appeal for the store; therefore, feature reference images can serve as "seed" images for the store. A multimodal large model can be used to perform a deep scan of the "seed" images, outputting a multi-dimensional (e.g., 500+ dimensions) tag matrix, including elements such as fabric, design points, and lighting / atmosphere. Then, the image tags of specified dimensions are aggregated to generate the store's "store DNA" features, i.e., object features. "Store DNA" features include image tags across multiple dimensions, such as user positioning features, style proposition features, design element features (e.g., fabric, design points), matching formula features, and photography proposition features (e.g., lighting / atmosphere). The "store DNA" features will be passed down as global variables to all downstream nodes.

[0107] Step S203: Based on the object features of the target object and the first product features corresponding to the first product image, update the first product displayed in the first product image to obtain multiple candidate product images; evaluate the indicators of each candidate product image through a pre-trained indicator evaluation model, and determine the second product image from the multiple candidate product images based on the evaluation results.

[0108] In the clothing image generation scenario, the first product is the clothing to be modified, the modified clothing is the second product, and the image used to display the second product is the second product image. The first product features are used to represent the element types of the clothing to be modified, and may include at least one of the following: sleeve length, fabric, collar type, color, etc.

[0109] When updating the first product, elements can be selected from the target object's material library, such as design point materials (e.g., specific prints, sleeve styles, etc.), and then the selected elements are used to update the first product. In the material library, each material image includes clothing matching the target object's features, matching garments and / or accessories, a mannequin, and a display scene. When modifying the clothing to be modified, the clothing features are first matched with the material images to obtain candidate material images that match the clothing features. Specifically, if the style shown in the material image matches the style of the clothing to be modified, the material image is considered to match the clothing features. Optionally, the similarity between the clothing features and the product features (e.g., tops, bottoms, and accessories) shown in the material image is calculated. If the similarity reaches a preset similarity threshold, the material image is considered to match the clothing features. After obtaining the candidate material images, design point materials (i.e., target elements), such as specific prints and sleeve styles, are extracted from the candidate material images, and the extracted design point materials are used to modify the clothing to obtain the modified clothing.

[0110] When multiple target elements are extracted, one element or a combination of multiple elements can be selected from these elements to modify the first product, resulting in multiple candidate product images. For example, using the design point material "puff sleeves" to modify the top, one modified top is obtained; then using the design point material "specific print" to modify the top, another modified top is obtained. Each top corresponds to one candidate product image.

[0111] The metrics include at least one of the following: click-through rate, conversion rate, exposure rate, favorites rate, add-to-cart rate, and user feedback information. Step S204: Obtain multiple candidate images of the complementary products that are paired with the second product.

[0112] Multiple candidate images of secondary products can be obtained from a resource library of the target object. The second product feature is matched with the resource images in the resource library to select the resource image that matches the second product feature as the candidate secondary product image. In the clothing image generation scenario, when matching the second product feature with the resource image, the similarity between the second product feature and the product features of the product shown in the resource image (such as tops, bottoms, and accessories) can be calculated. If the similarity reaches a preset similarity threshold, the resource image is determined to match the second product feature, and the matched resource image is identified as the candidate secondary product image.

[0113] Step S205: Calculate the multimodal cross-fit degree between the by-product features, second product features, and object features corresponding to each candidate by-product image; determine the target by-product image from multiple candidate by-product images based on the multimodal cross-fit degree.

[0114] The second product feature is used to indicate the type of elements that the modified garment has, such as at least one of the following: sleeve length, fabric, collar type, color, etc.

[0115] Step S206: Generate at least one candidate pairing image based on the second product image and the target by-product image; evaluate the second index of each candidate pairing image using an index evaluation model, and determine the target pairing image from the at least one candidate pairing image based on the evaluation results.

[0116] Candidate combination images are obtained by combining the target sub-product image and the second product image. For example, the second product image contains the updated top, and the target sub-product image contains the bottoms that match the updated top and other tops. When combining the images, the top in the target sub-product image can be replaced with the top in the second product image to obtain the candidate combination image.

[0117] Step S207: Obtain multiple candidate visual templates that match the object features, and determine the target visual template that matches the second product features from the multiple candidate visual templates based on the template feature matching relationship corresponding to the target object.

[0118] The template feature matching relationship is used to represent the matching relationship between the product features of the target object and the template features. Template features include at least one of the following: model features of the human body model and scene features of the display scene. Model features may include at least one of the following: style, body type, gender, facial features, facial occlusion, posture, etc. Scene features may include at least one of the following: outdoor shooting scene, studio shooting scene, lighting, composition, landscape shooting, camera position, light quality, etc.

[0119] Step S208: Based on the target matching image and the target visual template, generate multiple images of the target object to be published; the specified elements in each image to be published are different.

[0120] The specified elements include at least one of the following: the pose of the human model, the display scene, the partial style of the product, and the partial style of the derivative product. The images to be published include a first image and a second image. First, based on the target matching image and the target visual template, a first image of the second product is generated; then, the display features of the first image are determined, including at least one of the following: human model pose features, display scene features, partial product features, and partial derivative product features. Afterward, based on the display features of the first image, the specified elements in the first image are modified to obtain at least one second image. Optionally, the first image undergoes visual spatial analysis, and then image-to-text guidance technology is used to batch generate second images with completely consistent styles but different poses.

[0121] A pose library corresponding to the target object can be pre-configured, storing multiple human model poses that match the object features of the target object. Optionally, human model images (including various different poses) can be extracted from published product images of the target object and stored in the pose library. In addition, other human model images with high similarity to the human model images can be obtained from massive online images and stored in the pose library.

[0122] Step S209: Perform image detection on the image to be published. If the image detection is successful, output the image to be published to the target object's publishing pool for publication.

[0123] The image to be published can be inspected using a first image quality detection model and / or a second image quality detection model. The detection dimensions of the first and second image quality detection models have been described in detail in the above embodiments and will not be repeated here. In this embodiment, multiple verification stages can be configured, each stage including at least one of the following: after generating candidate product images, performing clothing structure consistency verification, quality degrading verification, and object feature matching verification on the candidate product images; after generating candidate matching images and / or the first image, performing click-through rate quantification prediction, subject integrity verification, composition verification, and low-quality AI image recognition on the candidate matching images and / or the first image; after generating the second image, performing consistency verification with the first image, posture rationality verification, clothing detail consistency verification, and image quality scoring on the second image. By configuring staged verification at different stages, candidate results that do not meet the requirements can be intercepted early, reducing subsequent invalid generation and redundant calculations.

[0124] Furthermore, an object feature matching scoring model can be used to score the matching degree between candidate product images and / or candidate combination images and object features. Taking candidate combination images as an example, the object feature matching scoring model extracts at least one of the following from the candidate combination images: user positioning features, style proposition features, design element features, combination formula features, shooting proposition features, scene features, human body model features, and composition features. Then, the extracted features are matched with the object features of the target object, and a matching score is output. Afterwards, multiple candidate combination images are sorted according to the matching scores, and the target combination image is selected based on the sorting results. If the target combination image cannot be selected based on the matching scores, for example, if the matching scores of all candidate combination images are low, the system can be triggered to regenerate the target combination image. The object feature matching scoring model can be trained based on published product images, user selection feedback information, image indicator feedback information, and consistency results of manually annotated styles.

[0125] As can be seen, the technical solution adopted in this application uses the object features of the target object as a data bus throughout, ensuring that the final generated image to be published conforms to the style and preferences of the target object, thus ensuring the visual consistency between the image to be published and the published product images. Furthermore, it achieves a complete intelligent closed loop of "extracting object features, guiding product image updates based on object features, guiding visual template selection based on object features, intelligently evaluating image defects, and automated publishing," breaking down the barriers between product image design and publishing, eliminating the need for manual intervention, and greatly improving the quality and efficiency of generated product images. In addition, by repeatedly calling the indicator evaluation model to evaluate the images, including the indicator evaluation of candidate product images and candidate combination images, it ensures that the final generated image to be published has a high indicator value, greatly improving the quality and conversion rate of the generated product images. Moreover, by automatically identifying image defects in the image to be published, it can automatically review and redraw defects such as clipping and multiple fingers that occasionally occur in AI generation, eliminating manual review work and achieving a truly fully automated product image generation effect.

[0126] Corresponding to the application scenarios and methods provided in the embodiments of this application, the embodiments of this application also provide an image generation apparatus.

[0127] Figure 3 A block diagram of an image generation apparatus provided in an embodiment of this application is shown, such as Figure 3 As shown, the image generation apparatus includes: Update module 31 is used to update the first product displayed in the first product image according to the object features of the target object and the first product features corresponding to the first product image, so as to obtain a second product image; the object features are determined based on the published product images of the target object; The acquisition module 32 is used to acquire a target secondary product image corresponding to the secondary product that is paired with the second product based on the object features and the second product features corresponding to the second product image, and to generate a target pairing image based on the second product image and the target secondary product image; The determining module 33 is used to determine the target visual template corresponding to the second product image based on the second product features and the object features; the target visual template includes a target human body model and / or a target display scene for displaying the second product; The generation module 34 is used to generate a publishable image of the target object based on the target matching image and the target visual template.

[0128] Optionally, when the updating module 31 updates the first product displayed in the first product image according to the object characteristics of the target object and the first product characteristics corresponding to the first product image to obtain the second product image, it performs the following steps: The first product is updated based on the object features and the first product features to obtain multiple candidate product images; The first indicator of each candidate product image is evaluated using a pre-trained indicator evaluation model to obtain a first evaluation result. Based on the first evaluation result, the second product image is determined from the plurality of candidate product images.

[0129] Optionally, when the update module 31 updates the first product based on the object features and the first product features to obtain multiple candidate product images, it performs the following steps: Based on the published product images, obtain material information that matches the characteristics of the object; the material information includes at least one of the following: product type, product style, product visual characteristics, product-related accessory characteristics, and visual template characteristics; The first product feature is matched with the material information to obtain candidate material information that matches the first product feature; Based on the candidate material information, target elements matching the features of the first product are generated, and the first product is updated using the target elements to obtain the candidate product image.

[0130] Optionally, when the update module 31 evaluates the first indicator of each candidate product image using a pre-trained indicator evaluation model and obtains a first evaluation result, it performs the following steps: Based on the image data of the published product images, determine the image data of the target elements corresponding to each candidate product image; the image data includes indicator data and / or product-related data. The image data of the target element, the object features, and the multiple candidate product images are input into the indicator evaluation model to obtain a first indicator value corresponding to each candidate product image; the first evaluation result includes the first indicator value.

[0131] Optionally, when the acquisition module 32 acquires the target by-product image corresponding to the by-product paired with the second product based on the object features and the second product features corresponding to the second product image, it performs the following steps: Acquire multiple candidate by-product images that match the characteristics of the second product; Calculate the multimodal cross-fit degree among the by-product features, the second product features, and the object features corresponding to each candidate by-product image; Based on the multimodal cross-fit, at least one target product image is determined from the plurality of candidate product images.

[0132] Optionally, when generating a target matching image based on the second product image and the target by-product image, the acquisition module 32 performs the following steps: Based on the target by-product image and the second product image, at least one candidate matching image is generated; The second indicator of the candidate matching image is evaluated using the indicator evaluation model to obtain a second evaluation result; Based on the second evaluation result, the target matching image is determined from the at least one candidate matching image. Optionally, when determining the target visual template corresponding to the second product image based on the second product features and the object features, the determining module 33 performs the following steps: Obtain multiple candidate visual templates that match the features of the object; Based on the template feature matching relationship corresponding to the target object, the target visual template that matches the second product feature is determined from the plurality of candidate visual templates; the template feature matching relationship is used to represent the matching relationship between the product feature of the target object and the template feature.

[0133] Optionally, the device further includes: The second determining module is used to determine the first visual template corresponding to the published product image; the first visual template is used to represent at least one of the following: a first human body model and a first display scene corresponding to the published product of the target object; The first extraction module is used to extract the product features of the published product corresponding to the published product image, and the template features of the first visual template; the template features include at least one of the following: the model features of the first human body model, and the scene features of the first display scene; A construction module is used to construct a matching relationship between the product image and the template features, which serves as the template feature matching relationship corresponding to the target object.

[0134] Optionally, when generating the image to be published of the target object based on the target matching image and the target visual template, the generation module 34 performs the following steps: Based on the target matching image and the target visual template, generate a first image of the second product; Determine the display features of the first image; the display features include at least one of the following: human model posture features, display scene features, product local features, and by-product local features; Based on the display features, specified elements in the first image are modified to obtain at least one second image; the specified elements include at least one of the following: the pose of the human model, the display scene, the partial style of the product, and the partial style of the derivative product; the image to be published includes the first image and the second image.

[0135] Optionally, the device further includes: The detection module is used to detect inferior materials in the image to be published by using a pre-trained first image quality detection model after generating the image to be published of the target object based on the target matching image and the target visual template, and to obtain the image detection result. The image scoring module is used to determine the image quality score of the image to be published by a pre-trained second image quality detection model when the image detection result meets the preset image requirements. The preset image requirements include at least one of the following: the confidence that the image to be published does not contain inferior elements is higher than or equal to a first confidence threshold, and the confidence that the image to be published is not a low-quality image is higher than or equal to a second confidence threshold. The publishing module is used to publish the image to be published when the image quality score is higher than or equal to a preset score threshold.

[0136] Optionally, the device further includes: The second acquisition module is used to, before updating the first product displayed in the first product image according to the object characteristics of the target object and the first product characteristics corresponding to the first product image to obtain the second product image, acquire an image that meets preset index conditions from the published product images according to the image indicators of the published product images, and use it as a feature reference image of the target object; the preset index conditions include at least one of the following: image exposure rate is greater than or equal to a preset exposure rate threshold, image click rate is greater than or equal to a preset click rate threshold, and the conversion rate of the published product is greater than or equal to a preset conversion rate threshold. The third determining module is used to determine the product image preference information of the target object; the product image preference data is used to represent the target object's preference information for product elements contained in the image to be published; The second generation module is used to generate multiple image labels of specified dimensions based on the product image preference information and the feature reference image; The combination module is used to combine multiple image labels to generate the object features.

[0137] According to the apparatus of this application embodiment, a second product image is obtained by updating the first product displayed in the first product image based on the object features of the target object and the first product features corresponding to the first product image; a target secondary product image corresponding to the secondary product paired with the second product is obtained based on the object features and the second product features corresponding to the second product image, and a target pairing image is generated based on the second product image and the target secondary product image; a target visual template corresponding to the second product image is determined based on the second product features and the object features, the target visual template including a target human body model and / or a target display scene for displaying the second product; and a to-be-published image of the target object is generated based on the target pairing image and the target visual template. Since the object features are determined based on the published product images of the target object, they can reflect the style and preferences of the target object's product images to a certain extent. Therefore, this application uses the object features of the target object as a data bus throughout, so that the final generated to-be-published image conforms to the style and preferences of the target object, ensuring the visual consistency between the to-be-published image of the target object and the published product images. In addition, it has achieved an intelligent and complete closed loop of "extracting object features, guiding product image updates based on object features, guiding visual template selection based on object features, and generating images to be published", breaking down the barriers between product image design and publishing, eliminating the need for manual intervention, and greatly improving the quality and efficiency of generating product images. The functions of each module in each device in the embodiments of this application can be found in the corresponding description in the above method, and they have corresponding beneficial effects, which will not be repeated here.

[0138] Figure 4This is a block diagram for implementing the electronic device provided in the embodiments of this application. Figure 4 As shown, the electronic device includes a memory 401 and a processor 402. The memory 401 stores a computer program that can run on the processor 402. When the processor 402 executes the computer program, it implements the method described in the above embodiments. The number of memories 401 and processors 402 can be one or more. In a specific implementation, the electronic device may also include a communication interface 403 for communicating with external devices and performing data exchange and transmission.

[0139] In practical implementation, if the memory 401, processor 402, and communication interface 403 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0140] Optionally, in a specific implementation, if the memory 401, processor 402 and communication interface 403 are integrated on a single chip, the memory 401, processor 402 and communication interface 403 can communicate with each other through an internal interface.

[0141] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method provided in this application.

[0142] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method provided in this application.

[0143] This application also provides a chip including a processor for calling and executing instructions stored in a memory, causing a communication device with the chip installed to perform the method provided in this application.

[0144] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.

[0145] It should be understood that the aforementioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. It is worth noting that the processor can be a processor supporting the Advanced Reduced Instruction Set Computing (ARM) architecture.

[0146] Further, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0147] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0148] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0149] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0150] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.

[0151] The logic and / or steps described in the flowchart or otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0152] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.

[0153] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.

[0154] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope described in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image generation method, comprising: Based on the object features of the target object and the first product features corresponding to the first product image, the first product displayed in the first product image is updated to obtain the second product image; the object features are determined based on the published product images of the target object. Based on the object features and the second product features corresponding to the second product image, obtain the target secondary product image corresponding to the secondary product that is paired with the second product, and generate a target pairing image based on the second product image and the target secondary product image; Based on the second product characteristics and the object characteristics, a target visual template corresponding to the second product image is determined; the target visual template includes a target human body model and / or a target display scene for displaying the second product. Based on the target image and the target visual template, generate a publishable image of the target object.

2. The method of claim 1, wherein, The step of updating the first product image displayed in the first product image to obtain the second product image based on the object features of the target object and the first product features corresponding to the first product image includes: The first product is updated based on the object features and the first product features to obtain multiple candidate product images; The first indicator of each candidate product image is evaluated using a pre-trained indicator evaluation model to obtain a first evaluation result. Based on the first evaluation result, the second product image is determined from the plurality of candidate product images.

3. The method of claim 2, wherein, The step of updating the first product based on the object features and the first product features to obtain multiple candidate product images includes: Based on the published product images, obtain material information that matches the characteristics of the object; the material information includes at least one of the following: product type, product style, product visual characteristics, product-related accessory characteristics, and visual template characteristics; The first product feature is matched with the material information to obtain candidate material information that matches the first product feature; Based on the candidate material information, target elements matching the features of the first product are generated, and the first product is updated using the target elements to obtain the candidate product image.

4. The method of claim 3, wherein, The first evaluation result is obtained by evaluating the first index of each candidate product image using a pre-trained index evaluation model, including: Based on the image data of the published product images, determine the image data of the target elements corresponding to each candidate product image; the image data includes indicator data and / or product-related data. The image data of the target element, the object features, and the multiple candidate product images are input into the indicator evaluation model to obtain a first indicator value corresponding to each candidate product image; the first evaluation result includes the first indicator value.

5. The method of claim 1, wherein, The step of obtaining the target by-product image corresponding to the by-product paired with the second product based on the object features and the second product features corresponding to the second product image includes: Acquire multiple candidate by-product images that match the characteristics of the second product; Calculate the multimodal cross-fit degree between the by-product features, the second product features, and the object features corresponding to each candidate by-product image; Based on the multimodal cross-fit, at least one target product image is determined from the plurality of candidate product images.

6. The method of claim 2, wherein, The step of generating a target matching image based on the second product image and the target by-product image includes: Based on the target by-product image and the second product image, at least one candidate combination image is generated; The second indicator of the candidate matching image is evaluated using the indicator evaluation model to obtain a second evaluation result; Based on the second evaluation result, the target matching image is determined from the at least one candidate matching image.

7. The method of claim 1, wherein, The step of determining the target visual template corresponding to the second product image based on the second product features and the object features includes: Obtain multiple candidate visual templates that match the features of the object; Based on the template feature matching relationship corresponding to the target object, the target visual template that matches the second product feature is determined from the plurality of candidate visual templates; the template feature matching relationship is used to represent the matching relationship between the product feature of the target object and the template feature.

8. The method of claim 7, wherein, Also includes: Determine the first visual template corresponding to the published product image; The first visual template is used to represent at least one of the following: a first human model and a first display scene corresponding to the published product of the target object; Extract the product features of the published product corresponding to the published product image, and the template features of the first visual template; the template features include at least one of the following: the model features of the first human body model, and the scene features of the first display scene; Construct a matching relationship between the product features and the template features, which serves as the template feature matching relationship corresponding to the target object.

9. The method of claim 1, wherein, The step of generating the target object's image to be published based on the target image and the target visual template includes: Based on the target matching image and the target visual template, generate a first image of the second product; Determine the display features of the first image; the display features include at least one of the following: human model posture features, display scene features, product local features, and by-product local features; Based on the display features, specified elements in the first image are modified to obtain at least one second image; the specified elements include at least one of the following: the pose of the human model, the display scene, the partial style of the product, and the partial style of the derivative product; the image to be published includes the first image and the second image.

10. The method of claim 1, wherein, After generating the image to be published for the target object based on the target matching image and the target visual template, the method further includes: The first image quality detection model, which is pre-trained, is used to detect inferior materials in the image to be published, and the image detection results are obtained. If the image detection result meets the preset image requirements, the image quality score of the image to be published is determined by a pre-trained second image quality detection model. The preset image requirements include at least one of the following: the confidence that the image to be published does not contain inferior elements is higher than or equal to a first confidence threshold, and the confidence that the image to be published is not a low-quality image is higher than or equal to a second confidence threshold. If the image quality score is higher than or equal to a preset score threshold, the image to be published is published.

11. The method of claim 1, wherein, Before updating the first product displayed in the first product image based on the object features of the target object and the first product features corresponding to the first product image to obtain the second product image, the method further includes: Based on the image metrics of the published product images, images that meet preset index conditions are obtained from the published product images and used as feature reference images of the target object; the preset index conditions include at least one of the following: image exposure rate is greater than or equal to a preset exposure rate threshold, image click rate is greater than or equal to a preset click rate threshold, and the conversion rate of the published product is greater than or equal to a preset conversion rate threshold. Determine the product image preference information of the target object; the product image preference data is used to represent the target object's preference information for product elements contained in the image to be published; Based on the product image preference information and the feature reference image, generate image labels with multiple specified dimensions; The multiple image labels are combined to generate the object features.

12. An image generation apparatus characterized by comprising: include: The update module is used to update the first product displayed in the first product image according to the object features of the target object and the first product features corresponding to the first product image, so as to obtain the second product image; the object features are determined based on the published product images of the target object; The acquisition module is used to extract the second product features corresponding to the second product image and acquire the target by-product image that matches the second product features; The by-product shown in the target by-product image is associated with the second product shown in the second product image; The determining module is used to determine a target visual template corresponding to the second product image based on the second product features and the object features; the target visual template includes a target human body model and / or a target display scene for displaying the second product; The generation module is used to generate a publishable image of the target object based on the second product image, the target by-product image, and the target visual template.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1 to 11.

14. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 1 to 11.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 11.