Commodity picture processing method, medium, computer equipment and program product
By obtaining the attribute information of the subject object in the product picture, using the large language model and the image generation model to generate and replace the subject object in the scene picture, the problem of high manual processing costs in the prior art is solved, and the scene pictures with consistent characteristics is automatically generated.
Patent Information
- Application Number
- CN202510121450.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-20
Smart Images

Figure CN120182431A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image processing, and in particular, to a method, medium, computer device, and program product for processing product images. Background Art
[0002] In e-commerce platforms, the visual effects of product images play a crucial role in attracting users' attention. High-quality and beautiful product images can quickly attract users' attention and enhance users' stickiness to e-commerce platforms. Therefore, it is necessary to optimize the product images on e-commerce platforms. In related technologies, open-source image processing models are usually used to process the product images on e-commerce platforms. Since the open-source image processing models are general-purpose models, the processing targets do not meet the requirements of e-commerce scenarios. Therefore, the product images processed based on the open-source image processing models often need to be further processed manually, resulting in high costs. Summary of the Invention
[0003] In a first aspect, an embodiment of this application provides a method for processing product images. The method includes: for any one of at least one target product image, obtaining attribute information of a main object in the target product image, where the attribute information includes contour information of the main object; generating scene description information of a target scene that matches the attribute information through a pre-trained large language model; generating a scene image through an image generation model, where the scene image includes the target scene described by the scene description information and a main object having a contour described by the contour information; and replacing the main object in the scene image with the main object in the target product image.
[0004] In the embodiment of this application, based on the large language model, scene description information that matches the attribute information of the main object in the target product image is generated, and based on the scene description information, a scene image is generated. Since the attribute information of the main object includes the contour information of the main object, the contour of the main object in the generated scene image is consistent with the contour of the main object in the target product image. In this way, the main object in the scene image can be replaced with the main object in the target product image, so that the features of the main object in the scene image are consistent with those of the main object in the target product image. Through the above method, a scene image of the main object with consistent features with the target product image can be automatically generated without manual synthesis, reducing the cost of the scene image synthesis process.
[0005] In a second aspect, an embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of this application is implemented.
[0006] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any embodiment of the present application is implemented.
[0007] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program which, when executed by a processor, implements the method described in any embodiment of the present application.
[0008] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Description of the Drawings
[0009] The drawings here are incorporated into the specification and form a part of the present application. These drawings show embodiments in line with the present application and, together with the specification, are used to explain the technical solutions of the present application.
[0010] Figure 1 is a flowchart of the method for processing product pictures in an embodiment of the present application.
[0011] Figure 2 is a schematic structural diagram of the second sub-model in an embodiment of the present application.
[0012] Figure 3 is an overall flowchart of the process for screening product pictures in an embodiment of the present application.
[0013] Figure 4 is a flowchart of the method for processing product pictures in another embodiment of the present application.
[0014] Figure 5 is an overall flowchart of the process for synthesizing scene pictures in an embodiment of the present application.
[0015] Figure 6 is a schematic diagram of the target product picture and the scene picture in an embodiment of the present application.
[0016] Figure 7 is a block diagram of the device for processing product pictures in an embodiment of the present application.
[0017] Figure 8 is a block diagram of the device for processing product pictures in another embodiment of the present application.
[0018] Figure 9 is a schematic diagram of the computer device in an embodiment of the present application. Detailed Embodiments
[0019] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0020] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Additionally, the term "at least one" as used herein represents any one of a plurality or any combination of at least two of a plurality.
[0021] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0022] In order to enable those skilled in the art of this technology to better understand the technical solutions in the embodiments of the present application and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more apparent and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0023] In the e-commerce scenario, some target product images will be selected from a large number of product images (hereinafter referred to as candidate product images) on the e-commerce platform for optimization. By optimizing the product images, the attractiveness of the product images to users can be improved, thereby enhancing the stickiness of users to the e-commerce platform. In the related art, an open-source image quality evaluation model is usually used to select images with higher quality from the numerous product images on the e-commerce platform. This screening method mainly screens images from an aesthetic perspective, and the optimization goal does not meet the requirements of the e-commerce scenario. Therefore, it is necessary to further manually screen the product images selected by the image quality evaluation model, resulting in a high cost.
[0024] Based on this, an embodiment of the present application proposes a screening scheme for product images that combines the traffic efficiency and aesthetic effect of product images. Among them, the traffic efficiency can effectively evaluate the attractiveness and dissemination effect of candidate product images on the e-commerce platform, and the image quality can reflect the aesthetic effect of the product images. Therefore, the present application can automatically screen out target product images with good aesthetic effects and meeting the requirements of the e-commerce platform, without manual screening, improving the image screening efficiency. The implementation details of the present application will be illustrated with reference to the accompanying drawings below.
[0025] Referring to Figure 1 , the present application provides a method for processing product images, the method comprising:
[0026] Step S12: Obtain multiple candidate product images of the e-commerce platform;
[0027] Step S14: Obtain the traffic parameters of the multiple candidate product images, and determine the traffic efficiency of the multiple candidate product images based on the traffic parameters of the multiple candidate product images; wherein, the traffic parameter of a candidate product image is used to characterize the access traffic of the user of the e-commerce platform to the candidate product image, and the traffic efficiency of a candidate product image is used to characterize the access traffic obtained by the candidate product image under a preset product image display cost;
[0028] Step S16: Obtain the quality of the multiple candidate product images through a pre-trained image quality evaluation model;
[0029] Step S18: Screen out at least one target product image from the multiple candidate product images based on the traffic efficiency and quality of the multiple candidate product images.
[0030] In step S12, all the full-scale product images of the e-commerce platform can be used as candidate product images. Or, the full-scale product images of the e-commerce platform can be preliminarily screened, and the preliminarily screened product images can be used as candidate product images. Among them, the preliminary screening can be implemented manually or automatically by software. For example, several candidate product images can be randomly selected from the full-scale product images of the e-commerce platform. Or, the images of products of a specific category on the e-commerce platform can be used as candidate product images. Or, the product images on the e-commerce platform that conform to a specific theme can be used as candidate product images. Other methods can also be used to obtain candidate product images, and the specific methods for obtaining candidate product images will not be listed one by one here.
[0031] In step S14, the traffic parameters of each candidate product image can be obtained. Among them, the traffic parameters of the candidate product image are used to characterize the access traffic of users on the e-commerce platform to the candidate product image, including but not limited to the access traffic brought by browsing the product image, the access traffic brought by the user clicking on the product image, the access traffic brought by the user purchasing the product corresponding to the product image, the access traffic brought by the user favoriting the product corresponding to the product image, and / or the access traffic brought by the user sharing the product corresponding to the product image.
[0032] In some embodiments, the types of traffic parameters can be greater than or equal to 1. In embodiments where the types of traffic parameters are greater than 1, the traffic parameters include traffic parameters at multiple stages on the product consumption link. Among them, the multiple stages on the product consumption link can include the traffic input stage and the traffic output stage. In the traffic input stage, the merchant needs to invest a certain cost to display the product image. Therefore, the traffic parameters in the traffic input stage can include the product image display cost. In the traffic output stage, the product image can be obtained by the user, thus bringing traffic to the merchant. The traffic output stage usually includes stages such as exposure, click, and purchase. Therefore, the traffic parameters in the traffic output stage include but are not limited to at least some of the exposure rate, click-through rate, conversion rate of the candidate product image, and the transaction amount of the product corresponding to the candidate product image. Among them, the exposure rate of any candidate product image can be determined by the ratio of the exposure volume of the candidate product image to the maximum exposure volume of each candidate product image. For example, assume there are 3 candidate product images, denoted as Image A, Image B, and Image C respectively. Among them, the exposure volume of Image A is 32, the exposure volume of Image B is 50, and the exposure volume of Image C is 33. Then the maximum exposure volume of the candidate product image is 50. Therefore, the exposure rate of Image A is 64%, the exposure rate of Image B is 100%, and the exposure rate of Image C is 66%. The click-through rate of any candidate product image can be determined by the ratio of the click volume of the candidate product image to the exposure volume of the candidate product image. For example, assume the click volume of Image B above is 20, then the click-through rate of Image B is 40%. The conversion rate of any candidate product image can be determined by the ratio of the purchase volume of the candidate product image to the exposure volume of the candidate product image. For example, assume the purchase volume of Image B above is 10, then the purchase rate of Image B is 20%. The transaction amount of the product corresponding to the candidate product image is the actual amount paid by the user when purchasing the product corresponding to the candidate product image.
[0033] It can be understood that the above is only an exemplary illustration and is not used to limit this application. In other examples, the traffic parameters can also include other types of parameters, and the calculation methods of the above various traffic parameters are not limited to the calculation methods described in the above embodiments.
[0034] After obtaining the traffic parameters of the candidate product pictures, the traffic efficiency of the candidate product pictures can be determined based on the traffic parameters of the candidate product pictures. The traffic efficiency refers to the contribution that can be provided for the achievement of the service target in the traffic, and can represent the access traffic obtained by the candidate product pictures under the preset product picture display cost. The traffic efficiency is positively correlated with the traffic parameters in the traffic output stage and negatively correlated with the traffic parameters in the traffic input stage. For example, the traffic efficiency can be determined based on the ratio between the traffic parameters in the traffic output stage and the traffic parameters in the traffic input stage. Further, when there are two or more traffic parameters in the traffic output stage, the product of multiple traffic parameters in the traffic output stage can be determined, and the traffic efficiency can be determined based on the ratio between the above product and the traffic parameters in the traffic input stage.
[0035] In an embodiment where the types of traffic parameters are greater than 1, adjustment coefficients corresponding to multiple stages can also be obtained. The adjustment coefficient corresponding to any stage is used to adjust the contribution degree of the traffic parameters in this stage to the traffic efficiency. The traffic parameters of the candidate product pictures in the corresponding stage are adjusted based on the adjustment coefficients corresponding to the multiple stages respectively, and the traffic efficiency of the multiple candidate product pictures is determined based on the product picture display cost and the adjusted traffic parameters of the multiple stages.
[0036] For example, following the previous example, when the traffic parameters of the candidate product pictures in multiple stages include the product picture display cost, exposure rate, click-through rate, conversion rate of the candidate product pictures, and the transaction amount of the product corresponding to the candidate product pictures, correspondingly, the adjustment coefficients corresponding to the multiple stages respectively include a cost adjustment coefficient, an exposure rate adjustment coefficient, a click-through rate adjustment coefficient, a conversion rate adjustment coefficient, and a transaction amount adjustment coefficient.
[0037] By obtaining the adjustment coefficients corresponding to each stage, the contribution degree of the traffic parameters in each stage to the traffic efficiency can be adjusted. For example, when the adjustment coefficient corresponding to the product exposure stage is large and the adjustment coefficient corresponding to the product click stage is small, the contribution degree of the traffic parameters (i.e., the exposure rate) in the product exposure stage to the traffic efficiency is large, while the contribution degree of the traffic parameters (i.e., the click-through rate) in the product click stage to the traffic efficiency is small.
[0038] In some embodiments, the product traffic of the candidate product pictures can be recorded as:
[0039]
[0040] Among them, value represents the traffic efficiency, ER represents the exposure rate, ctr represents the click-through rate, cvr represents the conversion rate, gmv represents the transaction amount, cpm is the cost per thousand impressions, that is, the fee that the merchant needs to pay to display the product picture to a thousand users, that is, the cost of displaying the product picture. x1, x2, x3, x4, and x5 respectively represent the exposure rate adjustment coefficient, click-through rate adjustment coefficient, conversion rate adjustment coefficient, transaction amount adjustment coefficient, and cost adjustment coefficient.
[0041] In some embodiments, the adjustment coefficients corresponding to multiple stages can be determined in the following manner: Obtain the traffic parameters of the multiple stages in the first time period of the candidate product picture within a preset period; Based on the initial adjustment coefficients corresponding to the multiple stages and the traffic parameters of the multiple stages in the first time period, determine the traffic efficiency of the candidate product picture within the preset period; Based on the traffic efficiency of the candidate product picture within the preset period and the initial adjustment coefficients corresponding to the multiple stages, determine the estimated traffic parameters of the multiple stages in the second time period of the candidate product picture within the preset period; Based on the difference between the estimated traffic parameters of the multiple stages in the second time period of the candidate product picture and the traffic parameters of the corresponding stages of the candidate product picture in the second time period, adjust the initial adjustment coefficients corresponding to the multiple stages to obtain the adjustment coefficients corresponding to the multiple stages.
[0042] For example, 30 days can be determined as a cycle. The first 25 days within the 30 days are determined as the first time period, and the last 5 days within the 30 days are determined as the second time period. Taking the calculation of flow efficiency through the above formula as an example, the initial values of x1, x2, x3, x4, and x5 (i.e., the initial adjustment coefficients corresponding to multiple stages) can be set first, and the flow parameters of each day in the first 25 days are substituted into the formula to obtain the flow efficiency corresponding to each of the first 25 days. The average value of the flow efficiency corresponding to each of the first 25 days is obtained to get the flow efficiency of the corresponding cycle. Since the last 5 days and the first 25 days belong to the same cycle, it can be considered that the flow efficiency of the last 5 days should be the same as the flow efficiency of this cycle calculated based on the above method. Substituting the initial values of x1, x2, x3, x4, and x5 and the flow efficiency of this cycle into the formula, the estimated flow parameters of the last 5 days can be obtained. The difference between the actual flow parameter of each day in the last 5 days and the above estimated flow parameter can be obtained, and the sum of the differences corresponding to the last 5 days is calculated to get the total difference corresponding to the last 5 days. The above total difference can reflect the estimation deviation caused by inaccurate setting of the initial values of x1, x2, x3, x4, and x5. Therefore, the initial values of x1, x2, x3, x4, and x5 can be adjusted based on the above total difference. Multiple iterations can be performed, and the above process is repeated each time until the iteration stop condition is met, such as the number of iterations reaches the preset upper limit, or the above total difference is less than the preset threshold, etc., and the x1, x2, x3, x4, and x5 obtained when the iteration stop condition is met are determined as the adjustment coefficients corresponding to the multiple stages respectively.
[0043] In step S16, the quality of the multiple candidate product pictures can be obtained by a pre-trained picture quality assessment model. In some embodiments, the picture quality assessment model can evaluate the quality of the candidate product pictures from multiple dimensions. Specifically, the picture quality assessment model can include multiple first sub-models, and different first sub-models are used to evaluate the quality scores of the candidate product pictures in different dimensions. The above multiple dimensions include but are not limited to clarity, color, composition, and / or authenticity, etc. For any one candidate product picture, each first sub-model can output a quality score for the candidate product picture, and the quality score is positively correlated with the quality of the candidate product picture in the corresponding dimension.
[0044] In other embodiments, the picture quality assessment model can include a second sub-model, and the second sub-model can evaluate the comprehensive quality score of the candidate product picture. As a way of implementation, the second sub-model can obtain the candidate product picture and the quality scores of the candidate product picture in the multiple dimensions (i.e., the quality scores output by each first sub-model), and obtain the comprehensive quality score of the candidate product picture based on the candidate product picture and the quality scores of the candidate product picture in the multiple dimensions. The quality of the candidate product picture can be determined based on the above comprehensive quality score.
[0045] Figure 2 Shows the structure of the second sub-model in some embodiments. The quality scores output by the first sub-model may include a quality score for evaluating the basic picture quality and a quality score for evaluating the picture display effect. Among them, the basic picture quality may include, but is not limited to, the blurriness, noise intensity, and / or compression intensity of the picture. The display effect can be obtained by judging the size of the main object, performing edge detection on the main object, obtaining the vulgarity score of the main object, and / or obtaining the aesthetic score of the main object, etc. Each quality score can be converted into an embedding representation through vectorization processing, and these embedding representations can be input into a multi-layer perceptron. The main object in the product picture refers to the most important and core product or item in the picture, usually the object that consumers are concerned about and purchase. In scenarios such as e-commerce platforms, advertising, and product displays, the main object of the product picture is the core product or service presented to customers through images.
[0046] The second sub-model can be implemented based on architectures such as ResNet or VIT. First, obtain the embedding representation corresponding to the candidate product picture, and then input the embedding representation corresponding to the candidate product picture into a multi-layer perceptron. It should be noted that although two multi-layer perceptrons are shown in the figure, in actual applications, only one multi-layer perceptron can be used. That is, the embedding representations corresponding to the quality scores in each dimension and the embedding representation of the candidate product picture can be output to the same multi-layer perceptron. The multi-layer perceptron can perform feature extraction on the embedding representations corresponding to the quality scores in each dimension and the embedding representation of the candidate product picture, and output the extracted features to a splicing module for splicing (concat). The spliced features can be output to a fusion module for feature fusion. In order to better fuse the features corresponding to the embedding representation of the picture and the features corresponding to the atomic features, the fusion module can be implemented based on the architecture of DCN_V2 (Deep&Cross Network). The fused features can be output to a classifier, and the classifier can output the comprehensive score of the candidate product picture based on the fused features.
[0047] The embodiments of the present application can comprehensively evaluate the quality of product pictures in multiple dimensions based on atomic features in multiple dimensions, and perform quality evaluation based on the overall candidate product picture, so as to evaluate the visual effect of the candidate product picture macroscopically, and thus obtain accurate and comprehensive quality evaluation results.
[0048] In addition, the comprehensive quality score and / or the quality scores in each dimension of the product picture obtained by the picture quality evaluation model can also be shown to the merchant for reference. If the score of the product picture is too low, the merchant can modify or re-upload the product picture.
[0049] In step S18, at least one target product image can be screened out from the multiple candidate product images based on the traffic efficiency and quality of the multiple candidate product images.
[0050] It should be noted that step S16 can be executed after step S14. For example, the traffic efficiency of each candidate product image can be obtained first based on step S14, and then the quality of several candidate product images with traffic efficiency higher than the preset efficiency threshold can be obtained through step S16, and at least one target product image can be screened out from the several candidate product images based on the quality of the several candidate product images. Or, step S16 can be executed before step S14. For example, the quality of each candidate product image can be obtained first through step S16, and then the traffic efficiency of several candidate product images with quality higher than the preset quality threshold can be obtained based on step S14, and at least one target product image can be screened out from the several candidate product images based on the traffic efficiency of the several candidate product images. Or, step S14 and step S16 can be executed in parallel, that is, the traffic efficiency and quality of each candidate product image can be obtained in parallel, and at least one target product image can be screened out from each candidate product image based on the traffic efficiency and quality of each candidate product image. Since the quality of the product image is obtained based on the image quality evaluation model, and the model operation requires a lot of resources, therefore, obtaining the traffic efficiency of each candidate product image first through step S14, and then obtaining the quality of several candidate product images with traffic efficiency higher than the preset efficiency threshold through step S16, this method can effectively reduce the number of candidate product images that the image quality evaluation model needs to process, thereby reducing the resource consumption in the image screening process.
[0051] In some embodiments, in addition to screening the target product image based on the traffic efficiency and quality of the candidate product image, the target product image can also be further screened according to the main object in the candidate product image. Specifically, at least one candidate target product image can be screened out from the multiple candidate product images based on the traffic efficiency and quality of the multiple candidate product images, and at least one target product image whose main object meets the preset conditions can be screened out from the at least one candidate target product image. Among them, the preset conditions can be determined based on at least one of the following: the edge complexity of the main object, the number of holes on the main object, the area of the main object, and the overlap degree between the main object and the character.
[0052] Edge complexity
[0053] Edge complexity refers to the degree of complexity of the boundary shape of the main object. Edge complexity is usually determined by factors such as the irregularity of the edge contour of the main object, the number of curves, and the degree of bending. If the edge of the main object presents many zigzag and irregular lines, then its edge complexity is high; while if the edge of the main object is a straight line or has few bends, the edge complexity is low. In some embodiments, the edge complexity of the main object can be determined based on the ratio of the perimeter of the edge contour of the main object to the area of the main object. In some embodiments, the preset conditions may include that the edge complexity of the main object is lower than a preset edge complexity threshold.
[0054] Number of holes
[0055] The number of holes refers to the number of voids or openings existing on the surface of the main object. The number of holes can reflect the integrity of the main object. The more the number of holes, the lower the integrity of the main object. In some embodiments, the preset conditions may include that the number of holes on the main object is less than a preset number threshold.
[0056] Area of the main object
[0057] The area of the main object can reflect the size of the space occupied by the main object in the candidate target commodity picture. If the area of the main object is too small, then the main object is usually not prominent enough in the commodity picture, difficult to attract user attention, and the optimization effect is usually not good enough. In some embodiments, the preset conditions may include that the area of the main object is greater than a preset area threshold, or the ratio between the area of the main object and the area of the candidate target commodity picture is greater than a preset threshold.
[0058] Overlap degree between the main object and the character
[0059] The overlap degree between the main object and the character refers to the degree of overlap between the main object and the character. The bounding box of the character and the bounding box of the main object can be recognized through OCR technology, and the overlap degree between the main object and the character can be determined based on the bounding box of the character and the bounding box of the main object. In some embodiments, the preset conditions may include that the overlap degree between the main object and the character is less than or equal to a preset overlap degree threshold. Optionally, the above overlap degree threshold can be 0, so that target commodity pictures in which the main object is not covered by characters can be screened out.
[0060] It can be understood that the above is only an exemplary illustration. In other examples, target commodity pictures in which the main object meets other conditions can also be screened out. For example, other conditions may include but are not limited to the shape, connectivity, edge smoothness, number of internal holes, average hole area, and / or average hole area ratio of the main object, etc.
[0061] Figure 3The overall process of the product picture screening process according to the embodiments of the present application is shown. In an e-commerce platform, the visual effect of the picture materials of products (also known as materials, that is, candidate product pictures) plays a crucial role in attracting and retaining users' attention. High-quality and beautiful picture materials can quickly attract consumers' attention and enhance users' willingness to purchase. In scenarios such as creative advertising placement (such as FB DPA, FB AEO, etc.), it is necessary to select appropriate product pictures from millions of product pools and generate the next-stage creative mock-up pictures and scene pictures through generative methods, in order to improve the visual effect of product pictures, attract consumers' attention, and then achieve the improvement of core indicators such as click-through rate and gmv.
[0062] In the related art, open-source models (such as MAN picture quality assessment, aesthetic scoring, etc.) are used for picture screening. However, the implementation effect of the open-source models in the e-commerce scenario is limited, and manual review is required to further confirm the material quality to ensure the overall implementation effect.
[0063] In response to the problem of insufficient picture material quality assessment in the e-commerce scenario, the present application constructs an integrated algorithm process, including recalling picture materials in combination with traffic efficiency, designing a comprehensive classification model in combination with indicators such as aesthetic scores and the embedded representation of pictures for fine screening of picture quality, and determining the availability of materials for the e-commerce scenario. This process can automatically output high-quality pictures that can be directly used for generating the next-stage creative mock-up pictures and scene pictures, improving the visual effect of product pictures while realizing the automation of the picture quality screening pipeline without manual intervention, greatly reducing the labor cost and improving the picture screening efficiency, and at the same time achieving a positive improvement in core indicators such as click-through rate and gmv_roi (the ratio of total transaction volume to input cost, that is, the traffic efficiency in the foregoing embodiments).
[0064] In this embodiment, material recall is first performed through traffic optimization. During this process, parameter tuning is carried out according to traffic parameters such as the click-through rate, conversion rate, and transaction amount of the materials, and materials with higher traffic efficiency are screened out. Then, material fine ranking is performed based on the image quality. During this process, according to the scores of individual dimensions such as the vulgarity score, aesthetics score, and blurriness score of each material screened in the previous step, as well as the image material itself, a comprehensive classification model (i.e., the second sub-model in the foregoing embodiment) is called to give a comprehensive quality score of the material. Then, material usability determination is performed. During this process, it can be first determined whether the background of the material meets the preset conditions. For example, it is determined whether the background is white, that is, it is judged whether the material is a white-background image. If so, it can be further determined whether there are characters outside the main object of the material. If not, the material can be directly determined to be usable. If there are, the material can be cropped, that is, the main object in the material is intercepted. If the background of the material does not meet the preset conditions (for example, the material is not a white-background image), the usability of the material can be first determined. As described in the foregoing embodiment, factors such as the edge complexity of the main object, the number of holes on the main object, the area of the main object, and the overlap degree between the main object and the characters can be obtained. If the characteristics of the main object in each of the above dimensions meet the requirements, it is determined that the material is usable, and thus the material can be cropped.
[0065] In some embodiments, a scene picture or a creative template picture can be generated based on the screened target commodity picture. A scene picture refers to a composite picture generated by combining the main object (such as a commodity, an item, a person, etc.) in the target commodity picture with a preset background or environment (scene). Such a picture not only shows the appearance of the commodity in the target commodity picture, but also makes the commodity appear more suitable for the actual use scene or marketing scene through the design of the background, thereby helping consumers better understand the use, applicable environment or purchase value of the commodity. Continuing with the previous example, for directly usable materials, a background picture can be directly generated, and a scene picture can be generated based on the material and the background picture. For the picture of the main object obtained by cropping, a scene picture can be generated based on the picture of the main object.
[0066] In some embodiments, the selected target product images can be sent to merchants for their use. Alternatively, the selected target product images can also be displayed on the product recommendation page of the e-commerce platform. Specifically, when a product recommendation request sent by the client is received, the target product images can be sent to the client for display in response to the product recommendation request. Or, the selected target product images can also be used for advertising promotion, marketing, or display on other channels or platforms outside the e-commerce platform. For example, the product images may be used for social media advertising, search engine advertising, email marketing, third-party e-commerce platforms, display ads (such as banner ads), etc. By optimizing the visual effects of these images, the click-through rate and purchase conversion rate of users can be increased, and the brand image and product attractiveness can be enhanced.
[0067] The image screening solution of this application has the following technical effects:
[0068] (1) Integrate traffic efficiency and aesthetic evaluation, construct an integrated algorithm process for image material recall, image quality fine screening, and material usability determination, achieve the landing effect of automatic image quality screening without manual intervention in the e-commerce scenario, and achieve a positive improvement in core indicators such as the click-through rate of product images and gmv_roi.
[0069] (2) Design a value formula according to the different stages of the conversion of image materials in the consumption link of the e-commerce scenario, optimize the parameters through efficiency data, determine the hyperparameter values (i.e., the adjustment coefficients corresponding to each stage), estimate the gmv_roi of the image material placement, and screen high-value images.
[0070] (3) Use a comprehensive classification model to fuse image atomic features such as the embedding representation (embedding) and aesthetic score of the image using dcn_v2, and achieve better results than ordinary image classification models in the e-commerce scenario.
[0071] (4) Propose a feasible solution for the e-commerce image scenario based on the contour features of the image, extract information such as the main body shape, complexity, and connectivity in the image to determine whether the cut-out main body is available.
[0072] As described above, the generated target product images can be further used for synthesizing scene images. In the related art, a general image editing model is usually used to edit the original image to obtain the output image. In the e-commerce scenario, the scene in the target product image can be edited and synthesized through a general image editing model to obtain the scene image. However, the image editing operation will change the visual features of the main object in the target product image, resulting in the inconsistency between the visual features of the main object and the features of the actual product.
[0073] Based on this, the present application provides a method for processing product pictures. It generates scene description information that matches the attribute information of the main object in the target product picture based on a large language model, and generates a scene picture based on the scene description information. Since the attribute information of the main object includes the contour information of the main object, the contour of the main object in the generated scene picture is consistent with the contour of the main object in the target product picture. Then, the main object in the scene picture is replaced with the main object in the target product picture, so that the features of the main object in the scene picture are consistent with those of the main object in the target product picture. The above method can automatically generate a scene picture of the main object with consistent features with the target product picture without manual synthesis, reducing the cost of the scene picture synthesis process. The implementation details of the present application will be illustrated with reference to the accompanying drawings below.
[0074] See Figure 4 , the present application provides a method for processing product pictures, and the method includes:
[0075] Step S22: For any one of at least one target product picture, obtain the attribute information of the main object in the target product picture, where the attribute information includes the contour information of the main object;
[0076] Step S24: Generate scene description information of the target scene that matches the attribute information through a pre-trained large language model;
[0077] Step S26: Generate a scene picture through a picture generation model, where the scene picture includes the target scene described by the scene description information and a main object with the contour described by the contour information;
[0078] Step S28: Replace the main object in the scene picture with the main object in the target product picture.
[0079] The target product picture in the embodiment of the present application can be screened from multiple candidate product pictures based on the method in the foregoing embodiment, or can be obtained based on other methods, and the present application does not limit this.
[0080] In step S22, the attribute information of the main object in the target product picture can be obtained. Among them, the main object in the target product picture can be the most important and core product or item in the picture, usually the object that consumers are concerned about and purchase. The attribute information of the main object is used to describe the characteristics of the main object. The attribute information can include the contour information of the main object, and the contour information is used to describe the contour characteristics of the main object. In addition, the attribute information of the main object can also include, but is not limited to, information about the type, color, shape, transparency, function, material, size, applicable object, applicable scene, and / or price and other characteristics of the main object.
[0081] In some embodiments, the target product image can be input into a Vision-Language Model (VLM) so that the vision-language model extracts the attribute information of the main object in the target product image based on the target product image. Among them, the vision-language model can adopt the LLaVA (Language and Vision Assistant) model or other models that simultaneously possess vision processing capabilities and language processing capabilities. In particular, for the contour information in the attribute information, a dedicated edge detection model can be used to extract the contour information of the main object in the target product image to accurately extract the contour information.
[0082] In step S24, at least some of the extracted attribute information can be input into a pre-trained Large Language Model (LLM). Since the large language model has been exposed to a large amount of data during training, and this data contains rich world knowledge, common sense reasoning, relationships between objects and environments, etc., the large language model can infer the environment, background, situation, etc. suitable for the main object based on the attribute information of the main object, thereby generating reasonable and practical scene description information. The scene description information includes but is not limited to the environment and background where the scene is located (such as time, weather, season, atmosphere, etc.), people and items in the scene, events occurring in the scene, visual information of the scene, cultural and social backgrounds, and / or perspectives and narrative angles.
[0083] For example, if the given main object is a "kettle", the large language model will, based on its common sense, infer that it may appear in a kitchen, at a dining table, or in an outdoor camping scene, thereby generating scene description information related to the kitchen, dining table, or outdoor camping scene. Similarly, if the main object is "skis", the large language model will infer that it is suitable for specific environments such as snow-capped mountains and ski resorts, thereby generating scene description information related to specific environments such as snow-capped mountains and ski resorts.
[0084] In step S26, the scene description information and the attribute information of the main object can be input into the image generation model, so that the image generation model generates a scene image including the target scene described by the scene description information and the main object with the contour described by the contour information. Among them, the attribute information input into the image generation model at least includes the contour information of the main object, and this contour information can be used as a constraint condition to make the contour of the main object in the scene image generated by the image generation model consistent with the contour of the main object in the target product image. For example, when the main object in the target product image is a kettle and the scene description information is description information related to the kitchen, the generated scene image is a picture with the kitchen as the background and including the kettle. And, the contour of the kettle in the scene image is consistent with the contour of the kettle in the target product image, but there may be certain differences in other visual features of the kettle in the scene image and the kettle in the target product image.
[0085] In some embodiments, the image generation model includes an SD (Stochastic Depth) model and a ControlNet model. Among them, the SD model is used to obtain the scene description information and generate the target scene and the main object in the scene image based on the scene description information. The ControlNet model is used to obtain the contour information of the main object and use this contour information as a constraint condition to constrain the contour of the main object in the scene image generated by the SD model.
[0086] In the related art, the SD inpainting scheme is usually adopted to generate the scene image. However, this method has a high dependence on resources and is not suitable for running on devices with limited resources. The embodiment of the present application adopts the SD model and the ControlNet model as the image generation model. Compared with the SD inpainting scheme, the dependence on data and computing power resources is low, and the implementation cost is low.
[0087] In some embodiments, the image generation model can be trained in a mask-based manner. In the mask-based training method, a first sample image can be obtained, the image area where the main object in the first sample image is located can be masked, and the image generation model can be used to restore this image area to obtain a predicted sample image. According to the difference between the predicted sample image and the first sample image, the image generation model can be trained.
[0088] Optionally, the picture area where the main object is located may only include the pixel points corresponding to the main object, and does not include the pixel points corresponding to other objects outside the main object. That is, the area of the picture area where the main object is located is equal to the area of the main object. Alternatively, optionally, in addition to including the pixel points corresponding to the main object, the picture area where the main object is located may also include the pixel points corresponding to other objects outside the main object. That is, the area of the picture area where the main object is located is greater than the area of the main object. The applicant found that when training a picture generation model in a mask-based manner, if only the pixel points corresponding to the main object are masked, the pictures generated by the trained picture generation model may have poor visual effects in the edge area of the main object. For example, the transition between the edge of the main object and the background area is not natural enough. By masking the pixel points corresponding to other objects outside the main object in the first sample picture, so that the area of the picture area where the main object is located is greater than the area of the main object, it is possible to make the pictures generated by the trained picture generation model have better visual effects in the edge area of the main object.
[0089] In some embodiments, the pictures generated by the picture generation model trained in the above manner may be more in line with the aesthetic standards of people, but may not meet the traffic requirements of the e-commerce scenario. Therefore, after training the picture generation model in the above manner, it is also possible to obtain a second sample picture with a traffic efficiency higher than a preset efficiency threshold, and fine-tune the trained picture generation model based on the second sample picture. In this way, the picture generation model can learn the general characteristics of pictures with high traffic efficiency in the e-commerce scenario, making the generated pictures more in line with the requirements of the e-commerce scenario.
[0090] In some embodiments, the number of picture generation models may be greater than or equal to 1. When the number of picture generation models is greater than 1, different picture generation models may be trained for the product pictures corresponding to the main objects of different categories. For example, the picture generation model trained for the product pictures corresponding to the products in the clothing category may be different from the picture generation model trained for the product pictures corresponding to the products in the electronic device category. Each trained picture generation model may be applicable to processing the product pictures corresponding to one or more categories of products. On this basis, after obtaining the target product picture, it is possible to first determine the target category to which the main object in the target product picture belongs, and then generate a scene picture based on the image generation model corresponding to the target category. The scene picture includes the target scene described by the scene description information and the main object belonging to the target category with the contour described by the contour information. For example, when the main object in the target product picture is a mobile phone, the picture generation model corresponding to the mobile phone can be used to generate the scene picture; while when the main object in the target product picture is clothes, the picture generation model corresponding to the clothes can be used to generate the scene picture.
[0091] In step S28, the main object in the scene picture can be replaced with the main object in the target product picture. On the one hand, since the outline of the main object in the scene picture is consistent with the outline of the main object in the target product picture, the main object in the target product picture can better fit the main object in the scene picture. On the other hand, after the replacement operation is performed, the visual features of the main object in the scene picture will be consistent with the visual features of the main object in the target product picture, thus being consistent with the visual features of the actual product.
[0092] In some embodiments, the perspective of the main object in the scene picture can also be adjusted. Adjusting the perspective of the main object in the scene picture means changing the display angle of the main object in the scene picture, thereby affecting its presentation, visual effect, and the audience's understanding and perception of it. In scenarios such as e-commerce, advertising design, and film and television production, adjusting the perspective can be used to highlight the features of the main object, enhance visual appeal, or improve the user experience. Adjusting the perspective of the main object includes, but is not limited to: mutual adjustment between the main object in the front view and the main object in the side view, mutual adjustment between the main object in the upward view and the main object in the downward view, mutual adjustment between the main object under a long focal length lens and the main object under a short focal length lens, etc.
[0093] Specifically, after replacing the main object in the scene picture with the main object in the target product picture, the perspective of the replaced main object in the scene picture can be adjusted. Or, the perspective of the main object in the scene picture and the perspective of the main object in the target product picture can be adjusted separately first, and then the main object in the adjusted scene picture can be replaced with the main object in the adjusted target product picture.
[0094] In some embodiments, after generating the scene picture, the first aesthetic score of the scene picture and the second aesthetic score of the target product picture used to generate the scene picture can also be compared. If the first aesthetic score is lower than the second aesthetic score, return to the step of obtaining the attribute information of the main object in the target product picture. Among them, the first aesthetic score and the second aesthetic score can be obtained using the picture quality evaluation model in the foregoing embodiments. By adopting this embodiment, the generated scene picture can have a higher aesthetic score than the original target product picture.
[0095] Figure 5The overall flow chart of the scene image synthesis process of the embodiment of the present application is shown. On e-commerce platforms, the visual effects of product images play a vital role in attracting user attention. High-quality and beautiful product images can quickly attract consumers' attention and enhance users' willingness to buy. At present, the product images on e-commerce platforms have the following problems:
[0096] (1) The product images are not of high quality and cannot directly reflect the product's purpose and usage scenarios. It is difficult to arouse user interest by recommending such products to users;
[0097] (2) Product images are of similar styles, making it difficult to meet the diverse interests and preferences of different user groups.
[0098] In order to truly solve the above business pain points, this application designs and implements a product image traffic-oriented solution based on a large model under the constraints of resources such as image annotation data and computing power, realizing an integrated process from image quality optimization to batch generation of scene images based on a large model, generation effect optimization, and then to image generation availability evaluation. This solution generates scene images that can show the value and usage scenarios of the product based on existing product images. Under the constraints of computing power and data resources, this application can stimulate user interest while improving the quality and aesthetics of the image, increase the content distribution limit while increasing the diversity of product image styles, and thus promote the improvement of traffic efficiency indicators.
[0099] Traditional scene image generation mostly uses SD-based inpainting solutions, but the existing open source SD inpainting models have average results in e-commerce scenarios. To achieve the goal of e-commerce scene graph production, further effect tuning is required, and SD-based full-parameter effect tuning relies on a large amount of high-quality commodity scene gallery data and computing power resources, which will generate a large amount of manual annotation and data procurement costs, making it difficult to achieve good results under resource constraints such as image annotation data and computing power.
[0100] The specific steps of this application may include:
[0101] (1) Image optimization: Based on the product efficiency consumption data, multiple candidate product images are initially screened out by means of parameter optimization. The quality is then optimized based on the image quality assessment model. The subject of the optimized image material is then cut out. Finally, the usability of the product subject (after cutting out) is determined based on the contour features of the subject object, thereby selecting the target product image.
[0102] (2)Efficient implementation of e-commerce scenarios under resource constraints: Based on the target product images generated in step (1), first obtain the content, color and other attribute information of the main object based on the LLaVA model, then generate scene description information suitable for the main object based on the above attribute information and the LLM model, and finally generate scene images under the control of ControlNet based on the image SD model. This solution has better effects than the SD inpainting-based solution, with low dependence on data and computing resources and low implementation costs. Since only the outline of the main object in the original target product image remains unchanged in the scene images generated at this stage, the main object in the original target product image is synthesized in the second stage to ensure that the main object in the scene image is the same as the main object in the target product image.
[0103] (3)Optimization of traffic-oriented effects: Further optimize the effects of the scene images generated in step (2) through scene pipeline adaptation and optimization of the perspective of the product main body. To further achieve low-cost model optimization under data and computing constraints, this application screens out sample images with high traffic efficiency (i.e., the second sample images), which are directly used for large model generation optimization after quality filtering. This dataset takes into account quality while better meeting the user preferences of e-commerce scenarios and does not rely on manual annotation resources. Based on this traffic-oriented dataset, low-cost category-specific fine-tuning based on SD+controlnet can be achieved to further optimize the generation effect of scene images.
[0104] (4)Evaluation of generation effects: After completing steps (1) to (3), this application evaluates the usability of the high-quality scene images and corresponding target product images output by the algorithm based on image aesthetics scores to further ensure the quality, aesthetics and usability of the generated scene images.
[0105] Schematic diagrams of the original target product images and the generated scene images are as Figure 6 shown. In the first stage, the image generation model generates first-stage scene images based on the scene description information and the attribute information of the main object. It can be seen that in the scene images generated at this stage, the outline of the main object is the same as the outline of the main object in the target product image, but other visual features are not exactly the same. In the second stage, the main object in the target product image is replaced into the scene image to obtain the second-stage scene image. It can be seen that in the scene images generated at this stage, the visual features of the main object are the same as the outline of the main object in the target product image.
[0106] The scene image generation solution of this application has the following advantages:
[0107] (1) Under the resource constraints of picture annotation data, computing power, etc., first generate scene description information suitable for commodities based on VLM and LLM, and then directly generate scene pictures based on the picture generation model (SD model) under the control of ControlNet. The direct picture generation effect based on SD+ControlNet is better than that of the SD inpainting scheme, with low dependence on data and computing power resources and low implementation cost.
[0108] (2) When directly generating scene pictures based on SD+ControlNet, only the outline of the main object remains unchanged, and the details of the main object may change. Therefore, in this application, the original target commodity picture and the first-stage scene picture are synthesized in the second stage to ensure that while generating an available scene for the commodity, the main body of the commodity is exactly the same as the original picture.
[0109] (3) To further achieve low-cost model tuning under data and computing power constraints, this application screens high-efficiency traffic-oriented pictures based on traffic efficiency data, and directly uses them for tuning the picture generation model after quality filtering. This data set takes into account quality and is more in line with the preferences of users in this scenario, and is completely independent of manual annotation resources. Based on this traffic-oriented data set, low-cost category-specific fine-tuning based on SD+ControlNet can be realized to further optimize the scene picture generation effect.
[0110] (4) To ensure the usability of the generated pictures, this application evaluates the usability of the scene pictures output by the algorithm and the original target commodity pictures based on image aesthetics scoring to ensure the final quality of the scene pictures.
[0111] See Figure 7 , this application also provides a commodity picture processing device, and the device includes:
[0112] The first acquisition module 102 is used to acquire multiple candidate commodity pictures of the e-commerce platform;
[0113] The determination module 104 is used to acquire the traffic parameters of the multiple candidate commodity pictures, and determine the traffic efficiency of the multiple candidate commodity pictures based on the traffic parameters of the multiple candidate commodity pictures; wherein, the traffic parameter of the candidate commodity picture is used to characterize the access traffic of the user of the e-commerce platform to the candidate commodity picture, and the traffic efficiency of the candidate commodity picture is used to characterize the access traffic obtained by the candidate commodity picture under the preset display cost of the commodity picture;
[0114] The second acquisition module 106 is used to acquire the quality of the multiple candidate commodity pictures through a pre-trained picture quality evaluation model;
[0115] A screening module 108, configured to screen at least one target product image from the multiple candidate product images based on the traffic efficiency and quality of the multiple candidate product images.
[0116] See Figure 8 , this application further provides a product image processing device, the device includes:
[0117] A third acquisition module 202, configured to acquire, for any one of the at least one target product image, attribute information of the main object in the target product image, where the attribute information includes contour information of the main object;
[0118] A first generation module 204, configured to generate scene description information of a target scene that matches the attribute information through a pre-trained large language model;
[0119] A second generation module 206, configured to generate a scene image through an image generation model, where the scene image includes the target scene described by the scene description information and the main object with the contour described by the contour information;
[0120] A replacement module 208, configured to replace the main object in the scene image with the main object in the target product image.
[0121] The functions or modules included in the device provided in this application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0122] An embodiment of this application further provides a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the program, it implements the method described in any one of the foregoing embodiments.
[0123] Figure 9 FIG. shows a more specific schematic hardware structure diagram of a computer device provided in an embodiment of this application. The device may include: a processor 302, a memory 304, an input / output interface 306, a communication interface 308, and a bus 310. Wherein, the processor 302, the memory 304, the input / output interface 306, and the communication interface 308 are communicatively connected to each other inside the device through the bus 310.
[0124] The processor 302 can be implemented in the form of a general - purpose central processing unit (CPU), a microprocessor, an application - specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The processor 302 can also include a graphics card, and the graphics card can be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0125] The memory 304 can be implemented in the form of a read - only memory (ROM), a random - access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 304 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present application through tools or firmware, the relevant program codes are stored in the memory 304 and are called and executed by the processor 302.
[0126] The input / output interface 306 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0127] The communication interface 308 is used to connect to the communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or can achieve communication through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).
[0128] The bus 310 includes a path for transmitting information between various components of the device (such as the processor 302, the memory 304, the input / output interface 306, and the communication interface 308).
[0129] It should be noted that although the above - mentioned device only shows the processor 302, the memory 304, the input / output interface 306, the communication interface 308, and the bus 310, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above - mentioned device may also only include the components necessary to implement the solutions of the embodiments of the present application, and does not necessarily include all the components shown in the figure.
[0130] An embodiment of the present application provides a computer program product, including a computer program, which implements the method described in any embodiment of the present application when executed by a processor.
[0131] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and the program implements the method described in any of the foregoing embodiments when executed by a processor.
[0132] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information accessible by a computer device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0133] Each embodiment in the present application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solution of the embodiments of the present application, the functions of the modules can be implemented in the same or multiple tools and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0134] The above is only the specific implementation manner of the embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the embodiments of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of the present application.
Claims
1. A method for processing a product image, the method comprising: For any one of the at least one target product image, acquiring attribute information of a main object in the target product image, wherein the attribute information includes contour information of the main object; Generate scene description information of the target scene matching the attribute information through a pre-trained large language model; Generate a scene picture by using a picture generation model, wherein the scene picture includes a target scene described by the scene description information and a main object having a contour described by the contour information; The main object in the scene image is replaced with the main object in the target product image.
2. According to the method of claim 1, before generating the scene picture by the picture generation model, the method further comprises: Get the first sample image; Masking the image region where the main object in the first sample image is located; wherein the area of the image region where the main object is located is larger than the area of the main object; Restoring the image region by using the image generation model to obtain a predicted sample image; The picture generation model is trained based on feature differences between the first sample picture and the predicted sample picture.
3. The method according to claim 2, after training the picture generation model based on the feature difference between the first sample picture and the predicted sample picture, the method further comprises: Acquire a second sample image whose flow efficiency is higher than a preset efficiency threshold; Fine-tune the trained image generation model based on the second sample image.
4. The method according to claim 1, further comprising: The viewing angle of the main object in the scene picture is adjusted.
5. The method according to claim 1, further comprising: comparing a first aesthetic score of the scene image and a second aesthetic score of a target product image used to generate the scene image; If the first aesthetic score is lower than the second aesthetic score, return to the step of obtaining attribute information of the main object in the target product image.
6. The method according to claim 1, wherein the attribute information of the subject object includes contour information of the subject object and other visual information of the subject object, wherein the other visual information is visual information other than the contour information; The acquiring of attribute information of the main object in the target product image includes: Acquire other visual information of the subject object through a visual language model; as well as The contour information of the main object is obtained through an edge detection model.
7. According to the method of claim 1, the picture generation model includes an SD model and a ControlNet model; the step of generating a scene picture through the picture generation model includes: Acquire scene description information through the SD model, and generate a target scene and a main object in a scene picture based on the scene description information; The contour information of the main object is obtained through the ControlNet model, and the contour information is used as a constraint condition to constrain the contour of the main object in the scene picture generated by the SD model.
8. The method according to claim 1, wherein the number of the image generation models is greater than or equal to 1; each image generation model corresponds to at least one commodity category; The step of generating a scene image through an image generation model includes: Determine the target category to which the main object in the target product image belongs; A scene picture is generated based on the image generation model corresponding to the target category, wherein the scene picture includes the target scene described by the scene description information and a main object having a contour described by the contour information and belonging to the target category.
9. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 8.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program.
11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Intelligent agent autonomous evolutionary algorithm for business scene and application system of intelligent agent autonomous evolutionary algorithm
CN120689523A