Commodity picture processing method, medium, computer equipment and program product

By obtaining and evaluating the traffic parameters and quality of product images on e-commerce platforms, we automatically filter out target product images with high traffic efficiency and aesthetic effects, solving the problem of high manual screening costs in the existing technology, and improving screening efficiency and product performance.

CN120182432APending Publication Date: 2025-06-20HANGZHOU ALIBABA INT INTERNET IND CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510121462.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, processing product images of e-commerce platforms based on an open source image processing model often requires further manual screening, resulting in high costs and does not meet the needs of e-commerce scenarios.

Method used

By obtaining the traffic parameters and quality of candidate product images of e-commerce platforms, based on the pre-trained image quality evaluation model, target product images with high traffic efficiency and aesthetic effects are selected.

Benefits of technology

It has realized the automatic screening of target product images that meet the needs of e-commerce platforms, which has improved the image screening efficiency, reduced the cost of manual screening, and increased the click-through rate and conversion rate of products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182432A_ABST
    Figure CN120182432A_ABST
Patent Text Reader

Abstract

The invention discloses a commodity picture processing method, a medium, computer equipment and a program product. The method comprises the following steps: acquiring a plurality of candidate commodity pictures of an e-commerce platform; acquiring flow parameters of the plurality of candidate commodity pictures, and determining flow efficiency of the plurality of candidate commodity pictures based on the flow parameters of the plurality of candidate commodity pictures; wherein the flow parameter of the candidate commodity picture is used for representing the access flow of a user of the e-commerce platform to the candidate commodity picture, and the flow efficiency of the candidate commodity picture is used for representing the access flow obtained by the candidate commodity picture under the preset commodity picture display cost; obtaining the quality of the plurality of candidate commodity pictures through a pre-trained picture quality evaluation model; and screening out at least one target commodity picture from the plurality of candidate commodity pictures based on the traffic efficiency and quality of the plurality of candidate commodity pictures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image processing, and in particular, to a method for processing product images, a medium, a computer device, and a program product. Background Art

[0002] In an e-commerce platform, the visual effect of product images plays a crucial role in attracting users' attention. High-quality and beautiful product images can quickly attract users' attention and enhance users' stickiness to the e-commerce platform. Therefore, it is necessary to optimize the product images on the e-commerce platform. In related technologies, an open-source image processing model is usually used to process the product images on the e-commerce platform. Since the open-source image processing model is a model for general fields and the processing target does not meet the requirements of the e-commerce scenario, the product images processed based on the open-source image processing model often need to be further processed manually, resulting in high costs. Summary of the Invention

[0003] In a first aspect, an embodiment of this application provides a method for processing product images, and the method includes: obtaining multiple candidate product images of an e-commerce platform; obtaining traffic parameters of the multiple candidate product images, and determining traffic efficiency of the multiple candidate product images based on the traffic parameters of the multiple candidate product images; where the traffic parameter of a candidate product image is used to represent the access traffic of a user of the e-commerce platform to the candidate product image, and the traffic efficiency of a candidate product image is used to represent the access traffic obtained by the candidate product image under a preset product image display cost; obtaining the quality of the multiple candidate product images through a pre-trained image quality evaluation model; and screening out at least one target product image from the multiple candidate product images based on the traffic efficiency and quality of the multiple candidate product images.

[0004] In the embodiment of this application, target product images are screened out from candidate product images based on two dimensions of traffic efficiency and quality of product images. Among them, the traffic parameter of a candidate product image is used to represent the access traffic of a user of the e-commerce platform to the candidate product image. Therefore, based on the traffic efficiency, the attractiveness and dissemination effect of candidate product images on the e-commerce platform can be effectively evaluated, which helps to select target product images that can maximize the attraction of potential users, thereby improving the click-through rate and conversion rate of products, and making the screened target product images match the requirements of the e-commerce platform; the image quality can reflect the aesthetic effect of the product image. In summary, the solution of the embodiment of this application can automatically screen out target product images with good aesthetic effects and meeting the requirements of the e-commerce platform, without manual screening, improving the image screening efficiency.

[0005] Second aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.

[0006] Third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any embodiment of the present application is implemented.

[0007] Fourth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.

[0008] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, rather than limiting the present application. Description of the Drawings

[0009] The drawings herein are incorporated into the specification and form a part of the present application. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application.

[0010] Figure 1 is a flowchart of the method for processing product pictures in an embodiment of the present application.

[0011] Figure 2 is a schematic structural diagram of the second sub-model in an embodiment of the present application.

[0012] Figure 3 is a general flowchart of the product picture screening process in an embodiment of the present application.

[0013] Figure 4 is a flowchart of the method for processing product pictures in another embodiment of the present application.

[0014] Figure 5 is a general flowchart of the scene picture synthesis process in an embodiment of the present application.

[0015] Figure 6 is a schematic diagram of the target product picture and the scene picture in an embodiment of the present application.

[0016] Figure 7 is a block diagram of the product picture processing device in an embodiment of the present application.

[0017] Figure 8 is a block diagram of the product picture processing device in another embodiment of the present application.

[0018] Figure 9 is a schematic diagram of the computer device in an embodiment of the present application. Detailed Embodiments

[0019] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0020] The terms used in the present application are for the purpose of describing particular embodiments only and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality.

[0021] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0022] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0023] In the e-commerce scenario, some target product images are selected from a large number of product images (hereinafter referred to as candidate product images) on the e-commerce platform for optimization. By optimizing the product images, the attractiveness of the product images to users can be improved, thereby enhancing the stickiness of users to the e-commerce platform. In the related art, an open-source image quality assessment model is usually used to select images with higher quality from the numerous product images on the e-commerce platform. This screening method mainly screens the images from an aesthetic perspective, and the optimization goal does not meet the requirements of the e-commerce scenario. Therefore, it is necessary to further manually screen the product images selected by the image quality assessment model, resulting in a higher cost.

[0024] Based on this, an embodiment of the present application proposes a commodity picture screening solution that integrates the traffic efficiency and aesthetic effect of commodity pictures. Among them, the traffic efficiency can effectively evaluate the attractiveness and dissemination effect of candidate commodity pictures on the e-commerce platform, and the picture quality can reflect the aesthetic effect of the commodity pictures. Therefore, the present application can automatically screen out target commodity pictures with better aesthetic effects and meeting the requirements of the e-commerce platform without manual screening, improving the picture screening efficiency. The implementation details of the present application will be illustrated with reference to the accompanying drawings below.

[0025] See Figure 1 , the present application provides a commodity picture processing method, the method comprising:

[0026] Step S12: Obtain multiple candidate commodity pictures of the e-commerce platform;

[0027] Step S14: Obtain the traffic parameters of the multiple candidate commodity pictures, and determine the traffic efficiency of the multiple candidate commodity pictures based on the traffic parameters of the multiple candidate commodity pictures; wherein, the traffic parameter of the candidate commodity picture is used to characterize the access traffic of the user of the e-commerce platform to the candidate commodity picture, and the traffic efficiency of the candidate commodity picture is used to characterize the access traffic obtained by the candidate commodity picture under the preset commodity picture display cost;

[0028] Step S16: Obtain the quality of the multiple candidate commodity pictures through a pre-trained picture quality evaluation model;

[0029] Step S18: Screen out at least one target commodity picture from the multiple candidate commodity pictures based on the traffic efficiency and quality of the multiple candidate commodity pictures.

[0030] In step S12, all the full-scale commodity pictures of the e-commerce platform can be used as candidate commodity pictures. Alternatively, the full-scale commodity pictures of the e-commerce platform can be preliminarily screened, and the preliminarily screened commodity pictures can be used as candidate commodity pictures. Among them, the preliminary screening can be implemented manually or automatically by software. For example, several candidate commodity pictures can be randomly selected from the full-scale commodity pictures of the e-commerce platform. Alternatively, the pictures of specific categories of commodities on the e-commerce platform can be used as candidate commodity pictures. Alternatively, the commodity pictures on the e-commerce platform that conform to a specific theme can be used as candidate commodity pictures. Other ways can also be used to obtain candidate commodity pictures, and the specific ways to obtain candidate commodity pictures will not be listed one by one here.

[0031] In step S14, the traffic parameters of each candidate product picture can be obtained. Among them, the traffic parameters of the candidate product picture are used to characterize the access traffic of the users of the e-commerce platform to the candidate product picture, including but not limited to the access traffic brought by browsing the product picture, the access traffic brought by the user clicking on the product picture, the access traffic brought by the user purchasing the product corresponding to the product picture, the access traffic brought by the user collecting the product corresponding to the product picture, and / or the access traffic brought by the user sharing the product corresponding to the product picture.

[0032] In some embodiments, the types of traffic parameters can be greater than or equal to 1. In the embodiments where the types of traffic parameters are greater than 1, the traffic parameters include the traffic parameters of multiple stages on the product consumption link. Among them, the multiple stages on the product consumption link can include the traffic input stage and the traffic output stage. In the traffic input stage, the merchant needs to invest a certain cost to display the product picture. Therefore, the traffic parameters in the traffic input stage can include the product picture display cost. In the traffic output stage, the product picture can be obtained by the user, thus bringing traffic to the merchant. The traffic output stage usually includes stages such as exposure, click, and purchase. Therefore, the traffic parameters in the traffic output stage include but are not limited to at least part of the exposure rate, click-through rate, conversion rate of the candidate product picture, and the transaction amount of the product corresponding to the candidate product picture. Among them, the exposure rate of any candidate product picture can be determined by the ratio of the exposure volume of the candidate product picture to the maximum exposure volume of each candidate product picture. For example, assume there are 3 candidate product pictures, denoted as Picture A, Picture B, and Picture C respectively. Among them, the exposure volume of Picture A is 32, the exposure volume of Picture B is 50, and the exposure volume of Picture C is 33. Then the maximum exposure volume of the candidate product picture is 50. Therefore, the exposure rate of Picture A is 64%, the exposure rate of Picture B is 100%, and the exposure rate of Picture C is 66%. The click-through rate of any candidate product picture can be determined by the ratio of the click volume of the candidate product picture to the exposure volume of the candidate product picture. For example, assume the click volume of Picture B above is 20, then the click-through rate of Picture B is 40%. The conversion rate of any candidate product picture can be determined by the ratio of the purchase volume of the candidate product picture to the exposure volume of the candidate product picture. For example, assume the purchase volume of Picture B above is 10, then the purchase rate of Picture B is 20%. The transaction amount of the product corresponding to the candidate product picture is the actual amount paid by the user when purchasing the product corresponding to the candidate product picture.

[0033] It can be understood that the above is only an exemplary illustration and is not used to limit this application. In other examples, the traffic parameters can also include other types of parameters, and the calculation methods of the above various traffic parameters are not limited to the calculation methods described in the above embodiments.

[0034] After obtaining the traffic parameters of the candidate product pictures, the traffic efficiency of the candidate product pictures can be determined based on the traffic parameters of the candidate product pictures. Traffic efficiency refers to the contribution that can be provided to the achievement of the service goal in the traffic, and can represent the access traffic obtained by the candidate product pictures under the preset product picture display cost. Traffic efficiency is positively correlated with the traffic parameters in the traffic output stage and negatively correlated with the traffic parameters in the traffic input stage. For example, the traffic efficiency can be determined based on the ratio between the traffic parameters in the traffic output stage and the traffic parameters in the traffic input stage. Further, when there are two or more traffic parameters in the traffic output stage, the product of multiple traffic parameters in the traffic output stage can be determined, and the traffic efficiency can be determined based on the ratio between the above product and the traffic parameters in the traffic input stage.

[0035] In an embodiment where the types of traffic parameters are greater than 1, adjustment coefficients corresponding to multiple stages can also be obtained. The adjustment coefficient corresponding to any stage is used to adjust the contribution degree of the traffic parameters in this stage to the traffic efficiency. The traffic parameters of the candidate product pictures in the corresponding stage are adjusted based on the adjustment coefficients corresponding to the multiple stages respectively, and the traffic efficiency of the multiple candidate product pictures is determined based on the product picture display cost and the adjusted traffic parameters of the multiple stages.

[0036] For example, following the previous example, when the traffic parameters of the candidate product pictures in multiple stages include the product picture display cost, exposure rate, click-through rate, conversion rate of the candidate product pictures, and the transaction amount of the product corresponding to the candidate product pictures, correspondingly, the adjustment coefficients corresponding to the multiple stages respectively include a cost adjustment coefficient, an exposure rate adjustment coefficient, a click-through rate adjustment coefficient, a conversion rate adjustment coefficient, and a transaction amount adjustment coefficient.

[0037] By obtaining the adjustment coefficients corresponding to each stage, the contribution degree of the traffic parameters in each stage to the traffic efficiency can be adjusted. For example, when the adjustment coefficient corresponding to the product exposure stage is large, while the adjustment coefficient corresponding to the product click stage is small, the traffic parameters (i.e., the exposure rate) in the product exposure stage have a greater contribution degree to the traffic efficiency, while the traffic parameters (i.e., the click-through rate) in the product click stage have a smaller contribution degree to the traffic efficiency.

[0038] In some embodiments, the product traffic of the candidate product pictures can be recorded as:

[0039]

[0040] Among them, value represents the traffic efficiency, ER represents the exposure rate, ctr represents the click-through rate, cvr represents the conversion rate, gmv represents the transaction amount, cpm is the cost per thousand impressions, that is, the fee that a merchant needs to pay to display the product picture to one thousand users, that is, the product picture display cost, and x1, x2, x3, x4, and x5 respectively represent the exposure rate adjustment coefficient, click-through rate adjustment coefficient, conversion rate adjustment coefficient, transaction amount adjustment coefficient, and cost adjustment coefficient.

[0041] In some embodiments, the adjustment coefficients corresponding to multiple stages can be determined in the following manner: Obtain the traffic parameters of the multiple stages in the first time period of the preset cycle for the candidate product picture; Based on the initial adjustment coefficients corresponding to the multiple stages and the traffic parameters of the multiple stages in the first time period, determine the traffic efficiency of the candidate product picture in the preset cycle; Based on the traffic efficiency of the candidate product picture in the preset cycle and the initial adjustment coefficients corresponding to the multiple stages, determine the estimated traffic parameters of the multiple stages in the second time period of the preset cycle for the candidate product picture; Based on the difference between the estimated traffic parameters of the multiple stages in the second time period of the candidate product picture and the traffic parameters of the corresponding stages of the candidate product picture in the second time period, adjust the initial adjustment coefficients corresponding to the multiple stages to obtain the adjustment coefficients corresponding to the multiple stages.

[0042] For example, 30 days can be determined as a cycle. The first 25 days out of the 30 days are determined as the first time period, and the last 5 days out of the 30 days are determined as the second time period. Taking the calculation of flow efficiency through the above formula as an example, the initial values of x1, x2, x3, x4, and x5 (i.e., the initial adjustment coefficients corresponding to multiple stages) can be set first, and the flow parameters of each day in the first 25 days are substituted into the formula to obtain the flow efficiency corresponding to each of the first 25 days. The average value of the flow efficiency corresponding to each of the first 25 days is obtained to get the flow efficiency of the corresponding cycle. Since the last 5 days and the first 25 days belong to the same cycle, it can be considered that the flow efficiency of the last 5 days should be the same as the flow efficiency of this cycle calculated based on the above method. Substituting the initial values of x1, x2, x3, x4, and x5 and the flow efficiency of this cycle into the formula, the estimated flow parameters of the last 5 days can be obtained. The difference between the actual flow parameter of each day in the last 5 days and the above estimated flow parameter can be obtained, and the sum of the differences corresponding to the last 5 days is calculated to get the total difference corresponding to the last 5 days. The above total difference can reflect the estimation deviation caused by inaccurate setting of the initial values of x1, x2, x3, x4, and x5. Therefore, the initial values of x1, x2, x3, x4, and x5 can be adjusted based on the above total difference. Multiple iterations can be performed, and the above process is repeated each time until the iteration stop condition is met, such as the number of iterations reaches the preset upper limit, or the above total difference is less than the preset threshold, etc., and the x1, x2, x3, x4, and x5 obtained when the iteration stop condition is met are determined as the adjustment coefficients corresponding to the multiple stages respectively.

[0043] In step S16, the quality of the multiple candidate product pictures can be obtained by a pre-trained picture quality evaluation model. In some embodiments, the picture quality evaluation model can evaluate the quality of the candidate product pictures from multiple dimensions. Specifically, the picture quality evaluation model can include multiple first sub-models, and different first sub-models are used to evaluate the quality scores of the candidate product pictures in different dimensions. The above multiple dimensions include but are not limited to clarity, color, composition, and / or authenticity, etc. For any candidate product picture, each first sub-model can output a quality score for the candidate product picture, and the quality score is positively correlated with the quality of the candidate product picture in the corresponding dimension.

[0044] In other embodiments, the picture quality evaluation model can include a second sub-model, and the second sub-model can evaluate the comprehensive quality score of the candidate product picture. As an implementation manner, the second sub-model can obtain the candidate product picture and the quality scores of the candidate product picture in the multiple dimensions (i.e., the quality scores output by each first sub-model), and obtain the comprehensive quality score of the candidate product picture based on the candidate product picture and the quality scores of the candidate product picture in the multiple dimensions. The quality of the candidate product picture can be determined based on the above comprehensive quality score.

[0045] Figure 2 The structure of the second sub-model in some embodiments is shown. The quality scores output by the first sub-model may include a quality score for evaluating the basic picture quality and a quality score for evaluating the picture display effect. Among them, the basic picture quality may include, but is not limited to, the blurriness, noise intensity, and / or compression intensity of the picture. The display effect can be obtained by judging the size of the main object, performing edge detection on the main object, obtaining the vulgarity score of the main object, and / or obtaining the aesthetic score of the main object, etc. Each quality score can be converted into an embedding representation through vector quantization processing, and these embedding representations can be input into a multi-layer perceptron. The main object in the product picture refers to the most important and core product or item in the picture, usually the object that consumers focus on and purchase. In scenarios such as e-commerce platforms, advertising, and product displays, the main object of the product picture is the core product or service presented to customers through images.

[0046] The second sub-model can be implemented based on architectures such as ResNet or VIT. First, obtain the embedding representation corresponding to the candidate product picture, and then input the embedding representation corresponding to the candidate product picture into a multi-layer perceptron. It should be noted that although two multi-layer perceptrons are shown in the figure, in actual applications, only one multi-layer perceptron can be used. That is, the embedding representations corresponding to the quality scores in each dimension and the embedding representation of the candidate product picture can be output to the same multi-layer perceptron. The multi-layer perceptron can extract features from the embedding representations corresponding to the quality scores in each dimension and the embedding representation of the candidate product picture, and output the extracted features to a splicing module for splicing (concat). The spliced features can be output to a fusion module for feature fusion. In order to better fuse the features corresponding to the embedding representation of the picture and the features corresponding to the atomic features, the fusion module can be implemented based on the architecture of DCN_V2 (Deep&Cross Network). The fused features can be output to a classifier, and this classifier can output the comprehensive score of the candidate product picture based on the fused features.

[0047] The embodiments of the present application can comprehensively evaluate the quality of product pictures in multiple dimensions based on atomic features in multiple dimensions, and perform quality evaluation based on the overall candidate product picture, which can evaluate the visual effect of the candidate product picture macroscopically, so as to obtain accurate and comprehensive quality evaluation results.

[0048] In addition, the comprehensive quality score of the product picture and / or the quality scores in each dimension obtained by the picture quality evaluation model can also be shown to the merchant for reference. If the score of the product picture is too low, the merchant can modify or re-upload the product picture.

[0049] In step S18, at least one target product image can be selected from the multiple candidate product images based on the traffic efficiency and quality of the multiple candidate product images.

[0050] It should be noted that step S16 can be executed after step S14. For example, the traffic efficiency of each candidate product image can be obtained first based on step S14, and then the quality of several candidate product images with traffic efficiency higher than the preset efficiency threshold can be obtained through step S16. Then, at least one target product image can be selected from the several candidate product images based on the quality of the several candidate product images. Or, step S16 can be executed before step S14. For example, the quality of each candidate product image can be obtained first through step S16, and then the traffic efficiency of several candidate product images with quality higher than the preset quality threshold can be obtained based on step S14. Then, at least one target product image can be selected from the several candidate product images based on the traffic efficiency of the several candidate product images. Or, step S14 and step S16 can be executed in parallel, that is, the traffic efficiency and quality of each candidate product image can be obtained in parallel, and at least one target product image can be selected from each candidate product image based on the traffic efficiency and quality of each candidate product image. Since the quality of the product image is obtained based on the image quality evaluation model, and the model operation requires a lot of resources, therefore, first obtaining the traffic efficiency of each candidate product image through step S14, and then obtaining the quality of several candidate product images with traffic efficiency higher than the preset efficiency threshold through step S16, this method can effectively reduce the number of candidate product images that the image quality evaluation model needs to process, thereby reducing the resource consumption in the image selection process.

[0051] In some embodiments, in addition to selecting the target product image based on the traffic efficiency and quality of the candidate product image, the target product image can also be further selected according to the main object in the candidate product image. Specifically, at least one candidate target product image can be selected from the multiple candidate product images based on the traffic efficiency and quality of the multiple candidate product images, and at least one target product image whose main object meets the preset conditions can be selected from the at least one candidate target product image. Among them, the preset conditions can be determined based on at least one of the following: the edge complexity of the main object, the number of holes on the main object, the area of the main object, the overlap degree between the main object and the character.

[0052] Edge complexity

[0053] Edge complexity refers to the complexity of the boundary shape of the main object. Edge complexity is usually determined by factors such as the irregularity of the edge contour of the main object, the number of curves, and the degree of bending. If the edge of the main object presents many twists and irregular lines, then its edge complexity is higher; while if the edge of the main object is a straight line or has few bends, the edge complexity is lower. In some embodiments, the edge complexity of the main object can be determined based on the ratio of the perimeter of the edge contour of the main object to the area of the main object. In some embodiments, the preset conditions may include that the edge complexity of the main object is lower than a preset edge complexity threshold.

[0054] Number of holes

[0055] The number of holes refers to the number of voids or openings existing on the surface of the main object. The number of holes can reflect the integrity of the main object. The more the number of holes, the lower the integrity of the main object. In some embodiments, the preset conditions may include that the number of holes on the main object is less than a preset number threshold.

[0056] Area of the main object

[0057] The area of the main object can reflect the space size occupied by the main object in the candidate target commodity picture. If the area of the main object is too small, then the main object is usually not prominent enough in the commodity picture, difficult to attract user attention, and the optimization effect is usually not good enough. In some embodiments, the preset conditions may include that the area of the main object is greater than a preset area threshold, or the ratio between the area of the main object and the area of the candidate target commodity picture is greater than a preset threshold.

[0058] Overlap degree between the main object and the character

[0059] The overlap degree between the main object and the character refers to the degree of overlap between the main object and the character. The bounding box of the character and the bounding box of the main object can be identified through OCR technology, and the overlap degree between the main object and the character can be determined based on the bounding box of the character and the bounding box of the main object. In some embodiments, the preset conditions may include that the overlap degree between the main object and the character is less than or equal to a preset overlap degree threshold. Optionally, the above overlap degree threshold can be 0, so that target commodity pictures in which the main object is not covered by characters can be screened out.

[0060] It can be understood that the above is only an exemplary illustration. In other examples, target commodity pictures in which the main object meets other conditions can also be screened out. For example, other conditions may include but are not limited to the shape, connectivity, edge smoothness, number of internal holes, average hole area, and / or average hole area ratio of the main object, etc.

[0061] Figure 3The overall process of the product picture screening process according to the embodiments of the present application is shown. In an e-commerce platform, the visual effect of the picture materials of products (also known as materials, that is, candidate product pictures) plays a crucial role in attracting and retaining users' attention. High-quality and beautiful picture materials can quickly attract consumers' attention and enhance users' willingness to purchase. In scenarios such as creative advertising placement (such as FB DPA, FB AEO, etc.), it is necessary to select appropriate product pictures from millions of product pools and generate the next-stage creative mock-up pictures and scene pictures through generative methods, in order to improve the visual effect of product pictures, attract consumers' attention, and then achieve the improvement of core indicators such as click-through rate and GMV.

[0062] In related technologies, open-source models (such as MAN picture quality assessment, aesthetic scoring, etc.) are used for picture screening. However, the implementation effect of open-source models in the e-commerce scenario is limited, and manual review is required to further confirm the material quality to ensure the overall implementation effect.

[0063] Aiming at the problem of insufficient evaluation of the quality of picture materials in the e-commerce scenario, the present application constructs an integrated algorithm process, including recalling picture materials in combination with traffic efficiency, designing a comprehensive classification model in combination with indicators such as aesthetic scores and the embedded representation of pictures for fine screening of picture quality, and determining the availability of materials for the e-commerce scenario. This process can automatically output high-quality pictures that can be directly used for the generation of the next-stage creative mock-up pictures and scene pictures, improving the visual effect of product pictures while realizing the automation of the picture quality screening pipeline without manual intervention, greatly reducing the labor cost and improving the picture screening efficiency, and at the same time achieving a positive improvement in core indicators such as click-through rate and GMV_ROI (the ratio of the total transaction amount to the input cost, that is, the traffic efficiency in the foregoing embodiments).

[0064] In this embodiment, material recall is first performed through traffic optimization. During this process, parameter tuning is carried out according to traffic parameters such as the click-through rate, conversion rate, and transaction amount of the materials, and materials with higher traffic efficiency are screened out. Then, material fine ranking is performed based on the picture quality. During this process, according to the ratings of individual dimensions such as the vulgarity rating, aesthetics rating, and blurriness rating of each material screened out in the previous step, as well as the picture material itself, a comprehensive classification model (i.e., the second sub-model in the foregoing embodiment) is called to give the comprehensive quality rating of the material. Then, material usability determination is performed. During this process, it can be first determined whether the background of the material meets the preset conditions. For example, it is determined whether the background is white, that is, it is judged whether the material is a white-background picture. If so, it can be further determined whether there are characters outside the main object of the material. If not, the material can be directly determined to be usable. If there are, the material can be cropped, that is, the main object in the material is intercepted. If the background of the material does not meet the preset conditions (for example, the material is not a white-background picture), the usability of the material can be first determined. As described in the foregoing embodiment, factors such as the edge complexity of the main object, the number of holes on the main object, the area of the main object, and the overlap degree between the main object and the characters can be obtained. If the characteristics of the main object in each of the above dimensions meet the requirements, it is determined that the material is usable, and thus the material can be cropped.

[0065] In some embodiments, a scene picture or a creative template picture can be generated based on the screened target commodity pictures. A scene picture refers to a composite picture generated by combining the main object (such as a commodity, an item, a person, etc.) in the target commodity picture with a preset background or environment (scene). Such a picture not only shows the appearance of the commodity in the target commodity picture, but also makes the commodity appear more suitable for the actual use scene or marketing scene through the design of the background, thereby helping consumers better understand the use, applicable environment, or purchase value of the commodity. Continuing with the previous example, for directly usable materials, a background picture can be directly generated, and a scene picture can be generated based on the material and the background picture. For the picture of the main object obtained by cropping, a scene picture can be generated based on the picture of the main object.

[0066] In some embodiments, the selected target product images can be sent to the merchant for the merchant to use. Alternatively, the selected target product images can also be displayed on the product recommendation page of the e-commerce platform. Specifically, when a product recommendation request sent by the client is received, the target product images can be sent to the client for display in response to the product recommendation request. Or, the selected target product images can also be used for advertising promotion, marketing, or display on other channels or platforms outside the e-commerce platform. For example, the product images may be used for social media advertising, search engine advertising, email marketing, third-party e-commerce platforms, display ads (such as banner ads), etc. By optimizing the visual effects of these images, the click-through rate and purchase conversion rate of users can be increased, and the brand image and product attractiveness can be enhanced.

[0067] The picture screening solution of the present application has the following technical effects:

[0068] (1) Integrating traffic efficiency and aesthetic evaluation, constructing an integrated algorithm process for picture material recall, picture quality fine screening, and material usability determination, achieving the landing effect of automatic picture quality screening without manual intervention in the e-commerce scenario, and achieving a positive improvement in core indicators such as the click-through rate of product pictures and gmv_roi.

[0069] (2) Designing a value formula according to different stages of the conversion of picture materials in the consumption link of the e-commerce scenario, and optimizing parameters through efficiency data to determine the hyperparameter values (i.e., the adjustment coefficients corresponding to each stage), estimating the gmv_roi of picture material placement, and screening high-value pictures.

[0070] (3) Using a comprehensive classification model to fuse picture atomic features such as the embedding representation and aesthetic score of pictures using DCN_v2, and achieving better results than ordinary picture classification models in the e-commerce scenario.

[0071] (4) Proposing a feasible solution for the e-commerce picture scenario based on the contour features of pictures, extracting information such as the main shape, complexity, and connectivity in the pictures to determine whether the cut-out main body is available.

[0072] As described above, the generated target product images can be further used to synthesize scene images. In the related art, a general picture editing model is usually used to edit the original picture to obtain the output picture. In the e-commerce scenario, the scene in the target product image can be edited and synthesized through a general picture editing model to obtain the scene image. However, the picture editing operation will change the visual features of the main object in the target product image, resulting in the inconsistency between the visual features of the main object and the features of the actual product.

[0073] Based on this, the present application provides a method for processing product images, which generates scene description information that matches the attribute information of the main object in the target product image based on a large language model, and generates a scene image based on the scene description information. Since the attribute information of the main object includes the contour information of the main object, the contour of the main object in the generated scene image is consistent with the contour of the main object in the target product image. Then, the main object in the scene image is replaced with the main object in the target product image, so that the features of the main object in the scene image are consistent with those of the main object in the target product image. The above method can automatically generate a scene image of the main object with consistent features with the target product image without manual synthesis, reducing the cost of the scene image synthesis process. The implementation details of the present application will be illustrated with reference to the accompanying drawings below.

[0074] See Figure 4 , the present application provides a method for processing product images, and the method includes:

[0075] Step S22: For any one of at least one target product image, obtain the attribute information of the main object in the target product image, where the attribute information includes the contour information of the main object;

[0076] Step S24: Generate scene description information of the target scene that matches the attribute information through a pre-trained large language model;

[0077] Step S26: Generate a scene image through an image generation model, where the scene image includes the target scene described by the scene description information and a main object with the contour described by the contour information;

[0078] Step S28: Replace the main object in the scene image with the main object in the target product image.

[0079] The target product image in the embodiment of the present application can be screened from multiple candidate product images based on the method in the foregoing embodiment, or can be obtained based on other methods, and the present application does not limit this.

[0080] In step S22, the attribute information of the main object in the target product image can be obtained. Among them, the main object in the target product image can be the most important and core product or item in the image, usually the object that consumers focus on and purchase. The attribute information of the main object is used to describe the characteristics of the main object. The attribute information can include the contour information of the main object, and the contour information is used to describe the contour characteristics of the main object. In addition, the attribute information of the main object can also include, but is not limited to, information about the type, color, shape, transparency, function, material, size, applicable object, applicable scene, and / or price and other characteristics of the main object.

[0081] In some embodiments, the target product image can be input into a Vision-Language Model (VLM) so that the vision-language model extracts the attribute information of the main object in the target product image based on the target product image. Among them, the vision-language model can adopt the LLaVA (Language and Vision Assistant) model or other models that simultaneously have vision processing capabilities and language processing capabilities. In particular, for the contour information in the attribute information, a dedicated edge detection model can be used to extract the contour information of the main object in the target product image to accurately extract the contour information.

[0082] In step S24, at least part of the extracted attribute information can be input into a pre-trained Large Language Model (LLM). Since the large language model has been exposed to a large amount of data during training, and this data contains rich world knowledge, common sense reasoning, relationships between objects and environments, etc., the large language model can infer the environment, background, situation, etc. suitable for the main object based on the attribute information of the main object, and thus generate reasonable and practical scene description information. The scene description information includes but is not limited to the environment and background where the scene is located (such as time, weather, season, atmosphere, etc.), people and items in the scene, events occurring in the scene, visual information of the scene, cultural and social backgrounds, and / or perspectives and narrative angles, etc.

[0083] For example, if the given main object is a "kettle", the large language model will, based on its common sense, infer that it may appear in a kitchen, at a dining table, or in an outdoor camping scene, and thus generate scene description information related to the kitchen, dining table, or outdoor camping scene. Similarly, if the main object is "skis", the large language model will infer that it is suitable for specific environments such as snow-capped mountains and ski resorts, and thus generate scene description information related to specific environments such as snow-capped mountains and ski resorts.

[0084] In step S26, the scene description information and the attribute information of the main object can be input into the image generation model, so that the image generation model generates a scene image including the target scene described by the scene description information and the main object with the contour described by the contour information. Among them, the attribute information input into the image generation model at least includes the contour information of the main object, and this contour information can be used as a constraint condition to make the contour of the main object in the scene image generated by the image generation model consistent with the contour of the main object in the target commodity image. For example, when the main object in the target commodity image is a kettle and the scene description information is description information related to the kitchen, the generated scene image is an image with the kitchen as the background and including the kettle. And, the contour of the kettle in the scene image is consistent with the contour of the kettle in the target commodity image, but there may be certain differences in other visual features of the kettle in the scene image and the kettle in the target commodity image.

[0085] In some embodiments, the image generation model includes an SD (Stochastic Depth) model and a ControlNet model. Among them, the SD model is used to obtain the scene description information and generate the target scene and the main object in the scene image based on the scene description information. The ControlNet model is used to obtain the contour information of the main object and use this contour information as a constraint condition to constrain the contour of the main object in the scene image generated by the SD model.

[0086] In the related art, the SD inpainting scheme is usually adopted to generate the scene image. However, this method has a high dependence on resources and is not suitable for running on devices with limited resources. The embodiment of the present application adopts the SD model and the ControlNet model as the image generation model. Compared with the SD inpainting scheme, the dependence on data and computing power resources is low, and the implementation cost is low.

[0087] In some embodiments, the image generation model can be trained in a mask-based manner. In the mask-based training method, a first sample image can be obtained, the image area where the main object in the first sample image is located can be masked, and the image generation model can be used to restore this image area to obtain a predicted sample image. According to the difference between the predicted sample image and the first sample image, the image generation model can be trained.

[0088] Optionally, the picture area where the main object is located may only include the pixel points corresponding to the main object, and does not include the pixel points corresponding to other objects outside the main object, that is, the area of the picture area where the main object is located is equal to the area of the main object. Or, optionally, in addition to including the pixel points corresponding to the main object, the picture area where the main object is located may also include the pixel points corresponding to other objects outside the main object, that is, the area of the picture area where the main object is located is greater than the area of the main object. The applicant has found that when training a picture generation model in a mask-based manner, if only the pixel points corresponding to the main object are masked, the pictures generated by the trained picture generation model may have poor visual effects in the edge area of the main object. For example, the transition between the edge of the main object and the background area is not natural enough. By masking the pixel points corresponding to other objects outside the main object in the first sample picture, so that the area of the picture area where the main object is located is greater than the area of the main object, it is possible to make the pictures generated by the trained picture generation model have better visual effects in the edge area of the main object.

[0089] In some embodiments, the pictures generated by the picture generation model trained in the above manner may be relatively in line with the aesthetic standards of people, but may not meet the traffic requirements of the e-commerce scenario. Therefore, after training the picture generation model in the above manner, it is also possible to obtain a second sample picture with a traffic efficiency higher than a preset efficiency threshold, and fine-tune the trained picture generation model based on the second sample picture. In this way, the picture generation model can learn the general characteristics of pictures with high traffic efficiency in the e-commerce scenario, making the generated pictures more in line with the requirements of the e-commerce scenario.

[0090] In some embodiments, the number of picture generation models may be greater than or equal to 1. When the number of picture generation models is greater than 1, different picture generation models may be trained for the product pictures corresponding to the main objects of different categories. For example, the picture generation model trained for the product pictures corresponding to the products in the clothing category may be different from the picture generation model trained for the product pictures corresponding to the products in the electronic device category. Each trained picture generation model may be applicable to processing the product pictures corresponding to one or more categories of products. On this basis, after obtaining the target product picture, it is possible to first determine the target category to which the main object in the target product picture belongs, and then generate a scene picture based on the image generation model corresponding to the target category. The scene picture includes the target scene described by the scene description information and the main object that has the contour described by the contour information and belongs to the target category. For example, when the main object in the target product picture is a mobile phone, the picture generation model corresponding to the mobile phone can be used to generate the scene picture; while when the main object in the target product picture is clothes, the picture generation model corresponding to the clothes can be used to generate the scene picture.

[0091] In step S28, the main object in the scene picture can be replaced with the main object in the target product picture. On the one hand, since the outline of the main object in the scene picture is consistent with the outline of the main object in the target product picture, the main object in the target product picture can fit well with the main object in the scene picture. On the other hand, after the replacement operation is performed, the visual features of the main object in the scene picture will be consistent with the visual features of the main object in the target product picture, and thus be consistent with the visual features of the actual product.

[0092] In some embodiments, the perspective of the main object in the scene picture can also be adjusted. Adjusting the perspective of the main object in the scene picture means changing the display angle of the main object in the scene picture, thereby affecting its presentation, visual effect, and the audience's understanding and perception of it. In scenarios such as e-commerce, advertising design, and film production, adjusting the perspective can be used to highlight the features of the main object, enhance visual appeal, or improve the user experience. Adjusting the perspective of the main object includes but is not limited to: mutual adjustment between the main object in the front view and the main object in the side view, mutual adjustment between the main object in the upward view and the main object in the downward view, mutual adjustment between the main object under a long focal length lens and the main object under a short focal length lens, etc.

[0093] Specifically, after replacing the main object in the scene picture with the main object in the target product picture, the perspective of the replaced main object in the scene picture can be adjusted. Or, the perspective of the main object in the scene picture and the perspective of the main object in the target product picture can also be adjusted respectively first, and then the main object in the adjusted scene picture is replaced with the main object in the adjusted target product picture.

[0094] In some embodiments, after generating the scene picture, the first aesthetic score of the scene picture and the second aesthetic score of the target product picture used to generate the scene picture can also be compared. If the first aesthetic score is lower than the second aesthetic score, the step of obtaining the attribute information of the main object in the target product picture is returned. Among them, the first aesthetic score and the second aesthetic score can be obtained by using the picture quality evaluation model in the foregoing embodiments. By adopting this embodiment, the generated scene picture can have a higher aesthetic score than the original target product picture.

[0095] Figure 5The overall flowchart of the scenario picture synthesis process of the embodiments of the present application is shown. In the e-commerce platform, the visual effect of product pictures plays a crucial role in attracting users' attention. High-quality and beautiful product pictures can quickly attract consumers' attention and enhance users' willingness to purchase. Currently, there are the following problems with the product pictures on the e-commerce platform:

[0096] (1) The quality of product pictures is not high enough to directly reflect the use and usage scenarios of the products. Recommending such materials to users is difficult to arouse users' interest;

[0097] (2) The styles of product pictures are homogenized, making it difficult to meet the diverse interest preferences of different user groups.

[0098] To truly solve the above business pain points, the present application designs and implements a large model-based product generation picture traffic guidance solution under resource constraints such as picture annotation data and computing power, realizing an integrated process from picture quality optimization to batch generation of scenario pictures based on the large model, optimization of generation effects, and then to evaluation of the usability of generated pictures. This solution generates scenario pictures that can display the value and usage scenarios of products based on existing product pictures. Under the constraints of computing power and data resources, the present application can arouse users' interest while improving the quality and aesthetics of pictures, increase the upper limit of content distribution while increasing the diversity of product picture styles, and thus promote the improvement of traffic efficiency indicators.

[0099] Most traditional scenario picture generation adopts the inpainting scheme based on SD, but the existing open-source SD inpainting model generally has poor implementation effects in the e-commerce scenario. To achieve the goal of implementing e-commerce scenario picture production, further effect optimization is required. However, the full-parameter effect optimization based on SD relies on a large amount of high-quality product scenario picture library data and computing power resources, which will incur a large amount of manual annotation and data procurement costs and is difficult to achieve good implementation effects under resource constraints such as picture annotation data and computing power.

[0100] The specific steps of the present application may include:

[0101] (1) Picture optimization: Based on product efficiency consumption data, multiple candidate product pictures are initially screened in a parameter optimization manner, then quality optimization is performed based on a picture quality evaluation model, then the main body of the selected picture materials is extracted, and finally, the usability of the product main body (after cropping) is determined based on the contour features of the main object, so as to screen out the target product pictures.

[0102] (2)Efficient implementation of e-commerce scenarios under resource constraints: Based on the target product images generated in step (1), first obtain attribute information such as the content and color of the main object based on the LLaVA model, then generate scene description information suitable for the main object based on the above attribute information and the LLM model, and finally generate scene images under the control conditions of the ControlNet based on the image SD model. This solution has better effects than the solution based on SD inpainting, and has low dependence on data and computing resources and low implementation costs. Since only the outline of the main object in the original target product image remains unchanged in the scene images generated at this stage, the main object in the original target product image is synthesized in the second stage to ensure that the main object in the scene image is the same as the main object in the target product image.

[0103] (3)Optimization of traffic-oriented effects: Further optimize the effects of the scene images generated in step (2) through scene pipeline adaptation and optimization of the perspective of the product main body. To further achieve low-cost model optimization under data and computing constraints, this application screens out sample images with high traffic efficiency (i.e., the second sample images), which are directly used for large model generation optimization after quality filtering. This dataset takes into account quality while being more in line with the user preferences of e-commerce scenarios and does not rely on manual annotation resources. Based on this traffic-oriented dataset, low-cost classification category fine-tuning based on SD+controlnet can be achieved to further optimize the scene image generation effect.

[0104] (4)Evaluation of generation effects: After completing steps (1) to (3), this application evaluates the generation availability of the high-quality scene images and the corresponding target product images output by the algorithm based on image aesthetics scores to further ensure the quality, aesthetics, and availability of the scene image generation.

[0105] Schematic diagrams of the original target product images and the generated scene images are as Figure 6 shown. In the first stage, the image generation model generates first-stage scene images based on the scene description information and the attribute information of the main object. It can be seen that in the scene images generated at this stage, the outline of the main object is the same as the outline of the main object in the target product image, but other visual features are not exactly the same. In the second stage, the main object in the target product image is replaced into the scene image to obtain the second-stage scene image. It can be seen that in the scene images generated at this stage, the visual features of the main object are the same as the outline of the main object in the target product image.

[0106] The scene image generation solution of this application has the following advantages:

[0107] (1) Under the resource constraints of picture annotation data, computing power, etc., first generate scene description information suitable for products based on VLM and LLM, and then directly generate scene pictures based on the picture generation model (SD model) under the control conditions of ControlNet. The direct picture generation effect based on SD+ControlNet is superior to the SD inpainting scheme, with low dependence on data and computing power resources and low implementation cost.

[0108] (2) When directly generating scene pictures based on SD+ControlNet, only the outline of the main object remains unchanged, and the details of the main object may change. Therefore, in this application, the original target product picture and the first-stage scene picture are synthesized in the second stage to ensure that while generating an available scene for the product, the main body of the product is exactly the same as the original picture.

[0109] (3) To further achieve low-cost model tuning under data and computing power constraints, this application filters high-efficiency traffic-oriented pictures based on traffic efficiency data, and directly uses them for tuning the picture generation model after quality filtering. This data set takes into account quality and is more in line with the preferences of users in this scenario, and is completely independent of manual annotation resources. Based on this traffic-oriented data set, low-cost category-specific fine-tuning based on SD+ControlNet can be achieved to further optimize the scene picture generation effect.

[0110] (4) To ensure the usability of the output pictures, this application evaluates the usability of the scene pictures output by the algorithm and the original target product pictures based on image aesthetics scores to ensure the final quality of the scene pictures.

[0111] See Figure 7 , this application also provides a product picture processing device, and the device includes:

[0112] The first acquisition module 102 is used to acquire multiple candidate product pictures on the e-commerce platform;

[0113] The determination module 104 is used to acquire the traffic parameters of the multiple candidate product pictures, and determine the traffic efficiency of the multiple candidate product pictures based on the traffic parameters of the multiple candidate product pictures; wherein, the traffic parameters of the candidate product pictures are used to characterize the access traffic of the users on the e-commerce platform to the candidate product pictures, and the traffic efficiency of the candidate product pictures is used to characterize the access traffic obtained by the candidate product pictures under the preset product picture display cost;

[0114] The second acquisition module 106 is used to acquire the quality of the multiple candidate product pictures through a pre-trained picture quality evaluation model;

[0115] A screening module 108, configured to screen at least one target product image from the multiple candidate product images based on the traffic efficiency and quality of the multiple candidate product images.

[0116] See Figure 8 , this application further provides a product image processing device, the device includes:

[0117] A third acquisition module 202, configured to acquire, for any one of the at least one target product image, attribute information of the main object in the target product image, where the attribute information includes contour information of the main object;

[0118] A first generation module 204, configured to generate scene description information of a target scene that matches the attribute information through a pre-trained large language model;

[0119] A second generation module 206, configured to generate a scene image through an image generation model, where the scene image includes the target scene described by the scene description information and the main object with the contour described by the contour information;

[0120] A replacement module 208, configured to replace the main object in the scene image with the main object in the target product image.

[0121] The functions or modules included in the device provided in this application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0122] An embodiment of this application further provides a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the program, it implements the method described in any one of the foregoing embodiments.

[0123] Figure 9 FIG. shows a more specific schematic hardware structure diagram of a computer device provided in an embodiment of this application. The device may include: a processor 302, a memory 304, an input / output interface 306, a communication interface 308, and a bus 310. Wherein, the processor 302, the memory 304, the input / output interface 306, and the communication interface 308 are communicatively connected to each other inside the device through the bus 310.

[0124] The processor 302 can be implemented in the form of a general - purpose central processing unit (CPU), a microprocessor, an application - specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The processor 302 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.

[0125] The memory 304 can be implemented in the form of a read - only memory (ROM), a random - access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 304 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present application through tools or firmware, the relevant program codes are stored in the memory 304 and are called and executed by the processor 302.

[0126] The input / output interface 306 is used to connect to the input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0127] The communication interface 308 is used to connect to a communication module (not shown in the figure) to realize the communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or can achieve communication through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0128] The bus 310 includes a path for transmitting information between various components of the device (such as the processor 302, the memory 304, the input / output interface 306, and the communication interface 308).

[0129] It should be noted that although the above - mentioned device only shows the processor 302, the memory 304, the input / output interface 306, the communication interface 308, and the bus 310, in the specific implementation process, this device may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the above - mentioned device may also only include the components necessary to implement the solutions of the embodiments of the present application, and does not necessarily include all the components shown in the figure.

[0130] An embodiment of the present application provides a computer program product, including a computer program which, when executed by a processor, implements the method described in any embodiment of the present application.

[0131] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any of the foregoing embodiments.

[0132] Computer-readable media include both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computer device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0133] Each embodiment in the present application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of the present application, the functions of the modules can be implemented in the same or multiple tools and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0134] The above is only the specific implementation manner of the embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the embodiments of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of the present application.

Claims

1. A method for processing a product image, the method comprising: Obtain multiple candidate product images from the e-commerce platform; Obtaining the traffic parameters of the plurality of candidate product images, and determining the traffic efficiency of the plurality of candidate product images based on the traffic parameters of the plurality of candidate product images; wherein the traffic parameters of the candidate product images are used to characterize the access traffic of users of the e-commerce platform to the candidate product images, and the traffic efficiency of the candidate product images is used to characterize the access traffic obtained by the candidate product images under a preset product image display cost; Obtaining the quality of the plurality of candidate product images through a pre-trained image quality assessment model; At least one target product image is selected from the plurality of candidate product images based on the flow efficiency and quality of the plurality of candidate product images.

2. According to the method of claim 1, the traffic parameters include traffic parameters of multiple stages in the commodity consumption link; the determining the traffic efficiency of the multiple candidate commodity images based on the traffic parameters of the multiple candidate commodity images comprises: Acquire adjustment coefficients corresponding to the multiple stages respectively, where the adjustment coefficient corresponding to any stage is used to adjust the contribution of the flow parameters of the stage to the flow efficiency; Adjusting the flow parameters of the candidate product images in the corresponding stages based on the adjustment coefficients respectively corresponding to the multiple stages; Based on the adjusted traffic parameters of the multiple stages, the traffic efficiency of the multiple candidate product images is determined.

3. According to the method of claim 2, the traffic parameters of the candidate product image at multiple stages include the exposure rate, click-through rate, conversion rate of the candidate product image, the transaction amount of the product corresponding to the candidate product image, and the product image display cost corresponding to the candidate product image; The adjustment coefficients corresponding to the multiple stages include an exposure rate adjustment coefficient, a click rate adjustment coefficient, a conversion rate adjustment coefficient, a transaction amount adjustment coefficient and a cost adjustment coefficient; in, The exposure rate adjustment coefficient, click rate adjustment coefficient, conversion rate adjustment coefficient and transaction amount adjustment coefficient are used to adjust the contribution of exposure rate, click rate, conversion rate and transaction amount to traffic efficiency respectively, and the cost adjustment coefficient is used to adjust the contribution of the product image display cost to traffic efficiency.

4. The method according to claim 3, further comprising: Obtaining the traffic parameters of the candidate product images at the multiple stages in the first time period within a preset period; Determining the flow efficiency of the candidate product image within the preset period based on the initial adjustment coefficients respectively corresponding to the multiple stages and the flow parameters of the multiple stages in the first time period; Determine estimated flow parameters of the candidate product image at the multiple stages in a second time period within the preset period based on the flow efficiency of the candidate product image within the preset period and the initial adjustment coefficients respectively corresponding to the multiple stages; Based on the difference between the estimated traffic parameters of the candidate product image in the multiple stages in the second time period and the traffic parameters of the candidate product image in the corresponding stages in the second time period, the initial adjustment coefficients corresponding to the multiple stages are adjusted to obtain the adjustment coefficients corresponding to the multiple stages.

5. According to the method of claim 1, the picture quality assessment model comprises a plurality of first sub-models for assessing the quality scores of the candidate product pictures in a plurality of dimensions, and a second sub-model for assessing the comprehensive quality scores of the candidate product pictures; the step of obtaining the quality of the plurality of candidate product pictures by using the pre-trained picture quality assessment model comprises: Inputting the candidate product images into the multiple first sub-models respectively, so as to obtain quality scores of the candidate product images in the multiple dimensions through the multiple first sub-models; Inputting the candidate product images and the quality scores of the candidate product images in the multiple dimensions into the second sub-model, so as to obtain the comprehensive quality scores of the candidate product images through the second sub-model; The quality of the candidate product image is determined based on the comprehensive quality score.

6. The method according to claim 1, wherein the selecting at least one target product image from the plurality of candidate product images based on the flow efficiency and quality of the plurality of candidate product images comprises: Selecting a number of candidate product images whose flow efficiency is higher than a preset flow efficiency threshold from the multiple candidate product images; Based on the qualities of the plurality of candidate product images, at least one target product image is selected from the plurality of candidate product images.

7. According to the method of claim 1, the selecting at least one target product image from the plurality of candidate product images based on the flow efficiency and quality of the plurality of candidate product images comprises: Based on the flow efficiency and quality of the multiple candidate product images, at least one candidate target product image is selected from the multiple candidate product images; At least one target product picture whose subject object meets a preset condition is selected from the at least one candidate target product picture; wherein the preset condition is determined based on at least one of the following: The edge complexity of the main object, the number of holes on the main object, the area of ​​the main object, and the overlap between the main object and the character.

8. The method according to claim 1, comprising: For any one of the at least one target product image, acquiring attribute information of a main object in the target product image, wherein the attribute information includes contour information of the main object; Generate scene description information of the target scene matching the attribute information through a pre-trained large language model; Generate a scene picture by using a picture generation model, wherein the scene picture includes a target scene described by the scene description information and a main object having a contour described by the contour information; The main object in the scene image is replaced with the main object in the target product image.

9. The method according to claim 8, before generating the scene picture by the picture generation model, the method further comprises: Get the first sample image; Masking the image region where the main object in the first sample image is located; wherein the area of ​​the image region where the main object is located is larger than the area of ​​the main object; Restoring the image region by using the image generation model to obtain a predicted sample image; The picture generation model is trained based on feature differences between the first sample picture and the predicted sample picture.

10. The method according to claim 9, after training the picture generation model based on the feature difference between the first sample picture and the predicted sample picture, the method further comprises: Acquire a second sample image whose flow efficiency is higher than a preset efficiency threshold; Fine-tune the trained image generation model based on the second sample image.

11. The method according to claim 8, further comprising: The viewing angle of the main object in the scene picture is adjusted.

12. The method according to claim 8, further comprising: comparing a first aesthetic score of the scene image and a second aesthetic score of a target product image used to generate the scene image; If the first aesthetic score is lower than the second aesthetic score, return to the step of obtaining attribute information of the main object in the target product image.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 12.

14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 12 when executing the computer program.

15. A computer program product, comprising a computer program, which implements the method according to any one of claims 1 to 12 when executed by a processor.

Citation Information

Cited By

  • Image processing method and device, image display method and device, equipment and storage medium

    CN120765355A