Product image processing method, device, storage medium and equipment
By training attention model to evaluate and beautify product pictures uploaded by e-commerce platform merchants, the problem of inconsistent picture quality is solved, and automated image quality improvement and commercial effect improvement are achieved.
Patent Information
- Application Number
- CN202111589259.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-12-23
AI Technical Summary
The quality of the product pictures uploaded by merchants on e-commerce platforms is uneven, which affects the platform's reputation and consumes a lot of human resources by manual review.
Through the trained attention model, multiple feature information of the product picture are processed, the total score and weight of the feature information are determined, and the pictures are beautified based on this.
Automatic image quality evaluation and beautification are realized, image quality is improved, thus improving product click-through rate and purchase rate, and the beautification results are more reasonable.
Smart Images

Figure CN114298930B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, storage medium, and device for processing product images. Background Art
[0002] With the development of e-commerce, more and more people are shopping online. Product images displayed on e-commerce platforms are crucial features of products, influencing consumer purchasing behavior. However, these product images are typically uploaded by merchants, and can suffer from poor quality or content that doesn't comply with regulations, potentially significantly impacting the platform's reputation and profitability. Given the large number of merchants, manual review of uploaded product images requires significant human resources and places a significant management burden on the platform.
[0003] Therefore, the market is in urgent need of a solution that can effectively manage product images. Summary of the Invention
[0004] According to a first aspect of an embodiment of this specification, a method for processing a product image is provided, comprising:
[0005] Acquire multiple feature information of the input product image; wherein one feature information is related to the image quality of the product image in a quality dimension;
[0006] The trained attention model is used to process the multiple feature information to determine the total image quality score of the product image in each quality dimension related to each feature information and the weight of each feature information, where the weight of each feature information is positively correlated with the contribution of the feature information to the total score;
[0007] Based on the total score and the weight, the product image is beautified in the quality dimension that needs beautification.
[0008] In some examples, each feature information is obtained by processing the product image through a detection model.
[0009] In some examples, the plurality of feature information is processed by a trained attention model to determine an overall image quality score of the product image in multiple quality dimensions and a weight of each feature information, including:
[0010] Get the feature vector corresponding to each feature information;
[0011] Concatenating the feature vectors corresponding to the plurality of feature information to obtain an input vector;
[0012] The input vector is input into the attention model to obtain the weight and the total score.
[0013] In some examples, the feature vectors corresponding to the plurality of feature information are concatenated to obtain an input vector, including:
[0014] Perform dimensionality reduction processing on each eigenvector to obtain the reduced dimensionality vector corresponding to each eigenvector;
[0015] Concatenate the reduced dimensionality vectors.
[0016] In some examples, the attention model includes a sparse self-attention model.
[0017] In some examples, based on the total score and the weight, beautification processing is performed on the quality dimensions of the product image that need to be beautified, including:
[0018] If the total score is lower than the quality score threshold, the quality dimension that needs to be beautified is determined based on the weight.
[0019] In some examples, quality dimensions that require improvement are determined based on weights, including:
[0020] Determine the quality dimension corresponding to the feature information whose weight exceeds the weight threshold as the quality dimension that needs to be beautified; or
[0021] The quality dimensions corresponding to several pieces of feature information with weights from high to low are determined as the quality dimensions that need to be beautified.
[0022] In some examples, the method further includes:
[0023] If the quality dimension that needs to be beautified includes a specified quality dimension, a prompt message is output to prompt that the product image violates the business rules.
[0024] In some examples, the method further includes:
[0025] After the beautification process, the process returns to the step of obtaining multiple feature information for the beautified product image.
[0026] According to a second aspect of the embodiments of this specification, a product image processing device is provided, including:
[0027] An acquisition module is used to acquire multiple feature information of the input product image; wherein one feature information is used to characterize the image quality of the product image in a quality dimension;
[0028] a determination module, configured to process the plurality of feature information using a trained attention model to determine an overall image quality score of the product image across multiple quality dimensions and a weight for each feature information, wherein the weight for each feature information is positively correlated with the contribution of the feature information to the overall score;
[0029] A beautification module is used to beautify the quality dimensions that need to be beautified in the product image based on the total score and the weight.
[0030] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, any one of the methods in the embodiments of the specification is implemented.
[0031] According to a fourth aspect of the embodiments of this specification, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any one of the methods in the embodiments of the specification is implemented.
[0032] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:
[0033] In the embodiments of this specification, a method, apparatus, storage medium and device for processing product images are disclosed. In this method, a plurality of feature information of a product image is processed by a trained attention model to determine the total score of the image quality of the product image in the quality dimension related to each feature information and the weight of each feature information, and then the product image is beautified based on the total score and the weight. In this way, it is possible to automatically perform quality assessment and beautification processing on the product images uploaded by merchants, thereby improving the image quality and thus improving the click-through rate and purchase rate of the corresponding products. Moreover, due to the use of the attention mechanism, attention is paid to the relationship between the various quality dimensions in the product image, thereby achieving targeted automatic beautification of the product image, making the beautification result more reasonable.
[0034] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.
[0036] Figure 1 This is a flowchart of a method for processing product images according to an exemplary embodiment of this specification;
[0037] Figure 2is a schematic diagram of an image quality API according to an exemplary embodiment of this specification;
[0038] Figure 3 This is a hardware structure diagram of a computer device where a product image processing device is located according to an exemplary embodiment of this specification;
[0039] Figure 4 This is a block diagram of a device for processing product images according to an exemplary embodiment of the present specification. DETAILED DESCRIPTION
[0040] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this specification, as detailed in the appended claims.
[0041] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. As used in this specification and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0042] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information without departing from the scope of this specification. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."
[0043] With the development of e-commerce, more and more people are shopping on the Internet. The product images displayed on the e-commerce business platform are an important feature of the products and affect consumers' purchasing behavior. However, these product images are usually uploaded by merchants, and there may be problems such as poor image quality or non-compliance with regulations, which can easily have a significant impact on the reputation and interests of the platform. Considering the large number of merchants, manual review of the product images uploaded by merchants requires a large amount of human resources and will also place a large management burden on the platform. Based on this, the embodiments of this specification provide a product image processing solution to solve the above problems.
[0044] Next, the embodiments of this specification are described in detail.
[0045] like Figure 1 As shown, Figure 1 This is a flowchart of a method for processing product images according to an exemplary embodiment of this specification. The method can be applied to a client of a supplier of products, or to a server of a business party providing services to the supplier. In some practical applications, the business party may be a business party of a food delivery platform, and correspondingly, the supplier may be a merchant on the food delivery platform.
[0046] The method comprises:
[0047] Step 101: Acquire multiple feature information of the input product image; wherein one feature information is related to the image quality of the product image in a quality dimension;
[0048] For merchants, they can display product information to consumers by entering product information on the business platform. When merchants enter product information, in addition to entering the product name and price, they usually also enter product pictures to attract consumers to click and buy. Since the image quality of product pictures will affect the click-through rate and purchase rate, it is necessary to conduct a quality assessment of the product pictures before showing them to consumers. It should be noted that the picture quality mentioned in this embodiment can include the aesthetic quality of the picture, that is, people's subjective evaluation of the visual experience of the picture, and can also include the degree of qualification of the picture in terms of business rules, that is, the objective evaluation of the picture in terms of business rules.
[0049] The quality dimensions mentioned in this step can be considered as a quality assessment component for product images, such as lighting and clarity. Correspondingly, each feature information can be considered as an assessment result corresponding to a specific assessment component, such as whether the lighting is dim or the image is clear. Of course, this assessment result can also be a score for a single dimension, such as a score representing the brightness of the lighting or a score representing the clarity of the image. The granularity of feature information expression can be limited based on the needs of the specific scenario. For example, in scenarios with more stringent requirements for product images, the granularity of feature information expression can be finer.
[0050] Since the higher the aesthetic quality of a product image, the easier it is to attract user attention, in some examples, the multiple feature information may include feature information characterizing the aesthetic quality of the image, such as feature information characterizing whether the light of the image is dim or uneven, feature information characterizing the clarity of the image, feature information characterizing whether the image has spots and defects, etc.
[0051] At the same time, if the source of the product image is illegal or the image content is unreasonable, it will also affect the image quality. Therefore, in other examples, the multiple feature information may include feature information related to the source of the image and / or feature information related to the image content.
[0052] Among them, feature information related to the source of the picture includes feature information indicating whether the picture is a reproduced picture (a picture obtained by photographing a picture using a photographing device), feature information indicating whether the picture is a spliced picture (a picture composed of at least two pictures spliced together), feature information indicating whether the picture has been processed twice (the picture has been modified and saved after being reproduced), etc.; it can also include feature information related to the content of the picture, such as feature information indicating whether the picture has a LOGO (including LOGO text and / or LOGO pattern), feature information indicating whether the items in the picture are contained in a container or are packaged, feature information indicating whether the items in the picture are consistent with the product name, etc.
[0053] In addition, the feature information related to the image content can be related to the specific category of the product. For example, if the product category is food, the feature information related to the image content may include feature information indicating whether the item in the image is directly edible; if the product category is toys, the feature information related to the image content may include feature information indicating whether the item in the image has sharp parts.
[0054] In some examples, each feature information is obtained by processing a product image through a detection model. The detection model can be a machine learning model for detecting the image quality of the product image in a quality dimension. Taking the feature information as the feature information that characterizes the clarity of the image as an example, the corresponding detection model can be a machine learning model for detecting the clarity of the image, which can be trained based on multiple sample images and the clarity label of each sample image, wherein the clarity label can be obtained based on the scoring results of multiple users scoring the clarity of the corresponding image. In addition, the detection model can be any one of the machine learning models such as the Convolutional Neural Networks (CNN) model and the Support Vector Machine (SVM) model, and this embodiment does not limit this. The types of models for detecting different dimensions can be the same or different.
[0055] Step 102: Process the multiple feature information using the trained attention model to determine the total image quality score of the product image in each quality dimension related to each feature information and the weight of each feature information, where the weight of each feature information is positively correlated with the contribution of the feature information to the total score;
[0056] After obtaining multiple feature information of the product image, this multiple feature information can be used to obtain an overall score representing the comprehensive quality of the product image. In some embodiments, the quality score of the product image in each quality dimension can be first obtained based on the feature information of each dimension, and the quality scores in each dimension can be directly averaged, or the quality scores in each dimension can be weighted and summed using fixed parameters obtained from the XGBoost (Extreme Gradient Boosting) algorithm to obtain the overall score.
[0057] The method of calculating the total score provided in the above example has the advantages of fast calculation speed and easy implementation. However, it does not take into account the relationship between the various quality dimensions of the product image. Therefore, its calculation result may not be comprehensive and accurate, and it is also unable to give the key reasons for the low total score of the current image. This embodiment adopts the attention model to solve this problem. The attention model is a neural network model based on the attention mechanism, which is widely used in deep learning tasks such as natural language processing, image recognition, and speech recognition. The attention mechanism is an algorithm that simulates a person's selective attention to certain things at a certain moment. It can find the correlation between the various parts of the data and allocate corresponding attention based on the importance of each part of the data. This attention is the weight. With the help of the attention mechanism, this embodiment can capture the correlation between multiple feature information of the product image, and dynamically calculate the weight to more reasonably calculate the total score of the product image, and the obtained weight can be used to indicate the key reasons affecting the total score.
[0058] In some examples, this step may include: obtaining a feature vector corresponding to each feature information; concatenating the feature vectors corresponding to the multiple feature information to obtain an input vector; and inputting the input vector into the attention model to obtain the weight and the total score. By extracting feature vectors from multiple feature information respectively and concatenating them, a vector sequence of feature vectors containing multiple feature information can be obtained, which is used as the input vector of the attention model, and then the weight and the total score output by the attention model can be obtained. For example, for 16 feature vectors corresponding to 16 feature information, the Concat function can be used to concatenate these feature vectors, and the connected feature vectors can be rearranged to form an input vector, wherein the product of the row dimension and the column dimension of the input vector can still be the sum of the dimensions of the 16 feature vectors.
[0059] Considering that the attention mechanism consumes significant video memory resources when implemented, in some examples, the input vector can be obtained by performing dimensionality reduction on each feature vector, obtaining the corresponding reduced-dimensionality vector for each feature vector, and then concatenating the reduced-dimensionality vectors. For example, for K features, N-dimensional features are first extracted from each feature as a feature vector. Each feature vector is then compressed to M dimensions to obtain a reduced-dimensionality vector. These K reduced-dimensionality vectors are then concatenated to obtain an M*K dimensional input vector that is fed into the attention model. This feature compression reduces the computation of correlations, meaning that each element is considered to be related only to a subset of elements in the sequence, thereby conserving video memory and accelerating computation. Alternatively, the ratio of the feature vector dimension to the reduced-dimensionality vector dimension can be the number of features, that is, N / M = K. This ensures that the input vector and the feature vector dimensions are consistent, facilitating comprehensive comparison and evaluation. Alternatively, N can be 512, and K can be 2 raised to the power of m, where m is an integer greater than or equal to 1 and less than or equal to 8. It should also be noted that the way to reduce the dimensionality of the feature vector can be to use deformable convolution. The receptive field of traditional convolution kernels is square, while in deformable convolution, the convolved receptive field is deformable, so the desired features can be extracted more accurately.
[0060] In some examples, the processing method of the attention model for the input vector may include: transforming the input vector to obtain a query matrix, a key matrix, and a value matrix; based on the query matrix and the key matrix, obtaining a weight matrix, wherein the weight matrix includes the weight of each feature information; based on the weight matrix and the value matrix, obtaining a total score. According to the different attention models adopted, the specific calculation process may also be different. For example, the weight matrix can be obtained by normalizing the similarity between the query matrix and the key matrix, and the similarity can be calculated by any of the methods such as cosine distance, dot product, perceptron, etc., and the normalization of the similarity can be achieved by sigmoid or softmax function.
[0061] In some cases, the attention model used can be a multi-head attention model. The multi-head attention model is similar to multiple convolution kernels. It can repeat the attention calculation multiple times and concatenate the results, thereby achieving multi-angle focus. The calculation process of the multi-head attention model can be described by the following formula:
[0062]
[0063] MultiHead(Q,K,V)=Concat(head1,...,head h )W o
[0064] In the above formula, Q, K, and V represent the query matrix, key matrix, and value matrix, respectively. are matrices used to linearly transform Q, K, and V, respectively. h is the number of heads, and Wo is the projection matrix used to map the output results to the target dimension. For example, in one embodiment, an 8-head attention model is used for processing, resulting in a total of 8 attention functions. Each attention function is responsible for only one eighth of the final output sequence, and each attention function is independent of each other, so it can realize 8 independent attention calculations, that is, achieve multi-angle focused attention.
[0065] In other cases, the attention model used may be a self-attention model. The self-attention model reduces reliance on external information and is more suitable for capturing the internal correlation of data or features. The calculation process of the self-attention model for the total score can be expressed based on the following formula:
[0066]
[0067] In the above formula, Q, K, and V represent the query matrix, key matrix, and value matrix respectively. k is the dimension of the input vector. In the self-attention model, Q, K, and V are all the results of linear transformations of the same input vector X. Therefore, the output is a sequence of vectors with the same dimension as X. This can directly capture the relationship between any two vectors in X and is easily parallelizable. Furthermore, when using the self-attention model, memory usage is positively correlated with the square of the sequence length. That is, if the sequence length doubles, the memory usage quadruples. To reduce memory usage, the self-attention model used can be a sparse self-attention model. The sparse self-attention model is based on the sparse self-attention mechanism. Traditional self-attention mechanisms require calculating the similarity between each output position and all input positions to obtain a dense similarity matrix, which is computationally complex. The core of the sparse self-attention mechanism is to decompose the dense similarity matrix into the product of two sparse similarity matrices, thereby reducing computational complexity and, consequently, computational cost and memory consumption.
[0068] In an optional embodiment, the training data of the attention model mentioned in this step may include multiple training samples and the total annotated score of each training sample; the attention model may be trained based on the following method: the training samples are processed by the attention model to obtain a predicted total score and a weight matrix; an initial loss function is constructed based on the total annotated score of the training sample and the predicted total score, and a target loss function is constructed based on the initial loss function and the weight matrix; the attention model is trained using the target loss function. The target loss function may be obtained by weighted summation of the initial loss function using a weight matrix. Of course, in other embodiments, the attention model may also be trained using other training methods, and this specification does not limit this.
[0069] Step 103: Based on the total score and the weight, beautify the product image in the quality dimension that needs beautification.
[0070] After determining the total score of the product image and the weight of each feature information, the product image can be beautified based on these two indicators. In some examples, whether the product image has a quality dimension that needs to be beautified can be judged based on the comparison result of the total score and the quality score threshold. For example, assuming that the quality score threshold is 5 points, if the total score of a product image is lower than 5 points, it can be judged that the product image has a quality dimension that needs to be beautified. If its total score is higher than or equal to 5 points, it can be judged that the product image does not have a quality dimension that needs to be beautified. The value of the quality score threshold can be set according to the needs of the specific scenario, and this embodiment does not limit this.
[0071] Because the weight of feature information represents its contribution to the overall score, when the overall score is low, feature information with a higher weight can be considered the key reason for the lower overall score. Based on this, when the overall score is below the quality score threshold, that is, when the product image has quality dimensions that need to beautified, the quality dimensions that need to beautified can be determined based on the weight. In one optional embodiment, the quality dimensions that need to beautify may include the quality dimensions corresponding to feature information with weights exceeding the weight threshold; in another optional embodiment, the quality dimensions that need to beautify may include the quality dimensions corresponding to several feature information with weights ranging from high to low. For example, the total image quality score of a product image in eight quality dimensions is lower than the quality score threshold. The eight feature information of the product image are recorded as feature 1, feature 2, feature 3, feature 4, feature 5, feature 6, feature 7, and feature 8, respectively, and their corresponding weights are 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, and 0.8, respectively. The quality dimensions that need to be beautified can be the quality dimensions corresponding to the feature information whose weights exceed the weight threshold of 0.55, that is, the quality dimensions corresponding to feature 6, feature 7, and feature 8, or they can be the quality dimensions corresponding to two feature information with weights from high to low, that is, the quality dimensions corresponding to feature 7 and feature 8. In this way, by determining the key reasons that lead to the poor overall quality of product images, the product images can be automatically beautified based on the key reasons, such as using a beautification algorithm corresponding to the quality dimension that needs to be beautified to process the product images. Compared with the unified operation of using preset filters to process each product image in the related technology, the method of this embodiment can be different for each image, and achieve targeted beautification processing, thereby effectively improving the image quality of the product images.
[0072] In actual applications, some quality dimensions of product images may not be convenient to beautify automatically. For example, when the LOGO text carried by a product image contains illegal text, if these illegal texts are deleted, the processed LOGO text is likely to contradict the original meaning. Based on this, in some examples, if the quality dimensions that need to beautify include specified quality dimensions, a prompt message is output to indicate that the product image violates the business rules. The specified quality dimensions here can be considered as quality dimensions that cannot be beautified. That is to say, when the quality dimensions that need to beautify include quality dimensions that cannot be beautified, the product image can be beautified without being processed, and a prompt message can be directly output to the user to prompt the user to change the product image. In addition, when the total score of a product image is extremely low, such as below the lower threshold, it indicates that the product image is very unqualified, wherein the lower threshold is less than the quality score threshold mentioned above. In this case, beautification processing of the product image will consume more resources, and the overall quality of the beautified product image is still likely to be poor. Therefore, in this case, the product image can also not be beautified, but a prompt message indicating that the product image is unavailable can be directly output to the user to reduce resource consumption.
[0073] In order to further verify the results of the beautification process, in some examples, after the beautification process, the step of obtaining multiple feature information of the beautified product image can be returned. That is, after the product image is beautified, the process of steps 101 to 103 can be executed again to determine whether the beautified product image has quality dimensions that need to be beautified. If the total score of the beautified product image is higher than or equal to the quality score threshold, it can be judged that the beautified product image is qualified, and the beautified product image can be published at this time; if the total score of the beautified product image is lower than the quality score threshold, it can be judged that the beautified product image is unqualified, and the step of returning can be continued until its total score is higher than or equal to the quality score threshold, or until the number of returns reaches the number threshold. In the case where the total score of the beautified product image is lower than the quality score threshold, a prompt message for prompting the user to change the product image can also be directly output. Through such a setting, it is possible to reasonably reduce resource consumption while automatically beautifying the product image.
[0074] The method of this embodiment processes multiple feature information of a product image through a trained attention model, determines the total score of the product image quality in each quality dimension related to the feature information and the weight of each feature information, and then beautifies the product image based on the total score and the weight. In this way, it is possible to automatically evaluate the quality and beautify the product images uploaded by merchants, thereby improving the image quality and thus increasing the click-through rate and purchase rate of the corresponding products. Moreover, due to the use of the attention mechanism, attention is paid to the relationship between the various quality dimensions in the product image, thereby achieving targeted automatic beautification of the product image, making the beautification results more reasonable.
[0075] The method of this embodiment can meet the needs of business parties to add or rectify new business rules. For example, when a business party proposes a new business rule, such as prohibiting the use of a certain manufacturer's logo and watermark in product images, it can import a batch of sample images related to the business rule to train the detection model and attention model. After training, the business rule will form the network's self-training capabilities and be integrated into the business scenario for the business party to use, so that the business rule can be implemented for all categories of goods, and business prompts such as warnings, removal from shelves, and bans can be taken for food images that violate the business rule.
[0076] In order to further illustrate the method of the embodiment of this specification, a specific embodiment is introduced below:
[0077] In this embodiment, the product image processing method of this specification is applied to the client of a merchant. The merchant is a merchant on a food delivery platform. Therefore, the product images in this embodiment are pictures of dishes. The client's processing process for dish images is as follows:
[0078] S201. Obtain dish image A uploaded by the merchant;
[0079] S202: Calling an image quality API to obtain multiple feature information representing the image quality of dish image A in multiple quality dimensions, a total score, and a weight of each feature information;
[0080] like Figure 2 As shown, Figure 2This is a schematic diagram of an image quality API shown in this embodiment, wherein the image quality API includes multiple detection models, each detection model is used to detect a feature information; in this embodiment, the specific classification of quality assessment content includes: image appearance score; whether the image content is blurred; whether the image lighting is dim or uneven; whether there are color spots and thorns; whether it is a re-photographed image; whether it is a puzzle; whether it is contained in a container or packaged; whether the image has been processed twice; whether the food in the image is edible directly; whether the image has a logo or text; whether the image has black edges, etc.; correspondingly, the multiple detection models in the image quality API include an appearance score model, a blur model, a light model, etc.; in actual application, the client can display each feature information to the user in the form of a score for each quality dimension, such as "lighting: poor, clarity: good, completeness: good, whether there are black and white edges: no, whether there are poster advertising products: no...";
[0081] The image quality API includes an attention model that processes multiple feature information to determine the overall score and weight. Specifically, a 512-dimensional feature vector is first extracted for each feature information. Then, a deformable convolution is performed to compress the feature dimension to 512 / N to obtain a reduced-dimensional vector. Multiple reduced-dimensional vectors are then concatenated to form the final input vector sent to the attention model. N is the number of feature information, so the dimension of the input vector is also 512. In this way, by compressing the features, the calculation of correlation is reduced.
[0082] The attention model used in this embodiment is a multi-head self-attention model. Analogous to multiple convolution kernels, this model repeats the Attention calculation multiple times and splices the results together, thereby achieving multi-angle focused attention. Moreover, based on the self-attention mechanism, the model's Q (query matrix), K (key matrix), and V (value matrix) are all the results of linear transformation of the same X (input vector). In this way, the output result is a vector sequence of the same length as X, which can directly capture the association between any two vectors in X and is easy to parallelize. Furthermore, considering that the computation time and memory usage of the self-attention mechanism are both at the square of n (n is the sequence length), this means that if the sequence length is doubled, the memory usage is quadrupled, and the computation time may also be quadrupled. Therefore, the attention model used in this embodiment is a sparse self-attention model, which can reduce memory consumption by reducing computational complexity.
[0083] The weight of each feature information is calculated based on Q and K; and the total score is calculated based on the weight of each feature information and V;
[0084] S203: Determine whether dish image A needs to be beautified based on the total score. If so, execute S204;
[0085] S204: Determine the quality dimensions that need to be enhanced based on the weights. Specifically, determine the quality dimensions corresponding to the top five feature information in descending order of weight as the quality dimensions that need to be enhanced. In actual applications, the client can display the total score and the quality dimensions that need to be enhanced to the user, such as "Overall quality score: 4; Key reasons: Lighting, watermark...";
[0086] S205: Determine whether the quality dimension that needs to be beautified includes the quality dimension that cannot be beautified. If so, execute S206; if not, execute S208.
[0087] S206. Beautify dish image A in the quality dimension that requires beautification processing using an image beautification API; the image beautification API includes multiple intelligent beautification algorithms, each of which is used to beautify an image in a quality dimension; specifically, the dish image A is processed using a corresponding intelligent beautification algorithm according to the quality dimension that requires beautification processing to obtain a beautified image B;
[0088] S207: Return to S202 for the beautified image B to determine the total score of the beautified image B and the weights of each feature information. In actual applications, the client can display the feature information, total score, and quality dimensions of the beautified image B to the user.
[0089] S208: Output prompt information for prompting that the current image violates business regulations.
[0090] The method of this embodiment can provide convenience for merchants. For merchants, they can see the comprehensive quality of pictures before and after beautification, and can also know the problems of dish pictures in time. The beautified dish pictures have improved picture quality, which can be converted into increased click-through rate and purchase rate. This method can also provide convenience for business parties, that is, merchant management and governance. For merchant management and governance, they can add corresponding adaptive tags to the system according to new needs and governance content, such as "poster advertising products are not allowed", and import samples into the training system according to the rules to generate corresponding rules. After the system automatically learns and processes according to the samples, it can implement rules for all categories of products, and take business prompts such as warnings, removal from shelves, and bans for dish pictures that violate the rules.
[0091] Corresponding to the embodiments of the aforementioned method, this specification also provides embodiments of a product image processing device and a terminal to which it is applied.
[0092] The embodiments of the product image processing device in this specification can be applied to computer devices, such as servers or terminal devices. The device embodiments can be implemented through software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of the file processing in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, if Figure 3 The figure is a hardware structure diagram of the computer device where the commodity image processing device of the embodiment of this specification is located, except Figure 3 In addition to the processor 310, memory 330, network interface 320, and non-volatile memory 340 shown, the server or electronic device where the device 331 is located in the embodiment may also include other hardware according to the actual function of the computer device, which will not be described in detail.
[0093] Accordingly, an embodiment of this specification further provides a computer storage medium, wherein the storage medium stores a program, and when the program is executed by a processor, the method in any of the above embodiments is implemented.
[0094] The embodiments of this specification may take the form of a computer program product implemented on one or more storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0095] like Figure 4 As shown, Figure 4 This is a block diagram of a product image processing device according to an exemplary embodiment of the present specification, the device comprising:
[0096] An acquisition module 41 is configured to acquire multiple feature information of the input product image; wherein one feature information is related to the image quality of the product image in a quality dimension;
[0097] Determination module 42, configured to process the plurality of feature information using the trained attention model to determine the total image quality score of the product image in each quality dimension related to each feature information and the weight of each feature information, wherein the weight of each feature information is positively correlated with the contribution of the feature information to the total score;
[0098] The beautification module 43 is configured to beautify the product image in the quality dimension that needs beautification based on the total score and the weight.
[0099] In some examples, each feature information is obtained by processing the product image through a detection model.
[0100] In some examples, the determining module includes:
[0101] An extraction module is used to obtain the feature vector corresponding to each feature information;
[0102] A splicing module, configured to splice the feature vectors corresponding to the plurality of feature information to obtain an input vector;
[0103] An input module is used to input the input vector into the attention model to obtain the weight and the total score.
[0104] In some examples, the splicing module includes:
[0105] The dimensionality reduction submodule is used to perform dimensionality reduction processing on each eigenvector to obtain the reduced dimensionality vector corresponding to each eigenvector;
[0106] The splicing submodule is used to splice the various dimensionality reduction vectors.
[0107] In some examples, the training data for the attention model includes multiple training samples and labels for each training sample. The attention model is trained based on the following method:
[0108] Processing the training samples through the attention model to obtain a predicted total score and a weight matrix;
[0109] constructing an initial loss function based on the total labeled score of the training sample and the total predicted score, and constructing a target loss function based on the initial loss function and the weight matrix;
[0110] The attention model is trained using the target loss function.
[0111] In some examples, the attention model includes a sparse self-attention model.
[0112] In some examples, the beautification module includes:
[0113] The determination submodule is configured to determine the quality dimension that needs to be beautified based on the weight if the total score is lower than the quality score threshold.
[0114] In some examples, the determination submodule is specifically configured to:
[0115] Determine the quality dimension corresponding to the feature information whose weight exceeds the weight threshold as the quality dimension that needs to be beautified; or
[0116] The quality dimensions corresponding to several pieces of feature information with weights from high to low are determined as the quality dimensions that need to be beautified.
[0117] In some cases, this also includes:
[0118] The output module is used to output a prompt message for prompting that the product image violates the business rules if the quality dimension that needs to be beautified includes the specified quality dimension.
[0119] In some cases, this also includes:
[0120] The return module is used to return the beautified product image to the above-mentioned acquisition module after beautification.
[0121] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0122] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0123] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0124] Other embodiments of the present invention will readily occur to those skilled in the art upon consideration of the present invention and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present invention being indicated by the following claims.
[0125] It should be understood that the present description is not limited to the exact structure that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present description is limited only by the appended claims.
[0126] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A method for processing product images, comprising: Acquire multiple feature information of the input product image; wherein one feature information is related to the image quality of the product image in a quality dimension; and each feature information is an evaluation result corresponding to a quality dimension; The plurality of feature information are processed using a trained attention model to determine an overall image quality score of the product image in each quality dimension associated with each feature information and a weight of each feature information, wherein the weight of each feature information is positively correlated with the contribution of the feature information to the overall score. The trained attention model is used to transform the plurality of feature information to obtain a query matrix, a key matrix, and a value matrix; a weight matrix is obtained based on the query matrix and the key matrix, the weight matrix including the weight of each feature information; and an overall score is obtained based on the weight matrix and the value matrix. If the total score is lower than the quality score threshold, the quality dimension corresponding to the feature information whose weight exceeds the weight threshold is determined as the quality dimension that needs to be beautified, or the quality dimensions corresponding to a preset number of feature information with weights from high to low are determined as the quality dimensions that need to be beautified, and the product image is beautified on the quality dimension that needs to be beautified.
2. The method according to claim 1, wherein each feature information is obtained by processing the product image through a detection model.
3. The method of claim 1, wherein the plurality of feature information is processed using a trained attention model to determine an overall image quality score of the product image in multiple quality dimensions and a weight for each feature information, comprising: Get the feature vector corresponding to each feature information; Concatenating the feature vectors corresponding to the plurality of feature information to obtain an input vector; The input vector is input into the attention model to obtain the weight and the total score.
4. The method according to claim 3, wherein the feature vectors corresponding to the plurality of feature information are concatenated to obtain an input vector, comprising: Perform dimensionality reduction processing on each eigenvector to obtain the reduced dimensionality vector corresponding to each eigenvector; Concatenate the reduced dimensionality vectors.
5. The method of claim 3, wherein the training data of the attention model includes a plurality of training samples and labels of each training sample; and the attention model is trained based on the following method: Processing the training samples through the attention model to obtain a predicted total score and a weight matrix; constructing an initial loss function based on the total labeled score of the training sample and the total predicted score, and constructing a target loss function based on the initial loss function and the weight matrix; The attention model is trained using the target loss function.
6. The method of claim 1, wherein the attention model comprises: Sparse self-attention model.
7. The method of claim 1 , further comprising: If the quality dimension that needs to be beautified includes a specified quality dimension, a prompt message is output to prompt that the product image violates the business rules.
8. The method of claim 1 , further comprising: After the beautification process, the process returns to the step of obtaining multiple feature information for the beautified product image.
9. A product image processing device, comprising: An acquisition module is configured to acquire multiple feature information of the input product image; wherein one feature information is related to the image quality of the product image in a quality dimension; and each feature information is an evaluation result corresponding to a quality dimension; a determination module, configured to process the plurality of feature information using a trained attention model to determine a total image quality score of the product image in each quality dimension associated with each feature information and a weight of each feature information, wherein the weight of each feature information is positively correlated with the contribution of the feature information to the total score; wherein the trained attention model is configured to transform the plurality of feature information to obtain a query matrix, a key matrix, and a value matrix; based on the query matrix and the key matrix, a weight matrix is obtained, wherein the weight matrix includes the weight of each feature information; and based on the weight matrix and the value matrix, a total score is obtained; A beautification module is used to determine the quality dimension corresponding to the feature information whose weight exceeds the weight threshold as the quality dimension that needs to be beautified if the total score is lower than the quality score threshold, or to determine the quality dimensions corresponding to a preset number of feature information with weights from high to low as the quality dimensions that need to be beautified, and beautify the product image on the quality dimension that needs to be beautified.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Image evaluation model generation method and device, image data processing method and device, computer equipment and storage medium
CN110378883A
Image detection method and device, information interaction method and device and electronic equipment
CN112036245A
Image and category-based attention target detection method, device and medium
CN112733944A
Image aesthetic quality evaluation method fused with multi-modal attention mechanism
CN113657380A