E-commerce product image generation method and system
By encoding and dividing the feature matrix of e-commerce images and combining the correlation and distance between the text encoding vector and the codebook index, the target index is determined to generate images, which solves the problems of detail loss and generation instability in the prior art, and achieves a more natural and coordinated e-commerce image generation.
Patent Information
- Application Number
- CN202510143943.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing e-commerce image generation methods may experience details loss or distortion when processing product images with rich details and complex structures, resulting in a gap between the generated image and the real product. The discrete potential space may introduce generation instability and artifact problems, affecting the natural sense of the image.
By encoding the e-commerce image, the image is divided into a background area and a foreground area, the number of neighbors is determined based on the area of the feature vector in the feature matrix, and the correlation and distance between the text encoding vector and the codebook index is determined to generate an image.
This method can effectively prevent excessive changes in the product appearance, ensure that the background generation is more inclined to user commands, ensure that the generated image is more coordinated and consistent globally, avoid visual fragmentation, and improve the naturalness and quality of the image.
Smart Images

Figure CN120014105A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to an e-commerce product image generation method and system. Background Art
[0002] The development of e-commerce has brought great convenience to people's shopping and consumption. As a carrier of product information, images play a vital role in attracting consumers and promoting sales. Exquisite product images can intuitively display the appearance, texture, function and other characteristics of the product, stimulating consumers' desire to buy. Traditional e-commerce images are mainly obtained through real shooting, cutout, rendering and other processes. On the one hand, this requires professional photography equipment and personnel, which is costly and time-consuming. On the other hand, cutout and splicing are to cut out the product from the background and then splice it with other materials, and the effect may not be natural enough.
[0003] Artificial intelligence is becoming more and more effective in image generation. Compared with traditional e-commerce image generation, e-commerce image generation methods based on artificial intelligence provide e-commerce platforms with an efficient, low-cost and flexible image production solution. However, current e-commerce image generation methods may lose details or distort them when processing product images with rich details and complex structures, resulting in a gap between the generated images and the real products; the discretized latent space may introduce certain generation instability and artifact problems, affecting the naturalness of the image. Summary of the invention
[0004] In view of the above-mentioned problems, a first aspect provides an e-commerce product image generation method, the method comprising: The e-commerce image is encoded to obtain a feature matrix, the e-commerce image is divided into a background area and a foreground area, and the number of neighbors of each feature vector is determined according to the corresponding area of the feature vector in the feature matrix on the e-commerce image; Encode the text corresponding to the e-commerce image to obtain an encoding vector, calculate the correlation between the encoding vector and each index in the codebook, and the distance from the feature vector to each index in the codebook; determine the index of the number of neighborhoods based on the correlation and the distance and the corresponding area of the feature vector on the e-commerce image; A target index is determined from the index of the number of nearest neighbors, a discrete feature matrix is obtained using the target index, and the discrete feature matrix is passed through a decoder to obtain a generated image.
[0005] Preferably, the method of determining the number of neighbors of each feature vector according to the area of the feature vector on the e-commerce image in the feature matrix is specifically: Map each feature vector in the feature matrix to a corresponding area on the e-commerce image; calculate the proportion of the background area in the area; The maximum number of nearest neighbors is obtained, and the product of the proportion and the maximum number of nearest neighbors is added to a minimum value and rounded up to obtain the number of nearest neighbors of the feature vector.
[0006] Preferably, the calculating the correlation between the coding vector and each index in the codebook is specifically: The features corresponding to the encoding vector and the index in the codebook are input into the trained neural network model to obtain the correlation between the encoding vector and each index.
[0007] Preferably, the index of the number of neighborhoods is determined based on the correlation, the distance, and the corresponding area of the feature vector on the e-commerce image, specifically: If the corresponding area of the feature vector on the e-commerce image only covers the foreground area, determining the index of the number of neighbors of the feature vector from the codebook according to the distance; If the corresponding area of the feature vector on the e-commerce image covers both the foreground area and the background area, determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance; If the corresponding area of the feature vector on the e-commerce image only covers the background area, the index of the number of neighborhoods is determined according to the feature vector whose index has been determined in the feature matrix.
[0008] Preferably, the determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance is specifically: For each index of the codebook, determine whether its correlation with the encoding vector is higher than the correlation threshold. If so, take the index as a candidate index; The candidate indexes are sorted in ascending order of the distance, and the candidate indexes with the number of nearest neighbors after sorting are used as the indexes of the number of nearest neighbors of the feature vector.
[0009] Preferably, the determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance is specifically: Calculate the ratio of the background area to the foreground area in the corresponding area of the feature vector on the e-commerce image; Determine the weight of the correlation and the weight of the distance according to the ratio, normalize the distance, and calculate the weighted value of the correlation and (1-normalized distance) according to the weight; The indexes in the codebook are sorted in descending order of the weighted values, and the index of the number of neighbors with the highest sorting order is used as the index of the feature vector.
[0010] Preferably, the determining of the index of the number of neighborhoods according to the feature vector whose index has been determined in the feature matrix is specifically: Randomly replace the feature vectors with the indexes in the feature matrix with the features corresponding to the indexes of the feature vectors; The feature vector of the undetermined index is predicted by using a mask method, and the index closest to the predicted feature is found in the codebook and added to the set of feature vectors of the undetermined index; The above process is repeated until the number of elements in the set of feature vectors whose indexes are not determined is equal to the number of neighborhoods; and the indexes in the set are used as the indexes of the number of neighborhoods.
[0011] Preferably, the determining of the target index from the index of the number of neighbors is specifically: The indexes of the number of neighbors of all feature vectors are clustered to obtain the distance from the index of the number of neighbors of the feature vector to each cluster, and the index with the shortest average distance to all clusters is used as the target index.
[0012] A second aspect provides an e-commerce product image generation system, the system comprising: A partitioning module is used to encode the e-commerce image to obtain a feature matrix, divide the e-commerce image into a background area and a foreground area, and determine the number of neighbors of each feature vector according to the corresponding area of the feature vector in the feature matrix on the e-commerce image; An index determination module is used to encode the text corresponding to the e-commerce image to obtain a coding vector, calculate the correlation between the coding vector and each index in the codebook, and the distance from the feature vector to each index in the codebook; determine the index of the number of neighborhoods based on the correlation and the distance and the corresponding area of the feature vector on the e-commerce image; The image generation module is used to determine the target index from the index of the number of nearest neighbors, obtain a discrete feature matrix using the target index, and pass the discrete feature matrix through a decoder to obtain a generated image.
[0013] Preferably, the method of determining the number of neighbors of each feature vector according to the area of the feature vector on the e-commerce image in the feature matrix is specifically: Map each feature vector in the feature matrix to a corresponding area on the e-commerce image; calculate the proportion of the background area in the area; The maximum number of nearest neighbors is obtained, and the product of the proportion and the maximum number of nearest neighbors is added to a minimum value and rounded up to obtain the number of nearest neighbors of the feature vector.
[0014] Preferably, the calculating the correlation between the coding vector and each index in the codebook is specifically: The features corresponding to the encoding vector and the index in the codebook are input into the trained neural network model to obtain the correlation between the encoding vector and each index.
[0015] Preferably, the index of the number of neighborhoods is determined based on the correlation, the distance, and the corresponding area of the feature vector on the e-commerce image, specifically: If the corresponding area of the feature vector on the e-commerce image only covers the foreground area, determining the index of the number of neighbors of the feature vector from the codebook according to the distance; If the corresponding area of the feature vector on the e-commerce image covers both the foreground area and the background area, determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance; If the corresponding area of the feature vector on the e-commerce image only covers the background area, the index of the number of neighborhoods is determined according to the feature vector whose index has been determined in the feature matrix.
[0016] Preferably, the determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance is specifically: For each index of the codebook, determine whether its correlation with the encoding vector is higher than the correlation threshold. If so, take the index as a candidate index; The candidate indexes are sorted in ascending order of the distance, and the candidate indexes with the number of nearest neighbors after sorting are used as the indexes of the number of nearest neighbors of the feature vector.
[0017] Preferably, the determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance is specifically: Calculate the ratio of the background area to the foreground area in the corresponding area of the feature vector on the e-commerce image; Determine the weight of the correlation and the weight of the distance according to the ratio, normalize the distance, and calculate the weighted value of the correlation and (1-normalized distance) according to the weight; The indexes in the codebook are sorted in descending order of the weighted values, and the index of the number of neighbors with the highest sorting order is used as the index of the feature vector.
[0018] Preferably, the determining of the index of the number of neighborhoods according to the feature vector whose index has been determined in the feature matrix is specifically: Randomly replace the feature vectors with the indexes in the feature matrix with the features corresponding to the indexes of the feature vectors; The feature vector of the undetermined index is predicted by using a mask method, and the index closest to the predicted feature is found in the codebook and added to the set of feature vectors of the undetermined index; The above process is repeated until the number of elements in the set of feature vectors whose indexes are not determined is equal to the number of neighborhoods; and the indexes in the set are used as the indexes of the number of neighborhoods.
[0019] Preferably, the determining of the target index from the index of the number of neighbors is specifically: The indexes of the number of neighbors of all feature vectors are clustered to obtain the distance from the index of the number of neighbors of the feature vector to each cluster, and the index with the shortest average distance to all clusters is used as the target index.
[0020] A third aspect provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method described in the first aspect when executed by a processor.
[0021] In the e-commerce image generation, the present invention obtains different numbers of indexes from the codebook for the foreground and background, and the more the foreground accounts for the feature vector in the area of the e-commerce image, the more the index is determined based on the distance, which can prevent the product appearance from changing too much, and the more the background accounts for, the more the index is determined based on the instruction, so that the generated background is more biased towards the user's command. When the final index is determined, the indexes of all feature vectors are clustered, and the index that best matches the overall style is selected, thereby ensuring that the image is more coordinated and consistent globally and avoiding a sense of visual fragmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a flow chart of Embodiment 1; Figure 2 Schematic diagram of the codebook structure. DETAILED DESCRIPTION
[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0024] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0025] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0026] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0027] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0028] Example 1 Figure 1 As shown, Figure 1 The steps of the product image generation method are shown, including: S1, encode the e-commerce image to obtain a feature matrix, divide the e-commerce image into a background area and a foreground area, and determine the number of neighbors of each feature vector according to the corresponding area of the feature vector in the feature matrix on the e-commerce image; The e-commerce image is encoded by an encoder, etc., and the feature information therein is extracted to form a feature matrix. The feature matrix is composed of multiple feature vectors. In one embodiment, the feature matrix is continuous, that is, the value can be a decimal; the encoder outputs a continuous potential representation. In order to generate image content more accurately, e-commerce images are usually divided into foreground areas and background areas. The foreground area refers to the product itself, while the background area is the part that is not related to the product, such as a solid color background or a shooting environment. There are many ways to divide the foreground area and the background area, including but not limited to deep learning segmentation models such as U-Net, DeepLab or traditional image segmentation methods such as GrabCut, segmentation based on color clustering, and manual division can also be used.
[0029] After obtaining the feature matrix and completing the image area division, the number of neighbors of each feature vector is determined. In one embodiment, since the foreground area is mainly the product itself, the product itself does not need to be changed too much in the generated e-commerce image, and the background area is generated according to the user's instructions, that is, the text corresponding to the e-commerce image. In order to make the background overall coordinated, the number of its neighbors is greater than the number of neighbors of the feature vector containing the foreground area. Optionally, the fewer foreground areas are included in the corresponding area of the feature vector on the e-commerce image, the greater the number of neighbors. For example, in the feature matrix, the corresponding area of the third feature vector on the e-commerce image does not contain the foreground area at all, then the number of neighbors of the third feature vector is 6, and the corresponding area of the tenth feature vector on the e-commerce image is half foreground area and half background area, then the number of neighbors of the tenth feature vector is 3. Among them, the corresponding area of the feature vector on the e-commerce image is determined according to the receptive field.
[0030] In yet another embodiment, the number of neighbors of each feature vector is determined according to the area of the feature vector on the e-commerce image in the feature matrix, specifically: Map each feature vector in the feature matrix to a corresponding area on the e-commerce image; calculate the proportion of the background area in the area; The maximum number of nearest neighbors is obtained, and the product of the proportion and the maximum number of nearest neighbors is added to a minimum value and rounded up as the number of nearest neighbors of the feature vector.
[0031] The receptive field can be used to calculate the window size of each eigenvector on the e-commerce image, and the window area on the e-commerce image can be determined by combining the position of the eigenvector in the feature matrix. This area is a corresponding area where the eigenvector is mapped to the e-commerce image. For each area where the eigenvector is located, the proportion of the background part in the area is calculated. For example, assuming that a eigenvector is mapped to an area in the image, half of which is the foreground (such as the product) and the other half is the background (such as the environment or background board), then the proportion of the background area is 50% or 0.5.
[0032] In order to determine the number of neighbors of each feature vector, a preset maximum number of neighbors is first obtained, for example, 10, and then the background proportion calculated above is multiplied by the maximum value, and a minimum value is added, preferably the minimum value is 0.001, to prevent the product from being 0 in some cases. If the background proportion of a certain area is 0.15, then the calculation process is 0.15×10=1.5, plus 0.001 to get 1.501, and finally through the rounding up operation, even if the decimal part is very small, it will be taken as 2, and the number of neighbors of the feature vector is determined to be 2.
[0033] S2, encoding the text corresponding to the e-commerce image to obtain an encoding vector, calculating the correlation between the encoding vector and each index in the codebook, and the distance from the feature vector to each index in the codebook; determining the index of the number of neighborhoods based on the correlation and the distance and the corresponding area of the feature vector on the e-commerce image; Encode the text information corresponding to the e-commerce image to obtain a code vector representing the text semantics. Compare the code vector with each index in the pre-built codebook, where the structure of the codebook is as follows: Figure 2As shown, 1, 2, ..., n are the indices of code in the codebook, and feature 1, feature 2, ..., feature n are the features corresponding to the indices. The correlation, distance, etc. are all calculated for the features corresponding to the indices. The correlation is the degree of match between the coding vector and the features represented by each index in the codebook, and the distance refers to the distance from the feature vector to the features represented by each index in the codebook, such as the Euclidean distance or the cosine distance. After obtaining the correlation between the coding vector and each index in the codebook and the distance between the feature vector and these indexes, the index of the number of neighborhoods is determined in combination with the specific area information corresponding to the feature vector on the e-commerce image.
[0034] In an optional embodiment, the calculating the correlation between the coding vector and each index in the codebook is specifically: The features corresponding to the encoding vector and the index in the codebook are input into the trained neural network model to obtain the correlation between the encoding vector and each index.
[0035] The encoding vector and these codebook features are input into a trained neural network model that has learned how to compare semantic relevance from the input data, thereby outputting a relevance score that reflects the degree of match between the text encoding and the corresponding feature in the codebook. In other words, each codebook index will obtain a numerical value. The higher the numerical value, the more similar the semantics or features between the encoding vector and the index feature are. This allows the system to select the most appropriate codebook index based on the text content, providing more accurate basic information for subsequent image generation or reconstruction. For example, if the text describes a blue sky, then after being processed by the neural network, the codebook index that highly matches the blue sky feature will obtain a higher relevance score and will be selected first in subsequent processing.
[0036] The foreground area mainly includes the product itself. The index of the number of neighbors is found in the codebook according to the distance, which can avoid over-generation of the product itself, while the background is mainly generated according to the text input by the user. In another embodiment, the index of the number of neighbors is determined based on the correlation and the distance and the corresponding area of the feature vector on the e-commerce image, specifically: If the corresponding area of the feature vector on the e-commerce image only covers the foreground area, determining the index of the number of neighbors of the feature vector from the codebook according to the distance; If a feature vector is mapped to the corresponding area on the e-commerce image, the feature vector is determined to be the index of the number of neighbors according to the distance between the feature vector and the feature corresponding to each index in the codebook. Preferably, the indexes are sorted in order of distance from small to large, and the index of the number of neighbors mentioned above is selected as the index corresponding to the feature vector. For example, for a feature vector 3, there are 12 indexes 1-12 in the codebook, and each index has a feature in the codebook. The Euclidean distance between the feature vector 3 and the feature of each index in the codebook is calculated. If the number of neighbors is 2, the first two indexes with the smallest Euclidean distance are selected as the index corresponding to the feature vector 3.
[0037] If the corresponding area of the feature vector on the e-commerce image covers both the foreground area and the background area, determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance; If a feature vector is mapped to a corresponding area on an e-commerce image that has both foreground areas and background areas, if only based on distance or correlation, it may cause product information to be lost or background generation to be unreasonable. Based on this, the index of the number of neighbors of the feature vector is determined from the codebook according to the correlation and distance. In an optional embodiment, the index of the number of neighbors of the feature vector is determined from the codebook according to the correlation and the distance, specifically: For each index of the codebook, determine whether its correlation with the encoding vector is higher than the correlation threshold. If so, take the index as a candidate index; The candidate indexes are sorted in ascending order of the distance, and the candidate indexes with the number of nearest neighbors after sorting are used as the indexes of the number of nearest neighbors of the feature vector.
[0038] First, each index in the codebook is screened to determine whether the correlation between each index and the text encoding vector exceeds a preset threshold. Only the indexes with correlation higher than the threshold are considered to have a sufficient degree of match with the text content and are thus included in the candidate set. In other words, only those codebook indexes with a strong semantic connection with the text can enter the next step of consideration. This process ensures that the index selected subsequently is semantically consistent with the text description. After obtaining the candidate indexes, they are sorted according to the distance between the feature vector and these candidate indexes, and the sorting order is from small to large, that is, the closer the distance, the higher the feature similarity. Subsequently, the system selects a certain number of candidate indexes that are in front of the sorting, that is, the predetermined number of neighbors, as the final index of the number of neighbors. This not only ensures that the candidate index is semantically highly correlated with the text, but also tries to select the index closest to the feature vector in the visual or feature space, so that the final neighbor index can comprehensively reflect the best match between text information and image features.
[0039] In yet another embodiment, determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance is specifically: Calculate the ratio of the background area to the foreground area in the corresponding area of the feature vector on the e-commerce image; calculate the ratio of the background area to the foreground area in the area, which reflects the relative proportion of background information and foreground information in the area where the current feature vector is located. If the ratio is large, it means that the background area accounts for a high proportion; conversely, if the ratio is small, it means that the foreground area accounts for a large proportion.
[0040] Determine the weight of the correlation and the weight of the distance according to the ratio, normalize the distance, and calculate the weighted value of the correlation and (1-normalized distance) according to the weight; The indexes in the codebook are sorted in descending order of the weighted values, and the index of the number of neighbors with the highest sorting order is used as the index of the feature vector.
[0041] According to the ratio of the foreground and background areas calculated above, two weights are determined: one is the weight for correlation (semantic matching between the coding vector and the codebook index), and the other is the weight for distance (numerical distance between the feature vector and the codebook index). The distance will be normalized first, and its value will be adjusted to a uniform range. Then, the degree of matching between each codebook index and the current feature vector is comprehensively evaluated by calculating the correlation weighted value = correlation weight × correlation + distance weight × (1-normalized distance). All codebook indexes are sorted from large to small according to the weighted value, and the top-ranked index is selected as the neighbor index of the feature vector. In this way, the accuracy of semantic matching can be guaranteed, and the similarity of visual features can be ensured, thereby providing a more accurate basis for image generation. Preferably, the sum of the correlation weight and the distance weight is 1, and the ratio of the two is equal to the ratio of the background area to the foreground area.
[0042] If the corresponding area of the feature vector on the e-commerce image only covers the background area, the index of the number of neighbors is determined according to the feature vector with the determined index in the feature matrix. When the image area corresponding to a certain feature vector belongs completely to the background area, directly using its correlation or distance with the codebook to determine the index may lead to unstable or inaccurate selection. Therefore, instead of selecting the index directly from the codebook, the index of the feature vector is determined by referring to the feature vector with the determined index in the feature matrix. Specifically, the feature vector with the determined index closest to the feature vector of the index to be determined is obtained, and the index corresponding to the feature vector with the closest distance is used as the starting index of the feature vector of the index to be determined, and then the index of the number of neighbors closest to the starting index feature is found from the codebook as the index of the feature vector of the index to be determined. For example, the feature vector of the index to be determined is 4, and the feature vector with the determined index closest to feature vector 4 is feature vector 6. Feature vector 6 has 3 indexes, and several indexes that are most similar to these 3 indexes are found from the codebook as the index of the index to be determined.
[0043] In yet another embodiment, the index of the feature vector with an undetermined index is determined by mask prediction. Specifically: Randomly replace the feature vectors with the indexes in the feature matrix with the features corresponding to the indexes of the feature vectors; The feature of the feature vector with undetermined index is predicted by using a mask method, and the index closest to the predicted feature is found in the codebook and added to the set of feature vectors with undetermined index; The above process is repeated until the number of elements in the set of feature vectors whose indexes are not determined is equal to the number of neighborhoods; and the indexes in the set are used as the indexes of the number of neighborhoods.
[0044] For feature vectors with determined indexes, randomly replace them with features of corresponding indexes, and then use mask prediction to predict the features of the masked part, that is, the features of feature vectors with undetermined indexes, and then find the index most similar to the predicted feature from the codebook, and add the most similar index to the set corresponding to the feature vector with undetermined index, and repeat until the number of elements in the set is the number of neighbors. For example, the feature matrix has 3 feature vectors, where the indexes corresponding to feature vectors 1 and 3 have been determined in the above steps, replace the features of feature vectors 1 and 3 with the features corresponding to the indexes of feature vectors 1 and 3 in the codebook, and then use mask prediction to predict the features of feature vector 2, find the index most similar to or closest to the predicted feature of feature vector 2 in the codebook, and add this closest index to the set corresponding to feature vector 2, and repeat several times to obtain the index of the number of neighbors of feature vector 2.
[0045] S3, determining a target index from the index of the number of nearest neighbors, obtaining a discrete feature matrix using the target index, and passing the discrete feature matrix through a decoder to obtain a generated image.
[0046] Each feature vector in the feature matrix has at least one corresponding index. An index is determined from at least one corresponding index as the target index. The corresponding feature vector in the feature matrix is replaced with the feature of the target index to obtain a discrete feature matrix. The discrete feature matrix is then passed through a decoder to obtain a generated image. The encoder and decoder preferably use the encoder and decoder in the VQ-VAE or VQ-GAN model.
[0047] In order to make the generated images consistent, in one embodiment, the target index is determined from the index of the number of neighbors, specifically: The indexes of the number of neighbors of all feature vectors are clustered to obtain the distance from the index of the number of neighbors of the feature vector to each cluster, and the index with the shortest average distance to all clusters is used as the target index.
[0048] Each feature vector has determined the index of the corresponding number of neighbors, where the index is the index of the code in the codebook. Using a clustering algorithm, all neighborhood indexes of all feature vectors are clustered. Clustering is clustering the features indexed in the codebook. For each neighborhood index of a feature vector, calculate its distance to each cluster. For example, feature vector 3 corresponds to 2 indexes. The distances from the first index to the three clusters are A11, A12, and A13, respectively. The distances from the second index to the three clusters are A21, A22, and A23, respectively. Since the total distance from the first index to the three clusters is the smallest, the first index is used as the target index.
[0049] Embodiment 2 provides a product image generation system, including: A partitioning module is used to encode the e-commerce image to obtain a feature matrix, divide the e-commerce image into a background area and a foreground area, and determine the number of neighbors of each feature vector according to the corresponding area of the feature vector in the feature matrix on the e-commerce image; An index determination module is used to encode the text corresponding to the e-commerce image to obtain a coding vector, calculate the correlation between the coding vector and each index in the codebook, and the distance from the feature vector to each index in the codebook; determine the index of the number of neighborhoods based on the correlation and the distance and the corresponding area of the feature vector on the e-commerce image; The image generation module is used to determine the target index from the index of the number of nearest neighbors, obtain a discrete feature matrix using the target index, and pass the discrete feature matrix through a decoder to obtain a generated image.
[0050] Optionally, the method of determining the number of neighbors of each feature vector according to the area of the feature vector on the e-commerce image in the feature matrix is specifically: Map each feature vector in the feature matrix to a corresponding area on the e-commerce image; calculate the proportion of the background area in the area; The maximum number of nearest neighbors is obtained, and the product of the proportion and the maximum number of nearest neighbors is added to a minimum value and rounded up to obtain the number of nearest neighbors of the feature vector.
[0051] Optionally, the calculating the correlation between the coding vector and each index in the codebook is specifically: The features corresponding to the encoding vector and the index in the codebook are input into the trained neural network model to obtain the correlation between the encoding vector and each index.
[0052] Optionally, the determining the index of the number of neighborhoods based on the correlation and the distance and a corresponding area of the feature vector on the e-commerce image is specifically: If the corresponding area of the feature vector on the e-commerce image only covers the foreground area, determining the index of the number of neighbors of the feature vector from the codebook according to the distance; If the corresponding area of the feature vector on the e-commerce image covers both the foreground area and the background area, determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance; If the corresponding area of the feature vector on the e-commerce image only covers the background area, the index of the number of neighborhoods is determined according to the feature vector whose index has been determined in the feature matrix.
[0053] Optionally, determining the index of the number of neighbors of the feature vector from a codebook according to the correlation and the distance is specifically: For each index of the codebook, determine whether its correlation with the encoding vector is higher than the correlation threshold. If so, take the index as a candidate index; The candidate indexes are sorted in ascending order of the distance, and the candidate indexes with the number of nearest neighbors after sorting are used as the indexes of the number of nearest neighbors of the feature vector.
[0054] Optionally, determining the index of the number of neighbors of the feature vector from a codebook according to the correlation and the distance is specifically: Calculate the ratio of the background area to the foreground area in the corresponding area of the feature vector on the e-commerce image; Determine the weight of the correlation and the weight of the distance according to the ratio, normalize the distance, and calculate the weighted value of the correlation and (1-normalized distance) according to the weight; The indexes in the codebook are sorted in descending order of the weighted values, and the index of the number of neighbors with the highest sorting order is used as the index of the feature vector.
[0055] Optionally, the determining the index of the number of neighborhoods according to the feature vector whose index has been determined in the feature matrix is specifically: Randomly replace the feature vectors with the indexes in the feature matrix with the features corresponding to the indexes of the feature vectors; The feature of the feature vector with undetermined index is predicted by using a mask method, and the index closest to the predicted feature is found in the codebook and added to the set of feature vectors with undetermined index; The above process is repeated until the number of elements in the set of feature vectors whose indexes are not determined is equal to the number of neighborhoods; and the indexes in the set are used as the indexes of the number of neighborhoods.
[0056] Optionally, determining the target index from the index of the number of neighbors is specifically: The indexes of the number of neighbors of all feature vectors are clustered to obtain the distance from the index of the number of neighbors of the feature vector to each cluster, and the index with the shortest average distance to all clusters is used as the target index.
[0057] Embodiment 3 provides a computer-readable storage medium, on which a computer program is stored, and it is characterized in that when the computer program is executed by a processor, it implements the method described in embodiment 1.
[0058] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0059] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0060] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A method for generating an e-commerce product image, characterized in that: The method comprises: The e-commerce image is encoded to obtain a feature matrix, the e-commerce image is divided into a background area and a foreground area, and the number of neighbors of each feature vector is determined according to the corresponding area of the feature vector in the feature matrix on the e-commerce image; Encode the text corresponding to the e-commerce image to obtain an encoding vector, calculate the correlation between the encoding vector and each index in the codebook, and the distance from the feature vector to each index in the codebook; determine the index of the number of neighborhoods based on the correlation and the distance and the corresponding area of the feature vector on the e-commerce image; A target index is determined from the index of the number of nearest neighbors, a discrete feature matrix is obtained using the target index, and the discrete feature matrix is passed through a decoder to obtain a generated image.
2. The method according to claim 1, characterized in that The method of determining the number of neighbors of each feature vector according to the area of the feature vector in the feature matrix on the e-commerce image is specifically as follows: Map each feature vector in the feature matrix to a corresponding area on the e-commerce image; calculate the proportion of the background area in the area; The maximum number of nearest neighbors is obtained, and the product of the proportion and the maximum number of nearest neighbors is added to a minimum value and rounded up to obtain the number of nearest neighbors of the feature vector.
3. The method according to claim 1, characterized in that The calculation of the correlation between the coding vector and each index in the codebook is specifically as follows: The features corresponding to the encoding vector and the index in the codebook are input into the trained neural network model to obtain the correlation between the encoding vector and each index.
4. The method according to claim 1, characterized in that The index of determining the number of neighborhoods based on the correlation, the distance, and the corresponding area of the feature vector on the e-commerce image is specifically: If the corresponding area of the feature vector on the e-commerce image only covers the foreground area, determining the index of the number of neighbors of the feature vector from the codebook according to the distance; If the corresponding area of the feature vector on the e-commerce image covers both the foreground area and the background area, determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance; If the corresponding area of the feature vector on the e-commerce image only covers the background area, the index of the number of neighborhoods is determined according to the feature vector whose index has been determined in the feature matrix.
5. The method according to claim 4, characterized in that The step of determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance is specifically: For each index of the codebook, determine whether its correlation with the encoding vector is higher than the correlation threshold. If so, take the index as a candidate index; The candidate indexes are sorted in ascending order of the distance, and the candidate indexes with the number of nearest neighbors after sorting are used as the indexes of the number of nearest neighbors of the feature vector.
6. The method according to claim 4, characterized in that The step of determining the index of the number of neighbors of the feature vector from the codebook according to the correlation and the distance is specifically: Calculate the ratio of the background area to the foreground area in the corresponding area of the feature vector on the e-commerce image; Determine the weight of the correlation and the weight of the distance according to the ratio, normalize the distance, and calculate the weighted value of the correlation and (1-normalized distance) according to the weight; The indexes in the codebook are sorted in descending order of the weighted values, and the index of the number of neighbors with the highest sorting order is used as the index of the feature vector.
7. The method according to claim 4, characterized in that The step of determining the index of the number of neighborhoods according to the feature vector whose index has been determined in the feature matrix is specifically: Randomly replace the feature vectors with the indexes in the feature matrix with the features corresponding to the indexes of the feature vectors; The feature of the feature vector with undetermined index is predicted by using a mask method, and the index closest to the predicted feature is found in the codebook and added to the set of feature vectors with undetermined index; The above process is repeated until the number of elements in the set of feature vectors whose indexes are not determined is equal to the number of neighborhoods; and the indexes in the set are used as the indexes of the number of neighborhoods.
8. The method according to claim 1, characterized in that The determining of the target index from the index of the number of neighbors is specifically: The indexes of the number of neighbors of all feature vectors are clustered to obtain the distance from the index of the number of neighbors of the feature vector to each cluster, and the index with the shortest average distance to all clusters is used as the target index.
9. An e-commerce product image generation system, characterized in that: The system comprises: A partitioning module is used to encode the e-commerce image to obtain a feature matrix, divide the e-commerce image into a background area and a foreground area, and determine the number of neighbors of each feature vector according to the corresponding area of the feature vector in the feature matrix on the e-commerce image; An index determination module is used to encode the text corresponding to the e-commerce image to obtain a coding vector, calculate the correlation between the coding vector and each index in the codebook, and the distance from the feature vector to each index in the codebook; determine the index of the number of neighborhoods based on the correlation and the distance and the corresponding area of the feature vector on the e-commerce image; The image generation module is used to determine the target index from the index of the number of nearest neighbors, obtain a discrete feature matrix using the target index, and pass the discrete feature matrix through a decoder to obtain a generated image.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 8.