Commodity selection method and related products thereof
Through the multi-modal graphic and text model, the classification label text and product images entered by the user are converted into embedded information, and the target products are selected efficiently, which solves the problem of low efficiency in selecting massive products in e-commerce scenarios, and realizes efficient and accurate product selection and management.
Patent Information
- Application Number
- CN202510179762.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-24
AI Technical Summary
In the e-commerce scenario, it is difficult for the existing technology to efficiently and accurately select target products that meet the newly added classification labels from a large number of products, resulting in high labor costs and low product management efficiency.
By obtaining the classification tag text input by the user, converting each product image in the product library into image embedding information based on the multi-modal graphic and text model, and efficiently matching the text embedding information to determine the target product.
It has achieved efficient and accurate selection of target products that meet the newly added classification labels from a large number of products, reducing labor costs and improving product management efficiency.
Smart Images

Figure CN120197075A_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the field of artificial intelligence technology. More specifically, this application relates to a method for selecting goods and related products thereof. Background Art
[0002] In the e-commerce scenario, a wide variety of goods belong to different product categories under different themes or festivals. Currently, for most goods, when they are newly listed, the merchant side provides the classification labels and classification levels of the new goods. However, the classification labels of these goods provided by the merchant side cannot meet the usage requirements of all business scenarios. For example, under the Christmas festival, the business operator only needs goods with Christmas characteristics, and the merchant side often does not divide and define the classification labels of such goods in a very detailed manner. If additional classification labels need to be added to the goods, a large amount of manual work is required to select suitable goods and add labels, which is time-consuming and laborious. And by automatically selecting goods through corresponding rules to add corresponding labels, the diversity of goods will be limited, and at the same time, the background maintenance rules are cumbersome and redundant.
[0003] In view of this, there is an urgent need to provide a method for selecting goods, so as to efficiently and accurately select target goods that meet the newly added classification labels from a large number of goods, meet the label usage requirements of various business scenarios, reduce labor costs, and improve the efficiency of goods management. Summary of the Invention
[0004] In order to solve at least one or more of the above-mentioned technical problems, this application proposes a method for selecting goods and related products thereof in multiple aspects. This method for selecting goods can efficiently and accurately select target goods that meet the newly added classification labels from a large number of goods, meet the label usage requirements of various business scenarios, reduce labor costs, and improve the efficiency of goods management.
[0005] In a first aspect, this application provides a method for selecting goods, including: obtaining a classification label text input by a user; determining text embedding information based on the classification label text; inputting each product image in a product library into a multi-modal text-image model to obtain image embedding information corresponding to each product image output by the multi-modal text-image model; wherein, the multi-modal text-image model is trained based on a product picture sample set and product information texts corresponding to each product picture sample in the product picture sample set; and determining target goods based on the text embedding information and the image embedding information corresponding to each product image.
[0006] In some embodiments, training a multi-modal graphic-text model based on a set of product image samples and product information texts corresponding to each product image sample in the set of product image samples includes: obtaining the set of product image samples and the product information texts corresponding to each product image sample in the set of product image samples; preprocessing each product image sample and each product information text respectively to obtain a product image training set and a product information training set; loading a pre-trained model and adding a dimensionality reduction layer to the pre-trained model to obtain an initial graphic-text model; and inputting the product image training set and the product information training set into the initial graphic-text model for training to obtain the multi-modal graphic-text model.
[0007] In some embodiments, inputting the product image training set and the product information training set into the initial graphic-text model for training includes: pairing each training product information in the product information training set with a training product image in the product image training set to obtain a plurality of positive example pairing combinations and a plurality of negative example pairing combinations; determining a contrast loss value based on the plurality of positive example pairing combinations and a contrast learning loss function of the initial graphic-text model; determining a matching loss value based on the plurality of positive example pairing combinations, the plurality of negative example pairing combinations, and a matching loss function of the initial graphic-text model; determining a semantic description loss value based on the product information training set and a semantic loss function of the initial graphic-text model; determining a target loss value based on the contrast loss value, the matching loss value, and the semantic description loss value; and determining whether the initial graphic-text model is trained completed based on the target loss value.
[0008] In some embodiments, determining a contrast loss value based on the plurality of positive example pairing combinations and a contrast learning loss function of the initial graphic-text model includes: respectively inputting the training product images in each positive example pairing combination into a visual encoder of the initial graphic-text model to obtain product image features output by the visual encoder; respectively inputting the training product information in each positive example pairing combination into a text encoder of the initial graphic-text model to obtain product information features output by the text encoder; inputting the product image features and the product information features into the dimensionality reduction layer to obtain reduced-dimensional image features and reduced-dimensional information features; determining a similarity matrix based on the reduced-dimensional image features and the reduced-dimensional information features; determining a first-direction contrast loss value and a second-direction contrast loss value based on the similarity matrix; and determining the contrast loss value based on the first-direction contrast loss value and the second-direction contrast loss value.
[0009] In some embodiments, determining a matching loss value based on the plurality of positive example pairing combinations, the plurality of negative example pairing combinations, and a matching loss function of the initial graphic-text model includes: determining intermediate interaction features corresponding to each pairing combination based on the plurality of positive example pairing combinations and the plurality of negative example pairing combinations; determining a matching probability corresponding to each pairing combination according to the intermediate interaction features corresponding to each pairing combination; and determining the matching loss value based on a preset classification label parameter and the matching probability corresponding to each pairing combination.
[0010] In some embodiments, determining a target product based on text embedding information and image embedding information corresponding to each product image includes: determining the cosine similarity between the text embedding information and the image embedding information corresponding to each product image; determining at least one candidate matching product according to the cosine similarity; sorting the at least one candidate matching product according to a preset reference factor, and determining the target product according to the sorting result.
[0011] In some embodiments, before obtaining the classification label text input by the user, the method further includes: obtaining a list of newly added products; wherein, the list of newly added products includes the newly added description text corresponding to each newly added product and the newly added image information corresponding to each newly added product; respectively inputting the newly added description text corresponding to each newly added product and the newly added image information corresponding to each newly added product into a multi-modal text-image model, to obtain the newly added text embedding information corresponding to each newly added product and the newly added image embedding information corresponding to each newly added product output by the multi-modal text-image model; pushing the newly added text embedding information corresponding to each newly added product and the newly added image embedding information corresponding to each newly added product to a product library indexing engine to update the product library index.
[0012] In some embodiments, before obtaining the classification label text input by the user, the method further includes: listening to the data stream of product status changes; determining the off-shelf products according to the data stream of product status changes; clearing the off-shelf text embedding information and off-shelf image embedding information corresponding to the off-shelf products in the product library index.
[0013] In a second aspect, the present application provides a device for product selection, including: a memory; and at least one processor configured to: obtain the classification label text input by the user; determine text embedding information based on the classification label text; input each product image in the product library into a multi-modal text-image model to obtain the image embedding information corresponding to each product image output by the multi-modal text-image model; wherein, the multi-modal text-image model is trained based on a product image sample set and product information texts corresponding to each product image sample in the product image sample set; determine a target product based on the text embedding information and the image embedding information corresponding to each product image.
[0014] In a third aspect, the present application provides a non-transitory machine-readable medium storing program code for product selection. When the program code is executed by at least one processor, the code guides the execution operations of the at least one processor. The program code includes: obtaining classification label text input by a user; determining text embedding information based on the classification label text; inputting each product image in a product library into a multi-modal text-image model to obtain image embedding information corresponding to each product image output by the multi-modal text-image model; wherein, the multi-modal text-image model is trained based on a product picture sample set and product information text corresponding to each product picture sample in the product picture sample set; determining a target product based on the text embedding information and the image embedding information corresponding to each product image.
[0015] The technical solution provided by the present application may include the following beneficial effects:
[0016] The product selection method and related products provided by the present application obtain the classification label text input by the user, and then determine the text embedding information based on the classification label text, so as to be able to extract the core semantic information of the new classification label text. On the other hand, by inputting each product image in the product library into the multi-modal text-image model, the image embedding information corresponding to each product image output by the multi-modal text-image model can be obtained. Among them, the multi-modal text-image model is trained based on a product picture sample set and product information text corresponding to each product picture sample in the product picture sample set, so that the multi-modal text-image model has the ability to accurately extract the product information features of each product in the product library.
[0017] Furthermore, the present application can determine the target product based on the text embedding information and the image embedding information corresponding to each product image. The efficient matching of the text embedding information and the image embedding information corresponding to each product image effectively improves the product selection efficiency of the target product that conforms to the classification label text, and can accurately and quickly complete product selection without manual intervention.
[0018] Generally speaking, the present application can efficiently and accurately select target products that conform to the new classification label from a large number of products, meet the label usage requirements of various business scenarios, reduce labor costs, and improve product management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become readily understandable. In the drawings, several embodiments of the present application are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0020] Figure 1 An exemplary flowchart of a product selection method according to some embodiments of the present application is shown;
[0021] Figure 2 Shows an exemplary flowchart of a product selection method according to other embodiments of the present application;
[0022] Figure 3 Shows an exemplary flowchart of a product selection method according to still other embodiments of the present application;
[0023] Figure 4 Shows a block diagram of the hardware configuration of a product selection device 400 that can implement the product selection method according to the embodiments of the present application. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. For the sake of simplicity and clarity of description, where appropriate, the same reference numerals may be repeated in the drawings to indicate corresponding or similar elements. In addition, the present application elaborates on many specific details in order to provide a thorough understanding of the embodiments described herein. However, those of ordinary skill in the art will understand that the embodiments described herein can be practiced without these specific details. In other cases, well-known methods, procedures, and components are not described in detail so as not to obscure the embodiments described herein. Moreover, this description should not be regarded as limiting the scope of the embodiments described herein. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0025] It should be understood that the possible terms "first" or "second" etc. in the claims, the description, and the drawings disclosed in the present application are used to distinguish different objects, rather than to describe a specific order. The terms "including" and "comprising" used in the description and claims of the present application indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0026] It should also be understood that the terms used in the description of the present application herein are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the description and claims of the present application, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should also be further understood that the term " / and" used in the description and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0027] As used in this specification and the claims, the term "if" may be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be construed, depending on the context, to mean "once determined" or "in response to determining" or "once [described condition or event] is detected" or "in response to detecting [described condition or event]".
[0028] In the e-commerce scenario, a wide variety of products should belong to different product categories under different themes or festivals. Currently, when most products are newly listed, the merchant side provides the classification labels and classification levels of the new products. However, the classification labels of these products provided by the merchant side cannot meet the usage requirements of all business scenarios. If additional classification labels need to be added to the products, a large amount of purely manual methods are required to select the appropriate products and add labels, which is time-consuming and laborious. By automatically selecting products according to corresponding rules to add corresponding labels, the diversity of products will be limited, and at the same time, the background maintenance rules are cumbersome and redundant.
[0029] In view of this, there is an urgent need to provide a product selection method, so as to efficiently and accurately select target products that meet the newly added classification labels from a large number of products, meet the label usage requirements of various business scenarios, reduce labor costs, and improve product management efficiency.
[0030] The following will describe in detail the specific implementation manners of the present application with reference to the accompanying drawings.
[0031] In step S101, the classification label text input by the user is obtained. In the embodiment of the present application, the classification label text may refer to the description text of the newly added classification label, such as "sleeveless dress suitable for summer, with prints on the dress", or may refer to the label text of the newly added classification label, such as "sleeveless printed dress". It can be understood that the form of the classification label text can be diverse. In actual applications, the form of the classification label text needs to be determined according to the actual application situation, and the present application does not make any restrictions in this regard.
[0032] In step S102, text embedding information is determined based on the classification label text. In the embodiments of the present application, the classification label text can be input into a natural language processing model for parsing to extract core semantic information (such as product type, material, pattern, etc.) to obtain the text embedding information. In Natural Language Processing (NLP), Text Embedding is the process of converting text data into a vector representation of a fixed dimension. These vectors can capture the semantic information of the text, so that texts with the same semantics but different expressions are mapped to similar positions in the vector space, while texts with different semantics maintain the corresponding distances.
[0033] In step S103, each product image in the product library is input into the multimodal text-image model to obtain the image embedding information corresponding to each product image output by the multimodal text-image model. In the embodiments of the present application, the foregoing multimodal text-image model can be trained based on a product picture sample set and the product information text corresponding to each product picture sample in the product picture sample set, so that the multimodal text-image model has the ability to accurately extract the product information features of each product in the product library. Among them, the product information text corresponding to each product picture sample can include but is not limited to: product title, product category information, material, filling, down content, upper wear type, lower wear type, hem shape, sleeve length, sleeve type, pattern, collar type, details, thinness and transparency, function type, whether it is fleece-lined, style type, length, boot shaft height type, color, product main picture information, etc. In addition, Image Embedding is a method of converting image data into a continuous, low-dimensional vector representation, and these vector representations are usually used in subsequent machine learning tasks such as classification, clustering, retrieval, etc. The purpose of Image Embedding is to convert high-dimensional image data into lower-dimensional data that is easier to process while retaining as much of the original image information as possible.
[0034] In step S104, the target product is determined based on the text embedding information and the image embedding information corresponding to each product image. In the embodiments of the present application, the text embedding information can be compared with the image embedding information corresponding to each product image, and then the target product can be determined according to the similarity shown by the comparison result, so that the product selection can be completed accurately and quickly without manual intervention.
[0035] The embodiment of the present application obtains the classification label text input by the user, and then determines the text embedding information based on the classification label text, so that the core semantic information of the newly added classification label text can be extracted. On the other hand, by inputting each product image in the product library into the multimodal graphic model, the image embedding information corresponding to each product image output by the multimodal graphic model can be obtained. Among them, the multimodal graphic model is trained based on the product image sample set and the product information text corresponding to each product image sample in the product image sample set, so that the multimodal graphic model has the ability to accurately extract the product information features of each product in the product library. Further, the present application can determine the target product based on the text embedding information and the image embedding information corresponding to each product image, and effectively improve the selection efficiency of the target product that meets the classification label text through the efficient matching of the text embedding information and the image embedding information corresponding to each product image, and the product selection can be completed accurately and quickly without manual intervention. In general, the present application can efficiently and accurately select the target product that meets the newly added classification label from a large number of products, meet the label usage requirements of various business scenarios, reduce labor costs, and improve product management efficiency.
[0036] In some embodiments, the training steps of the multimodal graphic model can be further designed. Figure 2 Let’s explain the training steps of the multimodal graph-text model in detail. Figure 2 See the exemplary flow chart of the commodity selection method of other embodiments of the present application. Figure 2 The commodity selection method shown in the embodiment of the present application may include:
[0037] In step S201, a product image sample set and product information text corresponding to each product image sample in the product image sample set are obtained. In the embodiment of the present application, images of various product scenes such as product style images and model images can be collected to form a product image sample set, wherein the product title, product category, attribute label (material, version, color, style, detailed description, etc.) and other information contained in each product image sample in the product image sample set can be used as the product information text corresponding to each product image sample in the product image sample set.
[0038] In step S202, each product image sample and each product information text are preprocessed respectively to obtain a product image training set and a product information training set. In the embodiments of the present application, the preprocessing may include operations such as screening and deduplication, text cleaning and encoding, and image preprocessing. Among them, the screening and deduplication operation may specifically be to remove duplicate data and abnormal data (such as invalid pictures, empty texts, etc.), and if there is noisy text or annotation errors, cleaning is also required. In addition, the text cleaning and encoding operation may specifically be to merge the product title, attribute information, etc. into one or more paragraphs of text, remove unnecessary special characters, and then a tokenizer / sub-word tokenizer (such as BPE, Byte Pair Encoding) compatible with the CLIP (Contrastive Language-Image Pre-training) model can be used to ensure that the input is consistent with the tokenization method during the model's pre-training. Furthermore, the image preprocessing operation may specifically be to first crop or scale each product image sample to the default input resolution of the model (such as 224×224), and then random horizontal flipping, random cropping, color jittering, etc. can be used to achieve image data augmentation, and the amplitude of augmentation can be adjusted as needed according to the actual application situation. After the preprocessing is completed, a product image training set and a product information training set can be obtained.
[0039] In step S203, a pre-trained model is loaded and a dimensionality reduction layer is added to the pre-trained model to obtain an initial text-image model. In the embodiments of the present application, the aforementioned pre-trained model may be a CLIP (Contrastive Language-Image Pre-training) pre-trained model, which is a multi-modal pre-trained model designed to be trained with a large number of "text-image" pairs to understand and match image content with corresponding natural language descriptions. The CLIP model embeds text and images into a common semantic space, making the representations of relevant text descriptions and image content close to each other in this space, while those that are not relevant are far apart. This pre-trained model includes a Visual Encoder and a Text Encoder, and usually open-source models such as ViT-B / 32 and ViT-B / 16 can be used.
[0040] On the other hand, in the embodiments of the present application, to meet the requirement of the final output of a 128-dimensional embedding, a linear mapping layer (i.e., a dimensionality reduction layer) can be newly created based on the original output of the pre-trained model (such as 512 dimensions), and it can be represented by the following formula one:
[0041] z = Linear 512→128 (x) (Formula One)
[0042] Among them, x is the original output of the pre-trained model, and z is the output of the newly created linear mapping layer. In some special application scenarios, if more flexible hierarchical dimensionality reduction is required, it can be reduced to 256 dimensions first and then to 128 dimensions as shown in the Matryoshaka Embedding Model, or it can be directly reduced to 128 dimensions at once.
[0043] In step S204, the product image training set and the product information training set are input into the initial text-image model for training to obtain a multi-modal text-image model. In the embodiments of the present application, each training product information in the product information training set can be first paired with the training product images in the product image training set to obtain a plurality of positive example pairing combinations and a plurality of negative example pairing combinations. Among them, a positive example pairing combination refers to the true pairing of a certain training product image with its corresponding training product information, and a negative example pairing combination refers to the pairing of a certain training product image with any training product information other than its corresponding training product information.
[0044] Then, the contrast loss value can be determined based on the contrast learning loss function of the plurality of positive example pairing combinations and the initial text-image model. Specifically, the training product images in each positive example pairing combination can be respectively input into the visual encoder of the initial text-image model to obtain the product image features output by the visual encoder, and the training product information in each positive example pairing combination can be respectively input into the text encoder of the initial text-image model to obtain the product information features output by the text encoder. Then, the product image features and the product information features are input into the dimensionality reduction layer to obtain the dimensionality-reduced image features and the dimensionality-reduced information features, and then the similarity matrix is determined based on the dimensionality-reduced image features and the dimensionality-reduced information features. Among them, the similarity matrix can be calculated by the following formula two:
[0045]
[0046] where s ij is the similarity matrix, is the dimensionality-reduced image feature obtained by dimensionality reduction of the i-th product image feature, is the dimensionality-reduced information feature obtained by dimensionality reduction of the j-th product information feature, and the denominator is the product of the vector norms (i.e., cosine similarity).
[0047] Furthermore, the first-direction contrast loss value and the second-direction contrast loss value can be determined based on the similarity matrix. Among them, the first-direction contrast loss value can be the contrast loss value in the image-to-text direction and can be calculated by the following formula three:
[0048]
[0049] where τ is the temperature coefficient, and this term encourages the product image feature Vi is higher than that with any unmatched product information feature t i in terms of similarity. j
[0050] In addition, the second-direction contrast loss value can be the contrast loss value in the text-to-image direction and can be calculated by the following formula four:
[0051]
[0052] Furthermore, the contrast loss value can be determined based on the first-direction contrast loss value and the second-direction contrast loss value. Among them, the contrast loss value can be calculated by the following formula five:
[0053]
[0054] where N is the total number of positive example pairing combinations.
[0055] Next, the matching loss value can be determined based on multiple positive example pairing combinations, multiple negative example pairing combinations, and the matching loss function of the initial text-image model. Specifically, the intermediate interaction feature corresponding to each pairing combination can be determined based on multiple positive example pairing combinations and multiple negative example pairing combinations, where the aforementioned intermediate interaction feature can be extracted by the initial text-image model. Furthermore, the matching probability corresponding to each pairing combination can be determined according to the intermediate interaction feature corresponding to each pairing combination. For example, the aforementioned matching probability can be output by a binary classification head, where the aforementioned binary classification head can adopt the network structure of a multi-layer perceptron (MLP) and adopt the Sigmoid function as the activation function, so as to map the output value to between (0, 1) through the Sigmoid function, thereby obtaining the matching probability.
[0056] Furthermore, the matching loss value can be determined based on the preset classification label parameter and the matching probability corresponding to each pairing combination. Specifically, the single-sample loss value of each pairing combination can be calculated by the following formula six:
[0057]
[0058] where is the single-sample loss value, y is the preset classification label parameter and y ∈ {0, 1}, is the matching probability. Furthermore, the average or sum of the obtained multiple single-sample loss values can be taken as the final matching loss value.
[0059] Moreover, the semantic description loss value can be determined based on the product information training set and the semantic loss function of the initial text-image model.
[0060] Among them, the semantic loss function can be expressed by the following formula seven:
[0061]
[0062] Among them, Ω = {i|M i = 1} is the set of all masked positions; M i is a Bernoulli variable, which is the value obtained by independently sampling each position i ∈ {1, …, L} of the text sequence of the training commodity information with a fixed probability p (usually p = 0.15) (L is the length of the text sequence). Among them, if M i = 1, it means that the i-th basic unit (token) in the text sequence is selected for masking (that is, the i-th basic unit is replaced with a mask symbol); if M i = 0, it means that the i-th basic unit in the text sequence is not masked. l i is the real word at the i-th position in the text sequence, represents the probability of predicting to generate the real word at position i.
[0063] Furthermore, based on the contrast loss value, the matching loss value, and the semantic description loss value, the target loss value is determined. In the embodiments of the present application, the contrast loss value, the matching loss value, and the semantic description loss value can be weighted and combined according to weights to obtain the target loss value. For example, it can be calculated by the following formula eight:
[0064]
[0065] Among them, α, β, and γ are all hyperparameters, and their relative importance can be adjusted according to actual specific application requirements. For example: if the contrast retrieval is the main goal, α can be made larger; if the text quality is good and the need for word-level understanding is high, β can be increased; if precise discrimination of the overall image-text matching is required (not just similarity ranking), γ can be increased accordingly.
[0066] Finally, based on the target loss value, it can be determined whether the initial image-text model is trained. In the embodiments of the present application, the gradient can be calculated for , and then all trainable parameters (including the visual encoder, the text encoder, and the dimensionality reduction layer, etc.) in the initial image-text model can be adjusted and updated according to the gradient direction. The optimizer for updating the parameters can choose AdamW or Adam, and the learning rate can be set to e -5 ~e -6 to prevent large-scale damage to the pre-trained weights. When the target loss value converges below the target value, it can be determined that the initial image-text model is trained, and a multi-modal image-text model is obtained.
[0067] In some embodiments, to ensure the timeliness and integrity of the product library, it is necessary to dynamically manage the product library. And the cosine similarity between the text embedding information and the image embedding information corresponding to each product image can be compared to determine the target product. The following will be combined with Figure 3 to describe the process of determining the target product in detail. Figure 3 FIG. shows an exemplary flowchart of the product selection method according to still some embodiments of the present application. Please refer to Figure 3 , the product selection method shown in the embodiments of the present application may include:
[0068] In step S301, the product library is dynamically managed. In the embodiments of the present application, dynamically managing the product library may include, but is not limited to, updating newly added products to the product library and dynamically managing each product in the product library according to the product status of each product in the product library.
[0069] Among them, updating the newly added products to the product library may specifically include: First, a list of newly added products can be obtained. Among them, the list of newly added products contains the newly added description text corresponding to each newly added product and the newly added picture information corresponding to each newly added product. Exemplarily, the list of newly added products can be obtained in the data warehouse tool HIVE of the product management platform. It can be understood that the acquisition method of the list of newly added products is diverse, and in actual applications, the acquisition method of the list of newly added products needs to be determined according to the actual application situation. The present application does not make any restrictions in this regard.
[0070] Further, the newly added description text corresponding to each newly added product and the newly added picture information corresponding to each newly added product are respectively input into the multi-modal text and image model to obtain the newly added text embedding information corresponding to each newly added product and the newly added image embedding information corresponding to each newly added product output by the multi-modal text and image model. Then, the newly added text embedding information corresponding to each newly added product and the newly added image embedding information corresponding to each newly added product are pushed to the product library index engine to update the product library index. Exemplarily, for example, the distributed stream processing platform Kafka can be used to push the newly added text embedding information corresponding to each newly added product and the newly added image embedding information corresponding to each newly added product to the Milvus distributed index engine. Then, the data is segmented and index managed according to the country and site to realize the update of the product library index corresponding to the country and site, ensuring that the latest product information can be accessed.
[0071] In addition, the dynamic management of each product in the product library according to the product status of each product in the product library may specifically include: First, it is possible to monitor the data stream of product status changes. For example, it can be monitored through the distributed stream processing platform Kafka. Then, determine the products to be taken off the shelf according to the data stream of product status changes. Furthermore, clear the off-shelf text embedding information and off-shelf image embedding information corresponding to the products to be taken off the shelf in the product library index, so as to ensure the effectiveness of the target products.
[0072] In step S302, obtain the classification label text input by the user, and determine the text embedding information based on the classification label text. In the embodiment of the present application, the content of step S302 is substantially the same as the content of steps S101 and S102, and will not be elaborated here.
[0073] In step S303, input each product image in the product library into the multi-modal text-image model to obtain the image embedding information corresponding to each product image output by the multi-modal text-image model. In the embodiment of the present application, the content of step S303 is substantially the same as the content of step S103, and will not be elaborated here.
[0074] In step S304, determine the target product based on the text embedding information and the image embedding information corresponding to each product image. In the embodiment of the present application, the cosine similarity between the text embedding information and the image embedding information corresponding to each product image can be determined. The cosine similarity can be used to measure the similarity between the two vectors of the text embedding information and the image embedding information corresponding to each product image, so as to realize the evaluation of the similarity between the text embedding information and the image embedding information corresponding to each product image. Then, at least one candidate matching product can be determined according to the cosine similarity. For example, the top ten products ranked by cosine similarity can be recalled as candidate matching products. Furthermore, at least one candidate matching product is sorted according to the preset reference factors, and the target product is determined according to the sorting result. Among them, the aforementioned preset reference factors may include but are not limited to product historical evaluations, product sales volumes, and user preference data, etc. Finally, the target product with the highest matching degree is determined.
[0075] Corresponding to the foregoing embodiment of the application function implementation method, the present application also provides a device for product selection and a corresponding embodiment.
[0076] Figure 4 The block diagram showing the hardware configuration of the product selection device 400 that can implement the product selection method of the embodiment of the present application. As Figure 4 shown, the product selection device 400 may include a processor 410 and a memory 420. In Figure 4In the device 400 for merchandise selection, only the constituent elements relevant to this embodiment are shown. Thus, it is obvious to those of ordinary skill in the art that the device 400 for merchandise selection may also include common constituent elements different from those shown in Figure 4 For example, a fixed-point arithmetic unit.
[0077] The device 400 for merchandise selection may correspond to a computing device with various processing functions. For example, functions such as generating a neural network, training or learning a neural network, quantizing a floating-point neural network into a fixed-point neural network, or retraining a neural network. For example, the device 400 for merchandise selection may be implemented as various types of devices, such as a personal computer (PC), a server device, a mobile device, etc.
[0078] The processor 410 controls all functions of the device 400 for merchandise selection. For example, the processor 410 controls all functions of the device 400 for merchandise selection by executing a program stored in the memory 420 on the device 400 for merchandise selection. The processor 410 may be implemented by a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), an artificial intelligence processor chip (IPU), etc. provided in the device 400 for merchandise selection. However, the present application is not limited thereto.
[0079] In some embodiments, the processor 410 may include an input / output (I / O) unit 411 and a computing unit 412. The I / O unit 411 may be used to receive various data, such as the classification label text input by the user. Exemplarily, the computing unit 412 may be used to determine text embedding information based on the classification label text received via the I / O unit 411; input each merchandise image in the merchandise library into a multimodal text-image model to obtain the image embedding information corresponding to each merchandise image output by the multimodal text-image model; wherein, the multimodal text-image model is trained based on a merchandise picture sample set and merchandise information text corresponding to each merchandise picture sample in the merchandise picture sample set; determine a target merchandise based on the text embedding information and the image embedding information corresponding to each merchandise image. This target merchandise may be output by the I / O unit 411, for example. The output data may be provided to the memory 420 for other devices (not shown) to read and use, or may be directly provided to other devices for use.
[0080] The memory 420 is hardware for storing various data processed in the device 400 for merchandise selection. For example, the memory 420 can store the processed data and the data to be processed in the device 400 for merchandise selection. The memory 420 can store the data sets involved in the merchandise selection method that the processor 410 has processed or is to process, such as the classified label text input by the user, etc. In addition, the memory 420 can store the applications, drivers, etc. to be driven by the device 400 for merchandise selection. For example, the memory 420 can store various programs related to the merchandise selection method to be executed by the processor 410. The memory 420 can be DRAM, but the present application is not limited thereto. The memory 420 can include at least one of volatile memory or non-volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, phase change RAM (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), ferroelectric RAM (FRAM), etc. The volatile memory can include dynamic RAM (DRAM), static RAM (SRAM), synchronous DRAM (SDRAM), PRAM, MRAM, RRAM, ferroelectric RAM (FeRAM), etc. In an embodiment, the memory 420 can include at least one of a hard disk drive (HDD), a solid state drive (SSD), a high density flash (CF) card, a secure digital (SD) card, a micro secure digital (Micro-SD) card, a mini secure digital (Mini-SD) card, an extreme digital (xD) card, caches, or a memory stick.
[0081] In summary, the specific functions implemented by the memory 420 and the processor 410 of the device 400 for merchandise selection provided in the embodiments of this specification can be explained in contrast to the foregoing embodiments in this specification, and can achieve the technical effects of the foregoing embodiments, and will not be elaborated here.
[0082] In this embodiment, the processor 410 can be implemented in any suitable manner. For example, the processor 410 can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and a form embedded microcontroller, and so on.
[0083] It should also be understood that any module, unit, component, server, computer, terminal, or device that executes instructions as exemplified herein may include or otherwise access a computer-readable medium, such as a storage medium, a computer storage medium, or a data storage device (removable) and / or non-removable), such as a magnetic disk, an optical disk, or a magnetic tape. A computer storage medium may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.
[0084] The foregoing may be better understood in accordance with the following clauses:
[0085] Clause A1. A method for selecting a commodity, comprising: obtaining a classification label text input by a user; determining text embedding information based on the classification label text; inputting each commodity image in a commodity library into a multimodal graphic-text model to obtain image embedding information corresponding to each commodity image output by the multimodal graphic-text model; wherein the multimodal graphic-text model is trained based on a commodity picture sample set and commodity information texts corresponding to each commodity picture sample in the commodity picture sample set; and determining a target commodity based on the text embedding information and the image embedding information corresponding to each commodity image.
[0086] Clause A2. The method for selecting a commodity according to Clause A1, wherein the training of the multimodal graphic-text model based on the commodity picture sample set and the commodity information texts corresponding to each commodity picture sample in the commodity picture sample set comprises: obtaining the commodity picture sample set and the commodity information texts corresponding to each commodity picture sample in the commodity picture sample set; respectively preprocessing each commodity picture sample and each commodity information text to obtain a commodity picture training set and a commodity information training set; loading a pre-trained model and adding a dimensionality reduction layer to the pre-trained model to obtain an initial graphic-text model; and inputting the commodity picture training set and the commodity information training set into the initial graphic-text model for training to obtain the multimodal graphic-text model.
[0087] Clause A3. The method for selecting products according to Clause A2, wherein the step of inputting the product image training set and the product information training set into the initial image-text model for training includes: pairing each training product information in the product information training set with the training product images in the product image training set to obtain a plurality of positive example pairing combinations and a plurality of negative example pairing combinations; determining a contrast loss value based on the plurality of positive example pairing combinations and the contrastive learning loss function of the initial image-text model; determining a matching loss value based on the plurality of positive example pairing combinations, the plurality of negative example pairing combinations, and the matching loss function of the initial image-text model; determining a semantic description loss value based on the product information training set and the semantic loss function of the initial image-text model; determining a target loss value based on the contrast loss value, the matching loss value, and the semantic description loss value; and determining whether the training of the initial image-text model is completed based on the target loss value.
[0088] Clause A4. The method for selecting products according to Clause A3, wherein the step of determining a contrast loss value based on the plurality of positive example pairing combinations and the contrastive learning loss function of the initial image-text model includes: respectively inputting the training product images in each positive example pairing combination into the visual encoder of the initial image-text model to obtain the product image features output by the visual encoder; respectively inputting the training product information in each positive example pairing combination into the text encoder of the initial image-text model to obtain the product information features output by the text encoder; inputting the product image features and the product information features into the dimensionality reduction layer to obtain reduced-dimensional image features and reduced-dimensional information features; determining a similarity matrix based on the reduced-dimensional image features and the reduced-dimensional information features; determining a first-direction contrast loss value and a second-direction contrast loss value based on the similarity matrix; and determining the contrast loss value based on the first-direction contrast loss value and the second-direction contrast loss value.
[0089] Clause A5. The method for selecting products according to Clause A3, wherein the step of determining a matching loss value based on the plurality of positive example pairing combinations, the plurality of negative example pairing combinations, and the matching loss function of the initial image-text model includes: determining the intermediate interaction features corresponding to each pairing combination based on the plurality of positive example pairing combinations and the plurality of negative example pairing combinations; determining the matching probability corresponding to each pairing combination according to the intermediate interaction features corresponding to each pairing combination; and determining the matching loss value based on the preset classification label parameters and the matching probability corresponding to each pairing combination.
[0090] Clause A6. The method for selecting a product according to Clause A1, wherein determining the target product based on the text embedding information and the image embedding information corresponding to each product image includes: determining the cosine similarity between the text embedding information and the image embedding information corresponding to each product image; determining at least one candidate matching product according to the cosine similarity; sorting the at least one candidate matching product according to a preset reference factor, and determining the target product according to the sorting result.
[0091] Clause A7. The method for selecting a product according to Clause A1, wherein before obtaining the classification label text input by the user, the method further includes: obtaining a list of new products; wherein the list of new products includes the new description text corresponding to each new product and the new picture information corresponding to each new product; respectively inputting the new description text corresponding to each new product and the new picture information corresponding to each new product into the multimodal graphic and text model to obtain the new text embedding information corresponding to each new product and the new image embedding information corresponding to each new product output by the multimodal graphic and text model; pushing the new text embedding information corresponding to each new product and the new image embedding information corresponding to each new product to the product library indexing engine to update the product library index.
[0092] Clause A8. The method for selecting a product according to Clause A7, wherein before obtaining the classification label text input by the user, the method further includes: monitoring the data stream of product status changes; determining the off-shelf products according to the data stream of product status changes; clearing the off-shelf text embedding information and off-shelf image embedding information corresponding to the off-shelf products in the product library index.
[0093] Clause A9. A device for product selection, including: a memory; and at least one processor configured to: obtain the classification label text input by the user; determine the text embedding information based on the classification label text; input each product image in the product library into the multimodal graphic and text model to obtain the image embedding information corresponding to each product image output by the multimodal graphic and text model; wherein the multimodal graphic and text model is trained based on a product picture sample set and the product information text corresponding to each product picture sample in the product picture sample set; determine the target product based on the text embedding information and the image embedding information corresponding to each product image.
[0094] Clause A10. A non-transitory machine-readable medium storing program code for merchandise selection, which, when executed by at least one processor, guides the execution operations of the at least one processor. The program code includes: obtaining classification label text input by a user; determining text embedding information based on the classification label text; inputting each merchandise image in a merchandise library into a multimodal text-image model to obtain image embedding information corresponding to each merchandise image output by the multimodal text-image model; wherein the multimodal text-image model is trained based on a merchandise picture sample set and merchandise information text corresponding to each merchandise picture sample in the merchandise picture sample set; and determining a target merchandise based on the text embedding information and the image embedding information corresponding to each merchandise image.
Claims
1. A commodity selection method, characterized in that: include: Get the category label text entered by the user; Determining text embedding information based on the classification label text; Input each product image in the product library into the multimodal image-text model to obtain image embedding information corresponding to each product image output by the multimodal image-text model; wherein the multimodal image-text model is trained based on a product image sample set and product information text corresponding to each product image sample in the product image sample set; The target product is determined based on the text embedding information and the image embedding information corresponding to each product image.
2. The commodity selection method according to claim 1, characterized in that: The training of the multimodal graphic-text model based on the product image sample set and the product information text corresponding to each product image sample in the product image sample set includes: Acquire the product image sample set and the product information text corresponding to each product image sample in the product image sample set; Preprocess each product image sample and each product information text respectively to obtain a product image training set and a product information training set; Load the pre-trained model and add a dimension reduction layer to the pre-trained model to obtain an initial graph-text model; The product image training set and the product information training set are input into the initial image-text model for training to obtain the multimodal image-text model.
3. The commodity selection method according to claim 2, characterized in that: The step of inputting the product image training set and the product information training set into the initial image-text model for training comprises: Pairing each training product information in the product information training set with a training product image in the product image training set to obtain a plurality of positive example pairing combinations and a plurality of negative example pairing combinations; Determining a contrastive loss value based on the contrastive learning loss function of the plurality of positive example pairing combinations and the initial image-text model; Determine a matching loss value based on the matching loss function of the plurality of positive example pairing combinations, the plurality of negative example pairing combinations, and the initial graph-text model; Determine a semantic description loss value based on the product information training set and the semantic loss function of the initial image-text model; Determine a target loss value based on the contrast loss value, the matching loss value, and the semantic description loss value; Determine whether the initial graph-text model is trained based on the target loss value.
4. The commodity selection method according to claim 3, characterized in that: The determining of the contrast loss value based on the contrast learning loss function of the plurality of positive example pairing combinations and the initial image-text model comprises: Inputting the training product images in each positive example pairing combination into the visual encoder of the initial image-text model to obtain product image features output by the visual encoder; Inputting the training product information in each positive example pairing combination into the text encoder of the initial image-text model to obtain product information features output by the text encoder; Inputting the product image features and the product information features into the dimension reduction layer to obtain reduced dimension image features and reduced dimension information features; Determine a similarity matrix based on the reduced dimension image features and the reduced dimension information features; Determine a first direction contrast loss value and a second direction contrast loss value based on the similarity matrix; The contrast loss value is determined based on the first direction contrast loss value and the second direction contrast loss value.
5. The commodity selection method according to claim 3, characterized in that: The determining of the matching loss value based on the matching loss function of the plurality of positive example pairing combinations, the plurality of negative example pairing combinations and the initial graph-text model comprises: Determine an intermediate interaction feature corresponding to each pairing combination based on the multiple positive example pairing combinations and the multiple negative example pairing combinations; Determine the matching probability corresponding to each pairing combination according to the intermediate interaction features corresponding to each pairing combination; The matching loss value is determined based on preset classification label parameters and the matching probability corresponding to each pairing combination.
6. The commodity selection method according to claim 1, characterized in that: The determining of the target product based on the text embedding information and the image embedding information corresponding to each product image comprises: Determining the cosine similarity between the text embedding information and the image embedding information corresponding to each product image; Determine at least one candidate matching product according to the cosine similarity; The at least one candidate matching commodity is sorted according to a preset reference factor, and the target commodity is determined according to the sorting result.
7. The commodity selection method according to claim 1, characterized in that: Before obtaining the classification label text input by the user, the method further includes: Obtain a list of newly added products; wherein the list of newly added products includes newly added description text corresponding to each newly added product and newly added picture information corresponding to each newly added product; Inputting the newly added description text corresponding to each newly added product and the newly added picture information corresponding to each newly added product into the multimodal graphic model respectively, and obtaining the newly added text embedding information corresponding to each newly added product and the newly added image embedding information corresponding to each newly added product output by the multimodal graphic model; The newly added text embedding information corresponding to each newly added product and the newly added image embedding information corresponding to each newly added product are pushed to the product library index engine to update the product library index.
8. The commodity selection method according to claim 7, characterized in that: Before obtaining the classification label text input by the user, the method further includes: Monitor the data stream of product status changes; Determine the product to be removed from the shelves according to the product status change data stream; The removed text embedded information and the removed image embedded information corresponding to the removed commodity are cleared from the commodity library index.
9. A device for selecting goods, characterized in that: include: Memory; as well as at least one processor configured to: Get the category label text entered by the user; Determining text embedding information based on the classification label text; Input each product image in the product library into the multimodal image-text model to obtain image embedding information corresponding to each product image output by the multimodal image-text model; wherein the multimodal image-text model is trained based on a product image sample set and product information text corresponding to each product image sample in the product image sample set; The target product is determined based on the text embedding information and the image embedding information corresponding to each product image.
10. A non-transitory machine-readable medium having a program code for commodity selection stored thereon, wherein when the program code is executed by at least one processor, the code guides the execution operation of the at least one processor, the program code comprising: Get the category label text entered by the user; Determining text embedding information based on the classification label text; Input each product image in the product library into the multimodal image-text model to obtain image embedding information corresponding to each product image output by the multimodal image-text model; wherein the multimodal image-text model is trained based on a product image sample set and product information text corresponding to each product image sample in the product image sample set; The target product is determined based on the text embedding information and the image embedding information corresponding to each product image.