Commodity identification method and device based on multi-modal model similarity matching
Through the multimodal model similarity matching method, combined with image and text features, the scalability and data requirements of the product recognition method in the prior art are solved, and a more comprehensive product recognition effect and higher recognition accuracy are achieved.
Patent Information
- Application Number
- CN202510383500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the commodity identification method has defects such as poor scalability, limited classification products and high data requirements. In particular, the model needs to be retrained and the data requirements are high when using classification models.
The multimodal model similarity matching method is adopted to obtain the image and text features of the product, calculate the similarity scores between the image and pre-stored features, and combine the similarity scores of the image and text features to determine the product category, use the CLIP model to extract image features and combine the BERT model to extract text features, and use cosine similarity calculation and two-way cross entropy loss function to optimize the model.
It achieves a richer and comprehensive product recognition effect, reduces dependence on a single mode, improves the robustness and generalization capabilities of the model, reduces data demand and misidentification rate, and improves recognition accuracy and user experience.
Smart Images

Figure CN120337126A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of commodity identification, and particularly to a commodity identification method and device based on multi-modal model similarity matching. Background Art
[0002] With the development of Internet of Things technology, intelligent vending cabinets have become more and more popular. Customers unlock the cabinet door by means of mobile phone scanning code, face recognition, etc., freely select the commodities in the cabinet, and the cabinet captures the video of the user's purchase process through a camera, uses a detection model to detect the images of the commodities purchased by the user from the video, and then crops these commodity images for transmission to an identification model, which is responsible for classifying and identifying the types of commodities to obtain the commodity information purchased by the user for settlement.
[0003] In the current related technologies, classification models (such as ResNet, shuffleNet, etc.) are usually used to classify commodity pictures to determine the categories of purchased commodities. However, there are the following problems with using such classification methods: poor scalability, and each time a new category of commodity is added, a new version of the model needs to be retrained and updated; the commodities that can be classified are limited. There are thousands of types of commodities, and it is not very realistic to directly train a classification model with so many categories, and the classification effect will not be very ideal; high data requirements, for each category of commodity, a large amount of data needs to be collected and labeled to improve the accuracy of the classification model. Therefore, there is an urgent need for a new commodity identification method to avoid the above problems existing in the prior art. Summary of the Invention
[0004] The present invention provides a commodity identification method and device based on multi-modal model similarity matching, which are used to solve the defects of poor scalability, limited classifiable commodities, and high data requirements existing in the prior art when using a classification model to classify commodity pictures, achieve a richer and more comprehensive commodity identification effect, and at the same time reduce the demand for commodity data.
[0005] The present invention provides a commodity identification method based on multi-modal model similarity matching, including:
[0006] Obtain an image of a purchased commodity, input the image of the purchased commodity into a preset multi-modal model, and obtain the image features of the purchased commodity output by the multi-modal model;
[0007] Calculate the figure-figure similarity scores between the image features of the purchased commodity and the pre-stored image features of each sold commodity in a preset feature library; the feature library stores pre-stored image features and pre-stored text features of multiple sold commodities;
[0008] Calculate the figure-text similarity scores between the image features and the pre-stored text features of each sold commodity in the feature library;
[0009] Based on the similarity scores of the figure - figure and figure - text of the image features and each sold commodity in the feature library, calculate the overall similarity score of the image features and each sold commodity in the feature library;
[0010] Based on the overall similarity score of the image features and each sold commodity in the feature library, determine the category of the purchased commodity corresponding to the image features.
[0011] According to a commodity recognition method based on multi - modal model similarity matching provided by the present invention, it further includes a training method of the multi - modal model:
[0012] Based on multi - modal training sample pairs, extract the sample image features and sample text features of the training sample pairs; the training sample pairs include positive sample pairs and negative sample pairs, and both the positive sample pairs and the negative sample pairs include the text description of the sold commodity and a picture in the commodity registration picture set.
[0013] Calculate the similarity scores of the sample image features and the sample text features;
[0014] Use a bidirectional cross - entropy loss function to fine - tune the multi - modal model so that the similarity score of the positive sample pairs is maximized and the similarity of the negative sample pairs is minimized.
[0015] According to a commodity recognition method based on multi - modal model similarity matching provided by the present invention, the commodity registration picture set includes multiple pictures of the sold commodity at different angles, lights, and positions; the text description includes literal descriptions of the name, weight, color, packaging, and manufacturer of the sold commodity.
[0016] According to a commodity recognition method based on multi - modal model similarity matching provided by the present invention, it further includes a construction method of the feature library:
[0017] Obtain the commodity registration picture set and the text description of each sold commodity;
[0018] Input the commodity registration picture set and the text description of each sold commodity into the trained multi - modal model, and obtain the image features and text features of each sold commodity output by the multi - modal model;
[0019] Use the image features corresponding to the sold commodity as the pre - stored image features, use the text features corresponding to the sold commodity as the pre - stored text features, and construct the feature library based on the pre - stored image features and pre - stored text features of all sold commodities.
[0020] A commodity recognition method based on multi-modal model similarity matching provided by the present invention calculates the similarity scores between the image features of the purchased commodity and the pre-stored image features of each sold commodity in a preset feature library, including:
[0021] Calculate the first cosine similarity between the image features and the pre-stored image features of each sold commodity in the feature library, and use the first cosine similarity as the similarity score;
[0022] Calculate the image-text similarity scores between the image features and the pre-stored text features of each sold commodity in the feature library, including:
[0023] Calculate the second cosine similarity between the image features and the pre-stored text features of each sold commodity in the feature library, and use the second cosine similarity as the image-text similarity score.
[0024] A commodity recognition method based on multi-modal model similarity matching provided by the present invention calculates the overall similarity scores between the image features and each sold commodity in the feature library based on the similarity scores and the image-text similarity scores, including:
[0025] For each sold commodity in the feature library, set a first weight for the similarity score between the image features and the sold commodity, set a second weight for the image-text similarity score between the image features and the sold commodity, calculate the weighted sum of the similarity score and the image-text similarity score, and use the weighted sum as the overall similarity score between the image features and the sold commodity.
[0026] A commodity recognition method based on multi-modal model similarity matching provided by the present invention determines the category of the purchased commodity corresponding to the image features based on the overall similarity scores between the image features and each sold commodity in the feature library, including:
[0027] Arrange the overall similarity scores between the image features and each sold commodity in the feature library, and select the sold commodity corresponding to the highest overall similarity score as the category of the purchased commodity corresponding to the image features.
[0028] The present invention also provides a commodity recognition device based on multi-modal model similarity matching, including:
[0029] A feature extraction module for obtaining an image of a purchased commodity, inputting the image of the purchased commodity into a preset multi-modal model, and obtaining the image features of the purchased commodity output by the multi-modal model;
[0030] A similarity calculation module is used to calculate the figure - figure similarity scores between the image features of the purchased product and the pre - stored image features of each sold product in a preset feature library. The feature library stores the pre - stored image features and pre - stored text features of multiple sold products. Calculate the figure - text similarity scores between the image features and the pre - stored text features of each sold product in the feature library. Based on the figure - figure similarity scores and the figure - text similarity scores between the image features and the pre - stored image features of each sold product in the feature library, calculate the overall similarity scores between the image features and the pre - stored image features of each sold product in the feature library.
[0031] A result determination module is used to determine the category of the purchased product corresponding to the image features based on the overall similarity scores between the image features and the pre - stored image features of each sold product in the feature library.
[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the commodity recognition method based on multi - modal model similarity matching as described in any one of the above.
[0033] The present invention also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the commodity recognition method based on multi - modal model similarity matching as described in any one of the above.
[0034] The commodity recognition method and device based on multi-modal model similarity matching provided by the present invention obtain the image of the purchased commodity, input the image of the purchased commodity into a preset multi-modal model, and obtain the image features of the purchased commodity output by the multi-modal model; calculate the figure-figure similarity scores between the image features of the purchased commodity and the pre-stored image features of each sold commodity in the preset feature library; the feature library stores the pre-stored image features and pre-stored text features of multiple sold commodities; calculate the image-text similarity scores between the image features and the pre-stored text features of each sold commodity in the feature library; based on the figure-figure similarity scores and the image-text similarity scores between the image features and the pre-stored text features of each sold commodity in the feature library, calculate the overall similarity scores between the image features and each sold commodity in the feature library; based on the overall similarity scores between the image features and each sold commodity in the feature library, determine the category of the purchased commodity corresponding to the image features. The method provided by the present invention introduces data of two modalities, image and text, for similarity matching, so as to obtain richer and more comprehensive data representations. By combining information of different modalities, the over-reliance on a single modality is reduced, thereby increasing the robustness of the model. Multi-modal similarity matching can make full use of the correlation between different modality data, improve the representation ability and generalization ability of the model, and thus improve the accuracy of similarity matching. In addition, the multi-modal model adopted by the present invention can learn richer feature representations, which helps to improve the generalization ability of the model on unseen data and achieve a richer and more comprehensive commodity recognition effect. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 It is a schematic flowchart of the commodity recognition method based on multi-modal model similarity matching provided by the present invention;
[0037] Figure 2 It is a schematic structural diagram of the commodity recognition device based on multi-modal model similarity matching provided by the present invention;
[0038] Figure 3 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0039] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] The following combines Figure 1 to describe the commodity recognition method based on multi-modal model similarity matching of the present invention.
[0041] As Figure 1 shown, the commodity recognition method based on multi-modal model similarity matching provided by the present invention includes the following steps:
[0042] S1. Obtain an image of the purchased commodity, input the image of the purchased commodity into a preset multi-modal model, and obtain the image features of the purchased commodity output by the multi-modal model.
[0043] Preferably, in this step, the multi-modal model adopted by the present invention is the CLIP (Contrastive Language–Image Pretraining) model. The CLIP model is a multi-modal pre-training model proposed by OpenAI, aiming to associate images and texts so that they can be compared and matched in the same feature space. The CLIP model includes an image encoder and a text encoder. In the present invention, the image encoder part of the CLIP model used is ViT-base, that is, the Vision Transformer (ViT) architecture is used as the specific implementation. ViT relies on the Self-Attention mechanism to capture global and local features in the image. In the CLIP model, ViT-base serves as the image encoder, responsible for converting the input image into high-dimensional feature vectors (i.e., image features), and these vectors can be used for subsequent similarity calculations.
[0044] In this step, the purchased commodity is the commodity selected by the user. In an alternative embodiment of the present invention, before step S1, it further includes: obtaining a video of the user purchasing a commodity, and cropping out the image of the purchased commodity from the video. Specifically, first extract the frames containing the commodity from the video of the purchase process. Video processing tools (such as OpenCV) can be used to read the video frame by frame and select the frames containing the commodity. For example, if the user shows the commodity in the video, the key frames can be located by detecting the user's gestures or specific positions of the commodity. The extracted frames may contain backgrounds or other irrelevant information, so the image needs to be cropped and preprocessed to ensure that only the commodity part is retained. Image segmentation techniques (such as Mask R-CNN) can be used to accurately extract the commodity area.
[0045] In an alternative embodiment of the present invention, the commodity recognition method based on multi-modal model similarity matching further includes a training method for the multi-modal model:
[0046] Step 1: Based on the multi-modal training sample pairs, extract the sample image features and sample text features of the training sample pairs.
[0047] Specifically, the training sample pairs include positive sample pairs and negative sample pairs. Both the positive sample pairs and negative sample pairs include the text description of the sold commodity and a picture in the commodity registration picture set. If the commodity picture and the commodity description are of the same commodity, then the sample pair is a positive sample; if they are of different commodities, then the sample pair is a negative sample. Use the image feature extraction model (ViT-base model) of the multi-modal model (CLIP model) to extract the image features of the commodity pictures, and use the text feature extraction model (BERT-base, a model based on the Bidirectional Encoder Representations from Transformers architecture) of the multi-modal model to extract the text features of the commodity pictures.
[0048] Step 2: Calculate the similarity scores of the sample image features and sample text features.
[0049] Specifically, use cosine similarity to measure the matching degree between the image features and text features. For a batch of image-text pairs, the multi-modal model will generate a similarity matrix, where each element represents the similarity between an image feature and a text feature. The goal is to make the similarity scores of the matching image-text pairs (positive sample pairs) as high as possible, while the similarity scores of the non-matching image-text pairs (negative sample pairs) as low as possible.
[0050] Step 3: Use the bidirectional cross-entropy loss function to fine-tune the multi-modal model to maximize the similarity scores of the positive sample pairs and minimize the similarity scores of the negative sample pairs.
[0051] Specifically, for each image-text pair, the multi-modal model will calculate the losses in two directions:
[0052] Loss from image to text: Calculate the similarity between the image features and all text features, and use cross-entropy loss to optimize it to maximize the similarity of the matching text features.
[0053] Loss from text to image: Calculate the similarity between the text features and all image features, and use cross-entropy loss to optimize it to maximize the similarity of the matching image features.
[0054] The final total loss is the average of the losses in these two directions. Through backpropagation and optimization algorithms (such as Adam), the multi-modal model continuously adjusts its parameters to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs, thus completing the training of the model.
[0055] S2. Calculate the similarity score between the image features of the purchased commodity and the pre-stored image features of each sold commodity in the preset feature library.
[0056] Specifically, after obtaining the image features of the purchased commodity, calculate the similarity between the image features of the purchased commodity and the image features of each commodity in the feature library. The similarity is the similarity between image features and image features. The higher the similarity, the closer the two images are in the feature space and the more similar the content is. Preferably, the specific method for calculating the similarity in the present invention is: calculate the first cosine similarity between the image features and the pre-stored image features of each sold commodity in the feature library, and use the first cosine similarity as the similarity score. Among them, the feature library stores pre-stored image features and pre-stored text features of multiple sold commodities. The sold commodities in the feature library should include all commodities that customers can purchase from systems such as vending machines. For example, if the sold commodities include three types: a, b, and c, then the feature library should at least include the pre-stored image features and pre-stored text features of these three commodities a, b, and c.
[0057] In an alternative embodiment of the present invention, the commodity recognition method based on multi-modal model similarity matching further includes a method for constructing a feature library:
[0058] Step 1: Obtain the commodity registration picture set and text description of each sold commodity.
[0059] Among them, "each sold commodity" refers to each commodity included in the sales system; the commodity registration picture set includes but is not limited to multiple pictures of the sold commodity at different angles, lights, and positions; the text description includes but is not limited to the name, weight, color, packaging, and literal description of the manufacturer of the sold commodity. Taking the soda of brand A as an example, the registration picture set includes multiple pictures of brand A soda at different angles (such as the front, side, and top), lights (such as bright and dim), and positions (such as on the table and in hand), and the text description includes the name of brand A soda (such as "Brand A lemon-flavored soda"), weight (such as "330ml"), color (such as "green"), packaging (such as "aluminum can"), and literal description of the manufacturer (such as "Company A").
[0060] Step 2: Input the commodity registration picture set and text description of each sold commodity into the trained multi-modal model, and obtain the image features and text features of each sold commodity output by the multi-modal model.
[0061] Specifically, the image part model ViT-base of the trained multi-modal model (CLIP model) is used to extract the image features of the images in the registered image set, and the text part model of the multi-modal model is used to extract the text features of the product descriptions.
[0062] Step 3: Use the image features corresponding to the sold products as the pre-stored image features, and use the text features corresponding to the sold products as the pre-stored text features, and construct a feature library based on the pre-stored image features and pre-stored text features of all sold products.
[0063] S3. Calculate the image-text similarity scores between the image features and the pre-stored text features of each sold product in the feature library.
[0064] Specifically, after obtaining the image features of the purchased product, calculate the image-text similarity between the image features of the purchased product and the image features of each product in the feature library. Calculating the cosine similarity between the image features of the purchased product and the text features of each product in the feature library is a cross-modal similarity calculation method used to measure the semantic association degree between an image and text. The image-text similarity score represents the matching degree between the image of the purchased product and the text description of the product in the feature library. Preferably, the specific method for calculating the image-text similarity in the present invention is: calculate the second cosine similarity between the image features and the pre-stored text features of each sold product in the feature library, and use the second cosine similarity as the image-text similarity score.
[0065] It should be understood that there is no strict sequence for steps S2 and S3, and they can be carried out simultaneously or separately in sequence.
[0066] S4. Based on the image-image similarity scores and image-text similarity scores between the image features and each sold product in the feature library, calculate the overall similarity scores between the image features and each sold product in the feature library.
[0067] Specifically, for each sold product in the feature library, set a first weight for the image-image similarity score between the image features and the sold product, and set a second weight for the image-text similarity score between the image features and the sold product, calculate the weighted sum of the image-image similarity score and the image-text similarity score, and use the weighted sum as the overall similarity score between the image features and the sold product. For example, assume that the image-image similarity score S II′ = 0.8 between the image I of the purchased product and the image I' of the sold product, and the image-image similarity score S IT = 0.6 between the image I of the purchased product and the image T of the sold product. The weight of the image-image similarity score is set as w1 = 0.7, and the weight of the image-text similarity score is set as w2 = 0.3. Then the overall similarity score between the image features of the purchased product and the sold product is: S 整体=0.7·0.8 + 0.3·0.6 = 0.56 + 0.18 = 0.74。
[0068] S5. Determine the category of the purchased product corresponding to the image feature based on the overall similarity scores between the image feature and each sold product in the feature library.
[0069] Specifically, based on the overall similarity scores between the image feature of the purchased product and each sold product in the feature library, arrange the overall similarity scores between the image feature and each sold product in the feature library, and select the sold product corresponding to the highest overall similarity score as the category of the purchased product corresponding to the image feature. In an alternative embodiment of the present invention, the overall similarity scores between the image feature and each sold product in the feature library can be arranged in descending order, and the sold product with the highest overall similarity score can be selected as the category of the purchased product by means of strict matching, or the three sold products with the highest overall similarity scores can be selected by means of loose matching, and the category of the purchased product is one of these three sold products.
[0070] The following takes the soda of a brand as an example to give an overall description of the product recognition method based on multi-modal model similarity matching provided by the present invention:
[0071] Suppose only the soda of brand B and the juice of brand C are stored in the feature library, the image feature of the soda of brand B is I1, and the image feature of the juice of brand C is I2; the text feature of the soda of brand B is T1, and the text feature of the juice of brand C is T2. A user has purchased a bottle of soda, and the image feature vector extracted by the multi-modal model is A.
[0072] Calculate the first cosine similarity between the image feature A and the pre-stored image features of each sold product in the feature library. Suppose the first cosine similarity between the image feature A and the soda of brand B is S I1 =0.90, the first cosine similarity between the image feature A and the soda of brand C is S I2 =0.40, the second cosine similarity between the image feature A and the soda of brand B is S T1 =0.85, the first cosine similarity between the image feature A and the soda of brand C is S T2 =0.30. Suppose the first weight w1 of the figure similarity score is 0.6, and the second weight w2 of the figure-text similarity score is 0.4. The overall similarity score = w1×S Ii +w2×S Ti , and calculate the overall similarity score between the image feature A and the soda of brand B:
[0073] 0.6×0.90 + 0.4×0.85 = 0.54 + 0.34 = 0.88;
[0074] The overall similarity score between the calculated image feature A and the C-brand soda is obtained:
[0075] 0.6×0.40 + 0.4×0.30 = 0.24 + 0.12 = 0.36.
[0076] As can be seen from the above calculation, the overall similarity score between the image feature A and the B-brand soda is the highest. Therefore, it can be determined that the purchased product is the B-brand soda.
[0077] In summary, the commodity recognition method based on multi-modal model similarity matching provided by the present invention obtains the image of the purchased commodity, inputs the image of the purchased commodity into a preset multi-modal model to obtain the image feature of the purchased commodity output by the multi-modal model; calculates the figure-to-figure similarity score between the image feature of the purchased commodity and the pre-stored image features of each sold commodity in the preset feature library; the feature library stores the pre-stored image features and pre-stored text features of multiple sold commodities; calculates the figure-to-text similarity score between the image feature and the pre-stored text features of each sold commodity in the feature library; based on the figure-to-figure similarity score and the figure-to-text similarity score between the image feature and each sold commodity in the feature library, calculates the overall similarity score between the image feature and each sold commodity in the feature library; based on the overall similarity score between the image feature and each sold commodity in the feature library, determines the category of the purchased commodity corresponding to the image feature. The method provided by the present invention designs a large commodity recognition model based on multi-modal model similarity matching to solve problems such as rich commodity types, high similarity of commodity packaging, and high misrecognition rate in the commodity recognition scenario. By introducing data of two modalities, image and text, for similarity matching, a richer and more comprehensive data representation can be obtained. At the same time, multi-modal can reduce the over-reliance on a certain modality by combining information of different modalities, thereby increasing the robustness of the model. In addition, multi-modal similarity matching can make full use of the correlation between data of different modalities to improve the representation ability and generalization ability of the model, thereby improving the accuracy of similarity matching. In addition, the multi-modal model also adopts a network structure based on transformer to construct our multi-modal model. Due to having more parameters and deeper layers, the large model can learn richer feature representations, which helps to improve the generalization ability of the model on unseen data, which is particularly important in commodity recognition because the sold commodities are constantly updated and changed. By using multi-modal and large model, the accuracy of the commodity recognition model has been greatly improved, which is beneficial to improving the user experience, enhancing the user's satisfaction and trust, and at the same time can reduce the manual intervention and error correction operations caused by misrecognition, and can effectively reduce the operation cost.
[0078] Based on the same inventive concept, the present invention also provides a commodity recognition device based on multi-modal model similarity matching. The commodity recognition device based on multi-modal model similarity matching provided by the present invention will be described below. The commodity recognition device based on multi-modal model similarity matching described below can be correspondingly referred to the commodity recognition method based on multi-modal model similarity matching described above.
[0079] As Figure 2 shown, the commodity recognition device based on multi-modal model similarity matching provided by the present invention includes a feature extraction module 21, a similarity calculation module 22, and a result determination module 23.
[0080] The feature extraction module 21 is used to obtain an image of the purchased commodity, input the image of the purchased commodity into a preset multi-modal model, and obtain the image features of the purchased commodity output by the multi-modal model.
[0081] The similarity calculation module 22 is used to calculate the figure-figure similarity scores between the image features of the purchased commodity and the pre-stored image features of each sold commodity in the preset feature library; the feature library stores the pre-stored image features and pre-stored text features of multiple sold commodities; calculate the figure-text similarity scores between the image features and the pre-stored text features of each sold commodity in the feature library; based on the figure-figure similarity scores and figure-text similarity scores between the image features and the pre-stored image features of each sold commodity in the feature library, calculate the overall similarity scores between the image features and the pre-stored image features of each sold commodity in the feature library.
[0082] The result determination module 23 is used to determine the category of the purchased commodity corresponding to the image features based on the overall similarity scores between the image features and the pre-stored image features of each sold commodity in the feature library.
[0083] As Figure 2 shown, in an optional embodiment of the present invention, the commodity recognition device based on multi-modal model similarity matching provided by the present invention further includes a data processing module 20, which is used to obtain a video of the user's purchased commodity and crop the image of the purchased commodity from the video.
[0084] Figure 3 Illustrates a schematic physical structure diagram of an electronic device. As Figure 3 shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete communication with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the commodity recognition method based on multi-modal model similarity matching provided by each of the above methods. The method includes:
[0085] Obtain an image of a purchased commodity, input the image of the purchased commodity into a preset multimodal model, and obtain the image features of the purchased commodity output by the multimodal model;
[0086] Calculate the similarity score between the image features of the purchased commodity and the pre-stored image features of each sold commodity in a preset feature library; multiple pre-stored image features and pre-stored text features of sold commodities are stored in the feature library;
[0087] Calculate the similarity score between the image features and the pre-stored text features of each sold commodity in the feature library;
[0088] Based on the similarity scores between the image features and the pre-stored image features and the similarity scores between the image features and the pre-stored text features of each sold commodity in the feature library, calculate the overall similarity score between the image features and each sold commodity in the feature library;
[0089] Based on the overall similarity score between the image features and each sold commodity in the feature library, determine the category of the purchased commodity corresponding to the image features.
[0090] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0091] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the commodity recognition method based on multimodal model similarity matching provided by the above-mentioned various methods. The method includes:
[0092] Obtain an image of a purchased commodity, input the image of the purchased commodity into a preset multimodal model, and obtain the image features of the purchased commodity output by the multimodal model;
[0093] Calculate the figure - figure similarity score between the image features of the purchased commodity and the pre - stored image features of each sold commodity in the preset feature library; the feature library stores pre - stored image features and pre - stored text features of multiple sold commodities;
[0094] Calculate the figure - text similarity score between the image features and the pre - stored text features of each sold commodity in the feature library;
[0095] Based on the figure - figure similarity score and the figure - text similarity score between the image features and the pre - stored text features of each sold commodity in the feature library, calculate the overall similarity score between the image features and each sold commodity in the feature library;
[0096] Based on the overall similarity score between the image features and each sold commodity in the feature library, determine the category of the purchased commodity corresponding to the image features.
[0097] In another aspect, the present invention also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the commodity recognition method based on multi - modal model similarity matching provided by the above - mentioned various methods. The method includes:
[0098] Obtain an image of a purchased commodity, input the image of the purchased commodity into a preset multi - modal model, and obtain the image features of the purchased commodity output by the multi - modal model;
[0099] Calculate the figure - figure similarity score between the image features of the purchased commodity and the pre - stored image features of each sold commodity in the preset feature library; the feature library stores pre - stored image features and pre - stored text features of multiple sold commodities;
[0100] Calculate the figure - text similarity score between the image features and the pre - stored text features of each sold commodity in the feature library;
[0101] Based on the figure - figure similarity score and the figure - text similarity score between the image features and the pre - stored text features of each sold commodity in the feature library, calculate the overall similarity score between the image features and each sold commodity in the feature library;
[0102] Based on the overall similarity score between the image features and each sold commodity in the feature library, determine the category of the purchased commodity corresponding to the image features.
[0103] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0104] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A commodity recognition method based on multi-modal model similarity matching, characterized in that Including: Obtain an image of a purchased commodity, input the image of the purchased commodity into a preset multimodal model, and obtain the image features of the purchased commodity output by the multimodal model; Calculate the figure-to-figure similarity score between the image features of the purchased commodity and the pre-stored image features of each sold commodity in a preset feature library; the feature library stores pre-stored image features and pre-stored text features of multiple sold commodities; Calculate the figure-to-text similarity score between the image features and the pre-stored text features of each sold commodity in the feature library; Based on the figure-to-figure similarity score and the figure-to-text similarity score between the image features and the pre-stored image features of each sold commodity in the feature library, calculate the overall similarity score between the image features and each sold commodity in the feature library; Based on the overall similarity score between the image features and each sold commodity in the feature library, determine the category of the purchased commodity corresponding to the image features.
2. The commodity recognition method based on multi-modal model similarity matching according to claim 1, wherein, It also includes the training method of the multimodal model: Based on multimodal training sample pairs, extract the sample image features and sample text features of the training sample pairs; The training sample pairs include positive sample pairs and negative sample pairs, and both the positive sample pairs and the negative sample pairs include the text description of the sold commodity and a picture in the commodity registration picture set; Calculate the similarity score between the sample image features and the sample text features; Use the bidirectional cross-entropy loss function to fine-tune the multimodal model to maximize the similarity score of the positive sample pairs and minimize the similarity of the negative sample pairs.
3. The commodity recognition method based on multimodal model similarity matching according to claim 2, wherein The commodity registration picture set includes multiple pictures of the sold commodity at different angles, lights, and positions; the text description includes literal descriptions of the name, weight, color, packaging, and manufacturer of the sold commodity.
4. The method for identifying goods based on multi-modal model similarity matching according to claim 3, wherein, It also includes the construction method of the feature library: Obtain the commodity registration picture set and the text description of each sold commodity; Input the commodity registration picture set and the text description of each sold commodity into the trained multimodal model, and obtain the image features and the text features of each sold commodity output by the multimodal model; Use the image features corresponding to the sold commodity as the pre-stored image features, use the text features corresponding to the sold commodity as the pre-stored text features, and construct the feature library based on the pre-stored image features and the pre-stored text features of all sold commodities.
5. The commodity recognition method based on multimodal model similarity matching according to claim 1, characterized in that Calculating the figure-to-figure similarity score between the image features of the purchased commodity and the pre-stored image features of each sold commodity in a preset feature library includes: Calculate the first cosine similarity between the image features and the pre-stored image features of each sold commodity in the feature library, and use the first cosine similarity as the figure-to-figure similarity score; Calculating the figure-to-text similarity score between the image features and the pre-stored text features of each sold commodity in the feature library includes: Calculate the second cosine similarity between the image features and the pre-stored text features of each sold commodity in the feature library, and use the second cosine similarity as the figure-to-text similarity score.
6. The commodity recognition method based on multimodal model similarity matching according to claim 5, characterized in that, Based on the image features, the figure - figure similarity scores and the figure - text similarity scores of each sold commodity in the feature library, calculate the overall similarity score between the image features and each sold commodity in the feature library, including: For each sold commodity in the feature library, set a first weight for the figure - figure similarity score between the image features and the sold commodity, set a second weight for the figure - text similarity score between the image features and the sold commodity, calculate the weighted sum of the figure - figure similarity score and the figure - text similarity score, and use the weighted sum as the overall similarity score between the image features and the sold commodity.
7. The method for identifying a product based on multi-modal model similarity matching according to any one of claims 1-6, characterized in that, Based on the overall similarity score between the image features and each sold commodity in the feature library, determine the category of the purchased commodity corresponding to the image features, including: Arrange the overall similarity scores between the image features and each sold commodity in the feature library, and select the sold commodity corresponding to the highest overall similarity score as the category of the purchased commodity corresponding to the image features.
8. A commodity recognition device based on multi-modal model similarity matching, characterized in that, Including: A feature extraction module, configured to obtain an image of a purchased commodity, input the image of the purchased commodity into a preset multi - modal model, and obtain the image features of the purchased commodity output by the multi - modal model; A similarity calculation module, configured to calculate the figure - figure similarity score between the image features of the purchased commodity and the pre - stored image features of each sold commodity in a preset feature library; the feature library stores pre - stored image features and pre - stored text features of multiple sold commodities; calculate the figure - text similarity score between the image features and the pre - stored text features of each sold commodity in the feature library; based on the figure - figure similarity score and the figure - text similarity score between the image features and each sold commodity in the feature library, calculate the overall similarity score between the image features and each sold commodity in the feature library; A result determination module, configured to determine the category of the purchased commodity corresponding to the image features based on the overall similarity score between the image features and each sold commodity in the feature library.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the commodity recognition method based on multi - modal model similarity matching according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the commodity recognition method based on multi - modal model similarity matching according to any one of claims 1 to 7.
Citation Information
Cited By
Goods and commodity matching method, equipment, medium and product
CN120996911A