Commodity identification method, device and equipment and storage medium
By using object detection models and OCR technology, combined with multi-angle white background images and text content, the system automatically identifies product labels, solving the problems of non-standard numbering and diverse quantities in SKU identification, and improving the efficiency and accuracy of product management and marketing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-24
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies suffer from issues such as non-standard numbering and errors in SKU identification, leading to low efficiency in product management. Furthermore, the large number of SKUs and the emergence of new numbers pose even greater challenges.
The system employs object detection and feature extraction models to obtain detection regions from images to be identified. It then uses a zero-shot object detection model and OCR recognition technology, combined with multi-angle white background images and text content, to automatically identify product labels.
It improved the efficiency and accuracy of merchandise management and marketing, and ensured the accuracy of SKU identification and diversified management.
Smart Images

Figure CN116824269B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure relate to the field of computer vision technology and related technical fields, and more specifically, to a product identification method, apparatus, device, and storage medium. Background Technology
[0002] With the rapid development of content marketing, the explosive growth of product information in various forms has made product management one of the essential problems to be solved in the content marketing field. In product management, the SKU (Stock Keeping Unit) is a unique identifier for a product, effectively helping companies track and manage their goods. SKU identification, as an indispensable part of product management, has gradually become a crucial issue of great concern in the content marketing field.
[0003] However, in the actual process of SKU identification, various problems are often encountered. For example, the SKU numbers of some products are not standardized, and there may be typos, errors, etc., making it difficult to accurately identify them using traditional methods. At the same time, the large number of SKUs and the continuous emergence of new SKU numbers also bring greater challenges to SKU management.
[0004] Given the problems with existing technologies, there is an urgent need for a product identification method to improve the efficiency and accuracy of product management in the content marketing field. Summary of the Invention
[0005] The embodiments described herein provide a product identification method, apparatus, device, and storage medium that address the problems existing in the prior art.
[0006] Firstly, based on the content of this disclosure, a product identification method is provided, including:
[0007] Based on the object detection model, the detection area corresponding to the detection tag identification group is obtained from the image to be identified, wherein the image to be identified is a marketing image of the product to be detected, and the detection tag identification group includes at least one tag identification in the target product category to which the product to be detected belongs;
[0008] Determine the target white background image corresponding to the detection area, wherein the target white background image is any one of the white background images of the products included in the target product category to which the product to be detected belongs, and the white background image is an image obtained by processing the product image, and the same product includes white background images from different angles;
[0009] Based on the target white background image, the target label identifier of the image to be identified is determined.
[0010] In some embodiments of this disclosure, determining the target white background image corresponding to the detection area includes:
[0011] Based on the feature extraction model, the target feature vector corresponding to the detection region is extracted;
[0012] The similarity between the target feature vector and the feature vector corresponding to each white background image is compared sequentially. The white background image is a white background image of a product image of a product in the target product category to which the product to be detected belongs.
[0013] The white background image corresponding to the feature vector with the highest similarity is selected as the target white background image corresponding to the detection region.
[0014] In some embodiments of this disclosure, before determining the target label identifier of the image to be identified based on the target white background image, the method further includes:
[0015] Construct a first association table between white background images and first preset label identifiers;
[0016] The step of determining the target label identifier of the image to be identified based on the target white background image includes:
[0017] Select the first preset label identifier corresponding to the target white background image from the first association table as the target label identifier of the image to be identified.
[0018] In some embodiments of this disclosure, when determining the target white background image corresponding to the detection area, the method further includes:
[0019] Determine the target text content corresponding to the detection area;
[0020] The step of determining the target label identifier of the image to be identified based on the target white background image includes:
[0021] Based on the target white background image and the target text content, the target label identifier of the image to be identified is determined.
[0022] In some embodiments of this disclosure, before determining the target label identifier of the image to be identified based on the target white background image, the method further includes:
[0023] Construct a first association table between white background images and first preset label identifiers;
[0024] Construct a second association table between preset text content and second preset tag identifier, wherein the preset text content is the text content of the product image of the product to be detected, which is included in the target product category;
[0025] The step of determining the target label identifier of the image to be identified based on the target white background image and the target text content includes:
[0026] Select the first target preset label identifier corresponding to the target white background image from the first association table;
[0027] Select the second target preset tag identifier corresponding to the target text content from the second association table;
[0028] The target label of the image to be identified is determined based on the relationship between the first target preset label and the second target preset label.
[0029] In some embodiments of this disclosure, determining the target label identifier of the image to be identified based on the relationship between the first target preset label identifier and the second target preset label identifier includes:
[0030] When the first target preset label identifier and the second target preset label identifier are the same, the first target preset label identifier or the second target preset label identifier is determined as the target label identifier of the image to be identified;
[0031] When the first target preset label identifier and the second target preset label identifier are different, the second target preset label identifier is determined to be the target label identifier of the image to be identified.
[0032] In some embodiments of this disclosure, obtaining the detection region corresponding to the detection label group from the image to be identified based on the target detection model includes:
[0033] Based on the target detection model, the location information of the detection area corresponding to the detection label identification group in the image to be identified is determined;
[0034] The detection area is obtained by processing the image to be identified based on the location information of the detection area.
[0035] Secondly, according to the present disclosure, a product identification device is provided, comprising:
[0036] The detection area acquisition module is used to acquire the detection area corresponding to the detection label identification group from the image to be identified based on the target detection model. The image to be identified is a marketing image of the product to be detected, and the detection label identification group includes at least one label identification in the target product category to which the product to be detected belongs.
[0037] The target white background image determination module is used to determine the target white background image corresponding to the detection area. The target white background image is any one of the white background images of the products included in the target product category to which the product to be detected belongs. The white background image is an image obtained by processing the product image. The same product includes white background images from different angles.
[0038] The target label identification module is used to determine the target label identifier of the image to be identified based on the target white background image.
[0039] Thirdly, according to the present disclosure, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method as described in any of the above embodiments.
[0040] Fourthly, according to the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when executed by a processor, the computer program implements the steps of the method as described in any of the above embodiments.
[0041] The product identification method, apparatus, device, and storage medium provided in this disclosure first obtain a detection area corresponding to a detection tag group from an image to be identified based on a target detection model. The image to be identified is a marketing image of the product to be identified, and the detection tag group includes at least one tag from the target product category to which the product belongs. Then, a target white-background image corresponding to the detection area is determined. Finally, based on the target white-background image, the target tag of the image to be identified is determined. That is, only multi-angle white-background images and corresponding tag icons are needed. The zero-shot target detection model enables automatic tagging and identification of marketing images, improving the efficiency and accuracy of product management and marketing.
[0042] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description
[0043] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein:
[0044] Figure 1 This is a schematic flowchart of a product identification method provided in an embodiment of this disclosure;
[0045] Figure 2 This disclosure provides a schematic diagram of the structure for obtaining the detection region by processing an image to be identified based on a target detection model;
[0046] Figure 3 This is a flowchart illustrating another product identification method provided in this embodiment of the disclosure;
[0047] Figure 4 This is a schematic diagram of the structure of a commodity identification device provided in an embodiment of this disclosure;
[0048] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure.
[0049] In the accompanying diagram, markers with the same last two digits correspond to the same elements. It should be noted that the elements in the diagram are schematic and not drawn to scale. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.
[0051] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, the statement of “connecting” or “coupling” two or more parts together shall mean that these parts are directly joined together or joined through one or more intermediate components.
[0052] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0053] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B exist simultaneously, or B exists. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0054] Furthermore, in all embodiments of this disclosure, terms such as “first” and “second” are used only to distinguish one component (or part of a component) from another component (or another part of a component).
[0055] In the description of this application, unless otherwise stated, "multiple" means two or more (including two), and similarly, "multiple groups" means two or more (including two groups).
[0056] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0057] Based on the problems existing in the current technology Figure 1 This is a flowchart illustrating a product identification method provided in an embodiment of this disclosure, such as... Figure 1 As shown, the specific process of the product identification method includes:
[0058] S110. Based on the target detection model, obtain the detection region corresponding to the detection label group from the image to be identified.
[0059] The image to be identified is a marketing image of the product to be detected, and the detection label group includes at least one label from the target product category to which the product to be detected belongs.
[0060] In the prior art, when a user manages goods in a digital asset management system, the marketing images of the goods to be detected are uploaded to the digital asset management system, which may lead to inaccurate recognition of the marketing images. Based on the problems existing in the prior art, the goods recognition method provided in this disclosure first obtains the detection area corresponding to the detection tag identification group from the image to be recognized based on the target detection model. The target detection model is a language-guided zero-shot detection model.
[0061] Zero-shot detection models are trained using training set data, enabling them to classify objects in the test set. However, there is no overlap between the categories in the training set and the categories in the test set. Therefore, it is necessary to use category descriptions (such as language guidance) to establish a connection between the training set and the test set, thus making the model effective.
[0062] In the specific implementation process, the zero-shot detection model guided by language first obtains the detection label corresponding to the product belonging to the same product category as the product to be detected. That is, the input is a white background image of the product belonging to the same product category as the product to be detected. The white background image of the product is an image obtained by processing the product image. There can be multiple white background images of a product, that is, white background images corresponding to product images from different perspectives. By inputting the white background images of the product into the zero-shot detection model, the label corresponding to the product is obtained.
[0063] After obtaining the labels corresponding to the products, users can select one or more labels from the labels to form a detection label group, which serves as a language guide for the zero-shot target detection model to recognize the image. In other words, the target detection model obtains the detection area corresponding to the detection label group from the image to be recognized by detecting the label group.
[0064] In specific implementations, the detection label identification group may include one detection label identification, two detection label identifications, or multiple detection label identifications. This disclosure does not specifically limit this. In specific implementation processes, users can select one or more label identifications from the label identifications corresponding to the product to form a detection label identification group to ensure the accuracy of the extracted detection area.
[0065] As a specific embodiment, when the detection label group includes two detection labels, namely detection label 1 and detection label 2, the target detection model first obtains the detection area corresponding to detection label 1 from the image to be identified, and then obtains the detection area corresponding to detection label 2 from the image to be identified.
[0066] For example, when the image to be identified input into the object detection model includes various skincare products such as toner, face cream, and facial cleanser, and the detection label group includes a label 'bottleORmakeup', then the detection area obtained by the object detection model is as follows: Figure 2 The area shown is enclosed in a solid line.
[0067] In the above embodiments, the zero-shot detection model can be trained by the user or an external or open-source pre-trained model such as GLIP or Grounding-DINO can be used.
[0068] As a specific implementation method, based on the object detection model, the detection region corresponding to the detection label identification group is obtained from the image to be identified, including: determining the location information of the detection region corresponding to the detection label identification group in the image to be identified based on the object detection model; and processing the image to be identified according to the location information of the detection region to obtain the detection region.
[0069] Specifically, the image to be identified is first input into the object detection model. The object detection model identifies the image to be identified and obtains the location information of the detection area corresponding to the detection label identifier included in the detection label identifier group (the location information includes the first location information, the second location information, the third location information, and the fourth location information). Then, based on the location information of the detection area, the detection area is obtained by cropping from the image to be identified.
[0070] It should be noted that, in the above embodiments, the location information of the detection area includes first location information, second location information, third location information, and fourth location information. Therefore, the cropped detection area is a quadrilateral. In other possible embodiments, the specific number of location information is not limited, and users can customize the settings according to the image to be identified.
[0071] S120. Determine the target white background image corresponding to the detection area.
[0072] Among them, the target white background image is any one of the white background images of products included in the target product category to which the product to be detected belongs. The white background image is an image obtained by processing the product image, and the same product includes white background images from different angles.
[0073] After step S110 is completed and the detection area corresponding to the detection label group is obtained from the image to be identified, the target white background image corresponding to the detection area is determined, and then the target label of the image to be identified is determined according to the association between the target white background image and the label.
[0074] As a specific implementation method, determining the target white background image corresponding to the detection area includes: extracting the target feature vector corresponding to the detection area based on the feature extraction model; comparing the similarity between the target feature vector and the feature vectors corresponding to each white background image in turn; and selecting the white background image corresponding to the feature vector with the highest similarity as the target white background image corresponding to the detection area.
[0075] Before executing step S120, a vector library is constructed, which includes feature vectors of white background images and preset labels corresponding to the white background images. Therefore, after the target feature vector corresponding to the detection area is extracted, the feature vectors of each white background image in the vector library are obtained, and the target feature vector corresponding to the detection area is matched with the feature vectors of the white background images in the vector library for similarity. The white background image corresponding to the feature vector with the highest similarity is selected as the target white background image of the detection area.
[0076] Specifically, the feature extraction model includes multiple cascaded convolutional layers and linear layers. The implementation method for comparing the similarity between the target feature vector and the feature vectors corresponding to each white background image can be: calculating the similarity between the target feature vector and the feature vectors corresponding to each white background image using the triplet loss function, or calculating the similarity between the target feature vector and the feature vectors corresponding to each white background image using the cross-entropy loss function.
[0077] S140. Based on the target white background image, determine the target label identifier of the image to be identified.
[0078] As a specific implementation method, based on step S120, after the target white background image is determined, the label corresponding to the target white background image is selected as the target label of the image to be identified through the correspondence between the target white background image and the label. The correspondence between the white background image and the label is pre-set, that is, a first association table between the white background image and the first preset label is constructed.
[0079] Before executing step S120, the following steps are performed: By constructing a vector library, the vector library includes feature vectors of white background images and preset label identifiers corresponding to white background images. That is, the vector library includes feature vectors of white background images and preset label identifiers corresponding to white background images. Therefore, based on the vector library, a first association table between white background images and first preset label identifiers can be determined. At this time, the first preset label identifier corresponding to the target white background image is selected from the first association table as the target label identifier of the image to be identified.
[0080] The product identification method provided in this disclosure first obtains a detection region corresponding to a detection tag group from the image to be identified based on an object detection model. The image to be identified is a marketing image of the product to be identified, and the detection tag group includes at least one tag from the target product category to which the product belongs. Then, a target white-background image corresponding to the detection region is determined. Finally, based on the target white-background image, the target tag of the image to be identified is determined. That is, only multi-angle white-background images and corresponding tag icons are required. The zero-shot object detection model enables automatic tagging and identification of marketing images, improving the efficiency and accuracy of product management and marketing.
[0081] Based on the above embodiments, Figure 3 This is a flowchart illustrating another product identification method provided in this disclosure. This disclosure is based on the above embodiments, such as... Figure 3 As shown, when performing step S120, the following is also included:
[0082] S130. Determine the target text content corresponding to the detection area.
[0083] The target text content is any one of the text contents of the product images included in the target product category to which the product to be detected belongs.
[0084] As a specific implementation method, determining the target text content corresponding to the detection area includes: extracting the text content corresponding to the detection area based on the OCR extraction model; comparing the similarity between the text content corresponding to the detection area and each preset text content in turn; and selecting the preset text content with the highest similarity as the target text content of the detection area.
[0085] In the specific implementation process, when the marketing images corresponding to different products have a high degree of similarity, the result of determining the target label of the image to be identified based solely on the target white background image corresponding to the detection area may be inaccurate. Therefore, when obtaining the target white background image corresponding to the detection area, it is also necessary to determine the target text content corresponding to the detection area.
[0086] Before executing step S120, a first association table between the white background image and the first preset label identifier is constructed.
[0087] Before executing step S130, a second association table between preset text content and second preset label identifier is constructed, wherein the preset text content is the text content of the product image of the product included in the target product category to which the product to be detected belongs.
[0088] When the product identification method includes step S130, the specific implementation of step S140 is as follows:
[0089] S141. Based on the target white background image and the target text content, determine the target label identifier of the image to be identified.
[0090] Based on the target white background image and the target text content, the target label identifier of the image to be identified is determined, including: selecting the first target preset label identifier corresponding to the target white background image from the first association table; selecting the second target preset label identifier corresponding to the target text content from the second association table; and determining the target label identifier of the image to be identified according to the relationship between the first target preset label identifier and the second target preset label identifier.
[0091] Specifically, a first association table is constructed between white background images and first preset labels, and a second association table is constructed between preset text content and second preset labels. Then, by sequentially comparing the similarity between the target feature vector corresponding to the detection area and the feature vector corresponding to each white background image, the white background image corresponding to the feature vector with the highest similarity is selected as the target white background image corresponding to the detection area. Based on the first association table between the white background image and the first preset label, the first target preset label is determined. Similarly, by sequentially comparing the similarity between the text content corresponding to the detection area and each preset text content, the preset text content with the highest similarity is selected as the target text content corresponding to the detection area. Based on the second association table between the preset text content and the second preset label, the second target preset label is determined.
[0092] When the first target preset label and the second target preset label are the same, the first target preset label or the second target preset label is determined as the target label of the image to be identified; when the first target preset label and the second target preset label are different, the second target preset label is determined as the target label of the image to be identified. That is, by performing text recognition on the image to be identified, the label determined based on the target white background image is verified, thereby improving the accuracy of the determined label of the image to be identified.
[0093] The product identification method provided in this disclosure first obtains a detection region corresponding to a detection tag group from the image to be identified based on an object detection model. The image to be identified is a marketing image of the product to be identified, and the detection tag group includes at least one tag from the target product category to which the product belongs. Then, a target white background image corresponding to the detection region and target text content corresponding to the detection region are determined. Finally, based on the target white background image and target text content, the target tag of the image to be identified is determined. That is, only multi-angle white background images and corresponding tag icons are required. By combining a zero-shot object detection model and an OCR recognition model, automatic tagging and recognition of marketing images can be achieved, improving the efficiency and accuracy of product management and marketing.
[0094] Based on the above embodiments, this disclosure also provides a product identification device, such as... Figure 4 As shown, the product identification device includes:
[0095] The detection area acquisition module 410 is used to acquire the detection area corresponding to the detection label identification group from the image to be identified based on the target detection model. The image to be identified is a marketing image of the product to be detected, and the detection label identification group includes at least one label identification in the target product category to which the product to be detected belongs.
[0096] The target white background image determination module 420 is used to determine the target white background image corresponding to the detection area. The target white background image is any one of the white background images of the products included in the target product category to which the product to be detected belongs. The white background image is an image obtained by processing the product image. The same product includes white background images from different angles.
[0097] The target label identification module 430 is used to determine the target label identification of the image to be identified based on the target white background image.
[0098] The product recognition device provided in this embodiment firstly, a detection area acquisition module, based on a target detection model, acquires a detection area corresponding to a detection tag identification group from the image to be recognized. The image to be recognized is a marketing image of the product to be detected, and the detection tag identification group includes at least one tag identification from the target product category to which the product to be detected belongs. Then, a target white background image determination module determines the target white background image corresponding to the detection area. Finally, a target tag identification module determines the target tag identification of the image to be recognized based on the target white background image. That is, only multi-angle white background images and corresponding tag identifications are needed; the zero-shot target detection model enables automatic tagging and recognition of marketing images, improving the efficiency and accuracy of product management and marketing.
[0099] In a specific implementation, the target white background image determination module includes: a feature vector extraction unit, a comparison unit, and a target white background image determination unit;
[0100] The feature vector extraction unit is used to extract the target feature vector corresponding to the detection region based on the feature extraction model.
[0101] The comparison unit is used to compare the similarity between the target feature vector and the feature vector corresponding to each white background image in turn. The white background image is a white background image of the product images of the products included in the target product category to which the product to be detected belongs.
[0102] The target white background image determination unit is used to select the white background image corresponding to the feature vector with the highest similarity as the target white background image corresponding to the detection region.
[0103] In a specific implementation, the product identification device further includes a first association table construction unit;
[0104] The first association table construction unit is used to construct a first association table between white background images and first preset label identifiers;
[0105] At this point, the target label identification module is implemented as follows:
[0106] Select the first preset label identifier corresponding to the target white background image from the first association table as the target label identifier of the image to be identified.
[0107] In a specific implementation, the product identification device also includes a target text content determination module;
[0108] The target text content determination module is used to determine the target text content corresponding to the detection area;
[0109] At this point, the specific implementation methods of the target label identification module include:
[0110] Based on the target white background image and the target text content, the target label identifier of the image to be identified is determined.
[0111] In a specific implementation, the product identification device further includes a second association table construction unit;
[0112] The second association table construction unit is used to construct a second association table between preset text content and second preset tag identifier, wherein the preset text content is the text content of the product image of the product to be detected, which is included in the target product category;
[0113] At this point, based on the target white background image and the target text content, the specific implementation methods for determining the target label identifier of the image to be identified include:
[0114] Select the first target preset label identifier corresponding to the target white background image from the first association table;
[0115] Select the second target preset label identifier corresponding to the target text content from the second association table;
[0116] The target label of the image to be identified is determined based on the relationship between the first target preset label and the second target preset label.
[0117] In a specific implementation, the target label of the image to be identified is determined based on the relationship between the first target preset label and the second target preset label, including:
[0118] When the first target preset label identifier and the second target preset label identifier are the same, the first target preset label identifier or the second target preset label identifier is determined as the target label identifier of the image to be identified;
[0119] When the first target preset label and the second target preset label are different, the second target preset label is determined as the target label of the image to be identified.
[0120] In a specific implementation, the detection area acquisition module includes a location information determination unit and a detection area determination unit;
[0121] The location information determination unit is used to determine the location information of the detection area corresponding to the detection label identification group in the image to be identified based on the target detection model;
[0122] The detection region determination unit is used to process the image to be recognized based on the location information of the detection region to obtain the detection region.
[0123] This application also provides a computer device. Please refer to the following for details. Figure 5 , Figure 5 This is a basic structural block diagram of the computer device in this embodiment.
[0124] The computer device includes a memory 510 and a processor 520 that are interconnected via a system bus. It should be noted that only a computer device with components 510-520 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0125] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0126] The memory 510 includes at least one type of readable storage medium, including non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. RAM may include static RAM or dynamic RAM. In some embodiments, the memory 510 may be an internal storage unit of a computer device, such as the hard disk or RAM of the computer device. In other embodiments, the memory 510 may also be an external storage device of the computer device, such as a plug-in hard drive, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard. Of course, the memory 510 may include both internal and external storage units of the computer device. In this embodiment, the memory 510 is typically used to store the operating system and various application software installed on the computer device, such as the program code of the methods described above. Furthermore, the memory 510 may also be used to temporarily store various types of data that have been output or will be output.
[0127] The processor 520 is typically used to perform the overall operation of a computer device. In this embodiment, the memory 510 is used to store program code or instructions, including computer operation instructions. The processor 520 is used to execute the program code or instructions stored in the memory 510 or to process data, such as program code that runs the methods described above.
[0128] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus system can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.
[0129] Another embodiment of this application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads computer-readable program code stored in the computer-readable medium, enabling the processor to execute the functional actions specified in each step or combination of steps in the above method; and to generate means for implementing the functional actions specified in each block or combination of blocks in the block diagram.
[0130] Computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any suitable combination thereof, wherein the memory is used to store program code or instructions, the program code including computer operation instructions, and the processor is used to execute the program code or instructions of the above-described methods stored in the memory.
[0131] The definitions of memory and processor can be found in the description of the foregoing computer device embodiments, and will not be repeated here.
[0132] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0133] In the various embodiments of this application, the functional units or modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0134] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0135] Unless otherwise expressly indicated by the context, the singular form of words used herein and in the appended claims includes the plural form, and vice versa. Thus, when referring to the singular, the plural form of the corresponding term is generally included. Similarly, the terms “comprising” and “including” shall be interpreted as including rather than exclusively. Likewise, the terms “including” and “or” shall be interpreted as including unless such interpretation is expressly prohibited herein. Where the term “example” is used herein, particularly when it follows a set of terms, the “example” is merely exemplary and illustrative and should not be considered exclusive or extensive.
[0136] Further aspects and scope of adaptation become apparent from the description provided herein. It should be understood that various aspects of this application may be implemented individually or in combination with one or more other aspects. It should also be understood that the descriptions and specific embodiments herein are for illustrative purposes only and are not intended to limit the scope of this application.
[0137] Several embodiments of this disclosure have been described in detail above. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of this disclosure without departing from the spirit and scope of this disclosure. The scope of protection of this disclosure is defined by the appended claims.
Claims
1. A product identification method, characterized in that, include: Based on the zero-shot object detection model, the detection region corresponding to the detection tag identification group is obtained from the image to be identified, wherein the image to be identified is a marketing image of the product to be detected, and the detection tag identification group includes at least one tag identification in the target product category to which the product to be detected belongs; Determine the target white background image corresponding to the detection area, wherein the target white background image is any one of the white background images of the products included in the target product category to which the product to be detected belongs, and the white background image is an image obtained by processing the product image, and the same product includes white background images from different angles; Based on the target white background image, determine the target label identifier of the image to be identified; The step of obtaining the detection region corresponding to the detection label group from the image to be identified based on the zero-shot object detection model includes: Based on the zero-shot target detection model, obtain the detection label identifier corresponding to the product that belongs to the same product category as the product to be detected; Select one or more labels from the detection labels corresponding to products belonging to the same product category as the product to be tested to form a detection label group; The detection label group is composed of one or more selected labels and serves as a language guide for the zero-shot visual detection model to recognize the image to be identified, and the detection region corresponding to the detection label group is obtained from the image to be identified.
2. The method according to claim 1, characterized in that, Determining the target white background image corresponding to the detection area includes: Based on the feature extraction model, the target feature vector corresponding to the detection region is extracted; The similarity between the target feature vector and the feature vector corresponding to each white background image is compared sequentially. The white background image is a white background image of a product image of a product in the target product category to which the product to be detected belongs. The white background image corresponding to the feature vector with the highest similarity is selected as the target white background image corresponding to the detection region.
3. The method according to claim 1, characterized in that, Before determining the target label identifier of the image to be identified based on the target white background image, the method further includes: Construct a first association table between white background images and first preset label identifiers; The step of determining the target label identifier of the image to be identified based on the target white background image includes: Select the first preset label identifier corresponding to the target white background image from the first association table as the target label identifier of the image to be identified.
4. The method according to claim 1, characterized in that, When determining the target white background image corresponding to the detection area, the method further includes: Determine the target text content corresponding to the detection area; The step of determining the target label identifier of the image to be identified based on the target white background image includes: Based on the target white background image and the target text content, the target label identifier of the image to be identified is determined.
5. The method according to claim 4, characterized in that, Before determining the target label identifier of the image to be identified based on the target white background image, the method further includes: Construct a first association table between white background images and first preset label identifiers; Construct a second association table between preset text content and second preset tag identifier, wherein the preset text content is the text content of the product image of the product to be detected, which is included in the target product category; The step of determining the target label identifier of the image to be identified based on the target white background image and the target text content includes: Select the first target preset label identifier corresponding to the target white background image from the first association table; Select the second target preset tag identifier corresponding to the target text content from the second association table; The target label of the image to be identified is determined based on the relationship between the first target preset label and the second target preset label.
6. The method according to claim 5, characterized in that, The step of determining the target label identifier of the image to be identified based on the relationship between the first target preset label identifier and the second target preset label identifier includes: When the first target preset label identifier and the second target preset label identifier are the same, the first target preset label identifier or the second target preset label identifier is determined as the target label identifier of the image to be identified; When the first target preset label identifier and the second target preset label identifier are different, the second target preset label identifier is determined to be the target label identifier of the image to be identified.
7. The method according to claim 4, characterized in that, The zero-shot target detection model obtains the detection region corresponding to the detection label group from the image to be identified, including: Based on the zero-shot target detection model, the location information of the detection area corresponding to the detection label identification group in the image to be identified is determined; The detection area is obtained by processing the image to be identified based on the location information of the detection area.
8. A product identification device, characterized in that, include: The detection region acquisition module is used to acquire the detection region corresponding to the detection label identification group from the image to be identified based on the zero-shot target detection model. The image to be identified is a marketing image of the product to be detected, and the detection label identification group includes at least one label identification in the target product category to which the product to be detected belongs. The target white background image determination module is used to determine the target white background image corresponding to the detection area. The target white background image is any one of the white background images of the products included in the target product category to which the product to be detected belongs. The white background image is an image obtained by processing the product image. The same product includes white background images from different angles. The target label identification module is used to determine the target label identification of the image to be identified based on the target white background image; The step of obtaining the detection region corresponding to the detection label group from the image to be identified based on the zero-shot object detection model includes: Based on the zero-shot target detection model, obtain the detection label identifier corresponding to the product that belongs to the same product category as the product to be detected; Select one or more labels from the detection labels corresponding to products belonging to the same product category as the product to be tested to form a detection label group; The detection label group is composed of one or more selected labels and serves as a language guide for the zero-shot visual detection model to recognize the image to be identified, and the detection region corresponding to the detection label group is obtained from the image to be identified.
9. A computer device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method and equipment for automatically identifying commodities in image set
CN113065447A
Target commodity big data accurate identification method and system based on simple photographing
CN115661833A