Data Evaluation Method, System and Computer Readable Storage Medium
By using the improved Mixmatch model and target classifier in the commodity data acquisition center, combining a small amount of labeled data and unlabeled data, the problems of long data acquisition cycle and high cost in the prior art are solved, and efficient and accurate data evaluation and acquisition strategy determination are achieved.
Patent Information
- Application Number
- CN202210710095.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-06-22
AI Technical Summary
The prior art relies on manual annotation in commodity data collection, resulting in long cycles, high work costs and lack of targeting, which cannot effectively solve the problem of data imbalance.
By selecting a small number of images from each category of products to be learned as an annotation set based on the product database, training the improved Mixmatch model with the unlabeled set, generating a target classifier, and calculating the score of the products to be learned based on the output confidence, and determining the data acquisition strategy.
It realizes that without a large amount of manual labeling, quickly locate product categories with insufficient data, save human resources, improve data acquisition efficiency, and shorten the AI modeling cycle.
Smart Images

Figure CN115035344B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer software and fast-moving consumer goods technology, and particularly relates to a data evaluation method, system, and computer-readable storage medium. Background Art
[0002] With the continuous expansion of artificial intelligence in the fast-moving consumer goods field, the large-scale application of object detection in commodity recognition has gradually become a trend, which greatly improves the work efficiency of salespersons' store inspections and helps enterprises quickly understand the details of product distribution and sales in terminal stores. However, establishing an accurate commodity detection model requires a sufficient amount of real-scene data to train the model, which requires data collectors to go to various shopping malls and stores to collect commodity picture data before model building. For better learning effects, it is generally required that various commodity category data be sufficient and relatively balanced. However, there are a large number of commodity types, and enterprises will also make preferential selections according to the sales popularity of commodities when distributing goods, resulting in extremely unbalanced collected data or duplicate collection. The common current practice is to first collect a batch of data, then manually screen and label it, and then count the quantity of each commodity SKU, and re-collect data for the insufficient SKUs. And this practice often has several drawbacks:
[0003] 1) Manually screening pictures wastes a lot of manpower and has a high working cost;
[0004] 2) The labeled samples are preference behaviors and lack pertinence. For example, pictures of best-selling commodities are marked in large quantities, but data on scarce products are still insufficient, thus wasting labeling resources;
[0005] 3) The efficiency of manual labeling is extremely low. The store inspectors cannot get timely feedback and sometimes have to conduct secondary collection, seriously affecting the normal progress of work.
[0006] In summary, there is an urgent need for a new data evaluation method to solve the above problems. Summary of the Invention
[0007] The purpose of this application is to provide a data evaluation method, system, and computer-readable storage medium to solve the problems of long cycle, high working cost, and lack of pertinence caused by the existing commodity data relying on manual labeling.
[0008] To achieve the above purpose, this application provides a data evaluation method, including:
[0009] Based on a commodity database, select n images from each type of commodity to be learned as a labeled set; where n << m, and m is the number of modeling requirements;
[0010] Perform commodity detection and segmentation on the scene images containing the commodities, and use the segmented sub-images as an unlabeled set;
[0011] Train an improved Mixmatch model using the labeled set and the unlabeled set until the model converges to generate a target classifier;
[0012] Input the unlabeled set into the target classifier and calculate the score of the product to be learned based on the output confidence;
[0013] Determine a data collection strategy based on the relationship between the score of the product to be learned and the required quantity for modeling.
[0014] Further, calculating the score of the product to be learned based on the output confidence includes:
[0015] Determine the two categories with the highest confidence in the output confidence, and denote the confidences as P A and P B respectively. Denote the first preset value and the second preset value as T 1 and T 2 respectively. Assume P A >P B ;
[0016] If P A >T 1 , then the score of the product category corresponding to P A is 1, and the scores of other categories are set to 0;
[0017] If P A ≤T 1 , and P A is a negative sample category, then the scores of all product categories are 0;
[0018] If P A +P B ≤T 2 , then the scores of all product categories are 0;
[0019] If P A +P B >T 2 , then the scores of the product categories corresponding to P A and P B are respectively:
[0020]
[0021]
[0022] where O A and O B are the normalized score outputs of the categories where P A and P B are located respectively; R A and R BSharpening intermediate results of confidence scores respectively; t is the temperature coefficient.
[0023] Further, determining the data acquisition strategy according to the relationship between the score situation of the commodity to be learned and the number of modeling requirements includes:
[0024] Calculate the total score of each commodity category. If the total score of a certain commodity category is less than 90% of the number of modeling requirements, mark that such commodities need to supplement commodity images;
[0025] If the total score of a certain commodity category is higher than 2 times the number of modeling requirements, mark that the number of images of this category of commodities meets the standard.
[0026] Further, determining the data acquisition strategy according to the relationship between the score situation of the commodity to be learned and the number of modeling requirements further includes:
[0027] Judge whether the balance degree of the commodity category meets the preset requirements according to the score situation of the commodity category;
[0028] If the balance degree of the commodity category does not meet the preset requirements, screen the scene images, including:
[0029] Determine the score set D of the sub-images segmented from the scene image:
[0030] D = {d 1 , d 2 ... d m};
[0031] Wherein,
[0032]
[0033] In the formula, d i is the i-th scene image, represents the score of the j-th sub-image segmented from the i-th scene image without including the negative sample category;
[0034] Calculate the variance of D, denoted as var, and set the third preset value as T 3 ;
[0035] Loop through each item d of D i , loop m times; after removing the current d i item, obtain the new sets D 1 , D 2 ,... D m ; and calculate the variances of D 1 , D 2 ,... D m respectively, denoted as var 1 , var 2 … varm ;
[0036] Use var to calculate the differences with var 1 , var 2 … var m respectively, obtaining m variance differences, selecting the largest variance difference among them, and denoting it as var index ;
[0037] When var index ≥ T 3 , determine the corresponding d index items and remove them, add the index-th image to the set S, and then return to execute the step of calculating the variance of D; where the set S is the scene images that do not need to enter the manual annotation process;
[0038] When var index < T 3 , the screening process ends.
[0039] Furthermore, before training the improved Mixmatch model using the labeled set and the unlabeled set, it further includes:
[0040] Obtaining the improved Mixmatch model, including:
[0041] Rotate the image after MixUP augmentation of the input image by 180 degrees;
[0042] Use the Object classifier to perform image classification and calculate the image classification loss;
[0043] Use the Rotation predictor to perform rotation angle classification prediction and calculate the rotation angle classification loss;
[0044] Determine the loss function of the improved Mixmatch model according to the image classification loss and the rotation angle classification loss.
[0045] Furthermore, the obtaining of the improved Mixmatch model further includes:
[0046] Executing the pseudo-label negative sample augmentation strategy, including:
[0047] Input the unlabeled set into the target classifier for prediction to obtain the confidence values of the positive and negative samples;
[0048] For the positive samples with confidence less than the fourth preset value, increase the confidence value of the negative samples and decrease the confidence value of the positive samples:
[0049]
[0050] In the formula, P is the model prediction value, Ppos Represents the predicted value of the positive sample model, P neg Represents the predicted value of the negative sample model, and T is the confidence threshold.
[0051] Furthermore, the implementation of the pseudo-label negative sample enhancement strategy further includes:
[0052] Determine the confidence distribution fluctuation difference according to the KL divergence. For samples with a KL divergence value greater than the divergence threshold, increase the confidence value of the negative samples and decrease the confidence value of the positive samples:
[0053]
[0054] In the formula, T kl is the divergence threshold, and P i is the prediction result of the model for the i-th randomly enhanced data.
[0055] Furthermore, the commodity detection and segmentation of the scene image containing the commodity includes:
[0056] Use a detector to perform commodity detection and segmentation on the scene image containing the commodity, and the detector uses the VFNet algorithm.
[0057] This application also provides a data evaluation system, including:
[0058] An annotation set generation unit, which is used to select n images from each type of commodity to be learned based on the commodity database as the annotation set; where n << m, and m is the number of modeling requirements;
[0059] An unannotated set generation unit, which is used to perform commodity detection and segmentation on the scene image containing the commodity, and use the segmented sub-images as the unannotated set;
[0060] A target classifier generation unit, which is used to train an improved Mixmatch model using the annotation set and the unannotated set until the model converges to generate a target classifier;
[0061] A classification score calculation unit, which is used to input the unannotated set into the target classifier and calculate the score of the commodity to be learned according to the output confidence;
[0062] An acquisition strategy determination unit, which is used to determine the data acquisition strategy according to the relationship between the score of the commodity to be learned and the number of modeling requirements.
[0063] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the data evaluation method described in any one of the above.
[0064] Compared with the prior art, the beneficial effects of this application are:
[0065] 1) Only a small amount of labeled data is required for modeling, without the need to label a large amount of other commodity data, saving human labeling resources;
[0066] 2) High accuracy, effectively distinguishing the data of this product from the data of other commodities;
[0067] 3) Strong data evaluation ability, which can effectively locate the commodity categories with insufficient modeling data and quickly feedback to the collector for supplementation;
[0068] 4) Fast response, timely feedback to guide the collection work of the store inspector, improving work efficiency and shortening the AI modeling cycle. Brief Description of the Drawings
[0069] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the implementation will be briefly introduced below. Obviously, the drawings in the following description are only some implementations of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0070] Figure 1 is a schematic flowchart of a data evaluation method provided by an embodiment of the present application;
[0071] Figure 2 is a schematic flowchart of the process of training the VFNet detector provided by an embodiment of the present application;
[0072] Figure 3 is a schematic structural diagram of the Mixmatch model provided by an embodiment of the present application;
[0073] Figure 4 is a schematic structural diagram of the improved Mixmatch model provided by an embodiment of the present application;
[0074] Figure 5 (a) is an image of the collected Coke sample provided by an embodiment of the present application;
[0075] Figure 5 (b) is a sub-image of the cut Coke sample provided by an embodiment of the present application;
[0076] Figure 6 is a schematic structural diagram of a data evaluation system provided by an embodiment of the present application. Detailed Embodiments
[0077] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.
[0078] It should be understood that the step numbers used in the text are only for convenient description and do not limit the order of execution of the steps.
[0079] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0080] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0081] The term " / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0082] To help understanding, first, the relevant terms involved in the embodiments of the present application are explained:
[0083] VFNet algorithm: A new dense object detector built on the FCOS architecture, abbreviated as VarifocalNet or VFNet.
[0084] Semi-supervised learning: A learning method that combines supervised learning and unsupervised learning, which can make good use of unlabeled data, thereby reducing the dependence on large labeled data sets.
[0085] Mixmatch model: A model built based on the MixMatch algorithm, where the MixMatch algorithm is a semi-supervised learning algorithm that can estimate low-entropy labels for unlabeled samples obtained by data augmentation and use MixUp to mix labeled and unlabeled data.
[0086] SKU: Stock Keeping Unit, that is, the basic unit for measuring inventory in and out, and now has been extended to the abbreviation of the unified product number. Each product corresponds to a unique SKU number.
[0087] At present, in the fast-moving consumer goods field, target detection technology is increasingly widely used. In order to help enterprises quickly understand the details of product placement and sales in terminal stores, it is usually necessary to establish an accurate product detection model. The prerequisite for building a product detection model is to provide sufficient and large amounts of data in real scenarios for model training, and data collection has always been a major problem in actual applications. Usually, data collectors will first collect a batch of data from each terminal store, then manually screen and label it, and then count the quantity of each SKU. For insufficient SKUs, data collection will be carried out again. On the one hand, both the labeling process and the counting process will inevitably consume a large amount of time, manpower, and material resources, and are at risk of making mistakes at any time. And once the cycle of feedback data from the collection link to the store inspectors is too long, it will lead to the store inspectors having to conduct secondary data collection in order to complete their work, which will also cause duplicate collection and waste of resources. Therefore, the embodiments of the present invention aim to propose a data evaluation method, which is modeled based on semi-supervised learning, realizes learning and evaluation of a large amount of collected data, and then quickly feeds back to guide data collectors to timely supplement the SKUs with insufficient modeling quantities, reduces manual image selection and repeated labeling work, thereby shortening the data collection cycle and improving the AI modeling efficiency.
[0088] Please refer to Figure 1 , in an embodiment of the present application, a data evaluation method is provided. As Figure 1 shown, the data evaluation method includes steps S10 to S50. The specific steps are as follows:
[0089] S10. Based on the product database, select n images from each type of product to be learned as the labeled set; where n << m, and m is the modeling requirement quantity.
[0090] In this step, first determine the categories of the products to be learned, and then select n images from the product database for each category of products to be learned as the labeled set. Where n << m, and m is the modeling requirement quantity. Since the method of the present application only needs to learn a small amount of labeled sets, the number n of the labeled set images is much smaller than the modeling requirement quantity m.
[0091] It should be noted that in a specific implementation manner, since the labeled set only has a small number of product images, it is necessary to ensure that the sample comprehensiveness of the collected product images is relatively good. For example, a certain product has multiple different faces. For example, a packaged tissue paper is usually a cuboid and has 6 faces. For such a product, when obtaining data, images of 6 faces need to be collected as the corresponding labeled set. Usually, the collected data set is denoted as (X, Y), where X is the set of product images to be learned, and Y is its corresponding category code.
[0092] As a preferred embodiment, when constructing the annotation set, for some competing products that are relatively similar to the product, a small amount of competing product data also needs to be supplemented as negative samples. Among them, when the overall appearance of the products is very similar, there are only differences in some areas, and the area of the different areas is small and not easily distinguishable by the human eye, the products can be considered similar. At this time, if one of them is used as the product, the other is used as a competing product. It should be noted that the negative samples here refer to the data that are not in the modeling learning list but require the model to accurately distinguish as competing products.
[0093] S20. Perform product detection and segmentation on the scene image containing the product, and use the segmented sub-images as the unlabeled set.
[0094] In this step, in order to construct the unlabeled set, it is first necessary to perform product detection and segmentation on the scene image containing the above product. Among them, the scene image can be a batch of images collected by the acquisition personnel in real application scenarios such as shelves, end caps, or refrigerators with many products. The specific selection of the scene can be adjusted according to actual needs. It can be understood that these scene images often contain not only product images but also other irrelevant background images. Therefore, in order to reduce the interference of irrelevant background and other information and reduce the memory of the image, after obtaining the scene image, this step needs to perform product detection and segmentation on these scene images, and then directly use the segmented sub-images as the unlabeled set. Usually, the number of unlabeled images in this step will be much larger than the number of images in the labeled set in step S10. That is, in this step, it is required to collect as many scene images as possible, segment as many sub-images as possible to generate the unlabeled set and then perform modeling to ensure the learning effect and efficiency.
[0095] In a preferred embodiment, the VFNet detector is used to perform product detection and segmentation on the scene image containing the product. It should be noted that the VFNet detector is a general object detector. In order to obtain an object detector with better recognition effect, a large number of real pictures obtained in fast-moving consumer goods scenarios such as shelves, end caps, and refrigerators in each terminal store are usually collected. Please refer to Figure 2 , Figure 2 provides a process of training an object detector using shelf pictures. As shown in Figure 2 , by these labeled shelf pictures, input them into the product detection network, that is, train the original VFNet network until the model converges to generate an object detector, and this object detector can segment all product images in the scene image. Due to the complexity of the real scene, such as the differences in product packaging, oblique shooting, and lighting, etc., it is required that the training set cover various industries such as beverages, foods, and daily chemicals, and contain pictures under various imaging conditions. When training this VFNet detector, it is preferably trained with tens of thousands of picture materials, and the trained VFNet detector is used for full product detection and segmentation to achieve an ideal segmentation effect.
[0096] S30. Train an improved Mixmatch model using the labeled set and the unlabeled set until the model converges to generate a target classifier.
[0097] In a specific embodiment, before performing step S30, to ensure that the final data evaluation performance of the established model is relatively good, it is usually necessary to process the commodity samples including the labeled set and the unlabeled set to eliminate some data that may cause interference.
[0098] Specifically, this embodiment eliminates some commodity images from the labeled set and the unlabeled set, including the following content:
[0099] When two commodities have different specifications, retain one of the specifications; where
[0100] The method for determining whether two commodities belong to different specifications is: under the same imaging conditions, if the pixel difference area between two commodities is less than 5% of the full image area, the two commodities are determined to have different specifications.
[0101] By eliminating some data in this embodiment, the interference of the data can be greatly reduced, and the final data evaluation performance can be improved.
[0102] Furthermore, after processing the labeled set and the unlabeled set, use the processed labeled set and unlabeled set to train an improved Mixmatch model until the model converges to generate a target classifier.
[0103] It should be noted that the key to the solution of this application lies in the classification performance of the few-shot model. If the classification result is poor, the final statistical result will have a serious deviation. Assuming that the existing Mixmatch model is directly used, two problems usually occur when using semi-supervised learning technology to solve commodity data evaluation:
[0104] First, there are many types of negative sample commodities, but the labeled negative sample data is only a very small part of them. There are still many negative samples in the unlabeled set, and the correlation between negative samples is relatively low. For example, the outer packaging forms vary greatly.
[0105] Second, in the few-shot learning process, the model is prone to overfitting and lacks the generalization ability for complex scenarios.
[0106] Therefore, to solve these two problems, two measures will be taken in the following embodiments. One is to add a self-supervised learning module to the existing Mixmatch model to further improve the feature extraction ability of the few-shot model. The other is to execute a negative sample enhancement strategy to accurately and reliably identify negative samples, and be able to identify thousands of other commodities as negative sample commodity classes to ensure that the target classifier also has strong generalization performance after few-shot training.
[0107] In a specific embodiment, obtaining an improved Mixmatch model includes the following steps:
[0108] 3.1) Rotate the image after MixUP enhancement of the input image by 180 degrees;
[0109] 3.2) Use the Object classifier to perform image classification and calculate the image classification loss;
[0110] 3.3) Use the Rotation predictor to perform rotation angle classification prediction and calculate the rotation angle classification loss;
[0111] 3.4) Determine the loss function of the improved Mixmatch model according to the image classification loss and the rotation angle classification loss.
[0112] For the sake of helping understanding, the relevant principles of the Mixmatch model will be described first. Please refer to Figure 3 , Figure 3 provides the network structure of the Mixmatch model. When performing few-shot modeling, Mixmatch adopts the fakelabel strategy to process unlabeled samples, averages the results of two random data augmentation branches, reduces the noise of fakelabel, and then sharpens the output labels, making the recognition results of unlabeled samples more inclined to a specific result. Mixup data augmentation can also better process data with certain noise and alleviate the problem of model overfitting, which is also particularly important in few-shot learning. However, although the above Mixmatch model can achieve good results in some few-shot learning scenarios, it is for a small amount of labeled data that can represent the general form of the current category. For few-shot learning in the evaluation application of commodity data, there are still many unlabeled negative samples. For this situation, the above Mixmatch model is no longer applicable. Therefore, the Mixmatch model is improved in this embodiment.
[0113] Specifically, in this embodiment, a self-supervised learning module is added to the Mixmatch model to further improve the feature extraction ability of the few-shot model. It includes rotating the input picture and predicting the rotation angle of the picture through an angle prediction head. As Figure 4As shown in the figure, we rotate the image I0 after MixUP augmentation of the input data by 180 degrees, denoted as I180. Then the data input to the Backbone model is (I0, I180), and the corresponding rotation categories are (0 degrees, 180 degrees). After feature extraction by the Backbone, it is input to the Object classifier for product category classification to calculate the classification loss, and then input to the Rotation predictor for rotation angle classification prediction to calculate the classification loss of the rotation angle. Finally, the loss function of the improved Mixmatch model is: the sum of the image classification loss and the rotation angle classification loss.
[0114] In a specific embodiment, obtaining the improved Mixmatch model further includes:
[0115] Implementing the pseudo-label negative sample augmentation strategy, including:
[0116] Inputting the unlabeled set into the target classifier for prediction to obtain the confidence values of the positive and negative samples;
[0117] For positive samples with a confidence level less than the fourth preset value, increase the confidence value of the negative samples and decrease the confidence value of the positive samples:
[0118]
[0119] In the formula, P is the model prediction value, P pos represents the positive sample model prediction value, P neg represents the negative sample model prediction value, and T is the confidence threshold.
[0120] Furthermore, the implementation of the pseudo-label negative sample augmentation strategy further includes:
[0121] Determining the confidence distribution fluctuation difference according to the KL divergence. For samples with a KL divergence value greater than the divergence threshold, increase the confidence value of the negative samples and decrease the confidence value of the positive samples:
[0122]
[0123] In the formula, T kl is the divergence threshold, and P i is the prediction result of the model for the i-th randomly augmented data.
[0124] This step mainly implements the pseudo-label negative sample enhancement strategy. It can be understood that since thousands of unlabeled negative samples are predicted by the classification model, their classification results (pseudo-labels) may be misclassified as positive sample classes, and the pseudo-label noise is continuously error-propagated with the iterative training of the Mixmatch network, affecting the final model classification effect. Therefore, this embodiment aims to strengthen the pseudo-labels and adjust the model output of possible negative samples. Specifically, it includes:
[0125] 1) For samples that are discriminated as positive samples but have a low confidence level (confidence level less than the third preset value), increase the confidence value of the negative samples and decrease the confidence value of the positive samples, specifically calculated by formula (1).
[0126] 2) Considering that the unlabeled dataset follows a long-tail distribution, that is, in the unlabeled dataset, the proportion of the products to be learned is large, and the number of a large number of other product classes is small. Due to the small number of other products, the model cannot fully learn these products, so for the negative samples obtained by multiple random data augmentations, the confidence levels of various categories in their model outputs fluctuate greatly. We use the KL divergence to characterize the difference in confidence distribution fluctuations, and for samples with a KL divergence value greater than the divergence threshold, increase the confidence value of the negative samples and decrease the confidence value of the positive samples, specifically calculated by formula (2).
[0127] In this embodiment, through the above pseudo-label negative sample enhancement strategy, the negative classes misclassified as positive samples can be gradually pulled closer to the negative sample space, removing the noise in the positive sample space, thereby improving the final classification effect of the target classifier.
[0128] S40: Input the unlabeled set into the target classifier, and calculate the score of the product to be learned according to the output confidence level.
[0129] After step S30 is executed, the target classifier can be obtained. In this step, mainly input the unlabeled set into the target classifier, and then calculate the score of the product to be learned based on the output classification result according to the confidence level.
[0130] In a specific implementation manner, the calculating the score of the product to be learned according to the output confidence level includes:
[0131] Determine the two categories with the highest confidence levels in the output confidence levels, and denote the confidence levels as P A 、P B , and denote the first preset value and the second preset value as T 1 、T 2 , assuming P A >P B ;
[0132] If P A >T 1 , then PA The corresponding product category score is 1, and the scores of other categories are set to 0;
[0133] If P A ≤T 1 , and P A is a negative sample class, then the scores of all product categories are 0;
[0134] If P A +P B ≤T 2 , then the scores of all product categories are 0;
[0135] If P A +P B >T 2 , then the scores of the product categories corresponding to P A and P B are respectively:
[0136]
[0137]
[0138] In the formula, O A and O B are respectively the normalized score outputs of the categories where P A and P B are located; R A and R B are respectively the sharpened intermediate results of the confidence scores; t is the temperature coefficient.
[0139] S50. Determine the data acquisition strategy according to the relationship between the score situation of the product to be learned and the number of modeling requirements.
[0140] Specifically, this step includes the following content:
[0141] 5.1) Calculate the total score of each product category. If the total score of a certain product category is less than 90% of the number of modeling requirements, mark that this category of product needs to supplement product images;
[0142] 5.2) If the total score of a certain product category is higher than 2 times the number of modeling requirements, mark that the number of images of this category of product meets the standard.
[0143] Through the above steps, it is possible to effectively locate the product categories with insufficient modeling data and quickly feedback to the collector for supplementation.
[0144] In another preferred embodiment, after determining the data acquisition strategy, although the various types of commodity data to be learned meet the basic modeling requirement quantity, there may still be a relatively serious class imbalance problem. To balance the sample classes, this embodiment also needs to screen the collected image data to reduce duplicate labeling.
[0145] Specifically, determining the data acquisition strategy according to the relationship between the score of the commodity to be learned and the modeling requirement quantity further includes the following:
[0146] 5.3) According to the score situation of the commodity category, determine whether the balance degree of the commodity category meets the preset requirements;
[0147] If the balance degree of the commodity category does not meet the preset requirements, screen the scene image, including:
[0148] 5.4) Determine the score set D of the sub-images segmented from the scene image:
[0149] D = {d 1 , d 2 ... d m} (5)
[0150] Wherein,
[0151]
[0152] In the formula, d i is the i-th scene image, represents the score of the j-th sub-image segmented from the i-th scene image excluding the negative sample class;
[0153] 5.5) Calculate the variance of D, denoted as var, and set the third preset value as T 3 ;
[0154] 5.6) Loop through each item d i of D for m times; after removing the current d i item, obtain the new sets D 1 , D 2 ,... D m ; and calculate the variances of D 1 , D 2 ,... D m respectively, denoted as var 1 , var 2 … var m ;
[0155] 5.7) Use var to compare with var 1 , var 2 … var mTake the difference to obtain m variance differences, select the largest variance difference among them, and denote it as var index ;
[0156] 5.8) When var index ≥T 3 , determine the corresponding d index item and eliminate it, add the index-th image to the set S, and then return to execute the step of calculating the variance of D; where the set S is the scene images that do not need to enter the manual annotation process;
[0157] 5.9) When var index <T 3 , the screening process ends.
[0158] In a preferred embodiment, the data evaluation method provided by the present application can also be extended to the evaluation method of new product coverage. The following takes the data evaluation of the store shelf scene as an example for elaboration:
[0159] First, N pictures of the store shelf can be obtained by shooting. All the sub-images segmented from each shelf picture are used as a set, and the output n cls of each category in this set is counted. If the category is greater than or equal to 1, the output of the current shelf picture is set to 1; otherwise, the score result of this category is output, where the output is:
[0160] N cls =∑min(n cls , 1) (7)
[0161] In the formula, n cls is the evaluation result of each category.
[0162] Furthermore, the calculation method of new product coverage is the ratio r of the statistical result of a certain category of all shelves to the total number of shelf images.
[0163]
[0164] Among them, N is the number of tested images.
[0165] To help understanding, in an embodiment, taking cola as the commodity to be classified for illustration:
[0166] 1) Obtain a trained general commodity detection network, preferably the VFNet detector.
[0167] 2) Assuming that you want to classify 1500ML Pepsi, you can select 50 1500ML Pepsi images as the annotation set, denoted as (X, Y). Among them, X is the set of 50 images, and Y is the corresponding category code; in addition, the 50 images are required to include images of all sides of Pepsi as much as possible. Since Coca-Cola and Pepsi are similar, several Coca-Cola images are taken as negative sample classes in this step.
[0168] 3) Collectors go to offline supermarkets to take photos and collect a batch of image data, such as Figure 5 As shown in (a), the VFNet detector in 1) is then used to cut and obtain a set of sub-images, which is denoted as the unlabeled set U, as shown in Figure 5 (b) as shown.
[0169] 4) Input the labeled set (X, Y) and the unlabeled set U in 2) into the improved Mixmatch model for semi-supervised training until the model converges to obtain the target classifier.
[0170] 5) Use the target classifier to predict each small image in U, and calculate the category score according to the confidence of the output result; the process of calculating the score is shown in step S40.
[0171] 6) Count the scores of all small pictures of U. If the modeling requirement for 1500ML Pepsi is 500, and the score is less than 500*0.9, that is, 450 pictures, it will prompt that the data volume is insufficient and feedback collectors are asked to supplement the data; if the score is higher than 1000 points, the data volume meets the requirement and no further supplement is needed.
[0172] In summary, the data evaluation method provided in this application can at least achieve the following effects:
[0173] (1) Modeling only requires a small amount of labeled data, and there is no need to label a large amount of other product data, saving human labeling resources;
[0174] (2) High accuracy, effectively distinguishing data about this product from data about other products;
[0175] (3) Strong data evaluation capabilities, which can effectively locate product categories with insufficient modeling data and quickly provide feedback to collectors for supplementation;
[0176] (4) Respond quickly and provide timely feedback to guide store inspectors in their data collection work, thereby improving work efficiency and shortening the AI modeling cycle.
[0177] See also Figure 6 , a certain embodiment of the present application further provides a data evaluation system, comprising:
[0178] The annotation set generation unit 01 is configured to select n images from each type of commodity to be learned based on the commodity database as the annotation set; where n << m, and m is the number of modeling requirements.
[0179] The unlabeled set generation unit 02 is configured to perform commodity detection and segmentation on the scene images containing the commodities, and use the segmented sub-images as the unlabeled set.
[0180] The target classifier generation unit 03 is configured to train an improved Mixmatch model using the annotation set and the unlabeled set until the model converges, and generate a target classifier.
[0181] The classification score calculation unit 04 is configured to input the unlabeled set into the target classifier, and calculate the score of the commodity to be learned according to the output confidence.
[0182] The acquisition strategy determination unit 05 is configured to determine a data acquisition strategy according to the relationship between the score of the commodity to be learned and the number of modeling requirements.
[0183] It can be understood that the data evaluation system provided in this embodiment is used to execute the data evaluation method described in any of the above embodiments, and achieve the same effect, which will not be further elaborated here.
[0184] In another exemplary embodiment, a computer-readable storage medium including a computer program is further provided. When the computer program is executed by a processor, the steps of the data evaluation method described in any of the above embodiments are implemented. For example, the computer-readable storage medium may be the above-mentioned memory including a computer program, and the above computer program may be executed by the processor of the terminal device to complete the data evaluation method described in any of the above embodiments, and achieve the same technical effect as the above method.
[0185] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods when implementing it in actual applications. For example, multiple units or page components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0186] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0188] The above-mentioned integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units stored in a storage medium include several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present application.
Claims
1. A data evaluation method, characterized in that, it includes: Based on the commodity database, select n images from each type of commodity to be learned as the annotation set; where n << m, and m is the number of modeling requirements; Perform commodity detection and segmentation on the scene images containing the commodities, and use the segmented sub-images as the unlabeled set; Use the annotation set and the unlabeled set to train an improved Mixmatch model until the model converges to generate a target classifier; Input the unlabeled set into the target classifier, and calculate the score of the commodity to be learned according to the output confidence; Determine the data acquisition strategy according to the relationship between the score of the commodity to be learned and the number of modeling requirements; Among them, calculating the score of the product to be learned according to the output confidence includes: determining the two categories with the highest confidence in the output confidence, and respectively denoting the confidences as P A and P B . The first preset value and the second preset value are respectively denoted as T 1 and T 2 . Assume P A >P B ; If P A > T 1 , then the score of the product category corresponding to P A is 1, and the scores of other categories are set to 0; If P A ≤T 1 and P A is a negative sample class, then the scores of all product categories are 0; If P A + P B ≤ T 2 , then the scores of all product categories are 0; If P A + P B > T 2 , then the commodity category scores corresponding to P A and P B are respectively: Where, O A and O B are the normalized score outputs of the categories where P A and P B are located respectively; R A and R B are the sharpened intermediate results of the confidence scores respectively; t is the temperature coefficient; Among them, the determining the data acquisition strategy according to the relationship between the score of the commodity to be learned and the number of modeling requirements includes: Calculate the total score of each commodity category. If the total score of a certain commodity category is less than 90% of the number of modeling requirements, mark that this type of commodity needs to supplement commodity images; If the total score of a certain commodity category is higher than 2 times the number of modeling requirements, mark that the number of images of this type of commodity meets the standard.
2. The data evaluation method according to claim 1, characterized in that, The determining the data acquisition strategy according to the relationship between the score of the commodity to be learned and the number of modeling requirements further includes: Judge whether the balance degree of the commodity category meets the preset requirements according to the score of the commodity category; If the balance degree of the commodity category does not meet the preset requirements, screen the scene images, including: Determine the score set D of the sub-images segmented from the scene images: D = {d 1 , d 2 ... d m}; Among them, where d i is the i-th scene image, represents the score of segmenting the j-th sub-image from the i-th scene image without including the negative sample category; Calculate the variance of D, denoted as var, and set the third preset value as T 3 ; For each item d in loop D i , loop m times; after removing the current d i item, a new set D 1 , D 2 ,... D m ; and calculate the variances of D 1 , D 2 ,... D m respectively, denoted as var 1 , var 2 … var m ; Use var to calculate the differences with var 1 , var 2 … var m respectively, obtaining m variance differences. Select the largest variance difference among them and denote it as var index ; When var index ≥T 3 , determine the corresponding d index items and eliminate them, add the index-th image to the set S, and then return to execute the step of calculating the variance of D; where the set S is the scene images that do not need to enter the manual annotation process. When var index <T 3 the screening process ends.
3. The data evaluation method according to claim 1, characterized in that, Before training the improved Mixmatch model using the annotation set and the unlabeled set, it further includes: Obtain the improved Mixmatch model, including: Rotate the image after MixUP enhancement of the input image by 180 degrees; Use the Object classifier to perform image classification and calculate the image classification loss; Use the Rotation predictor to perform rotation angle classification prediction and calculate the rotation angle classification loss; Determine the loss function of the improved Mixmatch model according to the image classification loss and the rotation angle classification loss.
4. The data evaluation method according to claim 3, characterized in that, The obtaining the improved Mixmatch model further includes: Execute the pseudo-label negative sample enhancement strategy, including: Input the unlabeled set into the target classifier for prediction to obtain the confidence values of positive and negative samples; For positive samples with a confidence less than the fourth preset value, increase the confidence value of negative samples and decrease the confidence value of positive samples: Where P is the model prediction value, P pos represents the positive sample model prediction value, P neg represents the negative sample model prediction value, and T is the confidence threshold.
5. The data evaluation method according to claim 4, characterized in that, The executing the pseudo-label negative sample enhancement strategy further includes: Determine the confidence distribution fluctuation difference according to the KL divergence. For samples with a KL divergence value greater than the divergence threshold, increase the confidence value of negative samples and decrease the confidence value of positive samples: Where P is the model prediction value, P pos represents the positive sample model prediction value, P neg represents the negative sample model prediction value, T kl is the divergence threshold, P i is the prediction result of the model for the i-th randomly augmented data, P j is the prediction result of the model for the j-th randomly augmented data.
6. The data evaluation method according to claim 1, characterized in that, Performing commodity detection and segmentation on the scene image containing the commodity includes: Using a detector to perform commodity detection and segmentation on the scene image containing the commodity, and the detector adopts the VFNet algorithm.
7. A data evaluation system Characterized in that It includes: An annotation set generation unit, configured to select n images from each type of commodity to be learned based on a commodity database as an annotation set; where n << m, and m is the number of modeling requirements; An unannotated set generation unit, configured to perform commodity detection and segmentation on the scene image containing the commodity, and use the segmented sub-images as an unannotated set; A target classifier generation unit, configured to train an improved Mixmatch model using the annotation set and the unannotated set until the model converges to generate a target classifier; A classification score calculation unit, configured to input the unannotated set into the target classifier and calculate the score situation of the commodity to be learned according to the output confidence; A data acquisition strategy determination unit, configured to determine a data acquisition strategy according to the relationship between the score situation of the commodity to be learned and the number of modeling requirements; Among them, calculating the score of the product to be learned according to the output confidence includes: determining the two categories with the highest confidence in the output confidence, and respectively denoting the confidence as P A and P B . The first preset value and the second preset value are respectively denoted as T 1 and T 2 . Assume P A >P B ; If P A > T 1 , then the score of the product category corresponding to P A is 1, and the scores of other categories are set to 0; If P A ≤T 1 and P A is a negative sample class, then the scores of all product categories are 0; If P A + P B ≤ T 2 , then the scores of all product categories are 0; If P A +P B >T 2 , then the commodity category scores corresponding to P A and P B are respectively: where O A and O B are the normalized score outputs of the categories where P A and P B are located; R A and R B are the sharpened intermediate results of the confidence scores; t is the temperature coefficient; Among them, determining the data acquisition strategy according to the relationship between the score situation of the commodity to be learned and the number of modeling requirements includes: Calculating the total score of each commodity category. If the total score of a certain commodity category is less than 90% of the number of modeling requirements, mark that this type of commodity needs to supplement commodity images; If the total score of a certain commodity category is higher than twice the number of modeling requirements, mark that the number of images of this type of commodity meets the standard.
8. A computer-readable storage medium, on which a computer program is stored Characterized in that When the computer program is executed by a processor, it implements the data evaluation method according to any one of claims 1-6.
Citation Information
Patent Citations
Dataset construction method and device, mobile terminal and readable storage medium
CN108764372A
Semi-supervised text classification model training method, text classification method, system, device and medium
CN111723209A