A High-Value Data Mining Method and Device Based on Retail Scenarios
By using multiple classification models to classify and judge retail product pictures in retail scenarios, high-value data are automatically determined, and the model is updated through annotation and iterative training, the problem of new packaging identification and high manual screening costs in the retail industry is solved, and automated and sustainable new product data mining and identification are achieved.
Patent Information
- Application Number
- CN202210155875.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-02-21
AI Technical Summary
In the retail industry, when providers receive new packaging products, it is difficult for providers to effectively identify new packaging, resulting in damage to the identification model. The existing methods require manual screening of new product data, which is costly and inefficient, and cannot support automatic sustainable scenarios.
By obtaining retail product pictures, using preset multiple classification models to classify the pictures, determine whether the average difference between the maximum category probability and the probability of the same category meets the preset conditions, and then automatically determine that the picture is high-value data, and update the classification model through labeling and iterative training.
It realizes automatic, proactive and sustainable new product data mining, reduces labor costs, improves data mining efficiency, and supports online models to actively discover new products and quickly obtain new product recognition capabilities.
Smart Images

Figure CN114528342B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of product packaging, and more particularly, to a high-value data mining method and device based on a retail scenario. Background Art
[0002] Currently, in a retail scenario, a brand owner usually uploads pictures to a provider so that the provider can identify the SKUs in the pictures through a deep learning algorithm and return them to the brand owner, thereby completing the intelligent channel verification of the brand owner. However, in the retail industry, whether it is a competing product or the product itself, the packaging is updated relatively quickly. This directly leads to an inevitable negative impact on the recognition model when the provider continuously receives new packaging. In the existing solution, the brand owner will notify the AI in advance to retrain the recognition model based on the new product data collected manually, so that the recognition model has the ability to identify products with new packaging. However, in practice, it is found that this method still requires manual screening of new product data, resulting in high labor costs and low efficiency; in addition, in this method, the AI is only in a passive position and cannot effectively perform active recognition, so that the existing method cannot be competent in an automatic and sustainable scenario. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a high-value data mining method and device based on a retail scenario, which can avoid manual screening of new product data, thereby realizing active, automatic, and sustainable new product data mining, and further reducing labor costs and improving mining efficiency.
[0004] The first aspect of the embodiments of the present application provides a high-value data mining method based on a retail scenario, including:
[0005] Obtain retail product pictures;
[0006] Classify the retail product pictures through a plurality of preset classification models to obtain a plurality of maximum class probabilities corresponding to the plurality of classification models one by one;
[0007] Determine whether the plurality of maximum class probabilities are all less than a preset class probability;
[0008] When the plurality of maximum class probabilities are all less than the preset class probability, determine that the retail product picture is high-value data.
[0009] Further, the method further includes:
[0010] When the plurality of maximum class probabilities are all greater than or equal to the preset class probability, obtain a plurality of same-class class probabilities corresponding to the plurality of classification models one by one; the plurality of same-class class probabilities correspond to the same classification category;
[0011] Determine whether the average difference of the probabilities of the multiple homogeneous categories is greater than or equal to a preset average difference threshold;
[0012] When the average difference of the probabilities of the multiple homogeneous categories is greater than or equal to the average difference threshold, determine that the retail commodity picture is high-value data.
[0013] Further, the method further includes:
[0014] When the average difference of the probabilities of the multiple homogeneous categories is less than the average difference threshold, determine the class center and class radius of the retail commodity picture according to each classification model;
[0015] Perform calculations according to a preset calculation formula, the retail commodity picture, the class center, the class radius, and preset parameters to obtain a calculation result;
[0016] When the calculation result indicates that the feature distance of the retail commodity picture is greater than or equal to a preset distance, determine that the retail commodity picture is high-value data; the preset distance is the product of the class radius and the preset parameters.
[0017] Further, the determination formula for the class center is:
[0018] Where,
[0019] μ i Represents the class center;
[0020] I ij Represents the jth image with classification category i in the training set;
[0021] E represents the convolutional layer module in the classification model;
[0022] E(I ij ) represents the feature value of the jth image with classification category i;
[0023] n i Represents the number of images with classification category i in the training set.
[0024] Further, the determination formula for the class radius is:
[0025] Where,
[0026] d i Represents the class radius;
[0027] μ i Represents the class center;
[0028] E represents the convolutional layer module in the classification model;
[0029] E(I ij) represents the eigenvalue of the j-th image with the classification category i;
[0030] ||E(I ij )-μ i || represents the feature distance of the j-th image with the classification category i, and the feature distance is the vector two-norm.
[0031] Further, the calculation formula is:
[0032] ||E(I pd )-μ i ||≥C 2 d i ; where
[0033] I pd represents a retail commodity picture;
[0034] ||E(I pd )-μ i || represents the feature distance of the retail commodity picture, and the feature distance is the vector two-norm;
[0035] C 2 d i represents a preset distance;
[0036] d i represents a class radius;
[0037] C 2 represents a preset parameter.
[0038] Further, the method further includes:
[0039] Labeling the high-value data to obtain labeled data;
[0040] Iteratively training the classification model according to the labeled data to obtain a new classification model.
[0041] The second aspect of the embodiments of the present application provides a high-value data mining device based on a retail scenario. The high-value data mining device based on a retail scenario includes:
[0042] An acquisition unit for acquiring retail commodity pictures;
[0043] A classification unit for classifying the retail commodity pictures through a plurality of preset classification models to obtain a plurality of maximum class probabilities corresponding to the plurality of classification models one by one;
[0044] A judgment unit for judging whether the plurality of maximum class probabilities are all less than a preset class probability;
[0045] A determination unit, configured to determine that the retail commodity picture is high-value data when all the multiple maximum category probabilities are less than the preset category probability.
[0046] A third aspect of the embodiments of the present application provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the high-value data mining method based on the retail scenario according to any one of the first aspects of the embodiments of the present application.
[0047] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores computer program instructions. When the computer program instructions are read and run by a processor, the high-value data mining method based on the retail scenario according to any one of the first aspects of the embodiments of the present application is executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0049] Figure 1 A flowchart of a high-value data mining method based on the retail scenario provided by the embodiments of the present application;
[0050] Figure 2 A structural diagram of a high-value data mining device based on the retail scenario provided by the embodiments of the present application;
[0051] Figure 3 A determination process of the classification category of a retail commodity picture provided by the embodiments of the present application;
[0052] Figure 4 A structural diagram of a classification model provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.
[0054] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance.
[0055] Embodiment 1
[0056] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a high-value data mining method based on a retail scenario provided in this embodiment. Among them, the high-value data mining method based on the retail scenario includes:
[0057] S101. Obtain retail product images.
[0058] In this embodiment, it is not yet determined whether the retail product images are high-value data.
[0059] S102. Classify the retail product images through a plurality of preset classification models to obtain a plurality of maximum class probabilities corresponding one-to-one to the plurality of classification models.
[0060] In this embodiment, if there are three classification models, name these three classification models as Classification Model 1, Classification Model 2, and Classification Model 3. At this time, Classification Model 1 can give the class probability P1 corresponding to the first category, the class probability P2 corresponding to the second category, and the class probability P3 corresponding to the third category; Classification Model 2 can give the class probability P4 corresponding to the first category, the class probability P5 corresponding to the second category, and the class probability P6 corresponding to the third category; Classification Model 3 can give the class probability P7 corresponding to the first category, the class probability P8 corresponding to the second category, and the class probability P9 corresponding to the third category.
[0061] In this embodiment, the plurality of maximum class probabilities can be P2, P6, and P8. Among them, P2 is larger than P1 and P3, so it is proposed; P6 is larger than P5 and P4, so it is proposed; P8 is larger than P7 and P9, so it is proposed.
[0062] S103. Determine whether the plurality of maximum class probabilities are all less than a preset class probability. If so, execute step S109; if not, execute step S104.
[0063] In this embodiment, if the plurality of maximum class probabilities are all less than the preset class probability, then the retail product image is considered a high-value image. The reason is that no classification model can effectively identify the retail product image.
[0064] In this embodiment, a better effect can be obtained by taking the preset class probability as 0.42.
[0065] In this embodiment, the method can correspond to n categories, and the classification probability is generally an n-dimensional vector.
[0066] S104. Obtain a plurality of same-class class probabilities corresponding one-to-one to the plurality of classification models; the plurality of same-class class probabilities correspond to the same classification category.
[0067] S105. Determine whether the average difference of multiple probabilities of the same category is greater than or equal to a preset average difference threshold. If so, execute step S109; if not, execute step S106.
[0068] In this embodiment, assume that the commodity category is category i, and the probabilities predicted by three classification models for this category are P1, P2, and P3 respectively. Then the arithmetic mean is:
[0069]
[0070] At this time, if |P1 i - u i | + |P2 i - u i | + |P3 i - u i | ≥ C 1 , it is considered that the retail commodity picture is high-value data. In practice, C 1 can take a value of 0.5.
[0071] S106. Determine the class center and class radius of the retail commodity picture according to each classification model.
[0072] In this embodiment, the formula for determining the class center is:
[0073] Where,
[0074] μ i represents the class center:
[0075] I ij represents the jth image in the training set with the classification category of i;
[0076] E represents the convolutional layer module in the classification model:
[0077] E(I ij ) represents the eigenvalue of the jth image with the classification category of i;
[0078] n i represents the number of images with the classification category of i in the training set.
[0079] In this embodiment, the formula for determining the class radius is:
[0080] Where,
[0081] d i represents the class radius;
[0082] μ i represents the class center;
[0083] E represents the convolutional layer module in the classification model;
[0084] E(I ij ) represents the eigenvalue of the j-th image with classification category i;
[0085] ||E(I ij ) - μ i || represents the feature distance of the j-th image with classification category i, and the feature distance is the vector two-norm.
[0086] In this embodiment, the pictures in the training set are passed through the convolutional layer in the classification model to obtain the image Embedding of the training set, and the following features of the training set category are calculated:
[0087] Class center μ i : where I ij represents the j-th image in category i in the training set, E represents the convolutional layer module in the classification, and E(I ij ) represents the Embedding of the j-th image in category i, and n i represents the number of images in category i in the training set;
[0088] Class radius d i : where || || represents the vector two-norm.
[0089] In this embodiment, the above calculations can be obtained after the model is trained. During the service provision, for each classification model, if it is predicted that the product picture is category i, and if there is ||E(I pd ) - μ i || ≥ C 2 d i , (where I pd represents the product picture), then the product picture is high-value data. In practice, C 2 is taken as 1.2.
[0090] S107. Calculate according to the preset calculation formula, retail product picture, class center, class radius, and preset parameters to obtain the calculation result.
[0091] In this embodiment, the calculation formula is:
[0092] ||E(I pd ) - μ i | ≥ C 2 d i ; where,
[0093] I pd represents the retail product picture;
[0094] ||E(I pd ) - μ i|| represents the feature distance of the retail commodity picture, and the feature distance is the vector two-norm;
[0095] C 2 d i represents a preset distance;
[0096] d i represents the class radius;
[0097] C 2 represents a preset parameter.
[0098] In this embodiment, C 2 can be taken as 1.2
[0099] S108. Determine whether the calculation result indicates that the feature distance of the retail commodity picture is greater than or equal to the preset distance. If so, execute step S109; if not, end this process.
[0100] In this embodiment, the preset distance is the product of the class radius and the preset parameter.
[0101] In this embodiment, the preset parameter is C 2 .
[0102] S109. Determine that the retail commodity picture is high-value data.
[0103] In this embodiment, the high-value data is high-value pictures, and high-value is used to represent data that can identify new products.
[0104] S110. Label the high-value data to obtain labeled data.
[0105] In this embodiment, it can be understood that the workload of labeling the high-value data above is much smaller than the workload of manually selecting new-packaged commodities.
[0106] S111. Iteratively train the classification model according to the labeled data to obtain a new classification model.
[0107] Please refer to Figure 3 , Figure 3 shows a process for determining the classification category of retail commodity pictures. In this process, first obtain retail pictures, then identify the commodity frame through AI recognition and crop the commodity frame to obtain a frameless picture, and further classify the frameless picture to obtain the category probabilities of multiple classification models, and calculate the mean value to determine the highest probability as the commodity category.
[0108] Please refer to Figure 4 , Figure 4 shows a structural schematic diagram of a classification model. Figure 4Among them, the classification model mainly obtains the image Embedding (usually a 512-dimensional vector) through the convolutional layer, and then sends the Embedding into the fully connected layer and passes through the softmax layer to obtain an n-dimensional vector, thereby obtaining the class probability.
[0109] In this embodiment, the execution subject of this method can be a computing device such as a computer or a server, and no limitation is made in this embodiment.
[0110] In this embodiment, the execution subject of this method can also be a smart device such as a smart phone or a tablet computer, and no limitation is made in this embodiment.
[0111] It can be seen that implementing the high-value data mining method based on the retail scenario described in this embodiment can perform online calculation at a small cost to select high-value pictures in real time, and then label the high-value pictures, and then use the labeled data to train the classification model, so that the online model can actively discover new products and quickly and low-costly obtain the new product recognition ability, thereby optimizing the user experience.
[0112] Embodiment 2
[0113] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a high-value data mining device based on the retail scenario provided in this embodiment. As Figure 2 shown, the high-value data mining device based on the retail scenario includes:
[0114] An acquisition unit 210, configured to acquire retail commodity pictures;
[0115] A classification unit 220, configured to classify the retail commodity pictures through a plurality of preset classification models to obtain a plurality of maximum class probabilities corresponding one-to-one to the plurality of classification models;
[0116] A judgment unit 230, configured to judge whether the plurality of maximum class probabilities are all less than a preset class probability;
[0117] A determination unit 240, configured to determine that the retail commodity picture is high-value data when the plurality of maximum class probabilities are all less than the preset class probability.
[0118] As an optional implementation manner, the acquisition unit 210 is further configured to acquire a plurality of same-class class probabilities corresponding one-to-one to the plurality of classification models when the plurality of maximum class probabilities are all greater than or equal to the preset class probability; the plurality of same-class class probabilities correspond to the same classification category;
[0119] The judgment unit 230 is further configured to judge whether the average difference of the plurality of same-class class probabilities is greater than or equal to a preset average difference threshold;
[0120] The determination unit 240 is further configured to determine that the retail commodity picture is high-value data when the average difference of multiple probabilities of the same category is greater than or equal to the average difference threshold.
[0121] As an alternative implementation, the high-value data mining device based on the retail scenario further includes:
[0122] The determination unit 240 is further configured to, when the average difference of multiple probabilities of the same category is less than the average difference threshold, determine the class center and class radius of the retail commodity picture according to each classification model;
[0123] The calculation unit 250 is configured to calculate according to a preset calculation formula, the retail commodity picture, the class center, the class radius, and preset parameters to obtain a calculation result;
[0124] The determination unit 240 is further configured to determine that the retail commodity picture is high-value data when the calculation result indicates that the feature distance of the retail commodity picture is greater than or equal to a preset distance; the preset distance is the product of the class radius and the preset parameters.
[0125] As an alternative implementation, the formula for determining the class center is:
[0126] Wherein,
[0127] μ i represents the class center;
[0128] I ij represents the j-th image with the classification category i in the training set;
[0129] E represents the convolutional layer module in the classification model;
[0130] E(I ij ) represents the feature value of the i-th image with the classification category i;
[0131] n i represents the number of images with the classification category i in the training set.
[0132] As an alternative implementation, the formula for determining the class radius is:
[0133] Wherein,
[0134] d i represents the class radius;
[0135] μ i represents the class center∶
[0136] E represents the convolutional layer module in the classification model;
[0137] E(I ijrepresents the eigenvalue of the j-th image with the classification category i;
[0138] ||E(I ij ) - μ i || represents the feature distance of the j-th image with the classification category i, and the feature distance is the vector two-norm.
[0139] As an alternative implementation, the calculation formula is:
[0140] ||E(I pd ) - μ i || ≥ C 2 d i ; where,
[0141] I pd represents the retail commodity picture;
[0142] ||E(I pd ) - μ i || represents the feature distance of the retail commodity picture, and the feature distance is the vector two-norm;
[0143] C 2 d i represents the preset distance;
[0144] d i represents the class radius;
[0145] C 2 represents the preset parameter.
[0146] As an alternative implementation, the high-value data mining device based on the retail scenario further includes:
[0147] The annotation unit 260 is used to annotate the high-value data to obtain the annotated data;
[0148] The training unit 270 is used to iteratively train the classification model according to the annotated data to obtain a new classification model.
[0149] In the embodiments of the present application, the explanation of the high-value data mining device based on the retail scenario can refer to the description in Embodiment 1, and no further elaboration will be made in this embodiment.
[0150] It can be seen that implementing the high-value data mining device based on the retail scenario described in this embodiment can perform online calculations at a relatively low cost to select high-value pictures in real time, and then annotate the high-value pictures to train the classification model with the annotated data, so that the online model can actively discover new products and quickly and low-costly obtain new product recognition capabilities, thereby optimizing the user experience.
[0151] An embodiment of the present application provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the high-value data mining method based on the retail scenario in Embodiment 1 of the present application.
[0152] An embodiment of the present application provides a computer-readable storage medium, which stores computer program instructions. When the computer program instructions are read and run by a processor, the high-value data mining method based on the retail scenario in Embodiment 1 of the present application is executed.
[0153] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0154] In addition, in each embodiment of the present application, the various functional modules may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0155] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0156] The above are only the embodiments of this application and are not used to limit the protection scope of this application. For those skilled in the art, this application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application. It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0157] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in this application and should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0158] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
Claims
1. A high-value data mining method based on a retail scenario, characterized in that, it includes: Obtain retail product pictures; Classify the retail product pictures through a plurality of preset classification models to obtain a plurality of maximum class probabilities corresponding one-to-one to the plurality of classification models; Judge whether the plurality of maximum class probabilities are all less than a preset class probability; When the plurality of maximum class probabilities are all less than the preset class probability, determine that the retail product picture is high-value data; Wherein, the method further includes: When the plurality of maximum class probabilities are all greater than or equal to the preset class probability, obtain a plurality of same-class class probabilities corresponding one-to-one to the plurality of classification models; the plurality of same-class class probabilities correspond to the same classification category; Judge whether the average difference of the plurality of same-class class probabilities is greater than or equal to a preset average difference threshold; When the average difference of the plurality of same-class class probabilities is greater than or equal to the average difference threshold, determine that the retail product picture is high-value data; Wherein, the method further includes: When the average difference of the plurality of same-class class probabilities is less than the average difference threshold, determine the class center and class radius of the retail product picture according to each classification model; Calculate according to a preset calculation formula, the retail product picture, the class center, the class radius and preset parameters to obtain a calculation result; When the calculation result indicates that the feature distance of the retail product picture is greater than or equal to a preset distance, determine that the retail product picture is high-value data; the preset distance is the product of the class radius and the preset parameter; Wherein, the formula for determining the class center is: Among them, μ i represents the class center; I ij denotes the j-th image in the training set with the classification category i; E represents the convolutional layer module in the classification model; E(I ij ) represents the eigenvalue of the j-th image with the classification category i; n i represents the number of images in the training set with the classification category i.
2. The high-value data mining method based on a retail scenario according to claim 1, characterized in that, The formula for determining the class radius is: Among them, d i represents the class radius; μ i represents the class center; E represents the convolutional layer module in the classification model; E(I ij ) represents the eigenvalue of the j-th image with the classification category i; ||E(I ij )-μ i || represents the feature distance of the j-th image with the classification category i, and the feature distance is the vector two-norm.
3. The high-value data mining method based on a retail scenario according to claim 1, characterized in that, The calculation formula is: ||E(I pd ) - μ i || ≥ C 2 d i ; wherein, I pd represents a picture of a retail product; ||E(I pd )-μ i || represents the feature distance of the retail commodity picture, and the feature distance is the vector two-norm; C 2 d i represents a preset distance; d i represents the class radius; C 2 Represents preset parameters.
4. The high-value data mining method based on a retail scenario according to claim 1, characterized in that, The method further includes: Label the high-value data to obtain labeled data; Iteratively train the classification model according to the labeled data to obtain a new classification model.
5. A high-value data mining device based on a retail scenario, characterized in that, The high-value data mining device based on a retail scenario includes: An acquisition unit for obtaining retail product pictures; A classification unit for classifying the retail product pictures through a plurality of preset classification models to obtain a plurality of maximum class probabilities corresponding one-to-one to the plurality of classification models; A judgment unit for judging whether the plurality of maximum class probabilities are all less than a preset class probability; A determination unit for determining that the retail product picture is high-value data when the plurality of maximum class probabilities are all less than the preset class probability; Wherein, the acquisition unit is further configured to obtain a plurality of same-class class probabilities corresponding one-to-one to the plurality of classification models when the plurality of maximum class probabilities are all greater than or equal to the preset class probability; the plurality of same-class class probabilities correspond to the same classification category; The determination unit is further configured to determine whether the average difference of multiple probabilities of the same category is greater than or equal to a preset average difference threshold; The determination unit is further configured to determine that the retail commodity picture is high-value data when the average difference of multiple probabilities of the same category is greater than or equal to the average difference threshold; Wherein, the high-value data mining device based on the retail scenario further includes: The determination unit is further configured to, when the average difference of multiple probabilities of the same category is less than the average difference threshold, determine the class center and class radius of the retail commodity picture according to each classification model; The calculation unit is configured to perform calculations according to a preset calculation formula, the retail commodity picture, the class center, the class radius, and preset parameters to obtain a calculation result; The determination unit is further configured to determine that the retail commodity picture is high-value data when the calculation result indicates that the feature distance of the retail commodity picture is greater than or equal to a preset distance; the preset distance is the product of the class radius and the preset parameters; Wherein, the determination formula of the class center is: Among them, μ i represents the class center; I ij denotes the j-th image in the training set with the classification category i; E represents the convolutional layer module in the classification model; E(I ij ) represents the eigenvalue of the j-th image with the classification category of i; n i represents the number of images in the training set with the classification category of i.
6. An electronic device, Characterized in that, The electronic device includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the high-value data mining method based on the retail scenario according to any one of claims 1 to 4.
7. A readable storage medium, Characterized in that, The readable storage medium stores computer program instructions, and when the computer program instructions are read and run by a processor, the high-value data mining method based on the retail scenario according to any one of claims 1 to 4 is executed.
Citation Information
Patent Citations
Sample rejection method, device and equipment, and storage medium
CN112308131A