Commodity data cleaning method and its device, equipment, medium, and product

By using classifiers and clustering algorithms to clean product data, the problem of simple and crude data cleaning methods in the prior art is solved, the quality of training data and the training effect of neural network models are improved, and it is especially suitable for product data cleaning in the e-commerce field.

CN113918554BActive Publication Date: 2025-05-16BUSINESS LINE COMMERCIAL PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111271713.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-05-16
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

When the prior art cleans the training data required by neural network models, the method is simple and crude, and the value factor of the data to the model is not fully considered, resulting in limited training effects. Especially in the e-commerce field, the source of label information of product data is different, so how to effectively clean and label these data has become a problem.

Method used

The classifier is used to classify the deep semantic information of the product data. According to the error information between the predicted classification label and the original classification label, the data that needs to be cleaned is determined, and the data that is inconsistent with the original classification label is deleted through the clustering algorithm to obtain the purified training set.

Benefits of technology

Through fine data cleaning methods, the quality of training data is improved, the impact of noise data is reduced, and the training effect of neural network models is improved. Especially in the e-commerce field, the data cleaning problem caused by different sources of product data label information is effectively solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113918554B_ABST
    Figure CN113918554B_ABST
Patent Text Reader

Abstract

The present application discloses a commodity data cleaning method and its device, equipment, medium, and product, the method comprising: determining a commodity data set, the commodity data set is divided into a training set and a data set, including a plurality of commodity data carrying original classification labels; using a classifier trained by the training set to classify the commodity data in the test set, and obtaining a predicted classification label for each commodity data; according to the error information between the predicted classification label and the original classification label, determining a plurality of original classification labels with the most erroneous predictions, and extracting the commodity data under these original classification labels from the training set as the data set to be cleaned; clustering the commodity data in the data set to be cleaned, and deleting the commodity data whose clustering results are inconsistent with the original classification labels from the training set. The present application can effectively clean a large amount of commodity data to obtain a high-quality training set, so that the classifier trained by it can obtain a classification prediction effect with a higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of e-commerce information technology, and in particular to a commodity data cleaning method and its corresponding device, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] With the development of deep learning, the scale of various neural network models is getting larger and larger, and the scale of training data is also getting larger and larger. As the scale expands, the noise data in the training data is also gradually increasing. A large amount of noise data will seriously affect the effect of the model. Traditional manual labeling is difficult to efficiently process such a large amount of data. If the data is not effectively cleaned, the noise data will seriously affect the training effect of the neural network model.

[0003] In the prior art, the training data required for the neural network model is cleaned by applying simple means to detect incomplete data, erroneous data, duplicate data, etc. This is very simple and crude, and does not fully consider the value factors of the data itself to the neural network model. Therefore, it has little effect on improving the value of the training data relative to the model.

[0004] The training data required for model training, especially data with supervisory labels, and whether the corresponding labels are accurate and effective will greatly affect the training effect of the model, and specifically affect the model's learning ability. Therefore, more attention should be paid to this in the data cleaning stage. The relevant solutions proposed by the industry for cleaning labeled training data are mainly based on clustering algorithms, which simply verify the clustering results and remove training data whose labeling information does not match the clustering results. Although this method improves the quality of the data to a considerable extent, it is not sophisticated enough and fails to be combined with the model to clean the relevant data. Therefore, there is still room for improvement.

[0005] The need for data cleaning is particularly evident in the e-commerce field. In the e-commerce field, the product data corresponding to the massive amount of goods generally have label information corresponding to them one by one. However, when the label information of these product data comes from different sources or is generated with different standards, how to effectively label these product data becomes a greater problem.

[0006] In summary, how to clean the training data consisting of product data so that it can adapt to the needs of the neural network model and become effective training data is worth exploring in the e-commerce field. Summary of the invention

[0007] The primary purpose of the present application is to solve at least one of the above problems and to provide a commodity data cleaning method and its corresponding device, computer equipment, computer-readable storage medium, and computer program product.

[0008] In order to meet the various objectives of this application, this application adopts the following technical solutions:

[0009] A commodity data cleaning method provided for one of the purposes of this application includes the following steps:

[0010] Determine a commodity data set, wherein the commodity data set includes a plurality of commodity data carrying original classification labels, and the commodity data is divided into a training set and a data set;

[0011] Using a classifier trained by the training set to classify the deep semantic information of the commodity data in the test set, to obtain a predicted classification label corresponding to each commodity data;

[0012] According to the error information between the predicted classification label and the original classification label of the commodity data, a plurality of original classification labels with a predetermined number of the most incorrect predictions are determined, and the commodity data under these original classification labels are extracted from the training set as the data set to be cleaned;

[0013] The product data in the cleansing data set are clustered, and the product data whose clustering results are inconsistent with the original classification labels are deleted from the training set to obtain a purified training set consisting of the remaining product data.

[0014] In a further embodiment, the deep semantic information of the commodity data in the test set is classified by using a classifier trained by the training set to obtain a predicted classification label corresponding to each commodity data, including the following steps:

[0015] Using a feature extraction model to obtain deep semantic information of the commodity data in the test set, the deep semantic information includes text feature information and / or image feature information of the commodity data;

[0016] A classifier is used to calculate the classification probability corresponding to the mapping of the deep semantic information to each of the original classification labels to determine the predicted classification label to which the commodity data corresponding to the deep semantic information belongs, and the predicted classification label is the original classification label with the highest classification probability for the commodity data.

[0017] In a specific embodiment, a feature extraction model is used to obtain deep semantic information of the commodity data in the test set, wherein the deep semantic information includes text feature information and / or image feature information of the commodity data, and the following steps are included:

[0018] Using a first feature extraction model to extract deep semantic information from the product titles of the product data in the test set to obtain corresponding text feature information;

[0019] Using a second feature extraction model to extract deep semantic information from the product images of the product data in the test set to obtain corresponding image feature information;

[0020] The text feature information and the image feature information of the same product data are spliced ​​into graphic feature information representing the deep semantic information of the product data.

[0021] In a further embodiment, according to the error information between the predicted classification label and the original classification label of the commodity data, a plurality of original classification labels with a predetermined number of the most incorrect predictions are determined, and the commodity data under these original classification labels are extracted from the training set as the data set to be cleaned, including the following steps:

[0022] A confusion matrix is ​​calculated based on the predicted classification labels of the commodity data in the test set and their original classification labels, wherein each element in the confusion matrix is ​​used to represent the statistical number of commodity data corresponding to the original classification label that are predicted to be a certain predicted classification label;

[0023] From the elements in the confusion matrix whose original classification labels are inconsistent with the predicted classification labels, for a predetermined number of target elements with the largest statistical number, determine the original classification labels corresponding to these target elements;

[0024] The commodity data carrying the original classification labels corresponding to the target elements are extracted from the training set to form a data set to be cleaned.

[0025] In a further embodiment, the commodity data in the to-be-cleaned data set are clustered, commodity data whose clustering results are inconsistent with the original classification labels are deleted from the training set, and a purified training set consisting of the remaining commodity data is obtained, including the following steps:

[0026] Using a preset clustering algorithm, clustering is performed according to the deep semantic information of the commodity data in the data set to be cleaned, to obtain a plurality of commodity data clusters obtained by clustering;

[0027] Counting the maximum number of original classification labels possessed by the commodity data in each commodity data cluster, and deleting the commodity data in the commodity data cluster that does not carry the maximum number of original classification labels to achieve data cleaning of the commodity data cluster;

[0028] A cleansed training set is determined, where the cleansed training set includes commodity data in each commodity data cluster that has completed data cleaning.

[0029] In the extended embodiment, after clustering the commodity data in the cleaned data set, deleting the commodity data whose clustering results are inconsistent with the original classification labels from the training set, and obtaining the purified training set consisting of the remaining commodity data, the following steps are included:

[0030] Using the purified training set as a new training set to train the classifier;

[0031] After the classifier is trained, the method is executed in a loop starting from the step of using the classifier trained by the training set to classify the deep semantic information of the commodity data in the test set;

[0032] Wherein, in the step of determining a predetermined number of original classification labels having the most erroneous predictions, the predetermined number is reset.

[0033] In a further embodiment, the commodity data cleaning method of the present application further includes the following subsequent steps:

[0034] Counting the prediction accuracy of the classifier obtained from each training, where the prediction accuracy is the accuracy rate at which the deep semantic information of the commodity data in the test set is correctly classified by the classifier into its original classification label;

[0035] The training set used by the classifier with the highest prediction accuracy is determined to be the optimal training set, and the classifier is used to respond to the classification requirements of the commodity data and determine the corresponding classification of the commodity data.

[0036] A commodity data cleaning device provided to meet one of the purposes of the present application includes: a data determination module, a classification prediction module, a false positive screening module, and a cluster cleaning module, wherein the data determination module is used to determine a commodity data set, the commodity data set includes multiple commodity data carrying original classification labels, and the commodity data is divided into a training set and a data set; the classification prediction module is used to classify the deep semantic information of the commodity data in the test set using a classifier trained by the training set to obtain a predicted classification label corresponding to each commodity data; the false positive screening module is used to determine a plurality of original classification labels with a predetermined number of the most false predictions based on error information between the predicted classification label and the original classification label of the commodity data, and extract the commodity data under these original classification labels from the training set as the data set to be cleaned; the cluster cleaning module is used to cluster the commodity data in the data set to be cleaned, delete the commodity data whose clustering results are inconsistent with the original classification labels from the training set, and obtain a purified training set composed of the remaining commodity data.

[0037] In a further embodiment, the classification prediction module includes: a semantic extraction submodule, which is used to use a feature extraction model to obtain deep semantic information of the product data in the test set, and the deep semantic information includes text feature information and / or image feature information of the product data; a classification mapping submodule, which is used to use a classifier to calculate the classification probability corresponding to the deep semantic information mapped to each of the original classification labels, so as to determine the predicted classification label to which the product data corresponding to the deep semantic information belongs, and the predicted classification label is the original classification label with the largest classification probability for the product data.

[0038] In a specific embodiment, the semantic extraction submodule includes: a text extraction unit, which is used to use a first feature extraction model to extract deep semantic information from the product titles of the product data in the test set to obtain corresponding text feature information; an image extraction unit, which is used to use a second feature extraction model to extract deep semantic information from the product images of the product data in the test set to obtain corresponding image feature information; and a feature fusion unit, which is used to splice the text feature information and image feature information of the same product data into graphic feature information representing the deep semantic information of the product data.

[0039] In a further embodiment, the false positive screening module includes: a false positive statistics submodule, which is used to count a confusion matrix based on the predicted classification labels and the original classification labels of the commodity data in the test set, and each element in the confusion matrix is ​​used to represent the statistical number of commodity data corresponding to the original classification label being predicted to be a certain predicted classification label; a false positive determination submodule, which is used to determine the original classification labels corresponding to a predetermined number of target elements with the largest statistical number among the elements in the confusion matrix whose original classification labels are inconsistent with the predicted classification labels; and a to-be-cleaned extraction submodule, which is used to extract commodity data carrying the original classification labels corresponding to the target elements from the training set to form a data set to be cleaned.

[0040] In a further embodiment, the clustering cleaning module includes: a title clustering submodule, which is used to adopt a preset clustering algorithm to cluster the product data in the data set to be cleaned according to the deep semantic information of the product data to be clustered, and obtain multiple product data clusters obtained by clustering; a cleaning execution submodule, which is used to count the maximum number of original classification labels possessed by the product data in each product data cluster, and delete the product data in the product data cluster that does not carry the maximum number of original classification labels to achieve data cleaning of the product data cluster; a purification data submodule, which is used to determine a purification training set, which includes the product data in each of the product data clusters that have completed data cleaning.

[0041] In an extended embodiment, the commodity data cleaning device of the present application includes: a re-training module, which is used to use the purified training set as a new training set to train the classifier. After the re-training module is run, it triggers the classification prediction module, the false detection screening module, and the cluster cleaning module of the device to run in a cycle again, wherein the predetermined number in the false detection screening module is reset.

[0042] In a further embodiment, the commodity data cleaning device of the present application also includes the following subsequent operating structure: an accuracy statistics module, which is used to count the prediction accuracy of the classifier obtained from each training, and the prediction accuracy is the accuracy of the deep semantic information of the commodity data in the test set being correctly classified by the classifier to its original classification label; a selection and determination module, which is used to determine that the training set used by the classifier with the highest prediction accuracy is the optimal training set, and use the classifier to respond to the classification requirements of the commodity data and determine the corresponding classification of the commodity data.

[0043] A computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory, wherein the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the commodity data cleaning method described in the present application.

[0044] A computer-readable storage medium is provided to meet another purpose of the present application, which stores a computer program implemented according to the commodity data cleaning method in the form of computer-readable instructions. When the computer program is called and executed by a computer, the steps included in the method are executed.

[0045] A computer program product provided to meet another purpose of the present application includes a computer program / instruction, which, when executed by a processor, implements the steps of the method described in any embodiment of the present application.

[0046] Compared with the prior art, the advantages of this application are as follows:

[0047] First, when cleaning the commodity data, the present application adopts the classification results provided by the classifier, and determines the commodity data that needs to be cleaned from the training set as the data set to be cleaned according to the error information in the classification results. Then, based on the data set to be cleaned, the commodity data that needs to be deleted is determined with the help of clustering means. Accordingly, excessive cleaning of the commodity data in the training set can be avoided, thereby making the data cleaning operation more effective and accurate.

[0048] Secondly, before performing data cleaning, the present application determines the prediction error information through the predicted classification results of the classifier, and then determines some of the original classification labels based on the error information, and only performs clustering and data cleaning on the product data corresponding to these original classification labels. It is not difficult to understand that, relatively speaking, only some categories of product data in the training set are cleaned, and there is no need to clean the entire set. Therefore, it can not only reduce the computational pressure of clustering based on the full data of the entire training set, but also significantly improve the data cleaning speed.

[0049] In addition, the purified training set determined in the present application is suitable for more efficiently training new classifier instances, which can facilitate faster convergence of the training process of the classifier instances and improve the prediction accuracy of the classifier. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0051] Figure 1 This is a flowchart of a typical embodiment of the commodity data cleaning method of the present application;

[0052] Figure 2 A flowchart of the process of obtaining deep semantic information of commodity data and classifying it in an embodiment of the present application;

[0053] Figure 3 A block diagram of a neural network model used in the embodiments of the present application;

[0054] Figure 4 A schematic diagram of the process of obtaining image and text feature information using a neural network model in an embodiment of the present application;

[0055] Figure 5 is a principle block diagram of another neural network model used in the embodiments of the present application;

[0056] Figure 6 A flowchart of a process for determining a data set to be cleaned according to error information between a predicted classification label and an original classification label in an embodiment of the present application;

[0057] Figure 7 A schematic diagram of a specific cleaning process for the data to be cleaned in an embodiment of the present application;

[0058] Figure 8 This is a flow chart of the process of iteratively cleaning data of a training set in an embodiment of the present application;

[0059] Fig. 9A flowchart of a process for optimizing a training set and a classifier generated by an iterative cleaning process in an embodiment of the present application;

[0060] Fig.10 This is a principle block diagram of the commodity data cleaning device of this application;

[0061] Fig.11 A schematic diagram of the structure of a computer device used in this application. DETAILED DESCRIPTION

[0062] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.

[0063] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0064] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.

[0065] It will be understood by those skilled in the art that the "client", "terminal" and "terminal device" used herein include both devices with wireless signal receivers, which are devices with only wireless signal receivers without transmission capabilities, and devices with receiving and transmitting hardware, which are devices with receiving and transmitting hardware capable of two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, which have single-line displays or multi-line displays or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service, personal communication system), which can combine voice, data processing, fax and / or data communication capabilities; PDA (Personal Digital Assistant, personal digital assistant), which may include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar and / or GPS (Global Positioning System, global positioning system) receiver; conventional laptop and / or palmtop computers or other devices, which have and / or include a conventional laptop and / or palmtop computer or other device with and / or including a radio frequency receiver. The "client", "terminal" and "terminal device" used herein may be portable, transportable, installed in a vehicle (air, sea and / or land), or suitable for and / or configured to run locally, and / or in a distributed form, at any other location on the earth and / or in space. The "client", "terminal" and "terminal device" used herein may also be a communication terminal, an Internet terminal, a music / video playing terminal, for example, a PDA, a MID (Mobile Internet Device) and / or a mobile phone with a music / video playing function, or a smart TV, a set-top box and other devices.

[0066] The hardware referred to by the names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann principle, such as a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit calls the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input and output devices to complete specific functions.

[0067] It should be pointed out that the concept of "server" referred to in this application can also be extended to the case of server clusters. According to the network deployment principle understood by those skilled in the art, the servers should be logically divided. In physical space, these servers can be independent of each other but can be called through interfaces, or integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility, and should not use it to restrict the implementation of the network deployment method of this application.

[0068] Unless expressly specified, one or more technical features of the present application can be deployed on a server for implementation and accessed by a client through a remote call to obtain an online service interface provided by the server, or can be directly deployed and run on a client for access.

[0069] The neural network models referenced or may be referenced in this application, unless expressly specified, can be deployed on a remote server and remotely called on the client, or can be deployed and directly called on a client with sufficient device capabilities. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.

[0070] Unless explicitly specified, the various data involved in this application can be stored remotely on a server or on a local terminal device, as long as it is suitable for being called by the technical solution of this application.

[0071] Those skilled in the art should be aware that, although the various methods of the present application are described based on the same concept and thus present commonality to each other, unless otherwise specified, these methods can be independently executed. Similarly, for each embodiment disclosed in the present application, they are all proposed based on the same inventive concept, therefore, concepts with the same expression, and concepts that are appropriately changed for convenience despite different expressions, should be understood as equivalent.

[0072] Unless the mutually exclusive relationship between the embodiments to be disclosed in this application is explicitly stated, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct a new embodiment, as long as such combination does not deviate from the creative spirit of this application and can meet the needs of the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0073] A commodity data cleaning method of the present application can be programmed as a computer program product and deployed in a service cluster for execution, so that the method can be executed by accessing an interface opened after the computer program product is running and performing human-computer interaction with the computer program product through a graphical user interface.

[0074] An application scenario exemplified in the present application is an application scenario related to classification task training in the e-commerce field. Due to the need for certain standard classification in the e-commerce field based on the product data corresponding to the product object, including but not limited to product title, product image, product details, product attributes, etc., a product data set composed of prepared product data is used to train the classifier. Therefore, these product data sets can be cleaned with the help of the technical solution of the present application to determine an effective purified training set, and the classifier trained with the purified training set is determined as the classifier required for production, so that the classifier can accurately classify the product data in the e-commerce platform according to the said standard.

[0075] The classification implemented by the classifier can be, for example, the classification of product safety attributes, the classification of products mapped to the category tree of the e-commerce platform, etc. Different classification tasks can be determined according to different classification standards and corresponding training can be performed. Accordingly, the classification labels in the product data set are also classification labels corresponding to the classification standards, so that the classifier can implement supervised learning based on the classification labels and acquire the corresponding classification capabilities after learning. Those skilled in the art should be aware of this.

[0076] See also Figure 1 In a typical embodiment, the commodity data cleaning method of the present application includes the following steps:

[0077] Step S1100: determine a commodity data set, wherein the commodity data set includes a plurality of commodity data carrying original classification labels, and the commodity data is divided into a training set and a data set:

[0078] First, a commodity data set is prepared, wherein the commodity data set contains commodity data required for training the exemplary neural network model of the present application, and the commodity data all carry classification labels with pre-established corresponding mapping relationships, which are referred to as original classification labels herein. The original classification labels are suitable for providing the trained neural network model with the supervision labels required for supervised learning, so as to supervise the training process of the neural network model and promote the neural network model to be trained to a convergence state.

[0079] The product data described above, for the exemplary application scenarios of this application, includes product titles, product images, etc., and the various types of information included therein can be determined based on the specific data on which the neural network model is based. For example, it may also include product details, product attributes and other data.

[0080] The commodity data may be obtained from a commodity database of an e-commerce platform, or from a website that is open and free to obtain data on the Internet, and may be flexibly selected by technicians in this field.

[0081] To facilitate data cleaning, the commodity data in the commodity data set is divided into a training set and a prediction set according to a certain preset ratio. The training set is used to train the neural network model, so it needs to be cleaned in this application; the prediction set is used to implement prediction verification on the neural network model after each training. In order to facilitate the comparison of multiple training results on the same prediction set, the commodity data in the prediction set is relatively fixed and does not need to be cleaned.

[0082] Step S1200: Use the classifier trained by the training set to classify the deep semantic information of the commodity data in the test set to obtain the predicted classification label corresponding to each commodity data:

[0083] The neural network model can be pre-trained with the training set. The neural network model generally includes two parts. The first part is a feature extraction model for representation learning of the commodity data. The second part includes a classifier. The feature extraction model performs representation learning on the relevant specific data in the commodity data to obtain the corresponding deep semantic information of each commodity data. The classifier then classifies the fully connected results of the deep semantic information to map it to a new classification result.

[0084] The feature extraction model is generally a pre-trained model, so it is not difficult to understand that the training of the classification task implemented by the corresponding neural network model in this application is mainly for the training implemented by the classifier. The classification space mapped by the classifier to perform the classification task is a classification space composed of the original classification labels in the commodity data set. When the classification probability corresponding to each classification label in the classification space is calculated for the deep semantic information of a commodity data through a classifier such as Softmax, the classification label with the largest classification probability can be determined as the predicted classification label of the commodity object. Accordingly, the deep semantic information of each commodity data in the prediction set is classified one by one through the classifier, and the predicted classification label corresponding to each commodity data in the prediction set can be determined.

[0085] Step S1300: According to the error information between the predicted classification label and the original classification label of the commodity data, a plurality of original classification labels with a predetermined number of the most incorrect predictions are determined, and commodity data under these original classification labels are extracted from the training set as the data set to be cleaned:

[0086] After each commodity data in the prediction set obtains its corresponding prediction classification label, the prediction classification label of the same commodity data may not be consistent with its original classification label, that is, although the classifier has been trained with the training set, it is still possible to make a wrong prediction because the commodity data in the training set has not been cleaned. In this regard, using the error information contained in the wrong prediction, the statistical number of commodity data under each original classification label that is wrongly predicted as other original classification labels can be determined, and this statistical number represents the number of inter-class mispredictions. It is not difficult to understand that the larger the statistical number of an original classification label mapped to a predicted classification label, the greater the inter-class confusion between the original classification label and the predicted classification label. Therefore, it is best to clean up the commodity data under this original classification label.

[0087] In the present application, all original classification labels can be mapped to the statistical quantities of all predicted classification labels respectively, and the original classification labels corresponding to the data with the largest statistical quantity are selected, and these original classification labels are determined as the target classifications that need data cleaning.

[0088] Then, we can go back to the training set, determine the product data corresponding to these target categories, that is, the original classification labels determined according to the statistical quantity, extract these product data from the training set to construct a data set to be cleaned for data cleaning.

[0089] Step S1400: cluster the commodity data in the data set to be cleaned, delete the commodity data whose clustering results are inconsistent with the original classification labels from the training set, and obtain a clean training set consisting of the remaining commodity data:

[0090] After determining the data set to be cleaned, the commodity data therein can be clustered with the aid of any known clustering algorithm, including but not limited to k-means clustering algorithm, mean shift clustering algorithm, DBSCAN clustering algorithm, expectation maximization (EM) clustering using Gaussian mixture model (GMM), hierarchical clustering algorithm, spectral clustering algorithm, etc. When clustering specifically, it can be performed based on some specific data in the commodity data, for example, clustering can be performed based on the commodity title or commodity attributes, or based on the deep semantic information of one or more specific data, and those skilled in the art can flexibly implement this.

[0091] The principle of clustering the data set to be cleaned in the present application is to calculate the similarity information, or distance information, between the commodity data at each location, and thereby regard the commodity data with a close distance or relatively similar distance as the same category of data, and regard the commodity data with a far distance or relatively dissimilar distance as different category of data, so as to implement data cleaning on this basis.

[0092] After clustering is completed, some product data that are considered to be different types of data are deleted, thereby obtaining a purified training set. It is not difficult to understand that when the data of the purified training set is used to train the instance of the classifier, the classification accuracy of the corresponding classifier instance can be improved.

[0093] Through this typical embodiment, it can be seen that the present application has many advantages, such as:

[0094] First, when cleaning the commodity data, the present application adopts the classification results provided by the classifier, and determines the commodity data that needs to be cleaned from the training set as the data set to be cleaned according to the error information in the classification results. Then, based on the data set to be cleaned, the commodity data that needs to be deleted is determined with the help of clustering means. Accordingly, excessive cleaning of the commodity data in the training set can be avoided, thereby making the data cleaning operation more effective and accurate.

[0095] Secondly, before performing data cleaning, the present application determines the prediction error information through the predicted classification results of the classifier, and then determines some of the original classification labels based on the error information, and only performs clustering and data cleaning on the product data corresponding to these original classification labels. It is not difficult to understand that, relatively speaking, only some categories of product data in the training set are cleaned, and there is no need to clean the entire set. Therefore, it can not only reduce the computational pressure of clustering based on the full data of the entire training set, but also significantly improve the data cleaning speed.

[0096] In addition, the purified training set determined in the present application is suitable for more efficiently training new classifier instances, which can facilitate faster convergence of the training process of the classifier instances and improve the prediction accuracy of the classifier.

[0097] See also Figure 2 In a further embodiment, the step S1200, using the classifier trained by the training set to classify the deep semantic information of the commodity data in the test set to obtain the predicted classification label corresponding to each commodity data, includes the following steps:

[0098] Step S1210: Using a feature extraction model to obtain deep semantic information of the commodity data in the test set, the deep semantic information includes text feature information and / or image feature information of the commodity data:

[0099] Please combine Figure 3 The network model shown is the neural network model used in this application. The neural network model includes a feature extraction model and a classifier, wherein the feature extraction model is responsible for extracting corresponding deep semantic information from the commodity data, and the classifier is responsible for performing classification prediction based on the deep semantic information.

[0100] For a more specific embodiment, please refer to Figure 4 The process shown and Figure 5 The specific neural network model shown, according to the process, step S1210 may include the following steps:

[0101] Step S1211: Use the first feature extraction model to extract deep semantic information from the product titles of the product data in the test set to obtain corresponding text feature information:

[0102] The first feature extraction model is a pre-trained text feature extraction model, and excellent existing models such as Bert, Electra, and Albert can be used. Such models are suitable for representation learning of text information. Therefore, the product title in the product data can be used as input, and the deep semantic information therein can be extracted by such models to obtain the corresponding text feature information, which is usually expressed in vector form. In this embodiment, since the classifier has been pre-trained, the product data in the test set of this application is used to extract deep semantic information.

[0103] Step S1212: Use the second feature extraction model to extract deep semantic information from the product images in the product data in the test set to obtain corresponding image feature information:

[0104] The second feature extraction model is a pre-trained image feature extraction model, which can be implemented using models with relatively good performance in the prior art, such as the Resnet series model, the EfficientNet model, etc. Such models are suitable for image representation learning. Therefore, the product images in the product data, usually the default images, can be used as input, and such models can extract the deep semantic information therein to obtain the corresponding image feature information, which is also usually represented in vector form.

[0105] Step S1213: splice the text feature information and the image feature information of the same product data into image and text feature information representing the deep semantic information of the product data:

[0106] like Figure 5 As shown in the figure, the text feature information and image feature information obtained by the two feature extraction models are further spliced ​​into a text-image feature information, which integrates the deep semantic information of the two types of data, the text and the image, and is more helpful to improve the reliability of the classifier's judgment and classification. Especially for the e-commerce field, the product title fails to effectively describe the product appearance, and the product image is often difficult to summarize the characteristics of the product itself. Therefore, the text-image feature information obtained by combining the two types of data has rich semantic representation capabilities and is more helpful for the classifier's judgment.

[0107] It should be pointed out that in other alternative embodiments, if the text feature extraction model or the image feature extraction model is used alone, and the text feature information or the image feature information is used alone for classification, it is also feasible and does not affect the embodiment of the creative spirit of the present application.

[0108] Step S1220: Use a classifier to calculate the classification probability corresponding to the mapping of the deep semantic information to each of the original classification labels, so as to determine the predicted classification label to which the commodity data corresponding to the deep semantic information belongs, wherein the predicted classification label is the original classification label with the largest classification probability for the commodity data:

[0109] After the commodity data has obtained deep semantic information through the feature extraction model, it can be classified by the classifier of this application. The classifier maps the result of fully connecting the vector corresponding to the deep semantic information to the classification space, calculates the classification probability corresponding to each classification label in the classification space, and then determines the classification label with the largest classification probability as the predicted classification label of the model. The classification space is composed of the original classification labels carried in the commodity data set. Therefore, the number of original classification labels in the commodity data set corresponds to the number of predicted classification labels in the classification space of the classifier.

[0110] The embodiments and their alternative embodiments given here use a test set to test a classifier that has been pre-trained using a training set, so as to facilitate the subsequent acquisition of its corresponding prediction error information. In addition, these embodiments also reveal the network structure of the neural network model in which the classifier of the present application is located. The introduction of such a network structure embodies the reliability of the classifier of the present application in performing classification tasks. It can make classification judgments through the deep semantic information of the commodity data. Therefore, it is easier to guide the technical personnel of the present application to refer to and implement it, thereby enhancing the practicality of the technical solution of the present application.

[0111] See also Figure 6 In a further embodiment, the step S1300, according to the error information between the predicted classification label and the original classification label of the commodity data, determines a plurality of original classification labels with a predetermined number of the most wrong predictions, and extracts the commodity data under these original classification labels from the training set as the data set to be cleaned, including the following steps:

[0112] Step S1310: A confusion matrix is ​​calculated based on the predicted classification labels of the commodity data in the test set and their original classification labels. Each element in the confusion matrix is ​​used to represent the statistical number of commodity data corresponding to the original classification label that is predicted to be a certain predicted classification label:

[0113] In order to conveniently determine the error information after the classifier classifies the commodity data in the test set, it is necessary to collect statistics on the distribution of the classification of the commodity data in the test set to each predicted classification label. In order to facilitate data presentation, a confusion matrix can be used to represent the statistical data. Specifically, Figure 7 As shown, the row coordinates of the confusion matrix represent each original classification label, and the column coordinates of the confusion matrix represent each predicted classification label, wherein each row of data represents the statistical number of all commodity data in the test set corresponding to an original classification label mapped to each predicted classification label. Based on this, for each full commodity data sample corresponding to an original classification label, it can be known that, except for the elements mapped to the predicted classification label that is the same as the original classification label, the remaining elements all represent the statistical number of commodity data in the full commodity data sample that is incorrectly predicted to each other classification label, and correspondingly represent the number of inter-class error mappings. Since the original classification label and the predicted classification label are essentially composed of classification labels in the same classification space, the confusion matrix is ​​a square matrix, and the diagonal from the upper left corner to the lower right corner is the element that the original classification label is consistent with the predicted classification label, indicating that the classifier has correctly classified and mapped the commodity data corresponding to the statistical number of these elements.

[0114] Step S1320: From the elements in the confusion matrix whose original classification labels are inconsistent with the predicted classification labels, for a predetermined number of target elements whose statistical number is the largest, determine the original classification labels corresponding to these target elements:

[0115] In the confusion matrix, each original classification label generally has multiple inter-class error mappings. Therefore, all original classification labels have the same situation. For each original classification label, among the several elements with inter-class error mappings, the element with the largest statistical number indicates that the commodity data corresponding to the original classification label is easy to confuse the classifier, and is easily determined by the classifier as the predicted classification label corresponding to the element. Therefore, the commodity data under the original classification labels corresponding to this part of the statistical number should be cleaned first.

[0116] In order to determine the part of the commodity data that needs to be cleaned, the Top_K algorithm can be used to select the K target elements with the largest statistical number from the entire confusion matrix, except for the diagonal elements from the upper left corner to the lower right corner. The K value here can be flexibly set by those skilled in the art, and the subsequent embodiments will also reveal that the value can be adjusted multiple times to optimize the cleaning effect. The Top_K algorithm is used to determine some but not all of the target elements that are incorrectly mapped, mainly to avoid excessive cleaning of commodity data and improve data cleaning efficiency.

[0117] After determining the K target elements with the largest statistical quantity, the original classification labels corresponding to these target elements can be determined, so as to determine the data set to be cleaned according to these original classification labels.

[0118] Step S1330: extract commodity data carrying the original classification labels corresponding to the target elements from the training set to form a data set to be cleaned:

[0119] To construct the data set to be cleaned, the corresponding commodity data can be obtained from the training set, specifically, the commodity data carrying the original classification label is extracted from the training set to construct the data set to be cleaned. It is not difficult to understand that the commodity data in the data set to be cleaned is the main commodity data that is easily confused by the classifier and causes incorrect classification mapping.

[0120] This embodiment uses a confusion matrix to count the classification distribution of the commodity data in the test set after being classified by the classifier, and then selects some easily confused classification labels according to the statistical quantity obtained by the statistics, and determines the corresponding commodity data from the training set according to these classification labels to construct a data set to be cleaned. This can effectively avoid excessive cleaning of the commodity data in the training set, thereby avoiding affecting the validity of the commodity data. At the same time, it also helps to avoid full cleaning of the commodity data in the training set, which can reduce the computing pressure of computer equipment.

[0121] See also Figure 7 In a further embodiment, the step S1400, clustering the commodity data in the data set to be cleaned, deleting commodity data whose clustering results are inconsistent with the original classification labels from the training set, and obtaining a cleansed training set consisting of the remaining commodity data, includes the following steps:

[0122] Step S1410: using a preset clustering algorithm, clustering is performed according to the deep semantic information of the commodity data in the data set to be cleaned, and a plurality of commodity data clusters obtained by clustering are obtained:

[0123] As mentioned above, the present application can use a variety of known clustering algorithms to cluster the commodity data in the cleaned data set. For example, K-means can be used to implement supervised clustering, and the number of clustered categories is the same as the number of original classification labels corresponding to the target elements determined from the confusion matrix.

[0124] When clustering, this embodiment calculates distance information based on the deep semantic information of the commodity data in the data set to be cleaned as the input of the K-means algorithm, thereby achieving clustering. The deep semantic information is also the deep semantic information of the commodity data used by the classifier of this application for classification, such as the text feature information, image feature information or graphic feature information described in the aforementioned embodiments. Therefore, in the process of representing training and testing of the training set and the test set, the deep semantic information extracted by the feature extraction model can also be stored, and the corresponding call can be directly made in this step.

[0125] It can be understood that after the commodity data in the data to be cleaned is clustered by the clustering algorithm, multiple commodity data clusters will be formed, and each commodity data cluster contains multiple.

[0126] Step S1420: Count the maximum number of original classification labels owned by the commodity data in each commodity data cluster, and delete the commodity data in the commodity data cluster that does not carry the maximum number of original classification labels to achieve data cleaning of the commodity data cluster;

[0127] The clustering standard between the commodity data in the commodity data cluster is determined by the clustering algorithm according to its own statistical logic of the distance information between the data. Therefore, most of the commodity data will have the same original classification label, while a small number of commodity data will have original classification labels that are different from the former and may be more discrete. Based on this, the commodity data with the same original classification label in each commodity data cluster can be retained, and the commodity data with other original classification labels can be deleted, thereby achieving data cleaning of the commodity data in each commodity data cluster.

[0128] Step S1430: determine a cleansed training set, where the cleansed training set includes the commodity data in each commodity data cluster that has completed data cleaning:

[0129] Since the commodity data in each commodity data cluster are all commodity data carrying the same original classification label, purification is achieved. Therefore, the commodity data in all commodity data clusters can be reconstructed into a training set, which is a purified training set.

[0130] The purified training set can be used to train another instance of the classifier. In theory, the accuracy of the classifier obtained by the training can be improved in classifying and judging commodity data.

[0131] In this embodiment, the data to be cleaned is cleaned with reference to the clustering information provided by the clustering algorithm, the product data that cannot be correctly clustered is deleted, and a new training set is constructed from the remaining product data. The product data in the training set has been purified and can be used as an effective training set for classifier instance training, which can significantly improve the classification accuracy implemented by the classifier.

[0132] This embodiment performs clustering based on deep semantic information of commodity data. The deep semantic information is the result of deep representation learning of the semantic information of commodity data. Therefore, a better clustering effect can be achieved.

[0133] See also Figure 8 In the extended embodiment, after the step S1400 of clustering the commodity data in the cleaned data set, deleting the commodity data whose clustering results are inconsistent with the original classification labels from the training set, and obtaining the cleaned training set consisting of the remaining commodity data, the following steps are included:

[0134] Step S1500: Use the purified training set as a new training set to train the classifier:

[0135] In this embodiment, by applying a loop mechanism, the purified training set obtained in the above-mentioned embodiments can be repeatedly optimized multiple times, so that the classification information of the commodity data in the final purified training set is more accurate. Specifically, the purified training set can be first used as a new training set to train the instance of the classifier, so that the classifier can theoretically improve the classification accuracy. Afterwards, the process of steps S1200 to S1400 of this application is re-executed to iteratively obtain the purified training set corresponding to the further cleaned commodity data. In this process, in the step S1300, specifically in the step of determining the presence of a predetermined number of multiple original classification labels with the most false predictions, the predetermined number is reset. For example, when performing data cleaning for the first time, the Top_K algorithm is used to determine K original classification labels that need to perform data cleaning, and the corresponding to the first data cleaning The data set to be cleaned is constructed accordingly. When performing data cleaning for the second time, when the Top_K algorithm is used, the value of K can be adjusted to a smaller value, thereby determining a smaller standardized data set to be cleaned, and so on, the training set is continuously optimized.

[0136] It can be understood that through multiple such iterative optimization of the training set, the purified training set finally obtained will become more and more accurate and efficient for classifier training.

[0137] See also Fig. 9 In a further embodiment, the commodity data cleaning method of the present application further includes the following subsequent steps:

[0138] Step S1600: Count the prediction accuracy of the classifier obtained from each training, where the prediction accuracy is the accuracy of the deep semantic information of the commodity data in the test set being correctly classified by the classifier into its original classification label:

[0139] According to the process disclosed in the previous embodiment, in the process of multiple iterations of optimizing the training set, multiple classifier instances trained with purified training sets of different specifications will be derived, and the training of these classifier instances will naturally be completed. These classifiers can be selected and used in the production stage to classify commodity data. However, the prediction accuracy of each classifier may vary, and it does not entirely depend on the number of data cleaning times of the training set. Therefore, it is necessary to select the best classifier for the production stage by examining the prediction accuracy of each classifier instance.

[0140] In order to count the prediction accuracy of the classifier (instance) obtained from each training, each classifier instance can still be tested with the test set so as to unify the standards based on each classifier instance. Similarly, after testing the test set for each classifier instance, a corresponding confusion matrix can be obtained. The concept of confusion matrix and its implementation can refer to the description of the embodiments of the foregoing application.

[0141] Since each row of data in the confusion matrix represents the statistical number of commodity data for which an original classification label is mapped to all classification labels in the classification space, the prediction accuracy for the original classification label can be determined by using the ratio of the statistical number of correct predicted classification labels to the sum of all statistical numbers in the row. By normalizing and summing the prediction accuracy corresponding to each original classification label, the overall prediction accuracy of the corresponding classifier instance can be determined.

[0142] Step S1700: determine that the training set used by the classifier with the highest prediction accuracy is the optimal training set, and use the classifier to respond to the classification requirements of the commodity data and determine the corresponding classification of the commodity data.

[0143] On the basis of determining the overall prediction accuracy corresponding to each classifier instance, the overall prediction accuracy of each classifier instance is compared with each other, and the classifier instance with the highest prediction accuracy can be determined as the optimal classifier. The training set (purified training set) used when training the optimal classifier is the optimal training set. It can be understood that this optimal training set has not only been used to train the classifier instance, but can also be used in other classification tasks. The optimal classifier is theoretically sufficient for use in the production stage to respond to the classification needs of commodity data and determine the corresponding classification of commodity data.

[0144] This embodiment realizes continuous optimization of the training set through an iterative mechanism. For each version of the purified training set generated during the optimization process, the prediction accuracy of the classifiers trained therefrom is also used for evaluation, and the best classifier instance and the best purified training set are evaluated, and finally the best classifier and the best purified training set are generated and can be put into practical use. This further effectively avoids the situation where the training set is over-cleaned, and in the process of cleaning the training set, a classifier with the highest prediction accuracy is also produced, achieving the best of both worlds.

[0145] See also Fig.10 A commodity data cleaning device provided to meet one of the purposes of the present application comprises: a data determination module 1100, a classification prediction module 1200, a false positive screening module 1300, and a cluster cleaning module 1400, wherein the data determination module 1100 is used to determine a commodity data set, the commodity data set comprises a plurality of commodity data carrying original classification labels, and the commodity data is divided into a training set and a data set; the classification prediction module 1200 is used to classify the deep semantic information of the commodity data in the test set by using a classifier trained by the training set to obtain a predicted classification label corresponding to each commodity data; the false positive screening module 1300 is used to determine a plurality of original classification labels with a predetermined number of the most false predictions according to error information between the predicted classification label and the original classification label of the commodity data, and extract the commodity data under these original classification labels from the training set as the data set to be cleaned; the cluster cleaning module 1400 is used to cluster the commodity data in the data set to be cleaned, delete the commodity data whose clustering results are inconsistent with the original classification labels from the training set, and obtain a purified training set consisting of the remaining commodity data.

[0146] In a further embodiment, the classification prediction module 1200 includes: a semantic extraction submodule, which is used to use a feature extraction model to obtain deep semantic information of the product data in the test set, and the deep semantic information includes text feature information and / or image feature information of the product data; a classification mapping submodule, which is used to use a classifier to calculate the classification probability corresponding to each of the original classification labels mapped to the deep semantic information, so as to determine the predicted classification label to which the product data corresponding to the deep semantic information belongs, and the predicted classification label is the original classification label with the largest classification probability for the product data.

[0147] In a specific embodiment, the semantic extraction submodule includes: a text extraction unit, which is used to use a first feature extraction model to extract deep semantic information from the product titles of the product data in the test set to obtain corresponding text feature information; an image extraction unit, which is used to use a second feature extraction model to extract deep semantic information from the product images of the product data in the test set to obtain corresponding image feature information; and a feature fusion unit, which is used to splice the text feature information and image feature information of the same product data into graphic feature information representing the deep semantic information of the product data.

[0148] In a further embodiment, the false positive screening module 1300 includes: a false positive statistics submodule, which is used to calculate a confusion matrix based on the predicted classification labels and the original classification labels of the commodity data in the test set, and each element in the confusion matrix is ​​used to represent the statistical number of commodity data corresponding to the original classification label being predicted to be a certain predicted classification label; a false positive determination submodule, which is used to determine the original classification labels corresponding to a predetermined number of target elements with the largest statistical number from the elements in the confusion matrix whose original classification labels are inconsistent with the predicted classification labels; and a to-be-cleaned extraction submodule, which is used to extract commodity data carrying the original classification labels corresponding to the target elements from the training set to form a data set to be cleaned.

[0149] In a further embodiment, the cluster cleaning module 1400 includes: a title clustering submodule, which is used to adopt a preset clustering algorithm to cluster the product data in the data set to be cleaned according to the deep semantic information, and obtain multiple product data clusters obtained by clustering; a cleaning execution submodule, which is used to count the maximum number of original classification labels possessed by the product data in each product data cluster, and delete the product data in the product data cluster that does not carry the maximum number of original classification labels to achieve data cleaning of the product data cluster; a purification data submodule, which is used to determine a purification training set, which includes the product data in each product data cluster that has completed data cleaning.

[0150] In an extended embodiment, the commodity data cleaning device of the present application includes: a re-training module, which is used to use the purified training set as a new training set to train the classifier. After the re-training module is run, the classification prediction module 1200, the false detection screening module 1300, and the cluster cleaning module 1400 of the present device are triggered to run in a cycle again, wherein the predetermined number in the false detection screening module 1300 is reset.

[0151] In a further embodiment, the commodity data cleaning device of the present application also includes the following subsequent operating structure: an accuracy statistics module, which is used to count the prediction accuracy of the classifier obtained from each training, and the prediction accuracy is the accuracy of the deep semantic information of the commodity data in the test set being correctly classified by the classifier to its original classification label; a selection and determination module, which is used to determine that the training set used by the classifier with the highest prediction accuracy is the optimal training set, and use the classifier to respond to the classification requirements of the commodity data and determine the corresponding classification of the commodity data.

[0152] In order to solve the above technical problems, the present application also provides a computer device. Fig.11 As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database may store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a commodity data cleaning method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the commodity data cleaning method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0153] In this embodiment, the processor is used to execute Fig.10 The memory stores the program code and various data required to execute the above modules or submodules. The network interface is used to transmit data between user terminals or servers. The memory in this embodiment stores the program code and data required to execute all modules / submodules in the commodity data cleaning device of this application, and the server can call the program code and data of the server to execute the functions of all submodules.

[0154] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the commodity data cleaning method of any embodiment of the present application.

[0155] The present application also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any embodiment of the present application when executed by one or more processors.

[0156] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments of the present application can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0157] In summary, the present application can effectively clean massive amounts of commodity data to obtain high-quality training sets, so that the classifiers for classification tasks trained by the application can obtain classification prediction effects with higher accuracy.

[0158] Those skilled in the art will appreciate that the various operations, methods, steps, measures, and schemes in the processes discussed in this application may be alternated, altered, combined, or deleted. Further, other steps, measures, and schemes in the various operations, methods, and processes discussed in this application may also be alternated, altered, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and schemes in the prior art that are similar to those disclosed in this application may also be alternated, altered, rearranged, decomposed, combined, or deleted.

[0159] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A commodity data cleaning method, characterized in that: The steps include: Determine a commodity data set, wherein the commodity data set includes a plurality of commodity data carrying original classification labels, and the commodity data is divided into a training set and a test set; The deep semantic information of the commodity data in the test set is classified by using a classifier trained by the training set to obtain a predicted classification label corresponding to each commodity data, including: using a feature extraction model to obtain the deep semantic information of the commodity data in the test set, the deep semantic information includes text feature information and / or image feature information of the commodity data; using a classifier to calculate the classification probability corresponding to each of the original classification labels mapped to the deep semantic information, so as to determine the predicted classification label to which the commodity data corresponding to the deep semantic information belongs, the predicted classification label being the original classification label with the largest classification probability for the commodity data; According to the error information between the predicted classification label and the original classification label of the commodity data, a plurality of original classification labels with a predetermined number of the most incorrect predictions are determined, and the commodity data under these original classification labels are extracted from the training set as the data set to be cleaned; The product data in the cleansing data set are clustered, and the product data whose clustering results are inconsistent with the original classification labels are deleted from the training set to obtain a purified training set consisting of the remaining product data.

2. The commodity data cleaning method according to claim 1, characterized in that: Using a feature extraction model to obtain deep semantic information of the commodity data in the test set, the deep semantic information includes text feature information and / or image feature information of the commodity data, including the following steps: Using a first feature extraction model to extract deep semantic information from the product titles of the product data in the test set to obtain corresponding text feature information; Using a second feature extraction model to extract deep semantic information from the product images of the product data in the test set to obtain corresponding image feature information; The text feature information and the image feature information of the same product data are spliced ​​into graphic feature information representing the deep semantic information of the product data.

3. The commodity data cleaning method according to claim 1, characterized in that: According to the error information between the predicted classification label and the original classification label of the commodity data, a plurality of original classification labels with a predetermined number of the most incorrect predictions are determined, and the commodity data under the original classification labels are extracted from the training set as the data set to be cleaned, including the following steps: A confusion matrix is ​​calculated based on the predicted classification labels of the commodity data in the test set and their original classification labels, wherein each element in the confusion matrix is ​​used to represent the statistical number of commodity data corresponding to the original classification label that are predicted to be a certain predicted classification label; From the elements in the confusion matrix whose original classification labels are inconsistent with the predicted classification labels, for a predetermined number of target elements with the largest statistical number, determine the original classification labels corresponding to these target elements; The commodity data carrying the original classification labels corresponding to the target elements are extracted from the training set to form a data set to be cleaned.

4. The commodity data cleaning method according to claim 1, characterized in that: Clustering the commodity data in the cleansing data set, deleting commodity data whose clustering results are inconsistent with the original classification labels from the training set, and obtaining a purified training set consisting of the remaining commodity data, including the following steps: Using a preset clustering algorithm, clustering is performed according to the deep semantic information of the commodity data in the data set to be cleaned, to obtain a plurality of commodity data clusters obtained by clustering; Counting the maximum number of original classification labels possessed by the commodity data in each commodity data cluster, and deleting the commodity data in the commodity data cluster that does not carry the maximum number of original classification labels to achieve data cleaning of the commodity data cluster; A cleansed training set is determined, where the cleansed training set includes commodity data in each commodity data cluster that has completed data cleaning.

5. The commodity data cleaning method according to any one of claims 1 to 4, characterized in that: After clustering the commodity data in the cleansing data set, deleting commodity data whose clustering results are inconsistent with the original classification labels from the training set, and obtaining the purification training set consisting of the remaining commodity data, the following steps are included: Using the purified training set as a new training set to train the classifier; After the classifier is trained, the method is executed in a loop starting from the step of using the classifier trained by the training set to classify the deep semantic information of the commodity data in the test set; Wherein, in the step of determining a predetermined number of original classification labels having the most erroneous predictions, the predetermined number is reset.

6. The commodity data cleaning method according to claim 5, characterized in that: The method also includes the following subsequent steps: Counting the prediction accuracy of the classifier obtained from each training, where the prediction accuracy is the accuracy rate at which the deep semantic information of the commodity data in the test set is correctly classified by the classifier into its original classification label; The training set used by the classifier with the highest prediction accuracy is determined to be the optimal training set, and the classifier is used to respond to the classification requirements of the commodity data and determine the corresponding classification of the commodity data.

7. A computer device comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 6 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and system for label-free data classification and predication

    CN107679734A

  • Data cleaning method and device, storage medium and electronic equipment

    CN113342792A