Data classification method, device, equipment, readable storage medium and product
Patent Information
- Application Number
- CN202211307793.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-10-25
AI Technical Summary
[0017]本公开的第四个方面是提供一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,所述计算机执行指令被处理器执行时用于实现如第一方面所述的方法。
Smart Images

Figure CN116150614B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and more particularly to a data classification method, apparatus, device, readable storage medium, and product. Background Technology
[0002] With the development of technology, neural network models are increasingly being applied in various technical fields. To obtain a neural network model, it is first necessary to use pre-defined training, testing, and validation sets to train the model. However, if the pre-defined model fits the existing data too closely, overfitting will occur. Therefore, it is necessary to properly classify the existing data.
[0003] Existing data classification methods typically divide a dataset D into k mutually exclusive subsets of similar size. Each subset maintains the consistency of the data distribution as much as possible, i.e., it is obtained from D through stratified sampling. Each time, the union of k-1 subsets is used as the training set, and the remaining subset is used as the test set; in this way, k groups of classifications can be obtained, k training and testing cycles can be performed, and finally the mean of the k evaluation results is returned.
[0004] However, classifying datasets using the above method often requires K training iterations of the model, resulting in low data classification efficiency and consequently affecting the model's training efficiency. Summary of the Invention
[0005] This disclosure provides a data classification method, apparatus, device, readable storage medium, and product to address the technical problem that existing data classification methods are inefficient and affect the training efficiency of models.
[0006] The first aspect of this disclosure is to provide a data classification method, including:
[0007] Obtain the dataset to be classified, wherein the dataset to be classified includes multiple data to be classified, the target entropy value and the predicted category corresponding to each data to be classified, and the labeled category corresponding to some data to be classified;
[0008] Identify multiple target classification data in the dataset to be classified, wherein the target classification data are the data to be classified whose predicted category is different from the labeled category;
[0009] Based on the target entropy value corresponding to each target classification data, perform iterative classification operations on multiple target classification data until the preset iteration termination condition is met, and obtain the classified target training set, target test set, and target validation set.
[0010] A second aspect of this disclosure is to provide a data classification apparatus, comprising:
[0011] The acquisition module is used to acquire the dataset to be classified, wherein the dataset to be classified includes multiple data to be classified, the target entropy value and the predicted category corresponding to each data to be classified, and the labeled category corresponding to some data to be classified;
[0012] A determination module is used to determine multiple target classification data in the dataset to be classified, wherein the target classification data are the data to be classified whose predicted category is different from the labeled category;
[0013] The classification module is used to perform iterative classification operations on multiple target classification data according to the target entropy value corresponding to each target classification data until the preset iteration termination condition is met, so as to obtain the classified target training set, target test set and target validation set.
[0014] A third aspect of this disclosure is to provide an electronic device, including: a memory and a processor;
[0015] Memory; memory for storing executable instructions of the processor;
[0016] The processor is used to invoke program instructions in the memory to execute the method as described in the first aspect.
[0017] A fourth aspect of this disclosure is to provide a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method as described in the first aspect.
[0018] A fifth aspect of this disclosure is to provide a computer program product including a computer program that, when executed by a processor, implements the method as described in the first aspect.
[0019] The data classification method, apparatus, device, readable storage medium, and product disclosed herein, after acquiring a dataset to be classified, identify target classification data in the dataset whose predicted category differs from the labeled category, and perform iterative classification operations on multiple target classification data based on the target entropy value corresponding to the target classification data. Iterative classification only on the target classification data effectively reduces the computational load in the dataset classification process and improves the efficiency of dataset classification. Furthermore, by using the target entropy value to determine whether each data point to be classified needs iterative classification, the accuracy of dataset classification can be improved, thereby enhancing the training accuracy of subsequent classification models. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings.
[0021] Figure 1 This is a schematic diagram of the system architecture upon which this disclosure is based;
[0022] Figure 2 A flowchart illustrating the data classification method provided in this embodiment of the disclosure;
[0023] Figure 3 A flowchart illustrating a data classification method provided in yet another embodiment of this disclosure;
[0024] Figure 4 A flowchart illustrating a data classification method provided in yet another embodiment of this disclosure;
[0025] Figure 5 A schematic diagram of the structure of the data classification device provided in the embodiments of this disclosure;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. All other embodiments obtained based on the embodiments of this disclosure are within the scope of protection of this disclosure.
[0028] In view of the technical problems mentioned above, such as the low efficiency of existing data classification methods and their impact on model training efficiency, this disclosure provides a data classification method, apparatus, device, readable storage medium, and product.
[0029] It should be noted that the data classification methods, apparatus, devices, readable storage media, and products disclosed herein can be used in any scenario of dataset classification during model training.
[0030] Existing technology provides a data classification method in a federated learning scenario, which can determine whether the data distribution of the original data provided by each federated learning participant is consistent; train a federated classification model using the original data provided by each federated learning participant and model test data; input the original data belonging to each federated learning participant into the federated classification model, and the input data of the federated classification model is the probability of the model test data; select a specified number of model input data according to the predicted probability from high to low as the validation set provided by the federated learning participant to verify the model performance, and use the remaining model input data as the training set provided by the federated learning participant to train the model. This invention can find data samples with the most similar data distribution to the test dataset in the data provided by each federated learning participant as the validation set for model training. By cooperating with multiple federated classification models and their respective original data, multiple models are trained to obtain the probabilities of each test data, and the dataset is classified according to the probability results. However, in a fast iteration scenario, training multiple models is time-consuming and labor-intensive, which is not conducive to the rapid iteration of models and data.
[0031] In addressing the aforementioned technical problems, the inventors discovered through research that a classification model can be pre-trained, and a dataset containing multiple data points to be classified, along with labeled categories corresponding to at least some of these data points, can be pre-built. The data to be classified is then input into the classification model, and the predicted category and target entropy value corresponding to each data point in the dataset are identified based on the model's output. This enables iterative classification of target data whose predicted category differs from the labeled category, eliminating the need for multiple training iterations of the classification model. Iterative classification of only the target data effectively reduces the computational load during dataset classification, thus improving classification efficiency.
[0032] Figure 1 This is a schematic diagram of the system architecture upon which this disclosure is based, such as Figure 1 As shown, the system architecture upon which this disclosure is based includes at least: a data server 11 and a server 12, wherein the data server 11 can be a cloud server or a server cluster, and stores a large amount of data. The server 12 may be equipped with a data classification device, which can be written in languages such as C / C++, Java, Shell or Python.
[0033] Based on the above system architecture, server 12 can obtain the dataset to be classified from data server 11 and determine the target classification data in the dataset whose predicted category differs from the labeled category. This allows iterative classification of the target classification data based on the target entropy value corresponding to the target classification data.
[0034] Figure 2This is a flowchart illustrating the data classification method provided in the embodiments of this disclosure, as shown below. Figure 2 As shown, the method includes:
[0035] Step 201: Obtain the dataset to be classified, wherein the dataset to be classified includes multiple data to be classified, the target entropy value and predicted category corresponding to each data to be classified, and the labeled category corresponding to some data to be classified.
[0036] In this embodiment, the execution entity is a data classification device, which can be coupled to a server. This server can communicate with a data server storing the dataset to be classified, thereby enabling the acquisition of the dataset. Then, it can accurately classify the data within the dataset.
[0037] In this embodiment, to obtain accurate target training sets, target test sets, and target validation sets before training the classification model, it is first necessary to acquire the dataset to be classified. This dataset includes multiple data points to be classified, a target entropy value and predicted category for each data point, and labeled categories for some of the data points. The target entropy value and predicted category for each data point are calculated based on the probability that the data points belong to each labeled category after inputting multiple data points into a preset classification model. The labeled categories are obtained by pre-labeling the data according to a preset labeling method.
[0038] Step 202: Determine multiple target classification data in the dataset to be classified, wherein the target classification data are the data to be classified whose predicted category is different from the labeled category.
[0039] In this embodiment, if the predicted category is different from the labeled category corresponding to the data to be classified, it indicates that the current classification of the data to be classified is inaccurate and needs to be reclassified. Therefore, after obtaining the dataset to be classified, the data to be classified whose predicted category is different from the labeled category can be filtered to obtain the target classification data, and then reclassification can be performed only on the target classification data.
[0040] Step 203: Perform iterative classification operations on multiple target classification data according to the target entropy value corresponding to each target classification data until the preset iteration termination condition is met, and obtain the classified target training set, target test set and target validation set.
[0041] In this embodiment, the target entropy value can accurately characterize whether the predicted category obtained by the classification model is accurate. Optionally, the smaller the target entropy value, the more accurate the predicted category; conversely, the larger the target entropy value, the less accurate the predicted category. Therefore, in order to perform classification operations on target classification data, iterative classification operations can be performed on multiple target classification data based on the target entropy value corresponding to each target classification data.
[0042] Optionally, different processing methods can be used for different target entropy values corresponding to the target classification data to perform iterative classification operations on the target classification data until the preset iteration termination condition is met, so as to obtain the classified target training set, target test set and target validation set.
[0043] It should be noted that the preset iteration termination conditions include, but are not limited to, the probability that the predicted category output by the classification model matches the labeled category corresponding to the target classification data exceeds a preset probability threshold; or, the number of iterations exceeds a preset number threshold; or, the iteration time exceeds a preset time threshold; or, the target entropy value corresponding to the target classification data is lower than a preset entropy value threshold, etc.
[0044] In practical applications, any of the above-mentioned iteration termination conditions can be used for iterative classification of the dataset to be classified. Alternatively, the iteration termination conditions can be adjusted according to actual needs, and this disclosure does not impose any restrictions on this.
[0045] The data classification method, apparatus, device, readable storage medium, and product disclosed herein, after acquiring a dataset to be classified, identify target classification data in the dataset whose predicted category differs from the labeled category, and perform iterative classification operations on multiple target classification data based on the target entropy value corresponding to the target classification data. Iterative classification only on the target classification data effectively reduces the computational load in the dataset classification process and improves the efficiency of dataset classification. Furthermore, by using the target entropy value to determine whether each data point to be classified needs iterative classification, the accuracy of dataset classification can be improved, thereby enhancing the training accuracy of subsequent classification models.
[0046] Figure 3 This is a flowchart illustrating a data classification method provided in yet another embodiment of the present disclosure. Based on any of the above embodiments, such as... Figure 3 As shown, step 201 includes:
[0047] Step 301: Obtain the original dataset, which includes multiple data to be classified and at least some of the labeled categories corresponding to the data to be classified.
[0048] Step 302: Input the data to be classified in the original dataset into a preset classification model to obtain the output data of the classification model. The output data includes the probability that the data to be classified belongs to each labeled category. The classification model is obtained by pre-training the preset training model using the initial training set, initial test set and initial validation set obtained by pre-classifying the dataset to be classified.
[0049] Step 303: Determine the target entropy value and predicted category corresponding to each data to be classified based on the output data, and generate the dataset to be classified based on the target entropy value, predicted category and the labeled category corresponding to at least some of the data to be classified.
[0050] In this embodiment, to perform the classification operation on the data to be classified, an original dataset is first obtained. This original dataset includes multiple data points to be classified, and each data point can be labeled with a category using a preset keyword matching method. Since the preset keyword matching method may not be able to label all the data to be classified, some data points in the original dataset may include corresponding labeled categories, while some data points may not include corresponding labeled categories.
[0051] Furthermore, the data to be classified in the original dataset can be pre-classified to obtain an initial training set, an initial test set, and an initial validation set. These initial training, test, and validation sets are then used to perform preliminary training on a pre-defined model to obtain a classification model capable of predicting the category of the data to be classified. The data to be classified in the original dataset is then input into the pre-defined classification model, and the output data of the model is obtained. Specifically, the output data of the classification model can include the probability that the data to be classified belongs to each category.
[0052] Accordingly, after obtaining the output data from the classification model, the predicted category corresponding to each piece of data to be classified can be determined based on this output data. Furthermore, the target entropy value corresponding to each piece of data to be classified can also be determined based on this output data. This target entropy value accurately characterizes whether the predicted category obtained by the classification model is accurate. Optionally, the smaller the target entropy value, the more accurate the predicted category; conversely, the larger the target entropy value, the less accurate the predicted category. Thus, a dataset to be classified can be generated based on the target entropy value, the predicted category, and at least some of the labeled categories corresponding to the data to be classified.
[0053] The data classification method provided in this embodiment, after obtaining the original dataset, predicts the category of the data to be classified in the dataset using a preset classification model, thereby obtaining the target entropy value and predicted category for each data to be classified. Then, iterative classification operations can be performed on target classification data whose predicted category does not match the labeled category based on the target entropy value. Performing iterative classification only on the target classification data effectively reduces the computational load in the dataset classification process and improves the efficiency of dataset classification.
[0054] Furthermore, based on any of the above embodiments, before step 302, the method further includes:
[0055] Data to be classified is extracted from the original dataset using a preset sampling method, and the data to be classified is added to the preset training set, test set, and validation set according to a preset ratio to obtain the initial training set, initial test set, and initial validation set.
[0056] In this case, each of the data to be classified is sampled once.
[0057] In this embodiment, after obtaining the original dataset, a preliminary classification operation can be performed on the data in the original dataset using a preset sampling method, wherein each data to be classified is sampled once. Considering that some categories may have a small amount of data, in order to make the final prediction result of the model to be trained more realistically reflect the prediction ability, a test set is allocated first, then a validation set, and finally a training set, so as to achieve the purpose of uniform sample distribution.
[0058] Optionally, in order to ensure that each data point to be classified is sampled only once, sampling without replacement can be used to perform the initial classification operation on the original dataset.
[0059] Furthermore, different data volume ratios can be set for the preset training set, test set, and validation set. Therefore, the data to be classified can be added to the preset training set, test set, and validation set according to the preset ratio to obtain the initial training set, initial test set, and initial validation set.
[0060] For example, a sampling method without replacement can be used to select 70% of the data from each category for the training set, 20% for the test set, and 10% for the validation set. Alternatively, this preset ratio can be adjusted according to actual needs, and this disclosure does not impose any restrictions on it.
[0061] The data classification method provided in this embodiment uses sampling without replacement to classify the data to be classified in the original dataset into an initial training set, an initial test set, and an initial validation set according to a preset ratio. This allows for the training of a preset model based on the initial training set, initial test set, and initial validation set, and then enables iterative classification of the data to be classified based on the trained classification model. This provides a foundation for simplifying the dataset classification process and improving the efficiency of dataset classification.
[0062] Furthermore, based on any of the above embodiments, step 203 includes:
[0063] The label category corresponding to the probability value in the output data that meets the preset condition is determined as the prediction category.
[0064] The target entropy value is obtained by calculating the output data using a preset entropy algorithm.
[0065] In this embodiment, the output data of the classification model may specifically include the probability that the data to be classified belongs to each category. Therefore, the labeled category corresponding to the probability value in the output data that meets a preset condition can be determined as the predicted category. For example, the labeled category with the highest probability value can be determined as the predicted category corresponding to the data to be classified, thereby improving the accuracy of the predicted category.
[0066] Furthermore, the target entropy value can be obtained by calculating the output data using a preset entropy algorithm. This preset entropy algorithm can be illustrated in Formula 1:
[0067] S=-∑p(x i log2p(x) i (1)
[0068] Where, x i p(x) represents the true class of sample i. i ) represents the probability values corresponding to each category output by the classification model.
[0069] In log2p(x i In the algorithm, log is a monotonic function. Using log can simplify calculations. Logarithms can convert multiplication into addition and division into subtraction. Differentiation can be performed separately. Therefore, using logarithms in entropy algorithms can greatly simplify calculations and improve the efficiency of calculating the target entropy value.
[0070] The data classification method provided in this embodiment improves the accuracy of predicted categories by determining the labeled categories corresponding to the probabilities that satisfy preset conditions in the output data as predicted categories. This allows for accurate classification of the dataset based on the accurate predicted categories. Furthermore, by calculating the target entropy value, a targeted classification operation can be performed on the data to be classified using appropriate iterative methods based on the target entropy value, further improving the accuracy of dataset classification.
[0071] Furthermore, based on any of the above embodiments, after step 204, the method further includes:
[0072] The classification model is iteratively trained using the target training set, target test set, and target validation set.
[0073] In this embodiment, after completing the iterative classification operation of the data to be classified in the dataset to be classified, and obtaining the target training set, target test set and target validation set after accurate classification, in order to improve the prediction accuracy of the classification model, the classification model can be further iteratively trained using the target training set, target test set and target validation set.
[0074] The data classification method provided in this embodiment improves the accuracy of the classification model's category prediction by accurately classifying the data to be classified in the dataset to be classified, obtaining the target training set, target test set, and target validation set, and then using the target training set, target test set, and target validation set to train the classification model again.
[0075] Figure 4 This is a flowchart illustrating a data classification method provided in yet another embodiment of the present disclosure. Based on any of the above embodiments, step 301 includes:
[0076] Step 401: Obtain multiple raw data sets.
[0077] Step 402: Perform clustering operations on the multiple original data to obtain multiple clustering results, wherein each clustering result corresponds to at least one original data.
[0078] Step 403: For each clustering result, determine the target keyword corresponding to each original data point of the clustering result by keyword matching, and perform category labeling operation on each original data point according to the target keyword to obtain the original dataset.
[0079] In this embodiment, in order to perform the classification operation on the dataset, the original data can first be labeled to obtain the original dataset.
[0080] Specifically, multiple sets of raw data can be obtained first, including user-initiated questions in actual applications.
[0081] For example, the original data can be shown in Table 1:
[0082] Table 1
[0083]
[0084]
[0085] Furthermore, clustering operations can be performed on multiple original data sets to obtain multiple clustering results, where each clustering result includes multiple original data sets. Optionally, clustering operations on multiple original data sets can be performed using a Gaussian mixture model, or any other clustering method can be used to perform clustering operations on multiple original data sets; this disclosure does not impose any restrictions on this.
[0086] For each clustering result, a target keyword can be determined for each original data point corresponding to that clustering result using a preset keyword matching method. Based on this target keyword, at least one original data point is labeled with its category to obtain the original dataset.
[0087] The data classification method provided in this embodiment generates an original dataset by performing clustering and category labeling operations on the original data. Then, the target classification data can be determined based on the original dataset, and the target classification data can be iteratively classified to achieve accurate classification of the data to be classified.
[0088] Furthermore, based on any of the above embodiments, step 303 includes:
[0089] Obtain a preset keyword mapping vocabulary, wherein the keyword mapping vocabulary includes multiple annotation categories, and each annotation category corresponds to at least one keyword.
[0090] For each set of raw data, a matching operation is performed between the raw data and the keywords in the keyword mapping vocabulary to obtain the target keywords corresponding to the raw data.
[0091] The original data is categorized according to the labeling categories mapped to the target keywords.
[0092] In this embodiment, a keyword mapping terminology can be pre-set. This terminology includes multiple label categories, which may include application conditions, application procedures, fee standards, login passwords, billing dates, etc. Each label category in the terminology may correspond to multiple different keywords.
[0093] To achieve automatic annotation of the raw data, the first step is to obtain a pre-defined keyword mapping vocabulary. Then, the keywords in the keyword mapping vocabulary are matched against the raw data to determine the corresponding keywords for each piece of raw data.
[0094] For each piece of raw data, it can be labeled according to the labeling category mapped to its corresponding keyword. After labeling all the raw data, the raw dataset is obtained.
[0095] It is understandable that the amount of raw data is large and the language expression is relatively rich. Therefore, some raw data may not include keywords from the keyword mapping thesaurus. Therefore, this part of the raw data can be left unannotated.
[0096] The preset labeling categories include, but are not limited to, application conditions, application procedures, fee standards, login passwords, and billing dates. Accordingly, the labeled dataset to be classified is shown in Table 2:
[0097] Table 2
[0098]
[0099]
[0100] The data classification method provided in this embodiment can accurately label the original data corresponding to each clustering result based on the pre-established keyword mapping vocabulary, thereby enabling the subsequent creation of an original dataset based on the labeled original data and the iterative classification operation of the data to be classified.
[0101] Furthermore, based on any of the above embodiments, step 203 includes:
[0102] For each target classification data, the target entropy value corresponding to the target classification data is compared with a preset entropy threshold, and the target classification data is iteratively classified based on the comparison result.
[0103] In this embodiment, the target entropy value corresponding to each piece of data to be classified can be determined based on the output data of the classification model. This target entropy value can accurately characterize whether the predicted category obtained by the classification model is accurate. Optionally, the smaller the target entropy value, the more accurate the predicted category is, while the larger the target entropy value, the less accurate the predicted category is.
[0104] Therefore, to achieve accurate classification of the data to be classified, an entropy threshold can be preset. The target entropy value is compared with the entropy threshold, and different methods are used to iteratively classify the target data based on the comparison results.
[0105] Optionally, the entropy threshold includes a first entropy threshold and a second entropy threshold, wherein the second entropy threshold is greater than the first entropy threshold. For example, in a practical application, the first entropy threshold can be 1, and the second entropy threshold can be 2.
[0106] Optionally, based on any of the above embodiments, the step of comparing the target entropy value corresponding to the target classification data with a preset entropy threshold, and performing an iterative classification operation on the target classification data according to the comparison result, includes:
[0107] For each target classification data, if the target entropy value corresponding to the target classification data is less than the first entropy threshold, then the predicted category is determined as the labeled category of the target classification data; the target classification data is added to the dataset to be classified, and the process of extracting the unclassified data from the dataset to be classified using a preset sampling method and adding the unclassified data to the preset training set, test set, and validation set according to a preset ratio is repeated until the preset iteration termination condition is met, thereby obtaining the classified target training set, target test set, and target validation set.
[0108] In this embodiment, for each target classification data, if the target entropy value corresponding to the target classification data is less than the first entropy threshold, it indicates that the prediction accuracy of the target classification data is relatively high. Therefore, during iterative classification, the predicted category of the classification model can be determined as the labeled category of the target classification data. The target classification data is added to the dataset to be classified, and the data in the dataset to be classified is re-sampled using a preset method and added to the preset training set, test set, and validation set according to a preset ratio. The classification model is then used to further predict the target classification data until the predicted category output by the classification model is the same as the labeled category of the target classification data, thus completing the iterative classification operation of the target classification data.
[0109] Optionally, based on any of the above embodiments, the step of comparing the target entropy value corresponding to the target classification data with a preset entropy threshold, and performing an iterative classification operation on the target classification data according to the comparison result, includes:
[0110] For each target classification data, if the target entropy value corresponding to the target classification data is greater than the second entropy threshold, the target classification data is sent to the terminal device; the labeled category corresponding to the target classification data fed back by the terminal device is obtained, and the labeled category corresponding to the target classification data fed back by the terminal device is determined as the labeled category of the target classification data; the target classification data is added to the dataset to be classified, and the process of extracting the data to be classified from the dataset to be classified using a preset sampling method and adding the data to be classified to the preset training set, test set, and validation set according to a preset ratio is repeated until the preset iteration termination condition is met, thereby obtaining the classified target training set, target test set, and target validation set.
[0111] In this embodiment, if the target entropy value corresponding to the target classification data is greater than the second entropy threshold, it indicates that the accuracy of the predicted category is low. In this case, the target classification data can be sent to the terminal device so that technicians can manually label the target classification data based on their experience.
[0112] Furthermore, the labeled category corresponding to the target classification data fed back by the terminal device can be obtained, and this labeled category is determined as the labeled category of the target classification data. This target classification data is added to the dataset to be classified, and the data in the dataset to be classified is then resampled without replacement and added to the preset training set, test set, and validation set according to a preset ratio. The classification model is then used to predict the target classification data until a preset iteration termination condition is met, resulting in the classified target training set, target test set, and target validation set.
[0113] Optionally, based on any of the above embodiments, the target classification data includes first target classification data with labeled categories and second target classification data without labeled categories.
[0114] The step of comparing the target entropy value corresponding to the target classification data with a preset entropy threshold, and performing iterative classification operations on the target classification data based on the comparison result, includes:
[0115] For each first target classification data, if the target entropy value corresponding to the first target classification data is greater than the first entropy threshold and less than the second entropy threshold, then the target classification data is added to the dataset to be classified, and the process of extracting the data to be classified from the dataset to be classified using a preset sampling method and adding the data to be classified to the preset training set, test set and validation set according to a preset ratio is repeated until the preset iteration termination condition is met, and the classified target training set, target test set and target validation set are obtained.
[0116] For each second target classification data, if the target entropy value corresponding to the second target classification data is greater than the first entropy threshold and less than the second entropy threshold, then the labeling category of the second target classification data is not marked, and the process returns to the steps of extracting the unclassified data from the unclassified dataset using a preset sampling method, and adding the unclassified data to the preset training set, test set and validation set according to a preset ratio, until the preset iteration termination condition is met, and the classified target training set, target test set and target validation set are obtained.
[0117] In this embodiment, the target classification data includes first target classification data with labeled categories and second target classification data without labeled categories. For each first target classification data, if the target entropy value corresponding to the target classification data is greater than a first entropy threshold and less than a second entropy threshold, the labeled category of the first target classification data does not need to be adjusted. The target classification data is added to the dataset to be classified, and the data in the dataset to be classified is resampled without replacement and added to the preset training set, test set, and validation set according to a preset ratio. Furthermore, a classification model is used to predict the target classification data until a preset iteration termination condition is met, thereby obtaining the classified target training set, target test set, and target validation set.
[0118] For each second target classification data point, if the target entropy value corresponding to the target classification data is greater than the first entropy threshold and less than the second entropy threshold, then the label category of the second target classification data is not assigned, and the target classification data is added to the dataset to be classified. The data in the dataset to be classified is then re-sampled using a preset method and added to the preset training set, test set, and validation set according to a preset ratio. Furthermore, a classification model is used to predict the target classification data until a preset iteration termination condition is met, resulting in the classified target training set, target test set, and target validation set.
[0119] The data classification method provided in this embodiment compares the target entropy value with a preset entropy threshold, thereby enabling targeted iterative classification operations on the data to be classified using different processing methods, which improves the accuracy of dataset classification.
[0120] Figure 5 This is a schematic diagram of the structure of the data classification device provided in the embodiments of this disclosure, such as... Figure 5As shown, the device includes: an acquisition module 51, a determination module 52, and a classification module 53. The acquisition module 51 acquires a dataset to be classified, which includes multiple data points to be classified, a target entropy value corresponding to each data point, a predicted category, and labeled categories corresponding to some of the data points. The determination module 52 determines multiple target classification data points in the dataset to be classified, wherein the target classification data points are data points whose predicted categories differ from their labeled categories. The classification module 53 iteratively classifies the multiple target classification data points based on the target entropy value corresponding to each target classification data point until a preset iteration termination condition is met, thereby obtaining a classified target training set, a target test set, and a target validation set.
[0121] Further, based on any of the above embodiments, the acquisition module is configured to: acquire an original dataset, the original dataset including multiple data to be classified and at least some of the labeled categories corresponding to the data to be classified; input the data to be classified in the original dataset into a preset classification model to obtain the output data of the classification model, the output data including the probability that the data to be classified belongs to each labeled category, the classification model being obtained by pre-training a preset training model using an initial training set, an initial test set, and an initial validation set obtained by pre-classifying the dataset to be classified; determine the target entropy value and predicted category corresponding to each data to be classified based on the output data; and generate the dataset to be classified based on the target entropy value, the predicted category, and the labeled categories corresponding to at least some of the data to be classified.
[0122] Further, based on any of the above embodiments, the acquisition module is configured to: acquire multiple raw data; perform clustering operations on the multiple raw data to obtain multiple clustering results, wherein each clustering result corresponds to at least one raw data; for each clustering result, determine the target keyword corresponding to each raw data corresponding to the clustering result by keyword matching; and perform category labeling operations on each raw data according to the target keyword to obtain the original dataset.
[0123] Furthermore, based on any of the above embodiments, the acquisition module is configured to: acquire a preset keyword mapping terminology, wherein the keyword mapping terminology includes multiple annotation categories, and each annotation category corresponds to at least one keyword; and perform a matching operation between the original data and the keywords in the keyword mapping terminology for each piece of original data to obtain the target keyword corresponding to the original data.
[0124] Furthermore, based on any of the above embodiments, the apparatus further includes: an extraction module, configured to extract data to be classified from the original dataset using a preset sampling method, and add the data to be classified to preset training sets, test sets, and validation sets according to preset ratios, thereby obtaining the initial training set, initial test set, and initial validation set. Each piece of data to be classified is sampled once.
[0125] Furthermore, based on any of the above embodiments, the determining module is configured to: determine the labeled category corresponding to the probability value satisfying a preset condition in the output data as the predicted category corresponding to the data to be classified; and calculate the target entropy value corresponding to the data to be classified by performing a preset entropy algorithm on the output data.
[0126] Furthermore, based on any of the above embodiments, the classification module is used to: for each target classification data, compare the target entropy value corresponding to the target classification data with a preset entropy threshold, and perform iterative classification operations on the target classification data according to the comparison result.
[0127] Furthermore, based on any of the above embodiments, the preset entropy threshold includes a first entropy threshold and a second entropy threshold. The second entropy threshold is greater than the first entropy threshold.
[0128] Further, based on any of the above embodiments, the classification module is used to: for each target classification data, if the target entropy value corresponding to the target classification data is less than the first entropy threshold, then determine the predicted category as the labeled category of the target classification data. The target classification data is added to the dataset to be classified, and the process of extracting the data to be classified from the dataset using a preset sampling method, and adding the data to be classified to preset training sets, test sets, and validation sets according to preset proportions, is repeated until a preset iteration termination condition is met, thereby obtaining the classified target training set, target test set, and target validation set.
[0129] Further, based on any of the above embodiments, the classification module is configured to: for each target classification data, if the target entropy value corresponding to the target classification data is greater than the second entropy threshold, then send the target classification data to the terminal device; obtain the label category corresponding to the target classification data fed back by the terminal device, and determine the label category corresponding to the target classification data fed back by the terminal device as the label category of the target classification data; add the target classification data to the dataset to be classified, return to execute the steps of extracting the data to be classified from the dataset to be classified using a preset sampling method, and adding the data to be classified to the preset training set, test set, and validation set according to a preset ratio, until the preset iteration termination condition is met, and obtain the classified target training set, target test set, and target validation set.
[0130] Further, based on any of the above embodiments, the target classification data includes first target classification data with labeled categories and second target classification data without labeled categories. The classification module is used to: for each first target classification data, if the target entropy value corresponding to the first target classification data is greater than the first entropy threshold and less than the second entropy threshold, then add the target classification data to the dataset to be classified, return to execute the steps of extracting the data to be classified from the dataset to be classified using a preset sampling method, and adding the data to be classified to a preset training set, test set, and validation set according to a preset ratio, until a preset iteration termination condition is met, thereby obtaining the classified target training set, target test set, and target validation set. For each second target classification data, if the target entropy value corresponding to the second target classification data is greater than the first entropy threshold and less than the second entropy threshold, then the labeling category processing of the second target classification data is not performed, and the process returns to the steps of extracting the unclassified data from the unclassified dataset using a preset sampling method, and adding the unclassified data to the preset training set, test set and validation set according to a preset ratio, until the preset iteration termination condition is met, and the classified target training set, target test set and target validation set are obtained.
[0131] Furthermore, based on any of the above embodiments, the apparatus further includes: a training module, used to perform iterative training operations on the classification model using the target training set, the target test set, and the target validation set.
[0132] To implement the above embodiments, this disclosure also provides an electronic device, including a processor and a memory.
[0133] The memory stores computer-executed instructions.
[0134] The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any of the above embodiments.
[0135] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this disclosure, such as... Figure 6 As shown, the electronic device 600 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0136] like Figure 6 As shown, electronic device 600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0137] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0138] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0139] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0140] Another embodiment of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the method described in any of the above embodiments.
[0141] Another embodiment of this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any of the above embodiments.
[0142] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0143] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0144] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0145] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0146] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure.
Claims
1. A data classification method, characterized in that, include: Obtain a dataset to be classified, wherein the dataset includes multiple data to be classified, a target entropy value corresponding to each data to be classified, a predicted category, and a labeled category corresponding to a portion of the data to be classified. The target entropy value corresponding to each data to be classified is obtained by fusing and calculating the probability of each data to be classified belonging to each category using a preset entropy value algorithm. The labeled category corresponding to a portion of the data to be classified is the labeled category mapped by the keywords corresponding to the data to be classified in a preset keyword mapping thesaurus. The labeled category includes one of the following: processing conditions, processing procedure, fee standard, login password, and billing date. Identify multiple target classification data in the dataset to be classified, wherein the target classification data are the data to be classified whose predicted category is different from the labeled category; For each target classification data, the target entropy value corresponding to the target classification data is compared with a preset entropy threshold; the preset entropy threshold includes a first entropy threshold and a second entropy threshold, and the second entropy threshold is greater than the first entropy threshold; The labeled categories of the target classification data are updated according to the comparison results. The updated target classification data is added to the dataset to be classified. Iterative classification operations are performed on the dataset to be classified until the preset iteration termination condition is met, and the classified target training set, target test set, and target validation set are obtained. The target classification data includes first target classification data with labeled categories and second target classification data without labeled categories. When the target entropy value is less than the first entropy threshold, the labeled category of the target classification data is the predicted category; When the target entropy value is greater than the second entropy value threshold, the labeled category of the target classification data is the labeled category fed back by the terminal device; When the target entropy value is greater than the first entropy threshold and less than the second entropy threshold, if the target classification data is the first type of target classification data, the labeling category of the target classification data is the labeling category of the first target classification data; if the target classification data is the second type of target classification data, the target classification data does not have a labeling category.
2. The method according to claim 1, characterized in that, The process of obtaining the dataset to be classified includes: Obtain the original dataset, which includes multiple data to be classified and at least some of the labeled categories corresponding to the data to be classified; The data to be classified in the original dataset is input into a preset classification model to obtain the output data of the classification model. The output data includes the probability that the data to be classified belongs to each labeled category. The classification model is obtained by pre-training the preset training model using an initial training set, an initial test set, and an initial validation set obtained by pre-classifying the dataset to be classified. The target entropy value and predicted category corresponding to each data to be classified are determined based on the output data, and the dataset to be classified is generated based on the target entropy value, the predicted category, and the labeled category corresponding to at least some of the data to be classified.
3. The method according to claim 2, characterized in that, The process of obtaining the original dataset includes: Obtain multiple raw data sets; Clustering operations are performed on the multiple original data to obtain multiple clustering results, wherein each clustering result corresponds to at least one original data. For each clustering result, the target keyword corresponding to each original data is determined by keyword matching. The original data is then labeled with the category based on the target keyword to obtain the original dataset.
4. The method according to claim 3, characterized in that, The step of determining the target keywords corresponding to each original data point in the clustering results through keyword matching includes: Obtain a preset keyword mapping vocabulary, wherein the keyword mapping vocabulary includes multiple annotation categories, and each annotation category corresponds to at least one keyword; For each set of raw data, a matching operation is performed between the raw data and the keywords in the keyword mapping vocabulary to obtain the target keywords corresponding to the raw data.
5. The method according to claim 2, characterized in that, Before inputting the data to be classified from the original dataset into the preset classification model, the method further includes: Data to be classified is extracted from the original dataset using a preset sampling method, and the data to be classified is added to the preset training set, test set, and validation set according to a preset ratio to obtain the initial training set, initial test set, and initial validation set. In this case, each of the data to be classified is sampled once.
6. The method according to claim 2, characterized in that, Based on the output data, the predicted category corresponding to each data to be classified is determined, including: The label category corresponding to the probability value in the output data that meets the preset condition is determined as the predicted category corresponding to the data to be classified.
7. The method according to claim 1, characterized in that, The step of updating the labeled category of the target classification data according to the comparison result, adding the updated target classification data to the dataset to be classified, and performing iterative classification operations on the dataset to be classified includes: For each target classification data, if the target entropy value corresponding to the target classification data is less than the first entropy threshold, then the predicted category is determined as the labeled category of the target classification data; the updated target classification data is added to the dataset to be classified, and the process of extracting the data to be classified from the dataset to be classified using a preset sampling method and adding the data to be classified to the preset training set, test set and validation set according to a preset ratio is repeated until the preset iteration termination condition is met, and the classified target training set, target test set and target validation set are obtained.
8. The method according to claim 1, characterized in that, The step of updating the labeled category of the target classification data according to the comparison result, adding the updated target classification data to the dataset to be classified, and performing iterative classification operations on the dataset to be classified includes: For each target classification data, if the target entropy value corresponding to the target classification data is greater than the second entropy value threshold, then the target classification data is sent to the terminal device; The process involves obtaining the labeled category corresponding to the target classification data fed back by the terminal device, determining the labeled category corresponding to the target classification data fed back by the terminal device as the labeled category of the target classification data, adding the target classification data to the dataset to be classified, returning to execute the steps of extracting the unclassified data from the dataset to be classified using a preset sampling method, and adding the unclassified data to the preset training set, test set, and validation set according to a preset ratio, until the preset iteration termination condition is met, and obtaining the classified target training set, target test set, and target validation set.
9. The method according to claim 1, characterized in that, The step of updating the labeled category of the target classification data according to the comparison result, adding the updated target classification data to the dataset to be classified, and performing iterative classification operations on the dataset to be classified includes: For each first target classification data, if the target entropy value corresponding to the first target classification data is greater than the first entropy threshold and less than the second entropy threshold, then the target classification data is added to the dataset to be classified, and the process returns to the step of extracting the unclassified data from the dataset to be classified using a preset sampling method, and adding the unclassified data to the preset training set, test set and validation set according to a preset ratio, until the preset iteration end condition is met, and the classified target training set, target test set and target validation set are obtained; For each second target classification data, if the target entropy value corresponding to the second target classification data is greater than the first entropy threshold and less than the second entropy threshold, then the labeling category of the second target classification data is not marked, the second target classification data is added to the dataset to be classified, and the process of using a preset sampling method to extract the data to be classified from the dataset to be classified, and adding the data to be classified to the preset training set, test set and validation set according to a preset ratio is repeated until the preset iteration termination condition is met, and the classified target training set, target test set and target validation set are obtained.
10. A data classification device, characterized in that, include: The acquisition module is used to acquire the dataset to be classified, wherein the dataset to be classified includes multiple data to be classified, a target entropy value corresponding to each data to be classified, a predicted category, and a labeled category corresponding to a portion of the data to be classified. The target entropy value corresponding to each data to be classified is obtained by fusing and calculating the probability of each data to be classified belonging to each category using a preset entropy value algorithm. The labeled category corresponding to a portion of the data to be classified is the labeled category mapped by the keywords corresponding to the data to be classified in a preset keyword mapping thesaurus. The labeled category includes one of the following: processing conditions, processing procedure, fee standard, login password, and billing date. A determination module is used to determine multiple target classification data in the dataset to be classified, wherein the target classification data are the data to be classified whose predicted category is different from the labeled category; The classification module is used to compare the target entropy value corresponding to each target classification data with a preset entropy threshold; the preset entropy threshold includes a first entropy threshold and a second entropy threshold, and the second entropy threshold is greater than the first entropy threshold. The labeled categories of the target classification data are updated according to the comparison results. The updated target classification data is added to the dataset to be classified. Iterative classification operations are performed on the dataset to be classified until a preset iteration termination condition is met, resulting in a classified target training set, target test set, and target validation set. The target classification data includes first target classification data with labeled categories and second target classification data without labeled categories. Wherein, when the target entropy value is less than the first entropy threshold, the labeled category of the target classification data is the predicted category; when the target entropy value is greater than the second entropy threshold, the labeled category of the target classification data is the labeled category fed back by the terminal device; when the target entropy value is greater than the first entropy threshold and less than the second entropy threshold, if the target classification data is the first type of target classification data, the labeled category of the target classification data is the labeled category of the first target classification data; if the target classification data is the second type of target classification data, the target classification data has no labeled category.
11. An electronic device, characterized in that, include: Memory, processor; Memory; Memory used to store the processor's executable instructions; The processor is used to invoke program instructions in the memory to execute the method as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-9.
Citation Information
Patent Citations
Labeling method and device based on multi-label classification, equipment and storage medium
CN112632278A