Data Classification Method, Apparatus, Device, and Storage Medium

By calculating the classification distance threshold of the sample distance set, clustering the classification data, and dynamically determining the number of category sets, the problems of low clustering efficiency and difficulty in human selection of parameters in the prior art are solved, and more efficient and accurate data classification is achieved.

CN114358102BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111060489.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-10
Publication Date
2025-06-27
Estimated Expiration
2041-09-10

AI Technical Summary

Technical Problem

The existing data clustering technology is complex, resulting in low text clustering efficiency, and requires artificial determination of the initial clustering center and number of classifications, which can easily lead to poor clustering results.

Method used

By obtaining the sample distance set, the classification distance threshold is calculated, and the classification data is clustered according to the threshold, and the number of categories is dynamically determined.

Benefits of technology

It improves the accuracy of data classification, avoids the error of artificially selecting the initial clustering center and classification number, and simplifies the clustering process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358102B_ABST
    Figure CN114358102B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a data classification method, apparatus, device, and storage medium. The method includes: obtaining a sample distance set, where the sample distance set includes distances between every two sample data in a sample data set; obtaining a classification distance threshold according to the sample distance set; clustering the data to be classified in a data set to be classified according to the classification distance threshold to obtain a plurality of category sets, where the data set to be classified includes the sample data set. By adopting the above classification method of the present application, when clustering the data set to be classified according to the classification distance threshold, the number of category sets can be dynamically determined according to the distribution of the data to be classified in the data set to be classified, so as to obtain a plurality of category sets, thereby effectively improving the accuracy of data classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a data classification method, apparatus, device, and storage medium. Background Art

[0002] With the development of Internet technology, a large amount of data has emerged on the Internet (such as news, short videos, short essays, comments, or user characteristics, etc.). Effectively classifying the above data can help understand the current popular trends and better analyze the interests of users.

[0003] Existing data clustering techniques mainly directly cluster all data, such as clustering methods based on structured text, the K-Means method, hierarchical clustering method, and self-organizing mapping clustering, etc. However, the complexity of the above algorithms is relatively high, resulting in low clustering efficiency for text. Summary of the Invention

[0004] In view of this, embodiments of this application propose a data classification method, apparatus, device, and storage medium, which can dynamically determine the number of category sets according to the data distribution, and thus can effectively improve the accuracy of data classification.

[0005] In a first aspect, an embodiment of this application provides a data classification method, the method including: obtaining a sample distance set, where the sample distance set includes the distances between every two sample data in a sample data set; obtaining a classification distance threshold according to the sample distance set; clustering the data to be classified in a data set to be classified according to the classification distance threshold to obtain a plurality of category sets, where the data set to be classified includes the sample data set.

[0006] In a second aspect, an embodiment of this application provides a data classification apparatus, including a distance obtaining module, a threshold obtaining module, and a data classification module. The distance obtaining module is configured to obtain a sample distance set, where the sample distance set includes the distances between every two sample data in a sample data set; the threshold obtaining module is configured to obtain a classification distance threshold according to the sample distance set; the data classification module is configured to cluster the data to be classified in a data set to be classified according to the classification distance threshold to obtain a plurality of category sets, where the data set to be classified includes the sample data set.

[0007] In a possible implementation, the data classification module is further configured to calculate the distance between the target data to be classified and each of the category sets when the target data to be classified is obtained from the set of data to be classified; and when there is a category set whose distance from the target data to be classified is less than the classification distance threshold, store the target data to be classified in this category set; and when there is no category set whose distance from the target data to be classified is less than the classification distance threshold, create a new category set and store the target data to be classified in the newly created category set, thereby obtaining multiple category sets.

[0008] In a possible implementation, the data classification module is further configured to create a new category set and store the target data to be classified obtained from the set of data to be classified in this category set when there is no category set, and to calculate the distance between the target data to be classified and each of the category sets when there are category sets.

[0009] In a possible implementation, the data classification module includes an eigenvalue obtaining sub-module and a first distance obtaining sub-module. The eigenvalue obtaining sub-module is configured to obtain the eigenvalue corresponding to the category set according to the eigenvalues of the data to be classified included in the category set; the first distance obtaining sub-module is configured to obtain the distance between the target data to be classified and the category set according to the eigenvalue of the target data to be classified and the eigenvalue corresponding to the category set.

[0010] In a possible implementation, the eigenvalue obtaining sub-module is further configured to calculate the mean value of the eigenvalues of the data to be classified included in the category set to obtain the eigenvalue corresponding to this category set.

[0011] In a possible implementation, the apparatus further includes an eigenvalue updating module, and the eigenvalue updating module is configured to update the eigenvalue corresponding to the category set to which the target data to be classified belongs according to the eigenvalue of the target data to be classified.

[0012] In a possible implementation, the data classification module further includes a category set acquisition sub-module, a second distance calculation sub-module, and a classification sub-module. The category set acquisition sub-module is configured to obtain the category set with the highest priority from various category sets according to the priority order of the category sets; the second distance calculation sub-module is configured to calculate the distance between the category set with the highest priority and the data to be classified; the classification sub-module is configured to confirm that there is a category set whose distance from the target data to be classified is less than the classification distance threshold when the calculated distance is less than the classification distance threshold, and store the target data to be classified into this category set; the classification sub-module is further configured to, when the calculated distance is not less than the classification distance threshold, delete the category set with the highest priority from the priority order to obtain the updated priority order of various category sets, and when there is no category set with the highest priority, confirm that there is no category set whose distance from the target data to be classified is less than the classification distance threshold, and create a new category set and store the target data to be classified into this category set.

[0013] In a possible implementation, the data classification module is further configured to store the target data to be classified into the category set with the minimum distance from the target data to be classified when there are at least two category sets whose distances from the target data to be classified are less than the classification distance threshold; and to store the target data to be classified into this category set when there is one category set whose distance from the target data to be classified is less than the classification distance threshold.

[0014] In a possible implementation, the threshold acquisition module includes a curve fitting sub-module and a threshold acquisition sub-module. The curve fitting sub-module is configured to fit the distances included in the sample distance set by using a Gaussian mixture model fitting function to obtain a probability density function curve; the threshold acquisition sub-module is configured to obtain the target distance corresponding to the probability value satisfying the specified condition in the probability density function curve, and obtain the classification distance threshold according to the target distance.

[0015] In a possible implementation, the threshold acquisition sub-module is further configured to obtain the target distance corresponding to the minimum probability value in the probability density function curve, and use the target distance as the classification distance threshold.

[0016] In a possible implementation, the distance acquisition module includes an eigenvalue acquisition sub-module and a distance acquisition sub-module. The eigenvalue acquisition module is configured to acquire the eigenvalues of each sample data in the sample dataset; the distance acquisition sub-module is configured to calculate the distances between the eigenvalues of each sample data and the eigenvalues of each of the remaining sample data corresponding to this sample data to obtain a sample distance set including the distances between every two sample data.

[0017] In a possible implementation manner, the distance acquisition sub-module is further configured to calculate, by using the Manhattan distance calculation formula, the distance between the feature value of each sample data and the feature value of each sample data in the remaining sample data corresponding to the sample data, so as to obtain a sample distance set including the distances between every two sample data.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory; one or more programs are stored in the memory and are configured to be executed by the processor to implement the above method.

[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, and when the program code is run by a processor, the above method is executed.

[0020] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device obtains the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.

[0021] A data classification method, device, equipment and storage medium provided by an embodiment of the present application, by obtaining a sample distance set, where the sample distance set includes the distances between every two sample data in a sample data set; obtaining a classification distance threshold according to the sample distance set; and clustering the to-be-classified data in a to-be-classified data set according to the classification distance threshold to obtain a plurality of class sets, where the to-be-classified data set includes the sample data set, can dynamically determine the number of class sets according to the classification distance threshold and the distribution of the to-be-classified data in the to-be-classified data set, and further can effectively improve the accuracy of data classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0023] Figure 1 Shows a schematic system architecture diagram of a data classification system provided by an embodiment of the present application;

[0024] Figure 2 Shows a schematic flowchart of a data classification method proposed by an embodiment of the present application;

[0025] Figure 3 It shows a schematic diagram of a probability density curve proposed in an embodiment of the present application;

[0026] Figure 4 It shows a schematic diagram of a classification result provided in an embodiment of the present application;

[0027] Figure 5 It shows a schematic flowchart of another data classification method provided in an embodiment of the present application;

[0028] Figure 6 It shows Figure 5 a schematic flowchart of step S230 in

[0029] Figure 7 It shows Figure 5 another schematic flowchart of step S230 in

[0030] Figure 8 It shows a schematic flowchart of yet another data classification method proposed in an embodiment of the present application;

[0031] Figure 9 It shows a connection block diagram of a data classification device provided in an embodiment of the present application;

[0032] Figure 10 It shows a structural block diagram of an electronic device for executing the method of the embodiment of the present application. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.

[0034] In recent years, with the rapid development of the Internet, people's daily lives have become more dependent on the network. At the same time, a large amount of business data has been generated. Such as user portrait data, video feature data, image feature data, document feature data, user location data, auto insurance data, and web access data, etc. For these data, they usually need to be divided according to certain criteria so that the data in the same set obtained by the division has as large a similarity as possible.

[0035] Taking user portrait data as an example, during the construction of the user portrait label model, after extracting user features and normalizing the feature data, there are many scenarios for label construction based on clustering, such as promotion sensitivity clustering, comment sensitivity clustering, user loyalty clustering, etc. Under the corresponding user features, the user set is divided into different classes or clusters, so that the similarity of user features within the same cluster is as large as possible or the feature distance is as small as possible, while the difference in user features not in the same cluster is also as large as possible.

[0036] Currently, the commonly used classification method is usually K-means clustering (k-means algorithm). Since the K-means algorithm can prune the tree based on the categories of fewer known clustering samples to determine the classification of some samples, in addition, the algorithm itself has an optimization iteration function, and it iteratively corrects the pruning on the already obtained clusters to determine the clustering of some samples, optimizing the unreasonable parts of the initial supervised learning sample classification and overcoming the inaccuracy of clustering with a small number of samples, so it is widely used.

[0037] When using the k-means algorithm for clustering, first, the parameter k is set, and then the n previously input data objects are divided into k clusters so that the obtained clusters satisfy: the objects in the same cluster have a high similarity; while the objects in different clusters have a small similarity. The clustering similarity is calculated using a "central object" (center of gravity) obtained from the means of the objects in each cluster. The basic steps of the k-means algorithm include: (1) arbitrarily select k objects from the n data objects as the initial clustering centers; (2) calculate the distance between each object and these central objects according to the mean (central object) of each clustering object, and re-divide the corresponding objects according to the minimum distance; (3) recalculate the mean (central object) of each clustering with changes; (4) calculate the standard measure function. When certain conditions are met, such as the function converges, the algorithm terminates; if the conditions are not met, return to step (2).

[0038] The inventor has found through research that when using the K-means algorithm for clustering, the number k of cluster centers needs to be given in advance. However, in practice, it is very difficult to estimate this k value. In many cases, it is not known in advance how many categories the given dataset should be divided into most appropriately. Secondly, in the K-means algorithm, it is necessary to artificially determine the initial cluster centers and determine an initial partition based on the initial cluster centers. Different initial cluster centers may lead to completely different clustering results. Once the initial cluster centers are not selected well, it may not be possible to obtain an effective clustering result. Furthermore, the K-means algorithm is sensitive to outliers and cannot detect outliers, while outliers sometimes have a great impact on the accuracy of the cluster centers. Moreover, when using the K-means algorithm for clustering, it is necessary to continuously adjust the sample classification and continuously calculate the new cluster centers after adjustment. The convergence is slow and the clustering time complexity of O(knt) is relatively high. When the data volume is very large, the time overhead of the algorithm is very large, and during the clustering process, it is necessary to scan all the data to be classified multiple times.

[0039] In view of this, an embodiment of the present application provides a data classification method. The method obtains a sample distance set, which includes the distances between every two sample data in the sample dataset; obtains a classification distance threshold according to the sample distance set; and clusters the data to be classified (business data) in the data set to be classified according to the classification distance threshold to obtain multiple category sets, where the data set to be classified includes the sample dataset. By adopting the above classification method, when clustering the data to be classified in the data set to be classified, it is not necessary to artificially determine the initial cluster centers or confirm the number of classifications. Instead, after calculating the classification distance threshold, multiple category sets are obtained by clustering the data set to be classified according to the classification distance threshold, which can avoid the adverse impact on the final clustering result in the case of incorrect selection of the initial cluster centers.

[0040] Figure 1 The schematic diagram of an exemplary system architecture to which the technical solution of the embodiment of the present invention can be applied is shown.

[0041] As Figure 1 shown, the system architecture may include a server and a terminal device (where the terminal device may be one or more of a smart phone, a tablet computer, and a portable computer configured with a camera component. Of course, it may also be a desktop computer, a television, etc. configured with a camera component). The terminal device and the server can be connected through a network, that is, the network is used as a medium to provide a communication link between the terminal device and the server. The network may include various connection types, such as a wired communication link, a wireless communication link, etc.

[0042] It should be understood that Figure 1The numbers of the terminal devices, networks, and servers in [it] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. For example, the server can be a server cluster composed of multiple servers, etc.

[0043] In an embodiment of the present invention, a user can send a data processing request for classifying service data to a server through a terminal device. The data processing request can include the service data to be processed or the address of the service data. After receiving the data processing request, the server can extract the service data and perform clustering on the extracted service data by executing the classification steps included in the above data classification method to obtain a classification result including multiple category sets and return it to the terminal.

[0044] It should be noted that the data classification method provided by the embodiments of the present invention is generally executed by a server. Correspondingly, the data classification device is generally set in the server. However, in other embodiments of the present invention, the terminal device can also have a similar function to cooperate with the server to execute the data classification method provided by the embodiments of the present invention.

[0045] Next, each embodiment of the present application will be specifically described with reference to the accompanying drawings.

[0046] Please refer to Figure 2 , Figure 2 shown is a data classification method proposed in an embodiment of the present application. This method can be applied to an electronic device as shown in Figure 1 shown. The method includes:

[0047] Step S110: Obtain a sample distance set, which includes the distances between every two sample data in the sample data set.

[0048] Among them, the sample data set refers to a set composed of multiple sample data, that is, the sample data set includes multiple sample data. Each sample data can refer to the feature data of a certain target, such as the feature data of a user, the feature data of a video, the feature data of an image, or the feature data of a document, etc. The above feature data can specifically be attribute information. Correspondingly, the sample data set can include the attribute information of multiple different users, can also include the attribute information of multiple different videos, can also include the attribute information of multiple images, and can also include the attribute information of multiple documents, etc. It should be understood that the above feature data can also be the feature value of a certain attribute.

[0049] Exemplarily, if the sample data included in the sample dataset is the attribute information of a user, the sample data may specifically include one or more of the characteristic data such as the user's age, gender, asset information, income information, or work status, etc. If the sample data is the attribute information of a file (such as a video, document, or news), the sample data may specifically include one or more of the characteristic data such as the file's score, category, required viewing duration, click-through rate, like count, and comment count, etc. If the sample data is the attribute information of an image, the sample data may specifically include one or more of the characteristic data such as the image's category (such as people, animals, plants, food, and buildings, etc.), main color, and clarity, etc.

[0050] The sample data in the sample distance set can be a set number of data to be classified randomly selected from the dataset to be classified, or the first set number of data to be classified generated by the data system received by the dataset to be classified. The above set number can be any number such as 100, 200, or 500, etc., and can be set according to actual needs.

[0051] The above data system can be any system capable of generating data that needs to be classified. For example, the data system can be a banking system, a location positioning system (a system for finding potential markets), an auto insurance system, and a web browsing system, etc. The corresponding sample data can be historical data such as user portrait data, user location data, auto insurance data, and web access data, etc. pre-stored in the database corresponding to the data system, or real-time data generated during the operation of the data system.

[0052] In an implementable manner, the obtaining method of the sample distance set can be: for each sample data, calculate the distance between this sample data and the remaining sample data in the sample distance set except this sample data. In this way, the distance between every two sample data can be obtained, that is, a sample distance set including the distance between every two sample data is obtained.

[0053] Considering that the sample data can refer to the characteristic information of a certain target, therefore, when obtaining the sample distance set, the characteristic value corresponding to the characteristic information of each sample data can be specifically obtained, and the distance value between every two sample data can be obtained according to the characteristic values corresponding to each sample data.

[0054] Therefore, the obtaining method of the distance between every two sample data in the sample distance set can be specifically: obtain the characteristic value of each sample data in the sample data, and calculate the distance between the characteristic value of each sample data and the characteristic value of each sample data in the remaining sample data corresponding to this sample data, so as to obtain a sample distance set including the distance between every two sample data.

[0055] In another feasible implementation, the way to obtain the distance between every two sample data in the sample distance set can be as follows: Based on the feature data corresponding to each sample data, obtain the feature vector corresponding to the sample data, and use a distance calculation formula to calculate the distance for the feature vector, so as to obtain the distance between every two sample data.

[0056] The above distance calculation formula can be an Euclidean distance calculation formula, a Manhattan distance calculation formula, a Mahalanobis distance calculation formula, etc.

[0057] Considering that different features have different influences on data classification, in this embodiment, the way to calculate the distance between the feature value of each sample data and the feature value of each sample data in the remaining sample data corresponding to the sample data, and obtain the sample distance set including the distance between every two sample data can also be to calculate the distance between the feature value of each sample data and the feature value of each sample data in the remaining sample data corresponding to the sample data based on the weight coefficients corresponding to different features.

[0058] Among them, the feature value of the sample data can specifically refer to the value obtained after parameterizing the attribute information in the sample data.

[0059] Step S120: Obtain a classification distance threshold according to the sample distance set.

[0060] Among them, there are various ways to obtain the sample distance threshold according to the sample distance set.

[0061] In a feasible implementation, the sample distance threshold can be calculated by using the method of calculating the mean value of the sample distance set.

[0062] In this method, specifically, the mean value of each distance in the sample distance set can be obtained, and the obtained mean value can be used as the classification distance threshold. Or the mean value of each distance in the sample distance set can be obtained, and the obtained mean value can be multiplied by a preset coefficient to obtain the classification distance threshold. Among them, the preset coefficient can be obtained based on the user's pre-setting, and can be any constant such as 0.9, 0.95, 0.98, 1.02, etc. It can also be to select the distance value whose middle distance value in the sample distance set is within the preset distance range, and obtain the target distance value by taking the mean value of the selected distance values, and then take the mean value of the target distance value, so as to obtain the classification distance threshold according to the obtained mean value.

[0063] In another feasible implementation, it can also be to use a Gaussian mixture model fitting function to fit the distances included in the sample distance set to obtain a probability density function curve; obtain the target distance corresponding to when the probability value in the probability density function curve satisfies the specified condition, and obtain the classification distance threshold according to the target distance.

[0064] In this way, the above Gaussian mixture model can be a two-dimensional Gaussian mixture model. The specific target distance corresponding to the obtained probability value satisfying the specified condition can be the target distance corresponding to the minimum obtained probability value, or a target distance can be confirmed from the distances corresponding to the points in the probability density function curve with values less than the preset value as the classification distance threshold.

[0065] Among them, the two-dimensional mixture Gaussian model is an extension of a single Gaussian probability density function. For example: there is a set of observed data X, and the data set X includes n data, that is , if the distribution in the corresponding d-dimensional space is not ellipsoidal, then it is not suitable to describe the probability density function of these data points with a single Gaussian density function. At this time, it is assumed that each point is generated by a single Gaussian distribution (the specific parameters are unknown), and this batch of data is generated by two single Gaussian models in total. Specifically, a certain data belongs to which single Gaussian model is unknown, and the proportion of each single Gaussian model in the mixture model is unknown. Mix all the data points from different distributions together, and this distribution is called a two-dimensional Gaussian mixture distribution.

[0066] As Figure 3 shown, the distances included in the sample distance set are fitted using the Gaussian mixture model fitting function to obtain a probability density function curve. The abscissa of this probability density curve is the distance, and the ordinate is the quantity value, that is, the corresponding sample quantity at this distance value.

[0067] As Figure 3 point A in shows, the abscissa of point A is the target distance corresponding to the minimum value point in the probability density function curve, that is, the abscissa of the minimum value point of the probability density function curve of the two-dimensional Gaussian mixture distribution can be used as the distance threshold for clustering. It can be simply understood that the set of intra-class distances follows a normal distribution with a small mean and a large variance, and the set of inter-class distances follows a normal distribution with a large mean and a small variance. The intersection point of the probability density function curves of the two normal distributions is the minimum value point.

[0068] Step S130: According to the classification distance threshold, cluster the data to be classified in the data set to be classified to obtain multiple category sets. The data set to be classified includes a sample data set.

[0069] As an implementable manner, the way to cluster the data set to be classified according to the classification distance threshold to obtain multiple category sets can be as follows: Select at least two data to be classified from the data set to be classified, where the distance between every two of the at least two data to be classified is greater than the classification distance threshold, and establish at least two category sets based on the at least two data to be classified respectively. Among them, each category set corresponds to storing one data to be classified. Subsequently, for the obtained data to be classified, calculate the distance between the data to be classified and the center of each category set, and compare the calculated distance with the classification distance threshold. If there is a distance between a category set that is less than the classification distance threshold, store the data to be classified into this category set. If not, create a new category set and store the data to be classified into the newly created category set, and return to the step of calculating the distance between the obtained data to be classified and the center of each category set.

[0070] As another implementable manner, the way to cluster the data set to be classified according to the classification distance threshold to obtain multiple category sets can also be: Select a data to be classified from the data set to be classified as the clustering center, take the data to be classified whose distance to this clustering center is less than the classification distance threshold as a category set, and the data in the data set to be classified other than the category set as a new data set to be classified, and return to execute the step of selecting any data to be classified from the data set to be classified as the clustering center until all the data in the data set to be classified are completed for classification to obtain multiple category sets.

[0071] In this way, the way to select a data to be classified from the data to be classified as the clustering center can be to randomly select a data to be classified from the data set to be classified; it can also be to establish an N-dimensional coordinate system based on the number of features (N) of the data to be classified to obtain the positions of each data to be classified in the N-dimensional coordinate system, so as to determine a target data to be classified according to the positions of each data to be classified in the N-dimensional coordinate system, and the position of the target data to be classified in the N-dimensional coordinate system is the position in the area with the highest data distribution density.

[0072] As Figure 4 shown, it is the classification result obtained by classifying multiple data to be classified in the classification data set according to the classification distance threshold obtained in Figure 3 . It can be seen from the figure that the inter-class distances between various category sets (such as category set 1, category set 2, and category set 3 in Figure 1 ) are relatively large and are usually greater than the classification distance threshold, and the intra-class distances between the data in the same category set (such as category set 1, category set 2, or category set 3) are relatively small and are usually less than the classification distance threshold.

[0073] By adopting the data classification method of the present application, a sample distance set is obtained, and the sample distance set includes the distances between every two sample data in the sample data set; according to the sample distance set, a classification distance threshold is obtained; according to the classification distance threshold, the data to be classified in the data set to be classified is clustered to obtain a plurality of category sets, and the data set to be classified includes the sample data set. It is possible to dynamically determine the number of category sets according to the data distribution situation, and thus effectively improve the accuracy of data classification.

[0074] Please refer to Figure 5 , an embodiment of the present application provides a data classification method, and the method includes the following steps:

[0075] Step S210: Obtain a sample distance set, and the sample distance set includes the distances between every two sample data in the sample data set.

[0076] Step S220: Obtain a classification distance threshold according to the sample distance set.

[0077] Step S230: If a target data to be classified is obtained from the data set to be classified, calculate the distances between the target data to be classified and each category set.

[0078] It should be understood that the step of calculating the distances between the target data to be classified and each category set should be executed when the category sets are stored in the electronic device.

[0079] In an implementable manner, the way of calculating the distances between the target data to be classified and each category set can be: for each category set, calculate the distances between the target data to be classified and all the data to be classified included in the category set, and obtain the distance between the target data to be classified and the category set according to the distances.

[0080] In this implementation manner, the way of obtaining the distance between the target data to be classified and the category set according to the distances can be: taking the average value of the distances to obtain the distance between the target data to be classified and the category set. It can also be: sorting the distances, and selecting the distance value in the middle after sorting as the distance between the target data to be classified and the category set.

[0081] In another implementable manner, the way of calculating the distances between the target data to be classified and each category set can also be, for each category set, determining a target feature based on the features of the data to be classified in the category set, and calculating the distance between the target feature and the feature of the target data to be classified, and this distance is the distance between the target data to be classified and the category set.

[0082] That is to say, please refer to Figure 6 , in this implementation manner, the above step S230 includes:

[0083] Step S231: Obtain the eigenvalue corresponding to the category set according to the eigenvalue of the data to be classified included in the category set.

[0084] In this implementation manner, the way to obtain the eigenvalue corresponding to the category set based on the features of the data to be classified in the category set can be to calculate the mean value of the features of the data to be classified in the category set, and this mean value is the eigenvalue corresponding to the category set. It can also be: select the eigenvalue with the highest occurrence frequency from the features of the data to be classified in the category set, which is the eigenvalue corresponding to the category set. It can also be: perform parameterization processing on the features of the data to be classified in the category set, and perform weighted summation on the features of each data to be classified after parameterization processing to obtain the weighted value corresponding to each data to be classified, and select the target value with the highest occurrence frequency from the weighted values, and calculate the mean value of the features of the data to be classified corresponding to this target value to obtain the eigenvalue corresponding to the category set.

[0085] In an implementable manner, the above step S231 can specifically be: calculate the mean value of the eigenvalues of the data to be classified included in the category set to obtain the eigenvalue corresponding to the category set.

[0086] Step S232: Obtain the distance between the target data to be classified and the category set according to the eigenvalue of the target data to be classified and the eigenvalue corresponding to the category set.

[0087] Among them, the above step S232 can specifically be: calculate the eigenvalue of the target data to be classified and the eigenvalue corresponding to the category set using the distance calculation formula to obtain the distance between the target data to be classified and the category set. Among them, the above distance calculation formula can be the Euclidean distance calculation formula, or the Manhattan distance calculation formula, or the Mahalanobis distance calculation formula, etc.

[0088] It should be understood that the distance calculation formulas used for the distance between the above sample data and the distance between the target data to be classified and the category set are the same.

[0089] Exemplarily, in this embodiment, when the distance between the sample data is the Manhattan distance, the distance between the target data to be classified and the category set is also the Manhattan distance. For the data to be classified with n types of eigenvalues (n-dimensional data to be classified), the calculation method of the Manhattan distance between two data to be classified x and y is as follows:

[0090] where r = 1. If r in the above formula is set to 2, it is the calculation method of Euclidean distance, that is, Euclidean distance needs to calculate the sum of squares and square roots, both of which are calculation methods with relatively slow speeds. In this application, by using Manhattan distance, only simple numerical addition and subtraction operations are required, and its calculation complexity is much lower than that of Euclidean distance, thus greatly reducing the calculation overhead and improving the performance and speed of data clustering.

[0091] Step S240: Detect whether there is a distance between the category set and the target data to be classified that is less than the classification distance threshold.

[0092] Among them, when there is a distance between the category set and the target data to be classified that is less than the classification distance threshold, it can be considered that there is similarity or consistency between the data in this category set and the target data to be classified, and they belong to the same category of data. If there is no distance between the category set and the target data to be classified that is less than the classification distance threshold, it can be considered that the target data to be classified does not have similarity or consistency with the data that has not been classified.

[0093] In an implementable manner, when detecting whether there is a distance between the category set and the target data to be classified that is less than the classification distance threshold, it can be after calculating the distance between a set of data to be classified and the target data to be classified, detecting whether the calculated distance is less than the classification threshold, and when the detection result is yes, confirming that there is a distance between the category set and the target data to be classified that is less than the classification distance threshold and performing subsequent classification steps.

[0094] The above detection of whether there is a distance between the category set and the target data to be classified that is less than the classification distance threshold can also be after calculating the distances between all target category sets and the target data to be classified, respectively comparing the calculated distances with the classification distance threshold to confirm whether there is a distance between the category set and the target data to be classified that is less than the classification distance threshold.

[0095] If there is a distance between the category set and the target data to be classified that is less than the classification distance threshold, execute Step S250: Store the target data to be classified into this category set, and return to execute Step S230: If the target data to be classified is obtained from the set of data to be classified, calculate the distance between the target data to be classified and each category set.

[0096] If there is no distance between the category set and the target data to be classified that is less than the classification distance threshold, execute Step S260: Create a new category set, and store the target data to be classified into the newly created category set to obtain multiple category sets, and return to execute Step S230: If the target data to be classified is obtained from the set of data to be classified, calculate the distance between the target data to be classified and each category set.

[0097] In one implementable manner, refer to Figure 7 , the above step S230: calculating the distance between the target data to be classified and each category set may specifically include:

[0098] Step S234: Obtain the category set with the highest priority from each category set according to the priority order of each category set.

[0099] Among them, the priority order of each category set can be confirmed based on the generation time of the category set, that is, the earlier the generation time of the category set, the higher the priority of the category set. The priority order of each category set can also be confirmed based on the amount of data included in each category set, that is, the more data included in the category set, the higher the priority of the category set.

[0100] Step S235: Calculate the distance between the category set with the highest priority and the data to be classified.

[0101] Regarding the calculation of the distance between the category set and the data to be classified, refer to the previous description, and the steps are not elaborated here one by one.

[0102] Step S236: Detect whether the calculated distance is less than the classification distance threshold.

[0103] If the calculated distance is less than the classification distance threshold, confirm that there is a category set whose distance from the target data to be classified is less than the classification distance threshold, and execute S250.

[0104] If the calculated distance is not less than the classification distance threshold, then execute step S237: delete the category set with the highest priority from the priority order to obtain the updated priority order of each category set, and return to execute step S234 until there is no category set with the highest priority, confirm that there is no category set whose distance from the target data to be classified is less than the classification distance threshold, and execute step S260: create a new category set and store the target data to be classified in this category set.

[0105] In another embodiment, if in the above step S240, after calculating the distances between all the obtained target category sets and the target data to be classified, the calculated distances are respectively compared with the classification distance threshold, there may be a situation where the distances between at least one (such as two or three) category sets and the target data to be classified are less than the classification distance threshold. Therefore, the above step S250 may specifically be: if there are at least two category sets whose distances from the target data to be classified are less than the classification distance threshold, store the target data to be classified in the category set with the minimum distance from the target data to be classified; if there is one category set whose distance from the target data to be classified is less than the classification distance threshold, store the target data to be classified in this category set.

[0106] Exemplarily, when classifying the target data to be classified a1, it is assumed that there are category sets A1, A2, and A3, and the priorities of A1, A2, and A3 decrease in sequence for illustration.

[0107] When classifying the target data to be classified a1, obtain the category set A1 with the highest priority, calculate the distance between the target data to be classified a1 and the category set A1. If the distance is less than the classification distance threshold, store a1 in the category set A1. If the distance is not less than the classification distance threshold, delete the category set A1 from the priority order to obtain the updated priority order. Obtain the category set A2 with the highest priority from the updated priority order, and calculate the distance between a1 and the category set A2. If the distance is less than the classification distance threshold, store a1 in the category set A2. If it is not less than the classification distance threshold, in a similar manner as above, continue to delete A2 from the priority order to obtain the updated priority order, and obtain the category set A3 with the highest priority from the updated priority order, and continue to calculate the distance between the target data to be classified a1 and the category set A3. If the distance is less than the classification distance threshold, store a1 in the category set A3. If it is not less than the classification distance threshold, delete the set A3 from the priority order to obtain the updated priority order. At this time, there is no category set with the highest priority in the updated priority order. Therefore, a new category set A4 needs to be created, and the target data to be classified is stored in the category set A4.

[0108] After completing the classification of a1, if it is necessary to classify the new data to be classified a2 obtained from the data set to be classified, the existing category set at this time includes A1, A2, A3, and A4. If the priority order of the category sets A1, A2, A3, and A4 decreases in turn, and when it is necessary to classify the target data to be classified a2, it can be compared and divided in a similar way to the target data to be classified a1. First, calculate the distance between the target data to be classified a2 and the category set A1 and compare it with the classification distance threshold. If it is less than the classification distance threshold, the target data to be classified a2 is stored in the category set A1. If it is not less than the classification threshold, calculate the distance between the target data to be classified a2 and the category set A2, and so on, and the classification of the data to be classified a2 can be completed. Similarly, by using a similar classification method, the classification of all the data to be classified can be completed.

[0109] After completing the classification of the target data to be classified, that is, after storing the target data to be classified in the category set, the method further includes:

[0110] According to the feature value of the target data to be classified, update the feature value corresponding to the category set to which the target data to be classified belongs.

[0111] Among them, according to the feature value of the target data to be classified, the way to update the feature value corresponding to the category set to which the target data to be classified belongs can be to calculate the average value of the feature values of all the data to be classified included in the category set to which the target data to be classified belongs, and obtain the updated feature value corresponding to the category set to which the target data to be classified belongs. Another way to update the feature value corresponding to the category set to which the target data to be classified belongs can be to select the feature value with the highest occurrence frequency from the features of the data to be classified in the category set obtained after adding the target data to be classified, as the updated feature value of the category set.

[0112] By updating the feature value corresponding to the category set to which the target data to be classified belongs, the classification based on the distance calculated from the updated feature value obtained in the subsequent classification process can be made more reliable.

[0113] By adopting the above steps S210 - S250, after obtaining the classification distance threshold according to the sample distance set, when classifying the target data to be classified obtained from the data set to be classified, if there is no category set, a new category set is created and the target data to be classified is stored in this category set, and when there is a category set, it is determined whether a new category set needs to be created or directly divided according to the distance between the target data to be classified and this category set. Thus, when clustering the data to be classified in the data set to be classified, it is not necessary to artificially determine the initial clustering center, but the number of clusters is intelligently determined by automatically analyzing the data distribution, and at the same time, the adverse impact on the final clustering result caused by the mistake in selecting the initial clustering center is avoided. Compared with other clustering algorithms, usually more categories are generated after clustering using this algorithm. This is because the number of data points contained in some categories is very small, and in most scenarios, these categories can be regarded as outlier categories, thus achieving the effect of automatically detecting outliers, while the categories containing more data points are regarded as categories with practical significance. In addition, after determining the distance threshold in this application, during the classification process, there is no need for iteration, and clustering can be completed by scanning the data once. Therefore, this solution can also be used to cluster real - time data, making the application scope of the algorithm wider.

[0114] Please refer to Figure 8 As shown, exemplarily, taking the data to be classified as the user portrait data in the banking system, and the user portrait data specifically includes three characteristic information: user asset information, user salary information, and user consumption level information for classification as an example for illustration.

[0115] When clustering the user portrait data, the following steps can be specifically executed:

[0116] Step S301: Obtain the characteristic values of the sample user portrait data.

[0117] Specifically, the sample user portrait data stored in the data set to be classified corresponding to the banking system can be obtained, and the characteristic information in the user portrait data is respectively parameterized to obtain the characteristic values corresponding to the sample user portrait data.

[0118] Among them, the parameter values corresponding to different characteristic information should be different. For example, user A has 2 properties in city a, an annual income of 200,000, and an average monthly consumption of 10,000, then the characteristic value of user A can be (4, 2, 2); user B has 2 properties in county b of city b, an annual income of 500,000, and an average monthly consumption of 20,000, then the characteristic value of user B can be (1, 5, 3); user C has no house, an annual income of 200,000, and an average monthly consumption of 3,000, then the characteristic value of user C can be (0, 2, 0.4).

[0119] Step S302: Using the Manhattan distance calculation formula, calculate a sample distance set including the distances between every two sample user portrait data.

[0120] Specifically, it can be to use the Manhattan distance calculation formula to calculate the distance between the eigenvalue of each sample user portrait data and the eigenvalue of each sample user portrait data in the remaining sample user portrait data corresponding to this sample user portrait data, so as to obtain a sample distance set including the distances between every two sample user portrait data.

[0121] Step S303: Perform Gaussian fitting on the sample distance set to obtain a curve.

[0122] Specifically, the Gaussian mixture model fitting function can be used to fit the distances included in the sample distance set to obtain a probability density function curve.

[0123] Step S304: Obtain the minimum value in the curve as the classification distance threshold.

[0124] Specifically, the target distance corresponding to the minimum probability value in the probability density function curve can be obtained, and the classification distance threshold can be obtained according to the target distance.

[0125] When classifying the user portrait data to be classified, step S305 can be executed: obtain the target user portrait data to be classified from the data set to be classified; and step S306 can be executed: detect whether there is a category set.

[0126] If there is no category set, then step S307 is executed: create a new category set, store the target user portrait data to be classified in the newly created category set, and return to execute step S305: obtain the target user portrait data to be classified from the data set to be classified.

[0127] If there is a category set, then step S308 is executed: obtain the category set with the highest priority from each category set according to the establishment order of each category set, and calculate the distance between the category set with the highest priority and the data to be classified.

[0128] Step S309: Detect whether the distance between the category set with the highest priority obtained by calculation and the data to be classified is less than the classification distance threshold.

[0129] If the calculated distance is less than the classification distance threshold, confirm that there is a distance between the category set and the target user portrait data to be classified that is less than the classification distance threshold, and execute step S310: store the target user portrait data to be classified in this category set.

[0130] If the calculated distance is not less than the classification distance threshold, perform step S311: Delete the category set with the highest priority from the priority order to obtain the updated priority order of each category set. Then return to execute step S308: According to the priority order of each category set, obtain the category set with the highest priority from each category set, until there is no category set with the highest priority, it can be confirmed that there is no category set whose distance from the target user portrait data to be classified is less than the classification distance threshold, and then create a new category set and store the target user portrait data to be classified in this category set.

[0131] After storing the target user portrait data to be classified in the category set, execute step S312: According to the eigenvalue of the target user portrait data to be classified, update the eigenvalue corresponding to the category set to which the target user portrait data belongs, and then return to execute step S306: Obtain the target user portrait data from the data set to be classified, until all the data to be classified in the data set to be classified are completed, the clustering task of all the user portrait data to be classified in the data set to be classified is achieved.

[0132] Please refer to Figure 9 , this application provides a data classification device 400, and the data classification device 400 includes a distance acquisition module 410, a threshold acquisition module 420, and a data classification module 430.

[0133] The distance acquisition module is used to acquire a sample distance set, and the sample distance set includes the distances between every two sample data in the sample data set.

[0134] Among them, the distance acquisition module 410 includes an eigenvalue acquisition sub-module and a distance acquisition sub-module.

[0135] The eigenvalue acquisition module is used to acquire the eigenvalue of each sample data in the sample data set; the distance acquisition sub-module is used to calculate the distance between the eigenvalue of each sample data and the eigenvalue of each sample data in the remaining sample data corresponding to the sample data, so as to obtain a sample distance set including the distances between every two sample data.

[0136] In this implementation manner, the distance acquisition sub-module is further used to calculate the distance between the eigenvalue of each sample data and the eigenvalue of each sample data in the remaining sample data corresponding to the sample data by using the Manhattan distance calculation formula, so as to obtain a sample distance set including the distances between every two sample data.

[0137] The threshold acquisition module 420 is used to obtain a classification distance threshold according to the sample distance set.

[0138] Among them, the threshold acquisition module 420 includes a curve fitting sub-module and a threshold acquisition sub-module.

[0139] A curve fitting sub-module is used to fit the distances included in the sample distance set by using a Gaussian mixture model fitting function to obtain a probability density function curve. A threshold obtaining sub-module is used to obtain a target distance corresponding to when the probability value in the probability density function curve meets a specified condition, and obtain a classification distance threshold according to the target distance.

[0140] In one implementation manner, the threshold obtaining sub-module is further used to obtain a target distance corresponding to when the probability value in the probability density function curve is the minimum value, and use the target distance as the classification distance threshold.

[0141] The data classification module 430 is used to cluster the data to be classified in the data set to be classified according to the classification distance threshold to obtain a plurality of category sets, and the data set to be classified includes a sample data set.

[0142] In one implementable manner, the data classification module 430 is further used to calculate the distance between the target data to be classified and each category set when obtaining the target data to be classified from the data set to be classified; and when there is a category set whose distance from the target data to be classified is less than the classification distance threshold, store the target data to be classified into this category set; and when there is no category set whose distance from the target data to be classified is less than the classification distance threshold, create a new category set and store the target data to be classified into the newly created category set to obtain a plurality of category sets.

[0143] In one implementable manner, the data classification module 430 is further used to create a new category set when there is no category set, and store the target data to be classified obtained from the data set to be classified into this category set, and is used to calculate the distance between the target data to be classified and each category set when there is a category set.

[0144] In one implementable manner, the data classification module 430 includes an eigenvalue obtaining sub-module and a first distance obtaining sub-module. The eigenvalue obtaining sub-module is used to obtain the eigenvalue corresponding to the category set according to the eigenvalues of the data to be classified included in the category set; the first distance obtaining sub-module is used to obtain the distance between the target data to be classified and the category set according to the eigenvalue of the target data to be classified and the eigenvalue corresponding to the category set.

[0145] In this implementation manner, the eigenvalue obtaining sub-module is further used to calculate the mean value of the eigenvalues of the data to be classified included in the category set to obtain the eigenvalue corresponding to this category set.

[0146] In this embodiment, the data classification module 430 further includes a category set acquisition sub-module, a second distance calculation sub-module, and a classification sub-module. The category set acquisition sub-module is configured to obtain the category set with the highest priority from various category sets according to the priority order of the category sets; the second distance calculation sub-module is configured to calculate the distance between the category set with the highest priority and the data to be classified; the classification sub-module is configured to confirm that there is a category set with a distance less than the classification distance threshold from the target data to be classified when the calculated distance is less than the classification distance threshold, and store the target data to be classified into this category set; the classification sub-module is further configured to, when the calculated distance is not less than the classification distance threshold, delete the category set with the highest priority from the priority order to obtain the updated priority order of the category sets, and when there is no category set with the highest priority, confirm that there is no category set with a distance less than the classification distance threshold from the target data to be classified, and create a new category set and store the target data to be classified into this category set.

[0147] In this embodiment, the data classification module 430 is further configured to store the target data to be classified into the category set with the minimum distance from the target data to be classified when there are at least two category sets with a distance less than the classification distance threshold from the target data to be classified; and to store the target data to be classified into this category set when there is one category set with a distance less than the classification distance threshold from the target data to be classified.

[0148] In an implementable embodiment, the data classification device 400 further includes a feature value update module, and the feature value update module is configured to update the feature value corresponding to the category set to which the target data to be classified belongs according to the feature value of the target data to be classified.

[0149] It should be noted that the device 400 embodiment in this application corresponds to the foregoing method embodiment. The specific principle in the device 400 embodiment can be referred to the content in the foregoing method embodiment, and will not be elaborated here.

[0150] Next, Figure 10 an electronic device 100 provided by this application will be described.

[0151] Please refer to Figure 10 , based on the data classification method provided in the foregoing embodiment, another electronic device 100 provided in the embodiment of this application includes a processor 102 that can execute the foregoing method. The electronic device 100 can be a server 10 or a terminal device, and the terminal device can be a smart phone, a tablet computer, a computer, a portable computer, or other devices.

[0152] The electronic device 100 further includes a memory 104. Among them, a program that can execute the content in the foregoing embodiments is stored in the memory 104, and the processor 102 can execute the program stored in the memory 104.

[0153] Among them, the processor 102 may include one or more cores for processing data and a message matrix unit. The processor 102 connects various parts within the entire electronic device 100 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by calling data stored in the memory 104, it executes various functions of the electronic device 100 and processes data. Optionally, the processor 102 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 102 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the foregoing modem may not be integrated into the processor 102 and may be implemented separately through a communication chip.

[0154] The memory 104 may include a random access memory (RAM) and may also include a read-only memory. The memory 104 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function, instructions for implementing the following various method embodiments, etc. The data storage area may also store data obtained by the electronic device 100 during use (such as data to be recommended and operation methods), etc.

[0155] The electronic device 100 may further include a network module and a screen. The network module is used to receive and send electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices, such as communicating with an audio playback device. The network module may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The network module can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through wireless networks. The above-mentioned wireless networks may include cellular phone networks, wireless local area networks or metropolitan area networks. The screen can display interface content and perform data interaction.

[0156] In some embodiments, the electronic device 100 may further include: a peripheral interface 106 and at least one peripheral device. The processor 102, the memory 104 and the peripheral interface 106 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral interface through a bus, signal lines or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency component 108, a positioning component 112, a camera 114, an audio component 116, a display screen 118, and a power supply 122, etc.

[0157] The peripheral interface 106 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 102 and the memory 104. In some embodiments, the processor 102, the memory 104 and the peripheral interface 106 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 102, the memory 104 and the peripheral interface 106 can be implemented on a separate chip or circuit board, and the embodiments of the present application do not limit this.

[0158] The radio frequency component 108 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency component 108 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency component 108 converts electrical signals into electromagnetic signals for transmission, or converts the received electromagnetic signals into electrical signals. Optionally, the radio frequency component 108 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency component 108 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, metropolitan area networks, intranets, generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the radio frequency component 108 may further include circuits related to NFC (Near Field Communication), which is not limited in this application.

[0159] The positioning component 112 is used to locate the current geographical location of the electronic device to implement navigation or LBS (Location-Based Service). The positioning component 112 can be a positioning component based on the GPS (Global Positioning System), Beidou system, or Galileo system.

[0160] The camera 114 is used to capture images or videos. Optionally, the camera 114 includes a front camera and a rear camera. Generally, the front camera is set on the front panel of the electronic device 100, and the rear camera is set on the back of the electronic device 100. In some embodiments, there are at least two rear cameras, which can be any one of the main camera, depth camera, wide-angle camera, and telephoto camera, to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera 114 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0161] The audio component 116 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 102 for processing, or input to the radio frequency component 108 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 102 or the radio frequency component 108 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio component 116 may further include a headphone jack.

[0162] The display screen 118 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 118 is a touch display screen, the display screen 118 also has the ability to collect touch signals on or above the surface of the display screen 118. The touch signals can be input to the processor 102 as control signals for processing. At this time, the display screen 118 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 118, which is arranged on the front panel of the electronic device 100; in other embodiments, there may be at least two display screens 118, which are respectively arranged on different surfaces of the electronic device 100 or in a foldable design; in still other embodiments, the display screen 118 may be a flexible display screen, which is arranged on the curved surface or the folding surface of the electronic device 100. Even, the display screen 118 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 118 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0163] The power supply 122 is used to supply power to each component in the electronic device 100. The power supply 122 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 122 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired circuit, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0164] The embodiments of the present application also provide a computer-readable storage medium. Program codes are stored in the computer-readable medium, and the program codes can be called by a processor to execute the methods described in the above method embodiments.

[0165] The computer-readable storage medium may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has a storage space for program codes for executing any method steps in the above methods. These program codes can be read out from one or more computer program products or written into one or more computer program products. The program codes can be compressed in an appropriate form, for example.

[0166] The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods described in the above various alternative implementation manners.

[0167] In summary, a data classification method, apparatus, device, and storage medium provided by the present application. The method includes: obtaining a sample distance set, where the sample distance set includes distances between every two sample data in a sample data set; obtaining a classification distance threshold according to the sample distance set; clustering the data to be classified in a data set to be classified according to the classification distance threshold to obtain a plurality of category sets, and the data set to be classified includes the sample data set. By adopting the above data classification method of the present application, clustering the data set to be classified according to the classification distance threshold to obtain a plurality of category sets can be realized, and thus the accuracy of data classification can be effectively improved.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A data classification method, characterized in that, The method includes: Obtain a sample distance set, which includes the distances between every two sample data in the sample data set, and the sample data includes feature data of a video, feature data of an image, or feature data of a document; Obtain a classification distance threshold according to the sample distance set; Cluster the data to be classified in the data set to be classified according to the classification distance threshold to obtain a plurality of category sets, where the data set to be classified includes the sample data set, and the data to be classified in the data set to be classified includes at least one of historical data and real-time data generated during the operation of the data system; The clustering of the data to be classified in the data set to be classified according to the classification distance threshold to obtain a plurality of category sets includes: If target data to be classified is obtained from the data set to be classified, calculate the distances between the target data to be classified and each of the category sets; If there is a category set whose distance from the target data to be classified is less than the classification distance threshold, store the target data to be classified in this category set, and return to execute the step of if target data to be classified is obtained from the data set to be classified, calculate the distances between the target data to be classified and each of the category sets; The calculating the distances between the target data to be classified and each of the category sets includes: According to the priority order of each category set, obtain the category set with the highest priority from each category set, where the priority of each category set is positively correlated with at least one of the quantity it stores and the established duration; Calculate the distance between the category set with the highest priority and the target data to be classified; If the calculated distance is less than the classification distance threshold, confirm that there is a category set whose distance from the target data to be classified is less than the classification distance threshold, and execute the step of storing the target data to be classified in this category set; If the calculated distance is not less than the classification distance threshold, delete the category set with the highest priority from the priority order to obtain the updated priority order of each category set, and return to execute the step of according to the priority order of each category set, obtain the category set with the highest priority from each category set; 2. The data classification method according to claim 1, characterized in that, The clustering of the data to be classified in the data set to be classified according to the classification distance threshold to obtain a plurality of category sets further includes: If there is no category set whose distance from the target data to be classified is less than the classification distance threshold, create a new category set, store the target data to be classified in the newly created category set to obtain a plurality of category sets, and return to execute the step of if target data to be classified is obtained from the data set to be classified, calculate the distances between the target data to be classified and each of the category sets; 3. The data classification method according to claim 2, characterized in that, Before calculating the distances between the data to be classified and each category set, the method further includes: If there is no category set, create a new category set, store the target data to be classified obtained from the data set to be classified into this category set, and return to execute the step of calculating the distance between the target data to be classified and each category set if the target data to be classified is obtained from the data set to be classified; If there is a category set, execute the step of calculating the distance between the target data to be classified and each category set.

4. The data classification method according to claim 2, wherein The calculating the distance between the target data to be classified and each category set includes: Obtaining the corresponding eigenvalue of the category set according to the eigenvalues of the data to be classified included in the category set; Obtaining the distance between the target data to be classified and the category set according to the eigenvalue of the target data to be classified and the eigenvalue corresponding to the category set.

5. The data classification method according to claim 4, wherein The obtaining the corresponding eigenvalue of the category set according to the eigenvalues of the data to be classified corresponding to the category set includes: Calculating the mean value of the eigenvalues of the data to be classified included in the category set to obtain the eigenvalue corresponding to this category set.

6. The data classification method according to claim 4, wherein After storing the target data to be classified into the category set, the method further includes: Updating the eigenvalue corresponding to the category set to which the target data to be classified belongs according to the eigenvalue of the target data to be classified.

7. The data classification method according to claim 2, characterized in that, The clustering the data to be classified in the data set to be classified according to the classification distance threshold to obtain multiple category sets further includes: If there is no category set with the highest priority, confirm that there is no distance between the category set and the target data to be classified that is less than the classification distance threshold, and execute the step of creating a new category set and storing the target data to be classified into this category set.

8. The data classification method according to claim 2, wherein The storing the target data to be classified into the category set if there is a distance between the category set and the target data to be classified that is less than the classification distance threshold includes: If there are at least two category sets with a distance less than the classification distance threshold from the target data to be classified, store the target data to be classified into the category set with the minimum distance from the target data to be classified; If there is one category set with a distance less than the classification distance threshold from the target data to be classified, store the target data to be classified into this category set.

9. The data classification method according to any one of claims 1 to 8, characterized in that The obtaining the classification distance threshold according to the sample distance set includes: Using a Gaussian mixture model fitting function to fit the distances included in the sample distance set to obtain a probability density function curve; Obtaining the target distance corresponding to when the probability value in the probability density function curve satisfies the specified condition, and obtaining the classification distance threshold according to the target distance.

10. The data classification method according to claim 9, characterized in that The obtaining the target distance corresponding to when the probability value in the probability density function curve satisfies the specified condition and obtaining the classification distance threshold according to the target distance includes: Obtaining the target distance corresponding to when the probability value in the probability density function curve is the minimum value, and using the target distance as the classification distance threshold.

11. The data classification method according to any one of claims 1 to 8, characterized in that Obtaining the sample distance set includes: Obtaining the eigenvalues of each sample data in the sample data set; Calculate the distance between the eigenvalue of each sample data and the eigenvalue of each sample data in the remaining sample data corresponding to the sample data, to obtain a sample distance set including the distance between every two sample data.

12. The data classification method according to claim 11, characterized in that, The calculating the distance between the eigenvalue of each sample data and the eigenvalue of each sample data in the remaining sample data corresponding to the sample data, and obtaining the sample distance set corresponding to each sample data, includes: Using the Manhattan distance calculation formula, calculate the distance between the eigenvalue of each sample data and the eigenvalue of each sample data in the remaining sample data corresponding to the sample data, to obtain a sample distance set including the distance between every two sample data.

13. A data classification device, characterized in that, The device includes: A distance acquisition module, configured to acquire a sample distance set, where the sample distance set includes the distance between every two sample data in the sample data set, and the sample data includes the feature data of a video, the feature data of an image, or the feature data of a document; A threshold obtaining module, configured to obtain a classification distance threshold according to the sample distance set; A data classification module, configured to cluster the data to be classified in the data set to be classified according to the classification distance threshold, to obtain a plurality of category sets, the data set to be classified includes the sample data set, and the data to be classified in the data set to be classified includes at least one of historical data and real-time data generated during the operation of the data system; The data classification module is further configured to, if target data to be classified is acquired from the data set to be classified, calculate the distance between the target data to be classified and each of the category sets; if there is a category set whose distance from the target data to be classified is less than the classification distance threshold, store the target data to be classified in the category set; The data classification module is further configured to, according to the priority order of each category set, obtain the category set with the highest priority from each category set, where the priority of each category set is positively correlated with at least one of the quantity it stores and the established duration; calculate the distance between the category set with the highest priority and the target data to be classified; if the calculated distance is less than the classification distance threshold, confirm that there is a category set whose distance from the target data to be classified is less than the classification distance threshold, then store the target data to be classified in the category set; if the calculated distance is not less than the classification distance threshold, delete the category set with the highest priority from the priority order to obtain the updated priority order of each category set.

14. An electronic device, characterized in that, Including a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, where, when the program code is run by the processor, the method according to any one of claims 1-12 is executed.

16. A computer program product, comprising a computer program / instructions, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Data clustering method and apparatus, computer readable medium and electronic device

    CN107067045A

  • Text clustering method, text clustering device, server and storage medium

    CN107992596A