Object classification method and device, computer equipment and storage medium

By using the scaling coefficient to calibrate the spacing between user objects and subclass centers, the classification unfairness and accuracy problems caused by differences in subclass center spacing in cluster analysis are solved, and a more fair and accurate object classification is achieved.

CN120372368APending Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410107285.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In clustering analysis, when a certain category of data objects is relatively scattered in a multi-dimensional space, the data objects of different subclasses are also scattered, which may lead to the misclassification of data objects located at the category boundaries, resulting in a decrease in classification unfairness and accuracy.

Method used

By obtaining the scaling coefficients of multiple risk categories, the first spacing between the user object and the center of the subclass is calculated, and the second spacing is calibrated using the scaling coefficient to determine the category to which the user object belongs to ensure the fairness and accuracy of the classification.

Benefits of technology

It effectively avoids misclassification problems caused by the difference in average spacing between multiple subclass centers within different risk categories, and improves the fairness and accuracy of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372368A_ABST
    Figure CN120372368A_ABST
Patent Text Reader

Abstract

The invention provides an object classification method and device, computer equipment and a storage medium, belongs to the technical field of computers, and can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic and auxiliary driving. The method comprises the following steps: for any to-be-classified user object, obtaining scaling coefficients of a plurality of risk categories; for any risk category, a plurality of first intervals of the user object under the risk category are determined, and the first intervals are the distances between the user object data of the user object and the subclass center of each subclass of the risk category; determining a plurality of second intervals of the user object under the risk category according to the plurality of first intervals and scaling coefficients of the risk category; and determining the risk category to which the minimum second interval of the user object in the plurality of risk categories belongs as the category to which the user object belongs. According to the technical scheme, the problem of unfair classification caused by the difference of the average distance between the sub-class centers can be avoided, so that the classification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to an object classification method, apparatus, computer device, and storage medium. Background Art

[0002] Cluster analysis based on clustering algorithms can play an important role in many fields. For example, in the financial field, cluster analysis can help financial institutions assess the risks of customers to provide more reasonable financial services; in the medical field, cluster analysis can help doctors classify diseases to formulate more targeted treatment plans, etc.

[0003] Currently, during cluster analysis, the data of the objects to be classified are placed in a multi-dimensional space, and different object data are classified into corresponding categories according to the degree of closeness of the spatial relationships. Further, the data of objects in different categories can be classified into multiple sub-categories to better represent the data characteristics and reflect the distribution of the data of multiple objects.

[0004] However, in the above process of cluster analysis, when the data objects of a certain category are relatively scattered in the multi-dimensional space, the data objects of different sub-categories in this category are also relatively scattered. This may cause the data objects located at the boundary of this category to be classified into other categories, resulting in unfair classification problems for the data objects of this category and a decrease in the overall classification accuracy. Summary of the Invention

[0005] Embodiments of this application provide an object classification method, apparatus, computer device, and storage medium, which can avoid the misclassification problem of objects at the boundary caused by the difference in the average distance between multiple sub-category centers within different categories, thereby improving the fairness and accuracy of classification. The technical solutions are as follows:

[0006] On the one hand, an object classification method is provided, and the method includes:

[0007] For any user object to be classified, obtain the scaling coefficients of multiple risk categories. The scaling coefficient of any risk category is the ratio of the reference distance to the average sub-category distance of the risk category. The average sub-category distance is the mean of the distances between the sub-category centers of multiple sub-categories within the risk category, and the sub-category center is the clustering center of the sub-category obtained by clustering within the risk category;

[0008] For any risk category, determine multiple first distances of the user object under the risk category. The first distance is the distance between the user object data of the user object and the sub-category centers of each sub-category of the risk category;

[0009] Determine multiple second spacings of the user object under the risk category according to the multiple first spacings and the scaling coefficient of the risk category;

[0010] Determine the risk category to which the smallest second spacing of the user object under the multiple risk categories belongs as the category to which the user object belongs.

[0011] On the other hand, an object classification device is provided, and the device includes:

[0012] An acquisition module, configured to, for any user object to be classified, acquire scaling coefficients of multiple risk categories, where the scaling coefficient of any risk category is the ratio of a reference spacing to the subclass average spacing of the risk category, and the subclass average spacing is the mean of the spacings between subclass centers of multiple subclasses within the risk category, and the subclass center is the clustering center of the subclasses clustered within the risk category;

[0013] A first determination module, configured to, for any risk category, determine multiple first spacings of the user object under the risk category, where the first spacing is the distance between the user object data of the user object and the subclass centers of each subclass of the risk category;

[0014] An operation module, configured to determine multiple second spacings of the user object under the risk category according to the multiple first spacings and the scaling coefficient of the risk category;

[0015] A second determination module, configured to determine the risk category to which the smallest second spacing of the user object under the multiple risk categories belongs as the category to which the user object belongs.

[0016] In some embodiments, the first determination module includes:

[0017] A first determination unit, configured to, for any risk category, determine the subclass centers of each subclass of the risk category;

[0018] A second determination unit, configured to, for the subclass center of any subclass, determine the spacing between the subclass center of the subclass and the user object data of the user object as the first spacing corresponding to the subclass under the risk category.

[0019] In some embodiments, the device further includes:

[0020] A third determination module, configured to, for any user training data set, determine the subclass centers of each subclass of the multiple risk categories, where the user training data set includes user sample data of the multiple risk categories;

[0021] A fourth determination module, configured to, for any risk category, determine the subclass average spacing of the risk category based on the subclass centers of each subclass of the risk category.

[0022] A fifth determination module, configured to determine the scaling factors of the multiple risk categories based on the subclass average spacings of the multiple risk categories.

[0023] In some embodiments, the third determination module is configured to, for any user training dataset, cluster the user sample data in the user training dataset to obtain the multiple risk categories; for any risk category, cluster the user sample data within the risk category to obtain subclasses with a risk quantity; for any subclass, determine the clustering center of the user sample data within the subclass as the subclass center of the subclass.

[0024] In some embodiments, the fourth determination module is configured to, for any risk category, determine the spacing between the subclass centers of each subclass of the risk category, where the number of subclasses is the risk quantity; determine the ratio of the sum of the subclass spacings of the risk category to the number of subclass spacings of the risk category as the subclass average spacing of the risk category, where the sum of the subclass spacings of the risk category is the sum of the spacings between the subclass centers of the subclasses with the risk quantity, and the number of subclass spacings is the number of spacings between the subclass centers of the subclasses with the risk quantity.

[0025] In some embodiments, the fourth determination module is configured to, for any risk category, determine the spacing between the subclass centers of each subclass of the risk category, where the number of subclasses is the risk quantity; determine a spacing matrix based on the spacing between the subclass centers of each subclass, where the number of rows and columns of the spacing matrix is the risk quantity, and the elements in the spacing matrix are the spacings between the subclass centers of each subclass of the risk category; determine the ratio of the sum of all elements in the spacing matrix to the risk quantity as the subclass average spacing of the risk category.

[0026] In some embodiments, the fifth determination module is configured to, for any risk category, determine the ratio of a reference spacing to the subclass average spacing of the risk category as the scaling factor of the risk category, where the reference spacing is the subclass average spacing of a reference category, and the reference category is any risk category among the multiple risk categories.

[0027] In some embodiments, the fifth determination module is configured to add the average distances between subclasses of the multiple risk categories to obtain the total sum of the average distances between subclasses; determine the ratio of the total sum of the average distances between subclasses to the number of risk categories as the reference distance; for any risk category, determine the ratio of the reference distance to the average distance between subclasses of the risk category as the scaling coefficient of the risk category.

[0028] On the other hand, a computer device is provided, which includes a processor and a memory. The memory is used to store at least one segment of computer program, and the at least one segment of computer program is loaded and executed by the processor to implement the object classification method in the embodiments of the present application.

[0029] On the other hand, a computer-readable storage medium is provided, in which at least one segment of computer program is stored, and the at least one segment of computer program is loaded and executed by a processor to implement the object classification method in the embodiments of the present application.

[0030] On the other hand, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the object classification method in the embodiments of the present application.

[0031] The present application provides an object classification method, which calibrates the distance between user object data and subclass centers by using a scaling coefficient, so as to ensure the classification fairness between different risk categories when classifying user objects. This method can avoid the misclassification problem of objects at the boundary caused by the difference in the average distances between multiple subclass centers within different risk categories, thereby improving the fairness and accuracy of classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 is a schematic diagram of the implementation environment of an object classification method provided according to an embodiment of the present application;

[0034] Figure 2 is a flowchart of an object classification method provided according to an embodiment of the present application;

[0035] Figure 3 is a flowchart of another object classification method provided according to an embodiment of the present application;

[0036] Figure 4 It is a block diagram of an object classification device provided according to an embodiment of the present application;

[0037] Figure 5 It is a block diagram of another object classification device provided according to an embodiment of the present application;

[0038] Figure 6 It is a schematic structural diagram of a terminal provided according to an embodiment of the present application;

[0039] Figure 7 It is a schematic structural diagram of a server provided according to an embodiment of the present application. Detailed implementation manners

[0040] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0041] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor are the quantity and execution order limited.

[0042] In the present application, the term "at least one" means one or more, and the meaning of "multiple" means two or more.

[0043] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the user training dataset and scaling factor involved in the present application are obtained under full authorization.

[0044] It should be noted that the present application involves artificial intelligence technology (Artificial Intelligence, AI). Artificial intelligence technology is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results of theory, method, technology, and application systems. When classifying objects in the present application, it involves big data processing technology and pre-trained model technology in artificial intelligence technology, etc.

[0045] Figure 1 It is a schematic diagram of the implementation environment of an object classification method provided according to an embodiment of the present application. See Figure 1, the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication means, which is not limited in this application.

[0046] In some embodiments, the terminal 101 includes, but is not limited to, various types of terminals such as mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, and aircraft. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc. An application program can be installed and run on the terminal 101. This application program can obtain a scaling coefficient and calibrate a first distance, and determine the risk category to which the user object to be classified belongs according to the obtained second distance. This application program is associated with the server 102, and the server 102 provides background services to the terminal 101.

[0047] In some embodiments, the server 102 is an independent physical server, and can also be a server cluster or a distributed system composed of multiple physical servers, and can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0048] In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing.

[0049] Figure 2 is a flowchart of an object classification method provided according to an embodiment of this application. This method is executed by a computer device. Refer to Figure 2 , this method includes the following steps:

[0050] 201. For any user object to be classified, obtain the scaling coefficients of multiple risk categories. The scaling coefficient of any risk category is the ratio of the reference distance to the subclass average distance of the risk category. The subclass average distance is the mean of the distances between the subclass centers of multiple subclasses within the risk category, and the subclass center is the clustering center of the subclasses clustered within the risk category.

[0051] In an embodiment of the present application, for any user object to be classified, the terminal obtains scaling coefficients for multiple risk categories related to the user object. Herein, the user object refers to a user conducting financial activities in the financial field, and the risk category refers to a category divided according to the credit rating of the user object. The scaling coefficient is used to calibrate the distance between the user object data and the subclass center to ensure classification fairness. For each risk category, multiple subclasses are re-clustered within the risk category, and the average value of the distances between the subclass centers of the multiple subclasses is called the subclass average distance, which can reflect the distribution status of the multiple subclasses in the risk category.

[0052] 202. For any risk category, determine multiple first distances of the user object under the risk category. The first distance is the distance between the user object data of the user object and the subclass centers of each subclass in the risk category.

[0053] In an embodiment of the present application, for any risk category, the terminal determines the distances between the user object data and the respective subclass centers under the risk category, denoted as the first distances. It should be noted that multiple methods such as Euclidean distance, Manhattan distance, and cosine similarity can be used to determine the first distances. This embodiment of the present application does not limit this, as long as the method for calculating all the first distances is the same.

[0054] 203. Determine multiple second distances of the user object under the risk category according to the multiple first distances and the scaling coefficient of the risk category.

[0055] In an embodiment of the present application, the terminal determines multiple second distances of the user object under the risk category through operations according to the multiple first distances and the scaling coefficient of the risk category. In each risk category, the above operations are repeated until multiple first distances and multiple second distances of the user object under all risk categories are determined.

[0056] 204. Determine the risk category to which the user object belongs as the risk category to which the smallest second distance of the user object under multiple risk categories belongs.

[0057] In an embodiment of the present application, the terminal can directly determine the smallest second distance from all the second distances under multiple risk categories, and then determine the risk category to which the user object belongs; it can also first determine the smallest second distance under each risk category, and then determine the smallest second distance from the smallest second distances of multiple risk categories, and then determine the risk category to which the user object belongs. This embodiment of the present application does not limit this.

[0058] An embodiment of the present application provides an object classification method. By using a scaling coefficient to calibrate the distance between user object data and subclass centers, it can ensure the classification fairness between different risk categories when classifying user objects. This method can avoid the misclassification problem of objects at the boundary caused by the difference in the average distance between multiple subclass centers within different risk categories, thereby improving the fairness and accuracy of classification.

[0059] It should be noted that the object classification method provided by the present application is used to determine the target category to which the object to be classified belongs. Determining the risk category to which the user object to be classified belongs is only an example of the present application. The present application does not limit the usage scenario of the object classification method. In other words, the present application does not limit the content of the object and the target category.

[0060] For example, when using the object classification method in the medical field, the object is a patient object, the target category is a disease category, and the object data is the disease data of the patient object; when using the object classification method in the financial field, the object is a user object, the target category is a risk category, and the object data is the credit rating data of the user object; when using the object classification method in the marketing field, the object is an enterprise object, the target category is a solution category, and the object data is the solution requirement data of the enterprise object.

[0061] Figure 3 is a flowchart of another object classification method provided by an embodiment of the present application, and this method is executed by a computer device. When using the object classification method in the financial field, in the following solution, the training data set is a user training data set, the sample data is user sample data, the target category is a risk category, the object is a user object, and the object data is the credit rating data of the user object. Refer to Figure 3 This method includes the following steps:

[0062] 301. For any training data set, determine the subclass centers of each subclass of multiple target categories. The training data set includes sample data of multiple target categories, and the subclass center is the clustering center of the subclass obtained by clustering within the target category.

[0063] In the embodiments of the present application, for any training data set, the terminal determines the subclass centers of each subclass of multiple target classes in the training data set. The training data set is used to train a classification model, and the training data set includes sample data of multiple target classes. The classification model is used to classify an object to be classified. The training data set can be an existing public data set, a data set automatically generated according to a neural network model, or a data set integrated according to relevant field information. The embodiments of the present application do not limit the field, content, form, and source of the sample data in the training data set. The multiple target classes can be obtained by clustering the sample data in the training data set according to a clustering algorithm, or by identifying the classes of the sample data in the training data set according to existing classification information. Clustering refers to dividing data with high similarity into the same class and dividing data with high dissimilarity into different classes according to the similarity principle. The embodiments of the present application do not limit the division method of the multiple target classes. The sample data refers to the relevant data of a sample object. The sample object can be various subjects in multiple fields, and the embodiments of the present application do not limit this.

[0064] For example, in the medical field, the sample object can be patient objects of different disease types; in the financial field, the sample object can be user objects with different credit ratings; in the marketing field. The sample object can be enterprise objects with different solution requirements, etc. Correspondingly, the training data set in the medical field can include sample data of various types of diseases of different sample objects; the training data set in the financial field can include sample data of various credit ratings of different sample objects; the training data set in the marketing field can include sample data of various solution requirements of different sample objects.

[0065] In some embodiments, the terminal can determine the subclass center by performing intra-class clustering. Correspondingly, for any training data set, the terminal clusters the sample data in the training data set to obtain multiple target classes; for any target class, the terminal clusters the sample data within the target class to obtain the target number of subclasses; for any subclass, the terminal determines the clustering center of the sample data within the subclass as the subclass center of the subclass. By first clustering the sample data in the training data set and then performing intra-class clustering on the sample data within the multiple target classes obtained by clustering, the subclass centers of each subclass of the multiple target classes can be determined, which is convenient for subsequent calibration of the average subclass spacing and object classification.

[0066] Taking the example of clustering the sample data in a training data set into y target classes and clustering the sample data in each target class into k subclasses, the following steps 1 to 4 are described. Where y and k are both positive integers, and k is the above-mentioned target number.

[0067] Step 1: Perform data preprocessing on the sample data in the training dataset. It should be noted that when performing data preprocessing on the sample data in the training dataset, data cleaning, missing value handling, and outlier handling and other processing methods can be used to ensure the high quality and consistency of the sample data; the sample data can also be standardized or normalized to make different data features comparable. The embodiments of the present application do not limit the specific methods of data preprocessing.

[0068] Among them, data cleaning is a method for inspecting and proofreading data. The purpose of data cleaning is to filter out data that does not meet the requirements. The data that does not meet the requirements mainly includes three categories: incomplete data, incorrect data, and duplicate data. Data cleaning requires deleting duplicate data, correcting incorrect data, and ensuring data consistency. Missing value handling refers to handling the data lost during the data collection and collation process. A missing value refers to the incomplete attribute value of data in a dataset due to the lack of information. There are various reasons for the generation of missing values, which can be mechanical reasons, such as memory damage, or mechanical failures in data collection; or human reasons, such as invalid relevant problem information, or omission of data entry. There are various methods for handling missing values, which can be deleting the relevant data containing missing values, which may seriously affect the objectivity and comprehensiveness of the dataset in a dataset with a small amount of data; or imputing the missing values, mainly including methods such as mean imputation, similar mean imputation, multiple imputation, and maximum likelihood estimation. Outlier handling refers to handling abnormal data during the data collection and collation process. An outlier refers to data that is significantly different from other data in a dataset, such as data that deviates from the normal range. There are various reasons for the generation of outliers, including measurement errors, data input errors, system failures, or real special situations, etc. Since the existence of outliers in data analysis may cause result deviations, it is necessary to handle them. There are various common methods for handling outliers, which can be directly deleting the relevant data containing outliers; or treating the outliers as missing values and using the methods for handling missing values for processing; or correcting the outliers according to the average value of the previous and next two observations. The embodiments of the present application do not elaborate on the above methods.

[0069] Step 2: Perform class labeling on each sample data in the training dataset, that is, assign the corresponding class label y to each sample data. The purpose of performing class labeling on the sample data is to divide the sample data in the training dataset according to different target classes, and to prepare for subsequent intra-class clustering within different target classes. For example, for the training dataset in the medical field, according to the existing classification information, add class labels indicating different disease types to different sample data; for the training dataset in the financial field, according to the existing classification information, add class labels indicating different credit ratings to different sample data; for the training dataset in the marketing field, according to the existing classification information, add class labels indicating different solution requirements to different sample data.

[0070] Step 3: Perform intra-class clustering on the sample data within each target class respectively, that is, cluster the sample data within each target class into k subclasses. Among them, k is a pre-set parameter. It should be noted that during the intra-class clustering process, methods such as cross-validation (Cross Validation) and silhouette coefficient (Silhouette Coefficient) can be used to determine the appropriate value of k; the appropriate clustering algorithm can be determined according to the characteristics and requirements of the sample data and reasonable parameters can be set. For example, use clustering algorithms such as K-means algorithm (K-means clustering algorithm), DBSCAN algorithm (Density-Based Spatial Clustering of Applications with Noise) or hierarchical clustering (Hierarchical Clustering) to achieve intra-class clustering. For the convenience of description, after completing intra-class clustering in the target class, assign a unique subclass label to each obtained subclass for description and analysis in subsequent steps.

[0071] Among them, cross-validation is a method of cutting sample data into smaller subclasses. First, the sample data is grouped into a training set and a validation set. The classification model is trained with the training set, and the trained classification model is tested with the validation set, which is used as a performance indicator to evaluate the classification model. The silhouette coefficient can be used to evaluate the impact of different algorithms or different running methods on the clustering results based on the same sample data. The silhouette coefficient comprehensively considers the similarity of the sample data within its own category and the dissimilarity with the nearest other categories. The K-means algorithm is a clustering algorithm based on Euclidean distance. The K-means algorithm uses distance as the standard for measuring the similarity between data, that is, the smaller the distance between data, the higher the similarity between data, and thus the more likely they are to be classified into the same category. The DBSCAN algorithm is a density-based clustering algorithm. The DBSCAN algorithm can determine multiple dense regions of the data and determine the above-mentioned multiple dense regions as different clustering categories. Hierarchical clustering creates a hierarchical nested tree by calculating the similarity between different categories. Hierarchical clustering assumes that there is a hierarchical structure between categories and clusters the data into hierarchical categories. The embodiments of the present application will not elaborate on the above methods.

[0072] Step 4: After performing intra-class clustering on the sample data within each target category respectively, determine the clustering center of the sample data within the subclass as the subclass center of the subclass. Among them, the subclass center can be represented by the mean of the sample data within the subclass. For ease of description, denote the subclass label of the kth subclass in the target category y as y k , correspondingly, denote the subclass center of the above subclass as y k c .

[0073] 302. For any target category, based on the subclass centers of each subclass of the target category, determine the subclass average spacing of the target category. The subclass average spacing is the mean of the spacings between the subclass centers of multiple subclasses within the target category.

[0074] In the embodiments of the present application, the terminal determines the subclass average spacings of multiple target categories respectively according to the subclass centers of each subclass of the multiple target categories determined in the above steps. It should be noted that the subclass average spacing is a description method adopted for ease of description. Similarly, the subclass average spacing can be expressed as the subclass center average spacing, the subclass core average spacing, and the subclass center core average spacing, etc., and the meaning it indicates is the mean of the spacings between the subclass centers of multiple subclasses. The embodiments of the present application do not limit this. For ease of description, denote the subclass average spacing as

[0075] First, for any target category, when determining the spacing between subclass centers, calculate the distances between each subclass center and other subclass centers respectively. It should be noted that the measurement methods that can be used to calculate the spacing between subclass centers include various methods such as Euclidean distance, Manhattan distance, and cosine similarity, which can be determined according to the characteristics and requirements of the sample data, and the embodiments of the present application do not limit this. For example, when using the Euclidean distance to calculate the spacing between subclass centers, the expression form of the L2 norm is used, and the relevant formula is as follows:

[0076] d(x,y k c )=||x - y k c ||2

[0077] where x is any subclass center, and y k c is another subclass center, and d(x,y k c ) is the spacing between the above two subclass centers. The L2 norm is defined as the square root of the sum of the squares of all elements.

[0078] Secondly, for any target category, after determining the spacing between subclass centers, the mean value of the spacing between subclass centers can be used as the subclass average spacing, so as to reflect the distribution of subclass centers within the target category and provide a basis for subsequent spacing calibration.

[0079] In some embodiments, the terminal directly determines the subclass average spacing according to the spacing between each subclass center of the target category. Correspondingly, for any target category, the terminal determines the spacing between the subclass centers of each subclass based on the subclass centers of each subclass of the target category, and the number of subclasses is the target number; the terminal determines the ratio of the sum of the subclass spacings of the target category to the number of subclass spacings of the target category as the subclass average spacing of the target category. The sum of the subclass spacings of the target category is the sum of the spacings between the subclass centers of the subclasses with the target number, and the number of subclass spacings is the number of spacings between the subclass centers of the subclasses with the target number. By first determining the spacing between each subclass center of the target category and then determining the ratio of the sum of the spacings between subclass centers to the number of spacings as the subclass center spacing of the target category, it is convenient to determine the scaling factor according to the subclass average spacings of multiple categories later.

[0080] In some embodiments, the terminal may use a spacing matrix to record the spacing between the centers of each subclass. Correspondingly, for any target category, the terminal determines the spacing between the centers of each subclass based on the centers of each subclass of the target category, and the number of subclasses is the target number; the terminal determines a spacing matrix based on the spacing between the centers of each subclass, and the number of rows and columns of the spacing matrix is the target number, and the elements in the spacing matrix are the spacing between the centers of each subclass of the target category; the terminal determines the ratio of the sum of all elements in the spacing matrix to the target number as the average subclass spacing of the target category. By using a spacing matrix to record the spacing between the centers of each subclass, the average subclass spacing of any target category can be quickly determined according to the elements in the spacing matrix, which is convenient for subsequently determining the scaling factor based on the average subclass spacing of multiple categories.

[0081] It should be noted that since the average subclass spacing exists as an intermediate value in the subsequent process of determining the scaling factor, the embodiments of the present application do not limit the method for determining the average subclass spacing, as long as it can reflect the differences in the average subclass spacing of different categories.

[0082] For example, when there are 4 subclasses in the target category, that is, the k value is 4. The spacings between the centers of the subclasses are a, b, c, d, e, f respectively, and the spacing matrix is At this time, the methods for determining the spacing between the centers of the subclasses include but are not limited to any one of the following Method 1 to Method 4.

[0083] Method 1: Determine the ratio of the sum of the spacings between the centers of the subclasses to the number of spacings as the spacing between the centers of the subclasses of the target category. At this time, the number of spacings between the centers of the subclasses is 6, and the spacing between the centers of the subclasses of the target category is (a + b + c + d + e + f) / 6.

[0084] Method 2: Determine the ratio of the sum of all elements in the spacing matrix to the target number as the average subclass spacing of the target category. At this time, the target number is 4, and the spacing between the centers of the subclasses of the target category is 2*(a + b + c + d + e + f) / 4.

[0085] Method 3: Determine the ratio of the sum of all elements in the spacing matrix to the number of elements as the average subclass spacing of the target category. At this time, the number of elements is 4*4 = 16, and the spacing between the centers of the subclasses of the target category is 2*(a + b + c + d + e + f) / 16.

[0086] Method 4: When considering that the elements at the diagonal positions in the spacing matrix are all 0, in order to avoid the interference of elements with a value of 0 on the result, the ratio of the sum of all elements in the spacing matrix to the number of target elements is determined as the average spacing of subclasses of the target category. At this time, the number of target elements is 4 * 3 = 12, and the average spacing of subclasses of the target category is 2 * (a + b + c + d + e + f) / 12.

[0087] Optionally, the relevant formulas of the above Method 3 are summarized as follows:

[0088]

[0089] Among them, is the average spacing of subclasses of the target category, x is the center of any subclass in the above target category, y k c is the center of another subclass in the above target category, d(x, y k c ) is the spacing between the above two subclass centers, and k is the number of targets, that is, the number of rows and columns of the spacing matrix.

[0090] It should be noted that after determining the average spacing of subclasses of multiple target categories, the average spacing of subclasses of multiple target categories can be statistically analyzed. For example, using various forms such as tables, bar charts, and line charts to display the average spacing of subclasses of multiple target categories helps to understand the differences in the average spacing of subclasses between multiple target categories and provides a basis for subsequent distance calibration. The embodiments of the present application do not limit the display method of the average spacing of subclasses.

[0091] 303. Based on the average spacing of subclasses of multiple target categories, determine the scaling coefficients of multiple target categories. The scaling coefficient of any target category is the ratio of the reference spacing to the average spacing of subclasses of the target category.

[0092] In the embodiments of the present application, the terminal determines the scaling coefficients of multiple target categories according to the average spacing of subclasses of multiple target categories. It should be noted that the embodiments of the present application do not limit the method for determining the scaling coefficient, as long as after calibrating the average spacing of subclasses of multiple target categories using the scaling coefficient, the average spacing of subclasses of the above multiple target categories is the same. For the convenience of description, the scaling coefficient is denoted as t.

[0093] In some embodiments, the terminal determines the scaling factor of a target category based on the average spacing between subcategories of the target category and a reference spacing. Accordingly, for any target category, the terminal determines the ratio of the reference spacing to the average spacing between subcategories of the target category as the scaling factor of the target category. The reference spacing is the average spacing between subcategories of a reference category, and the reference category is any one of the multiple target categories. By obtaining the scaling factors of multiple target categories based on the reference spacing and the average spacing between subcategories of the multiple target categories, the spacing between the object data to be classified and the subcategory center can be calibrated using the scaling factor in subsequent steps to ensure classification fairness.

[0094] In some embodiments, the terminal determines the reference spacing based on multiple average spacings between subcategories. Accordingly, the terminal adds up the average spacings between subcategories of multiple target categories to obtain the total sum of the average spacings between subcategories; the terminal determines the ratio of the total sum of the average spacings between subcategories to the number of target categories as the reference spacing; for any target category, the terminal determines the ratio of the reference spacing to the average spacing between subcategories of the target category as the scaling factor of the target category. By using the ratio of the total sum of the average spacings between subcategories to the number of targets as the reference spacing and then determining the scaling factors of multiple target categories, it is convenient to use the scaling factor to calibrate the spacing between the object data to be classified and the subcategory center in subsequent steps to ensure classification fairness.

[0095] For example, there are three target categories M, N, and Q in the training dataset, and the average spacings between subcategories of the above three target categories are m, n, and q. At this time, the methods for determining the scaling factor of the target category include but are not limited to any one of the following Method 1 and Method 2.

[0096] Method 1: Select any one of the multiple target categories as the reference category, and then determine the scaling factor. At this time, take target category M as the reference category and m as the reference spacing, then the scaling factors of the above three target categories are 1, m / n, and q / n respectively.

[0097] Method 2: Determine the reference spacing based on multiple average spacings between subcategories, and then determine the scaling factor. At this time, the reference spacing is (m + n + q) / 3, then the scaling factors of the above three target categories are (m + n + q) / 3m, (m + n + q) / 3n, and (m + n + q) / 3q respectively.

[0098] In some embodiments, after determining the initial scaling factors of each target category, a reference category can be determined from multiple target categories, and the initial scaling factors of other target categories are compared with the initial scaling factor of the reference category, and then the adjusted scaling factor of each target category is determined.

[0099] Optionally, the relevant formulas of Method 2 are summarized as follows:

[0100]

[0101] Among them, is the average distance between subclasses of the target category, is the average distance between subclasses of the target category for which the scaling factor is to be determined, |y| is the number of target categories, is the scaling factor of the target category for which the scaling factor is to be determined.

[0102] 304. For any object to be classified, obtain the scaling factors of multiple target categories.

[0103] In the embodiments of the present application, when an object to be classified needs to be classified, the terminal obtains the scaling factors of multiple target categories. Among them, the object to be classified can be various types of objects. For example, the object to be classified can be a user object for which the risk category needs to be determined, or an enterprise object for which the solution category needs to be determined, or an animal object for which the gender needs to be determined, etc. It should be noted that the terminal can request and obtain the scaling factor from the server, or directly obtain the scaling factor from local storage. The embodiments of the present application do not limit the acquisition method of the scaling factor.

[0104] 305. For any target category, determine multiple first distances of the object under the target category. The first distance is the distance between the object data of the object and the subclass centers of each subclass of the target category.

[0105] In the embodiments of the present application, for any target category, the terminal determines the distance between the object data and each subclass center under the target category, denoted as the first distance. It should be noted that when calculating multiple first distances, the measurement methods that can be used include various methods such as Euclidean distance, Manhattan distance, and cosine similarity. It needs to be determined according to the characteristics and requirements of the actual data. The embodiments of the present application do not limit this, as long as the method for calculating all first distances is the same.

[0106] In some embodiments, the terminal determines multiple first distances according to each subclass center. Correspondingly, for any target category, the terminal determines the subclass centers of each subclass of the target category; for the subclass center of any subclass, the terminal determines the distance between the subclass center of the subclass and the object data of the object as the first distance corresponding to the subclass under the target category. By first determining the subclass centers of each subclass and then determining the distances between each subclass center and the object data, it is convenient to calibrate the multiple first distances subsequently to ensure classification fairness.

[0107] 306. Multiply the multiple first distances by the scaling factor of the target category respectively to obtain multiple second distances of the object under the target category.

[0108] In the embodiments of the present application, the terminal multiplies multiple first spacings by the scaling coefficients of the target categories respectively to obtain multiple second spacings of the object under the target categories. In each target category, the above operation is repeated until multiple first spacings and multiple second spacings of the object under all target categories are determined. It should be noted that after calibrating the multiple first spacings under each target category with the scaling coefficient of that target category, the obtained multiple second spacings can be used as the classification basis for object data.

[0109] 307. Determine the target category to which the smallest second spacing of the object under multiple target categories belongs as the category to which the object belongs.

[0110] In the embodiments of the present application, the terminal compares all the second spacings of the object under multiple target categories to determine the smallest second spacing, that is, to determine the subclass center closest to the object data, and assigns the object data to the target category to which the subclass center belongs, so as to classify the object data more fairly and accurately. Optionally, the terminal compares the multiple second spacings in each target category to determine the smallest second spacing in that target category, that is, to determine the subclass center closest to the object data in that target category, and then compares the multiple smallest second spacings under multiple target categories to determine the smallest second spacing among them, that is, to determine the closest subclass center from the multiple subclass centers closest to the object data under multiple target categories, and assigns the object data to the target category to which the subclass center belongs, so as to classify the object data more fairly and accurately.

[0111] Optionally, the relevant formula for the smallest second spacing determined in each target category is summarized as follows:

[0112]

[0113] Where x is the object data, y k c is the subclass center in the target category, and d(x, y k c ) is the spacing between the calibrated object data and the subclass center, that is, the second spacing, is the smallest second spacing in the target category.

[0114] Optionally, the relevant formula for determining the smallest second spacing from the multiple smallest second spacings under multiple target categories is summarized as follows:

[0115]

[0116] Where is the smallest second spacing in the target category, is the category to which the object belongs.

[0117] The method provided by the embodiment of the present application for automatically calibrating the average spacing of subclasses to achieve fair classification of objects has various beneficial effects. It can not only better adapt to the distribution of real data, improve the classification accuracy of the classification model, but also eliminate the unfairness between different classes, and at the same time can be widely applied to various clustering-based interpretable algorithms. First, by clustering within the class, the subclasses in each class can be divided according to the characteristics of the sample data, so as to better reflect the distribution of real data, which helps to improve the performance of the classification model in various classification tasks and provide more accurate classification results for users. Second, in traditional classification methods, due to the fixed number of subclasses, the average spacing of subclasses in some classes may be relatively large, resulting in unfairness when the classification model processes the object data of the above classes. By automatically adjusting the subclass spacing according to the actual data distribution, the fairness between different classes can be ensured, which helps to improve the reliability of the classification model in various application scenarios. In addition, by applying this method to the existing classification system, the performance and accuracy of the system can be improved. At the same time, this method has good interpretability, which is convenient for users to understand and interpret the output of the classification model, thereby improving the trust and acceptance of users in the classification model.

[0118] For example, in the medical field, this method can help doctors classify diseases more accurately, so as to provide more effective treatment plans for patients; in the financial field, this method can help financial institutions evaluate the risks of customers more accurately and improve customer satisfaction; in the marketing field, this method can help enterprises segment customers more accurately, so as to formulate more targeted marketing strategies.

[0119] The embodiment of the present application provides an object classification method. By using a scaling factor to calibrate the spacing between object data and subclass centers, the classification fairness between different target classes can be ensured when classifying objects. This method can avoid the misclassification problem of objects at the boundary caused by the difference in the average spacing between multiple subclass centers within different target classes, thereby improving the fairness and accuracy of classification.

[0120] Figure 4 It is a block diagram of an object classification device provided by the embodiment of the present application. This device is used to execute the steps when the above object classification method is executed. Refer to Figure 4 As shown in the figure, the object classification device includes: an acquisition module 401, a first determination module 402, an operation module 403, and a second determination module 404.

[0121] An acquisition module 401 is configured to acquire scaling coefficients of multiple risk categories for any user object to be classified. The scaling coefficient of any risk category is the ratio of a reference interval to the average sub-interval of the risk category. The average sub-interval is the mean of the distances between the sub-centers of multiple sub-classes within the risk category, and the sub-center is the clustering center of the sub-classes clustered within the risk category.

[0122] A first determination module 402 is configured to determine, for any risk category, multiple first intervals of the user object under the risk category. The first interval is the distance between the user object data of the user object and the sub-centers of the respective sub-classes of the risk category.

[0123] An operation module 403 is configured to determine multiple second intervals of the user object under the risk category according to the multiple first intervals and the scaling coefficient of the risk category.

[0124] A second determination module 404 is configured to determine the risk category to which the smallest second interval of the user object under multiple risk categories belongs as the category to which the user object belongs.

[0125] In some embodiments, Figure 5 is a block diagram of another object classification device provided according to an embodiment of the present application. Refer to Figure 5 As shown, the first determination module 402 includes:

[0126] A first determination unit 501 is configured to determine, for any risk category, the sub-centers of the respective sub-classes of the risk category.

[0127] A second determination unit 502 is configured to determine, for the sub-center of any sub-class, the interval between the sub-center of the sub-class and the user object data of the user object as the first interval corresponding to the sub-class under the risk category.

[0128] In some embodiments, the device further includes:

[0129] A third determination module 503 is configured to determine, for any user training data set, the sub-centers of the respective sub-classes of multiple risk categories. The user training data set includes user sample data of multiple risk categories.

[0130] A fourth determination module 504 is configured to determine, for any risk category, the average sub-interval of the risk category based on the sub-centers of the respective sub-classes of the risk category.

[0131] A fifth determination module 505 is configured to determine the scaling coefficients of multiple risk categories based on the average sub-intervals of multiple risk categories.

[0132] In some embodiments, the third determination module 503 is configured to, for any user training dataset, cluster the user sample data in the user training dataset to obtain multiple risk categories; for any risk category, cluster the user sample data within the risk category to obtain subcategories with a risk quantity; for any subcategory, determine the clustering center of the user sample data within the subcategory as the subcategory center of the subcategory.

[0133] In some embodiments, the fourth determination module 504 is configured to, for any risk category, determine the distance between the subcategory centers of each subcategory based on the subcategory centers of each subcategory of the risk category, where the number of subcategories is the risk quantity; determine the ratio of the sum of the subcategory distances of the risk category to the number of subcategory distances of the risk category as the average subcategory distance of the risk category, where the sum of the subcategory distances of the risk category is the sum of the distances between the subcategory centers of the subcategories with the risk quantity, and the number of subcategory distances is the number of distances between the subcategory centers of the subcategories with the risk quantity.

[0134] In some embodiments, the fourth determination module 504 is configured to, for any risk category, determine the distance between the subcategory centers of each subcategory based on the subcategory centers of each subcategory of the risk category, where the number of subcategories is the risk quantity; determine a distance matrix based on the distances between the subcategory centers of each subcategory, where the number of rows and columns of the distance matrix is the risk quantity, and the elements in the distance matrix are the distances between the subcategory centers of each subcategory of the risk category; determine the ratio of the sum of all elements in the distance matrix to the risk quantity as the average subcategory distance of the risk category.

[0135] In some embodiments, the fifth determination module 505 is configured to, for any risk category, determine the ratio of a reference distance to the average subcategory distance of the risk category as the scaling coefficient of the risk category, where the reference distance is the average subcategory distance of a reference category, and the reference category is any one of the multiple risk categories.

[0136] In some embodiments, the fifth determination module 505 is configured to add up the average subcategory distances of multiple risk categories to obtain the total sum of the average subcategory distances; determine the ratio of the total sum of the average subcategory distances to the number of risk categories as the reference distance; for any risk category, determine the ratio of the reference distance to the average subcategory distance of the risk category as the scaling coefficient of the risk category.

[0137] An object classification device provided by the present application calibrates the distance between user object data and subcategory centers by using a scaling coefficient, so as to ensure the classification fairness between different risk categories when classifying user objects. This device can avoid the misclassification problem of objects at the boundary caused by the difference in the average distance between multiple subcategory centers within different risk categories, thereby improving the fairness and accuracy of classification.

[0138] It should be noted that: when the object classification device provided in the above embodiments runs an application program, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the terminal is divided into different functional modules to complete all or part of the functions described above. In addition, the object classification device provided in the above embodiments and the embodiments of the object classification method belong to the same concept. For the specific implementation process, see the method embodiments and will not be elaborated here.

[0139] Figure 6 It is a schematic structural diagram of a terminal provided according to an embodiment of the present application. The terminal 600 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 600 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0140] Generally, the terminal 600 includes: a processor 601 and a memory 602.

[0141] The processor 601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0142] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 601 to implement the object classification method provided in the method embodiments of the present application.

[0143] In some embodiments, the terminal 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, and a power supply 608.

[0144] The peripheral device interface 603 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0145] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. In some embodiments, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 604 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0146] The display screen 605 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on or above the surface of the display screen 605. The touch signals can be input to the processor 601 as control signals for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, which is disposed on the front panel of the terminal 600; in other embodiments, there may be at least two display screens 605, which are respectively disposed on different surfaces of the terminal 600 or are in a foldable design; in other embodiments, the display screen 605 may be a flexible display screen, which is disposed on the curved surface or the folding surface of the terminal 600. Even further, the display screen 605 can be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 605 can be prepared using materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0147] The camera module 606 is used to capture images or videos. In some embodiments, the camera module 606 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which can be any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement functions such as background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting functions or other fused shooting functions. In some embodiments, the camera module 606 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0148] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 601 for processing, or input to the radio frequency circuit 604 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 607 may further include a headphone jack.

[0149] The power supply 608 is used to supply power to each component in the terminal 600. The power supply 608 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 608 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0150] In some embodiments, the terminal 600 further includes one or more sensors 609. The one or more sensors 609 include but are not limited to: an acceleration sensor 610, a gyroscope sensor 611, a pressure sensor 612, an optical sensor 613, and a proximity sensor 614.

[0151] The acceleration sensor 610 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 600. For example, the acceleration sensor 610 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 601 can control the display screen 605 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 610. The acceleration sensor 610 can also be used for collecting game or user movement data.

[0152] The gyroscope sensor 611 can detect the body direction and rotation angle of the terminal 600. The gyroscope sensor 611 can cooperate with the acceleration sensor 610 to collect the 3D actions of the user on the terminal 600. According to the data collected by the gyroscope sensor 611, the processor 601 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0153] The pressure sensor 612 can be disposed on the side frame of the terminal 600 and / or the lower layer of the display screen 605. When the pressure sensor 612 is disposed on the side frame of the terminal 600, it can detect the holding signal of the user on the terminal 600, and the processor 601 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 612. When the pressure sensor 612 is disposed on the lower layer of the display screen 605, the processor 601 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 605. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0154] The optical sensor 613 is used to collect the ambient light intensity. In one embodiment, the processor 601 can control the display brightness of the display screen 605 according to the ambient light intensity collected by the optical sensor 613. Optionally, when the ambient light intensity is high, the display brightness of the display screen 605 is increased; when the ambient light intensity is low, the display brightness of the display screen 605 is decreased. In another embodiment, the processor 601 can also dynamically adjust the shooting parameters of the camera module 609 according to the ambient light intensity collected by the optical sensor 613.

[0155] The proximity sensor 614, also known as a distance sensor, is disposed on the front panel of the terminal 600. The proximity sensor 614 is used to collect the distance between the user and the front of the terminal 600. In one embodiment, when the proximity sensor 614 detects that the distance between the user and the front of the terminal 600 is gradually decreasing, the processor 601 controls the display screen 605 to switch from the lit state to the off state; when the proximity sensor 614 detects that the distance between the user and the front of the terminal 600 is gradually increasing, the processor 601 controls the display screen 605 to switch from the off state to the lit state.

[0156] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the terminal 600, and it may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout.

[0157] Figure 7It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 700 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 701 and one or more memories 702. Among them, at least one computer program is stored in the memory 702, and the at least one computer program is loaded and executed by the processor 701 to implement the object classification method provided by each of the above method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0158] An embodiment of the present application also provides a computer-readable storage medium, in which at least one segment of computer program is stored, and the at least one segment of computer program is loaded and executed by a processor to implement the object classification method in the above embodiment. For example, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0159] An embodiment of the present application also provides a computer program product, including a computer program, and the computer program is executed by a processor to implement the object classification method in the embodiment of the present application.

[0160] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0161] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for object classification, characterized in that The method includes: For any user object to be classified, obtain the scaling coefficients of multiple risk categories. The scaling coefficient of any risk category is the ratio of the reference spacing to the subclass average spacing of the risk category. The subclass average spacing is the mean of the spacings between the subclass centers of multiple subclasses within the risk category, and the subclass center is the clustering center of the subclasses clustered within the risk category. For any risk category, determine multiple first spacings of the user object under the risk category. The first spacing is the distance between the user object data of the user object and the subclass centers of each subclass of the risk category. According to the multiple first spacings and the scaling coefficient of the risk category, determine multiple second spacings of the user object under the risk category. Determine the risk category to which the smallest second spacing of the user object under the multiple risk categories belongs as the category to which the user object belongs.

2. The method according to claim 1, wherein The step of, for any risk category, determining multiple first spacings of the user object under the risk category includes: For any risk category, determine the subclass centers of each subclass of the risk category. For the subclass center of any subclass, determine the spacing between the subclass center of the subclass and the user object data of the user object as the first spacing corresponding to the subclass under the risk category.

3. The method according to claim 1, wherein The method further includes: For any user training dataset, determine the subclass centers of each subclass of the multiple risk categories. The user training dataset includes the user sample data of the multiple risk categories. For any risk category, based on the subclass centers of each subclass of the risk category, determine the subclass average spacing of the risk category. Based on the subclass average spacings of the multiple risk categories, determine the scaling coefficients of the multiple risk categories.

4. The method according to claim 3, wherein The step of, for any user training dataset, determining the subclass centers of each subclass of the multiple risk categories includes: For any user training dataset, cluster the user sample data in the user training dataset to obtain the multiple risk categories. For any risk category, cluster the user sample data within the risk category to obtain a certain number of subclasses. For any subclass, determine the clustering center of the user sample data within the subclass as the subclass center of the subclass.

5. The method according to claim 3, wherein The step of, for any risk category, based on the subclass centers of each subclass of the risk category, determining the subclass average spacing of the risk category includes: For any risk category, based on the subclass centers of each subclass of the risk category, determine the spacing between the subclass centers of each subclass. The number of subclasses is a certain number. Determine the ratio of the sum of the subclass spacings of the risk category to the number of subclass spacings of the risk category as the subclass average spacing of the risk category. The sum of the subclass spacings of the risk category is the sum of the spacings between the subclass centers of the certain number of subclasses, and the number of subclass spacings is the number of spacings between the subclass centers of the certain number of subclasses.

6. The method according to claim 3, wherein For any risk category, determining the subclass average spacing of the risk category based on the subclass centers of each subclass of the risk category includes: For any risk category, determining the spacing between the subclass centers of each subclass based on the subclass centers of each subclass of the risk category, where the number of subclasses is the number of risks; Based on the spacing between the subclass centers of each subclass, determining a spacing matrix, where the number of rows and columns of the spacing matrix is the number of risks, and the elements in the spacing matrix are the spacing between the subclass centers of each subclass of the risk category; Determining the ratio of the sum of all elements in the spacing matrix to the number of risks as the subclass average spacing of the risk category.

7. The method according to claim 3, characterized in that, Determining the scaling factors of the multiple risk categories based on the subclass average spacings of the multiple risk categories includes: For any risk category, determining the ratio of a reference spacing to the subclass average spacing of the risk category as the scaling factor of the risk category, where the reference spacing is the subclass average spacing of a reference category, and the reference category is any one of the multiple risk categories.

8. The method according to claim 3, wherein Determining the scaling factors of the multiple risk categories based on the subclass average spacings of the multiple risk categories includes: Adding the subclass average spacings of the multiple risk categories to obtain the total subclass average spacing; Determining the ratio of the total subclass average spacing to the number of risk categories as the reference spacing; For any risk category, determining the ratio of the reference spacing to the subclass average spacing of the risk category as the scaling factor of the risk category.

9. An object classification device, characterized in that, The apparatus includes: An acquisition module, configured to acquire the scaling factors of multiple risk categories for any user object to be classified. The scaling factor of any risk category is the ratio of a reference spacing to the subclass average spacing of the risk category. The subclass average spacing is the mean of the spacings between the subclass centers of multiple subclasses within the risk category, and the subclass center is the clustering center of the subclasses clustered within the risk category; A first determination module, configured to determine, for any risk category, multiple first spacings of the user object under the risk category, where the first spacing is the distance between the user object data of the user object and the subclass centers of each subclass of the risk category; An operation module, configured to determine multiple second spacings of the user object under the risk category according to the multiple first spacings and the scaling factor of the risk category; A second determination module, configured to determine the risk category to which the smallest second spacing of the user object under the multiple risk categories belongs as the category to which the user object belongs.

10. A computer device, characterized in that, The computer device includes a processor and a memory. The memory is configured to store at least one segment of computer program, and the at least one segment of computer program is loaded and executed by the processor to perform the object classification method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store at least one segment of computer program, and the at least one segment of computer program is used to perform the object classification method according to any one of claims 1 to 8.