Credit user clustering method and device, electronic equipment and storage medium
By combining the K-mean and density clustering algorithms, the number of cluster clusters is determined using the elbow rule and the core points are selected, the problem of low accuracy of the single K-mean clustering algorithm is solved, and the accuracy of credit user clustering is improved.
Patent Information
- Application Number
- CN202510558179.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
AI Technical Summary
The existing single K mean clustering algorithm has the problem of low accuracy of clustering results in credit user clustering.
Combining the K-mean clustering algorithm and density clustering algorithm, the number of cluster clusters for K-mean clustering is determined through the elbow rule, and the core points of the density clustering algorithm are selected based on the neighborhood distance threshold and the neighborhood point threshold, and iterative clustering is performed to obtain the clustering results of credit users.
The accuracy of credit user clustering is improved, and the advantages of density clustering algorithm processing arbitrary shape data are fully utilized, which makes up for the limitations of the K-mean clustering algorithm on data shapes, and achieves more refined clustering results.
Smart Images

Figure CN120470348A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology or other related technical fields, and in particular to a method for clustering credit users, a device thereof, an electronic device, and a storage medium. Background Art
[0002] In today's data-driven financial environment, financial institutions and other lending institutions face unprecedented challenges and opportunities. With the development of FinTech and the digitization of consumer behavior, financial institutions have access to vast amounts of data on users' transaction records, credit scores, social network information, and more. This data is not only massive in volume but also diverse in form, encompassing multiple dimensions, including users' financial status, spending habits, and social influence, providing financial institutions with rich and in-depth insights.
[0003] However, processing and analyzing such large amounts of heterogeneous data to extract valuable insights is a complex task. Traditional statistical methods and simple rules often fall short in capturing complex patterns within the data. In this context, cluster analysis, an unsupervised learning method, has gained popularity due to its ability to automatically identify structures and patterns within datasets. Clustering technology allows financial institutions to categorize credit customers based on their shared attributes and behavioral patterns, enabling refined risk management and personalized services.
[0004] Through clustering, financial institutions can categorize credit users into different groups based on their credit risk characteristics, such as high-risk, medium-risk, and low-risk groups. Based on the clustering results, they can adopt different approval standards and monitoring mechanisms for groups with different risk levels, thereby reducing bad debt rates and improving the security and liquidity of credit assets.
[0005] In the related art, users are clustered using a single K-means clustering algorithm, which iteratively searches for the centroids of K clusters in the data to minimize the sum of the squared distances from the data points within the cluster to the centroids. Although it is widely used due to its high computational efficiency, the uncertainty of the number of clusters K in the K-means clustering algorithm leads to uncertainty in the clustering results. If the K value is not selected appropriately, the clustering results may be less accurate.
[0006] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0007] The embodiments of the present invention provide a method for clustering credit users, an apparatus thereof, an electronic device, and a storage medium, in order to at least solve the technical problem in related technologies of clustering credit users based on a single K-means clustering algorithm, which results in low accuracy of clustering results.
[0008] According to one aspect of an embodiment of the present invention, a method for clustering credit users is provided, comprising: acquiring multidimensional credit data of each credit user to obtain credit sample data, and preprocessing the credit sample data; determining the number of clusters of the K-means clustering algorithm based on the elbow rule, and iteratively clustering the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm; calculating the neighborhood distance threshold and neighborhood point number threshold of each cluster according to the first clustering result; selecting the core point of the density clustering algorithm based on the neighborhood distance threshold and neighborhood point number threshold of each cluster, and clustering the sample points in the preprocessed credit sample data based on each core point to obtain the clustering result of the credit user.
[0009] Furthermore, the step of preprocessing the multidimensional credit data includes: calculating the distance value between two sample points based on the credit sample data, and calculating the circle radius based on the distance values between all the sample points; calculating the decision threshold value based on the total number of the sample points; calculating the density value of each sample point according to the circle radius, and comparing the density value of each sample point with the decision threshold value to obtain a comparison result; determining the outlier sample point based on the comparison result, and deleting the outlier sample point from the credit sample data.
[0010] Furthermore, the step of determining the number of clusters of the K-means clustering algorithm based on the elbow rule includes: step one, determining the value range of the number of clusters based on the total number of sample points in the credit sample data; step two, randomly selecting a number of clusters within the value range, and clustering the preprocessed credit sample data based on the selected number of clusters using the K-means clustering algorithm; step three, calculating the intra-class distance corresponding to the number of clusters; step four, repeating the above steps two to three, selecting a different number of clusters each time, until all the number of clusters within the value range are selected, and the intra-class distance corresponding to each number of clusters is obtained; step five, drawing an elbow diagram based on the number of clusters and the intra-class distance, and selecting the number of clusters corresponding to the inflection point of the elbow diagram as the number of clusters K of the K-means clustering algorithm.
[0011] Furthermore, the steps of iteratively clustering the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm include: step one, constructing the objective function of the K-means clustering algorithm; step two, randomly selecting K sample points from the preprocessed credit sample data as cluster center points based on the number of clusters K, to obtain K cluster center points; step three, for each sample point in the preprocessed credit sample data, calculating the distance value between the sample point and each cluster center point, and assigning the sample point to the cluster with the smallest distance value from the cluster center point, to obtain K clusters; step four, calculating the objective function value based on the objective function and the K clusters; step five, repeating the above steps two to four, iteratively updating the cluster center points and clusters until the objective function value reaches the minimum value, stopping the iteration, and obtaining the first clustering result based on the cluster center points and clusters of the last iteration.
[0012] Furthermore, the step of calculating the neighborhood distance threshold and neighborhood point count threshold of each cluster based on the first clustering result includes: calculating the average distance of sample points in each cluster to obtain the neighborhood distance threshold of each cluster; calculating the average density of sample points in each cluster to obtain the neighborhood point count threshold of each cluster.
[0013] Furthermore, the step of selecting the core point of the density clustering algorithm based on the neighborhood distance threshold and the neighborhood point number threshold of each clustering cluster includes: selecting each sample point from the preprocessed credit sample data as the target sample point in turn, and calculating the distance value between the target sample point and other sample points, wherein the other sample points are sample points in the credit sample data other than the target sample point; determining the cluster where the target sample point is located according to the first clustering result, and obtaining the neighborhood distance threshold and the neighborhood point number threshold corresponding to the target sample point; comparing the distance value between the target sample point and the other sample points with the neighborhood distance threshold corresponding to the target sample point, selecting other sample points with a distance value less than the neighborhood distance threshold as the neighborhood samples of the target sample point, and adding the neighborhood samples to the neighborhood sample set of the target sample point; when the number of sample points in the neighborhood sample set of the target sample point is greater than or equal to the neighborhood point number threshold corresponding to the target sample point, taking the target sample point as the core point of the density clustering algorithm, and adding the core point to the core point set.
[0014] Furthermore, the step of clustering the sample points in the preprocessed credit sample data based on each core point includes: step one, initializing the set of unvisited sample points in the preprocessed credit sample data, selecting a target core point from the core point set, and constructing a cluster corresponding to the target core point based on the neighborhood samples of the target core point; step two, constructing a core point queue based on the neighborhood samples of the target core point, wherein the core point queue contains neighborhood core points, and the neighborhood core points represent sample points whose neighborhood samples of the target core point are core points; step three, traversing the neighborhood core points in the core point queue, adding the neighborhood samples of the neighborhood core points to the cluster corresponding to the target core point, and updating the core point queue based on the traversed neighborhood core points and the neighborhood samples of the neighborhood core points until the core point queue is an empty set, ending the traversal, obtaining the cluster after clustering, and updating the set of unvisited sample points; step four, repeating the above steps one to three until the set of unvisited sample points is an empty set, and completing the clustering of the sample points in the credit sample data.
[0015] According to another aspect of an embodiment of the present invention, a clustering device for credit users is also provided, including: an acquisition unit, used to acquire multidimensional credit data of each credit user, obtain credit sample data, and preprocess the credit sample data; a first clustering unit, used to determine the number of clusters of the K-means clustering algorithm based on the elbow rule, and iteratively cluster the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm; a calculation unit, used to calculate the neighborhood distance threshold and neighborhood point number threshold of each cluster cluster based on the first clustering result; a second clustering unit, used to select the core point of the density clustering algorithm based on the neighborhood distance threshold and neighborhood point number threshold of each cluster cluster, and cluster the sample points in the preprocessed credit sample data based on each core point to obtain the clustering result of the credit user.
[0016] Furthermore, the acquisition unit includes: a first calculation module, used to calculate the distance value between two sample points based on the credit sample data, and calculate the circle radius based on the distance values between all the sample points; a second calculation module, used to calculate the decision threshold value based on the total number of the sample points; a first comparison module, used to calculate the density value of each of the sample points according to the circle radius, and compare the density value of each of the sample points with the decision threshold value to obtain a comparison result; a first deletion module, used to determine the outlier sample point based on the comparison result, and delete the outlier sample point from the credit sample data.
[0017] Furthermore, the first clustering unit includes: a first determination module, used for step one, determining the value range of the number of clusters based on the total number of sample points in the credit sample data; a first clustering module, used for step two, randomly selecting the number of clusters within the value range, and clustering the preprocessed credit sample data based on the selected number of clusters using the K-means clustering algorithm; a third calculation module, used for step three, calculating the intra-class distance corresponding to the number of clusters; a first repetition module, used for step four, repeating the above steps two to three, selecting a different number of clusters each time, until all the number of clusters within the value range are selected, and the intra-class distance corresponding to each number of clusters is obtained; a first drawing module, used for step five, drawing an elbow diagram based on the number of clusters and the intra-class distance, and selecting the number of clusters corresponding to the inflection point of the elbow diagram as the number of clusters K of the K-means clustering algorithm.
[0018] Furthermore, the first clustering unit includes: a first construction module, used for step one, to construct the objective function of the K-means clustering algorithm; a first selection module, used for step two, to randomly select K sample points as cluster center points from the preprocessed credit sample data based on the number of clusters K, to obtain K cluster center points; a first allocation module, used for step three, for each sample point in the preprocessed credit sample data, to calculate the distance value between the sample point and each cluster center point, and to allocate the sample point to the cluster with the smallest distance value from the cluster center point, to obtain K clusters; a fourth calculation module, used for step four, to calculate the objective function value based on the objective function and the K clusters; a first iteration module, used for step five, to repeat the above steps two to four, to iteratively update the cluster center points and cluster clusters until the objective function value reaches the minimum value, to stop the iteration, and to obtain the first clustering result based on the cluster center points and cluster clusters of the last iteration.
[0019] Furthermore, the calculation unit includes: a fifth calculation module, used to calculate the average distance of sample points in each cluster to obtain the neighborhood distance threshold of each cluster; a sixth calculation module, used to calculate the average density of sample points in each cluster to obtain the neighborhood point count threshold of each cluster.
[0020] Furthermore, the second clustering unit includes: a seventh calculation module, which is used to select each sample point from the preprocessed credit sample data as a target sample point in turn, and calculate the distance value between the target sample point and other sample points, wherein the other sample points are sample points in the credit sample data other than the target sample point; a second determination module, which is used to determine the cluster cluster where the target sample point is located according to the first clustering result, and obtain the neighborhood distance threshold and the neighborhood point number threshold corresponding to the target sample point; a second comparison module, which is used to compare the distance value between the target sample point and the other sample points with the neighborhood distance threshold corresponding to the target sample point, select other sample points with a distance value less than the neighborhood distance threshold as the neighborhood samples of the target sample point, and add the neighborhood samples to the neighborhood sample set of the target sample point; a first adding module, which is used to use the target sample point as the core point of the density clustering algorithm when the number of sample points in the neighborhood sample set of the target sample point is greater than or equal to the neighborhood point number threshold corresponding to the target sample point, and add the core point to the core point set.
[0021] Furthermore, the second clustering unit includes: a first construction module for step 1, initializing the set of unvisited sample points in the pre-processed credit sample data, selecting a target core point from the core point set, and constructing a cluster corresponding to the target core point based on the neighborhood samples of the target core point; a second construction module for step 2, constructing a core point queue based on the neighborhood samples of the target core point, wherein the core point queue contains neighborhood core points, and the neighborhood core points represent sample points whose neighborhood samples of the target core point are core points; a first traversal module , used for step three, traverse the neighborhood core points in the core point queue, add the neighborhood samples of the neighborhood core points to the cluster cluster corresponding to the target core point, and update the core point queue based on the traversed neighborhood core points and the neighborhood samples of the neighborhood core points, until the core point queue is an empty set, end the traversal, obtain the cluster cluster after clustering, and update the unvisited sample point set; the second repetition module, used for step four, repeats the above steps one to three until the unvisited sample point set is an empty set, and completes the clustering of the sample points in the credit sample data.
[0022] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any of the above-mentioned credit user clustering methods.
[0023] According to another aspect of an embodiment of the present invention, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the above-mentioned credit user clustering methods.
[0024] In this application, the following steps are performed: multidimensional credit data of each credit user is obtained to obtain credit sample data, and the credit sample data is preprocessed, and the number of clusters of the K-means clustering algorithm is determined based on the elbow rule, and the sample points in the preprocessed credit sample data are iteratively clustered based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm, and then the neighborhood distance threshold and neighborhood point number threshold of each cluster are calculated based on the first clustering result, and finally the core point of the density clustering algorithm is selected based on the neighborhood distance threshold and neighborhood point number threshold of each cluster, and the sample points in the preprocessed credit sample data are clustered based on each core point to obtain the clustering result of the credit user.
[0025] In this application, the number of clusters of the K-means clustering algorithm is determined based on the elbow rule, which can more accurately reflect the true distribution of the data, thereby improving the accuracy of the clustering results. According to the number of clusters, the K-means clustering algorithm is used to perform the first clustering of the credit samples. Then, based on the results of the first clustering, the relevant parameters of the density clustering algorithm, namely the neighborhood distance threshold and the neighborhood distance threshold, are calculated, thereby improving the quality of parameter selection. Based on the first clustering results, the core point of the density clustering algorithm is identified, and clustering expansion is performed with the core point as the center to achieve the second clustering of the credit samples and obtain the clustering results. Combining the two clustering algorithms can fully utilize the advantages of the density clustering algorithm in processing data of arbitrary shapes, make up for the limitations of the K-means clustering algorithm on data shape, and use high-quality parameters to perform refined second clustering, achieving the technical effect of improving the accuracy of clustering results. This solves the technical problem of low clustering result accuracy in the related technology of clustering credit users based on a single K-means clustering algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0027] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a credit user clustering method is shown;
[0028] Figure 2is a flow chart of an optional credit user clustering method according to an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of an optional credit user clustering process according to an embodiment of the present invention;
[0030] Figure 4 is a schematic diagram of an optional device for clustering credit users according to an embodiment of the present invention;
[0031] Figure 5 This is a hardware structure block diagram of an optional electronic device (or mobile device) for executing a credit user clustering method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] It should be noted that the credit user clustering method and device in this application can be used in the field of big data technology. When credit users are clustered based on an improved clustering algorithm, they can also be used in any field other than the field of big data technology. When credit users are clustered based on an improved clustering algorithm, the application field of the credit user clustering method and device in this application is not limited.
[0035] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation portals for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.
[0036] The following embodiments of the present invention can be applied to various credit user clustering systems, applications, and devices. This invention combines the K-means clustering algorithm and the density clustering algorithm. Using the results of the K-means clustering algorithm, the parameters of the density clustering algorithm are optimized, resulting in a more precise neighborhood distance threshold and neighborhood point count threshold. This allows the density clustering algorithm to identify its core points, which are then expanded into clusters centered around these core points to yield the final credit user clustering results. This fully leverages the density clustering algorithm's advantage in processing data of arbitrary shapes, overcoming the K-means clustering algorithm's limitations on data shape. This improves the accuracy of clustering for complex data shapes and enhances the accuracy of the clustering results.
[0037] The present invention adopts the elbow rule to determine the appropriate number of clusters for the K-means clustering algorithm to perform rough clustering, and then uses the density clustering algorithm to perform secondary processing on the data, so that the algorithm can perform clustering more accurately.
[0038] The present invention will be described in detail below with reference to various embodiments.
[0039] Example 1
[0040] According to an embodiment of the present invention, an embodiment of a method for clustering credit users is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0041] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a clustering method for credit users. Figure 1As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0042] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0043] Memory 104 can be used to store software programs and modules for application software, such as the program instructions / data storage device corresponding to the credit user clustering method described in the embodiments of the present application. Processor 102 executes the software programs and modules stored in memory 104 to perform various functional applications and data processing, thereby implementing the credit user clustering method described above. Memory 104 may include high-speed random access memory (RAM) and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, memory 104 may further include memory located remotely from processor 102, which can be connected to computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0044] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0045] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0046] Under the above operating environment, this application provides Figure 2 The credit user clustering method shown in FIG. 1 is implemented by a credit user clustering system.
[0047] Figure 2 is a flow chart of an optional credit user clustering method according to an embodiment of the present invention. Figure 2 As shown, the method includes the following steps:
[0048] Step S201: Acquire multi-dimensional credit data of each credit user, obtain credit sample data, and pre-process the credit sample data.
[0049] In the above step S201, the user group that needs to be evaluated is first determined, for example, it can be the credit customer group of a financial institution. The embodiment of the present invention can also be applied to various fields, such as clustering policyholders, medical patients, social users, etc.
[0050] In some embodiments, clustering credit user groups can help financial institutions identify user groups with similar credit characteristics, thereby implementing differentiated risk assessments for different groups. For example, by dividing users into high-risk, medium-risk, and low-risk groups, credit policies, loan interest rates, and approval processes can be adjusted more accurately, effectively controlling the non-performing loan rate and improving the quality of credit assets. Specifically, when clustering credit users, multi-dimensional data about credit users is first collected from various available resources, including but not limited to credit scores, income levels, educational backgrounds, loan histories, and repayment behaviors. This information constitutes a comprehensive portrait of the credit user and is the basis for subsequent analysis and clustering. The collected multi-dimensional credit data is integrated into a unified data set, with each credit user corresponding to a sample data, containing all the above-collected feature information.
[0051] When integrating data, you can clean the multi-dimensional data to remove missing values, outliers, and duplicate data to ensure data quality, and you can also standardize the multi-dimensional data.
[0052] For credit sample data, preprocessing can include removing outlier points, that is, removing points in the dataset that significantly deviate from other observations. Identifying and eliminating outlier points significantly improves the accuracy and robustness of subsequent cluster analysis.
[0053] Furthermore, the steps of preprocessing the multidimensional credit data include: calculating the distance value between two sample points based on the credit sample data, and calculating the circle radius based on the distance values between all sample points; calculating the decision threshold value based on the total number of sample points; calculating the density value of each sample point according to the circle radius, and comparing the density value of each sample point with the decision threshold value to obtain a comparison result; determining the outlier sample points based on the comparison result, and deleting the outlier sample points from the credit sample data.
[0054] Specifically, an appropriate metric (such as Euclidean distance, Manhattan distance, etc.) is used to calculate the distance between any two sample points in the credit sample dataset, and then the maximum distance value is selected to calculate the circle radius. The selection of the circle radius is a key parameter for subsequent outlier detection, which is used to define the neighborhood range of a sample point. The calculation formula for the circle radius is defined as: Among them, d(x i ,y i ) represents the distance value between the i-th sample and the j-th sample in the credit sample data.
[0055] Then, a decision threshold is calculated based on the total number of sample points, that is, the threshold for determining whether a sample point is an outlier. The calculation formula of the decision threshold is expressed as: τ = sqrt(n), where n is the total number of sample points.
[0056] Finally, using the circle radius as the neighborhood range, the number of points in each sample point's neighborhood is calculated. This value reflects the local density of the sample point. A higher density value means that the sample point is located in a dense area of the data set, while a lower density value may indicate that it is at the edge or sparse area of the data. The density value of each sample point is compared with the decision threshold value. If the density value of the sample point is higher than the decision threshold, the decision condition is met: d(x i )=c(x i ,dd)≥τ,d(x i ) represents the sample point x i If the density is , then the sample point will be marked as a non-outlier point, otherwise it will be regarded as an outlier. The cluster point will be deleted from the credit sample data to obtain the preprocessed credit sample data.
[0057] Step S202 : determining the number of clusters of the K-means clustering algorithm based on the elbow rule, and iteratively clustering the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm.
[0058] It should be noted that the K-means clustering algorithm requires the number of clusters K to be specified in advance. The selection of the K value often depends on experience or multiple experiments. If the K value is not selected properly, the clustering results may be unsatisfactory. In this embodiment of the present invention, the elbow rule is used to optimize the K-means clustering algorithm to determine the appropriate number of clusters K and improve the accuracy of the clustering results. At the same time, a weighted calculation method of sample points is used to iteratively cluster to determine the optimal clustering scheme and obtain the first clustering result.
[0059] Furthermore, the steps of determining the number of clusters of the K-means clustering algorithm based on the elbow rule include: step one, determining the value range of the number of clusters based on the total number of sample points in the credit sample data; step two, randomly selecting a number of clusters within the value range, and clustering the preprocessed credit sample data based on the selected number of clusters using the K-means clustering algorithm; step three, calculating the intra-class distance corresponding to the number of clusters; step four, repeating the above steps two to three, selecting a different number of clusters each time, until all the number of clusters within the value range are selected, and the intra-class distance corresponding to each number of clusters is obtained; step five, drawing an elbow diagram based on the number of clusters and the intra-class distance, and selecting the number of clusters corresponding to the inflection point of the elbow diagram as the number of clusters K of the K-means clustering algorithm.
[0060] Specifically, when determining the number of clusters K of the K-means clustering algorithm, first, based on the total number of sample points in the credit sample data, the value range of the number of clusters K is determined to be Determining the range of values helps avoid choosing too few or too many clusters. The former may lead to insufficient pattern recognition, while the latter may lead to overfitting, that is, the model is too complex and loses the ability to capture the general characteristics of the dataset.
[0061] Next, within the set range, a K value is randomly selected and the K-means clustering algorithm is used to cluster the preprocessed credit sample data. By randomly selecting a K value for preliminary clustering, we can explore the clustering structure under different cluster numbers, providing diversified clustering results for subsequent intra-cluster distance calculation and elbow rule application. For each clustering result under the K value, the intra-cluster distance is calculated: Among them, m i represents the mean of the sample data in the i-th cluster, p ij Represents the jth sample data in the i-th cluster, j = 1, 2, ..., n, C iRepresents the i-th cluster. The intra-cluster distance is an indicator for evaluating the closeness of clustering.
[0062] Repeat the above steps until all K values within the range have been selected, obtaining the intra-class distance corresponding to each K value. Finally, plot each K value and the corresponding intra-class distance on a graph to form an elbow plot. The elbow plot uses the K value as the horizontal axis and the corresponding intra-class distance as the vertical axis. A trend can be observed in which the intra-class distance changes with the K value, and there is an inflection point in this trend. The K value corresponding to this inflection point is used as the optimal number of clusters. This K value balances clustering effectiveness and model complexity, avoiding the problem of over- or under-clustering, and ensuring that the model remains simple while effectively reflecting the inherent structure of the credit user population.
[0063] In the above steps, by systematically determining and optimizing the number of clusters in the K-means clustering algorithm, the reliability and effectiveness of the clustering results were significantly improved. This process not only avoided the poor clustering performance that could result from blindly selecting a K value, but also, through continuous iteration and evaluation of intra-cluster distances, ensured that the final number of clusters selected would reveal the inherent clustering structure of the dataset while avoiding over-complication of the model. Using the elbow method to determine the optimal K value enabled K-means clustering to more accurately capture the segmented characteristics of users in the credit sample data, providing a solid foundation for subsequent, deeper data analysis and risk management strategies.
[0064] Furthermore, the steps of iteratively clustering the sample points in the preprocessed credit sample data based on the number of clusters to obtain the first clustering result based on the K-means clustering algorithm include: step one, constructing the objective function of the K-means clustering algorithm; step two, randomly selecting K sample points from the preprocessed credit sample data as cluster center points based on the number of clusters K, to obtain K cluster center points; step three, for each sample point in the preprocessed credit sample data, calculating the distance value between the sample point and the center point of each cluster, and assigning the sample point to the cluster with the smallest distance value from the cluster center point, to obtain K clusters; step four, calculating the objective function value based on the objective function and the K clusters; step five, repeating the above steps two to four, iteratively updating the cluster center points and clusters until the objective function value reaches the minimum value, stopping the iteration, and obtaining the first clustering result based on the cluster center points and clusters of the last iteration.
[0065] Specifically, after determining the number of clusters K, the credit sample data is clustered using the determined K value. First, the objective function of the K-means clustering algorithm is constructed. According to the importance of each data point in the clustering process, the x i Assign a corresponding weight ω i , define the objective function as: Among them, S is the set of all K clusters, S i is the set of data points assigned to the i-th cluster, K is the number of clusters, ω i ||x-μ i || 2 is the center point μ of the sample point x and its cluster i The weighted value of the square of the Euclidean distance between them.
[0066] In the process of iterative clustering, K data points are randomly selected as the initial cluster centers. For each data point in the data set, the distance between it and all cluster centers is calculated, and then it is assigned to the cluster represented by the nearest cluster center. Then, the objective function is calculated based on the current cluster until the objective function value reaches the minimum and no longer decreases after multiple rounds of iterations. Then, the iteration is stopped, and the sample points of each cluster at this time are recorded to obtain the first clustering result.
[0067] By iteratively assigning each data point to the nearest weighted cluster center, the clustering results can be gradually optimized over multiple iterations, making the data points within clusters closer together and the differences between clusters greater. The introduction of weights allows the algorithm to treat data points of varying importance in the dataset more fairly and accurately, avoiding oversensitivity to certain low-quality or insignificant data points, thereby improving the robustness and effectiveness of clustering.
[0068] Step S203: Calculate the neighborhood distance threshold and neighborhood point count threshold of each cluster according to the first clustering result.
[0069] In some embodiments, the credit sample data is preliminarily clustered using the K-means clustering algorithm to divide the sample points into K clusters. Then, the sample points of the K clusters are used to calculate the relevant parameters of the density clustering algorithm, including the neighborhood distance threshold and the neighborhood point count threshold. The density clustering algorithm relies on the selection of the neighborhood distance threshold and the neighborhood point count threshold. Inappropriate neighborhood distance threshold and neighborhood point count threshold will lead to poor clustering quality. Therefore, the calculation of the neighborhood distance threshold and the neighborhood point count threshold can be optimized through the results of the first clustering to improve the final clustering quality.
[0070] Furthermore, the step of calculating the neighborhood distance threshold and neighborhood point count threshold of each cluster according to the first clustering result includes: calculating the average distance of sample points in each cluster to obtain the neighborhood distance threshold of each cluster; calculating the average density of sample points in each cluster to obtain the neighborhood point count threshold of each cluster.
[0071] Specifically, after the first clustering result is determined, for each cluster, the distances between all sample points in the cluster are calculated, and then the average of these distances is calculated. This average distance will be used as the neighborhood distance threshold Eps of the cluster to define the neighborhood range of the sample point. The density of the sample points in each cluster is calculated. The density is usually defined as the number of sample points in the neighborhood of each sample point (defined by Eps). Then, the average value of the density of all sample points in the cluster is calculated, and the average value will be used as the neighborhood point count threshold MinPts of the cluster. The neighborhood point count threshold MinPts is set to distinguish between core points and boundary points. In the density clustering algorithm, if a point has at least MinPts neighboring sample points in the Eps neighborhood, the point is considered to be a core point. By calculating the average density of sample points in the cluster as MinPts, it can be ensured that the algorithm takes into account the density distribution of the cluster when identifying core points, and avoids the excessive or insufficient influence of the artificially set MinPts value on the cluster structure.
[0072] Step S204: Select core points of the density clustering algorithm based on the neighborhood distance threshold and neighborhood point number threshold of each cluster, and cluster the sample points in the preprocessed credit sample data based on each core point to obtain the clustering results of the credit users.
[0073] In the above step S204, after obtaining the first clustering result, the calculated neighborhood distance threshold (Eps) and neighborhood point number threshold (MinPts) of each cluster are used to further analyze the sample points in each cluster. Specifically, for any sample point in each cluster, check whether there are at least MinPts other sample points in its Eps neighborhood. If this condition is met, the point is regarded as a core point of the density clustering algorithm. The selection of core points is the key to the density clustering algorithm. By combining the two parameters Eps and MinPts with the first clustering result, it is possible to more accurately identify which points are the areas with higher density in the cluster. These core points are crucial for the formation of subsequent clusters. They define the range and shape of the clusters.
[0074] After identifying the core points, we begin clustering them using a density-based clustering algorithm. For each sample point in the preprocessed credit sample data, we check whether it lies within the Eps neighborhood of any core point. If so, the point is assigned to the cluster represented by that core point, and we continue to expand this cluster until no additional sample points can be added to any cluster.
[0075] In this embodiment of the present invention, starting from the core points, the density clustering algorithm can effectively cluster points in high-density areas together, forming tighter clusters with more similar characteristics. This helps to more accurately understand the risk characteristics of different credit user groups and provides a basis for formulating personalized credit strategies.
[0076] Furthermore, the step of selecting the core point of the density clustering algorithm based on the neighborhood distance threshold and neighborhood point number threshold of each clustering cluster includes: selecting each sample point as the target sample point from the preprocessed credit sample data in turn, and calculating the distance value between the target sample point and other sample points, wherein the other sample points are sample points other than the target sample point in the credit sample data; determining the cluster where the target sample point is located according to the first clustering result, and obtaining the neighborhood distance threshold and neighborhood point number threshold corresponding to the target sample point; comparing the distance value of the target sample point and the other sample points with the neighborhood distance threshold corresponding to the target sample point, selecting other sample points with distance values less than the neighborhood distance threshold as neighborhood samples of the target sample point, and adding the neighborhood samples to the neighborhood sample set of the target sample point; when the number of sample points in the neighborhood sample set of the target sample point is greater than or equal to the neighborhood point number threshold corresponding to the target sample point, taking the target sample point as the core point of the density clustering algorithm, and adding the core point to the core point set.
[0077] Specifically, for each sample point in the credit sample data, the cluster to which the sample point belongs is determined by the first clustering result, and the target sample point x is found by using the distance metric. i Eps-neighborhood N Eps (x i ):
[0078]
[0079] Among them, N Eps (x i ) is the credit sample data X′ and the target sample point x i The distance is less than the sample point x i The set of all sample points corresponding to Eps.
[0080] If the target sample point x i The number of neighborhood sample points in the Eps-neighborhood|N Eps (p)|≥MinPts, then the sample point x i Add to the core point set Ω: Ω=Ω∪{x i}.
[0081] In this way, all core points of the density clustering algorithm are screened out to obtain the core point set.
[0082] The embodiment of the present invention significantly improves the efficiency and quality of core point identification through parameter adjustment based on the first clustering result and neighborhood analysis of sample points, providing a solid foundation for subsequent core point-based clustering expansion.
[0083] Furthermore, the step of clustering the sample points in the preprocessed credit sample data based on each core point includes: step one, initializing the set of unvisited sample points in the preprocessed credit sample data, selecting a target core point from the core point set, and constructing a cluster corresponding to the target core point based on the neighborhood samples of the target core point; step two, constructing a core point queue based on the neighborhood samples of the target core point, wherein the core point queue contains neighborhood core points, and the neighborhood core point indicates that the neighborhood samples of the target core point are sample points of the core point; step three, traversing the neighborhood core points in the core point queue, adding the neighborhood samples of the neighborhood core points to the cluster corresponding to the target core point, and updating the core point queue based on the traversed neighborhood core points and the neighborhood samples of the neighborhood core points until the core point queue is an empty set, ending the traversal, obtaining the cluster after clustering, and updating the set of unvisited sample points; step four, repeating the above steps one to three until the set of unvisited sample points is an empty set, and completing the clustering of the sample points in the credit sample data.
[0084] Specifically, when performing density clustering on user credit data using a density clustering algorithm, the set of unvisited sample points is first initialized to ensure that all sample points are initially unassigned or unvisited. Next, a point is randomly selected from the identified core point set as the target core point. Based on the neighborhood sample set of this point (i.e., all sample points within the Eps distance), initial clusters are constructed and marked as visited. By selecting a core point and defining its surrounding neighborhood samples as the basis for clustering, these steps lay the foundation for subsequent cluster expansion.
[0085] Next, we continue to use the target core point's neighborhood samples to filter out points that are both core points and neighborhood samples of the target core point. These points are added to a new core point queue. The core point queue will contain all core sample points that can further expand the cluster. The core point queue is constructed to identify which neighborhood samples are also core points. Their presence means that the cluster can continue to expand outward until the entire high-density area is covered.
[0086] A neighboring core point is taken from the core point queue, and the cluster expansion process is repeated, incorporating its neighboring samples into the cluster corresponding to the target core point. At the same time, these newly added sample points are marked as visited to prevent duplicate counting. If these neighboring samples contain other core points that have not been traversed, they are also added to the core point queue. This process continues until the core point queue is empty, that is, until all core points in the queue have been visited and their neighboring samples have been added to the cluster. Through this traversal and update process, the cluster is expanded from the initial single core point to include all directly or indirectly connected core points and their neighboring samples, forming a complete high-density cluster.
[0087] Once the target core point cluster is complete and all associated neighborhood samples have been visited, the sample points that have been added to the cluster are removed from the set of unvisited sample points. Next, the next target core point is selected from the remaining set of core points, and the cluster construction process is repeated until all sample points in the preprocessed credit sample data have been classified into a cluster.
[0088] Through the above steps, multidimensional credit data of each credit user is obtained to obtain credit sample data, and the credit sample data is preprocessed. The number of clusters of the K-means clustering algorithm is determined based on the elbow rule, and the sample points in the preprocessed credit sample data are iteratively clustered based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm. Then, the neighborhood distance threshold and the neighborhood point number threshold of each cluster are calculated based on the first clustering result. Finally, the core point of the density clustering algorithm is selected based on the neighborhood distance threshold and the neighborhood point number threshold of each cluster, and the sample points in the preprocessed credit sample data are clustered based on each core point to obtain the clustering result of the credit user.
[0089] In this embodiment, the number of clusters in the K-means clustering algorithm is determined based on the elbow rule, which can more accurately reflect the true distribution of the data, thereby improving the accuracy of the clustering results. Based on the number of clusters, the K-means clustering algorithm is used to perform a first clustering of the credit sample. Then, based on the results of the first clustering, the relevant parameters of the density clustering algorithm, namely, the neighborhood distance threshold and the neighborhood distance threshold, are calculated. This improves the quality of parameter selection. Based on the first clustering results, the core point of the density clustering algorithm is identified, and cluster expansion is performed around the core point to achieve a second clustering of the credit sample and obtain a clustering result. Combining the two clustering algorithms can fully utilize the advantages of the density clustering algorithm in processing data of arbitrary shapes, compensate for the limitations of the K-means clustering algorithm on data shape, and use high-quality parameters to perform a refined second clustering, achieving the technical effect of improving the accuracy of the clustering results. This solves the technical problem of low clustering accuracy in the related art method of clustering credit users based on a single K-means clustering algorithm.
[0090] The following describes in detail another optional specific implementation.
[0091] Figure 3 is a schematic diagram of an optional credit user clustering process according to an embodiment of the present invention, such as Figure 3 As shown in Figure 2, the clustering process of credit users specifically includes:
[0092] Step 1: Obtain multi-dimensional data of credit user groups;
[0093] Assume that there are n samples in the data set X of the credit user group to be divided, that is, X={x1,x2,...,x n}, where each x i Contains data of various dimensions corresponding to the credit users to be classified.
[0094] Step 2: Eliminate outliers and construct credit sample data;
[0095] Process outliers and mark and remove interfering outliers such as noise data:
[0096] Define the circle radius as:
[0097] The decision threshold is: τ = sqrt(n),
[0098] The judgment condition is: d(x i )=c(x i ,dd)≥τ, where d(x i ) represents the object x i density.
[0099] The non-outlier points that meet the above equation are screened out to form a new data set X′, and the preprocessed credit sample data is obtained.
[0100] Step 3: Use the elbow method to determine the number of clusters K;
[0101] Determine the value range of the number of clusters K Cluster all samples in X' into K classes. The intra-class distance D in the dataset X' is: Among them, m i Represents the mean of the sample data in the i-th class, p ij Represents the sample data in the i-th class, j = 1, 2, ..., n, C i Denotes the i-th cluster. Calculate the intra-cluster distance for different values of K. Draw the elbow plot and determine the optimal number of clusters K corresponding to the inflection point.
[0102] Step 4: Assign weights to each sample point and construct the objective function;
[0103] According to the importance of each data point in the clustering process, i Assign a corresponding weight ω i , define the objective function as:
[0104]
[0105] Among them, S is the set of all K clusters, S i is the set of sample points assigned to the i-th cluster, K is the number of clusters, ω i ||x-μ i || 2 is the center point μ of the data point x and its cluster i The weighted value of the square of the Euclidean distance between them.
[0106] Step 5: Perform iterative clustering based on the objective function to obtain the first clustering result;
[0107] After multiple iterations, the That is, the weighted cluster square sum, each sample point is assigned to the nearest weighted cluster center by iteration, and the weighted cluster center is updated until convergence. The cluster division S={S1,S2,...,S k}.
[0108] Step 6: Calculate the Eps (corresponding to the above neighborhood distance threshold) and Minpts (corresponding to the above neighborhood point number threshold) of each cluster based on the first clustering result;
[0109] Based on the improved first clustering results above, calculate the Eps and MinPts of different clusters. Eps is the neighborhood radius of a specified sample, indicating that if the distance between two samples is less than or equal to this value, they are connected to each other; MinPts is a threshold parameter, indicating the minimum number of samples that must be included in the neighborhood of Eps of a sample.
[0110] Step 7, initialization parameters;
[0111] Initialize the core point set Ω = φ, initialize the cluster number k = 0, initialize the unvisited sample point set Γ = D, and divide the new cluster S′ = φ.
[0112] Step 8: determine the core points and obtain the core point set;
[0113] Find the sample point x by distance metric i Eps-neighborhood N Eps (x i ):
[0114]
[0115] Among them, N Eps (x i ) is the sample point x in the data set X′ i The distance between Eps and the number of points in the neighborhood is less than the set of all sample points. Eps (p)|≥MinPts, then the sample point x i Add to the core point set Ω: Ω=Ω∪{x i}, thus obtaining the core point set of the density clustering algorithm.
[0116] Step 9: Determine whether the core point set is an empty set. If so, the algorithm ends. If not, execute step 10.
[0117] Step 10: Initialize the core point set, update the category number, initialize the current cluster sample set, and update the unvisited sample set;
[0118] Step 11: Select a core point as the center point of the current cluster and perform extended clustering based on the core point;
[0119] In the core point set Ω, randomly select a core point o and initialize the core point queue Ω of the current cluster cur =o, update the cluster number k, initialize the current cluster sample set S' k ={o}, update the unvisited sample set Γ=Γ-{o}.
[0120] With the core point as the center, the neighborhood samples of the core point are added to the current cluster sample set. If there are other core points in the neighborhood samples, these core points are updated to the core point queue, and the other core points in the core point queue are traversed, and the neighborhood samples of other core points are also extended to the current cluster. Repeat this step until the core point queue is an empty set.
[0121] When traversing the core point queue to perform extended clustering of the current cluster, for the core point o′ that has been traversed, all Eps-neighborhoods N(o′) of the core point o′ are delineated by the neighborhood distance threshold Eps, and Δ=N(o′)∩Γ is set to update the sample set S′ of the current cluster. k =S′ k ∪Δ, and update the unvisited sample set Γ=Γ-Δ, update the core point queue Ω cur =Ω cur ∪(Δ∩Ω)-o′, use this sample point for verification to avoid missing sample points when expanding the cluster.
[0122] Step 12: determine whether the core point queue of the current cluster is an empty set. If not, return to step 10. If so, execute step 13.
[0123] Step 13, obtain the clustering result of the current cluster;
[0124] Step 14: determine whether the unvisited sample set is an empty set. If not, repeat steps 10 to 14. If so, execute step 15.
[0125] Step 15: Get the clustering results of credit users.
[0126] Through the above steps, iterative clustering is performed to obtain the final clustering result S′={S1′,S2′,...,S′ k}.
[0127] In the above example, the K-means clustering algorithm divides data based on centroids, tends to find spherical clusters, and is ineffective for clustering non-convex data. In contrast, the density clustering algorithm, based on the density of data points, can identify clusters of arbitrary shapes and effectively identify noise points in the dataset. Combining the two algorithms can fully leverage the advantages of the density clustering algorithm in processing arbitrary-shaped data, offset the limitations of the K-means clustering algorithm on data shapes, and thus improve the clustering accuracy of complex-shaped data.
[0128] At the same time, the K-means clustering algorithm requires the number of clusters K to be specified in advance. The selection of the K value often relies on experience or multiple experiments. If the K value is not chosen properly, the clustering results may be unsatisfactory. The density clustering algorithm does not require the number of clusters to be formed in advance; it automatically identifies clusters based on the density of data points. By combining these two algorithms, the present embodiment first uses the elbow rule to determine the appropriate number of clusters K for the K-means clustering algorithm to perform rough clustering. Then, the density clustering algorithm is used to perform secondary processing on the credit sample data, enabling more accurate clustering and improving the accuracy of the clustering results.
[0129] Furthermore, the clustering results of the K-means clustering algorithm are significantly affected by the choice of initial cluster centers; different initial values can lead to different clustering results. However, the results of the density clustering algorithm do not depend on the choice of initial values. Combining the two can, to a certain extent, reduce the sensitivity of the K-means clustering algorithm to initial values, making the final clustering results more stable and reliable. Furthermore, when spatial clustering density is uneven and the distances between clusters vary greatly, the difficulty in selecting the Eps and MinPts parameters in the density clustering algorithm can lead to poor clustering quality. By using the initial clustering results of the K-means clustering algorithm, we can adopt a parameter adaptation approach to improve the quality of the Eps and MinPts parameter selection, thereby improving the final clustering quality.
[0130] The embodiment of the present invention also marks the noise and outliers before performing K-means clustering, and performs K-means clustering on new data that are not outliers, so that it can focus more on clustering valid data, improve the quality of clustering results, and enhance the robustness of the algorithm to noise and outliers.
[0131] The following describes it in detail with reference to another embodiment.
[0132] Example 2
[0133] A credit user clustering device provided in this embodiment includes multiple implementation units, each implementation unit corresponds to each implementation step in the above-mentioned embodiment 1. Its specific implementation method and beneficial effects can be referred to the above-mentioned method embodiment and will not be repeated here.
[0134] Figure 4 is a schematic diagram of an optional device for clustering credit users according to an embodiment of the present invention, such as Figure 4 As shown, the credit user clustering device may include: an acquisition unit 41, a first clustering unit 42, a calculation unit 43, and a second clustering unit 44, wherein:
[0135] An acquisition unit 41 is used to acquire multi-dimensional credit data of each credit user, obtain credit sample data, and pre-process the credit sample data;
[0136] A first clustering unit 42 is configured to determine the number of clusters of the K-means clustering algorithm based on the elbow rule, and iteratively cluster the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm;
[0137] A calculation unit 43 is used to calculate a neighborhood distance threshold and a neighborhood point count threshold of each cluster according to the first clustering result;
[0138] The second clustering unit 44 is used to select the core point of the density clustering algorithm based on the neighborhood distance threshold and the neighborhood point number threshold of each cluster cluster, and cluster the sample points in the preprocessed credit sample data based on each core point to obtain the clustering result of the credit users.
[0139] The above-mentioned credit user clustering device obtains multidimensional credit data of each credit user through the acquisition unit 41, obtains credit sample data, and preprocesses the credit sample data; determines the number of clusters of the K-means clustering algorithm based on the elbow rule through the first clustering unit 42, and iteratively clusters the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm; calculates the neighborhood distance threshold and neighborhood point number threshold of each cluster according to the first clustering result through the calculation unit 43; selects the core point of the density clustering algorithm based on the neighborhood distance threshold and neighborhood point number threshold of each cluster, and clusters the sample points in the preprocessed credit sample data based on each core point to obtain the clustering result of the credit user through the second clustering unit 44.
[0140] In this embodiment, the number of clusters in the K-means clustering algorithm is determined based on the elbow rule, which can more accurately reflect the true distribution of the data, thereby improving the accuracy of the clustering results. Based on the number of clusters, the K-means clustering algorithm is used to perform a first clustering of the credit sample. Then, based on the results of the first clustering, the relevant parameters of the density clustering algorithm, namely, the neighborhood distance threshold and the neighborhood distance threshold, are calculated. This improves the quality of parameter selection. Based on the first clustering results, the core point of the density clustering algorithm is identified, and cluster expansion is performed around the core point to achieve a second clustering of the credit sample and obtain a clustering result. Combining the two clustering algorithms can fully utilize the advantages of the density clustering algorithm in processing data of arbitrary shapes, compensate for the limitations of the K-means clustering algorithm on data shape, and use high-quality parameters to perform a refined second clustering, achieving the technical effect of improving the accuracy of the clustering results. This solves the technical problem of low clustering accuracy in the related art method of clustering credit users based on a single K-means clustering algorithm.
[0141] Furthermore, the acquisition unit includes: a first calculation module, used to calculate the distance value between two sample points based on the credit sample data, and calculate the circle radius based on the distance value between all sample points; a second calculation module, used to calculate the decision threshold value based on the total number of sample points; a first comparison module, used to calculate the density value of each sample point according to the circle radius, and compare the density value of each sample point with the decision threshold value to obtain a comparison result; a first deletion module, used to determine the outlier sample point based on the comparison result, and delete the outlier sample point from the credit sample data.
[0142] Furthermore, the first clustering unit includes: a first determination module, used for step one, determining the value range of the number of clusters based on the total number of sample points in the credit sample data; a first clustering module, used for step two, randomly selecting the number of clusters within the value range, and clustering the preprocessed credit sample data based on the selected number of clusters using the K-means clustering algorithm; a third calculation module, used for step three, calculating the intra-class distance corresponding to the number of clusters; a first repetition module, used for step four, repeating the above steps two to three, selecting a different number of clusters each time, until all the number of clusters within the value range are selected, and the intra-class distance corresponding to each number of clusters is obtained; a first drawing module, used for step five, drawing an elbow diagram based on the number of clusters and the intra-class distance, and selecting the number of clusters corresponding to the inflection point of the elbow diagram as the number of clusters K of the K-means clustering algorithm.
[0143] Furthermore, the first clustering unit includes: a first construction module, used for step one, to construct the objective function of the K-means clustering algorithm; a first selection module, used for step two, to randomly select K sample points as cluster center points from the preprocessed credit sample data based on the number of clusters K, to obtain K cluster center points; a first allocation module, used for step three, for each sample point in the preprocessed credit sample data, to calculate the distance value between the sample point and the center point of each cluster, and to allocate the sample point to the cluster with the smallest distance value from the cluster center point, to obtain K clusters; a fourth calculation module, used for step four, to calculate the objective function value based on the objective function and the K clusters; a first iteration module, used for step five, to repeat the above steps two to four, to iteratively update the cluster center points and cluster clusters until the objective function value reaches the minimum value, to stop the iteration, and to obtain the first clustering result based on the cluster center points and cluster clusters of the last iteration.
[0144] Furthermore, the calculation unit includes: a fifth calculation module, which is used to calculate the average distance of sample points in each cluster to obtain the neighborhood distance threshold of each cluster; and a sixth calculation module, which is used to calculate the average density of sample points in each cluster to obtain the neighborhood point count threshold of each cluster.
[0145] Furthermore, the second clustering unit includes: a seventh calculation module, which is used to select each sample point as a target sample point from the preprocessed credit sample data in turn, and calculate the distance value between the target sample point and other sample points, wherein the other sample points are sample points in the credit sample data other than the target sample point; a second determination module, which is used to determine the cluster cluster where the target sample point is located according to the first clustering result, and obtain the neighborhood distance threshold and neighborhood point number threshold corresponding to the target sample point; a second comparison module, which is used to compare the distance value of the target sample point and other sample points with the neighborhood distance threshold corresponding to the target sample point, select other sample points with a distance value less than the neighborhood distance threshold as neighborhood samples of the target sample point, and add the neighborhood samples to the neighborhood sample set of the target sample point; a first adding module, which is used to use the target sample point as the core point of the density clustering algorithm when the number of sample points in the neighborhood sample set of the target sample point is greater than or equal to the neighborhood point number threshold corresponding to the target sample point, and add the core point to the core point set.
[0146] Furthermore, the second clustering unit includes: a first construction module, used for step one, initializing the set of unvisited sample points in the preprocessed credit sample data, selecting a target core point from the core point set, and constructing a cluster corresponding to the target core point based on the neighborhood samples of the target core point; a second construction module, used for step two, constructing a core point queue based on the neighborhood samples of the target core point, wherein the core point queue contains neighborhood core points, and the neighborhood core points represent sample points whose neighborhood samples of the target core point are the core points; a first traversal module, used for step three, traversing the neighborhood core points in the core point queue, adding the neighborhood samples of the neighborhood core points to the cluster corresponding to the target core point, and updating the core point queue based on the traversed neighborhood core points and the neighborhood samples of the neighborhood core points until the core point queue is an empty set, ending the traversal, obtaining the cluster after clustering, and updating the set of unvisited sample points; a second repetition module, used for step four, repeating the above steps one to three until the set of unvisited sample points is an empty set, and completing the clustering of the sample points in the credit sample data.
[0147] It should be noted that the acquisition unit 41, the first clustering unit 42, the calculation unit 43, and the second clustering unit 44 correspond to steps S201 to S204 in the first embodiment. The examples and application scenarios implemented by the above units and the corresponding steps are the same, but are not limited to the contents disclosed in the first embodiment. It should be noted that the above modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above modules or units can also be part of a device and can be run in the computer terminal 10 provided in the first embodiment.
[0148] The present invention is described below in conjunction with another optional embodiment.
[0149] Example 3
[0150] An embodiment of the present invention may further provide an electronic device, Figure 5 FIG. 1 is a hardware structure block diagram of an electronic device (or mobile device) for performing an optional method for clustering credit users according to an embodiment of the present invention, such as Figure 5 As shown, the electronic device may include: one or more ( Figure 5 Only one is shown) processor 502, memory 504, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0151] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0152] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain multidimensional credit data of each credit user, obtain credit sample data, and preprocess the credit sample data; determine the number of clusters of the K-means clustering algorithm based on the elbow rule, and iteratively cluster the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm; calculate the neighborhood distance threshold and neighborhood point number threshold of each cluster according to the first clustering result; select the core point of the density clustering algorithm based on the neighborhood distance threshold and neighborhood point number threshold of each cluster, and cluster the sample points in the preprocessed credit sample data based on each core point to obtain the clustering result of the credit user.
[0153] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: the steps of preprocessing the multidimensional credit data include: calculating the distance value between two sample points based on the credit sample data, and calculating the circle radius based on the distance value between all sample points; calculating the decision threshold value based on the total number of sample points; calculating the density value of each sample point according to the circle radius, and comparing the density value of each sample point with the decision threshold value to obtain a comparison result; determining the outlier sample point based on the comparison result, and deleting the outlier sample point from the credit sample data.
[0154] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: the step of determining the number of clusters of the K-means clustering algorithm based on the elbow rule includes: step one, determining the value range of the number of clusters based on the total number of sample points in the credit sample data; step two, randomly selecting a number of clusters within the value range, and clustering the preprocessed credit sample data based on the selected number of clusters using the K-means clustering algorithm; step three, calculating the intra-class distance corresponding to the number of clusters; step four, repeating the above steps two to three, selecting a different number of clusters each time, until all the number of clusters within the value range are selected, and the intra-class distance corresponding to each number of clusters is obtained; step five, drawing an elbow diagram based on the number of clusters and the intra-class distance, and selecting the number of clusters corresponding to the inflection point of the elbow diagram as the number of clusters K of the K-means clustering algorithm.
[0155] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: iteratively clustering the sample points in the preprocessed credit sample data based on the number of clusters, and obtaining the first clustering result based on the K-means clustering algorithm includes: step one, constructing the objective function of the K-means clustering algorithm; step two, randomly selecting K sample points from the preprocessed credit sample data as cluster center points based on the number of clusters K, to obtain K cluster center points; step three, for each sample point in the preprocessed credit sample data, calculating the distance value between the sample point and the center point of each cluster, and assigning the sample point to the cluster with the smallest distance value from the cluster center point, to obtain K clusters; step four, calculating the objective function value based on the objective function and the K clusters; step five, repeating the above steps two to four, iteratively updating the cluster center points and clusters until the objective function value reaches the minimum value, stopping the iteration, and obtaining the first clustering result based on the cluster center points and clusters of the last iteration.
[0156] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: the steps of calculating the neighborhood distance threshold and neighborhood point count threshold of each cluster based on the first clustering result include: calculating the average distance of the sample points in each cluster to obtain the neighborhood distance threshold of each cluster; calculating the average density of the sample points in each cluster to obtain the neighborhood point count threshold of each cluster.
[0157] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: the step of selecting the core point of the density clustering algorithm based on the neighborhood distance threshold and neighborhood point number threshold of each cluster cluster includes: selecting each sample point as the target sample point from the preprocessed credit sample data in turn, and calculating the distance value between the target sample point and other sample points, wherein the other sample points are sample points in the credit sample data other than the target sample point; determining the cluster where the target sample point is located according to the first clustering result, and obtaining the neighborhood distance threshold and neighborhood point number threshold corresponding to the target sample point; comparing the distance value of the target sample point and the other sample points with the neighborhood distance threshold corresponding to the target sample point, selecting other sample points with distance values less than the neighborhood distance threshold as neighborhood samples of the target sample point, and adding the neighborhood samples to the neighborhood sample set of the target sample point; when the number of sample points in the neighborhood sample set of the target sample point is greater than or equal to the neighborhood point number threshold corresponding to the target sample point, taking the target sample point as the core point of the density clustering algorithm, and adding the core point to the core point set.
[0158] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: the step of clustering the sample points in the preprocessed credit sample data based on each core point includes: step one, initializing the set of unvisited sample points in the preprocessed credit sample data, selecting a target core point from the core point set, and constructing a cluster cluster corresponding to the target core point based on the neighborhood samples of the target core point; step two, constructing a core point queue based on the neighborhood samples of the target core point, wherein the core point queue contains neighborhood core points, and the neighborhood core point indicates that the neighborhood samples of the target core point are sample points of the core point; step three, traversing the neighborhood core points in the core point queue, adding the neighborhood samples of the neighborhood core points to the cluster cluster corresponding to the target core point, and updating the core point queue based on the traversed neighborhood core points and the neighborhood samples of the neighborhood core points until the core point queue is an empty set, ending the traversal, obtaining the cluster cluster after clustering, and updating the set of unvisited sample points; step four, repeating the above steps one to three until the set of unvisited sample points is an empty set, and completing the clustering of the sample points in the credit sample data.
[0159] An embodiment of the present invention provides a method for clustering loan users. The number of clusters of the K-means clustering algorithm is determined based on the elbow rule, which can more accurately reflect the true distribution of the data, thereby improving the accuracy of the clustering results. Based on the number of clusters, the K-means clustering algorithm is used to perform a first clustering of the credit samples. Then, based on the results of the first clustering, the relevant parameters of the density clustering algorithm, namely, the neighborhood distance threshold and the neighborhood distance threshold, are calculated, thereby improving the quality of parameter selection. Based on the first clustering results, the core point of the density clustering algorithm is identified, and cluster expansion is performed with the core point as the center to achieve a second clustering of the credit samples and obtain a clustering result. Combining the two clustering algorithms can fully utilize the advantages of the density clustering algorithm in processing data of arbitrary shapes, compensate for the limitations of the K-means clustering algorithm on data shapes, and use high-quality parameters to perform a refined second clustering, achieving the technical effect of improving the accuracy of the clustering results. This solves the technical problem of low clustering result accuracy in the related art of clustering credit users based on a single K-means clustering algorithm.
[0160] It can be understood by those skilled in the art that Figure 5 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone, a tablet computer, a PDA, a mobile Internet device (MID), or a PAD. Figure 5 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 5 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 5 Different configurations shown.
[0161] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0162] The present invention is described below in conjunction with another optional embodiment.
[0163] Example 4
[0164] The embodiment of the present invention further provides a computer-readable storage medium. Optionally, in the embodiment of the present invention, the computer-readable storage medium can be used to store the program code executed by the credit user clustering method provided in the first embodiment.
[0165] Optionally, in an embodiment of the present invention, the above-mentioned storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0166] An embodiment of the present invention also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of a clustering method for credit users: obtaining multidimensional credit data of each credit user, obtaining credit sample data, and preprocessing the credit sample data; determining the number of clusters of the K-means clustering algorithm based on the elbow rule, and iteratively clustering the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm; calculating the neighborhood distance threshold and neighborhood point number threshold of each cluster according to the first clustering result; selecting the core point of the density clustering algorithm based on the neighborhood distance threshold and neighborhood point number threshold of each cluster, and clustering the sample points in the preprocessed credit sample data based on each core point to obtain a clustering result of the credit user.
[0167] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0168] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0169] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0170] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0171] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0172] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0173] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A credit user clustering method, characterized in that: include: Acquiring multidimensional credit data of each credit user to obtain credit sample data, and preprocessing the credit sample data; Determining the number of clusters of the K-means clustering algorithm based on the elbow rule, and iteratively clustering the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm; Calculating a neighborhood distance threshold and a neighborhood point count threshold of each cluster according to the first clustering result; The core points of the density clustering algorithm are selected based on the neighborhood distance threshold and the neighborhood point number threshold of each cluster cluster, and the sample points in the preprocessed credit sample data are clustered based on each core point to obtain the clustering results of the credit users.
2. The method according to claim 1, characterized in that The step of preprocessing the multidimensional credit data includes: Calculating a distance value between two sample points based on the credit sample data, and calculating a circle radius based on the distance values between all the sample points; Calculating a decision threshold based on the total number of sample points; Calculating a density value of each sample point according to the circle radius, and comparing the density value of each sample point with the decision threshold value to obtain a comparison result; Outlier sample points are determined based on the comparison results, and the outlier sample points are deleted from the credit sample data.
3. The method according to claim 1, characterized in that The steps to determine the number of clusters for the K-means clustering algorithm based on the elbow rule include: Step 1: determining a value range of the number of clusters based on the total number of sample points in the credit sample data; Step 2: randomly selecting a cluster number value within a value range, and clustering the pre-processed credit sample data based on the selected cluster number value using a K-means clustering algorithm; Step 3: Calculate the intra-class distance corresponding to the number of clusters; Step 4: Repeat steps 2 to 3 above, selecting a different number of clusters each time, until all the number of clusters within the range are selected, and obtain the intra-cluster distance corresponding to each number of clusters; Step 5: Draw an elbow diagram based on the number of clusters and the intra-cluster distance, and select the number of clusters corresponding to the inflection point of the elbow diagram as the number of clusters K of the K-means clustering algorithm.
4. The method according to claim 3, characterized in that The step of iteratively clustering the sample points in the pre-processed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm includes: Step 1: Construct the objective function of the K-means clustering algorithm; Step 2: randomly selecting K sample points from the pre-processed credit sample data as cluster centers based on the number of clusters K, to obtain K cluster centers; Step 3: for each sample point in the preprocessed credit sample data, calculate the distance between the sample point and the center point of each cluster, and assign the sample point to the cluster with the smallest distance from the center point of the cluster, to obtain K clusters; Step 4: Calculate the objective function value based on the objective function and the K clusters; Step 5: Repeat steps 2 to 4 above, iteratively update the cluster centers and clusters until the objective function value reaches a minimum, stop iteration, and obtain the first clustering result based on the cluster centers and clusters of the last iteration.
5. The method according to claim 1, wherein The step of calculating the neighborhood distance threshold and the neighborhood point number threshold of each cluster according to the first clustering result includes: Calculate the average distance of the sample points in each cluster to obtain the neighborhood distance threshold of each cluster; The average density of the sample points in each cluster is calculated to obtain the neighborhood point count threshold of each cluster.
6. The method according to claim 1, characterized in that The step of selecting the core point of the density clustering algorithm based on the neighborhood distance threshold and the neighborhood point number threshold of each cluster includes: Selecting each sample point from the pre-processed credit sample data as a target sample point in sequence, and calculating a distance value between the target sample point and other sample points, wherein the other sample points are sample points in the credit sample data other than the target sample point; Determine the cluster where the target sample point is located according to the first clustering result, and obtain the neighborhood distance threshold and the neighborhood point number threshold corresponding to the target sample point; Comparing the distance between the target sample point and the other sample points with the neighborhood distance threshold corresponding to the target sample point, selecting other sample points whose distance values are smaller than the neighborhood distance threshold as neighborhood samples of the target sample point, and adding the neighborhood samples to the neighborhood sample set of the target sample point; When the number of sample points in the neighborhood sample set of the target sample point is greater than or equal to the neighborhood point number threshold corresponding to the target sample point, the target sample point is used as the core point of the density clustering algorithm, and the core point is added to the core point set.
7. The method according to claim 6, characterized in that The step of clustering the sample points in the pre-processed credit sample data based on each core point includes: Step 1: Initialize the set of unvisited sample points in the preprocessed credit sample data, select a target core point from the core point set, and construct a cluster corresponding to the target core point based on the neighborhood samples of the target core point; Step 2: constructing a core point queue based on the neighborhood samples of the target core point, wherein the core point queue includes neighborhood core points, and the neighborhood core points represent sample points whose neighborhood samples of the target core point are core points; Step 3: traverse the neighborhood core points in the core point queue, add the neighborhood samples of the neighborhood core points to the cluster corresponding to the target core point, and update the core point queue based on the traversed neighborhood core points and the neighborhood samples of the neighborhood core points until the core point queue is empty. Then, the traversal ends, the cluster after clustering is obtained, and the set of unvisited sample points is updated. Step 4: Repeat steps 1 to 3 above until the set of unvisited sample points is an empty set, thus completing the clustering of the sample points in the credit sample data.
8. A device for clustering credit users, characterized in that: include: an acquisition unit, configured to acquire multi-dimensional credit data of each credit user, obtain credit sample data, and pre-process the credit sample data; a first clustering unit, configured to determine the number of clusters of a K-means clustering algorithm based on an elbow rule, and iteratively cluster the sample points in the preprocessed credit sample data based on the number of clusters to obtain a first clustering result based on the K-means clustering algorithm; a calculation unit, configured to calculate a neighborhood distance threshold and a neighborhood point count threshold of each cluster according to the first clustering result; The second clustering unit is used to select the core point of the density clustering algorithm based on the neighborhood distance threshold and the neighborhood point number threshold of each cluster cluster, and cluster the sample points in the preprocessed credit sample data based on each core point to obtain the clustering result of the credit user.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the credit user clustering method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The system comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the credit user clustering method described in any one of claims 1 to 7.