Group classification method and electronic device
By performing multi-factor clustering and risk value calculation on credentials in fintech and dynamically adjusting the cluster centers, the problem of inaccurate group classification in existing technologies is solved, and more accurate user group classification is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WEBANK (CHINA)
- Filing Date
- 2021-08-30
- Publication Date
- 2026-05-05
AI Technical Summary
In the fintech field, the inaccuracy of group classification results in existing technologies is mainly due to the fact that the threshold range corresponding to the risk category is set based on empirical values, which leads to a discrepancy between the actual situation and the classification results.
By clustering the first voucher based on multiple set factors, calculating the risk value of each voucher, and further clustering based on these risk values, the final group classification result is output. The clustering process is optimized by dynamically adjusting the random risk value and the cluster center.
It improves the accuracy of group classification results, making the results more reflective of the actual situation and enhancing the accuracy and rationality of user group classification in fintech.
Smart Images

Figure CN113743493B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more particularly to a group classification method and electronic device. Background Technology
[0002] With the development of computer technology, more and more technologies are being applied in the financial field, and the traditional financial industry is gradually transforming into fintech. However, due to the security and real-time requirements of the financial industry, fintech also places higher demands on technology. In the fintech field, in the scenario of user group classification, the risk category to which a voucher belongs is determined based on the correspondence between a set risk category and a set threshold range, as well as the set threshold range to which the key indicators of the user's voucher fall. Based on the set risk category to which the voucher belongs and the correspondence between the voucher and the user, the user group classification result is output. However, the set threshold range corresponding to the set risk category is set based on empirical values, which leads to the group classification result not matching the actual situation, resulting in inaccurate group classification results. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a group classification method and an electronic device to solve the technical problem of inaccurate group classification results in related technologies.
[0004] To achieve the above objectives, the technical solution of the present invention is implemented as follows:
[0005] This invention provides a group classification method, comprising:
[0006] Based on the value of each of the at least two set factors corresponding to the first voucher, multiple first vouchers are clustered to obtain the first clustering result corresponding to each set factor;
[0007] The weighted summation of the first risk values corresponding to each first voucher in each first clustering result yields the second risk value corresponding to each first voucher; the first risk value represents the random risk value corresponding to the cluster center of the cluster in the corresponding first clustering result of the first voucher;
[0008] Based on the second risk value corresponding to the first voucher, the plurality of first vouchers are clustered to obtain a second clustering result;
[0009] Based on the second clustering result and the user corresponding to the first credential, the group classification result is output.
[0010] In the above scheme, when clustering multiple first credentials, the method includes:
[0011] Based on the number of clusters and the total number of first vouchers, calculate the sorting number corresponding to each cluster center;
[0012] Based on the sorting number corresponding to each cluster core, all cluster cores are determined in the first voucher after sorting according to the first value; the first value includes the value of the set factor or the second risk value;
[0013] Calculate the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center;
[0014] Add the first voucher to the cluster containing the cluster center corresponding to the smallest first difference.
[0015] In the above scheme, when calculating the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center, the method includes:
[0016] The square of the difference between the square of the first value corresponding to the first voucher and the square of the first value corresponding to the cluster center is determined as the first difference value.
[0017] In the above scheme, after adding all the first credentials to the corresponding cluster, the method further includes:
[0018] Calculate the convergence threshold and the absolute difference for each cluster; where the absolute difference represents the absolute value of the difference between the first value corresponding to the cluster center and the corresponding first mean; the first mean represents the mean of the first values corresponding to all first vouchers in the corresponding cluster; the convergence threshold represents the quotient of the second difference and the total number of random risk values; the second difference represents the difference between the largest and smallest first values;
[0019] If the calculated absolute difference is greater than the convergence threshold, a new cluster center for each cluster is determined based on the first mean corresponding to each cluster, and the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center, as well as subsequent steps, are executed; or
[0020] If all calculated absolute differences are less than or equal to the convergence threshold, the plurality of first credentials are not re-clustered.
[0021] The method in the above scheme further includes:
[0022] If the number of clusters is less than the total number of random risk values corresponding to all first vouchers, and the second mean is greater than or equal to the convergence threshold, the minimum value between the second value and the total number of random risk values is determined as the new number of clusters. Based on the new number of clusters and the first value corresponding to each first voucher, the plurality of first vouchers are re-clustered; wherein...
[0023] The second mean represents the absolute value of the mean of the differences between any two adjacent first means in the first mean array; the second value is determined by the number of clusters and the difference between the second mean and the convergence threshold.
[0024] The method in the above scheme further includes:
[0025] If the number of clusters is greater than or equal to the total number of random risk values, or if the second mean is less than the convergence threshold, the clustering ends, and the first value and random risk value corresponding to all cluster centers are sorted according to the first sorting method.
[0026] Each random risk value in the sorted random risk values is assigned to the cluster center corresponding to the first value in the sorting sequence.
[0027] The random risk value assigned to each cluster center is determined as the risk value corresponding to each first voucher in the corresponding cluster.
[0028] In the above scheme, before clustering the first voucher based on the value of each of the at least two set factors corresponding to the first voucher, the method further includes:
[0029] A random risk value is generated for each first voucher using a predefined random number function;
[0030] The value of the setting factor corresponding to the first certificate with the same random risk value is determined by multiple threads.
[0031] In the above scheme, the first voucher represents the voucher for setting up installment business; the setting factors include at least two of the following:
[0032] Overdue days;
[0033] The first ratio represents the quotient of the number of overdue periods to the total number of periods;
[0034] The second ratio represents the quotient of overdue resource share to total resource share.
[0035] The present invention also provides an electronic device, comprising:
[0036] The first clustering unit is used to cluster multiple first vouchers based on the value of each of the at least two set factors corresponding to the first voucher, and to obtain the first clustering result corresponding to each set factor.
[0037] The calculation unit is used to perform a weighted summation of the first risk value corresponding to each first voucher in each first clustering result to obtain the second risk value corresponding to each first voucher; the first risk value represents the random risk value corresponding to the cluster center of the cluster in the corresponding first clustering result of the first voucher;
[0038] The second clustering unit is used to cluster the plurality of first vouchers based on the second risk value corresponding to the first voucher, and obtain the second clustering result;
[0039] The classification unit is used to output the group classification result based on the second clustering result and the user corresponding to the first credential.
[0040] The present invention also provides an electronic device, comprising: a processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, performs the steps of the above-described group classification method.
[0041] In this embodiment of the invention, multiple first vouchers are clustered based on the value of each of the at least two set factors corresponding to the first voucher, resulting in a first clustering result for each set factor. A weighted sum of the first risk values corresponding to each first voucher in each first clustering result is then performed to obtain a second risk value for each first voucher. The first risk value represents the random risk value corresponding to the cluster center of the cluster in the corresponding first clustering result. The multiple first vouchers are then clustered based on the second risk values corresponding to the first vouchers to obtain a second clustering result. Based on the second clustering result and the users corresponding to the first vouchers, a group classification result is output. Since the second risk value corresponding to each first voucher is obtained by weighted summation of the first risk values corresponding to each first voucher in each first clustering result, the accuracy of the second clustering result obtained based on the second risk value can be improved. The group classification result obtained from the second clustering result can accurately reflect the quality of the customer group, making the group classification result more reasonable and improving the accuracy of the group classification result. Attached Figure Description
[0042] Figure 1 This is a schematic diagram illustrating the implementation process of the group classification method provided in an embodiment of the present invention;
[0043] Figure 2 This is a schematic diagram illustrating the implementation process of the clustering method in the population classification method provided in this embodiment of the invention;
[0044] Figure 3 This is a schematic diagram illustrating the implementation process of the clustering method in a population classification method provided in another embodiment of the present invention;
[0045] Figure 4 This is a schematic diagram illustrating the implementation process of the clustering method in a population classification method provided in another embodiment of the present invention;
[0046] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention;
[0047] Figure 6This is a schematic diagram of the hardware composition structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0049] Figure 1 This is a schematic diagram illustrating the implementation flow of the group classification method provided in an embodiment of the present invention, wherein the execution subject of the flow is an electronic device such as a terminal or server. Figure 1 The group classification methods shown include:
[0050] Step 101: Based on the value of each of the at least two set factors corresponding to the first voucher, cluster the multiple first vouchers to obtain the first clustering result corresponding to each set factor.
[0051] Here, the electronic device acquires the first credential corresponding to the designated service for each user among multiple users to be classified, assigns a random risk value to each acquired first credential, and establishes a correspondence between the first credential and the corresponding random risk value. In practical applications, each acquired first credential is tagged, thereby writing a random risk value to each first credential. The random risk value is generated by a random number function.
[0052] The electronic device determines the value of a first setting factor corresponding to each first credential based on the value of the set indicator contained in each acquired first credential. Based on the value of the first setting factor corresponding to each first credential, multiple first credentials are clustered to obtain a first clustering result corresponding to the first setting factor. Thus, a first clustering result corresponding to each setting factor is obtained. Here, the first setting factor refers to any setting factor corresponding to a first credential. Each user corresponds to at least one first credential. Different setting factors are used to evaluate whether a first credential has risk from different dimensions.
[0053] After clustering all first vouchers, the random risk value corresponding to the first voucher in the final first clustering result is updated to the random risk value corresponding to the cluster center of the cluster in the first clustering result, thus obtaining the first risk value corresponding to the first voucher. This determines the first risk value corresponding to each first voucher in each first clustering result. The cluster center is also called the clustering center. It should be noted that after clustering, the random risk value corresponding to each cluster center for each set factor can be a random risk value assigned to the cluster center, or it can be determined based on the random risk values of all cluster centers corresponding to the corresponding set factor and the value of the corresponding set factor. For example, random risk values can be assigned to cluster centers sorted in ascending order of their corresponding set factor values.
[0054] It should be noted that the value of the setting factor can be the value of a setting indicator corresponding to the first voucher, or it can be determined by the values of each of the two setting indicators corresponding to the first voucher. The correlation between different setting factors is low.
[0055] In practical applications, the first voucher represents a customer's borrowing or repayment behavior. For example, in the case of an installment repayment business, the indicators set for the first voucher should at least include the total loan amount P, the total number of loan periods T, the number of overdue days D, the overdue amount O, and the number of overdue periods R, and may also include the loan interest rate I and the principal repaid M.
[0056] In some embodiments, the first voucher represents a voucher for setting up installment transactions; the setting factors include at least two of the following:
[0057] Overdue days;
[0058] The first ratio represents the quotient of the number of overdue periods to the total number of periods;
[0059] The second ratio represents the quotient of overdue resource share to total resource share. In practical applications, the first voucher is the voucher for installment repayment business; the first ratio is the percentage of overdue periods, obtained by dividing the number of overdue periods by the total number of loan periods; the second ratio is the percentage of overdue amount, obtained by dividing the overdue amount by the total loan amount.
[0060] To expedite the determination of the values of each set factor, in some embodiments, before clustering the first voucher based on the values of each of the at least two set factors corresponding to the first voucher, the method further includes:
[0061] A random risk value is generated for each first voucher using a predefined random number function;
[0062] The value of the setting factor corresponding to the first certificate with the same random risk value is determined by multiple threads.
[0063] Here, when the electronic device obtains the first credential corresponding to the user to be classified, it generates a random risk value for each first credential using a set random number function, and writes the generated random risk value to the corresponding first credential. Based on the total number of generated random risk values, multiple threads are started, and each thread determines the value of each set factor corresponding to all first credentials for a given random risk value. The total number of started threads is less than or equal to the total number of generated random risk values.
[0064] If the total number of threads started is less than the total number of generated random risk values, at least one thread calculates the value of each set factor corresponding to each first voucher with a first random risk value after calculating the value of each set factor corresponding to each first voucher with a second random risk value.
[0065] When the total number of threads started is equal to the total number of random risk values generated, each thread corresponds to a random risk value. That is, each thread calculates the value of the set factor corresponding to the first voucher with the corresponding random risk value.
[0066] In practical applications, the random number function is Random().
[0067] To improve the accuracy of the first clustering results corresponding to each set factor, such as Figure 2 As shown, in some embodiments, when clustering multiple first credentials, the method includes:
[0068] Step 201: Based on the number of clusters and the total number of first vouchers, calculate the sorting number corresponding to each cluster center.
[0069] The cluster number refers to the total number of clusters included in the clustering result. In determining the first clustering result, the initial cluster number is a set value. This cluster number is updatable during the clustering process; please refer to the relevant description in step 208 below for the update method.
[0070] Here, the electronic device substitutes the number of clusters and the total number of first vouchers into the first calculation formula to calculate the sorting number corresponding to each cluster core. The first calculation formula is used to calculate the sorting number corresponding to the cluster core, and its expression is:
[0071] c j V represents the sorting index corresponding to the j-th cluster center; j is a positive integer less than or equal to x; sizeThe first credential represents the total number of first credentials; x represents the number of clusters, and x is updatable; k1, k2, and k3 are all set constants, which can be completely different, partially the same, or completely the same. In practical applications, k1 and k2 are both 2; k3 is 1.
[0072] In practical applications, the value of the setting factor corresponding to each cluster center is written into array C, array C = [C1, ..., C2]. j ]. C j The value of the setting factor corresponding to the j-th cluster center.
[0073] Step 202: Based on the sorting number corresponding to each cluster core, determine all cluster cores in the first voucher after sorting according to the first value; the first value includes the value of the set factor.
[0074] Here, in the process of determining the first clustering result corresponding to the first set factor, the electronic device sorts all the first vouchers in ascending order of the value of the first set factor to obtain the first sequence corresponding to the first set factor; based on the sorting number corresponding to each cluster core, the first voucher corresponding to each sorting number is determined in the first sequence corresponding to the first set factor to obtain all the cluster cores corresponding to the first set factor.
[0075] It should be noted that, in some embodiments, all first vouchers can be sorted in descending order of the value of the first set factor to obtain the first sequence corresponding to the first set factor.
[0076] Step 203: Calculate the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center.
[0077] Here, in the process of determining the first clustering result corresponding to the first set factor, the electronic device calculates the first difference between the first set factor and each cluster core based on the value of the first set factor corresponding to the first voucher and the value of the first set factor corresponding to each cluster core in all the cluster cores corresponding to the first set factor.
[0078] To improve clustering efficiency, in some embodiments, when calculating the first difference between the first value corresponding to the first credential and the first value corresponding to each cluster center, the method includes:
[0079] The square of the difference between the square of the first value corresponding to the first voucher and the square of the first value corresponding to the cluster center is determined as the first difference value.
[0080] Here, the electronic device calculates the square of the value of the first set factor corresponding to the first credential, and also calculates the square of the value of the first set factor corresponding to each cluster center; the difference between the square of the value of the first set factor corresponding to the first credential and the square of the value of the first set factor corresponding to the first cluster center is determined as the first difference between the first credential and the first cluster center. The first cluster center represents any one of the cluster centers.
[0081] In practical applications, the formula OV is used. i-j =(V i 2 -C j 2 ) 2 The first difference between the first voucher and each cluster center is calculated to increase the first difference between each first voucher and each cluster center, thereby accelerating the clustering speed and improving the clustering efficiency.
[0082] Among them, OV i-j V represents the first difference between the i-th first voucher and the j-th cluster center, where i is a positive integer less than or equal to n, n is the total number of first vouchers, and j is a positive integer less than or equal to the number of clusters; D V represents any one of at least two set factors; i 2 C represents the square of the value of any set factor corresponding to the i-th first voucher; j 2 The square of the value of the first set factor corresponding to the j-th cluster center.
[0083] Step 204: Add the first voucher to the cluster where the cluster center corresponding to the smallest first difference is located.
[0084] Here, the electronic device calculates the first difference between the first credential and each cluster core using a first set factor. From the calculated first differences, it identifies the smallest first difference and adds the first credential to the cluster containing the cluster core with the smallest first difference. Following this method, all first credentials are added to the clusters containing their respective smallest first differences, thereby obtaining the first clustering result corresponding to the first set factor.
[0085] It should be noted that, by following steps 201 to 204, the electronic device can obtain the first clustering result corresponding to each set factor.
[0086] To improve the accuracy of the first clustering results, during the clustering process for all first vouchers, the electronic device needs to determine whether the clustering has converged after each clustering cycle. If the clustering has not converged, all first vouchers need to be re-clustered; if the clustering has converged, this round of clustering ends, and all first vouchers are not re-clustered. Figure 3As shown, in some embodiments, after adding all first credentials to the corresponding cluster, the method further includes:
[0087] Step 205: Calculate the convergence threshold and the absolute difference for each cluster; wherein, the absolute difference represents the absolute value of the difference between the first value corresponding to the cluster center and the corresponding first mean; the first mean represents the mean of the first values corresponding to all first vouchers in the corresponding cluster; the convergence threshold represents the quotient of the second difference and the total number of random risk values; the second difference represents the difference between the largest and smallest first values.
[0088] Here, we will take the first clustering result corresponding to the first defined factor as an example for illustration:
[0089] Based on the value of the first set factor corresponding to each first voucher, the electronic device adds all first vouchers to the corresponding cluster according to steps 203 to 204, thereby completing one clustering. In all clusters corresponding to the first set factor, the maximum value and the minimum value of the first set factor are determined, and the difference between the maximum value and the minimum value of the first set factor is calculated to obtain the second difference. The quotient of the second difference and the total number of generated random risk values is determined as the convergence threshold.
[0090] Based on the values of the first set factor corresponding to all first vouchers in each cluster, the mean of the first set factor is calculated, thus obtaining the first mean for each cluster. In practical applications, the formula is used. Calculate the first mean for each cluster. Where NC j The first mean of the j-th cluster is represented by sum(Y). j Y represents the summation of the set factor values corresponding to all first credentials in the j-th cluster; j.size This represents the total number of first credentials contained in the j-th cluster.
[0091] In practical applications, after calculating the first mean for each cluster, the first mean is written into the first array NC, resulting in the first mean array NC; NC = [NC1, ..., NC2]. j ,…,NC x ]. NC j The first mean is represented by the j-th cluster.
[0092] Having calculated the first mean for each cluster, the absolute value of the difference between the first set factor corresponding to the cluster center of each cluster and the first mean corresponding to the corresponding cluster is calculated to obtain the absolute difference for each cluster. In practical applications, formula C is used. abs(j) =abs(C j -NC j Calculate the absolute difference for each cluster. Where C abs(j)The absolute difference corresponding to the j-th cluster is represented by abs(C). j -NC j The absolute value of the difference between the value of the set factor corresponding to the cluster center of the j-th cluster and the corresponding first mean is used to represent the calculation of the cluster center of the j-th cluster.
[0093] After calculating the convergence threshold and the absolute difference corresponding to each cluster, each absolute difference is compared with the convergence threshold to obtain a comparison result. If the comparison result indicates that the calculated absolute difference is greater than the calculated convergence threshold, it indicates that the clustering has not converged, and all first vouchers need to be re-clustered based on the value corresponding to the first set factor, and step 206 is executed. If all calculated absolute differences are less than or equal to the convergence threshold, it indicates that the clustering has converged, and step 207 is executed.
[0094] Step 206: If the calculated absolute difference is greater than the convergence threshold, determine the new cluster center of the corresponding cluster based on the first mean of each cluster, and execute the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center, as well as subsequent steps.
[0095] Here, when the calculated absolute difference is greater than the calculated convergence threshold, the first voucher corresponding to the first mean can be determined as the new cluster center of the corresponding cluster. Alternatively, in the corresponding cluster, the value of the first set factor with the smallest difference from the corresponding first mean can be determined, and the first voucher corresponding to the value of the first set factor can be determined as the new cluster center of the corresponding cluster.
[0096] For example, based on the first mean corresponding to each cluster, the value of the first set factor that is equal to the corresponding first mean is searched in the corresponding cluster; if the value of the first set factor that is equal to the corresponding first mean is found, the first voucher corresponding to the value of the first set factor is determined as the new cluster center of the corresponding cluster; if the value of the first set factor that is equal to the corresponding first mean is not found, the value of the first set factor that has the smallest difference from the first mean is determined in the corresponding cluster, and the first voucher corresponding to the value of the first set factor is determined as the new cluster center.
[0097] Once the new cluster center for the corresponding cluster is determined, steps 201 to 205 are executed based on the new cluster center for each cluster and the value of the first set factor, thereby re-clustering all the first credentials.
[0098] Step 207: If all calculated absolute differences are less than or equal to the convergence threshold, do not re-cluster the plurality of first vouchers.
[0099] Here, clustering convergence is indicated when all calculated absolute differences are less than or equal to the convergence threshold. The electronic device no longer re-clusters all first vouchers based on the value of the first set factor, ends the current round of clustering, and obtains the first clustering result corresponding to each set factor.
[0100] At this point, the electronic device can determine the random risk value corresponding to each cluster core in the first clustering result as the first risk value corresponding to each first voucher in the corresponding cluster; alternatively, it can update the random risk value corresponding to each cluster core based on the values of the set factors and random risk values corresponding to all cluster cores in the first clustering result, and determine the updated random risk value as the first risk value corresponding to each first voucher in the corresponding cluster. For example, for the first clustering result corresponding to the first set factor, the values of the first set factor and random risk values corresponding to all cluster cores are sorted; each random risk value in the sorted random risk values is assigned to the cluster core corresponding to the value of the first set factor at the corresponding sorting number; and the random risk value assigned to each cluster core is determined as the first risk value corresponding to each first voucher in the corresponding cluster.
[0101] It should be noted that the number of clusters included in the first clustering result corresponding to each set factor can be the same or different.
[0102] To improve the accuracy of the first clustering results corresponding to the set factor, if the calculated absolute difference is less than or equal to the convergence threshold, the clustering process is judged based on the number of clusters. If clustering is not complete, the number of clusters is increased to perform the next round of clustering on all first vouchers. When clustering is complete, the first clustering results corresponding to the set factor are output. Figure 4 As shown, in some embodiments, if the first mean of all calculated absolute differences is less than or equal to the convergence threshold, the method further includes:
[0103] Step 208: If the number of clusters is less than the total number of random risk values corresponding to all first vouchers, and the second mean is greater than or equal to the convergence threshold, the minimum value between the second value and the total number of random risk values is determined as the new number of clusters. Based on the new number of clusters and the first value corresponding to each first voucher, the plurality of first vouchers are re-clustered; wherein,
[0104] The second mean represents the absolute value of the mean of the differences between any two adjacent first means in the first mean array; the second value is determined by the number of clusters and the difference between the second mean and the convergence threshold.
[0105] Here, in the process of determining the first clustering result corresponding to the first set factor, the electronic device obtains the first mean array corresponding to the first set factor from the first mean of each cluster in the clustering result corresponding to the first set factor; calculates the difference between every two adjacent first means in the first mean array, and calculates the mean of the calculated difference, and determines the absolute value of the calculated mean as the second mean corresponding to the first set factor.
[0106] In practical applications, the formula is used. Calculate the second mean corresponding to the set factor. Wherein, NC size The total number of elements included in the first mean array corresponding to the first set factor; NC j+1 Represents the (j+1)th element in the first mean array; NC j Find the j-th element in the first mean array; j+1 is less than or equal to x.
[0107] Given a determined second mean corresponding to the first set factor, it is determined whether this second mean is greater than or equal to a convergence threshold, thus obtaining a first judgment result corresponding to the first set factor. If the first judgment result indicates that the second mean is greater than or equal to the convergence threshold, and the number of clusters is less than the total number of random risk values, then the number of clusters needs to be increased. Based on the new number of clusters and the value corresponding to the first set factor, a new round of clustering is performed on all first vouchers. At this time, the electronic device calculates the difference between the second mean corresponding to the first set factor and the number of clusters, and determines the minimum value among the sum of this difference, the number of clusters, and the total number of random risk values as the second value corresponding to the first set factor.
[0108] The electronic device determines the new cluster number by taking the minimum value between the second value corresponding to the first set factor and the total number of random risk values.
[0109] When the electronic device determines a new number of clusters, it executes steps 201 to 206, or steps 201 to 205 and steps 207 to 208, based on the new number of clusters, thereby performing a new round of clustering on multiple first vouchers based on the value of the first set factor corresponding to the first voucher.
[0110] It should be noted that clustering of all first vouchers is complete when the number of clusters is greater than or equal to the total number of random risk values, or when the second mean is less than the convergence threshold.
[0111] In this embodiment, during the clustering of multiple first vouchers, a second value can be calculated based on the current clustering results. If the number of clusters is less than the total number of random risk values, and the second mean is greater than or equal to the convergence threshold, the minimum of the second value and the total number of random risk values is determined as the new number of clusters. All first vouchers are then re-clustered based on this new number of clusters. Therefore, the number of clusters can be dynamically adjusted based on the current clustering results. Compared to re-clustering based on randomly adjusted cluster numbers, this reduces the number of clustering operations and improves clustering efficiency.
[0112] Since the random risk value corresponding to each first voucher is randomly generated, in order for the random risk value corresponding to each cluster center to truly reflect the risk level of the corresponding first voucher, such as... Figure 4 As shown, in some embodiments, the method further includes:
[0113] Step 209: If the number of clusters is greater than or equal to the total number of random risk values, or if the second mean is less than the convergence threshold, the clustering is terminated, and the first value and risk value corresponding to all cluster centers are sorted according to the first sorting method.
[0114] Step 210: Assign each random risk value in the sorted random risk values to the cluster center corresponding to the first value of the sorting sequence number;
[0115] Step 211: Determine the random risk value assigned to each cluster center as the risk value corresponding to each first voucher in the corresponding cluster.
[0116] Here, the first sorting method represents either an ascending or descending sorting method. The first value is the value of any set factor. For example, in determining the first clustering result corresponding to the first set factor, the first value is the value of the first set factor; in determining the first clustering result corresponding to the second set factor, the first value is the value of the second set factor.
[0117] Here, in the process of determining the first clustering result corresponding to the first set factor, if the number of clusters is greater than or equal to the total number of random risk values, or if the second mean is less than the convergence threshold, the clustering ends and the first clustering result corresponding to the first set factor is obtained; the values of the first set factor and the random risk values corresponding to all cluster centers are sorted respectively; each random risk value in the sorted random risk values is assigned to the cluster center corresponding to the value of the first set factor at the corresponding sorting number; the random risk value assigned to each cluster center is determined as the first risk value corresponding to each first voucher in the corresponding cluster.
[0118] For example, the smallest random risk value is assigned to the cluster center corresponding to the minimum value among the values of the first set factor; the largest random risk value is assigned to the cluster center corresponding to the maximum value among the values of the first set factor.
[0119] In this embodiment of the invention, when clustering is completed, each random risk value in the sorted random risk values is assigned to the cluster center corresponding to the first value of the corresponding sorting number. The random risk value assigned to each cluster center is determined as the first risk value corresponding to each first voucher in the corresponding cluster. This can improve the accuracy of the first risk value corresponding to the first voucher, thereby improving the accuracy of the second risk value determined by the first risk value, and further improving the accuracy of the second clustering result and the accuracy of the group classification result.
[0120] Step 102: For each first voucher, the first risk value corresponding to each first cluster result is weighted and summed to obtain the second risk value corresponding to each first voucher; the first risk value represents the random risk value corresponding to the cluster center of the cluster in the corresponding first cluster result.
[0121] Here, after determining the first risk value corresponding to each first voucher in each first clustering result, the electronic device performs a weighted summation of the first risk values corresponding to each first voucher in each first clustering result to obtain the second risk value corresponding to each first voucher.
[0122] In practical applications, each set factor has the same set weight. Of course, in some embodiments, each set factor can also be assigned a corresponding set weight based on its importance.
[0123] In practical applications, the setting factors corresponding to the first voucher include the number of overdue days, the first ratio, and the second ratio; based on the first risk value corresponding to each first voucher in each first cluster result, the mean of all first risk values corresponding to each first voucher is calculated to obtain the corresponding second risk value.
[0124] Step 103: Cluster the multiple first vouchers based on the second risk value corresponding to the first voucher to obtain the second clustering result.
[0125] Here, when the electronic device determines the second risk value corresponding to each first voucher, it clusters all the first vouchers based on the second risk value corresponding to each first voucher to obtain the second clustering result.
[0126] To improve the accuracy of the second clustering results, in some embodiments, when clustering multiple first credentials, the method includes:
[0127] Based on the number of clusters and the total number of first vouchers, calculate the sorting number corresponding to each cluster center;
[0128] Based on the sorting number corresponding to each cluster core, all cluster cores are determined in the first voucher after sorting according to the second risk value;
[0129] Calculate the first difference between the second risk value corresponding to the first voucher and the second risk value corresponding to each cluster center;
[0130] Add the first voucher to the cluster containing the cluster center corresponding to the smallest first difference.
[0131] Here, the number of clusters is the maximum number of clusters contained in all the final first clustering results. The process of clustering all first vouchers based on the second risk value corresponding to the first voucher is similar to the process of clustering all first vouchers based on the value of the set factor corresponding to the first voucher. Please refer to the relevant descriptions of steps 201 to 204 above for the implementation process.
[0132] To improve the accuracy of the second clustering results, during the clustering process for all first credentials, the electronic device needs to determine whether the clustering has converged after each clustering operation. If the clustering has not converged, all first credentials need to be re-clustered; if the clustering has converged, all first credentials are not re-clustered. In some embodiments, after adding all first credentials to the corresponding cluster, the method further includes:
[0133] Calculate the convergence threshold and the absolute difference for each cluster; where the absolute difference represents the absolute value of the difference between the second risk value corresponding to the cluster center and the corresponding first mean; the first mean represents the mean of the second risk values corresponding to all first vouchers in the corresponding cluster; the convergence threshold represents the quotient of the second difference and the total number of random risk values; the second difference represents the difference between the maximum second risk value and the minimum second risk value.
[0134] If the calculated absolute difference is greater than the convergence threshold, a new cluster center is determined based on the first mean of each cluster, and the first difference between the calculated second risk value corresponding to the first voucher and the second risk value corresponding to each cluster center, as well as subsequent steps, are executed; or
[0135] If all calculated absolute differences are less than or equal to the convergence threshold, the plurality of first credentials are not re-clustered.
[0136] The implementation process of the above steps is described in steps 205 to 207 above, and will not be repeated here.
[0137] To improve the accuracy of the second clustering results, in some embodiments, the method further includes, when the first mean of all calculated absolute differences is less than or equal to the convergence threshold:
[0138] If the number of clusters is less than the total number of random risk values, and the second mean is greater than or equal to the convergence threshold, the minimum of the second mean and the total number of random risk values is determined as the new number of clusters. Based on the new number of clusters and the second risk value corresponding to each first voucher, the plurality of first vouchers are re-clustered; wherein,
[0139] The second mean represents the absolute value of the mean of the differences between any two adjacent first means in the first mean array; the second value is determined by the number of clusters and the difference between the second mean and the convergence threshold.
[0140] The process of re-clustering the multiple first credentials is described in step 208 above and will not be repeated here.
[0141] It should be noted that in practical applications, since a suitable number of clusters has already been determined in the process of determining the first clustering result, in order to improve clustering efficiency, the second clustering result is determined by directly using the maximum number of clusters among all the first clustering results, without adjusting the number of clusters.
[0142] In some embodiments, if the number of clusters is greater than or equal to the total number of random risk values, or if the second mean is less than the convergence threshold, the clustering is terminated, and the second risk values and random risk values corresponding to all cluster centers are sorted according to the first sorting method.
[0143] Each random risk value in the sorted random risk values is assigned to the cluster center corresponding to the second risk value at the corresponding sorting number; that is, the second risk value corresponding to each cluster center is updated to the random risk value at the corresponding sorting number.
[0144] The second risk value assigned to each cluster core is determined as the second risk value corresponding to each first voucher in the corresponding cluster.
[0145] Step 104: Based on the second clustering result and the user corresponding to the first credential, output the group classification result.
[0146] Here, the electronic device classifies the users corresponding to the first credential based on the correspondence between the user and the first credential, and based on the second clustering result, to obtain the group classification result.
[0147] In this embodiment of the invention, multiple first vouchers are clustered based on the value of each of the at least two set factors corresponding to the first voucher, resulting in a first clustering result for each set factor. A weighted sum of the first risk values corresponding to each first voucher in each first clustering result is then performed to obtain a second risk value for each first voucher. The first risk value represents the random risk value corresponding to the cluster center of the cluster in the corresponding first clustering result. The multiple first vouchers are then clustered based on the second risk values corresponding to the first vouchers to obtain a second clustering result. Based on the second clustering result and the users corresponding to the first vouchers, a group classification result is output. Since the second risk value corresponding to each first voucher is obtained by weighted summation of the first risk values corresponding to each first voucher in each first clustering result, the accuracy of the second clustering result obtained based on the second risk value can be improved. The group classification result obtained from the second clustering result can accurately reflect the quality of the customer group, making the group classification result more reasonable.
[0148] The following example uses the first voucher as the voucher for installment repayment. The setting factors corresponding to the first voucher include the number of overdue days, the percentage of overdue periods, and the percentage of overdue amount. This example illustrates the process of classifying customers into groups:
[0149] First, based on the value of the set index for each first voucher, calculate the value of each set factor corresponding to each first voucher:
[0150] The first voucher's defined indicators include at least the total loan amount P, the total number of loan periods T, the number of overdue days D, the overdue amount O, and the number of overdue periods R. The percentage of overdue periods is obtained by dividing the number of overdue periods by the total number of loan periods; the percentage of overdue amount is obtained by dividing the overdue amount by the total loan amount.
[0151] When the electronic device obtains the first credential corresponding to the user to be classified, it generates a random risk value corresponding to each first credential through a set random number function, and writes the generated random risk value into the corresponding first credential. Based on the total number of generated random risk values, a corresponding number of threads are started, and each thread determines the value of each set factor corresponding to all first credentials for a random risk value.
[0152] The electronic device will store all the generated random risk values into a risk value array Gn.
[0153] For example, thread 1 calculates the value of each set factor corresponding to the first voucher with a random risk value of 1, and stores the value of each set factor into the corresponding array.
[0154] Among them, the array V corresponding to the number of overdue days D =[V D(1) V D(2) ,…,V D(n)]; The array V corresponding to the percentage of overdue periods R =[V R(1) V R(2) ,…,V R(n) ]; Array V corresponding to the percentage of overdue amounts O =[V O(1) V O(2) ,…,V O(n) ]. V D(i) V represents the number of overdue days corresponding to the i-th first voucher; R(i) V represents the percentage of overdue amount corresponding to the i-th first voucher; O(i) This represents the percentage of overdue amount corresponding to the i-th first voucher. i is a positive integer less than or equal to n, where n is the total number of first vouchers.
[0155] Secondly, based on the number of overdue days, the proportion of overdue periods, and the proportion of overdue amounts corresponding to the first voucher, clustering is performed on the n first vouchers to obtain the first clustering result for each set factor:
[0156] The following example illustrates how to cluster n first vouchers based on the number of overdue days corresponding to the first voucher:
[0157] (1) Obtain the number of clusters x, where x represents the number of clusters contained in each first clustering result, or the number of clusters; in practical applications, the initial value of x is 2;
[0158] (2) Based on the formula Calculate the sorting number corresponding to each cluster center; V size Represents the total number of first credentials.
[0159] (3) Array V D =[V D(1) V D(2) ,…,V D(n) After sorting in ascending order, we get sort(V) D ), from sort(V D Take the cth element from each of the following: j indivual The number of overdue days corresponding to the cluster center of the j-th cluster is used as the basis for determining the number of overdue days, and the extracted overdue days are sequentially written into the array C = [C1, C2, ..., C...]. j ,…C x ].
[0160] (4) Initialize x arrays Y1, Y2, ..., Y j ,…,Y x Y is used to store the overdue days corresponding to the first voucher in each cluster of the clustering results. j Let x represent the set of overdue days for the j-th cluster; j is a positive integer less than or equal to x.
[0161] (5) Traverse array V D For all elements in the set, for the element V that has been traversed... D(i) Calculate V respectively D(i) The first difference between the values of the elements in array C and the values of the elements in array C.
[0162] To increase the first difference between each first document and each cluster core, thereby accelerating the clustering speed and improving clustering efficiency, the formula OV is used. i-j =(V D(i) 2 -C j 2 ) 2 Calculate the first difference.
[0163] (6) After calculating V D(i) Given the first difference between the values of each element in array C, find the smallest first difference min(OV). i-j V D(i) Belongs to min(OV) i-j The corresponding C j V D(i) Store min(OV) i-j The corresponding C j The array Y corresponding to the cluster j V D(i) The corresponding first voucher is added to C. j The cluster to which it belongs, thus completing the process for V D(i) The corresponding first voucher is classified.
[0164] (7) For array V D After classifying the first vouchers corresponding to all elements in the dataset, the formula is used... Calculate the arrays Y1, Y2, ..., Y respectively. j ,…,Y x The first mean, where sum(Y) j ) represents the condition for all arrays Y j Sum all elements in Y; j.size Characteristic array Y j The total number of elements contained in the array is calculated, and the calculated mean values are stored in the first mean array NC, NC = [NC1, ..., NC2]. j ,…,NC x ]; among them, NC j The first mean is represented by the j-th cluster.
[0165] (8) Based on array C and array NC, using formula C abs(j) =abs(C j -NC j) Calculate the absolute difference for each cluster, and based on the formula Calculate the convergence threshold, C. th The convergence threshold of the table; max(V) D ) Representation array V D The maximum value in, min(V) D ) Representation array V D The minimum value in the risk value array Gn is given by Gn.size, which represents the total number of elements in the risk value array Gn.
[0166] (9) Based on C abs(j) and C th To determine whether the current round of clustering has converged.
[0167] (10) When C abs(i) >C th If the clustering fails, the array NC is used as the new array C, and (4) to (9) are executed again.
[0168] (11) When all C abs(i) >C th At that time, based on x and Gn.size, and based on T avg and C th Determine whether x needs to be increased further.
[0169] in, T avg It represents the absolute value of the mean of the differences between any two adjacent elements in array NC.
[0170] When x < Gn.size, and T avg ≥C th In this case, it indicates that the current round of clustering is not completed, and x needs to be increased to execute (12);
[0171] When x ≥ Gn.size, or T avg <C th In this case, the characterization clustering has been completed, and (13) is executed;
[0172] (12) Calculate the new x for the next round of clustering, new x = min(x + (T) avg -C th ()+1,Gn.size), based on the calculated new x, execute (4) to (11) to perform the next round of clustering on the n first credentials.
[0173] (13) After clustering is completed, obtain the first clustering result corresponding to the overdue days and the final array Y1, Y2, ..., Y j ,…,Y xThe elements in the final array C are sorted in ascending order to obtain sort(C). The random risk values corresponding to each element in array C are sorted in ascending order to obtain sort(Gn). The j-th random risk value in sort(Gn) is used as the first risk value and assigned to all the first vouchers in the cluster corresponding to the j-th element in sort(C). That is, the first risk value of the first voucher in the corresponding cluster is obtained.
[0174] Here, the first risk value corresponding to all first vouchers can be written into the first risk value array V. C(D) .
[0175] (14) The clustering process for the overdue period ratio and the overdue amount ratio is the same as the clustering process for the overdue days. For the overdue period ratio and the overdue amount ratio, clustering is performed according to (1) to (13) above to obtain the first risk value array V corresponding to the overdue period ratio. C(R) The first risk value array V corresponding to the percentage of overdue amounts C(O) .
[0176] Again, for each first credential in V C(D) V C(R) and V C(O) The average of the first risk values is calculated to obtain the second risk value for each first voucher; based on the second risk values corresponding to the first vouchers, the n first vouchers are clustered to obtain the second clustering result.
[0177] Here, through the formula Calculate the second risk value corresponding to each first voucher. V C(i) V represents the second risk value corresponding to the i-th first certificate; C(D)(i) Representing the i-th first credential in V C(D) The corresponding first risk value; V C(R)(i) Representing the i-th first credential in V C(R) The corresponding first risk value; V C(O)(i) Representing the i-th first credential in V C(O) The corresponding first risk value.
[0178] The electronic device writes the second risk value corresponding to all first vouchers into the second risk value array V. C Based on the second risk value array V C Cluster all first credentials as follows:
[0179] (1) Initialize the number of clusters X, X = max(Y) D.size Y R.size Y O.size ); where Y D.sizeY represents the total number of clusters in the final first clustering result corresponding to the number of overdue days; R.size The total number of clusters in the final first clustering result corresponding to the percentage of overdue periods; Y O.size The total number of clusters in the final first clustering result corresponding to the proportion of overdue amount.
[0180] (2) Based on the formula Calculate the sorting number corresponding to each cluster center.
[0181] (3) Arrange the second risk value array V in ascending order. C Sort the elements in the array using sort(V) C ), from sort(V C Take the cth element from each of the following: j indivual The extracted second risk values are used as the cluster center of the j-th cluster and are sequentially written into the array C' = [C'1, C'2, ..., C'']. j ,…C' x ].
[0182] (4) Initialize X arrays Y1, Y2, ..., Y j ,…,Y X It is used to store the second risk value corresponding to the first credential in each cluster of the clustering results.
[0183] (5) Traverse the array: Traverse the second risk value array V. C For all elements in the set, for the element V that has been traversed... C(i) Calculate V respectively C(i) The first difference between the values of the elements in array C' and the values of the elements in array C'.
[0184] To increase the first difference between each first document and each cluster core, thereby accelerating the clustering speed and improving clustering efficiency, the formula OV is used. i-j =(V C(i) 2 -C' j 2 ) 2 Calculate the first difference.
[0185] (6) After calculating V C(i) Given the first difference between the values of each element in array C', take the smallest first difference min(OV). i-j V C(i) Belongs to min(OV) i-j The corresponding C' j V C(i) Store min(OV) i-j The corresponding C' jThe array Y corresponding to the cluster j V C(i) Add the corresponding first voucher to C' j The cluster to which it belongs, thus completing the process for V C(i) The corresponding first voucher is classified.
[0186] (7) For the second risk value array V C After classifying the first vouchers corresponding to all elements in the dataset, the formula is used... Calculate the arrays Y1, Y2, ..., Y respectively. j ,…,Y X The first mean, where sum(Y) j ) represents the condition for all arrays Y j Sum all elements in Y; j.size Characteristic array Y j The total number of elements contained in the array is used to calculate the mean values, which are then stored in the first mean array NC', where NC' = [NC'1, ..., NC']. j ,…,NC' X ]; where NC' j The first mean is represented by the j-th cluster.
[0187] (8) Based on array C' and array NC', use formula C' abs(j) =abs(C' j -NC' j ) Calculate the absolute difference for each cluster, and based on the formula Calculate the convergence threshold, C' th The convergence threshold of the table; max(V) C The second risk value array V represents the risk value. C The maximum value in, min(V) C The second risk value array V represents the risk value. C The minimum value in.
[0188] (9) Based on C' abs(j) and C' th To determine whether the current round of clustering has converged.
[0189] (10) When C' abs(j) >C' th If the clustering fails, the array NC' is used as the new array C', and (4) to (9) are executed again.
[0190] (11) When all C' abs(j) >C' th When the clustering has converged, the final array Y1, Y2, ..., Y is obtained. j ,…,Y XThe final second clustering result is obtained; the elements in the final array C' are sorted in ascending order to get sort(C'); the random risk values corresponding to each element in array C' are sorted in ascending order to get sort(Gn'); the j-th random risk value in sort(Gn') is used as the final second risk value and assigned to all the first vouchers in the cluster corresponding to the j-th element in sort(C').
[0191] In the second clustering result, each array Y j It corresponds to a cluster.
[0192] Finally, based on the second clustering result and the user corresponding to the first credential, the group classification result is output.
[0193] To implement the method of the embodiments of the present invention, the embodiments of the present invention also provide an electronic device, such as... Figure 5 As shown, the electronic device includes:
[0194] The first clustering unit 51 is used to cluster multiple first vouchers based on the value of each of the at least two set factors corresponding to the first voucher, and to obtain the first clustering result corresponding to each set factor.
[0195] The calculation unit 52 is used to perform a weighted summation of the first risk value corresponding to each first voucher in each first clustering result to obtain the second risk value corresponding to each first voucher; the first risk value represents the random risk value corresponding to the cluster center of the cluster in the corresponding first clustering result of the first voucher;
[0196] The second clustering unit 53 is used to cluster the plurality of first vouchers based on the second risk value corresponding to the first voucher to obtain a second clustering result;
[0197] Classification unit 54 is used to output group classification results based on the second clustering result and the user corresponding to the first credential.
[0198] In some embodiments, when the first clustering unit 51 and the second clustering unit 53 cluster multiple first credentials, they are specifically used for:
[0199] Based on the number of clusters and the total number of first vouchers, calculate the sorting number corresponding to each cluster center;
[0200] Based on the sorting number corresponding to each cluster core, all cluster cores are determined in the first voucher after sorting according to the first value; the first value includes the value of the set factor or the second risk value;
[0201] Calculate the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center;
[0202] Add the first voucher to the cluster containing the cluster center corresponding to the smallest first difference.
[0203] When the first clustering unit 51 performs the above steps, the first value includes the value of the setting factor; when the second clustering unit 53 performs the above steps, the first value includes the second risk value.
[0204] In some embodiments, the first clustering unit 51 and the second clustering unit 53 are specifically used for:
[0205] The square of the difference between the square of the first value corresponding to the first voucher and the square of the first value corresponding to the cluster center is determined as the first difference value.
[0206] In some embodiments, after adding all the first credentials to the corresponding cluster, the first clustering unit 51 and the second clustering unit 53 are further configured to:
[0207] Calculate the convergence threshold and the absolute difference for each cluster; where the absolute difference represents the absolute value of the difference between the first value corresponding to the cluster center and the corresponding first mean; the first mean represents the mean of the first values corresponding to all first vouchers in the corresponding cluster; the convergence threshold represents the quotient of the second difference and the total number of random risk values; the second difference represents the difference between the largest and smallest first values;
[0208] If the calculated absolute difference is greater than the convergence threshold, a new cluster center for each cluster is determined based on the first mean corresponding to each cluster, and the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center, as well as subsequent steps, are executed; or
[0209] If all calculated absolute differences are less than or equal to the convergence threshold, the plurality of first credentials are not re-clustered.
[0210] In some embodiments, the first clustering unit 51 and the second clustering unit 53 are further configured to:
[0211] If the number of clusters is less than the total number of random risk values corresponding to all first vouchers, and the second mean is greater than or equal to the convergence threshold, the minimum value between the second value and the total number of random risk values is determined as the new number of clusters. Based on the new number of clusters and the first value corresponding to each first voucher, the plurality of first vouchers are re-clustered; wherein...
[0212] The second mean represents the absolute value of the mean of the differences between any two adjacent first means in the first mean array; the second value is determined by the number of clusters and the difference between the second mean and the convergence threshold.
[0213] In some embodiments, the first clustering unit 51 and the second clustering unit 53 are further configured to:
[0214] If the number of clusters is greater than or equal to the total number of random risk values, or if the second mean is less than the convergence threshold, the clustering ends, and the first value and random risk value corresponding to all cluster centers are sorted according to the first sorting method.
[0215] Each random risk value in the sorted random risk values is assigned to the cluster center corresponding to the first value in the sorting sequence.
[0216] The random risk value assigned to each cluster center is determined as the risk value corresponding to each first voucher in the corresponding cluster.
[0217] In some embodiments, the electronic device further includes:
[0218] The generation unit is used to generate a random risk value corresponding to each first voucher through a set random number function;
[0219] The determination unit is used to determine the value of the setting factor corresponding to the first certificate with the same random risk value through multiple threads.
[0220] In some embodiments, the first voucher represents a voucher for setting up an installment plan; the setting factors include at least two of the following:
[0221] Overdue days;
[0222] The first ratio represents the quotient of the number of overdue periods to the total number of periods;
[0223] The second ratio represents the quotient of overdue resource share to total resource share.
[0224] In practical applications, the first clustering unit 51, the calculation unit 52, the second clustering unit 53, the classification unit 54, the generation unit, and the determination unit can be implemented by processors in electronic devices, such as central processing units (CPUs), digital signal processors (DSPs), microcontroller units (MCUs), or field-programmable gate arrays (FPGAs). Of course, the processor needs to run programs stored in memory to implement the functions of each of the above program modules.
[0225] It should be noted that the electronic device provided in the above embodiments is only illustrated by the division of the above-described program modules when performing group classification. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the electronic device provided in the above embodiments and the group classification method embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0226] Based on the hardware implementation of the above program modules, and in order to implement the group classification method of the present invention, the present invention also provides an electronic device. Figure 6 This is a schematic diagram of the hardware composition structure of the electronic device provided in the embodiments of the present invention, such as... Figure 6 As shown, electronic device 6 includes:
[0227] Communication interface 61 enables information exchange with other devices, such as network devices;
[0228] The processor 62 is connected to the communication interface 61 to enable information interaction with other devices and, when running a computer program, executes the group classification method provided by one or more of the above-mentioned technical solutions. The computer program is stored in the memory 63.
[0229] Of course, in practical applications, the various components in electronic device 6 are coupled together through bus system 64. It can be understood that bus system 64 is used to realize the connection and communication between these components. In addition to the data bus, bus system 64 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in... Figure 6 The general labeled all buses as Bus System 64.
[0230] In this embodiment of the invention, memory 63 is used to store various types of data to support the operation of electronic device 6. Examples of such data include any computer program used to operate on electronic device 6.
[0231] It is understood that memory 63 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 63 described in this embodiment of the invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0232] The methods disclosed in the above embodiments of the present invention can be applied to processor 62, or implemented by processor 62. Processor 62 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 62 or by instructions in the form of software. The processor 62 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 62 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory 63. Processor 62 reads the program in memory 63 and combines its hardware to complete the steps of the aforementioned method.
[0233] Optionally, when the processor 62 executes the program, it implements the corresponding processes implemented by the terminal in the various methods of the embodiments of the present invention. For the sake of brevity, these will not be described in detail here.
[0234] In an exemplary embodiment, the present invention also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a first memory 63 storing a computer program, which can be executed by a processor 62 of a terminal to complete the steps described in the foregoing method. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0235] In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0236] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0237] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing module, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0238] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0239] It should be noted that the technical solutions described in the embodiments of the present invention can be combined arbitrarily without conflict.
[0240] It should be noted that the term "at least two" in the embodiments of the present invention means any combination of at least two of a plurality of elements, such as at least two of A, B, and C, and can mean any two or more elements selected from the set consisting of A, B, and C.
[0241] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for classifying groups, characterized in that, include: Based on the value of each of the at least two set factors corresponding to the first voucher, multiple first vouchers are clustered to obtain a first clustering result corresponding to each set factor. The first voucher represents a voucher for a set installment business. The at least two set factors include at least two of the following: overdue days, a first ratio representing the quotient of the overdue period to the total number of periods, and a second ratio representing the quotient of the overdue resource share to the total resource share. Each first voucher has a random risk value, which is generated by a set random number function. The weighted summation of the first risk values corresponding to each first voucher in each first clustering result yields the second risk value corresponding to each first voucher; the first risk value represents the random risk value corresponding to the cluster center of the cluster in the corresponding first clustering result of the first voucher; The multiple first vouchers are clustered based on the second risk value corresponding to the first voucher to obtain a second clustering result, including: calculating the sorting number corresponding to each cluster core based on the number of clusters and the total number of first vouchers; determining all cluster cores among the first vouchers sorted according to a first value based on the sorting number corresponding to each cluster core, wherein the first value includes the second risk value; calculating the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster core; and adding the first voucher to the cluster containing the cluster core with the smallest first difference to obtain the second clustering result. Based on the second clustering result and the user corresponding to the first credential, the group classification result is output.
2. The method according to claim 1, characterized in that, When clustering multiple first vouchers based on the value of each of at least two set factors corresponding to the first voucher, the method includes: Based on the number of clusters and the total number of first vouchers, calculate the sorting number corresponding to each cluster center; Based on the sorting number corresponding to each cluster core, all cluster cores are determined in the first voucher after sorting according to the first value; the first value includes the value of the set factor; Calculate the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center; The first certificate is added to the cluster containing the cluster center corresponding to the smallest first difference to obtain the first clustering result.
3. The method according to claim 1 or 2, characterized in that, When calculating the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center, the method includes: The square of the difference between the square of the first value corresponding to the first voucher and the square of the first value corresponding to the cluster center is determined as the first difference value.
4. The method according to claim 1 or 2, characterized in that, After adding all the first credentials to the corresponding cluster, the method further includes: Calculate the convergence threshold and the absolute difference for each cluster; where the absolute difference represents the absolute value of the difference between the first value corresponding to the cluster center and the corresponding first mean; the first mean represents the mean of the first values corresponding to all first vouchers in the corresponding cluster; the convergence threshold represents the quotient of the second difference and the total number of random risk values; the second difference represents the difference between the largest and smallest first values; If the calculated absolute difference is greater than the convergence threshold, a new cluster center for each cluster is determined based on the first mean corresponding to each cluster, and the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster center, as well as subsequent steps, are executed; or If all calculated absolute differences are less than or equal to the convergence threshold, the plurality of first credentials are not re-clustered.
5. The method according to claim 4, characterized in that, The method further includes: If the number of clusters is less than the total number of random risk values corresponding to all first vouchers, and the second mean is greater than or equal to the convergence threshold, the minimum value between the second value and the total number of random risk values is determined as the new number of clusters. Based on the new number of clusters and the first value corresponding to each first voucher, the plurality of first vouchers are re-clustered; wherein... The second mean represents the absolute value of the mean of the differences between any two adjacent first means in the first mean array; the second value is determined by the number of clusters and the difference between the second mean and the convergence threshold.
6. The method according to claim 5, characterized in that, The method further includes: If the number of clusters is greater than or equal to the total number of random risk values, or if the second mean is less than the convergence threshold, the clustering ends, and the first value and random risk value corresponding to all cluster centers are sorted according to the first sorting method. Each random risk value in the sorted random risk values is assigned to the cluster center corresponding to the first value in the sorting sequence. The random risk value assigned to each cluster center is determined as the risk value corresponding to each first voucher in the corresponding cluster.
7. The method according to claim 1, characterized in that, Before clustering the first voucher based on the value of each of the at least two set factors corresponding to the first voucher, the method further includes: A random risk value is generated for each first voucher using a predefined random number function; The value of the setting factor corresponding to the first certificate with the same random risk value is determined by multiple threads.
8. An electronic device, characterized in that, include: The first clustering unit is used to cluster multiple first vouchers based on the value of each of the at least two set factors corresponding to the first voucher, to obtain a first clustering result corresponding to each set factor. The first voucher represents a voucher for a set installment business. The at least two set factors include at least two of the following: overdue days, a first ratio representing the quotient of the overdue period to the total number of periods, and a second ratio representing the quotient of the overdue resource share to the total resource share. Each first voucher has a random risk value, which is generated by a set random number function. The calculation unit is used to perform a weighted summation of the first risk value corresponding to each first voucher in each first clustering result to obtain the second risk value corresponding to each first voucher; the first risk value represents the random risk value corresponding to the cluster center of the cluster in the corresponding first clustering result of the first voucher; The second clustering unit is used to cluster the plurality of first vouchers based on the second risk value corresponding to the first voucher, and obtain a second clustering result, including: calculating the sorting number corresponding to each cluster core based on the number of clusters and the total number of first vouchers; determining all cluster cores in the first vouchers sorted according to a first value based on the sorting number corresponding to each cluster core, wherein the first value includes the second risk value; calculating the first difference between the first value corresponding to the first voucher and the first value corresponding to each cluster core; and adding the first voucher to the cluster where the cluster core with the smallest first difference is located. The classification unit is used to output the group classification result based on the second clustering result and the user corresponding to the first credential.
9. An electronic device, characterized in that, include: The processor and the memory used to store computer programs that can run on the processor. When the processor is used to run the computer program, it performs the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Brokerage customer risk preference classification method based on big data
CN106228399A
User clustering method, device and apparatus
CN112381163A