Method for determining customer groups and related equipment

By performing pseudo-eigenvector generation, feature matrix construction and coefficient matrix regularization processing on customer active data, the problem of low accuracy of customer group clustering in the prior art is solved, and more efficient customer group division is achieved.

CN114418652BActive Publication Date: 2025-05-02AGRICULTURAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210099322.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-05-02
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

In the prior art, the enhanced spectral clustering algorithm cannot effectively suppress the correlation between different clustering objects when determining customer groups, resulting in low accuracy of customer groups clustering.

Method used

By obtaining the active data of the customer, a pseudo-eigenvector is generated and an eigenma is generated based on these pseudo-eigenvectors. Then, the first coefficient matrix and the first functional formula are determined, and the first coefficient matrix is ​​regularized using regularization parameters to obtain the second coefficient matrix. Finally, the active data is clustered using the second coefficient matrix to determine the customer group to which the customer belongs.

Benefits of technology

The second coefficient matrix obtained by regularization is sparse, which can effectively suppress the correlation between different clustered objects and improve the accuracy of customer group clustering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114418652B_ABST
    Figure CN114418652B_ABST
Patent Text Reader

Abstract

The present invention provides a method for determining a customer group and related equipment, the method comprising: generating a pseudo feature vector according to active data of each customer; determining a first coefficient matrix according to each pseudo feature vector; determining a first function according to the first coefficient matrix and the feature matrix, and obtaining a regularization parameter; regularizing the first coefficient matrix in the first function according to the regularization parameter to obtain a target function, and solving the minimum solution of the target function to obtain a second coefficient matrix; performing spectral clustering on each active data according to the second coefficient matrix to obtain multiple clusters of data, and determining the customer group to which the customer corresponding to each cluster of data belongs according to the characteristics of each cluster of data. In the present invention, the second coefficient matrix has a grouping effect and sparsity, and the spectral clustering performed using the second coefficient matrix can suppress the correlation between objects of different clusters and increase the correlation between objects of the same cluster, thereby improving the clustering accuracy of the customer group to which the customer belongs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of customer classification, and in particular to a method for determining a customer group and related equipment. Background Art

[0002] In order to attract new customers to apply for credit cards or savings cards, banks will launch various activities to attract customers. However, the bank's services and activities are not only to attract new customers, but also to prevent customers from choosing to cancel their accounts due to dissatisfaction with the bank's services.

[0003] At present, we can determine the customer group to which the customer belongs through the customer's active data. If the customer belongs to the group of customers who actively cancel their accounts, we can give targeted discounts to retain the customer. We can also find variables that are highly correlated with active account cancellation, and we can focus on these variables in future work.

[0004] When determining the customer group to which a customer belongs, a clustering algorithm is needed for cluster prediction, such as the enhanced spectral clustering algorithm. However, in addition to expressing the correlation between objects in the same cluster, clustering methods also suppress the correlation between objects in different clusters. However, the enhanced spectral clustering algorithm focuses on enhancing the correlation between objects in the same cluster by deriving a coefficient matrix that can amplify the correlation within the cluster, without suppressing the correlation between objects in different clusters, resulting in low clustering accuracy for the customer group to which the customer belongs. Summary of the invention

[0005] The present invention provides a method for determining a customer group and related equipment, which are used to solve the problem of low clustering accuracy of the customer group to which a customer belongs.

[0006] In one aspect, the present invention provides a method for determining a customer group, comprising:

[0007] Acquire active data corresponding to each customer, and generate each pseudo feature vector according to each active data;

[0008] Generate a characteristic matrix according to each of the pseudo-eigenvectors, and determine a first coefficient matrix according to the characteristic matrix;

[0009] Determine a first functional formula according to the first coefficient matrix and the feature matrix, and obtain a regularization parameter, wherein the first functional formula is a functional formula corresponding to a spectral clustering method;

[0010] Regularizing the first coefficient matrix in the first functional formula according to the regularization parameter to obtain a target functional formula, and solving the target functional formula for a minimum solution to obtain a second coefficient matrix;

[0011] Spectral clustering is performed on each of the active data according to the second coefficient matrix to obtain multiple clusters of data, and the customer group to which the customer corresponding to each cluster of data belongs is determined according to the characteristics of each cluster of data, wherein the customers to which multiple active data in each cluster of data belong belong to the same customer group.

[0012] In one embodiment, the regularization parameter includes an L1 norm, and the step of performing regularization processing on the first coefficient matrix in the first functional formula according to the regularization parameter to obtain the target functional formula includes:

[0013] Regularizing the first coefficient matrix in the first functional formula according to the L1 norm to obtain a second functional formula;

[0014] Obtaining the constraint conditions and the Langrange multiplier corresponding to the L1 norm;

[0015] The second functional formula is transformed according to the constraint conditions and the Lagrange multipliers to obtain the objective functional formula.

[0016] In one embodiment, the regularization parameter includes a trace cable amount, and the step of performing regularization processing on the first coefficient matrix in the first functional formula according to the regularization parameter to obtain the target functional formula includes:

[0017] Determine the trace cable amount based on the correlation and uncorrelation between the pseudo-feature vectors in the feature vector;

[0018] Regularization processing is performed on the first coefficient matrix in the first functional formula according to the trace cable amount to obtain the target functional formula.

[0019] In one embodiment, the step of performing spectral clustering on each of the active data according to the second coefficient matrix to obtain multiple clusters of data includes:

[0020] determining a correlation matrix for the second coefficient matrix;

[0021] Spectral clustering is performed on each of the active data according to the correlation matrix to obtain multiple clusters of data.

[0022] In one embodiment, the step of generating a feature matrix according to each of the pseudo feature vectors comprises:

[0023] Performing whitening processing on each of the pseudo feature vectors;

[0024] The feature matrix is ​​generated according to each of the pseudo feature vectors after whitening processing.

[0025] In one embodiment, the step of generating each pseudo feature vector according to each active data includes:

[0026] Generate a similarity matrix according to each of the active data, and obtain a nearest neighbor graph of the transferability K;

[0027] Determine a diagonal matrix corresponding to a similarity matrix, and determine a weight matrix according to the similarity matrix and the diagonal matrix;

[0028] Perform enhanced iteration on the weight matrix to obtain various pseudo feature vectors.

[0029] On the other hand, the present invention also provides a device for determining a customer group, comprising:

[0030] An acquisition module, used for acquiring active data corresponding to each customer, and generating each pseudo feature vector according to each active data;

[0031] A generating module, used for generating a characteristic matrix according to each of the pseudo characteristic vectors, and determining a first coefficient matrix according to the characteristic matrix;

[0032] A determination module, configured to determine a first functional formula according to the first coefficient matrix and the feature matrix, and obtain a regularization parameter, wherein the first functional formula is a functional formula corresponding to a spectral clustering method;

[0033] A processing module, configured to perform regularization processing on the first coefficient matrix in the first functional formula according to the regularization parameter to obtain a target functional formula, and solve the minimum solution of the target functional formula to obtain a second coefficient matrix;

[0034] A clustering module is used to perform spectral clustering on each of the active data according to the second coefficient matrix to obtain multiple clusters of data, and determine the customer group to which the customer corresponding to each cluster of data belongs according to the characteristics of each cluster of data, wherein the customers to which multiple active data in each cluster of data belong belong to the same customer group.

[0035] On the other hand, the present invention also provides a device for determining a customer group, comprising: a memory and a processor;

[0036] The memory stores computer-executable instructions;

[0037] The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method for determining the customer group as described above.

[0038] On the other hand, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method for determining a customer group as described above.

[0039] On the other hand, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for determining a customer group as described above is implemented.

[0040] The method and related devices for determining customer groups provided by the present invention obtain active data corresponding to each customer, generate each pseudo feature vector according to each active data, generate a feature matrix according to each feature vector, determine a first coefficient matrix according to the feature matrix, obtain the first coefficient matrix and a first function corresponding to the feature matrix, and then perform regularization processing on the first coefficient matrix in the first function according to a regularization parameter to obtain a target function, solve the minimum solution of the target function to obtain a second coefficient matrix, and then perform spectral clustering on each active data according to the second coefficient matrix to obtain multiple clusters of data, and finally determine the customer group to which the customer corresponding to each cluster of data belongs according to the characteristics of each cluster of data. In the present invention, a regularization process is performed on the first coefficient matrix in the first functional formula based on a regularization parameter, so that the second coefficient matrix solved by the objective functional formula after the regularization process has sparseness, and the spectral clustering performed using the second coefficient matrix can suppress the correlation between objects of different clusters; further, the first functional formula is a functional formula of the spectral clustering method, and the second coefficient matrix has the grouping effect originally contained in the spectral clustering method, that is, the spectral clustering performed using the second coefficient matrix can increase the correlation between objects of the same cluster, thereby improving the clustering accuracy of the customer group to which the customer belongs. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0042] Figure 1 A system architecture diagram for implementing a method for determining a customer group according to the present invention;

[0043] Figure 2 It is a flowchart of a first embodiment of a method for determining a customer group of the present invention;

[0044] Figure 3 This is a detailed flow chart of step S40 in the second embodiment of the method for determining a customer group of the present invention;

[0045] Figure 4 This is a detailed flow chart of step S40 in the third embodiment of the method for determining a customer group of the present invention;

[0046] Figure 5 This is a detailed flow chart of step S50 in the fourth embodiment of the method for determining a customer group of the present invention;

[0047] Figure 6A schematic diagram of a module of a device for determining a customer group of the present invention;

[0048] Figure 7 A schematic diagram of the hardware structure of a device for determining a customer group of the present invention.

[0049] The above drawings show clear embodiments of the present disclosure, which will be described in more detail below. These drawings and text descriptions are not intended to limit the scope of the present disclosure in any way, but to illustrate the concepts of the present disclosure to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0050] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0051] The present invention provides a method for determining a customer group, which can be Figure 1 The system architecture shown is implemented. Customers have active data, which can be data that can characterize the customer's activity level in the recent period of time. Active data can be the customer's transaction amount and number of transactions, etc. Figure 1 , the customer group determination device 100 obtains the active data of each customer, and each customer is exemplified as customer A, customer B, customer B, customer C, customer D, and customer E. The customer group determination device 100 determines each pseudo feature vector based on the active data of each customer, thereby determining a coefficient matrix for spectral clustering based on the pseudo feature vector, and finally classifies each active data through the coefficient matrix, thereby outputting each customer belonging to the same customer group. For example, the customer group determination device 100 outputs that customer A, customer B, and customer C belong to the same customer group, and the customer group determination device 100 outputs that customer D and customer E belong to the same customer group. Technical personnel of a financial institution can determine a customer group with a tendency to close accounts by using the data features of each customer group output by the customer group determination device 100, thereby designing activities to retain such customers based on the data features of the customer group with a tendency to close accounts.

[0052] The terms used in the embodiments of the present invention are explained below.

[0053] Clustering: The analytical process of dividing a collection of data objects into several categories.

[0054] Multi-scale data: data involving different spaces or different times.

[0055] Adaptation: Adaptation is a process of continuously approaching a goal, and the path it follows is represented by a mathematical model.

[0056] Spectral clustering: It is an algorithm evolved based on graph theory knowledge. This method regards all data as points in space. The farther the distance, the lower the edge weight value. The purpose of spectral clustering is to divide the current space into different areas so that the weight of points between areas is the lowest and the weight of points within areas is the largest.

[0057] Whitening: Reduce the redundancy of input data. Taking image processing as an example, the correlation between adjacent pixels in the image is very strong, so they cannot be input in full during training. It is necessary to ensure that the correlation between data features is low and the variance is equal.

[0058] Regularization: Set the coefficients of some unimportant variables in the fitting equation to 0 to make the parameter matrix more sparse and prevent the model from overfitting the training data.

[0059] Convex function: For a function f(x), if the second derivative of the function with respect to x is greater than 0, then the function is called a convex function.

[0060] Frobenius norm of a matrix: F-norm for short, the F-norm of A is denoted by ||A|| F ,

[0061] L1 norm and L2 norm of a vector: The L1 norm is the sum of the absolute values ​​of the elements of the vector, and the L2 norm is the square root of the sum of the squares of the elements of the vector.

[0062] Lagrange multiplier method: is a method for finding the extreme value of a multivariate function when its variables are subject to one or more constraints. This method can transform an optimization problem with n variables and k constraints into a problem of solving a system of equations with n+k variables. This method introduces one or a group of new unknowns, namely Lagrange multipliers. Lagrange multipliers are the coefficients of each vector in the linear combination of the gradient in the transformed equations (constraint equations). For example, when requiring the maximum value of f(x,y) under the condition of g(x,y)=c, a new Lagrange multiplier variable λ can be introduced. At this time, we only need to require the extreme value of the following Lagrange function:

[0063] L(x,y,λ)=f(x,y)+λ*(g(x,y)-c)

[0064] Singular value decomposition: This method decomposes the matrix, but does not require the decomposed matrix to be a square matrix. Assuming that matrix A is an m*n matrix, the singular value decomposition of matrix A is defined as:

[0065] A=UΣV T

[0066] Among them, U is an m*m matrix, V is an n*n matrix, satisfying U T U=I,V T V = I. Σ is an m*n matrix, and only the elements on the main diagonal are not zero. The elements on the main diagonal are called singular values.

[0067] The following specific embodiments are used to describe in detail the technical solutions of the present invention and how the technical solutions of the present application solve the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0068] Reference Figure 2 , Figure 2 This is a first embodiment of the method for determining a customer group of the present invention. The method for determining a customer group includes the following steps:

[0069] Step S10, obtaining active data corresponding to each customer, and generating each pseudo feature vector according to each active data.

[0070] In this embodiment, the execution subject is a device for determining a customer group. For ease of description, the device is used below to refer to the device for determining a customer group. The device obtains the customer's active data, which can be data representing the customer's activity level, such as the transaction amount and number of transactions of the customer in the recent period. The device generates each pseudo-feature vector based on each active data, for example, a multi-dimensional pseudo-feature vector is generated based on the transaction amount and number of transactions in each active data. A pseudo-feature vector refers to a vector that cannot be accurately used to represent the customer's activity level.

[0071] Furthermore, the device can generate a pseudo feature vector based on the nearest neighbor graph of the transferability K. The device first generates a similarity matrix S through each active data, and obtains the nearest neighbor graph of the transferability K. The similarity matrix S is the initial value of the first coefficient matrix Z, and S can be randomly generated through each active data. K is the number of categories, which can be determined manually based on experience.

[0072] Given a set of objects X = {x1, x2, ..., x n}, let X be the vertex set and E be the edge set, then the transitive K nearest neighbor graph G K =(X,E) is an undirected graph. Specifically, if there exists a sequence such that adjacent objects in the sequence are each other's K nearest neighbors, then the edge (x i , x j )∈E. The reachability matrix W (weight matrix) can be used to represent the transitivity K nearest neighbor graph. If (x i, x j )∈E, then W ij =1; otherwise 0.

[0073] Based on this, the device determines the diagonal matrix D corresponding to the similarity matrix, D ii =∑ j S ij , and the weight matrix is ​​determined by the similarity matrix and the focus matrix. Weight matrix W = D -1 S. Use enhanced iteration on the weight matrix W to generate pseudo feature vectors

[0074] Step S20: Generate a characteristic matrix according to each pseudo-eigenvector, and determine a first coefficient matrix according to the characteristic matrix.

[0075] The device generates a feature matrix based on each pseudo feature vector. For example, p pseudo feature vectors together form a p×n feature matrix X. The qth column of X is the object x q The eigenvector x q That is, robust spectral clustering uses pseudo-feature vectors to reduce the dimension of objects. Robust spectral clustering uses formula (1) to represent the correlation between X:

[0076] X=XZ+O (1)

[0077] Among them, Z∈R n×n Represents a coefficient matrix where each element Z ij Both describe an object x i Describe another object x j The ability of O∈R p×n is the error matrix. The more similar two objects are, the more likely one object is to represent the other.

[0078] Based on formula 1 and the characteristic matrix, the first coefficient matrix Z can be obtained.

[0079] In addition, after each pseudo-feature vector is generated, whitening is performed on each pseudo-feature vector to reduce redundancy of the pseudo-feature vector, and then a feature matrix is ​​generated according to each pseudo-feature vector after the whitening process.

[0080] Step S30, determining a first functional formula according to the first coefficient matrix and the characteristic matrix, and obtaining a regularization parameter, the first functional formula is a functional formula corresponding to the spectral clustering method.

[0081] The device can determine a first functional formula according to the first coefficient matrix and the characteristic matrix. The first functional formula is a functional formula corresponding to the spectral clustering method, and the spectral clustering method is a robust spectral clustering method. The first functional formula is specifically:

[0082]

[0083] Among them, the first term in the first functional formula reduces the error matrix O, the second term is the Frobenius norm of Z, and the third term regularizes Z through the transitivity K nearest neighbor graph.

[0084] Step S40, regularizing the first coefficient matrix in the first functional formula according to the regularization parameter to obtain the target functional formula, and solving the minimum solution of the target functional formula to obtain the second coefficient matrix.

[0085] The device obtains a regularization parameter, which may be an L1 norm or a trace cable amount. The first coefficient matrix Z in equation 2 is regularized by the regularization parameter to obtain an objective function.

[0086] The device solves the objective function to obtain a solution that minimizes the objective function, and the solution is the minimum solution as the second coefficient matrix Z*. That is, the device obtains the minimum solution of the objective function to obtain the second coefficient matrix Z*.

[0087] Step S50, performing spectral clustering on each active data according to the second coefficient matrix to obtain multiple clusters of data, and determining the customer group to which the customer corresponding to each cluster of data belongs according to the characteristics of each cluster of data, wherein the customers to which multiple active data in each cluster of data belong belong to the same customer group.

[0088] After obtaining the second coefficient matrix Z*, the device can use the second coefficient matrix Z* as a parameter to perform spectral clustering on each active parameter, thereby outputting multiple clusters of data. Each cluster of data includes multiple active data, and the customers of the active data belonging to the same cluster are the same customer group, that is, the customers to which the multiple active data in each cluster of data belong belong to the same customer group. The device can determine the customer group of the customer corresponding to each cluster of data based on the characteristics of each cluster of data, that is, the device can determine the common characteristics based on each active data of each cluster of data, and based on the common characteristics, it can be determined whether the cluster data is a customer with a tendency to cancel accounts, thereby designing activities to retain such customers based on the data characteristics of the customer group with a tendency to cancel accounts.

[0089] In the technical solution provided in this embodiment, active data corresponding to each customer is obtained, and each pseudo feature vector is generated according to each active data, and a feature matrix is ​​generated according to each feature vector, and a first coefficient matrix is ​​determined according to the feature matrix, and the first coefficient matrix and the first function corresponding to the feature matrix are obtained, and then the first coefficient matrix in the first function is regularized according to the regularization parameter to obtain the target function, and the minimum solution of the target function is solved to obtain the second coefficient matrix, so as to perform spectral clustering on each active data according to the second coefficient matrix to obtain multiple clusters of data, and finally, the customer group to which the customer corresponding to each cluster of data belongs is determined according to the characteristics of each cluster of data. In the present invention, a regularization process is performed on the first coefficient matrix in the first functional formula based on a regularization parameter, so that the second coefficient matrix solved by the objective functional formula after the regularization process has sparseness, and the spectral clustering performed using the second coefficient matrix can suppress the correlation between objects of different clusters; further, the first functional formula is a functional formula of the spectral clustering method, and the second coefficient matrix has the grouping effect originally contained in the spectral clustering method, that is, the spectral clustering performed using the second coefficient matrix can increase the correlation between objects of the same cluster, thereby improving the clustering accuracy of the customer group to which the customer belongs.

[0090] Reference Figure 3 , Figure 3 This is a second embodiment of the method for determining a customer group of the present invention. Based on the first embodiment, step S40 includes:

[0091] Step S41, regularizing the first coefficient matrix in the first functional formula according to the L1 norm to obtain a second functional formula.

[0092] In this embodiment, the robust spectral clustering method generates a coefficient matrix Z with a grouping effect, so that highly related objects have similar vector representations in Z. However, in this embodiment, it is also hoped that the connection matrix of objects from different clusters has sparseness. In order to enhance sparsity, Z can be regularized using the L1 norm, that is, the first coefficient matrix in the first functional formula is regularized using the L1 norm to obtain the second functional formula:

[0093]

[0094] Among them, the second term in the second functional formula is the L1 norm of Z.

[0095] Step S42, obtaining the constraint conditions and the Lagrange multipliers corresponding to the L1 norm.

[0096] In order to avoid any x i (x i is an element of X) contains x iBy itself, add the constraint diag(Z)=0, where diag(Z) is the main diagonal element of Z. Constraint: stdiag(Z)=0.

[0097] Step S43, transforming the second functional formula according to the constraint conditions and the Lagrange multipliers to obtain the target functional formula.

[0098] The second functional form is convex, so the second functional form is equivalent to:

[0099]

[0100] stJ=Z-Diag(Z)

[0101] Among them, Diag(Z) returns a diagonal matrix with the main diagonal vector Z. Equation (4) can be solved by the inexact enhanced Lagrange multiplier method, that is, the second function is transformed by Lagrange multipliers to obtain:

[0102]

[0103] Where Y is the Lagrange multiplier and μ>0 is the penalty parameter, which can be set by yourself. By updating one variable and fixing the other variables alternately, L can be minimized. To update J, you can set have

[0104] J=(X T X+α2I+μI) -1 (X T X+α2W-Y+μZ-μ·Diag(Z)) (5)

[0105] To update Z, we set And concluded

[0106]

[0107] Among them, T η (·) is the shrinkage threshold operator acting on a given matrix, defined as T η (v)=(|v|-η) + sgn(v). If x is non-negative, the operator (x) + returns x; otherwise, it returns 0. If x is positive, the operator sgn(x) returns 1; if x is negative, the operator sgn(x) returns -1; otherwise, it returns 0. Algorithm 1 (the above calculation process is Algorithm 1) shows the process of generating the sparse coefficient matrix Z using the inexact enhanced Lagrange multiplier method. The minimum solution of Z in Equation 6 is the second coefficient matrix Z*.

[0108] Algorithm 1 uses the inexact enhanced Lagrange multiplier method to solve equation (3)

[0109]

[0110] In the technical solution provided in this embodiment, the device regularizes the first coefficient matrix of the first functional formula according to the L1 norm to obtain a second functional formula, obtains the constraint conditions and Lagrange multipliers corresponding to the L1 norm, and then transforms the second functional formula through the constraint conditions and Lagrange multipliers to obtain the target functional formula, thereby obtaining a second coefficient matrix with sparsity based on the target functional formula, so that when each active data is spectrally clustered, the correlation between objects in different clusters is suppressed.

[0111] Reference Figure 4 , Figure 4 This is a third embodiment of the method for determining a customer group of the present invention. Based on the first embodiment, step S40 includes:

[0112] Step S44, determining the trace cable amount based on the correlation and irrelevance between the pseudo feature vectors in the feature vector.

[0113] Step S45, regularizing the first coefficient matrix in the first functional formula according to the trace cable amount to obtain the target functional formula.

[0114] In this embodiment, although the coefficient matrix Z derived by L1 regularization formula (3) has sparseness for unrelated objects, regularization also weakens the connection between related objects, which is not conducive to the grouping effect. In order to construct a matrix that has a grouping effect for highly related objects and is sparse for unrelated objects, in this embodiment, Z is regularized by the trace lasso method. Let X be the feature matrix, whose qth column is the object x q The eigenvector x q , for x q Normalize so that for any 1≤q≤n, Let z be the coefficient vector corresponding to the object x, that is, the vector in the coefficient matrix Z, then the trace cable of z is defined as:

[0115] Ω(z)=||X Diag(z)|| * (7)

[0116] Among them, Diag(z) is a diagonal matrix with z as the main diagonal. The trace cable of X can be distinguished from other commonly used norms of X (such as L1 norm and L2 norm), which are:

[0117]

[0118] Among them, e q is the vector of the canonical basis. TX is considered as the correlation matrix of the objects. Then, if X T X=I, that is, the object is irrelevant, because x i and e i The orthogonality of , the trace cable quantity is equal to the L1 norm:

[0119]

[0120] On the other hand, if all objects are highly correlated and have the same eigenvector x, that is, X T X=11 T , then the trace Lasso is equivalent to the L2 norm:

[0121] ||XDiag(z)|| * =||xz T || * =||x||2||z||2=||z||2 (10)

[0122] After considering these two extreme cases, for other cases, the value of the trace cable is between the L1 norm and the L2 norm:

[0123] ||z||2≤||XDiag(z)|| * ≤||z||1 (11)

[0124] It can be understood that the trace cable amount can be determined based on the correlation and irrelevant relationship between each pseudo-feature vector in the feature vector, and Formula 11 is the trace cable amount.

[0125] The trace cable is applied to Z in equation (2) for regularization. Given an object x, the function is obtained:

[0126]

[0127] The above objective function is a convex function and can be solved by the inexact enhanced Lagrange multiplier method. First, the problem is converted to:

[0128]

[0129] The enhanced Lagrangian function of formula (13):

[0130]

[0131] Where λ1, λ2 and Y are Lagrange multipliers, and μ>0 is the penalty parameter. Another strategy can be used to update The update rule is as follows:

[0132]

[0133] In formula (14), the Diag() function is overloaded. This function returns a diagonal matrix whose main diagonal is the diagonal of the parameter matrix. diag() returns the main diagonal vector of the parameter matrix.

[0134]

[0135]

[0136] In addition, updating J is equivalent to finding the optimal solution of equation (17):

[0137]

[0138] Equation (17) is a convex function and has a solution that can be solved by singular value decomposition. Assume that for a matrix of rank r Perform singular value decomposition, where Σ=Diag([σ1,...,σ r ]), σ i is the i-th largest singular value of the matrix. The solution of equation (17) is:

[0139] set up (18)

[0140] The process of solving equation (12) is shown in Algorithm 2:

[0141] Algorithm 2 uses the inexact enhanced Lagrange multiplier method to solve equation (12)

[0142]

[0143] In the technical solution provided in this embodiment, the device determines the trace cable amount according to the correlation and irrelevant relationship between each pseudo-feature vector of the feature vector, and transforms the first functional formula according to the trace cable amount to obtain the target functional formula, thereby obtaining a second coefficient matrix with sparsity based on the target functional formula, so that when each active data is spectrally clustered, the correlation between objects in different clusters is suppressed.

[0144] Reference Figure 5 , Figure 5 This is a fourth embodiment of the method for determining a customer group of the present invention, based on any one of the first to third embodiments, step S50 includes:

[0145] Step S51, determining the correlation matrix of the second coefficient matrix.

[0146] Step S52: spectral clustering is performed on each active data according to the correlation matrix to obtain multiple clusters of data.

[0147] In this embodiment, when clustering is performed, X={x1, ..., xn} is given, and equation (12) is solved for each data. And the coefficient matrix Construct the optimal solution. In order to avoid any x i The expression contains x i itself, when determining , it is necessary to delete the i-th column vector from the behavior data X.

[0148] But Z * may be asymmetric and contain negative values. To address this problem, the correlation matrix can be calculated Because |Z * |,|(Z * ) T | has a grouping effect, so There is also a grouping effect. In addition, since the trace cable amount will be adaptively adjusted to the L1 norm or L2 norm, Z * Enhanced the sparsity of different clustering objects.

[0149] The device determines a correlation matrix according to the second coefficient matrix, and performs spectral clustering on each active data according to the correlation matrix to obtain multiple clusters of data.

[0150] Algorithm 3 outlines the spectral clustering of each active data as follows:

[0151] Algorithm 3 Multi-scale data adaptive spectral clustering algorithm based on correlation

[0152]

[0153] Among them, C is active data.

[0154] In the technical solution provided in this embodiment, the correlation matrix is ​​determined based on the second coefficient matrix, so that spectral clustering is performed on each active data according to the correlation matrix, so that the device accurately outputs the clustering result.

[0155] The present invention also provides a device for determining a customer group, referring to Figure 6 , the customer group determination device 600 includes:

[0156] An acquisition module 610 is used to acquire active data corresponding to each customer and generate each pseudo feature vector according to each active data;

[0157] A generating module 620, configured to generate a feature matrix according to each pseudo feature vector, and determine a first coefficient matrix according to the feature matrix;

[0158] A determination module 630 is used to determine a first functional formula according to the first coefficient matrix and the characteristic matrix, and obtain a regularization parameter, wherein the first functional formula is a functional formula corresponding to the spectral clustering method;

[0159] A processing module 640 is used to perform regularization processing on the first coefficient matrix in the first functional formula according to the regularization parameter to obtain the target functional formula, and solve the minimum solution of the target functional formula to obtain the second coefficient matrix;

[0160] The clustering module 650 is used to perform spectral clustering on each active data according to the second coefficient matrix to obtain multiple clusters of data, and determine the customer group to which the customer corresponding to each cluster of data belongs according to the characteristics of each cluster of data, wherein the customers to which multiple active data in each cluster of data belong belong to the same customer group.

[0161] In one embodiment, the customer group determination device 600 includes:

[0162] A processing module 640 is used to perform regularization processing on the first coefficient matrix in the first functional formula according to the L1 norm to obtain a second functional formula;

[0163] An acquisition module 610 is used to acquire the constraint conditions and the Lagrange multipliers corresponding to the L1 norm;

[0164] The processing module 640 is used to transform the second functional formula according to the constraint conditions and the Lagrange multipliers to obtain the target functional formula.

[0165] In one embodiment, the customer group determination device 600 includes:

[0166] A determination module 630, configured to determine the amount of the trace cable based on the correlation and irrelevance between the pseudo feature vectors in the feature vector;

[0167] The processing module 640 is used to perform regularization processing on the first coefficient matrix in the first functional formula according to the trace cable amount to obtain the target functional formula.

[0168] In one embodiment, the customer group determination device 600 includes:

[0169] A determination module 630, configured to determine a correlation matrix of the second coefficient matrix;

[0170] The clustering module 650 is used to perform spectral clustering on each active data according to the correlation matrix to obtain multiple clusters of data.

[0171] In one embodiment, the customer group determination device 600 includes:

[0172] A processing module 640 is used to perform whitening processing on each pseudo feature vector;

[0173] The generating module 620 is used to generate a feature matrix according to each pseudo feature vector after whitening processing.

[0174] In one embodiment, the customer group determination device 600 includes:

[0175] A generating module 620, for generating a similarity matrix according to each active data, and obtaining a nearest neighbor graph of the transitivity K;

[0176] A determination module 630 is used to determine a diagonal matrix corresponding to the similarity matrix, and determine a weight matrix according to the similarity matrix and the diagonal matrix;

[0177] The processing module 640 is used to perform enhanced iteration on the weight matrix to obtain each pseudo feature vector.

[0178] Figure 7 The figure is a hardware structure diagram of a device for determining a customer group according to an exemplary embodiment.

[0179] The customer group determination device 700 may include: a processor 701, such as a CPU, a memory 702, and a transceiver 703. Those skilled in the art will appreciate that Figure 7 The structure shown in the figure does not constitute a limitation on the device for determining the customer group, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently. The memory 702 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0180] The processor 701 may call a computer program stored in the memory 702 to complete all or part of the steps of the above-mentioned method for determining a customer group.

[0181] The transceiver 703 is used to receive information sent by an external device and send information to the external device.

[0182] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a device for determining a customer group, enables the device for determining a customer group to execute the above-mentioned method for determining a customer group.

[0183] A computer program product comprises a computer program. When the computer program is executed by a processor of a device for determining a customer group, the device for determining a customer group can execute the method for determining a customer group.

[0184] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0185] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for determining a customer group, characterized in that: include: Acquire active data corresponding to each customer, and generate each pseudo feature vector according to each active data; Generate a feature matrix X based on each of the pseudo feature vectors, and use the formula Determine a first coefficient matrix Z; wherein O is an error matrix; A first functional formula is determined according to the first coefficient matrix Z and the feature matrix X, and the first functional formula is: , wherein the first term in the first functional formula is used to reduce the error matrix O, the second term is the Frobenius norm of the Z, and the third term represents the regularization of the Z by the transitive K nearest neighbor graph; W is a weight matrix used to represent the transitive K nearest neighbor graph; the first functional formula is the functional formula corresponding to the spectral clustering method; Obtain a regularization parameter, perform regularization processing on the first coefficient matrix Z in the first functional formula according to the regularization parameter to obtain an objective functional formula; and solve the minimum solution of the objective functional formula to obtain a second coefficient matrix ; The second coefficient matrix As a parameter, spectral clustering is performed on each of the active data to obtain multiple clusters of data, and the customer group to which the customer corresponding to each cluster of data belongs is determined according to the characteristics of each cluster of data, wherein each cluster of data includes multiple active data, and the customers of the active data belonging to the same cluster are the same customer group.

2. The method for determining a customer group according to claim 1, characterized in that: The regularization parameter includes an L1 norm, and the step of performing regularization processing on the first coefficient matrix Z in the first functional formula according to the regularization parameter to obtain the target functional formula includes: The first coefficient matrix Z in the first functional formula is regularized according to the L1 norm to obtain a second functional formula ; Wherein, the second term in the second functional formula is the L1 norm of Z; Introduce the variable J, and according to the constraints corresponding to the L1 norm , the second functional formula is equivalent to: , ; Where Diag(Z) is a diagonal matrix with Z as the main diagonal; Get the Lagrange multiplier Y; The objective functional formula is obtained by transforming the equivalent second functional formula according to the constraint conditions and the Lagrange multipliers. , where μ is the penalty parameter.

3. The method for determining a customer group according to claim 1, characterized in that: The regularization parameter includes a trace cable amount, and the step of performing regularization processing on the first coefficient matrix in the first functional formula according to the regularization parameter to obtain the target functional formula includes: Based on the correlation and uncorrelation between the pseudo-eigenvectors in the eigenvector, the trace cable amount of the vector z in the Z is determined. ;in, , Diag(z) is a diagonal matrix with z as the main diagonal; The trace cable amount is applied to the first coefficient matrix Z in the first functional formula for regularization processing to obtain the objective functional formula ,x is the preset object.

4. The method for determining a customer group according to claim 1, characterized in that: The second coefficient matrix As a parameter, spectral clustering is performed on each of the active data to obtain multiple clusters of data, including: Determine the second coefficient matrix The correlation matrix of Spectral clustering is performed on each of the active data according to the correlation matrix to obtain multiple clusters of data.

5. The method for determining a customer group according to any one of claims 1 to 4, characterized in that: The step of generating a feature matrix X according to each of the pseudo feature vectors comprises: Performing whitening processing on each of the pseudo feature vectors; The feature matrix is ​​generated according to each of the pseudo feature vectors after whitening processing.

6. The method for determining a customer group according to any one of claims 1 to 4, characterized in that: The step of generating each pseudo feature vector according to each active data comprises: Generate a similarity matrix according to each of the active data, and obtain a nearest neighbor graph of the transferability K; Determine a diagonal matrix corresponding to a similarity matrix, and determine a weight matrix according to the similarity matrix and the diagonal matrix; Perform enhanced iteration on the weight matrix to obtain various pseudo feature vectors.

7. A device for determining a customer group, characterized in that: include: An acquisition module, used for acquiring active data corresponding to each customer, and generating each pseudo feature vector according to each active data; A generating module is used to generate a feature matrix X according to each of the pseudo feature vectors, and according to the formula Determine a first coefficient matrix Z; wherein O is an error matrix; A determination module is used to determine a first functional formula according to the first coefficient matrix Z and the feature matrix X, where the first functional formula is: , wherein the first term in the first functional formula is used to reduce the error matrix O, the second term is the Frobenius norm of the Z, and the third term represents the regularization of the Z by the transitive K nearest neighbor graph; W is a weight matrix, which is used to represent the transitive K nearest neighbor graph; the first functional formula is the functional formula corresponding to the spectral clustering method; obtain the regularization parameter; A processing module, configured to perform regularization processing on the first coefficient matrix Z in the first functional formula according to the regularization parameter to obtain a target functional formula, and solve the minimum solution of the target functional formula to obtain a second coefficient matrix ; Clustering module, used to transform the second coefficient matrix As a parameter, spectral clustering is performed on each of the active data to obtain multiple clusters of data, and the customer group to which the customer corresponding to each cluster of data belongs is determined according to the characteristics of each cluster of data, wherein each cluster of data includes multiple active data, and the customers of the active data belonging to the same cluster are the same customer group.

8. A device for determining a customer group, characterized in that: include: Memory and processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method for determining a customer group according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method for determining a customer group according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for determining a customer group according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method and device for analyzing user transaction behavior

    CN106127493A

  • Crowd classification method and device based on spectral clustering and medium

    CN110276382A