Prediction Method, Device, Equipment and Medium for Credit Risk
By using the local optimal block conjugation gradient method to reduce the dimensionality of the Graplace matrix in credit risk prediction, the problem of huge calculation volume when adding new customers or new information items in the existing technology is solved, and a fast and efficient credit risk prediction is achieved.
Patent Information
- Application Number
- CN202210477902.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-29
AI Technical Summary
In banks with more than tens of millions of customer data, when adding new customers or new information items, the existing technology needs to recalculate and classify all data, resulting in huge calculations, time-consuming and resource-consuming.
The local optimal block conjugation gradient method is used to reduce the dimensionality of the graph Laplace matrix, and the graph Laplace matrix is calculated by spectral clustering method, and clustered based on the feature matrix to quickly predict credit risk.
It greatly reduces the computational volume and training speed of neural networks, has fast computing speed and small memory occupancy, and can quickly respond to data changes.
Smart Images

Figure CN114707762B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data technology, and particularly to a method, apparatus, device, and medium for predicting credit risk. Background Art
[0002] Currently, for big data analysis of credit risk, the information of each customer is usually taken as a sample, and machine learning methods are used to classify these customers to obtain the credit risk level of each customer. However, for banks with customer data of more than ten million, when adding a new customer or a new piece of information, the machine needs to recalculate and classify all the data, which will generate a huge amount of computation, consuming time and resources. Summary of the Invention
[0003] In view of the above problems, the present disclosure provides a method, apparatus, device, medium, and program product for predicting credit risk.
[0004] According to a first aspect of the present disclosure, a method for predicting credit risk is provided, including the following steps:
[0005] Obtain the authorization of the customer for obtaining personal information;
[0006] When the authorization of the customer for obtaining personal information is obtained, obtain the personal information of n customers, where the personal information of each customer includes N information items, and the N information items are all related to credit risk, n is an integer greater than or equal to 1, and N is an integer greater than or equal to 2;
[0007] Quantify the personal information of the n customers to obtain a personal information matrix, where the personal information matrix is an n-row and N-column matrix, and each row of the personal information matrix represents the quantified personal information of a customer;
[0008] For the personal information matrix, use the spectral clustering method to calculate the Laplacian matrix corresponding to the personal information matrix, where the Laplacian matrix is an n-row and n-column matrix;
[0009] Use the locally optimal block conjugate gradient method to reduce the dimension of the Laplacian matrix to obtain a feature matrix corresponding to the Laplacian matrix, where the feature matrix is an n-row and b-column matrix, b is a positive integer and 1≤b<n;
[0010] Based on the feature matrix, use the clustering method to classify the n customers; and
[0011] Predict the credit risk of the n customers according to the classified n customers.
[0012] According to the prediction method in the embodiments of the present disclosure, by using the locally optimal block conjugate gradient method, the best gradient direction can be quickly searched, the dimensionality of N information items of each customer is reduced, and then the eigen-space of the Laplacian matrix can be quickly obtained, with fast calculation speed and small memory occupation, greatly reducing the calculation amount and training speed of the neural network.
[0013] According to some exemplary embodiments, the dimensionality reduction of the Laplacian matrix by using the locally optimal block conjugate gradient method to obtain a feature matrix corresponding to the Laplacian matrix specifically includes:
[0014] Using an iterative method to determine a search direction such that the search direction gradually coincides with the direction of each column vector in the to-be-determined feature matrix; and
[0015] Determining a feature matrix corresponding to the Laplacian matrix according to the determined search direction.
[0016] According to some exemplary embodiments, the using an iterative method to determine a search direction such that the search direction gradually coincides with the direction of each column vector in the to-be-determined feature matrix specifically includes:
[0017] Based on the Laplacian matrix, an intermediate matrix is obtained, where the intermediate matrix is an n-row and b-column matrix;
[0018] Calculating the eigenvalues and eigenvectors of the intermediate matrix; and
[0019] Generating a first sub-matrix according to the eigenvectors of the intermediate matrix, where each column vector in the first sub-matrix represents a search direction, and the search directions represented by each column vector in the first sub-matrix respectively correspond to the directions of each column vector in the to-be-determined feature matrix, and the first sub-matrix is an n-row and b-column matrix.
[0020] According to some exemplary embodiments, the using an iterative method to determine a search direction such that the search direction gradually coincides with the direction of each column vector in the to-be-determined feature matrix further specifically includes:
[0021] Generating a second sub-matrix according to the eigenvalues and eigenvectors of the intermediate matrix and the first sub-matrix, and the second sub-matrix is an n-row and b-column matrix;
[0022] Wherein, each column vector in the second sub-matrix represents a residual vector between the Laplacian matrix and the first sub-matrix.
[0023] According to some exemplary embodiments, the using an iterative method to determine a search direction such that the search direction gradually coincides with the direction of each column vector in the to-be-determined feature matrix further specifically includes:
[0024] Generate a third sub - matrix based on the eigenvectors of the intermediate matrix, the first sub - matrix, and the second sub - matrix, where the third sub - matrix is an n - row and b - column matrix;
[0025] Among them, the first sub - matrix, the second sub - matrix, and the third sub - matrix constitute a search matrix representing a search subspace.
[0026] According to some exemplary embodiments, the iterative method for determining the search direction such that the search direction gradually aligns with the direction of each column vector in the to - be - determined eigen - matrix further specifically includes:
[0027] Update the first sub - matrix using an iterative method according to the eigenvectors of the intermediate matrix and the search matrix.
[0028] Among them, during the iteration process, update the first sub - matrix according to the eigenvectors of the intermediate matrix in the previous iteration process and the search matrix in the previous iteration process to generate the first sub - matrix in the current iteration process.
[0029] According to some exemplary embodiments, the iterative method for determining the search direction such that the search direction gradually aligns with the direction of each column vector in the to - be - determined eigen - matrix further specifically includes:
[0030] During the iteration process, update the second sub - matrix according to the eigenvalues and eigenvectors of the intermediate matrix in the current iteration process and the first sub - matrix in the current iteration process to generate the second sub - matrix in the current iteration process.
[0031] According to some exemplary embodiments, each column vector in the third sub - matrix represents the difference between the bases of the search subspace in two adjacent iteration processes.
[0032] According to some exemplary embodiments, the iterative method for determining the search direction such that the search direction gradually aligns with the direction of each column vector in the to - be - determined eigen - matrix further specifically includes:
[0033] During the iteration process, update the third sub - matrix according to the eigenvectors of the intermediate matrix in the previous iteration process and the first and second sub - matrices in the previous iteration process to generate the third sub - matrix in the current iteration process.
[0034] According to some exemplary embodiments, the determining of the eigen - matrix corresponding to the Laplacian matrix according to the determined search direction specifically includes:
[0035] During the iteration process, when the vector in the i-th column of the updated second sub-matrix satisfies the first specified condition, the vector in the i-th column of the updated first sub-matrix is determined as a column of the eigenmatrix, where i is a positive integer and 1 ≤ i < b.
[0036] According to some exemplary embodiments, the vector in the i-th column of the updated second sub-matrix satisfying the first specified condition includes:
[0037] The norm of the vector in the i-th column of the updated second sub-matrix is less than a specified threshold.
[0038] According to some exemplary embodiments, obtaining the intermediate matrix based on the Laplacian matrix specifically includes:
[0039] Generating the intermediate matrix based on the Laplacian matrix, the search matrix, and the transpose matrix of the search matrix.
[0040] According to some exemplary embodiments, the first sub-matrix first used in the iteration process is a random matrix with n rows and b columns.
[0041] According to some exemplary embodiments, the method further includes: when at least one information item of at least one customer among the n customers changes, updating the eigenmatrix; and / or,
[0042] After obtaining the personal information of the (n + 1)-th customer, updating the eigenmatrix.
[0043] According to some exemplary embodiments, in the process of updating the eigenmatrix, the first sub-matrix first used in the iteration process is the eigenmatrix before update.
[0044] A second aspect of the present disclosure provides a credit risk prediction device, including:
[0045] A customer authorization acquisition module, configured to acquire the customer's authorization for acquiring personal information;
[0046] A personal information acquisition module, configured to: when obtaining the customer's authorization for acquiring personal information, acquire the personal information of n customers, where the personal information of each customer includes N information items, and the N information items are all related to credit risk, n is an integer greater than or equal to 1, and N is an integer greater than or equal to 2;
[0047] A personal information matrix acquisition module, configured to: quantify the personal information of the n customers to obtain a personal information matrix, where the personal information matrix is a matrix with n rows and N columns, and each row of the personal information matrix represents the quantified personal information of a customer;
[0048] The Laplacian matrix calculation module is configured to: for the personal information matrix, calculate the Laplacian matrix corresponding to the personal information matrix by using the spectral clustering method, where the Laplacian matrix is an n×n matrix;
[0049] The eigenmatrix acquisition module is configured to: perform dimensionality reduction on the Laplacian matrix by using the locally optimal block conjugate gradient method to obtain an eigenmatrix corresponding to the Laplacian matrix, where the eigenmatrix is an n×b matrix, b is a positive integer and 1≤b<n;
[0050] The classification module is configured to: classify the n customers based on the eigenmatrix by using a clustering method; and
[0051] The credit risk prediction module is configured to: predict the credit risks of the n customers according to the n classified customers.
[0052] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above method.
[0053] A fourth aspect of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above method.
[0054] A fifth aspect of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features and advantages of the present disclosure will become clearer. In the drawings:
[0056] Figure 1 A schematic application scenario diagram of the credit risk prediction method according to an embodiment of the present disclosure is shown;
[0057] Figure 2 A schematic flowchart of the credit risk prediction method according to an embodiment of the present disclosure is shown;
[0058] Figure 3 A schematic flowchart of dimensionality reduction of the Laplacian matrix by using the locally optimal block conjugate gradient method according to an embodiment of the present disclosure is shown;
[0059] Figure 4 A schematic block diagram of the credit risk prediction device according to an embodiment of the present disclosure is shown; and
[0060] Figure 5 A block diagram of an electronic device suitable for implementing a method for predicting credit risk according to an embodiment of the present disclosure is schematically shown. Detailed implementation manners
[0061] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0062] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0063] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0064] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0065] Currently, major banks usually predict the credit risk of customers by manual evaluation or machine prediction. Manual classification is greatly affected by subjective factors and it is difficult to form an objective evaluation standard. For classifying and predicting using machine learning, usually, the information of each customer is used as a sample, and then algorithms such as SVM (Support Vector Machine), neural network, or decision tree are used to obtain the credit risk level of each customer.
[0066] However, for banks with customer data of more than tens of millions, when adding each new customer or a new information item, the machine needs to recalculate and classify all the data, which will generate a huge amount of computation, consume time and resources, and may not be able to respond to the changing data in a timely manner.
[0067] Since manual judgment is relatively subjective, the prediction method provided in this application is improved based on machine learning. Among the multiple information items of customers, there are a lot of irrelevant or unimportant information, which interfere with the evaluation results of credit risk and will also increase the unnecessary computational burden. This application effectively avoids this information to quickly find useful information items, and then classifies the credit of financial customers or enterprises. Different from dimensionality reduction in the prior art, this application uses the locally optimal block conjugate gradient method to reduce the dimensionality of the Laplacian matrix, which can quickly approximate the eigen-space of the Laplacian matrix. Therefore, the prediction method of this application has the advantages of less memory occupation and fast calculation.
[0068] To facilitate the understanding of the technical solution of this application, the following will introduce the technical terms involved in this application.
[0069] Spectral clustering algorithm: An unsupervised machine learning method, the spectral clustering algorithm is based on the spectral graph theory in graph theory. Its essence is to transform the clustering problem into the optimal partitioning problem of a graph. Compared with traditional clustering algorithms, it has the advantages of being able to cluster in the sample space of any shape and converging to the global optimal solution.
[0070] Dimensionality reduction: A key step in the spectral clustering algorithm. In this application, it can reduce the input data to reduce the amount of calculation. For example, the amount of data drops from n×n to n×b, where b must be less than n.
[0071] Eigenvector and eigenvalue: The eigenvector of a matrix is one of the important concepts in matrix theory. The eigenvector of a linear transformation is a non-degenerate vector, and its direction remains unchanged under this transformation. The ratio by which this vector is scaled under this transformation becomes its eigenvalue. Mathematically, if the vector v and the transformation A satisfy Av = λv, then the vector v is an eigenvector of the transformation A, and λ is the corresponding eigenvalue.
[0072] Eigen-space: The space where the eigenvector is located, and each eigenvector corresponds to a unique coordinate in the eigen-space.
[0073] The following will describe the embodiments of this application with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary and not intended to limit the scope of the disclosure of this application. In the following detailed description, for the sake of explanation, many specific details are elaborated and a comprehensive explanation of the embodiments of this application is provided. However, one or more embodiments can also be implemented without these specific details. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion.
[0074] It should be noted that in the technical solution of this application, the acquisition, storage, and application of the customer's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.
[0075] Figure 1 FIG. schematically shows an application scenario diagram of a method for predicting credit risk according to an embodiment of the present disclosure.
[0076] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0077] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0078] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0079] The server 105 may be a server providing various services, such as a background management server that supports the websites browsed by users using the terminal devices 101, 102, 103 (only as an example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0080] It should be noted that the method for predicting credit risk provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the device for predicting credit risk provided by the embodiments of the present disclosure can generally be set in the server 105. The method for predicting credit risk provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the device for predicting credit risk provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0081] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in [[ ]] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0082] Figure 2 A flowchart of a method for predicting credit risk according to an embodiment of the present disclosure is schematically shown.
[0083] As Figure 2 shown, the prediction method of this embodiment includes operations S110 to S170.
[0084] In operation S110, obtain the authorization of the customer for obtaining personal information.
[0085] In an embodiment of the present disclosure, before obtaining the information of the customer, the consent or authorization of the customer needs to be obtained. For example, before operation S120, the customer can be notified and a request for obtaining the customer's personal information and other associated information can be sent to the customer. When the customer consents or authorizes the acquisition of personal information, operation S120 is executed.
[0086] In operation S120, when the authorization of the customer for obtaining personal information is obtained, obtain the personal information of n customers, where the personal information of each customer includes N information items, all of the N information items are related to credit risk, n is an integer greater than or equal to 1, and N is an integer greater than or equal to 2.
[0087] Before predicting the credit risk of a customer, the personal information related to the customer, that is, the so-called customer data, needs to be obtained. From the customer data, information such as the customer's business operation ability, profitability, debt repayment ability, development ability, customer quality, and credit status can usually be learned, and the credit of the customer is judged based on these related information.
[0088] The channels for obtaining the personal information of the customer can be diversified. For example, in a specific example, it may include the following text segments: "A certain company, established in 2015, turnover, earnings per share, net profit growth rate, performance fulfillment, ……". Relevant information on the institution attributes and operation conditions can be learned from the above data information. In some instances, the above data information can be further expanded. For example, the information of the institution can include: supplementary information of enterprise registration, information on enterprise registration changes, shareholder registration information, unit insurance participation information, enterprise legal person, enterprise financial data, enterprise tax information, enterprise tax registration information, enterprise provident fund payment information, etc.; other information such as: institution name, enterprise operating income, enterprise basic information, etc.
[0089] After obtaining the personal information of customers, the personal information is sorted into N information items, that is, each customer corresponds to N information items. However, among the N information items, there may be some information items that are related. For example, for a company or an individual with high income, the tax will increase accordingly. The two information items of income information and tax amount must be directly proportional. Therefore, the related information items can be merged to achieve the purpose of dimensionality reduction. This application involves dimensionality reduction. Therefore, in this operation, the obtained information item N is an integer greater than or equal to 2.
[0090] It can be understood that if multiple information items are inversely proportional, the same dimensionality reduction processing by merging can also be performed.
[0091] In operation S130, the personal information of n customers is quantified to obtain a personal information matrix. Among them, the personal information matrix is a matrix with n rows and N columns, and each row of the personal information matrix represents the quantified personal information of a customer.
[0092] Before dimensionality reduction, mathematical information processing is performed on the personal information of n customers.
[0093] The specific processing method can be adjusted according to the specific content in the information item. For example, for profitability (net profit, gross profit margin, etc.), it is represented by simple numbers. Another example is that the performance fulfillment situation is represented by the number of defaults, or by the ratio of the number of defaults to the total number of credit times.
[0094] The quantified customer data can be represented by the element X ij where i represents the i-th customer, and j represents the j-th information item of the i-th customer. Therefore, both i and j can be represented as nodes in an undirected graph.
[0095] In operation S140, for the personal information matrix, the spectral clustering method is used to calculate the graph Laplacian matrix corresponding to the personal information matrix. Among them, the graph Laplacian matrix is a matrix with n rows and n columns.
[0096] When calculating the graph Laplacian, it is necessary to first generate an adjacency matrix W using the quantified personal information. The adjacency matrix is a matrix representation of a graph. With the help of it, the structure of the graph can be conveniently stored, and the problems of the graph can be studied by means of linear algebra. In this application, it is for the solution of n customers. Therefore, the adjacency matrix W is an n*n matrix, that is, the adjacency matrix W is a matrix with an n*n dimension, where the matrix element is X ij obtained in operation S130. The element represents the weight of the edge (i, j). If there is no edge connection between two nodes, the corresponding element in the adjacency matrix is 0. Specifically, through the k-nearest neighbor method, the KNN algorithm is used to traverse all sample points, and only the k nearest points of each sample are retained as the nearest neighbors, that is, only the W ij> 0, and the rest are all set to 0.
[0097] The calculation formula is as follows:
[0098]
[0099]
[0100] In the above formula, X i is the data of the i-th row, that is, the personal information of the i-th customer, and X j is the data of the j-th row, that is, the personal information of the j-th customer.
[0101] It can be understood that formula (1) means that as long as X i is in the K-nearest neighbor set of X j , then W ij is retained; formula (2) means that the two points of X i and X j must be in each other's k-nearest neighbor sets to retain W ij .
[0102] For the calculation of W ij , the Euclidean distance can be used to measure the distance between any two points (X i and X j ). In the preset function process, common ones include polynomial kernel function, Gaussian kernel function, and Sigmoid kernel function.
[0103] As a specific embodiment of this application, the most commonly used Gaussian kernel function is adopted for calculation.
[0104]
[0105] After obtaining the adjacency matrix W through the above method, further calculate the diagonal matrix D, that is, the degree matrix.
[0106] Since this application belongs to an undirected graph, for an undirected graph, the weighted degree of a node is the sum of the weight values of all edges related to that node. The weighted degree of node i in the adjacency matrix W of the undirected graph is the sum of the elements in the i-th row of the adjacency matrix.
[0107] Calculate the corresponding diagonal matrix D according to the proximity matrix W, and the formula is:
[0108]
[0109] Finally, the obtained diagonal matrix D is:
[0110]
[0111] Given n nodes (n customers) in an undirected graph, with the adjacency matrix W and the diagonal matrix D, a graph Laplacian matrix can be defined based on the adjacency matrix W and the diagonal matrix D. That is, the graph Laplacian matrix A is defined as the difference between the diagonal matrix D and the adjacency matrix W:
[0112] A = D - W
[0113] In operation S150, the graph Laplacian matrix is dimensionally reduced using the locally optimal block conjugate gradient method to obtain an eigenmatrix corresponding to the graph Laplacian matrix. Here, the eigenmatrix is a matrix with n rows and b columns, where b is a positive integer and 1 ≤ b < n.
[0114] The obtained graph Laplacian matrix A is an n*n-dimensional matrix. In this operation, by using the locally optimal block conjugate gradient method to dimensionally reduce the graph Laplacian matrix, that is, changing from an n*n-dimensional graph Laplacian matrix to an n*b-dimensional eigenmatrix.
[0115] The locally optimal block conjugate gradient method can explore the optimal gradient direction, enabling rapid approximation to the eigenspace of the graph Laplacian A during machine training and use.
[0116] In operation S160, based on the eigenmatrix, a clustering method is used to classify the n customers.
[0117] For the n*b-dimensional eigenmatrix Q, each row of data is used as a sample, and the customers are classified using K-means. It should be clear that the eigenmatrix Q is a matrix with n rows and b columns. In this operation, for each of the n customers, the b-dimensional personal information after dimensional reduction corresponding to each customer is used as a sample for classification.
[0118] In the embodiments of the present disclosure, after obtaining the n*b-dimensional eigenmatrix using the dimensional reduction method described above, each row of data is used as a sample to classify the customers. The embodiments of the present disclosure are not limited to the above K-means clustering method, and training neural networks or decision trees (including classical decision tree methods and derivative methods such as random forests) can also be used to classify all customers.
[0119] It should be noted that in practical applications, according to experience, the finally obtained b columns should be the same as the number of categories after classifying the customers. Otherwise, the error will increase. That is, if the customers are to be divided into 7 categories, then there should be n*7 data in the eigenmatrix Q, that is, 7 eigenvectors of the graph Laplacian A are obtained, and the dimensional-reduced eigenmatrix Q is a matrix with n rows and 7 columns.
[0120] In operation S170, based on the classified n customers, the credit risks of the n customers are predicted.
[0121] Set each category to a different credit rating, that is, n customers can be divided into b categories, corresponding to b credit risks.
[0122] In a specific embodiment, after the customer classification is completed in operation S160, the performance of all customers is retrieved. Taking the classified groups as units, the performance of the customers in each group is summed up, and then sorted according to the obtained values. The group with the highest score, that is, the group with the best performance, can be set as level one, the second highest as level two, and so on.
[0123] According to the prediction method in the embodiments of the present disclosure, by using the locally optimal block conjugate gradient method, the best gradient direction can be quickly searched, the dimensionality of the N information items of each customer can be reduced, and then the eigen-space of the Laplacian matrix can be quickly obtained. The calculation speed is fast and the memory occupancy is small, which greatly reduces the calculation amount and training speed of the neural network.
[0124] Figure 3 Schematically shows a flowchart of reducing the dimensionality of the Laplacian matrix by using the locally optimal block conjugate gradient method according to the embodiments of the present disclosure.
[0125] As Figure 3 shown, the dimensionality reduction process of this embodiment includes operation S210 to operation S220.
[0126] In operation S210, an iterative method is used to determine the search direction, so that the search direction gradually becomes consistent with the direction of the vector of each column in the feature matrix to be determined.
[0127] For operation S210, first, based on the Laplacian matrix, an intermediate matrix is obtained, where the intermediate matrix is a matrix with n rows and b columns.
[0128] It can be understood that the Laplacian matrix A changes from an n*n dimension to an n*b dimension in the intermediate matrix B, that is, dimensionality reduction is achieved.
[0129] In the machine learning process of this step, the problem of solving partial differential equations will be involved. In actual applications, the finite element method can be used to simplify the solution process and obtain an approximate solution of the calculus equation.
[0130] Then, the eigenvalues and eigenvectors of the intermediate matrix are calculated.
[0131] Finally, according to the eigenvectors of the intermediate matrix, a first sub-matrix is generated, where each column vector in the first sub-matrix represents the search direction, and the search directions represented by each column vector in the first sub-matrix correspond to the directions of the vectors of each column in the feature matrix to be determined. The first sub-matrix is a matrix with n rows and b columns.
[0132] The algorithm used in the above solution process is the Rayleigh-Ritz method, which directly starts from the functional and finds the function that can minimize it.
[0133] Exemplarily, the application method of the Rayleigh-Ritz algorithm can be the following process:
[0134] Input: A ∈ R n*n , matrix;
[0135] S ∈ R n*b , matrix, (0 ≤ b < n);
[0136] Output: θ: a b-order diagonal matrix
[0137] Y: a b-order matrix
[0138] RR:
[0139] Calculate matrix B = S T AS
[0140] Find all eigenvectors y of B 1 , y 2 ,..., y b and all eigenvalues θ 1 , θ 2 ,..., θ b
[0141] Let Y = [y 1 y 2 …y b ,
[0142] Among them, matrix B is the intermediate matrix, the input matrix A is the graph Laplacian matrix, S is the search matrix, the output matrix θ is the eigenvalue matrix, the output Y is the eigenvector, and RR is the abbreviation of Rayleigh-Ritz, which represents the Rayleigh-Ritz algorithm.
[0143] The above algorithm can be interpreted as calculating an orthogonal basis in S ∈ R n*b to approximate the eigenspace corresponding to b eigenvectors (b is the number of categories to be obtained, which is specifically explained in operations S160 and S170), calculating the intermediate matrix B, and solving the eigenvectors Y and eigenvalues θ of the intermediate matrix B.
[0144] When calculating the intermediate matrix B, based on the graph Laplacian matrix, the search matrix, and the transpose matrix of the search matrix, generate the intermediate matrix, that is, B = S T AS.
[0145] The following will use the eigenvalues θ and eigenvectors Y obtained by the Rayleigh-Ritz algorithm to reduce the dimensionality of the graph Laplacian matrix.
[0146] Generate a second submatrix based on the eigenvalues and eigenvectors of the intermediate matrix and the first submatrix. The second submatrix is an n×b matrix. Among them, the vector of each column in the second submatrix represents the residual vector between the graph Laplacian matrix and the first submatrix.
[0147] Solve the second submatrix R according to the intermediate matrix B, eigenvector Y, and eigenvalue θ 0 , and the residual vector characterizes the desired accuracy.
[0148] In this process, Rayleigh-Ritz can be used again, using
[0149] RR(Y, θ) = RR(A, X 0 )
[0150] where X 0 is a random matrix, X 0 is an n×b matrix, that is, X 0 ∈R n*b , and then through
[0151] R 0 = AX 0 - X 0 θ 0
[0152] obtain the second submatrix R 0 , and the second submatrix R 0 can be understood as the direction of the search subspace residual.
[0153] It should be noted that the first submatrix used for the first time in the iterative process is an n×b random matrix, that is, the first submatrix is X 0 .
[0154] Generate a third submatrix based on the eigenvector of the intermediate matrix, the first submatrix, and the second submatrix. The third submatrix is an n×b matrix. Among them, the first submatrix, the second submatrix, and the third submatrix form a search matrix representing the search subspace.
[0155] Solve the third submatrix P according to the eigenvector Y of the intermediate matrix, the first submatrix X 0 and the second submatrix R 0 , and finally the first submatrix X 0 , the second submatrix R 0 and the third submatrix P 0 form a search matrix S representing the search subspace 0 0 , that is, it can be expressed as
[0156] S 0 = [X 0 , R 0 , P 0
[0157] According to the eigenvectors of the intermediate matrix and the search matrix, an iterative method is used to update the first sub-matrix. Among them, during the iteration process, according to the eigenvectors of the intermediate matrix in the previous iteration process and the search matrix in the previous iteration process, the first sub-matrix is updated to generate the first sub-matrix in the current iteration process.
[0158] During the iteration process, according to the eigenvalues and eigenvectors of the intermediate matrix in the current iteration process and the first sub-matrix in the current iteration process, the second sub-matrix is updated to generate the second sub-matrix in the current iteration process.
[0159] During the iteration process, according to the eigenvectors of the intermediate matrix in the previous iteration process and the first and second sub-matrices in the previous iteration process, the third sub-matrix is updated to generate the third sub-matrix in the current iteration process.
[0160] In the field of computers, k can be used to represent the number of iterations, usually noted in the subscript. During the iteration process, every time an iteration process is experienced, the eigenvectors of the intermediate matrix, the first sub-matrix, the values of the second sub-matrix, and the third sub-matrix are incremented by 1 based on the original subscript. Based on the relationship between the eigenvectors of the intermediate matrix, the first sub-matrix, the values of the second sub-matrix, and the third sub-matrix, the iteration process can be expressed as:
[0161] X k+1 = S k Y k
[0162] P k+1 = [0, R k , P k Y k
[0163] R k+1 = AX k+1 - X k+1 θ k+1
[0164] k = k + 1
[0165] Furthermore, among them, the vectors in each column of the third sub-matrix represent the difference between the bases of the search subspaces in two adjacent iteration processes.
[0166] In operation S220, according to the determined search direction, the eigenmatrix corresponding to the Laplacian matrix is determined.
[0167] For operation S220, during the iteration process, when the vector in the i-th column of the updated second sub-matrix satisfies the first specified condition, the vector in the i-th column of the updated first sub-matrix is determined as a column of the eigenmatrix, where i is a positive integer and 1 ≤ i < b.
[0168] Furthermore, the vector in the i-th column of the updated second sub-matrix satisfying the first specified condition includes: the norm of the vector in the i-th column of the updated second sub-matrix is less than a specified threshold.
[0169] For example, in the embodiments of the present disclosure, the eigenmatrix Q can be determined according to the following iteration process.
[0170] (1) Generate a random matrix X 0 ∈R n*b , (0 ≤ b < n);
[0171] (2) Use the Rayleigh-Ritz algorithm to calculate the eigenvectors and eigenvalues:
[0172] (3) Give the initial iteration values: R 0 := AX 0 - X 0 θ 0 , k := 0, Q := [], P 0 := [];
[0173] (4) When the number of columns of the eigenmatrix Q is less than b, perform the following iteration process:
[0174] Orthogonalize Q and R k orthonormally;
[0175] Let S k := [X k , R k , P k ,
[0176] X k+1 := S k Y k ;
[0177] P k+1 := [0, R k , P k Y k ;
[0178] R k+1 := AX k+1 - X k+1 θk +1 ;
[0179] k := k + 1;
[0180] If matrix R k+1 the norms of certain columns are less than a specified threshold ε, put the corresponding columns in X k+1 into the feature matrix Q; set the corresponding columns in X k+1 to random vectors, and then: X 0 := X k+1 , k := 0, and execute the above iterative process until the number of columns of the feature matrix Q is equal to b.
[0181] According to an embodiment of the present application, the prediction method further includes: when at least one information item of at least one customer among the n customers changes, updating the feature matrix; and / or, after obtaining the personal information of the (n + 1)-th customer, updating the feature matrix.
[0182] When adding a new information item, changing the field value corresponding to the information item of one of the customers, and adding a new customer, it is inevitable that the graph Laplacian matrix is updated, and thus the feature matrix is also updated.
[0183] In one embodiment, during the process of updating the feature matrix, the first sub-matrix first used in the iterative process is the feature matrix before update.
[0184] For the situation in the prior art where when adding a new information item, changing the field value corresponding to the information item of one of the customers, and adding a new customer, it is necessary to recalculate all the data.
[0185] Considering that even when adding new information items or a new customer, the various eigenvalues change little, and the newly calculated feature space must be close to the original feature space, so the new feature space can be obtained quickly. Based on the above concept, the prediction method of the present application is calculated on the original feature matrix. When the information of the customer changes, a new feature space of the graph Laplacian matrix can be established on the existing feature space of the graph Laplacian matrix, that is, directly replace X 0 = 0 in operation S150 with making X 0 = Q.
[0186] In one experiment, the calculation speed of the second calculation can be hundreds to thousands of times faster than that of the first calculation. It can be concluded that using the prediction method of the present application can reduce the machine calculation amount, thereby saving resources and reducing the calculation time.
[0187] Based on the above credit risk prediction method, the present disclosure also provides a credit risk prediction device. The following will be combined with Figure 4 to describe this device in detail. Figure 4A structural block diagram of a prediction device according to an embodiment of the present disclosure is schematically shown.
[0188] As Figure 4 shown, the prediction device 800 of this embodiment includes a customer authorization acquisition module 810, a personal information acquisition module 820, a personal information matrix acquisition module 830, a graph Laplacian matrix calculation module 840, a feature matrix acquisition module 850, a classification module 860, and a credit risk prediction module 870.
[0189] The customer authorization acquisition module 810 is configured to acquire the customer's authorization for acquiring personal information. In one embodiment, the customer authorization acquisition module 810 may be used to perform the operation S110 described above, which will not be elaborated here.
[0190] The personal information acquisition module 820 is configured to: when the customer's authorization for acquiring personal information is obtained, acquire the personal information of n customers, where the personal information of each customer includes N information items, all of the N information items are related to credit risk, n is an integer greater than or equal to 1, and N is an integer greater than or equal to 2. In one embodiment, the personal information acquisition module 820 may be used to perform the operation S120 described above, which will not be elaborated here.
[0191] The personal information matrix acquisition module 830 is configured to: quantify the personal information of n customers to obtain a personal information matrix, where the personal information matrix is an n×N matrix, and each row of the personal information matrix represents the quantified personal information of a customer. In one embodiment, the personal information matrix acquisition module may be used to perform the operation S130 described above, which will not be elaborated here.
[0192] The graph Laplacian matrix calculation module 840 is configured to: for the personal information matrix, calculate the graph Laplacian matrix corresponding to the personal information matrix by using the spectral clustering method, where the graph Laplacian matrix is an n×n matrix. In one embodiment, the graph Laplacian matrix calculation module may be used to perform the operation S140 described above, which will not be elaborated here.
[0193] The feature matrix acquisition module 850 is configured to: perform dimensionality reduction on the graph Laplacian matrix by using the locally optimal block conjugate gradient method to obtain a feature matrix corresponding to the graph Laplacian matrix, where the feature matrix is an n×b matrix, b is a positive integer and 1≤b<n. In one embodiment, the feature matrix acquisition module may be used to perform the operation S150 described above, which will not be elaborated here.
[0194] The classification module 860 is configured to: classify n customers based on the feature matrix by using a clustering method. In one embodiment, the classification module may be used to perform the operation S160 described above, which will not be elaborated here.
[0195] The credit risk prediction module 870 is configured to: predict the credit risks of n customers based on the n classified customers. In one embodiment, the credit risk prediction module 830 can be used to perform the operation S170 described above, which will not be elaborated here.
[0196] According to the prediction device in the embodiments of the present disclosure, the above prediction method can be executed. By using the locally optimal block conjugate gradient method, the best gradient direction can be quickly searched, the dimensionality of N information items of each customer can be reduced, and then the eigen-space of the graph Laplacian matrix can be obtained rapidly. The calculation speed is fast and the memory occupancy is small, which greatly reduces the computational amount and training speed of the neural network.
[0197] According to the embodiments of the present disclosure, any multiple of the customer authorization acquisition module 810, the personal information acquisition module 820, the personal information matrix acquisition module 830, the graph Laplacian matrix calculation module 840, the eigen-matrix acquisition module 850, the classification module 860, and the credit risk prediction module 870 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to the embodiments of the present disclosure, at least one of the customer authorization acquisition module 810, the personal information acquisition module 820, the personal information matrix acquisition module 830, the graph Laplacian matrix calculation module 840, the eigen-matrix acquisition module 850, the classification module 860, and the credit risk prediction module 870 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in a suitable combination of any several of them. Or, at least one of the customer authorization acquisition module 810, the personal information acquisition module 820, the personal information matrix acquisition module 830, the graph Laplacian matrix calculation module 840, the eigen-matrix acquisition module 850, the classification module 860, and the credit risk prediction module 870 can be at least partially implemented as a computer program module, and when the computer program module runs, it can execute the corresponding functions.
[0198] Figure 5 A block diagram of an electronic device suitable for implementing the prediction method of credit risk according to an embodiment of the present disclosure is schematically shown.
[0199] As Figure 5As shown, an electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), etc. The processor 901 may also include on-board memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of a method flow according to an embodiment of the present disclosure.
[0200] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of a method flow according to an embodiment of the present disclosure by executing the program in the ROM 902 and / or the RAM 903. It should be noted that the program may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also perform various operations of a method flow according to an embodiment of the present disclosure by executing the program stored in the one or more memories.
[0201] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, and the input / output (I / O) interface 905 is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read therefrom can be installed into the storage section 908 as needed.
[0202] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, a method according to an embodiment of the present disclosure is implemented.
[0203] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903.
[0204] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the credit risk prediction method provided by the embodiment of the present disclosure.
[0205] When the computer program is executed by the processor 901, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0206] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 909, and / or be installed from the removable medium 911. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0207] In such an embodiment, the computer program may be downloaded and installed from the network through the communication part 909, and / or be installed from the removable medium 911. When the computer program is executed by the processor 901, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0208] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the client computing device, partially on the client device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the client computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0209] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0210] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0211] The embodiments of the present disclosure have been described above. However, these embodiments are merely for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A method for predicting credit risk, characterized in that, it includes the following steps: Obtain the customer's authorization for obtaining personal information; When obtaining the customer's authorization for obtaining personal information, obtain the personal information of n customers, where the personal information of each customer includes N information items, and the N information items are all related to credit risk, n is an integer greater than or equal to 1, and N is an integer greater than or equal to 2; Quantify the personal information of the n customers to obtain a personal information matrix, where the personal information matrix is a matrix with n rows and N columns, and each row of the personal information matrix represents the quantified personal information of a customer; For the personal information matrix, use the spectral clustering method to calculate the graph Laplacian matrix corresponding to the personal information matrix, where the graph Laplacian matrix is a matrix with n rows and n columns; Use the locally optimal block conjugate gradient method to reduce the dimension of the graph Laplacian matrix to obtain a characteristic matrix corresponding to the graph Laplacian matrix, where the characteristic matrix is a matrix with n rows and b columns, b is a positive integer and 1 ≤ b < n, and the locally optimal block conjugate gradient method includes: Use an iterative method to determine the search direction so that the search direction gradually aligns with the direction of each column vector in the to-be-determined characteristic matrix, and determine the characteristic matrix corresponding to the graph Laplacian matrix based on the search direction; the determination of the search direction includes: obtaining an intermediate matrix based on the graph Laplacian matrix; calculating the eigenvalues and eigenvectors of the intermediate matrix; and generating a first sub-matrix according to the eigenvectors of the intermediate matrix, where each column vector in the first sub-matrix represents the search direction; Based on the characteristic matrix, use a clustering method to classify the n customers; and Predict the credit risk of the n customers according to the classified n customers.
2. The method according to claim 1, characterized in that, The use of an iterative method to determine the search direction so that the search direction gradually aligns with the direction of each column vector in the to-be-determined characteristic matrix further specifically includes: Generate a second sub-matrix according to the eigenvalues and eigenvectors of the intermediate matrix and the first sub-matrix, the intermediate matrix is a matrix with n rows and b columns, the search direction represented by each column vector in the first sub-matrix corresponds to the direction of each column vector in the to-be-determined characteristic matrix, the first sub-matrix is a matrix with n rows and b columns, and the second sub-matrix is a matrix with n rows and b columns; Wherein, each column vector in the second sub-matrix represents the residual vector between the graph Laplacian matrix and the first sub-matrix.
3. The method according to claim 2, characterized in that, The use of an iterative method to determine the search direction so that the search direction gradually aligns with the direction of each column vector in the to-be-determined characteristic matrix further specifically includes: Generate a third sub-matrix according to the eigenvectors of the intermediate matrix, the first sub-matrix, and the second sub-matrix, the third sub-matrix is a matrix with n rows and b columns; Wherein, the first sub-matrix, the second sub-matrix, and the third sub-matrix form a search matrix representing the search subspace.
4. According to the method described in claim 3, wherein, the iterative method is used to determine the search direction such that the search direction gradually aligns with the direction of the vector in each column of the feature matrix to be determined, and specifically further includes: updating the first sub-matrix by using an iterative method according to the eigenvector of the intermediate matrix and the search matrix, wherein, during the iteration process, the first sub-matrix is updated according to the eigenvector of the intermediate matrix in the previous iteration process and the search matrix in the previous iteration process to generate the first sub-matrix in the current iteration process.
5. According to the method described in claim 4, wherein, the iterative method is used to determine the search direction such that the search direction gradually aligns with the direction of the vector in each column of the feature matrix to be determined, and specifically further includes: during the iteration process, the second sub-matrix is updated according to the eigenvalue and eigenvector of the intermediate matrix in the current iteration process and the first sub-matrix in the current iteration process to generate the second sub-matrix in the current iteration process.
6. According to the method described in claim 5, wherein, the vector in each column of the third sub-matrix represents the difference between the bases of the search sub-spaces in two adjacent iteration processes.
7. According to the method described in claim 6, wherein, the iterative method is used to determine the search direction such that the search direction gradually aligns with the direction of the vector in each column of the feature matrix to be determined, and specifically further includes: during the iteration process, the third sub-matrix is updated according to the eigenvector of the intermediate matrix in the previous iteration process and the first and second sub-matrices in the previous iteration process to generate the third sub-matrix in the current iteration process.
8. According to the method described in claim 7, wherein, determining the feature matrix corresponding to the Laplacian matrix according to the determined search direction specifically includes: during the iteration process, when the vector in the i-th column of the updated second sub-matrix satisfies the first specified condition, the vector in the i-th column of the updated first sub-matrix is determined as a column of the feature matrix, where i is a positive integer and 1 ≤ i < b.
9. According to the method described in claim 8, wherein, the vector in the i-th column of the updated second sub-matrix satisfying the first specified condition includes: the norm of the vector in the i-th column of the updated second sub-matrix is less than the specified threshold.
10. According to the method described in claim 1, wherein, obtaining the intermediate matrix based on the Laplacian matrix specifically includes: generating the intermediate matrix based on the Laplacian matrix, the search matrix, and the transpose matrix of the search matrix.
11. According to the method described in claim 4, wherein, the first sub-matrix first used in the iteration process is a random matrix with n rows and b columns.
12. According to the method described in claim 11, wherein, the method further includes: when at least one information item of at least one of the n customers changes, updating the feature matrix; and / or, After obtaining the personal information of the (n + 1)-th customer, update the feature matrix.
13. The method according to claim 12, wherein, in the process of updating the feature matrix, the first sub-matrix used for the first time in the iterative process is the feature matrix before updating.
14. A prediction device for credit risk, wherein, comprising: a customer authorization acquisition module for acquiring the customer's authorization for obtaining personal information; a personal information acquisition module for: when obtaining the customer's authorization for obtaining personal information, acquiring the personal information of n customers, wherein the personal information of each customer includes N information items, and the N information items are all related to credit risk, n is an integer greater than or equal to 1, and N is an integer greater than or equal to 2; a personal information matrix acquisition module for: quantifying the personal information of the n customers to obtain a personal information matrix, wherein the personal information matrix is a matrix with n rows and N columns, and each row of the personal information matrix represents the quantified personal information of a customer; a graph Laplacian matrix calculation module for: for the personal information matrix, calculating the graph Laplacian matrix corresponding to the personal information matrix by using the spectral clustering method, wherein the graph Laplacian matrix is a matrix with n rows and n columns; a feature matrix acquisition module for: performing dimensionality reduction on the graph Laplacian matrix by using the locally optimal block conjugate gradient method to obtain the feature matrix corresponding to the graph Laplacian matrix, wherein the feature matrix is a matrix with n rows and b columns, b is a positive integer and 1 ≤ b < n, and the locally optimal block conjugate gradient method includes: determining the search direction by using an iterative method such that the search direction gradually coincides with the directions of the column vectors in the feature matrix to be determined, and determining the feature matrix corresponding to the graph Laplacian matrix based on the search direction; the determining of the search direction includes: obtaining an intermediate matrix based on the graph Laplacian matrix; calculating the eigenvalues and eigenvectors of the intermediate matrix; and generating a first sub-matrix according to the eigenvectors of the intermediate matrix, and each column vector in the first sub-matrix represents the search direction; a classification module for: classifying the n customers by using a clustering method based on the feature matrix; and a credit risk prediction module for: predicting the credit risks of the n customers according to the n classified customers.
15. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 13.
16. A computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 13.
17. A computer program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Method for Anomaly Detection in Time Series Data Based on Spectral Partitioning
US20150363699A1
Searching multidimensional indexes using associated clustering and dimension reduction information
US6134541A