Credit risk determination method and apparatus, and computer program product

By selecting the training data of the target client collection in the federated learning model, the problems of low communication efficiency and high evaluation cost due to data heterogeneity are solved, and efficient and secure credit risk assessment is achieved.

CN120374250APending Publication Date: 2025-07-25CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510443571.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The heterogeneity of the training data of the federated learning model leads to inefficient communication, which in turn increases the cost of credit risk assessment.

Method used

By obtaining the data to be analyzed by the target user, extracting the area to which it belongs, determining the client corresponding to the area, using the training data of the target client collection to train the federated learning model, outputting credit risk values, avoiding the heterogeneity of client data and reducing communication overhead.

Benefits of technology

It reduces the cost of credit risk assessment, improves the efficiency and security of credit risk assessment, and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374250A_ABST
    Figure CN120374250A_ABST
Patent Text Reader

Abstract

The invention discloses a credit risk determination method and apparatus, and a computer program product. The method comprises the steps of obtaining to-be-analyzed data of a target user; a target area to which the target user belongs is extracted from the identity information, a client corresponding to the target area in a federated learning model is determined, the federated learning model comprises a global model and local models of multiple clients, and the federated learning model is obtained by training of training data of a target client set; and inputting the to-be-analyzed data into a target local model of the client, and outputting a credit risk value of the target user, the target local model being obtained by training the model parameters of the global model and the training data of the client. Through the credit risk assessment method and device, the problem that in the related technology, due to heterogeneity of training data of a federated learning model, the communication efficiency of the federated learning model is low, and then the cost is high when credit risk assessment is conducted on the user is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular, to a method, apparatus, and computer program product for determining credit risk. Background Art

[0002] In related technologies, the centralized machine learning methods used for credit risk assessment include federated learning methods. As a branch of distributed machine learning, federated learning breaks through the limitations of centralized machine learning on data. Federated learning trains models locally on clients. While protecting the privacy of clients, each client collaborates to complete the training of the global model, effectively solving various problems faced by centralized machine learning, such as insufficient network bandwidth, insufficient computing resources, and data security threats. All the calculation parameters and results in the federated learning model are placed on the central server, and the client only needs to upload local data for global model update. However, this calculation method has high requirements for network bandwidth and machine computing power.

[0003] The data types, quantities, and distributions owned by different clients vary. This data difference will bring various challenges to federated learning. First, non-independent and identically distributed data may lead to a decline in the generalization ability of the model. The model cannot cover all data situations during training, resulting in a decline in performance when facing new data. Second, due to the inconsistent data distributions of the participants, it may affect the convergence speed of the model. Different participants may require different numbers of iterations during local training to reach the same convergence degree, which may lead to uncertainty in training time and an increase in computing costs.

[0004] In addition, non-independent and identically distributed data also brings difficulties to the update of the global model. In federated learning, the global model is updated by aggregating the local model parameters of each participant. However, when the data distributions of the participants are different, the fusion of local model parameters will face challenges, which may lead to a decline in the performance of the global model, thus affecting the overall training effect. Finally, non-independent and identically distributed data also increases the risk of privacy leakage. By observing the update situation of the global model, it may be possible to infer the data feature distribution of the participants, thus increasing the risk of privacy leakage.

[0005] Due to the heterogeneity of devices and differences in usage habits, there are differences in the data distribution of different clients, and this difference is dynamically changing. In related technologies, a two-sampling process is adopted to address the challenge of non-independent and identically distributed data in federated learning. In the first sampling process, clients are selected as much as possible such that they all follow a certain same distribution. In the second sampling process, model training is performed among the selected client combinations to ensure that the trained local models can accurately predict label classifications based on the client data characteristics, so as to conform to the joint probability distribution. Since the non-independent and identically distributed data and the proportion of the selected clients in the overall population are insufficient, and after two samplings, the feature space of the samples changes, that is, the feature space is different in different iteration rounds, which will cause the convergence of the global model to oscillate and may even converge, and both the convergence speed and the model accuracy will suffer unpredictable losses.

[0006] Aiming at the problem in related technologies that due to the heterogeneity of the training data of the federated learning model, the communication efficiency of the federated learning model is low, and further the cost of credit risk assessment for users is high, no effective solution has been proposed yet. Summary of the Invention

[0007] The main purpose of this application is to provide a method, device and computer program product for determining credit risk, so as to solve the problem in related technologies that due to the heterogeneity of the training data of the federated learning model, the communication efficiency of the federated learning model is low, and further the cost of credit risk assessment for users is high.

[0008] To achieve the above objective, according to one aspect of this application, a method for determining credit risk is provided. The method includes: obtaining the data to be analyzed of the target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; extracting the target region to which the target user belongs from the identity information, and determining the client corresponding to the target region in the federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by the training data of the target client set; inputting the data to be analyzed into the target local model of the client, and outputting the credit risk value of the target user, where the target local model is trained by the model parameters of the global model and the training data of the client.

[0009] Optionally, the federated learning model is trained as follows: For each iterative training in the federated learning model, obtain the training data of each client, perturb the training data through a random perturbation algorithm to obtain the perturbation vector of the client, where the training data includes the historical data to be analyzed and the historical credit risk values of the users of the client; generate a perturbation matrix based on the perturbation vectors of all clients, perform singular value decomposition on the perturbation matrix to obtain the target eigenvector and the target rank, where the number of rows of the perturbation matrix represents the number of clients, and the number of columns of the perturbation matrix represents the number of data categories of the training data of all clients; screen out multiple groups of target perturbation vectors from the perturbation matrix based on the target eigenvector and the target rank, determine the client set corresponding to each group of target perturbation vectors to obtain multiple groups of client sets; for each group of client sets, calculate the evaluation value based on the model parameters of the local model of each client in the client set and the model parameters of the global model to obtain the value evaluation value of the client set; determine the client set corresponding to the maximum value evaluation value among the multiple groups of client sets as the target client set, and train the federated learning model based on the target client set.

[0010] Optionally, screening out multiple groups of target perturbation vectors from the perturbation matrix based on the target eigenvector and the target rank includes: extracting a reference vector group from the perturbation matrix, and determining the vectors to be added as the vectors in the perturbation matrix except the reference vector group; for each target vector to be added, traverse one by one the other vectors to be added in the perturbation matrix except the target vector to be added, and calculate the rank of the matrix formed by each traversed vector to be added and the target vector to be added; when the rank is the target value, add the traversed vector to be added to the reference vector group until the rank of the matrix formed by the reference vector group is equal to the number of data categories, or when all the vectors to be added in the perturbation matrix are traversed, end the traversal; form a group of perturbation vectors by each target vector to be added and the reference vector group corresponding to the target vector to be added to obtain multiple groups of perturbation vectors; verify each group of perturbation vectors based on the target eigenvalue and the target rank, and determine each group of perturbation vectors that pass the verification as a group of target perturbation vectors.

[0011] Optionally, extracting a reference vector group from the perturbation matrix includes: for each perturbation vector in the perturbation matrix, calculate the norm of the perturbation vector; when the norm of the perturbation vector is greater than the preset value, determine the perturbation vector as a reference vector; form a reference vector group by all the reference vectors.

[0012] Optionally, verify each group of perturbation vectors based on the target eigenvalue and the target rank, and determine the groups of perturbation vectors that pass the verification as a group of target perturbation vectors, including: calculating the rank and eigenvalue of the matrix formed by each group of perturbation vectors, determining whether the rank is the same as the target rank, and determining whether the eigenvalue is the same as the target eigenvalue; removing the group of perturbation vectors when the rank is different from the target rank or the eigenvalue is different from the target eigenvalue; and determining the group of perturbation vectors as a group of target perturbation vectors when the rank is the same as the target rank and the eigenvalue is the same as the target eigenvalue.

[0013] Optionally, calculate the evaluation value based on the model parameters of the local models of each client in the client set and the model parameters of the global model, and obtain the value evaluation value of the client set, including: for each client in the client set, calculate the similarity between the model parameters of the local model of the client and the model parameters of the global model to obtain the similarity evaluation value of the client; determine the actual probability distribution and the target probability distribution of the training data of the client, and calculate the relative entropy based on the actual probability distribution and the target probability distribution to obtain the balance evaluation value of the client; calculate the sum of the reciprocal of the similarity evaluation value, the balance evaluation value, and the adjustment parameter, and calculate the reciprocal of the sum to obtain the sub-value evaluation value of the client; calculate the sum of the sub-value evaluation values of all clients in the client set to obtain the value evaluation value of the client set.

[0014] Optionally, calculate the similarity between the model parameters of the local model of the client and the model parameters of the global model to obtain the similarity evaluation value of the client, including: determining the first model parameters of the global model in the previous iteration round of the federated learning model; training the local model based on the first model parameters and the training data of the client to obtain the model parameters of the local model; obtaining the second model parameters of the global model in the current iteration round of the federated learning model, and calculating the cosine similarity between the second model parameters and the model parameters of the local model to obtain the similarity evaluation value of the client.

[0015] Optionally, train the federated learning model based on the target client set, including: determining the first model parameters of the global model in the previous iteration round of the federated learning model, training the local model based on the first model parameters and the training data of each client in the target client set to obtain the local model parameters of each client; aggregating the local model parameters of all clients in the target client set to obtain the model parameters of the global model in the current iteration round of the federated learning model.

[0016] To achieve the above object, according to another aspect of the present application, there is provided an apparatus for determining credit risk. The apparatus includes: an acquisition unit configured to acquire data to be analyzed of a target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; an extraction unit configured to extract a target region to which the target user belongs from the identity information, and determine a client corresponding to the target region in a federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by training data of a target client set; and an input unit configured to input the data to be analyzed into a target local model of the client and output a credit risk value of the target user, where the target local model is trained by model parameters of the global model and training data of the client.

[0017] To achieve the above object, according to another aspect of the present application, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the method for determining credit risk in various embodiments of the present application are implemented.

[0018] In the present application, the following steps are adopted: acquiring data to be analyzed of a target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; extracting a target region to which the target user belongs from the identity information, and determining a client corresponding to the target region in a federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by training data of a target client set; inputting the data to be analyzed into a target local model of the client and outputting a credit risk value of the target user, where the target local model is trained by model parameters of the global model and training data of the client. This solves the problem in the related art that due to the heterogeneity of the training data of the federated learning model, the communication efficiency of the federated learning model is low, and further the cost of credit risk assessment for users is high. By using the training data of the target client set to train the federated learning model instead of using the training data of all clients, the problem of heterogeneity of client data is avoided, thereby reducing the communication overhead during the training of the federated learning model, and further achieving the effect of reducing the cost of credit risk assessment for users. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the accompanying drawings:

[0020] Figure 1 is a flowchart of the method for determining credit risk according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of a method for training a federated learning model provided according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of randomly perturbing client data provided according to an embodiment of the present application;

[0023] Figure 4 is a schematic diagram of a device for determining credit risk provided according to an embodiment of the present application;

[0024] Figure 5 is a schematic diagram of an electronic device provided according to an embodiment of the present application. Detailed implementation manners

[0025] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data that have been authorized by the user or fully authorized by all parties.

[0029] It should be noted that the information collected is information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant regions, necessary confidentiality measures are taken, it does not violate public order and good customs, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0030] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 It is a flowchart of a method for determining credit risk provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0031] Step S101, obtain the data to be analyzed of the target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records.

[0032] In step S101, the identity information may include the user's full name, ID number, date of birth, occupation, etc. The asset data may cover the financial resources owned by the user, such as deposits, real estate investments, vehicles, stocks, and bonds. The consumption records may include the user's purchase habits, credit card usage, loan repayment records, bill payment history, etc.

[0033] Step S102, extract the target area to which the target user belongs from the identity information, and determine the client corresponding to the target area in the federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by the training data of the target client set.

[0034] In step S102, extract data related to the area from the identity information of the target user, such as the user's residential address, postal code, IP address, etc., in order to determine the geographical area where the user is located. Map the parsed geographical location information to a predefined area code or area identifier for matching in the federated learning model. For example, map the postal code or IP address to the level of the city, province, or region.

[0035] In the federated learning system, each client may correspond to a local institution of the company in different regions. When determining the credit risk of a user, in order to protect user privacy, only the data to be analyzed of the user is retained in the local client handling the business, that is, the client corresponding to the target area.

[0036] The federated learning model is composed of local models of multiple clients and a global model. Each client has a different local dataset. To avoid the heterogeneity problem of training data of different clients, the federated learning model in this embodiment does not use the federated averaging algorithm to select the training data of clients for training. Instead, it evaluates each client based on the principle of balanced training data distribution, filters out the target client set according to the value evaluation value of the client, and trains the federated learning model with the training data of the target client set.

[0037] Step S103: Input the data to be analyzed into the target local model of the client, and output the credit risk value of the target user. The target local model is trained by the model parameters of the global model and the training data of the client.

[0038] In step S103, at the beginning stage of federated learning, the central server or coordinator uses the available initial dataset (such as randomly initialized parameters) to create a global model. This global model will serve as the starting point of the federated learning process. The central server distributes the latest parameters of the global model to the clients participating in federated learning. After receiving these parameters, the clients apply them to their local models as the initial point of training.

[0039] The target clients use the received global model parameters and the locally stored training data to train the local models. The training process can include but is not limited to forward propagation, loss calculation, backward propagation, and parameter update. After training is completed, the local models of the target client set send the parameters back to the central server without directly uploading the user data. After receiving the updated parameters of all participating clients, the central server performs an aggregation process to update the parameters of the global model. The central server aggregates the parameter updates of all clients to optimize the global model. The optimized global model parameters will be used as the starting point for the next round of federated learning and will be distributed to all clients for further training and optimization. Federated learning is an iterative process. The parameters of the global model will go through multiple rounds of local training by clients and parameter aggregation by the central server until the global model reaches the preset convergence criteria or meets specific performance metrics.

[0040] The model parameters of the trained global model are distributed to each client again to guide the completion of the training of the local models of the clients. The target local model is also the model of the client after training is completed. After inputting the data to be analyzed of the target user into the target local model, the target local model can output the credit risk value of the target user.

[0041] After the global model is trained, it will be able to perform credit risk assessment on new user data. This model can be deployed in financial institutions, credit institutions, online trading platforms, etc. to provide real-time credit assessment services while protecting user privacy and data security.

[0042] The method for determining credit risk provided by the embodiments of this application obtains the data to be analyzed of the target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; extracts the target area to which the target user belongs from the identity information, and determines the client corresponding to the target area in the federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by the training data of the target client set; inputs the data to be analyzed into the target local model of the client and outputs the credit risk value of the target user, where the target local model is trained by the model parameters of the global model and the training data of the client, which solves the problem in the related art that due to the heterogeneity of the training data of the federated learning model, the communication efficiency of the federated learning model is low, and further the cost of performing credit risk assessment on users is relatively high. By using the training data of the target client set to train the federated learning model instead of using the training data of all clients, the problem of heterogeneity of client data is avoided, thereby reducing the communication overhead during the training of the federated learning model, and further achieving the effect of reducing the cost of performing credit risk assessment on users.

[0043] To evaluate the credit risk value of a user, a federated learning model needs to be trained. Optionally, in the method for determining credit risk provided by the embodiments of this application, the federated learning model is trained in the following manner: for each iterative training in the federated learning model, obtain the training data of each client, and perturb the training data through a random perturbation algorithm to obtain the perturbation vector of the client, where the training data includes the historical data to be analyzed and the historical credit risk value of the user of the client; generate a perturbation matrix based on the perturbation vectors of all clients, perform singular value decomposition on the perturbation matrix to obtain the target eigenvector and the target rank, where the number of rows of the perturbation matrix represents the number of clients, and the number of columns of the perturbation matrix represents the number of data categories of the training data of all clients; screen out multiple groups of target perturbation vectors from the perturbation matrix based on the target eigenvector and the target rank, determine the client set corresponding to each group of target perturbation vectors to obtain multiple groups of client sets; for each group of client sets, calculate the evaluation value based on the model parameters of the local model of each client in the client set and the model parameters of the global model to obtain the value evaluation value of the client set; determine the client set corresponding to the maximum value evaluation value among the multiple groups of client sets as the target client set, and train the federated learning model based on the target client set.

[0044] In some examples, due to the poor robust performance of the original training data of the client, in order to simplify the calculation, a random perturbation algorithm is used to add noise to the training data of the client. First, a classification function is used to perform non-independent and identically distributed data partitioning on the data, and these partitioned data are used for random perturbation. The perturbation vectors formed after perturbation are uploaded to the central server. By perturbing the training data, the robustness of the client data is enhanced, effectively dealing with the data classification problem in federated learning, and at the same time simplifying the calculation process.

[0045] Since non-independent and identically distributed data will cause an offset effect on the federated learning model. In federated learning, the federated averaging algorithm adopts a strategy of randomly selecting clients. If the data distribution differences between the selected clients are small, then the convergence of the global model is not affected to a great extent. However, if in a certain round of random selection, the selected clients are all in extreme data distribution situations, then the aggregation of these local models will bring a large degree of offset to the global model. In order to reduce the adverse effects brought by this random client selection method and ensure that the clients participating in each round do not have extreme data distributions, the training method of the federated learning model in this embodiment is a client selection method based on data distribution balance and data value evaluation. This method is a two-stage client selection method. Among the client combinations with relatively balanced data distributions, the client combination with the largest value contribution degree is given priority, avoiding the influence of extreme distribution data on the aggregated model. The client data selected in each round should maximize the value evaluation value.

[0046] For example, Figure 2 is a flowchart of the training method of the federated learning model provided according to an embodiment of the present application. As Figure 2 shown, (1) Use a random perturbation algorithm to perturb the client data. After forming the perturbation vector of the client, upload the perturbation vector to the central server, and generate a perturbation matrix M in the central server.

[0047] Randomly perturb the client data. Perturbation not only reduces the calculation amount but also eliminates some extreme data. For example, Figure 3 is a schematic diagram of randomly perturbing the client data provided according to an embodiment of the present application. As Figure 3As shown in the figure, assume that the data of n clients comes from the ABC dataset. A certain client i contains 6 categories in the ABC dataset. After counting the number of each category and converting it into an array Ci: [222, 3, 0, 0, 409, 28, 1, 0, 0, 35] as the input, it is perturbed by the client data distribution to obtain Ci': [1, 0, 0, 0, 1, 1, 0, 0, 1, 1]. All client data forms a perturbation vector after being perturbed by the random response and is sent to the central server. The central server will automatically combine the perturbation vectors into an n×m matrix M according to the algorithm. Among them, n represents the number of perturbed clients, and m represents the number of categories of data held by all clients. The central server will add a unique identification ID in front of each perturbation vector.

[0048] (2) The central server performs singular value decomposition on the matrix M according to the rules to calculate the eigenvectors and rank (i.e., Figure 2 the eigenvalues in), and uses the linear independence between vectors to extract the target perturbation vectors from the perturbation matrix and then combines them to generate multiple groups of client sets Si, and verifies the similarity between the selected client data distribution and the global data distribution. Otherwise, the currently selected client set Si will not be added to the pending set S, and the KL divergence of this combination is calculated.

[0049] For example, the central server aggregates the perturbation vectors to generate the matrix M, performs singular value decomposition on the matrix M, calculates the eigenvector matrix V, the singular value matrix B, and the rank of the matrix, and then selects the client m i to add to the set S i , and verifies that when the rank of the set S i is equal to the rank of the matrix M, the selection and addition of clients will stop. Calculate the KL divergence of the vectors in the set S i in the third-party proxy device, analyze the difference between the current client combination and the global distribution, and those that meet the conditions enter the next stage.

[0050] (3) If this combination is added to the set S, calculate the cosine similarity corresponding to this combination in the second stage, and determine the combination with the largest value evaluation value among all client selection schemes in this round.

[0051] For example, in the third-party proxy device, calculate the cosine similarity between the local model generated by each client vector in the set S i and the global model of the previous round one by one, substitute the key algorithm results in steps three and four into the above formula, and select the combination with the largest value evaluation value.

[0052] (4) The central server notifies the client combination that meets the requirements to perform the next round of training according to the result calculated in step (3).

[0053] In the first stage of this embodiment, a perturbed vector formed by perturbing client data is used to screen out a combination of clients whose client data is similar to the global data distribution, that is, a combination of clients with balanced distribution; in the second stage, value evaluation calculation is performed based on KL divergence and cosine similarity, that is, on the basis of the combination of clients selected in the first stage, a combination of clients with the largest value evaluation value is screened out. While ensuring that the model converges in a reliable direction, it maximally reduces the amount of non-compliant data participating in training, and effectively reduces the communication overhead during the training process of the federated learning model.

[0054] To screen out clients with balanced training data distribution, clients corresponding to the target perturbation vector can be screened out from the perturbation matrix based on the target feature vector and the target rank. Optionally, in the method for determining credit risk provided in the embodiments of the present application, screening out multiple groups of target perturbation vectors from the perturbation matrix based on the target feature vector and the target rank includes: extracting a reference vector group from the perturbation matrix, and determining the perturbation vectors other than the reference vector group in the perturbation matrix as vectors to be added; for each target vector to be added, traversing one by one the other vectors to be added in the perturbation matrix except the target vector to be added, and calculating the rank of the matrix formed by each traversed vector to be added and the target vector to be added; when the rank is the target value, adding the traversed vector to be added to the reference vector group until the rank of the matrix formed by the reference vector group is equal to the number of data categories, or when the vectors to be added in the perturbation matrix are traversed; a group of perturbation vectors is formed by each target vector to be added and the reference vector group corresponding to the target vector to be added, and multiple groups of perturbation vectors are obtained; each group of perturbation vectors is verified based on the target eigenvalue and the target rank, and each group of perturbation vectors that passes the verification is determined as a group of target perturbation vectors.

[0055] In some examples, a reference vector group is selected from the perturbation matrix. The reference vector group can be a vector with as many 1s as possible in the perturbation vector. The more 1s there are, the more uniform the data distribution of the vector is. All vectors in the perturbation matrix, except those already selected into the reference vector group, are used as vectors to be added.

[0056] For each target vector to be added, traverse all other vectors to be added in the perturbation matrix except this vector. Calculate the rank of the matrix formed when each traversed vector to be added is combined with the target vector to be added. The purpose of rank calculation is to evaluate whether adding a new vector can increase the linearly independent dimension of the vector group, that is, whether it can bring more information to the matrix, that is, whether it can make the distribution of training data more balanced.

[0057] If the rank of the formed matrix equals the target value (e.g., 2) of the number of data categories after adding a vector to be added, it indicates that the vector to be added is linearly independent of the target vector to be added, and it can provide more data categories of training data. Then add the vector to be added to the reference vector group. The traversal process will end when any of the following conditions is met: the rank of the matrix formed by the reference vector group equals the number of data categories, in which case the reference vector group already covers all data categories and no additional new samples are needed; or all vectors to be added in the perturbation matrix have been traversed. Each target vector to be added and its corresponding reference vector group form a set of perturbation vectors. Generate multiple sets of perturbation vectors, and each set of perturbation vectors is a representation of the data distribution of a potential client set.

[0058] For example, based on the reference vector group, extract a perturbation vector m in sequence i and add it to the reference vector group. Then, perform a traversal operation on the remaining vectors to be added in matrix M according to the exhaustive search algorithm, and calculate the rank of the target vector to be added m i and the vector to be added m that has been traversed j If R(m i m j ) = 2, then add m j to S i as well. Repeat this operation until the rank of the matrix in S i equals the number of classifications m or the traversal loop ends.

[0059] Based on the target eigenvalue and target rank, verify each set of perturbation vectors. The purpose of the verification is to check whether the selected vector group conforms to the eigenvectors and rank of the perturbation matrix. If it conforms, it means that the data distribution of the selected set of perturbation vectors is similar to the global data distribution of the perturbation matrix, ensuring that the selected vector group can effectively represent the information structure of the matrix. For the perturbation vector groups that pass the verification, determine them as the target perturbation vector groups.

[0060] In this embodiment, by screening out perturbation vector groups with balanced data distribution from the perturbation matrix to construct a balanced and efficient client set, the training effect of the federated learning model is improved. It effectively overcomes the challenges brought by uneven data distribution, while protecting data privacy and security. It enhances the performance of the model and improves the prediction ability and robustness of the federated learning model.

[0061] To ensure that the selected clients cover more data categories, a reference vector group can be selected based on the norm of the perturbation vectors. Optionally, in the credit risk determination method provided in the embodiments of this application, extracting the reference vector group from the perturbation matrix includes: for each perturbation vector in the perturbation matrix, calculating the norm of the perturbation vector; when the norm of the perturbation vector is greater than a preset value, determining the perturbation vector as a reference vector; and forming a reference vector group from all the reference vectors. Combine all the determined reference perturbation vectors into a reference vector group.

[0062] In some examples, a reference vector group is selected from the perturbation matrix M. The selection criterion can be a vector with as many 1s in the column as possible, that is, a vector where l > m / 2. l is the norm of each perturbation vector, and m / 2 is the preset value, which is also half of the number of data categories. The calculation formula for the norm of the perturbation vector is as follows:

[0063]

[0064] where a 1、 a 2、 a 3、 a m is the eigenvalue in the perturbation vector.

[0065] In this embodiment, by quantifying the information content of the perturbation vectors, vectors that have a significant impact on the data distribution can be effectively screened out, thereby constructing a reference vector group that can reflect the core characteristics of the data set. In the context of federated learning, these reference vector groups will be used for subsequent client set selection and model training optimization to improve the performance and stability of the global model, while reducing communication overhead and protecting data privacy.

[0066] For each group of screened perturbation vectors, it is necessary to further verify whether they can represent the characteristics of the perturbation matrix. Optionally, in the credit risk determination method provided in the embodiments of this application, verifying each group of perturbation vectors based on the target eigenvalue and target rank, and determining each group of perturbation vectors that pass the verification as a group of target perturbation vectors includes: calculating the rank and eigenvalue of the matrix formed by each group of perturbation vectors, determining whether the rank is the same as the target rank, and determining whether the eigenvalue is the same as the target eigenvalue; when the rank is different from the target rank, or the eigenvalue is different from the target eigenvalue, removing this group of perturbation vectors; when the rank is the same as the target rank and the eigenvalue is the same as the target eigenvalue, determining this group of perturbation vectors as a group of target perturbation vectors.

[0067] In some examples, singular value decomposition is performed on the matrix formed by each group of perturbation vectors to calculate the eigenvalues and rank of the group. The eigenvalues reflect the main information and energy distribution of the group of perturbation vectors, while the rank represents the linearly independent dimension of the group of perturbation vectors, that is, the size of the data space that the vector group can cover. The calculated eigenvalues and rank of each group of perturbation vector matrices are compared with the target eigenvalues and target rank. If the rank of the matrix formed by a certain group of perturbation vectors is the same as the target rank, and the eigenvalues of this group are the same as or close enough to the target eigenvalues, then this group of perturbation vectors passes the verification and is considered the target perturbation vector group. If the rank of the perturbation vector group is different from the target rank, or the eigenvalues are significantly different from the target eigenvalues, it indicates that this vector group cannot fully reflect the key attributes of the perturbation matrix. Therefore, this group of perturbation vectors should be removed from the candidate list to avoid its participation in the subsequent federated learning model training.

[0068] In this embodiment, the perturbation vector group is verified to screen out the target perturbation vector group, ensuring that the model training is based on the vector group that can best reflect the characteristics of the entire dataset, thereby improving the generalization ability and robustness of the model, while reducing communication overhead and data processing time.

[0069] To screen out the set of clients suitable for training the federated learning model, it is necessary to calculate the value evaluation value of each set of client sets. Optionally, in the method for determining credit risk provided in the embodiments of the present application, the evaluation value is calculated based on the model parameters of the local models of each client in the client set and the model parameters of the global model. The value evaluation value of the client set includes: for each client in the client set, calculate the similarity between the model parameters of the local model of the client and the model parameters of the global model to obtain the similarity evaluation value of the client; determine the actual probability distribution and target probability distribution of the training data of the client, and calculate the relative entropy based on the actual probability distribution and the target probability distribution to obtain the balance evaluation value of the client; calculate the sum of the reciprocal of the similarity evaluation value, the balance evaluation value, and the adjustment parameter, and calculate the reciprocal of the sum to obtain the sub-value evaluation value of the client; calculate the sum of the sub-value evaluation values of all clients in the client set to obtain the value evaluation value of the client set.

[0070] In some examples, the client data participating in the federated learning training exhibits the characteristics of non-independent and identically distributed. When randomly selecting clients in the federated averaging algorithm and selecting clients with extremely unbalanced local data distribution and large distribution deviation and relatively low data quality, this embodiment adopts an equilibrium evaluation method based on KL divergence to screen the client set from the current client selection scheme based on whether the distribution of client data is approximately the same as the distribution of global data. The equilibrium evaluation value can be calculated by the following formula:

[0071]

[0072]

[0073] Among them, the balance evaluation value can be calculated using the KL divergence. The KL divergence is used to represent the similarity between two data distributions. p(x) represents the actual probability distribution, and Q(x) represents the target probability distribution, that is, the theoretical probability distribution of the data. In the balance evaluation method based on the KL divergence, the true probability distribution of the client data distribution is p(x), and the theoretical probability distribution of the dataset is Q(x). In a dataset following a uniform distribution, assuming the category where the data is located is x and the total number of categories in the dataset is N, the theoretical probability distribution Q(x) of the data with category x is 1 / N. The theoretical probability distribution Q(x) is 1 / N for each classification. It is only necessary to calculate the true probability distribution p(x) corresponding to each classification to calculate the KL divergence between p(x) and Q(x). The calculation formula of p(x) is as follows:

[0074]

[0075] Among them, yi is the data volume corresponding to each classification, and pi is the probability corresponding to each classification. The calculation formula of the sub-value evaluation value of each client is as follows:

[0076]

[0077] Among them, the smaller the KL divergence, the higher the similarity degree between the two distributions. The value range calculated by the KL divergence is a positive number greater than 0. However, the calculation of the cosine similarity is the cosine distance between two models, and the larger it is, the smaller the distance between the two models. Therefore, the sum of the reciprocal of the cosine similarity and the KL divergence is selected as the denominator. At the same time, in order to prevent the occurrence of extreme situations, an ε is added as an adjustment parameter. ε is an extremely small positive number.

[0078] After calculating the sub-value evaluation value of each client, calculate the sum of the sub-value evaluation values of all clients in the client set to obtain the value evaluation value of the client set.

[0079] In this embodiment, by quantifying the value evaluation value, the federated learning system can intelligently select those clients with balanced data distributions and model parameters highly similar to the global model for training, thereby effectively improving the generalization ability and training efficiency of the model, and at the same time reducing the problem of model performance degradation caused by non-independent and identically distributed data.

[0080] To calculate the value evaluation value of the client set, it is necessary to calculate the similarity evaluation value of each client. Optionally, in the method for determining credit risk provided in the embodiments of the present application, calculating the similarity between the model parameters of the local model of the client and the model parameters of the global model to obtain the similarity evaluation value of the client includes: determining the first model parameters of the global model in the previous iteration round of the federated learning model; training the local model based on the first model parameters and the training data of the client to obtain the model parameters of the local model; obtaining the second model parameters of the global model in the current iteration round of the federated learning model, and calculating the cosine similarity between the second model parameters and the model parameters of the local model to obtain the similarity evaluation value of the client.

[0081] In some examples, in order to maintain the long-term effective participation of data owners in training, improve the convergence speed of the model, and fairly evaluate the quality of each data source and compensate the contribution of data owners to the training process is very important. High-quality data sources can not only reduce the number of communication rounds, but also improve the quality of the model. In the client selection process of the first stage, an equilibrium evaluation calculation is performed on the clients, and a combination of clients that meet the requirements is selected. However, the contribution degrees of the local models trained by different clients to the convergence of the global model are different. In order to reduce the communication rounds and optimize the convergence process of the model, it is necessary to select local models with a higher contribution degree to the global model, that is, select local models that are more similar to the global model.

[0082] In federated learning, the global model can be initialized as w 0 , so that the global model in the previous round in the t-th round of iteration is w t-1 , that is, the first model parameters. The i-th client selected in this round is under w t-1 The training of the global model generates (that is, the model parameters of the local model), W t That is, the second model parameters. According to these two parameters, the cosine similarity of the two models can be calculated using the formula as follows:

[0083]

[0084] The higher the cosine similarity, the closer the relationship between the two models, indicating that the contribution degree of the local model to the global model is higher. If the current local model is aggregated on the central server, it can enhance the convergence speed and accuracy of the global model. In particular, adding high-quality local models to the aggregation of the global model can effectively reduce the communication overhead.

[0085] In this embodiment, by precisely evaluating the degree of fit between the model parameter updates of each client in the current iteration round and the global model, it provides an important basis for subsequent model aggregation, client selection, and training optimization. The calculation of the similarity evaluation value helps identify the key clients that can improve the performance of the global model, and also helps detect and process abnormal or low-quality data sources, ensuring the efficiency of the federated learning process and the robustness of the model.

[0086] After determining the target client set, it is necessary to train the federated learning model using the target client set. Optionally, in the method for determining credit risk provided in the embodiments of the present application, training the federated learning model based on the target client set includes: determining the first model parameters of the global model of the federated learning model in the previous iteration round, training local models based on the first model parameters and the training data of each client in the target client set to obtain the local model parameters of each client; aggregating the local model parameters of all clients in the target client set to obtain the model parameters of the global model of the federated learning model in the current iteration round.

[0087] In some examples, during the process of federated learning, at the beginning of each round of iteration, the central server distributes the global model parameters w t-1 aggregated after the end of the previous round of iteration, where t represents the current iteration round. These parameters constitute the initial state of the federated learning model and are used for training and updating in the current round.

[0088] For each client i in the target client set S, it uses the distributed global model parameters w t-1 and its own local training data set to perform local model training. The training process involves steps such as forward propagation, loss calculation, backpropagation, and parameter update. Client i adjusts w t-1 according to the local data to obtain the updated local model parameters This process is carried out independently for each client, ensuring data privacy and parallelism. After completing local training, client i uploads the updated local model parameters to the central server. The uploaded data only contains the updates of the model parameters and does not contain the original training data, thus protecting data privacy.

[0089] The central server collects the local model parameters from all clients in the target client set S and performs parameter aggregation. Parameter aggregation can be calculating the weighted average of all local model parameters, and the weights can be a function of the client data volume or a dynamic adjustment based on model performance. The central server distributes the aggregated global model parameters W t to the clients participating in the training in the next round, and repeats the above process until the model converges, that is, the change in the global model parameters is lower than a preset threshold or reaches a predetermined number of iterations.

[0090] In this embodiment, by using the selected set of target clients, model training and parameter update are effectively carried out while protecting the privacy of client data. When dealing with large-scale and distributed data sets, it can overcome the imbalance and non-independent and identically distributed nature of data distribution, and improve the generalization ability and training efficiency of the model.

[0091] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0092] The embodiment of the present application also provides a device for determining credit risk. It should be noted that the device for determining credit risk in the embodiment of the present application can be used to execute the method for determining credit risk provided in the embodiment of the present application. The following introduces the device for determining credit risk provided in the embodiment of the present application.

[0093] Figure 4 It is a schematic diagram of the device for determining credit risk provided according to the embodiment of the present application. As Figure 4 shown, the device includes:

[0094] An acquisition unit 401, configured to acquire the data to be analyzed of the target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records;

[0095] An extraction unit 402, configured to extract the target area to which the target user belongs from the identity information, and determine the client corresponding to the target area in the federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by the training data of the target client set;

[0096] An input unit 403, configured to input the data to be analyzed into the target local model of the client and output the credit risk value of the target user, where the target local model is trained by the model parameters of the global model and the training data of the client.

[0097] The credit risk determination device provided by the embodiment of the present application, through the acquisition unit 401, acquires the data to be analyzed of the target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; the extraction unit 402 extracts the target area to which the target user belongs from the identity information, determines the client corresponding to the target area in the federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by the training data of the target client set; the input unit 403 inputs the data to be analyzed into the target local model of the client and outputs the credit risk value of the target user, where the target local model is trained by the model parameters of the global model and the training data of the client, which solves the problem in the related art that due to the heterogeneity of the training data of the federated learning model, the communication efficiency of the federated learning model is low, and further the cost of credit risk assessment for users is high. By using the training data of the target client set to train the federated learning model instead of using the training data of all clients, the problem of heterogeneity of client data is avoided, thereby reducing the communication overhead during the training of the federated learning model, and further achieving the effect of reducing the cost of credit risk assessment for users.

[0098] Optionally, in the credit risk determination device provided by the embodiment of the present application, the device further includes: a perturbation unit, configured to, for each iterative training in the federated learning model, acquire the training data of each client, and perturb the training data through a random perturbation algorithm to obtain the perturbation vector of the client, where the training data includes the historical data to be analyzed and the historical credit risk value of the user of the client; a generation unit, configured to generate a perturbation matrix based on the perturbation vectors of all clients, perform singular value decomposition on the perturbation matrix to obtain a target eigenvector and a target rank, where the number of rows of the perturbation matrix represents the number of clients, and the number of columns of the perturbation matrix represents the number of data categories of the training data of all clients; a screening unit, configured to screen out multiple groups of target perturbation vectors from the perturbation matrix based on the target eigenvector and the target rank, determine the client set corresponding to each group of target perturbation vectors, and obtain multiple groups of client sets; a calculation unit, configured to, for each group of client sets, calculate an evaluation value based on the model parameters of the local model of each client in the client set and the model parameters of the global model, and obtain the value evaluation value of the client set; a training unit, configured to determine the client set corresponding to the maximum value evaluation value among the multiple groups of client sets as the target client set, and train the federated learning model based on the target client set.

[0099] Optionally, in the credit risk determination device provided in the embodiments of the present application, the screening unit includes: an extraction module, configured to extract a reference vector group from the perturbation matrix, and determine the perturbation vectors other than the reference vector group in the perturbation matrix as vectors to be added; a traversal module, configured to, for each target vector to be added, traverse one by one the other vectors to be added in the perturbation matrix except the target vector to be added, and calculate the rank of the matrix formed by each traversed vector to be added and the target vector to be added; an addition module, configured to, when the rank is a target value, add the traversed vector to be added to the reference vector group until the rank of the matrix formed by the reference vector group is equal to the number of data categories or the traversal ends when all the vectors to be added in the perturbation matrix are traversed; a composition module, configured to form a group of perturbation vectors from each target vector to be added and the reference vector group corresponding to the target vector to be added, and obtain multiple groups of perturbation vectors; a verification module, configured to verify each group of perturbation vectors based on the target eigenvalue and the target rank, and determine each group of perturbation vectors that pass the verification as a group of target perturbation vectors.

[0100] Optionally, in the credit risk determination device provided in the embodiments of the present application, the extraction module includes: a first calculation sub-module, configured to calculate the norm of each perturbation vector in the perturbation matrix; a first determination sub-module, configured to, when the norm of the perturbation vector is greater than a preset value, determine the perturbation vector as a reference vector; a composition sub-module, configured to form a reference vector group from all the reference vectors.

[0101] Optionally, in the credit risk determination device provided in the embodiments of the present application, the verification module includes: a second calculation sub-module, configured to calculate the rank and eigenvalue of the matrix formed by each group of perturbation vectors, determine whether the rank is the same as the target rank, and determine whether the eigenvalue is the same as the target eigenvalue; a rejection sub-module, configured to, when the rank is different from the target rank or the eigenvalue is different from the target eigenvalue, reject the group of perturbation vectors; a second determination sub-module, configured to, when the rank is the same as the target rank and the eigenvalue is the same as the target eigenvalue, determine the group of perturbation vectors as a group of target perturbation vectors.

[0102] Optionally, in the credit risk determination device provided in the embodiments of the present application, the calculation unit includes: a first calculation module, configured to calculate, for each client in the client set, the similarity between the model parameters of the local model of the client and the model parameters of the global model, to obtain the similarity evaluation value of the client; a first determination module, configured to determine the actual probability distribution and the target probability distribution of the training data of the client, and calculate the relative entropy based on the actual probability distribution and the target probability distribution, to obtain the balance evaluation value of the client; a second calculation module, configured to calculate the sum of the reciprocal of the similarity evaluation value, the balance evaluation value, and the adjustment parameter, and calculate the reciprocal of the sum, to obtain the sub-value evaluation value of the client; a third calculation module, configured to calculate the sum of the sub-value evaluation values of all clients in the client set, to obtain the value evaluation value of the client set.

[0103] Optionally, in the credit risk determination device provided in the embodiments of the present application, the first calculation module includes: a third determination sub-module, configured to determine the first model parameters of the global model of the previous iteration round of the federated learning model; a training sub-module, configured to train the local model based on the first model parameters and the training data of the client, to obtain the model parameters of the local model; an acquisition sub-module, configured to acquire the second model parameters of the global model of the current iteration round of the federated learning model, and calculate the cosine similarity between the second model parameters and the model parameters of the local model, to obtain the similarity evaluation value of the client.

[0104] Optionally, in the credit risk determination device provided in the embodiments of the present application, the training unit includes: a second determination module, configured to determine the first model parameters of the global model of the previous iteration round of the federated learning model, and train the local model based on the first model parameters and the training data of each client in the target client set, to obtain the local model parameters of each client; an aggregation module, configured to aggregate the local model parameters of all clients in the target client set, to obtain the model parameters of the global model of the current iteration round of the federated learning model.

[0105] The credit risk determination device includes a processor and a memory. The above-mentioned acquisition unit 401, extraction unit 402, input unit 403, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above-mentioned program units stored in the memory.

[0106] The processor includes a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and the cost of credit risk assessment for users can be reduced by adjusting the kernel parameters.

[0107] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0108] An embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, a method for determining credit risk is implemented.

[0109] An embodiment of the present invention provides a processor for running a program, and when the program runs, a method for determining credit risk is executed.

[0110] Figure 5 It is a schematic diagram of an electronic device provided according to an embodiment of the present application. As Figure 5 shown, the electronic device 501 includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented: obtaining the data to be analyzed of the target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; extracting the target area to which the target user belongs from the identity information, and determining the client corresponding to the target area in the federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by the training data of the target client set; inputting the data to be analyzed into the target local model of the client, and outputting the credit risk value of the target user, where the target local model is trained by the model parameters of the global model and the training data of the client. The device herein may be a server, a PC, a PAD, a mobile phone, etc.

[0111] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with the following method steps: obtaining the data to be analyzed of the target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; extracting the target area to which the target user belongs from the identity information, and determining the client corresponding to the target area in the federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by the training data of the target client set; inputting the data to be analyzed into the target local model of the client, and outputting the credit risk value of the target user, where the target local model is trained by the model parameters of the global model and the training data of the client.

[0112] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0113] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0114] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0116] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0117] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0118] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0119] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0120] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for determining credit risk, characterized in that, Including: Obtain the data to be analyzed of the target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; Extract the target area to which the target user belongs from the identity information, and determine the client corresponding to the target area in the federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by the training data of the target client set; Input the data to be analyzed into the target local model of the client, and output the credit risk value of the target user, where the target local model is trained by the model parameters of the global model and the training data of the client.

2. The method according to claim 1, wherein The federated learning model is trained in the following manner: For each iterative training in the federated learning model, obtain the training data of each client, and perturb the training data through a random perturbation algorithm to obtain the perturbation vector of the client, where the training data includes the historical data to be analyzed and the historical credit risk value of the user of the client; Generate a perturbation matrix based on the perturbation vectors of all clients, perform singular value decomposition on the perturbation matrix to obtain the target eigenvector and the target rank, where the number of rows of the perturbation matrix represents the number of clients, and the number of columns of the perturbation matrix represents the number of data categories of the training data of all clients; Screen out multiple groups of target perturbation vectors from the perturbation matrix based on the target eigenvector and the target rank, determine the client set corresponding to each group of target perturbation vectors, and obtain multiple groups of client sets; For each group of client sets, calculate the evaluation value based on the model parameters of the local models of each client in the client set and the model parameters of the global model, and obtain the value evaluation value of the client set; Determine the client set corresponding to the maximum value evaluation value among the multiple groups of client sets as the target client set, and train the federated learning model based on the target client set.

3. The method according to claim 2, characterized in that, Screening out multiple groups of target perturbation vectors from the perturbation matrix based on the target eigenvector and the target rank includes: Extract a reference vector group from the perturbation matrix, and determine the vectors to be added as the vectors in the perturbation matrix other than the reference vector group; For each target vector to be added, traverse one by one the other vectors to be added in the perturbation matrix except the target vector to be added, and calculate the rank of the matrix formed by each traversed vector to be added and the target vector to be added; When the rank is the target value, add the traversed vector to be added to the reference vector group until the rank of the matrix formed by the reference vector group is equal to the number of data categories, or when the vectors to be added in the perturbation matrix are traversed; Each group of perturbation vectors is composed of a target vector to be added and the reference vector group corresponding to the target vector to be added, and multiple groups of perturbation vectors are obtained; Verify each group of perturbation vectors based on the target eigenvalue and the target rank, and determine each group of perturbation vectors that pass the verification as a group of target perturbation vectors.

4. The method according to claim 3, characterized in that, Extracting a reference vector group from the perturbation matrix includes: For each perturbation vector in the perturbation matrix, calculating the norm of the perturbation vector; When the norm of the perturbation vector is greater than a preset value, determining the perturbation vector as a reference vector; Forming the reference vector group from all the reference vectors.

5. The method according to claim 2, wherein Verifying each group of perturbation vectors based on the target eigenvalue and the target rank, and determining each group of perturbation vectors that pass the verification as a group of target perturbation vectors includes: Calculating the rank and eigenvalue of the matrix formed by each group of perturbation vectors, and determining whether the rank is the same as the target rank and whether the eigenvalue is the same as the target eigenvalue; When the rank is different from the target rank or the eigenvalue is different from the target eigenvalue, removing this group of perturbation vectors; When the rank is the same as the target rank and the eigenvalue is the same as the target eigenvalue, determining this group of perturbation vectors as a group of target perturbation vectors.

6. The method according to claim 2, wherein Calculating an evaluation value based on the model parameters of the local models of each client in the client set and the model parameters of the global model to obtain the value evaluation value of the client set includes: For each client in the client set, calculating the similarity between the model parameters of the local model of the client and the model parameters of the global model to obtain the similarity evaluation value of the client; Determining the actual probability distribution and the target probability distribution of the training data of the client, and calculating the relative entropy based on the actual probability distribution and the target probability distribution to obtain the balance evaluation value of the client; Calculating the sum of the reciprocal of the similarity evaluation value, the balance evaluation value and the adjustment parameter, and calculating the reciprocal of the sum to obtain the sub-value evaluation value of the client; Calculating the sum of the sub-value evaluation values of all clients in the client set to obtain the value evaluation value of the client set.

7. The method according to claim 6, wherein Calculating the similarity between the model parameters of the local model of the client and the model parameters of the global model to obtain the similarity evaluation value of the client includes: Determining the first model parameters of the global model in the previous iteration round of the federated learning model; Training the local model based on the first model parameters and the training data of the client to obtain the model parameters of the local model; Obtaining the second model parameters of the global model in the current iteration round of the federated learning model, and calculating the cosine similarity between the second model parameters and the model parameters of the local model to obtain the similarity evaluation value of the client.

8. The method according to claim 2, wherein Training the federated learning model based on the target client set includes: Determining the first model parameters of the global model in the previous iteration round of the federated learning model, and training the local model based on the first model parameters and the training data of each client in the target client set to obtain the local model parameters of each client; Aggregating the local model parameters of all clients in the target client set to obtain the model parameters of the global model in the current iteration round of the federated learning model.

9. A device for determining credit risk, characterized in that, Includes: An acquisition unit, configured to acquire the data to be analyzed of a target user, where the data to be analyzed includes at least one of the following: identity information, asset data, and consumption records; An extraction unit, configured to extract a target region to which the target user belongs from the identity information, and determine a client corresponding to the target region in a federated learning model, where the federated learning model includes a global model and local models of multiple clients, and the federated learning model is trained by training data of a target client set; An input unit, configured to input the data to be analyzed into a target local model of the client, and output a credit risk value of the target user, where the target local model is trained by model parameters of the global model and training data of the client.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for determining credit risk according to any one of claims 1 to 8.