Risk management method, risk model training method, device, equipment and medium
Through the risk training of federal learning and combined with encryption algorithms, the problem of inaccurate assessment in collateral management of commercial banks is solved, and more scientific and effective credit risk management is achieved, and loan risks are reduced.
Patent Information
- Application Number
- CN202210376473.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-11
AI Technical Summary
Commercial banks lack effective collateral management mechanisms, resulting in credit risks when issuing loans, and cannot scientifically and reasonably evaluate the collateral value, which is prone to value decay.
The risk federal model is trained using a federated learning algorithm, and the bank server and the big data bureau server jointly train collateral sample data to generate credit request feedback data and whitelists, and risk assessment is carried out through user characteristic data, combining homomorphic encryption and asymmetric encryption to protect data privacy.
It improves the scientificity and effectiveness of risk management, reduces loan risks, ensures the rationality and accuracy of the assessment results, and avoids the situation where the assessment price is higher than the real price.
Smart Images

Figure CN114971841B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, the field of big data technology or the field of finance, and particularly relates to a risk management method, a risk model training method, a device, a device, a medium and a program product. Background Art
[0002] At present, when commercial banks determine the value of collateral, they generally refer to the evaluation value of a third-party evaluation agency and then conduct a re-evaluation according to the relevant rules and regulations of the bank's internal collateral evaluation to determine the market value of the collateral. However, there is a lack of professional evaluation personnel within commercial banks, making it difficult to conduct an effective re-evaluation of the commodity value, and often they can only accept the collateral evaluation value of the third-party agency. Due to factors such as the lack of supervision of evaluation agencies and fierce market competition, it is easy to occur that the collateral that seems to be sufficient and effective at the time of loan issuance shows a significant value decline at the time of disposal.
[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in the prior art:.
[0004] Since commercial banks have not established an effective collateral management mechanism, they passively accept the conclusions of third-party agencies, which leaves hidden dangers at the time of loan issuance and cannot scientifically, reasonably and effectively avoid credit risks. Summary of the Invention
[0005] In view of the above problems, the embodiments of the present disclosure provide a risk management method, device, device, medium and program product that improve the scientificity, rationality and effectiveness of risk management. Further, the embodiments of the present disclosure also provide a risk model training method, device, device, medium and program product.
[0006] According to a first aspect of the present disclosure, there is provided a risk management method applied to a bank server, including: obtaining credit request data, where the credit request data includes a user identifier; obtaining corresponding user feature data based on the user identifier; inputting the user feature data into a pre-trained first risk federated model to obtain a risk assessment result; and generating credit request feedback data and a credit white list based on the risk assessment result. Wherein, the first risk federated model is trained based on a federated learning algorithm, and the training is jointly executed by the bank server and the big data bureau server, and the data used for training includes collateral sample data, and the collateral sample data is obtained from the big data bureau server.
[0007] According to an embodiment of the present disclosure, the user feature data includes user attribute information, user asset information, and product attribute information.
[0008] According to an embodiment of the present disclosure, the pre-trained first risk federated model is automatically updated based on a preset time period.
[0009] According to an embodiment of the present disclosure, after generating the credit white list, the method further includes: storing the credit white list and sending the credit white list to the big data bureau server.
[0010] A second aspect of the present disclosure provides a method for training a risk model based on federated learning, which includes: obtaining intersection identification data of first user identification data and second user identification data based on an asymmetric encryption algorithm, wherein the first user identification data is obtained from a bank server, and the second user identification data is obtained from a big data bureau server; initializing the first federated model parameters deployed on the bank server and the second federated model parameters deployed on the big data bureau server; training and updating the first federated model parameters and the second federated model parameters based on a homomorphic encryption algorithm until a preset training cut-off condition is reached, wherein the first federated model and the second federated model are jointly trained based on first feature data and second feature data, the first feature data and the second feature data are respectively obtained by the bank server and the big data bureau server based on the intersection identification data, the first feature data includes user credit sample data, and the user credit sample data includes user feature sample data and user credit labels, and the second feature data includes collateral sample data; and obtaining a first risk federated model and a second risk federated model, wherein the first risk federated model is the first federated model including the first federated model parameters at the end of training, and the second risk federated model is the second federated model including the second federated model parameters at the end of training.
[0011] According to an embodiment of the present disclosure, updating the first federal model parameters and the second federal model parameters based on the homomorphic encryption algorithm until a preset training cut-off condition is reached includes: the bank server and the big data bureau server respectively training the first federal model and the second federal model based on the first feature data and the second feature data, and calculating the first transfer data and the second transfer data; the bank server and the big data bureau server respectively performing homomorphic encryption on the first transfer data and the second transfer data, and interchanging and transmitting the first encrypted transfer data and the second encrypted transfer data after homomorphic encryption, wherein the public key used in the homomorphic encryption is obtained from a third-party server; the bank server training the first federal model based on the second encrypted transfer data to obtain first encrypted gradient information and first encrypted loss information, and the big data bureau server training the second federal model based on the first encrypted transfer data to obtain second encrypted gradient information; the bank server and the big data bureau server respectively sending the first encrypted gradient information, the first encrypted loss information, and the second encrypted gradient information to the third-party server; the third-party server decrypting the first encrypted gradient information, the first encrypted loss information, and the second encrypted gradient information based on the held private key, obtaining and sending the first decrypted gradient information and the first decrypted loss information to the bank server, and obtaining and sending the second decrypted gradient information to the big data bureau server; and the bank server updating the first federal model parameters based on the first decrypted gradient information and the first decrypted loss information, and the big data bureau server updating the second federal model parameters based on the second decrypted gradient information.
[0012] According to an embodiment of the present disclosure, before obtaining the intersection identification data, the method further includes a user screening step, including: obtaining training set users and validation set users, wherein the training set users include positive sample users and negative sample users, the positive sample users include users who applied for collateral loans and were approved and repaid on time within a first preset time period, the negative sample users include users who applied for collateral loans but were not approved for collateral, or users who applied for collateral loans and were approved but did not repay on time, the number of positive sample users and the number of negative sample users are set based on a preset ratio, and the validation set users include users who applied for collateral loans within a second preset time period.
[0013] According to an embodiment of the present disclosure, the first risk model and the second risk model are constructed based on the XGBoost algorithm.
[0014] According to an embodiment of the present disclosure, the collateral sample data includes at least one of movable property collateral data, immovable property collateral data, and intangible asset collateral data.
[0015] A third aspect of the present disclosure provides a risk management device deployed in a bank server, including: a first acquisition module configured to acquire credit request data, where the credit request data includes a user identifier; a second acquisition module configured to acquire corresponding user feature data based on the user identifier; a calculation module configured to input the user feature data into a pre-trained first risk federated model to obtain a risk assessment result, where the first risk federated model is trained based on a federated learning algorithm, and the training is jointly executed by the bank server and the big data bureau server, and the data used for training includes collateral sample data, and the collateral sample data is acquired from the big data bureau server; and a generation module configured to generate credit request feedback data and a credit white list based on the risk assessment result.
[0016] A fourth aspect of the present disclosure provides a risk model training system, including: an alignment device configured to obtain intersection identifier data of first user identifier data and second user identifier data based on an asymmetric encryption algorithm, where the first user identifier data is acquired from the bank server and the second user identifier data is acquired from the big data bureau server; an initialization device configured to initialize the first federated model parameters deployed in the bank server and the second federated model parameters deployed in the big data bureau server; a calculation device configured to perform training updates on the first federated model parameters and the second federated model parameters based on a homomorphic encryption algorithm until a preset training cut-off condition is reached, where the first federated model and the second federated model are jointly trained based on first feature data and second feature data, and the first feature data and the second feature data are respectively acquired by the bank server and the big data bureau server based on the intersection identifier data, the first feature data includes user credit sample data, and the user credit sample data includes user feature sample data and user credit labels, and the second feature data includes collateral sample data; and a generation device configured to obtain a first risk federated model and a second risk federated model, where the first federated risk model is the first federated model including the first federated model parameters at the end of training, and the second federated risk model is the second federated model including the second federated model parameters at the end of training.
[0017] A fifth aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above method.
[0018] A sixth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the above method.
[0019] The seventh aspect of the present disclosure further provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0020] The method provided by the embodiments of the present disclosure trains a first federal risk model deployed on a bank server through federated learning technology. During the process of obtaining the first federal risk model, collateral sample data from a big data bureau server is utilized. The first risk federal model processes new user credit request data to evaluate risks, improving the scientificity, rationality, and effectiveness of risk management. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0022] Figure 1 Schematically shows an application scenario diagram of a risk management method, device, equipment, medium, and program product according to an embodiment of the present disclosure.
[0023] Figure 2 Schematically shows a flowchart of a risk management method according to an embodiment of the present disclosure.
[0024] Figure 3 Schematically shows a flowchart of a method for synchronously storing a credit white list according to an embodiment of the present disclosure.
[0025] Figure 4 Schematically shows a flowchart of a risk model training method based on federated learning according to another embodiment of the present disclosure.
[0026] Figure 5 Schematically shows a schematic diagram of a method for aligning first user identification data and second user identification data.
[0027] Figure 6 Schematically shows a schematic diagram of a homomorphic encryption algorithm.
[0028] Figure 7 Schematically shows an exemplary system architecture of a method and device for training a risk model according to another embodiment of the present disclosure.
[0029] Figure 8 Schematically shows a flowchart of a method for updating first federal model parameters and second federal model parameters based on a homomorphic encryption algorithm according to another embodiment of the present disclosure.
[0030] Figure 9 Schematically shows a flowchart of a method for screening users according to an embodiment of the present disclosure.
[0031] Figure 10A structural block diagram of a risk management device according to an embodiment of the present disclosure is schematically shown.
[0032] Figure 11 A structural block diagram of a risk model training system according to another embodiment of the present disclosure is schematically shown.
[0033] Figure 12 A block diagram of an electronic device suitable for implementing a risk management method according to an embodiment of the present disclosure is schematically shown. Detailed implementation manners
[0034] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0035] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0036] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0037] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0038] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.
[0039] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.
[0040] It should be noted that the risk management methods, risk model training methods, devices, equipment, media, and program products provided in the embodiments of the present disclosure can be used in aspects related to risk management in artificial intelligence technology and big data technology, and can also be used in various fields other than artificial intelligence technology and big data technology, such as the financial field. The application fields of the risk management methods, risk model training methods, devices, equipment, media, and program products provided in the embodiments of the present disclosure are not limited.
[0041] When a debtor or a third party provides collateral to secure the realization of relevant claims of a commercial bank, they usually mortgage or pledge the collateral to the commercial bank, which is property or rights used to mitigate credit risks. Collateral management should follow: 1) The principle of legality. Collateral management should comply with laws and regulations; 2) The principle of effectiveness. The mortgage and pledge guarantee procedures are complete, the value of the collateral is reasonable and easy to liquidate in kind, and it has a good role in guaranteeing claims; 3) The principle of prudence. Fully consider the possible risk factors in the collateral itself, carefully formulate collateral management policies, and dynamically evaluate the value of the collateral and its risk mitigation effect; 4) The principle of subordination. The commercial bank's use of collateral to mitigate credit risks should be premised on a comprehensive assessment of the debtor's debt repayment ability. Currently, for the determination of the value of collateral by commercial banks, generally, the evaluation value of a third-party evaluation agency is referred to, and then a re-evaluation is carried out according to the relevant regulations of the bank's internal collateral evaluation, so as to determine the market value of the collateral. However, there is a lack of professional evaluation personnel within commercial banks, making it difficult to conduct an effective re-evaluation of the value of goods, and often they can only accept the collateral evaluation value of third-party agencies. Due to factors such as the lack of supervision of evaluation agencies and fierce market competition, it is easy to occur that the collateral that seems to be sufficient and effective at the time of loan issuance shows a significant decline in value when it comes to disposal. This is mainly manifested in: the risk of overestimating the value of the collateral. Borrowers hope to maximize the loan amount and reduce the default cost, so they have the motivation to inflate the value of the collateral, and there are situations where the evaluation value of the collateral is distorted and overestimated. Since commercial banks have not established an effective collateral management mechanism, it is easy to omit or fail to identify some important information in the evaluation report, resulting in passively accepting the conclusions of third-party agencies, thus leaving hidden dangers at the time of loan issuance and failing to meet the relevant regulations of collateral risk management. The scientificity, rationality, and effectiveness of risk management are poor.
[0042] Currently, machine learning based on big data has promoted the vigorous development of artificial intelligence (AI) technology. Among them, federated learning is a distributed machine learning technology and system, and its core idea is "data does not move, the model moves". Federated learning can jointly model multiple data sources to provide inference and prediction services. Each participating party does not exchange raw data, but only exchanges intermediate calculation results of model parameters, ensuring that the data of each party is not leaked. In the model inference stage, the trained federated learning model can be deployed at each participating party of the federated learning system, or a public interface can be provided for multiple parties to share.
[0043] In the process of obtaining the embodiments of the present disclosure, the inventors found that by docking with the Big Data Bureau and using the method of federated learning, the user feature data in commercial banks and the collateral information of users in the Big Data Bureau can be used as feature data to train a risk management model that can fully utilize the collateral information to evaluate the repayment ability of customers. After receiving a new credit request, a commercial bank can use this model to evaluate the user's credit request, so as to more reasonably, scientifically and effectively evaluate credit risks, avoid the situation where the evaluation price of the evaluation company is higher than the real price, and reduce the risk of lending by commercial banks.
[0044] An embodiment of the present disclosure provides a risk management method applied to a bank server, including: obtaining credit request data, where the credit request data includes a user identifier; obtaining corresponding user feature data based on the user identifier; inputting the user feature data into a pre-trained first risk federated model to obtain a risk assessment result; and generating credit request feedback data and a credit white list based on the risk assessment result, where the first risk federated model is trained based on a federated learning algorithm, and the training is jointly executed by the bank server and the Big Data Bureau server, and the data used for training includes collateral sample data, and the collateral sample data is obtained from the Big Data Bureau server.
[0045] The following will elaborate on the above operations around achieving at least one object of the present disclosure in conjunction with the accompanying drawings and their explanatory text.
[0046] Figure 1 A schematic application scenario diagram of a risk management method, device, device, medium and program product according to an embodiment of the present disclosure is shown.
[0047] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a bank server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the bank server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0048] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. For example, users can use the terminal devices 101, 102, 103 to send credit request data and receive credit request feedback data. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0049] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and so on.
[0050] The bank server 105 can be a server that provides functions such as receiving credit requests, model training, and application. For example, it can be a background management server (only for example) that supports the websites browsed by users using the terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0051] It should be noted that the risk management method provided by the embodiments of the present disclosure can generally be executed by the bank server 105. Correspondingly, the risk management device provided by the embodiments of the present disclosure can generally be set in the bank server 105. The risk management method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the bank server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the bank server 105. Correspondingly, the risk management device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the bank server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the bank server 105.
[0052] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in
[0053] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the scenario described below Figures 2 - 3 the risk management method of the embodiments of the present disclosure will be described in detail.
[0054] Figure 2 FIG. schematically shows a flowchart of the risk management method according to an embodiment of the present disclosure.
[0055] As Figure 2 shown, the risk management of this embodiment includes operations S210 to S240. This transaction processing method can be executed by a processor or any electronic device including a processor.
[0056] In operation S210, credit request data is obtained.
[0057] According to an embodiment of the present disclosure, the credit request data includes a user identifier. The user identifier may be a user identity ID, for example, a user ID number. It should be understood that a user may submit a loan application through mobile banking, bank branches, government financial service windows, and other channels, and when submitting a loan application, a bank server may obtain the credit request data. In an embodiment of the present disclosure, before obtaining the user's information, the user's consent or authorization may be obtained. For example, before operation S210, a request to obtain the user's information may be issued to the user. When the user agrees or authorizes that the user's information may be obtained, the operation S210 is performed.
[0058] In operation S220, corresponding user characteristic data is acquired based on the user identification.
[0059] According to an embodiment of the present disclosure, after obtaining the user identifier in the credit request data, the bank server can obtain the corresponding user feature data according to the user identifier. There is a mapping relationship between the user feature data and the user identifier. The user feature data can be the basic information data of the user held by the bank system and the financial asset data held by the user.
[0060] In operation S230, the user feature data is input into a pre-trained first risk federation model to obtain a risk assessment result.
[0061] According to an embodiment of the present disclosure, the acquired user feature data can be used to evaluate the customer's repayment ability. Furthermore, the user feature data can be used to input the pre-trained first risk federation model to evaluate the user's default probability. In an embodiment of the present disclosure, the first risk federation model is trained based on a federated learning algorithm, and the training is jointly performed by a bank server and a big data bureau server, wherein the data used for training includes collateral sample data, and the collateral sample data is obtained from the big data bureau server. After inputting the user feature data, the first risk federation model can be linked with the second risk federation model deployed in the big data bureau server to obtain a risk assessment result. Since the collateral sample data in the big data bureau server is used in the process of training the first risk federation model, the correlation between the model and the user's collateral information can be improved, so that the bank can have a more comprehensive understanding of the customer's collateral situation, thereby making a loan access decision. It avoids the situation where the evaluation price of the evaluation company is higher than the actual price, and credit risk management is carried out through a more reasonable, scientific and effective evaluation method to reduce the risk of commercial banks lending.
[0062] In operation S240, credit request feedback data and a credit whitelist are generated based on the risk assessment result.
[0063] According to an embodiment of the present disclosure, the bank server may complete loan access based on the risk assessment result, and generate credit request feedback data to feedback to the user. Further, a credit white list may also be generated. The credit white list contains pre-approved customer information for the business department to save and manage.
[0064] In some specific embodiments, the user feature data includes user attribute information, user asset information, and product attribute information. Among them, the user attribute information may include basic customer information, such as gender, age, region, occupation, education level, banking age, etc. The user asset information may include deposit balance, credit card limit, fund balance, wealth management fund balance, etc. The product attribute information may include information such as the product situation of bank financial products and fund products purchased by the user.
[0065] In some specific embodiments, the pre-trained first risk federated model is automatically updated based on a preset time period. The preset time period may be one day / one week / one month, and the preset time period can be flexibly adjusted based on the data volume of the bank system, the speed of data update, and the risk assessment requirements. In some examples, in order to ensure the accuracy of the data, one day can be used as the preset time period. For example, at a fixed time every day (such as 24:00), the latest user feature sample data can be incorporated to perform self-learning update training on the model to improve the accuracy of data processing. It should be understood that the self-learning update training is based on federated learning. The first risk federated model deployed on the bank server and the second risk federated model deployed on the big data bureau server synchronously update the model parameters.
[0066] In some specific embodiments, after generating the credit white list, the method further includes the step of synchronously storing the credit white list.
[0067] Figure 3 The flowchart of the method for synchronously storing the credit white list according to an embodiment of the present disclosure is schematically shown.
[0068] As Figure 3 shown, the risk management of this embodiment includes operation S310.
[0069] In operation S310, the credit white list is stored and the credit white list is sent to the big data bureau server.
[0070] According to specific embodiments of the present disclosure, after obtaining the credit white list, the banking system can store it and can also send the credit white list to the big data bureau server. Currently, the big data bureau server is usually deployed in the government affairs system. Government departments can access it through the API interface after being authorized by the big data bureau to access data. The big data platform can collect the government affairs data of each government department through system direct connection. It should be understood that each government department has accumulated a large amount of data of individuals and enterprises in various dimensions, but it is difficult for the big data bureau to collect the credit data of individuals and enterprises. According to the method of the specific embodiments of the present disclosure, the big data bureau server can obtain richer user credit information, further enrich and improve the database, achieve information sharing with the bank, and realize a win-win situation.
[0071] In an embodiment of the present disclosure, the first risk federated model is trained based on federated learning.
[0072] Another embodiment of the present disclosure provides a risk model training method based on federated learning.
[0073] Figure 4 Schematically shows a flowchart of a risk model training method based on federated learning according to another embodiment of the present disclosure.
[0074] As Figure 4 shown, the risk management of this embodiment includes operations S410 to S440.
[0075] In operation 410, based on the asymmetric encryption algorithm, obtain the intersection identification data of the first user identification data and the second user identification data. Among them, the first user identification data is obtained from the bank server, and the second user identification data is obtained from the big data bureau server.
[0076] According to an embodiment of the present disclosure, a method of vertical federated learning is used to train a model. Specifically, vertical federated learning is used when there are many overlapping users but few overlapping data features. After aligning the user identifiers, data with the same user identifiers but different features in the bank server and the big data bureau server are used for modeling. The data held by the bank server and the data held by the big data bureau server remain stored locally, and no raw data is exchanged during the modeling and model running processes. Two sets of model parameters are obtained and held separately by the bank and the big data bureau and used jointly to meet the requirements of data privacy protection for the bank and the big data bureau. It can be seen that before model training, the user identifiers required by the bank server need to be aligned with the user identifiers in the big data bureau server. User identifier alignment aims to complete the intersection calculation of user identifier data on the premise of protecting the data privacy of the participating parties (i.e., the bank server and the big data bureau server). After the calculation is completed, one or more of the participating parties can only obtain the correct intersection of the multi-party data sets, and will not obtain any information of other participating parties outside the intersection. In the embodiment of the present disclosure, an asymmetric encryption method is used to align the user identifiers of the bank server and the big data bureau server. The encryption process is irreversible, and no underlying data is leaked to the other party. Among them, the user identifiers required by the bank server are the first user identifier data, and the user identifiers included in the big data bureau server are the second user identifier data. After aligning the first user identifier data and the second user identifier data, intersection identifier data can be obtained.
[0077] Figure 5 Schematically shows the principle diagram of the method for aligning the first user identifier data and the second user identifier data.
[0078] As Figure 5 shown, Party A (Participant A) is the big data bureau server, and Party B (Participant B) is the bank server. In the data set X A of Participant A, it contains the second user identifiers u1, u2, u3, u4. In the data set X B of Participant B, it contains the first user identifiers u1, u2, u3, u5. The goal of alignment is to obtain the intersection user identifiers u1, u2, u3 without leaking any information about the second user identifier u4 and the first user identifier u5. For Participant A, the confidentiality operation consists of a hash mechanism and a randomly generated random number (ri). For Participant B, the confidentiality operation is implemented by a hash mechanism and its own generated private key (d).
[0079] A typical alignment process is as follows:
[0080] (1) Participant B generates n, e (public key), d (private key) through an asymmetric encryption algorithm (RSA algorithm), and sends the public key (n, e) to Participant A.
[0081] (2) Participant A encrypts its own data, uses the random number ri and the hashing mechanism to encrypt the data of users u1, u2, u3, u4, and sends the encrypted data YA to Participant B.
[0082] (3) After Participant B obtains YA, due to the principle of the hashing mechanism and the unknown ri, it is very difficult to reverse-engineer the user data of Participant A. Participant B takes the d-th power of YA to get ZA, and then encrypts its own user data through the hashing machine combined with the private key d to get ZB, and sends ZA and ZB to Participant A.
[0083] (4) After Participant A obtains ZB, it is also unable to reverse-engineer the user data of Participant B. When dividing the encrypted data ZA of its own user by ri and performing hashing processing at the same time, DA is obtained.
[0084] The essence of DA and ZB is the data obtained after performing the same operation on the data. Therefore, if the data sources are the same, the data obtained after the operation is also the same. Thus, based on the result of finding the intersection of DA and ZB, Participant A can determine what the common data of Participant A and Participant B are, and then send the result I to B, thereby completing the alignment of user identities and obtaining the intersection user identities u1, u2, u3.
[0085] In operation 420, initialize the first federated model parameters deployed on the bank server and the second federated model parameters deployed on the big data bureau server.
[0086] In the embodiments of the present disclosure, both the bank server and the big data bureau server are deployed with federated models, where the bank server is deployed with a first federated model and the big data bureau server is deployed with a second federated model. Before starting training, the parameters of the first federated model and the parameters of the second federated model can be initialized respectively. During the model training process, the two sets of model parameters are trained simultaneously, held by the bank server and the big data bureau server respectively, and jointly used during application.
[0087] In operation 430, train and update the first federated model parameters and the second federated model parameters based on the homomorphic encryption algorithm until a preset training cutoff condition is reached.
[0088] In the embodiments of the present disclosure, the method of homomorphic encryption is used to train and update the first federated model parameters and the second federated model parameters. Homomorphic encryption is a cryptographic technology based on the computational complexity theory of mathematical problems. Two ciphertexts obtained using the same homomorphic encryption algorithm can be added or multiplied without decrypting, and the result is the same as the result of directly adding or multiplying in the plaintext state and then encrypting.
[0089] Figure 6Schematically shows the schematic diagram of the homomorphic encryption algorithm.
[0090] As Figure 6 shown, a and b are the data of user a and user b respectively, c is the intermediate result data for passing from the relevant party to the irrelevant party, and op is the operator. Among them, the relevant party holds the plaintext data, and the plaintext intermediate data c is obtained by operating on the plaintext data of user a and the plaintext data of user b. Enc(a) is the ciphertext data obtained by subjecting the plaintext data of user a to the homomorphic encryption algorithm, Enc(b) is the ciphertext data obtained by subjecting the plaintext data of user b to the same homomorphic encryption algorithm, and op’ is the operator for the ciphertext data Enc(a) and Enc(b). Enc(a) and Enc(b) can be directly operated to obtain the ciphertext intermediate result Enc(c), and Enc(c) is the same as the result after encrypting the plaintext c. Thus, the privacy protection of the relevant party's data can be realized.
[0091] According to an embodiment of the present disclosure, the first federated model and the second federated model are jointly trained based on the first feature data and the second feature data, and the first feature data and the second feature data are respectively obtained by the bank server and the big data bureau server based on the intersection identification data. For example, after obtaining the intersection identification data, the user identification alignment between the bank server and the big data bureau server is completed. The user feature data having a mapping relationship with the intersection identification data on the bank side can be matched, that is, the first feature data. Similarly, the user feature data having a mapping relationship with the intersection identification data on the big data bureau side can be matched, that is, the second feature data. Among them, the first feature data includes user credit sample data, and the user credit sample data includes user feature sample data and user credit labels. The second feature data includes collateral sample data. The user credit label can be a label indicating whether the user is a good user. Since the second feature data includes collateral sample data, the collateral dimension can be mainly considered to confirm the user label, so as to realize a more objective, scientific, reasonable, and effective evaluation of the impact of the collateral on the credit risk. It should be understood that in the construction of the federated learning model, since the big data bureau side does not have the label of whether the past loan behavior is approved, it cannot independently construct the model. After cooperating with the bank, the bank side provides the label, so that both cooperating parties can complete the training of the federated model. The bank side reduces the credit risk, and the big data bureau side enriches the information of customer credit, achieving a win-win situation.
[0092] In an embodiment of the present disclosure, the model training can be terminated after the loss function converges. The training termination condition can be preset in advance. For example, the number of model iterations can be preset, and when the number is reached, the model training is terminated. It is also possible to preset to stop the model training when the model recognition accuracy reaches a certain threshold.
[0093] In operation 440, the first risk federated model and the second risk federated model are obtained.
[0094] According to an embodiment of the present disclosure, when the training is terminated, the first federated model parameters and the second model parameters corresponding to the training termination time can be obtained, whereby the first risk federated model and the second risk federated model can be determined. Among them, the first risk federated model is the first federated model including the first federated model parameters at the end of training, and the second risk federated model is the second federated model including the second federated model parameters at the end of training. It should be understood that the first risk federated model is deployed on the bank server and is used to input user feature data when applying the model. The second risk federated model is deployed on the big data bureau server and is used to jointly operate with the first risk federated model to obtain a risk assessment result when applying the model.
[0095] Figure 7 Schematically shows an exemplary system architecture of a method and apparatus for training a risk model according to another embodiment of the present disclosure. It should be noted that Figure 7 The shown is only an example of a system architecture to which another embodiment of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0096] As Figure 7 shown, the system architecture 700 according to this embodiment may include a bank server 105, a network 104, a big data bureau server 106, and a third-party server 107. The network 104 is used to provide a medium for communication links between the bank server 105, the big data bureau server 106, and the third-party server 107. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0097] The third-party server 107 can create a key pair and send the public key to the bank server 105 and the big data bureau server 106 respectively. The private key is held by the third-party server 107. Among them, the public key is used for the homomorphic encryption of the intermediate results used for interactive transmission when the bank server 105 and the big data bureau server 106 jointly train the first federated model and the second federated model. The bank server 105 can obtain the encrypted intermediate result a from the big data bureau server 106 and train the first federated model based on the intermediate result a and the first feature data held by itself to obtain the combined data a. The big data bureau server 106 can obtain the encrypted intermediate result b from the bank server and train the second federated model based on the intermediate result b and the second feature data held by itself to obtain the combined data b. The bank server 105 and the big data bureau server 106 respectively send the combined data a and the combined data b to the third-party server 107. The third-party server 107 can use the held private key to decrypt the combined data a and b, and then send the decryption result of the combined data a to the bank server 105 and the decryption result of the combined data b to the big data bureau server 106 respectively, so that the bank server 105 and the big data bureau server 106 can update their respective model parameters based on the decryption results.
[0098] Figure 8 FIG. schematically shows a flowchart of a method for updating the first federated model parameters and the second federated model parameters based on the homomorphic encryption algorithm according to another embodiment of the present disclosure.
[0099] As Figure 8 shown, the risk management of this another embodiment includes operations S810 to S860.
[0100] In operation S810, the bank server and the big data bureau server respectively train the first federated model and the second federated model based on the first feature data and the second feature data, and calculate the first transfer data and the second transfer data.
[0101] In the embodiments of the present disclosure, the bank server trains the first federated model based on the first feature data to obtain the intermediate result on the bank side - the first transfer data. The big data bureau server trains the second federated model based on the second feature data to obtain the intermediate result on the big data bureau side - the second transfer data. It should be noted that before training, the bank server and the big data bureau server respectively obtain the public key from the third-party server for homomorphic encryption of the first transfer data and the second transfer data.
[0102] In operation S820, the bank server and the big data bureau server respectively use the public key to perform homomorphic encryption on the first transfer data and the second transfer data, and interact and transfer the first encrypted transfer data and the second encrypted transfer data after homomorphic encryption to each other.
[0103] In operation S830, the bank server trains the first federated model based on the second encrypted transfer data to obtain the first encrypted gradient information and the first encrypted loss information. The big data bureau server trains the second federated model based on the first encrypted transfer data to obtain the second encrypted gradient information. It should be noted that when interacting with the intermediate results for the first time, both the bank server and the big data bureau server can only perform calculations using their respective first feature data and second feature data. And in the process of obtaining and updating the gradient information and loss information, in addition to applying the second encrypted transfer data, the bank server will also utilize its own first feature data, so that the gradient information and loss information can be obtained and updated by combining the data of both parties in federated learning. Similarly, in addition to applying the first encrypted transfer data, the big data bureau server will also utilize its own second feature data, so that the gradient information can be obtained and updated by combining the data of both parties in federated learning, in order to achieve the purpose of fully utilizing the data of both parties for model training without revealing privacy.
[0104] In operation S840, the bank server and the big data bureau server respectively send the first encrypted gradient information, the first encrypted loss information, and the second encrypted gradient information to the third-party server.
[0105] In operation S850, the third-party server decrypts the first encrypted gradient information, the first encrypted loss information, and the second encrypted gradient information based on the held private key, obtains and sends the first decrypted gradient information and the first decrypted loss information to the bank server, and obtains and sends the second decrypted gradient information to the big data bureau server.
[0106] In operation S860, the bank server updates the first federated model parameters based on the first decrypted gradient information and the first decrypted loss information, and the big data bureau server updates the second federated model parameters based on the second decrypted gradient information.
[0107] It should be noted that, in order to further protect data privacy, when the bank server sends the first encrypted gradient information and the first encrypted loss information to the third-party server, an additional mask can be added. Similarly, when the big data bureau server sends the second encrypted gradient information to the third-party server, an additional mask can also be added. Thus, when the bank server obtains the first decrypted gradient information and the first decrypted loss information from the third-party server, and when the big data bureau server obtains the second decrypted gradient information from the third-party server, the additional masks can be removed and decrypted again respectively to obtain the gradient information and / or loss information that can be used to update the first federated model parameters and the second federated model parameters respectively.
[0108] According to an embodiment of the present disclosure, before obtaining the intersection identification data, the method further includes a user screening step.
[0109] Figure 9 A flowchart of a method for screening users according to an embodiment of the present disclosure is schematically shown.
[0110] As Figure 9 shown, the risk management of this embodiment includes operation S910.
[0111] In operation S910, training set users and validation set users are obtained.
[0112] In the embodiments of the present disclosure, emphasis is placed on examining the impact of collateral on credit risk. When screening users, the user groups are mainly divided based on whether the collateral approval is passed and whether the loan is repaid on time after the approval is passed. Among them, the training set users include positive sample users and negative sample users. The positive sample users include users who apply for collateral loans and are approved and repay the loans on time within the first preset time period. The negative sample users include users who apply for collateral loans within the first preset time period but whose collateral approvals are not passed, or users who are approved for collateral loans but do not repay the loans on time. The number of positive sample users and the number of negative sample users are set based on a preset ratio. The validation set users include users who apply for collateral loans within the second preset time period. In some specific embodiments, the first preset time period can be one month, one quarter, or half a year. For users who apply for collateral loans and are approved, the repayment data of the users within one year can be obtained to evaluate whether the users repay the loans on time, so as to distinguish positive and negative sample users. The second preset time period can be the same as the first preset time period, and all users who apply for collateral loans within the second preset time period can be included in the scope of the validation set users. In a specific example, the positive sample users can be users who applied for collateral loans in January 2020, whose collateral was approved by the bank side, and who repaid the loans on time between January and December 2020. The negative sample users can be users who applied for collateral loans in January 2020 but whose collateral was not approved by the bank side and the bank did not grant the loans, and users who did not repay the loans on time between January and December 2020. The random sampling method can be used for negative sample users, and the ratio of the number of positive sample users to the number of negative sample users is set to 1:3. The validation set users can be users who applied for collateral loans in January 2021. It should be understood that for the validation set users, data on whether their collateral approvals are passed and whether they repaid the loans on time between January and December 2020 can be obtained. In a specific example of the present disclosure, through experimental verification, it can be more accurately predicted whether a user can maintain a good credit record based on the user's repayment situation within one year. By screening the above sample users, the association between collateral data and user credit risk can be established scientifically, reasonably, and accurately, effectively reducing the risk of bank credit.
[0113] In some embodiments, the collateral sample data includes at least one of movable collateral data, immovable collateral data, and intangible asset collateral data. Specifically, collateral data information including real estate, land use rights, transportation equipment, resource assets, intangible assets, long-term investments, current assets, and charging rights of users can be extracted from the big data bureau. For example, for real estate collateral, the location information can be intercepted according to the address information provided by the user, and the average price of houses with similar square meters in the location can be calculated as the valuation of real estate collateral. For collateral of land use rights type, the average of land use fees of similar locations and similar sizes can be statistically averaged as collateral valuation. For transportation equipment collateral, similar equipment of the same model and production year can be selected for comparison, and the average value can be taken as collateral valuation. It should be understood that after completing the user sample alignment, the user feature data and collateral sample data can be cleaned by feature engineering, including but not limited to mean filling, abnormal sample deletion, giving feature value meaning (such as converting the account opening date to the account opening period), normalizing the data amount, the amount of occurrence and other fields, and unifying the measurement to solve the problem of too large a gap in the amount feature.
[0114] In some embodiments, the first risk model and the second risk model are constructed based on the XGBoost algorithm. Xgboost is a tool for large-scale parallel Boosted Tree (boosting number algorithm), and is currently the fastest and best open source Boosted Tree toolkit. The Xgboost toolkit is more than 10 times faster than common toolkits, with very fast model training speed and prediction accuracy. In the big data classification scenario of the embodiment of the present disclosure, the use of the XgBoost algorithm is conducive to improving the model training speed and prediction accuracy. Based on experimental verification, the better model setting parameters in the embodiment of the present disclosure are shown in Table 1. The parameters set in this way further improve the accuracy of the model prediction of the embodiment of the present disclosure.
[0115]
[0116] Table 1
[0117] Based on the above risk management method, the present disclosure also provides a risk management device. Figure 10 The device is described in detail.
[0118] Figure 10 The structure block diagram of the risk management device according to the embodiment of the present disclosure is schematically shown.
[0119] like Figure 10 As shown, the risk management device 1000 of this embodiment includes a first acquisition module 1010 , a second acquisition module 1020 , a calculation module 1030 and a generation module 1040 .
[0120] Among them, the first acquisition module 1010 is configured to acquire credit request data, where the credit request data includes a user identifier.
[0121] The second acquisition module 1020 is configured to acquire corresponding user feature data based on the user identifier.
[0122] The calculation module 1030 is configured to input the user feature data into a pre-trained first risk federated model to obtain a risk assessment result, where the first risk federated model is trained based on a federated learning algorithm, and the training is jointly executed by a bank server and a big data bureau server. The data used for training includes collateral sample data, and the collateral sample data is obtained from the big data bureau server.
[0123] The generation module 1040 is configured to generate credit request feedback data and a credit white list based on the risk assessment result.
[0124] Figure 11 Schematically shows a structural block diagram of a risk model training system according to another embodiment of the present disclosure.
[0125] As Figure 11 shown, the risk model training system 1100 of this embodiment includes an alignment device 1110, an initialization device 1120, a calculation device 1130, and a generation device 1140.
[0126] Among them, the alignment device 1110 is configured to obtain intersection identification data of first user identification data and second user identification data based on an asymmetric encryption algorithm, where the first user identification data is obtained from a bank server, and the second user identification data is obtained from the big data bureau server.
[0127] The initialization device 1120 is configured to initialize the first federated model parameters deployed on the bank server and the second federated model parameters deployed on the big data bureau server.
[0128] The calculation device 1130 is configured to perform training updates on the first federated model parameters and the second federated model parameters based on a homomorphic encryption algorithm until a preset training cut-off condition is reached. The first federated model and the second federated model are jointly trained based on first feature data and second feature data. The first feature data and the second feature data are respectively obtained by the bank server and the big data bureau server based on the intersection identification data. The first feature data includes user credit sample data, and the user credit sample data includes user feature sample data and user credit labels. The second feature data includes collateral sample data.
[0129] The generating device 1140 is configured to obtain a first risk federated model and a second risk federated model, where the first risk model is the first federated model including the first federated model parameters at the end of training, and the second risk model is the second federated model including the second federated model parameters at the end of training.
[0130] According to an embodiment of the present disclosure, any multiple of the first obtaining module 1010, the second obtaining module 1020, the calculating module 1030, and the generating module 1040 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the first obtaining module 1010, the second obtaining module 1020, the calculating module 1030, and the generating module 1040 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the first obtaining module 1010, the second obtaining module 1020, the calculating module 1030, and the generating module 1040 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0131] Figure 12 A block diagram of an electronic device suitable for implementing a risk management method and a risk model training method according to an embodiment of the present disclosure is schematically shown.
[0132] As Figure 12 shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 901 may also include on-board memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0133] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0134] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, and the input / output (I / O) interface 905 is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the I / O interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage portion 908 as needed.
[0135] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0136] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the ROM 902 and / or RAM 903 and / or ROM 902 and RAM 903 described above.
[0137] An embodiment of the present disclosure further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present disclosure.
[0138] When the computer program is executed by the processor 901, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0139] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 909, and / or be installed from the removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0140] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or be installed from the removable medium 911. When the computer program is executed by the processor 901, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0141] In accordance with embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0143] Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0144] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A risk management method, applied to a bank server, characterized in that Including: Obtain credit request data, where the credit request data contains a user identifier; Obtain corresponding user feature data based on the user identifier; Input the user feature data into a pre-trained first risk federated model to obtain a risk assessment result; and Generate credit request feedback data and a credit white list based on the risk assessment result, where the first risk federated model is trained based on a federated learning algorithm, and the training is jointly executed by a bank server and a big data bureau server. The data used for training includes collateral sample data, and the collateral sample data is obtained from the big data bureau server. where the training includes: Obtain intersection identifier data of first user identifier data and second user identifier data based on an asymmetric encryption algorithm, where the first user identifier data is obtained from the bank server, and the second user identifier data is obtained from the big data bureau server; Initialize the first federated model parameters deployed on the bank server and the second federated model parameters deployed on the big data bureau server; Perform training updates on the first federated model parameters and the second federated model parameters based on a homomorphic encryption algorithm until a preset training cut-off condition is reached. The first federated model and the second federated model are jointly trained based on first feature data and second feature data. The first feature data is user feature data on the bank server side that has a mapping relationship with the intersection identifier data, and the second feature data is user feature data on the big data bureau server side that has a mapping relationship with the intersection identifier data. The first feature data includes user credit sample data, and the user credit sample data includes user feature sample data and user credit labels. The second feature data includes collateral sample data; and Obtain the first risk federated model and the second risk federated model, where the first risk federated model is the first federated model containing the first federated model parameters at the end of training, and the second risk federated model is the second federated model containing the second federated model parameters at the end of training.
2. The method according to claim 1, wherein, The user feature data includes user attribute information, user asset information, and product attribute information.
3. The method according to claim 1, wherein The pre-trained first risk federated model is automatically updated based on a preset time period.
4. The method according to claim 1, wherein, After generating the credit white list, the method further includes: Store the credit white list and send the credit white list to the big data bureau server.
5. A risk model training method based on federated learning, characterized in that, Including: Obtain intersection identifier data of first user identifier data and second user identifier data based on an asymmetric encryption algorithm, where the first user identifier data is obtained from the bank server, and the second user identifier data is obtained from the big data bureau server; Initialize the first federated model parameters deployed on the bank server and the second federated model parameters deployed on the big data bureau server; Training and updating the first and second federated model parameters based on a homomorphic encryption algorithm until a preset training cutoff condition is reached, where the first and second federated models are jointly trained based on first and second feature data. The first feature data is user feature data on the bank server side that has a mapping relationship with the intersection identification data, and the second feature data is user feature data on the big data bureau server side that has a mapping relationship with the intersection identification data. The first feature data includes user credit sample data, and the user credit sample data includes user feature sample data and user credit labels. The second feature data includes collateral sample data; and Obtaining a first risk federated model and a second risk federated model, where the first risk federated model is the first federated model containing the first federated model parameters at the end of training, and the second risk federated model is the second federated model containing the second federated model parameters at the end of training.
6. The method according to claim 5, wherein The training and updating of the first and second federated model parameters based on the homomorphic encryption algorithm until a preset training cutoff condition is reached includes: The bank server and the big data bureau server respectively train the first and second federated models based on the first and second feature data, and calculate first transfer data and second transfer data; The bank server and the big data bureau server respectively perform homomorphic encryption on the first transfer data and the second transfer data, and interactively transfer the first encrypted transfer data and the second encrypted transfer data after homomorphic encryption. The public key used in the homomorphic encryption is obtained from a third-party server; The bank server trains the first federated model based on the second encrypted transfer data to obtain first encrypted gradient information and first encrypted loss information, and the big data bureau server trains the second federated model based on the first encrypted transfer data to obtain second encrypted gradient information; The bank server and the big data bureau server respectively send the first encrypted gradient information, the first encrypted loss information, and the second encrypted gradient information to the third-party server; The third-party server decrypts the first encrypted gradient information, the first encrypted loss information, and the second encrypted gradient information based on the held private key, obtains and sends the first decrypted gradient information and the first decrypted loss information to the bank server, and obtains and sends the second decrypted gradient information to the big data bureau server; and The bank server updates the first federated model parameters based on the first decrypted gradient information and the first decrypted loss information, and the big data bureau server updates the second federated model parameters based on the second decrypted gradient information.
7. The method according to claim 5, wherein, Before obtaining the intersection identification data, the method further includes a user screening step, including: Obtain training set users and validation set users. Among them, the training set users include positive sample users and negative sample users. The positive sample users include users who applied for collateral loans and were approved and repaid on time within a first preset time period. The negative sample users include users who applied for collateral loans within a first preset time period but whose collateral approvals were not passed, or users who applied for collateral loans and were approved but did not repay on time. The number of positive sample users and the number of negative sample users are set based on a preset ratio. The validation set users include users who applied for collateral loans within a second preset time period.
8. The method according to claim 5, wherein The first risk federated model and the second risk federated model are constructed based on the XGBoost algorithm.
9. The method according to claim 5, wherein The collateral sample data includes at least one of movable property collateral data, immovable property collateral data, and intangible asset collateral data.
10. A risk management device, deployed in a bank server, characterized in that, It includes: A first acquisition module configured to acquire credit request data, where the credit request data includes a user identifier; A second acquisition module configured to acquire corresponding user feature data based on the user identifier; A calculation module configured to input the user feature data into a pre-trained first risk federated model to obtain a risk assessment result. Among them, the first risk federated model is trained based on the federated learning algorithm, and the training is jointly executed by the bank server and the big data bureau server. The data used for training includes collateral sample data, and the collateral sample data is obtained from the big data bureau server; and A generation module configured to generate credit request feedback data and a credit white list based on the risk assessment result, Among them, the training includes: Obtain the intersection identifier data of the first user identifier data and the second user identifier data based on the asymmetric encryption algorithm. Among them, the first user identifier data is obtained from the bank server, and the second user identifier data is obtained from the big data bureau server; Initialize the first federated model parameters deployed on the bank server and the second federated model parameters deployed on the big data bureau server; Based on the homomorphic encryption algorithm, train and update the first federated model parameters and the second federated model parameters until a preset training cut-off condition is reached. Among them, the first federated model and the second federated model are jointly trained based on the first feature data and the second feature data. The first feature data is the user feature data on the bank server side that has a mapping relationship with the intersection identifier data, and the second feature data is the user feature data on the big data bureau server side that has a mapping relationship with the intersection identifier data. The first feature data includes user credit sample data, and the user credit sample data includes user feature sample data and user credit labels. The second feature data includes collateral sample data; and Obtain the first risk federated model and the second risk federated model. Among them, the first risk federated model is the first federated model including the first federated model parameters at the end of training, and the second risk federated model is the second federated model including the second federated model parameters at the end of training.
11. A risk model training system, including: An alignment device configured to obtain intersection identification data of first user identification data and second user identification data based on an asymmetric encryption algorithm, wherein the first user identification data is obtained from a bank server, and the second user identification data is obtained from a big data bureau server; An initialization device configured to initialize first federated model parameters deployed on the bank server and second federated model parameters deployed on the big data bureau server; A calculation device configured to perform training and updating on the first federated model parameters and the second federated model parameters based on a homomorphic encryption algorithm until a preset training cut-off condition is reached, wherein a first federated model and a second federated model are jointly trained based on first feature data and second feature data, the first feature data is user feature data on the bank server side that has a mapping relationship with the intersection identification data, the second feature data is user feature data on the big data bureau server side that has a mapping relationship with the intersection identification data, the first feature data includes user credit sample data, and the user credit sample data includes user feature sample data and user credit labels, and the second feature data includes collateral sample data; and A generation device configured to obtain a first risk federated model and a second risk federated model, wherein the first risk federated model is the first federated model including the first federated model parameters at the end of training, and the second risk federated model is the second federated model including the second federated model parameters at the end of training.
12. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 9.
13. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to execute the method according to any one of claims 1 to 9.
14. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Enterprise risk evaluation method, device and equipment and computer readable storage medium
CN113011632A