Three-agriculture customer data risk control modeling method based on federal learning
Through the risk control modeling method of rural rural customers based on federal learning, the problems of insufficient data and data privacy protection of rural rural customers were solved, effective privacy protection, data silos were broken and communication expenses were reduced, and loan acquisition for rural rural customers was improved.
Patent Information
- Application Number
- CN202311479587.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-11-08
AI Technical Summary
There is less financial data for rural and rural customers, and a single institution is difficult to support risk control modeling, and financial institutions are difficult to evaluate the credit status of rural and rural customers, resulting in loan difficulties, and there is a risk of data leakage in federal learning technology.
The data risk control modeling method of rural and rural customers based on federated learning is adopted. Through data preprocessing, dimensionality reduction processing, model initialization and federated learning training, logistic regression is selected as the training model, and the training is combined with synchronous federated and asynchronous federated methods to ensure data privacy and reduce communication overhead.
Effective privacy protection, breaking data barriers, reducing communication overhead, improving model reusability and anti-interference capabilities, and solving the problem of loan difficulties for rural and rural customers.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data risk control technology, and specifically to a method for data risk control modeling of agriculture, rural areas and farmers based on federated learning. Background Art
[0002] Federated learning is an advanced machine learning method that allows model training in a distributed data environment without centralizing the data in one place. This is very important for protecting privacy, reducing data transmission costs, and complying with data regulations.
[0003] In the field of agriculture, rural areas and farmers, a large amount of farmer and agricultural data is involved. These data are often distributed in different regions or managed by different institutions, such as agricultural cooperatives, agricultural enterprises, etc. At the same time, agricultural data may contain a lot of sensitive information, such as farmers' personal information, land area, crop varieties, etc., which need to be properly protected.
[0004] Logistic regression is a statistical learning method for classification problems. It can be used to predict the probability of an event. In the field of financial risk control, logistic regression can be applied to tasks such as predicting whether a customer will fulfill their contract.
[0005] The above technology still has the following problems: 1. The financial data generated by rural customers is relatively small, and the information is relatively congested. The amount of data from rural customers of a single institution is difficult to support risk control modeling.
[0006] 2. It is difficult for banks and other financial institutions to use their existing data to assess the credit status of rural customers, resulting in difficulties for rural customers to obtain loans. Even if they pass the assessment, the amount that can be loaned is very low.
[0007] 3. Financial institutions have blacklist data of customers and have more basic information data of rural customers. However, banks and other financial institutions have strict data management regulations, which makes it difficult for the two institutions to exchange data, resulting in the problem of data islands.
[0008] 4. Federated learning technology is still developing rapidly, but there is still the problem of original data being cracked through reverse engineering. If the parties use the original data for training, there is still a risk of data leakage. Summary of the invention
[0009] In view of the deficiencies of the prior art, the present invention is implemented through the following technical solutions: a method for risk control modeling of rural customer data based on federated learning, comprising the following steps: Step 1: Determine the dimensions of data indicators involved in training based on the data sources of all parties involved in federated learning.
[0010] Step 2: Preprocess the data and agree on the format of training data. Each participant preprocesses their own data, including data cleaning, missing value processing, feature selection, etc. For discrete variables, random forest algorithm and K nearest neighbor algorithm are used alternately to fill in, and for continuous variables, mean value is used to fill in, to ensure that the data of each party is relatively consistent in terms of features.
[0011] Step 3: All parties need to perform dimensionality reduction on the data. For data with small differences in the same column, low variance filtering is used for processing, and reverse feature elimination is performed on regular data. The processed data is divided, 85% of which is used for training, and each party retains 15% of the data for verification.
[0012] Step 4: The participants negotiate the specific agreement on model training. Logistic regression is selected as the training model, and hyperparameters such as training cycle and learning rate are determined according to the amount of data.
[0013] Step 5: Participants initialize the model.
[0014] Step 6: Federated learning training uses a combination of synchronous federation and asynchronous federation to ensure that all participants have enough time to provide data for model training. The model training steps include: local model update, model parameter aggregation, model aggregation, and feedback update.
[0015] Step 7: Repeat the training process in step 6 until the predetermined number of training rounds is reached or the loss function has stabilized for a long time.
[0016] Step 8: Evaluate and verify the model. Use the reserved test data or verify in the actual scenario to evaluate the performance of the model. If the accuracy is higher than 96%, use this model. Otherwise, increase the training cycle in step 4 by 10% each time. If the accuracy is still lower than 96%, reduce the learning rate parameter in step 4 by / 10 each time. Repeat the training until the accuracy reaches 96%.
[0017] The present invention provides a data risk control modeling method for agriculture, rural areas and farmers based on federated learning. It has the following beneficial effects: 1. Effectively protect privacy. Farmers’ personal information is often sensitive information. Federated learning can be used to train models without transmitting data to a central server, thereby protecting data privacy.
[0018] 2. Break down data barriers and solve the problem of data silos. The data distribution of different institutions may be different. Federated learning allows model updates to be performed locally, avoiding the problem of centralizing data in one place.
[0019] 3. Reduced communication overhead: In federated learning, only the update information of model parameters will be transmitted to the central server instead of the original data, thus reducing the communication cost.
[0020] 4. Model reusability: Models trained using large data sets have good anti-interference capabilities and can be used repeatedly for a long time.
[0021] 5. During the data preprocessing process, use the agreed method to preprocess the data, change the original structure of the data and retain the features to the greatest extent. Do not use the original data for training, making it easier for the data party to accept and thus obtain more data sources. Implementation
[0022] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] The present invention provides a technical solution: according to the data source conditions of the parties participating in the federated learning, the dimensions of the data indicators involved in the training are determined.
[0024] Data preprocessing is performed, and the format of training data is agreed upon. Each participant preprocesses their own data, including data cleaning, missing value processing, feature selection, etc. For discrete variables, random forest algorithm and K nearest neighbor algorithm are used alternately to fill, and for continuous variables, mean value is used to fill, to ensure that the data of each party is relatively consistent in terms of features.
[0025] Each party needs to perform dimensionality reduction on the data. For data with small differences in the same column, use low variance filtering to process it, and perform reverse feature elimination on regular data. The processed data is divided, 85% of which is used for training, and each party retains 15% of the data for verification.
[0026] The parties negotiated a specific agreement on model training. Logistic regression was selected as the training model, and hyperparameters such as training cycle and learning rate were determined based on the amount of data.
[0027] The participants initialize the model. Effectively protect privacy. Farmers' personal information is often sensitive information. Federated learning can be used to train the model without transmitting data to the central server, thereby protecting data privacy.
[0028] Federated learning training uses a combination of synchronous federation and asynchronous federation to ensure that all participants have enough time to provide data for model training. The model training steps include: local model update, model parameter aggregation, model aggregation, and feedback update. Breaking down data barriers and solving the problem of data silos, the data distribution of different institutions may be different. Federated learning allows local model updates, avoiding the problem of centralizing data in one place.
[0029] Repeat the training process in step 6 until the predetermined number of training rounds is reached or the loss function has been stable for a long time. The reusability of the model, the model trained with a large data set has good anti-interference ability and can be used repeatedly for a long time.
[0030] During the data preprocessing process, the data is preprocessed using an agreed method to change the original structure of the data and retain the features to the greatest extent possible. The original data is not used for training, making it easier for the data party to accept and thus obtaining more data sources.
[0031] Perform model evaluation and verification. Use the reserved test data or verify in the actual scenario to evaluate the performance of the model. If the accuracy is higher than 96%, use this model. Otherwise, increase the training cycle in step 4 by 10% each time. If the accuracy is still lower than 96%, reduce the learning rate parameter in step 4 by / 10 each time. Repeat the training until the accuracy reaches 96%.
[0032] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0033] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for risk control modeling of rural customer data based on federated learning, characterized by: The following steps are included: Step 1: Determine the data indicator dimensions involved in training based on the data sources of all parties involved in federated learning; Step 2: Preprocess the data and agree on the format of training data. Each participant preprocesses their own data, including data cleaning, missing value processing, and feature selection. Step 3: All parties need to perform dimensionality reduction on the data. For data with small differences in the same column, low variance filtering is used for processing, and reverse feature elimination is performed on regular data. Step 4: The participants negotiate the specific agreement on model training; Step 5: Participants initialize the model; Step 6: Federated learning training, using a combination of synchronous federation and asynchronous federation to ensure that all participants have enough time to provide data for model training; Step 7: Repeat the training process in step 6 until the predetermined number of training rounds is reached or the loss function has been stable for a long time; Step 8: Perform model evaluation and validation, using reserved test data or validating in actual scenarios to evaluate the performance of the model.
2. According to claim 1, a method for risk control modeling of rural customer data based on federated learning is characterized by: In the step 2, the random forest algorithm and the K nearest neighbor algorithm are alternately used to fill in the discrete variables, and the mean is used to fill in the continuous variables, so as to ensure that the data of each party are relatively consistent in terms of features.
3. According to the method of risk control modeling of rural customer data based on federated learning in claim 1, it is characterized by: In step three, the processed data is divided, 85% of the data is used for training, and each party retains 15% of the data for verification.
4. According to claim 1, a method for risk control modeling of rural customer data based on federated learning is characterized by: In step 4, logistic regression is selected as the training model, and hyperparameters such as training cycle and learning rate are determined according to the amount of data.
5. According to claim 1, a method for risk control modeling of rural customer data based on federated learning is characterized by: The model training steps in step six include: local model update, model parameter aggregation, model aggregation, and feedback update.
6. According to the method of risk control modeling of rural customer data based on federated learning in claim 1, it is characterized by: If the accuracy rate in step eight is higher than 96%, this model is used. Otherwise, the training cycle in step four needs to be increased by 10% each time. If the accuracy rate is still lower than 96%, the learning rate parameter in step four is reduced by / 10 each time, and the training is repeated until the accuracy rate reaches 96%.
Citation Information
Patent Citations
Loan risk assessment method based on multi-source data federated learning
CN113240509A
Cross-industry data joint modeling method and system based on block chain and federal learning
CN113836809A
Crop variety yield prediction method and device based on distributed data
CN115564145A
Block chain and federated learning-based crop data yield prediction method
CN116341722A
Federated learning method and apparatus based on multi-source heterogeneous system
WO2021109647A1