Personalized credit scoring method in financial credit field
Through personalized credit scoring algorithms, combined with big data and machine learning technology, the shortcomings of traditional credit evaluation methods in individual differences and market changes are solved, more accurate credit evaluation and higher model adaptability are achieved, and the accuracy and security of loan approval are improved.
Patent Information
- Application Number
- CN202510086266.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional credit assessment methods cannot fully capture the individualized risks of borrowers and have weak ability to deal with individual differences and rapidly changing market conditions.
The personalized credit scoring algorithm is adopted, combined with big data analysis and machine learning technology, and through data acquisition and preprocessing, feature extraction, model construction and training, personalized fine-tuning and model evaluation and optimization, the accuracy and adaptability of credit evaluation are improved.
Significantly reduce the wrong loan rejection rate and over-credit rate in loan approval, enhance the adaptability and interpretability of the model, improve user experience and the competitiveness of financial institutions, and protect user privacy and data security.
Smart Images

Figure CN120336981A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of financial credit, and particularly relates to a personalized credit scoring method in the field of financial credit. Background Art
[0002] In the financial credit industry, traditional credit assessment methods mainly rely on limited historical transaction data and static pooling information. Although this model is stable, its ability to handle individual differences and rapidly changing market conditions is weak, and it cannot fully capture the personalized risks of borrowers. To solve this problem, the present invention proposes a personalized credit scoring algorithm that combines the latest big data analysis and machine learning technologies, aiming to provide more accurate credit assessments and assist financial institutions in making more intelligent loan decisions. Summary of the Invention
[0003] (I) Object of the Invention
[0004] In order to overcome the above deficiencies, the object of the present invention is to provide a personalized credit scoring method in the field of financial credit to solve the above technical problems.
[0005] (II) Technical Solution
[0006] To achieve the above object, the technical solution provided by the present application is as follows:
[0007] A personalized credit scoring method in the field of financial credit, the algorithm comprising the following steps:
[0008] S1 Data collection and preprocessing, including internal data integration and external data integration;
[0009] S2 Feature extraction, including statistical property analysis, trend index extraction, and diversified feature construction;
[0010] S3 Model construction and training, including basic model selection, cross-validation, feature selection and weight adjustment, and multi-model integration;
[0011] S4 Personalized fine-tuning, including stereotype and class customization, feedback learning mechanism, and risk dynamic mapping;
[0012] S5 Model evaluation and optimization, including accuracy testing, risk stability detection, interpretability analysis, and compliance auditing.
[0013] Preferably, in the said S1,
[0014] Internal data integration automatically collects customers' account information, repayment records, and credit reports, and performs structured processing;
[0015] External data integration utilizes the OAuth protocol to collect influence dimension data of social media performance and e-commerce transaction data through customer authorization, and uses the API to standardize data fields;
[0016] Data cleaning and feature engineering are also included in S1,
[0017] The data cleaning includes: detecting missing values and outliers by applying methods of data sampling and statistical analysis, and determining data filling or missing value imputation strategies; implementing data type verification, converting incorrect data types into the correct format, and filtering invalid data items;
[0018] The feature engineering includes: statistical property analysis, generating descriptive statistical data based on existing data and using it for the basic construction of features; extracting trend indicators, analyzing trend indicators of account activities, including the time series dynamics of account balance and debt increase or decrease; constructing diversified features, combining professional background, living habits to construct additional credit score influencing factors.
[0019] Preferably, in S2, statistical property analysis generates descriptive statistical data based on existing data, trend indicator extraction analyzes trend indicators of account activities, and diversified feature construction combines professional background, living habits or other external data to construct additional credit score influencing factors.
[0020] Preferably, in S3, the selection of the basic model includes the following steps:
[0021] A1 Select a decision tree as the basic model to construct an initial credit score model;
[0022] Use information gain to select the optimal splitting feature and splitting point, calculate the information gain of each feature, and select the feature with the largest information gain for splitting;
[0023] The calculation formula of information gain is:
[0024]
[0025] Among them, D is the data set, A is the feature, Entropy(D) is the entropy of the data set D, Values(A) is the value set of the feature A, and Dv is the subset of the data set D where the value of the feature A is v.
[0026] A2 Select a support vector machine (SVM);
[0027] For non-linearly separable data, map the data to a high-dimensional space through a kernel function to make it linearly separable in the high-dimensional space;
[0028] In the credit field, use the characteristics of borrowers as input, and classify borrowers into two categories: good credit and bad credit through the SVM model;
[0029] The SVM using a linear kernel function constructs the following classification hyperplane: ω T x + b = 0, where ω is the normal vector of the hyperplane, x is the input feature vector, and b is the bias term;
[0030] A3 selects a neural network as the basic model to construct an initial credit scoring model; constructs a multi-layer perceptron, including an input layer, a hidden layer, and an output layer. The input layer receives the features of the borrower, the hidden layer extracts features through non-linear transformation, and the output layer gives the prediction result of the credit score or default probability;
[0031] The calculation formula of the cross-entropy loss function is:
[0032]
[0033] where N is the number of samples, y i is the true label, is the predicted label.
[0034] Preferably, in S3, the cross-validation includes the following steps:
[0035] B1 Data partitioning:
[0036] Randomly partition the data set into a training set and a test set, and partition it according to a certain ratio. For cross-validation, further partition the training set into multiple subsets;
[0037] B2 Model training and evaluation:
[0038] For each type of basic model, the process of cross-validation is as follows:
[0039] Take turns selecting a subset as the validation set, and the remaining subsets as the training set;
[0040] Train the model on the training set and evaluate the performance of the model on the validation set. Use accuracy, recall, F1 score, and AUC to evaluate the performance of the model;
[0041] Repeat the above process until each subset has been used as a validation set once;
[0042] Calculate the average performance index of the cross-validation as the final evaluation result of the model;
[0043] B3 Model selection:
[0044] Compare the performance of different basic models in cross-validation, and select the model with the best test results as the basis for further optimization. Determine the best model according to specific business requirements and the importance of evaluation indicators.
[0045] Preferably, in the step S3, feature selection and weight adjustment include the following steps:
[0046] For the decision tree model, the feature selection step includes:
[0047] Evaluating feature importance based on information gain or Gini coefficient reduction; selecting key features for model training according to feature importance;
[0048] For the support vector machine (SVM) model, the feature selection and weight adjustment steps include:
[0049] Recursive feature elimination combined with SVM, starting from the full feature set, gradually deleting the least important features;
[0050] Based on kernel function-based feature mapping, mapping non-linearly separable data to a high-dimensional space for feature selection;
[0051] Implementing weight adjustment by optimizing the weights and bias terms of support vectors;
[0052] For the neural network model, the feature selection and weight adjustment steps include:
[0053] Automatic feature learning, learning the feature representation of input data by adjusting the weights of neuron connections;
[0054] Applying regularization methods such as L1 or L2 regularization for feature selection and weight adjustment;
[0055] Implementing weight adjustment using the backpropagation algorithm, calculating the weight gradient according to the error and updating the weights using an optimization algorithm.
[0056] Preferably, in the multi-model integration step, implementing the Stacking algorithm or directly using the outputs of each model as input parameters to establish a new integrated model, and improving the stability and accuracy of the overall credit evaluation ability through weight tuning.
[0057] Preferably, in the step S4,
[0058] Stereotyping and class customization include customer classification management and customized model optimization. Customer classification management divides customers into different categories according to needs and customizes a weight adjustment strategy for each category; the customized model optimization uses transfer learning or parameter fine-tuning techniques to perform additional customization on the model, increasing the sensitivity of the model to domain-specific risks;
[0059] The feedback learning mechanism establishes a real user feedback, complaint, and subsequent repayment behavior information feedback module, and establishes a corresponding update strategy to perform online learning and parameter update on the deployed model;
[0060] The risk dynamic mapping implements a dynamic risk control strategy, and dynamically adjusts the classification boundary, scoring threshold, and weight strategy in a globally and segmentally variable environment.
[0061] Preferably, in S5, the accuracy test uses statistical indicators such as ROC, accuracy, and precision to comprehensively evaluate the classification effect of the model; the risk stability detection uses backtesting the model over multiple cycles to detect the stability and adaptability under different environmental fluctuations; the interpretability analysis uses aspects such as perceptron decision boundary visualization, feature attribution, and constraint perception to build a cognitive bridge between the machine learning model and business experts; the compliance audit implements compliance criteria and best practices in model evaluation audits to ensure that the credit scoring model complies with legal and regulatory requirements.
[0062] Beneficial effects:
[0063] 1. Improve the accuracy of credit assessment: Through the personalized credit assessment model, the error rejection rate and over-credit rate of loan approvals are minimized.
[0064] 2. Enhance the adaptability of the model: Dynamically adjust and update the credit scoring model to adapt to changes in the market environment and regulatory requirements.
[0065] 3. Improve the user experience: The highly interpretable credit scoring model improves the work efficiency of financial reviewers and the user experience.
[0066] 4. Protect user privacy and data security: By implementing strict privacy and security measures, gain the trust of customers and avoid legal and regulatory risks.
[0067] 5. Enhance the competitiveness of the financial industry: Effective credit scoring algorithms enhance the market competitiveness and brand image of financial institutions.
[0068] In summary, with its innovative technical solution, the present invention is expected to significantly improve the accuracy and security of credit decisions in a complex financial market environment, bringing far-reaching impacts to the field of financial credit. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the specific embodiments and with reference to the attached Figure 1 , drawings. It should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0071] A personalized credit scoring method in the field of financial credit provided by the present invention, the algorithm comprising the following steps:
[0072] S1 Data collection and preprocessing, including internal data integration and external data integration;
[0073] S2 Feature extraction, including statistical characteristic analysis, trend index extraction and diversified feature construction;
[0074] S3 Model construction and training, including basic model selection, cross-validation, feature selection and weight adjustment, and multi-model integration;
[0075] S4 Personalized fine-tuning, including stereotyping and class customization, feedback learning mechanism and risk dynamic mapping;
[0076] S5 Model evaluation and optimization, including accuracy test, risk stability detection, interpretability analysis and compliance audit.
[0077] Preferably, in the S1,
[0078] Internal data integration automatically collects the customer's account information, repayment records and credit reports, and performs structured processing;
[0079] External data integration uses the OAuth protocol to collect influence dimension data of social media performance and e-commerce transaction data through customer authorization, and uses the API to standardize data fields;
[0080] Data cleaning and feature engineering are also included in the S1,
[0081] The data cleaning includes: detecting missing values and outliers by applying the methods of data sampling and statistical analysis, and determining the data filling or missing value imputation strategy; implementing data type verification, converting incorrect data types into correct formats, and filtering invalid data items;
[0082] The feature engineering includes: statistical characteristic analysis, generating descriptive statistical data based on existing data and using it for the basic construction of features; trend index extraction, analyzing the trend indexes of account activities, including the time series dynamics of account balance and debt increase and decrease; diversified feature construction, combining professional background and living habits to construct additional credit scoring influencing factors.
[0083] Preferably, in the S2, statistical characteristic analysis generates descriptive statistical data based on existing data, trend index extraction analyzes the trend indexes of account activities, and diversified feature construction combines professional background, living habits or other external data to construct additional credit scoring influencing factors.
[0084] Preferably, in the S3, the basic model selection includes the following steps:
[0085] Select the decision tree as the basic model to construct the initial credit scoring model;
[0086] Use information gain to select the optimal splitting feature and splitting point, calculate the information gain of each feature, and select the feature with the largest information gain for splitting;
[0087] The calculation formula of information gain is:
[0088]
[0089] where D is the data set, A is the feature, Entropy(D) is the entropy of the data set D, Values(A) is the value set of the feature A, and Dv is the subset of the data set D where the value of the feature A is v.
[0090] Select the support vector machine (SVM) c;
[0091] For non-linearly separable data, map the data to a high-dimensional space through a kernel function to make it linearly separable in the high-dimensional space;
[0092] In the credit field, take the features of the borrower as the input, and classify the borrower into two categories: good credit and bad credit through the SVM model;
[0093] The SVM using the linear kernel function constructs the following classification hyperplane: ω T x + b = 0, where ω is the normal vector of the hyperplane, x is the input feature vector, and b is the bias term;
[0094] Select the neural network as the basic model to construct the initial credit scoring model; construct a multi-layer perceptron, including an input layer, a hidden layer, and an output layer. The input layer receives the features of the borrower, the hidden layer extracts features through non-linear transformation, and the output layer gives the prediction result of the credit score or default probability;
[0095] The calculation formula of the cross-entropy loss function is:
[0096]
[0097] where N is the number of samples, y i is the true label, is the predicted label.
[0098] Preferably, in S3, the cross-validation includes the following steps:
[0099] B1 Data partitioning:
[0100] Randomly partition the data set into a training set and a test set, partition it according to a certain ratio, and further partition the training set into multiple subsets for cross-validation;
[0101] B2 Model Training and Evaluation:
[0102] For each type of basic model, the process of cross - validation is as follows:
[0103] Take turns to select a subset as the validation set, and the remaining subsets as the training set;
[0104] Train the model on the training set and evaluate the performance of the model on the validation set. Use accuracy, recall, F1 - score, and AUC to evaluate the performance of the model;
[0105] Repeat the above process until each subset has been used as a validation set once;
[0106] Calculate the average performance metrics of cross - validation as the final evaluation result of the model;
[0107] B3 Model Selection:
[0108] Compare the performance of different basic models in cross - validation, and select the model with the best test results as the basis for further optimization. Determine the best model according to specific business requirements and the importance of evaluation metrics.
[0109] Preferably, in S3, feature selection and weight adjustment include the following steps:
[0110] For the decision tree model, the feature selection step includes:
[0111] Evaluate feature importance based on information gain or reduction of Gini coefficient; select key features for model training according to feature importance;
[0112] For the support vector machine (SVM) model, the feature selection and weight adjustment steps include:
[0113] Recursive feature elimination combined with SVM, starting from the full feature set, gradually delete the least important features;
[0114] Feature mapping based on kernel function, map non - linearly separable data to a high - dimensional space for feature selection;
[0115] Adjust the weights and bias terms of support vectors through an optimization algorithm to achieve weight adjustment;
[0116] For the neural network model, the feature selection and weight adjustment steps include:
[0117] Automatic feature learning, adjust the weights of neuron connections to learn the feature representation of input data;
[0118] Apply regularization methods such as L1 or L2 regularization for feature selection and weight adjustment;
[0119] The backpropagation algorithm implements weight adjustment, calculates the weight gradient based on the error, and updates the weights using an optimization algorithm.
[0120] Preferably, in the multi-model integration step, the Stacking algorithm is implemented or the outputs of each model are directly used as input parameters to establish a new integrated model, and the overall credit evaluation ability is improved in terms of stability and accuracy through weight optimization.
[0121] Preferably, in S4,
[0122] Stereotyping and class customization include customer classification management and customized model optimization. Customer classification management divides customers into different categories according to needs and formulates a weight adjustment strategy for each category; for the customized model optimization, transfer learning or parameter fine-tuning techniques are used to perform additional customization on the model to increase the model's sensitivity to domain-specific risks.
[0123] The feedback learning mechanism establishes a real user feedback, complaint, and subsequent repayment behavior information feedback module, and establishes a corresponding update strategy to perform online learning and parameter update on the deployed model.
[0124] The risk dynamic mapping implements a dynamic risk control strategy, and dynamically adjusts the classification boundary, scoring threshold, and weight strategy in a changing environment of the global and segmented markets.
[0125] Preferably, in S5, for the accuracy test, statistical indicators such as ROC, accuracy, and precision are used to comprehensively evaluate the classification effect of the model; for the risk stability detection, the model is backtested over multiple cycles to detect the stability and adaptability under different environmental fluctuations; for the interpretability analysis, aspects such as perceptron decision boundary visualization, feature attribution, and constraint perception are used to build a cognitive bridge between the machine learning model and business experts; for the compliance audit, the compliance criteria and best practices in model evaluation audit are implemented to ensure that the credit scoring model meets legal and regulatory requirements.
[0126] The personalized credit scoring algorithm of the present invention significantly improves the accuracy of credit evaluation by reducing the error rejection rate and over-credit rate in the loan approval process; its dynamic adjustment ability enables the model to quickly adapt to changes in the market environment and regulatory requirements, enhancing the adaptability of the model; at the same time, due to the high interpretability of the model, the work efficiency of financial reviewers and the user experience are improved; in addition, strict privacy and security measures protect user data, win customer trust, and avoid legal risks; and it enhances the market competitiveness and brand image of financial institutions, and is expected to greatly improve the accuracy and security of credit decisions in the financial market, having a profound impact on the financial credit field.
[0127] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0128] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A personalized credit scoring method in the field of financial credit, characterized in that, The algorithm includes the following steps: S1 Data collection and preprocessing, including internal data integration and external data integration; S2 Feature extraction, including statistical property analysis, trend index extraction, and diversified feature construction; S3 Model construction and training, including basic model selection, cross-validation, feature selection and weight adjustment, and multi-model integration; S4 Personalized fine-tuning, including stereotyping and class customization, feedback learning mechanism, and risk dynamic mapping; S5 Model evaluation and optimization, including accuracy test, risk stability detection, interpretability analysis, and compliance audit.
2. The personalized credit scoring method in the field of financial credit according to claim 1, wherein, In S1, Internal data integration automatically collects the customer's account information, repayment records, and credit reports, and performs structured processing; External data integration uses the OAuth protocol to collect data on the influence dimensions of social media performance and e-commerce transaction data through customer authorization, and uses the API to standardize data fields; S1 also includes data cleaning and feature engineering, The data cleaning includes: using the methods of data sampling and statistical analysis to detect missing values and outliers, and determining data filling or missing value imputation strategies; implementing data type verification, converting incorrect data types into the correct format, and filtering invalid data items; The feature engineering includes: statistical property analysis, generating descriptive statistical data based on existing data and using it for the basic construction of features; trend index extraction, analyzing the trend indexes of account activities, including the time-series dynamics of account balance and debt increase or decrease; diversified feature construction, combining professional background and living habits to construct additional credit score influencing factors.
3. A personalized credit scoring method in the field of financial credit according to claim 1, characterized in that, In S2, statistical property analysis generates descriptive statistical data based on existing data, trend index extraction analyzes the trend indexes of account activities, and diversified feature construction combines professional background, living habits, or other external data to construct additional credit score influencing factors.
4. A personalized credit scoring method in the field of financial credit according to claim 1, characterized in that In S3, the basic model selection includes the following steps: A1 Select a decision tree as the basic model to construct an initial credit score model; Use information gain to select the optimal splitting feature and splitting point, calculate the information gain of each feature, and select the feature with the largest information gain for splitting; The calculation formula for information gain is: where D is the data set, A is the feature, Entropy(D) is the entropy of the data set D, Values(A) is the value set of the feature A, and Dv is the subset of the data set D where the value of the feature A is v. A2 Select a support vector machine (SVM); For non-linearly separable data, map the data to a high-dimensional space through a kernel function to make it linearly separable in the high-dimensional space; In the credit field, use the features of borrowers as input, and classify borrowers into two categories: good credit and bad credit through the SVM model; The SVM using a linear kernel function constructs the following classification hyperplane: w T · x + b = 0, where ω is the normal vector of the hyperplane, x is the input feature vector, and b is the bias term; A3 Select a neural network as the basic model to construct an initial credit score model; construct a multi-layer perceptron, including an input layer, a hidden layer, and an output layer. The input layer receives the features of borrowers, the hidden layer extracts features through non-linear transformation, and the output layer gives the prediction results of credit scores or default probabilities; The calculation formula for using the cross-entropy loss function is: where N is the number of samples, and y i is the true label, and is the predicted label.
5. A personalized credit scoring method in the field of financial credit according to claim 1, characterized in that, In S3, the cross-validation includes the following steps: B1 Data Partitioning: Randomly partition the dataset into a training set and a test set, with a certain ratio. For cross-validation, further partition the training set into multiple subsets; B2 Model Training and Evaluation: For each type of basic model, the process of cross-validation is as follows: Take turns selecting a subset as the validation set, and the remaining subsets as the training set; Train the model on the training set and evaluate the performance of the model on the validation set. Use accuracy, recall, F1-score, and AUC to evaluate the performance of the model; Repeat the above process until each subset has been used as a validation set once; Calculate the average performance metrics of cross-validation as the final evaluation result of the model; B3 Model Selection: Compare the performance of different basic models in cross-validation, and select the model with the best test results as the basis for further optimization. Determine the best model according to specific business requirements and the importance of evaluation metrics.
6. The personalized credit scoring method in the field of financial credit according to claim 1, characterized in that, In the above S3, feature selection and weight adjustment include the following steps: For the decision tree model, the feature selection steps include: Evaluate feature importance based on information gain or Gini coefficient reduction; select key features for model training according to feature importance; For the support vector machine (SVM) model, the feature selection and weight adjustment steps include: Recursive feature elimination combined with SVM, starting from the full feature set, gradually delete the least important features; Feature mapping based on kernel functions, map non-linearly separable data to a high-dimensional space for feature selection; Adjust the weights and bias terms of support vectors through an optimization algorithm to achieve weight adjustment; For the neural network model, the feature selection and weight adjustment steps include: Automatic feature learning, adjust the weights of neuron connections to learn the feature representation of input data; Apply regularization methods such as L1 or L2 regularization for feature selection and weight adjustment; Use the backpropagation algorithm to achieve weight adjustment, calculate the weight gradient according to the error and update the weights using an optimization algorithm.
7. A personalized credit scoring method in the field of financial credit according to claim 1, characterized in that, In the multi-model integration step, implement the Stacking algorithm or directly use the outputs of each model as input parameters to establish a new integrated model, and improve the stability and accuracy of the overall credit evaluation ability through weight tuning.
8. A personalized credit scoring method in the field of financial credit according to claim 1, characterized in that, In the above S4, Stereotyping and class customization include customer classification management and customized model optimization. Customer classification management divides customers into different categories according to needs, and customizes a weight adjustment strategy for each category; the customized model optimization uses transfer learning or parameter fine-tuning techniques to perform additional customization on the model, increasing the sensitivity of the model to domain-specific risks; The feedback learning mechanism establishes a real user feedback, complaint, and subsequent repayment behavior information feedback module, and establishes a corresponding update strategy to perform online learning and parameter update on the deployed model; The risk dynamic mapping implements dynamic risk control strategies, and dynamically adjusts the classification boundary, scoring threshold, and weight strategy in the changing environment of the global and segmented markets.
9. The personalized credit scoring method in the field of financial credit according to claim 1, characterized in that, In the above S5, accuracy testing uses statistical metrics such as ROC, accuracy, and precision to comprehensively evaluate the classification effect of the model; risk stability detection uses backtesting the model over multiple cycles to detect the stability and adaptability under different environmental fluctuations; Interpretability analysis uses perceptron decision boundary visualization, feature attribution, and constraint perception to build a cognitive bridge between machine learning models and business experts; compliance auditing implements compliance guidelines and best practices in model evaluation auditing to ensure that credit scoring models meet legal and regulatory requirements.
Citation Information
Cited By
AI Agent-driven credit risk early warning strategy automatic evaluation method and device, control equipment and computer readable storage medium
CN120807138A