User information processing method and device, equipment, program product and storage medium
By combining static models and dynamic models, using target meta-learners for feature fusion and prediction, the problem of insufficient accuracy and generalization capabilities of existing credit risk assessment methods under large-scale diversified data is solved, and more efficient credit risk assessment is achieved.
Patent Information
- Application Number
- CN202510357096.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
When facing large-scale and diversified data, existing credit risk assessment methods are difficult to deal with nonlinear relationships and high-dimensional data, resulting in poor generalization capabilities of the model and inaccurate evaluation results.
The method of combining static models and dynamic models with target meta-learners is adopted to extract static features through static models, dynamic models extract dynamic features, and the target meta-learners are used to fusion and prediction to optimize the credit risk assessment results.
It significantly improves the accuracy, interpretability and adaptability of credit risk assessment, reduces the deviation of a single model, and enhances the generalization ability of the model.
Smart Images

Figure CN120258964A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of big data and can also be applied to the field of fintech, and particularly relate to a method, apparatus, device, program product, and storage medium for processing user information. Background Art
[0002] With the development of information technology, banks have accumulated a large amount of user data, including transaction records, user behaviors, and social network information, etc. Such data is not only huge in quantity but also diverse in types, including structured data (such as database records) and unstructured data (such as texts, pictures, etc.). At the same time, with the rapid development of fintech, financial institutions are facing an increasingly complex environment and ever-changing customer needs. While providing diversified financial services, financial institutions are also facing unprecedented risk challenges. Among them, the credit risk of users, as one of the most critical risks in the financial field, is directly related to the stable operation and sustainable development of financial institutions.
[0003] Although credit risk assessment of users has been widely applied in the financial field, existing assessment methods still have some limitations. Traditional risk assessment models mainly rely on statistical methods (such as logistic regression or decision trees, etc.), and this method is difficult to cope with the challenges brought by large-scale and diverse data. When facing non-linear relationships and high-dimensional data, overfitting or underfitting is likely to occur, resulting in poor generalization ability of the model and inaccurate assessment results. Summary of the Invention
[0004] Embodiments of the present invention provide a method, apparatus, device, program product, and storage medium for processing user information, which can automatically and accurately extract static and dynamic features from a large amount of data with complex structures, improve the accuracy of the model's credit risk assessment of users, and further improve the efficiency of credit risk assessment.
[0005] In a first aspect, embodiments of the present invention provide a method for processing user information, including:
[0006] Obtaining to-be-processed data of a to-be-processed user, where the to-be-processed data includes basic data and historical financial data of the to-be-processed user;
[0007] Based on the to-be-processed data, a pre-trained static model, and a pre-trained dynamic model, obtaining static prediction data and dynamic prediction data corresponding to the to-be-processed data;
[0008] According to a pre-determined target meta-learner, performing feature fusion and feature prediction on the static prediction data and the dynamic prediction data to obtain output data corresponding to the to-be-processed data;
[0009] Determine the credit risk assessment information of the user to be processed based on the output data, and display the credit risk assessment information to the first user.
[0010] In a second aspect, an embodiment of the present invention provides a user information processing device, which includes:
[0011] A data acquisition module, configured to acquire the data to be processed of the user to be processed, where the data to be processed includes the basic data and historical financial data of the user to be processed;
[0012] A first processing module, configured to obtain static prediction data and dynamic prediction data corresponding to the data to be processed based on the data to be processed, a pre-trained static model, and a pre-trained dynamic model;
[0013] A second processing module, configured to perform feature fusion and feature prediction on the static prediction data and the dynamic prediction data according to a pre-determined target meta-learner to obtain output data corresponding to the data to be processed;
[0014] A result determination module, configured to determine the credit risk assessment information of the user to be processed based on the output data, and display the credit risk assessment information to the first user.
[0015] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the user information processing method according to any one of the embodiments of the present invention.
[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the user information processing method according to any one of the embodiments of the present invention.
[0017] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, which implements the user information processing method according to any one of the embodiments of the present invention when executed by a processor.
[0018] In an embodiment of the present invention, to-be-processed data of a to-be-processed user is obtained, where the to-be-processed data includes basic data and historical financial data of the to-be-processed user; based on the to-be-processed data, a pre-trained static model, and a pre-trained dynamic model, static prediction data and dynamic prediction data corresponding to the to-be-processed data are obtained; according to a pre-determined target meta-learner, feature fusion and feature prediction are performed on the static prediction data and the dynamic prediction data to obtain output data corresponding to the to-be-processed data; based on the output data, credit risk assessment information of the to-be-processed user is determined, and the credit risk assessment information is presented to a first user. The method of the embodiment of the present invention effectively extracts static information in the to-be-processed data through the static model; accurately extracts dynamic information in the to-be-processed data through the dynamic model. The target meta-learner further optimizes the prediction result by learning the combination of the prediction results of the static model and the dynamic model and the original features. At the same time, the meta-learner can automatically adjust the importance of each feature, reduce the bias of a single model, and improve the generalization ability of the model. The method of the present invention significantly improves the accuracy, interpretability, and adaptability of credit risk assessment by integrating the advantages of the static model and the dynamic model and using the target meta-learner for feature fusion and final prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 The first flowchart of a user information processing method provided by an embodiment of the present invention;
[0021] Figure 2 The structural schematic diagram of the target meta-learner provided by an embodiment of the present invention;
[0022] Figure 3 The schematic diagram of the cross-validation process of the target meta-learner provided by an embodiment of the present invention;
[0023] Figure 4 The second flowchart of a user information processing method provided by an embodiment of the present invention;
[0024] Figure 5 The structural schematic diagram of a user information processing device provided by an embodiment of the present invention;
[0025] Figure 6 The structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the convenience of description, only the parts related to the present invention are shown in the drawings, rather than all the structures.
[0027] Figure 1 FIG. 1 is a first flowchart of a user information processing method provided by an embodiment of the present invention. The method of the embodiment of the present invention can automatically and accurately extract static and dynamic features from a large amount of complex-structured data, improving the accuracy of the model's credit risk assessment of users and further enhancing the efficiency of credit risk assessment. The information collected in the method of the embodiment of the present invention is information and data authorized by the user or fully authorized by all parties, and the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application complies with relevant laws, regulations, and standards of relevant countries and regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or reject. This method can be executed by a user information processing device provided by an embodiment of the present invention, and the device can be implemented in software and / or hardware. The following embodiments will be described by taking the integration of the device in an electronic device as an example. The electronic device can be a server or a computer device, etc. Refer to Figure 1 , and the method can specifically include the following steps:
[0028] Step 101: Obtain the data to be processed of the user to be processed.
[0029] Among them, the user to be processed is the user for whom credit risk prediction needs to be performed. The data to be processed is the relevant data of the user to be processed, including the basic data and historical financial data of the user to be processed. The basic data includes the personal basic information of the user to be processed (such as age, gender, and marital status, etc.), financial information (such as income, debt, and deposit, etc.), and historical credit information (such as credit score and historical loan records, etc.). The historical financial data includes the time series data of the transaction records of the user to be processed (such as consumption amount, repayment date, and overdue days, etc.). Specifically, when the user to be processed handles a business that requires credit risk assessment in a financial institution, the server can obtain the data to be processed of the user to be processed in the local database, or directly receive the data to be processed of the user to be processed uploaded by the user to be processed or the staff.
[0030] Step 102: Obtain the static prediction data and dynamic prediction data corresponding to the data to be processed based on the data to be processed, the pre-trained static model, and the pre-trained dynamic model.
[0031] Among them, the static model is a machine learning model for processing static features. Static features are used to represent the fixed information in the data to be processed that does not change over time, such as the basic user data of the user to be processed. The static model in this solution is a pre-trained machine learning model that can process structured data, such as LightGBM, XGBoost, or Random Forest model, etc. LightGBM is an efficient machine learning algorithm based on the gradient boosting framework, with characteristics such as fast training speed, low memory usage, and efficient processing of large-scale data. XGBoost is an optimized distributed gradient boosting library, based on the gradient boosting algorithm, and supports various machine learning tasks (such as classification, regression, etc.). The dynamic model is a machine learning model for processing dynamic features. Dynamic features are used to represent the unfixed information in the data to be processed that changes over time, such as the historical financial data of the user. The dynamic model in this solution is a pre-trained deep learning model that can process time series data, such as Long Short-Term Memory (LSTM) or Gated Recurrent Unit (GRU). LSTM and GRU are good at capturing long-term dependencies and dynamic changes in time series data.
[0032] Specifically, after obtaining the data to be processed, the server can perform data cleaning, outlier removal, data formatting, time format unification, and numerical format unification on the data to be processed. For the static data (basic user data) in the data to be processed, the server can convert it into numerical features. For example, gender can be encoded as "male = 1, female = 0"; marital status can be encoded as "married = 1, unmarried = 0". After obtaining the numerical features, standardize or normalize the numerical features to obtain the static features of the data to be processed. For the dynamic data (historical financial data) in the data to be processed, calculate its statistics according to the pre-set statistical dimensions, such as calculating the average consumption amount within a period of time in the historical financial data; the maximum and minimum consumption amounts within a period of time; the growth rate of monthly consumption amount; the number or frequency of overdue within a period of time. Further, determine the dynamic features of the data to be processed according to the statistics of each statistical dimension. After obtaining the static features and dynamic features, integrate the static features and dynamic features into a complete feature set. At the same time, input the static features into the static model, and obtain the static prediction data corresponding to the data to be processed through the static model. Input the dynamic features into the dynamic model, and obtain the dynamic prediction data corresponding to the data to be processed through the dynamic model.
[0033] In an alternative embodiment, after obtaining the data to be processed of the user to be processed, data preprocessing and static feature analysis are performed on the data to be processed to obtain the static features of the data to be processed; dynamic feature analysis is performed on the data to be processed based on the time series information of the data to be processed to obtain the dynamic features of the data to be processed; the static features and the dynamic data are feature fused to obtain fused features; the static features are input into a static model to obtain static prediction data; wherein, the static model is a model based on a lightweight gradient boosting machine; the dynamic features are input into a dynamic model to obtain dynamic prediction data; wherein, the dynamic model is a model based on a long short-term memory network.
[0034] Step 103: Perform feature fusion and feature prediction on the static prediction data and the dynamic prediction data according to the pre-determined target meta-learner to obtain the output data corresponding to the data to be processed.
[0035] Among them, the target meta-learner is a model pre-trained by the server for obtaining the risk assessment information corresponding to the data to be processed. The meta-learner belongs to the ensemble learning algorithm, which fuses multiple models according to a certain strategy. In this solution, before obtaining the data to be processed of the user to be processed, the following steps A1 - A3 are further included:
[0036] Step A1: Obtain the historical training data of each user, and perform feature extraction on the historical training data to obtain training static features and training dynamic features.
[0037] Among them, the historical training data is determined according to the data to be processed of each user in the historical period and is used to train the initial meta-learner. Specifically, when the initial meta-learner needs to be trained, the server can obtain the data to be processed of each user in the historical period from the local database of the financial institution, perform data cleaning and sorting on it, etc., to obtain the historical training data of each user. One-hot encoding processing is performed on the historical training data to obtain training static features. Time series analysis is performed on the historical training data, and the statistical features of the time series data are extracted to obtain training dynamic features.
[0038] Step A2: Input the training static features into the static model to obtain the static prediction result output by the static model; input the training dynamic features into the dynamic model to obtain the dynamic prediction result output by the dynamic model.
[0039] The static model is a machine learning model used to process static features. In this solution, the static model is a lightweight gradient boosting machine. The dynamic model is a machine learning model used to process dynamic features. In this solution, the dynamic model is a long short-term memory network. Specifically, after obtaining the static features and training dynamic features, the training static features are input into the lightweight gradient boosting machine, and the static prediction results are obtained through the lightweight gradient boosting machine. The training dynamic features are input into the long short-term memory network, and the dynamic prediction results are obtained through the long short-term memory network.
[0040] Step A3: Fuse the static prediction results and the dynamic prediction results to obtain a fusion result, and optimize the parameters of the pre-established initial meta-learner based on the fusion result, the static prediction result, and the dynamic prediction result to obtain a target meta-learner.
[0041] Specifically, use feature fusion techniques (such as dynamic feature fusion) to adjust the importance of the static prediction results and the dynamic prediction results, and fuse them to obtain a fusion result. The initial meta-learner can learn the combination of the prediction results of the base learners (static model and dynamic model) and the original features (static features, dynamic features, and the fused features of static and dynamic features), and use methods such as cross-validation or grid search to optimize the model parameters of the initial meta-learner to improve the prediction performance, and obtain a target meta-learner that can accurately output user credit risk information. By generating prediction results through the base learners (static model and dynamic model), and performing feature fusion and parameter optimization through the meta-learner, the advantages of different models are utilized, and the accuracy and generalization ability of credit risk assessment are improved.
[0042] Exemplarily, Figure 2 is a schematic structural diagram of the target meta-learner provided by an embodiment of the present invention. As Figure 2 shown, the target meta-learner consists of n primary learners and a meta-learner. The target meta-learner trains the original features through the primary learners, uses the prediction values output by the primary learners as the input features of the meta-learner, and uses the original feature labels as new labels to form a new data feature and perform further training in the meta-learner. The meta-learner outputs classification or regression results based on the prediction values of the primary learners. In the process of meta-learner model fusion, to prevent overfitting caused by repeated learning of the training set, the K-fold cross-validation method is proposed. Figure 3 is a schematic diagram of the cross-validation process of the target meta-learner provided by an embodiment of the present invention. As Figure 3As shown, K = 5. The data set is divided into a training set and a test set. The training set is randomly divided into 5 parts without replacement, with 1 part used for prediction and the remaining 4 parts used for training according to the learner model. Model 1 and Model 2 are two primary learners. Output predictions of the first-level learner are made for each prediction data set in the training set to obtain 5 sets of predicted values. The newly generated predicted values are combined in a simple average manner to form a new feature data set. Finally, the meta-learner learns and trains based on the new feature data set and the newly generated test set, and outputs the final result.
[0043] Specifically, after obtaining the static prediction data and the dynamic prediction data, the static prediction data and the dynamic prediction data are used as new features, and feature fusion is performed with the static features and the dynamic features to obtain a new feature set. The meta-learner weights, combines, or adjusts the feature set, and outputs the final credit risk prediction result, that is, the output data corresponding to the data to be processed. In an optional implementation manner, the static prediction data, the dynamic prediction data, the static features, the dynamic features, and the fusion features are determined as the input data of the target meta-learner; feature fusion is performed on the input data through the target meta-learner to obtain an input feature set; the features in the input feature set are weighted and combined to obtain the output data corresponding to the data to be processed.
[0044] Step 104: Determine the credit risk assessment information of the user to be processed based on the output data, and display the credit risk assessment information to the first user.
[0045] Among them, the first user can be a staff member of a financial institution. The credit risk assessment information includes the credit score, default probability, risk level, key risk factors, and risk suggestions of the user to be processed, etc. After obtaining the output data of the meta-learner, the output data is sorted out to obtain the credit risk assessment information of the user to be processed, and the credit risk assessment information is displayed to the first user. After receiving the credit risk assessment information of the user to be processed, the first user can perform risk management and customer relationship management, etc. based on the credit risk assessment information of the user to be processed. In this solution, optionally, determining the credit risk assessment information of the user to be processed based on the output data includes: performing credit score conversion on the output data according to a pre-set credit score mapping relationship to obtain the credit score of the user to be processed; generating credit risk assessment information based on the credit score and the data to be processed of the user to be processed.
[0046] Among them, the score mapping relationship is a relationship predetermined by the server based on domain big data and other factors. Specifically, after obtaining the output data of the meta-learner, the output data is converted to a score. For example, if the output relationship includes the probability of default, the default probability p can be converted into a standardized credit score using the credit score mapping relationship: credit score = offset + coefficient × log(p / (1-p)); wherein the offset and coefficient are pre-set according to specific needs. After obtaining the credit score of the user to be processed, a detailed credit risk assessment report is generated based on the user's credit score and the data to be processed. By converting the output data of the meta-learner into a standardized credit score, the standardization, interpretability and decision-making efficiency of credit risk assessment are significantly improved.
[0047] The technical solution of this embodiment obtains the data to be processed of the user to be processed, wherein the data to be processed includes the basic data and historical financial data of the user to be processed; based on the data to be processed, the pre-trained static model and the pre-trained dynamic model, the static prediction data and the dynamic prediction data corresponding to the data to be processed are obtained; according to the predetermined target meta-learner, the static prediction data and the dynamic prediction data are subjected to feature fusion and feature prediction to obtain the output data corresponding to the data to be processed; the credit risk assessment information of the user to be processed is determined based on the output data, and the credit risk assessment information is displayed to the first user. The technical solution of this embodiment effectively extracts the static information in the data to be processed through the static model; accurately extracts the dynamic information of the data to be processed through the dynamic model, and the target meta-learner further optimizes the prediction result by learning the combination of the prediction results and the original features of the static model and the dynamic model. At the same time, the meta-learner can automatically adjust the importance of each feature, reduce the deviation of a single model, and improve the generalization ability of the model. The technical solution of this embodiment significantly improves the accuracy, interpretability and adaptability of credit risk assessment by integrating the advantages of the static model and the dynamic model and using the target meta-learner for feature fusion and final prediction.
[0048] Figure 4 This is a second flow chart of a user information processing method provided by an embodiment of the present invention. This embodiment is a refinement based on the above embodiment. The specific method can be as follows Figure 4 As shown, the method may include the following steps:
[0049] Step 401: Obtain the data to be processed of the user to be processed.
[0050] The data to be processed includes the basic data and historical financial data of the users to be processed.
[0051] Step 402: perform data preprocessing and static feature analysis on the data to be processed to obtain static features of the data to be processed.
[0052] Among them, static features are used to represent the fixed information in the data to be processed that does not change over time, such as the basic user data of the user to be processed. After obtaining the data to be processed, data preprocessing can be performed on the data to be processed, including checking whether there are missing values in the data. If there are many missing values in certain features, the feature can be deleted; if the missing values are few, it can be processed by filling (such as filling with the average value, median or mode). Check whether there are outliers in the data (such as negative income or age of 0, etc.). If there are outliers, correct or delete them. Convert dynamic time data (such as transaction date and repayment date) into timestamp or date format. Perform one-hot encoding on the static data in the data to be processed, such as the basic user data. For example, for categorical features (such as gender, marital status, and education level, etc.), convert them into numerical data. For ordered categorical features (such as education level: primary school = 1, middle school = 2, university = 3), directly use numbers to represent their order. Standardize and normalize the numerical features to obtain the static features of the data to be processed.
[0053] Step 403: Perform dynamic feature analysis on the data to be processed based on the time series information of the data to be processed to obtain the dynamic features of the data to be processed.
[0054] Among them, dynamic features are used to represent the unfixed information in the data to be processed that changes over time, such as the historical financial data of the user. After obtaining the data to be processed, extract the dynamic data (historical financial data) in the data to be processed, sort and organize the dynamic data in chronological order, such as converting it into a unified timestamp or date format. Use a preset time window to perform statistics on the dynamic data, and calculate the statistics (such as average value, maximum value, minimum value, standard deviation, median, and trend features) within each time window (such as monthly or quarterly). Convert the statistically processed data into a numerical feature vector to obtain the dynamic features of the data to be processed.
[0055] Step 404: Obtain the static prediction data and dynamic prediction data corresponding to the data to be processed based on the static features, dynamic features, pre-trained static model, and pre-trained dynamic model.
[0056] Among them, feature fusion is the process of combining static features and dynamic features to form a more comprehensive feature set. A static model is a machine learning model used to process static features. The static model in this solution is a lightweight gradient boosting machine. A dynamic model is a machine learning model used to process dynamic features. The dynamic model in this solution is a long short-term memory network. In this solution, optionally, based on static features, dynamic features, a pre-trained static model, and a pre-trained dynamic model, static prediction data and dynamic prediction data corresponding to the data to be processed are obtained, including: performing feature fusion on the static features and dynamic data to obtain fused features; inputting the static features into the static model to obtain static prediction data; and inputting the dynamic features into the dynamic model to obtain dynamic prediction data.
[0057] Specifically, after obtaining the static features and dynamic features, the static features and dynamic features are directly concatenated together to form a larger feature vector, that is, the fused features. The static features are input into the lightweight gradient boosting machine, and static prediction data is obtained through the lightweight gradient boosting machine. The dynamic features are input into the long short-term memory network, and dynamic prediction data is obtained through the long short-term memory network. The dynamic model can capture long-term dependencies in time series data, thereby better understanding the behavior patterns of the users to be processed. Capturing long-term dependencies in time series
[0058] Using the lightweight gradient boosting machine can quickly process large-scale data sets, while maintaining low memory usage and reducing the computational complexity of feature extraction. Using the long short-term memory network can effectively capture long-term dependencies in time series data and accurately extract the dynamic features of the data to be processed.
[0059] Step 405: Determine the static prediction data, dynamic prediction data, static features, dynamic features, and fused features as the input data of the target meta-learner.
[0060] Specifically, the static features, dynamic features, static prediction data, dynamic prediction data, and fused features are combined, and the combined data is determined as the input data of the target meta-learner. The input data can provide comprehensive information for the target meta-learner to help it learn how to integrate these features to generate more accurate credit risk assessment results.
[0061] Step 406: Perform feature fusion on the input data through the target meta-learner to obtain an input feature set; perform weighted combination on the features in the input feature set to obtain the output data corresponding to the data to be processed.
[0062] Specifically, after receiving the input data, the target meta-learner further fuses the static prediction data, dynamic prediction data, static features, dynamic features, and fusion features. These are concatenated into a new feature vector. By dynamically learning, the importance of each feature is adjusted to obtain a more accurate result. The target meta-learner can adjust the importance of each feature according to the weights learned during the training phase. For example, if the weight of the static prediction data is high, it indicates that the meta-learner believes the prediction result of the static model is more valuable for the final evaluation. The target meta-learner multiplies each feature by its corresponding weight and performs a weighted sum to generate the final output data.
[0063] Step 407: Determine the credit risk assessment information of the user to be processed based on the output data, and display the credit risk assessment information to the first user.
[0064] In the technical solution of this embodiment, the data to be processed of the user to be processed is obtained. Among them, the data to be processed includes the basic data and historical financial data of the user to be processed. The data to be processed is preprocessed and static feature analysis is performed to obtain the static features of the data to be processed. Based on the time series information of the data to be processed, dynamic feature analysis is performed on the data to be processed to obtain the dynamic features of the data to be processed. Based on the static features, dynamic features, pre-trained static model, and pre-trained dynamic model, the corresponding static prediction data and dynamic prediction data of the data to be processed are obtained. The static prediction data, dynamic prediction data, static features, dynamic features, and fusion features are determined as the input data of the target meta-learner. Through the target meta-learner, feature fusion is performed on the input data to obtain an input feature set; the features in the input feature set are weighted and combined to obtain the output data corresponding to the data to be processed. Based on the output data, the credit risk assessment information of the user to be processed is determined, and the credit risk assessment information is displayed to the first user. The technical solution of this embodiment ensures the quality and consistency of the input data through data preprocessing, providing a reliable data basis for subsequent model training. By analyzing time series data, the behavior patterns and change trends of users can be captured, making the credit risk assessment more timely. The meta-learner can automatically adjust the weights of each feature according to the training data, ensuring that the model can make optimal decisions in different situations and further improving the prediction accuracy. This solution combines the prediction results of the static model and the dynamic model, giving full play to the advantages of different models in processing different types of data and improving the overall prediction performance. Through multi-model fusion, the overfitting or underfitting problems that may exist in a single model are reduced, the adaptability of the model to new data is enhanced, and the generalization of the prediction is improved.
[0065] Figure 5 FIG. is a schematic structural diagram of a user information processing device provided by an embodiment of the present invention. This device is applicable to execute the user information processing method provided by the embodiment of the present invention. As Figure 5As shown in the figure, the device may specifically include:
[0066] A data acquisition module 501, configured to acquire the data to be processed of the user to be processed, where the data to be processed includes the basic data and historical financial data of the user to be processed;
[0067] A first processing module 502, configured to obtain static prediction data and dynamic prediction data corresponding to the data to be processed based on the data to be processed, a pre-trained static model, and a pre-trained dynamic model;
[0068] A second processing module 503, configured to perform feature fusion and feature prediction on the static prediction data and the dynamic prediction data according to a pre-determined target meta-learner to obtain output data corresponding to the data to be processed;
[0069] A result determination module 504, configured to determine credit risk assessment information of the user to be processed based on the output data and display the credit risk assessment information to a first user.
[0070] Optionally, the first processing module 502 is specifically configured to: perform data preprocessing and static feature analysis on the data to be processed to obtain static features of the data to be processed;
[0071] Perform dynamic feature analysis on the data to be processed based on the time series information of the data to be processed to obtain dynamic features of the data to be processed;
[0072] Obtain static prediction data and dynamic prediction data corresponding to the data to be processed based on the static features, the dynamic features, a pre-trained static model, and a pre-trained dynamic model.
[0073] Optionally, the first processing module 502 is further configured to: perform feature fusion on the static features and the dynamic data to obtain fusion features;
[0074] Input the static features into the static model to obtain the static prediction data; where the static model is a model based on a lightweight gradient boosting machine;
[0075] Input the dynamic features into the dynamic model to obtain the dynamic prediction data; where the dynamic model is a model based on a long short-term memory network.
[0076] Optionally, the second processing module 503 is specifically configured to: determine the static prediction data, the dynamic prediction data, the static features, the dynamic features, and the fusion features as input data of the target meta-learner;
[0077] Performing feature fusion on the input data through the target meta-learner to obtain an input feature set;
[0078] Performing weighted combination on the features in the input feature set to obtain the output data corresponding to the data to be processed.
[0079] Optionally, the result determination module 504 is specifically configured to: perform credit score conversion on the output data according to a preset credit score mapping relationship to obtain the credit score of the user to be processed;
[0080] Generating the credit risk assessment information based on the credit score and the data to be processed of the user to be processed.
[0081] Optionally, before obtaining the data to be processed of the user to be processed, the second processing module 503 is further configured to: obtain the historical training data of each user, and perform feature extraction on the historical training data to obtain training static features and training dynamic features;
[0082] Inputting the training static features into the static model to obtain a static prediction result output by the static model; inputting the training dynamic features into the dynamic model to obtain a dynamic prediction result output by the dynamic model;
[0083] Fusing the static prediction result and the dynamic prediction result to obtain a fusion result, and optimizing the parameters of a pre-established initial meta-learner based on the fusion result, the static prediction result, and the dynamic prediction result to obtain the target meta-learner.
[0084] The user information processing device provided by the embodiments of the present invention can execute the user information processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. The content not described in detail in this embodiment can be referred to the description in any method embodiment of the present invention.
[0085] The embodiments of the present invention also provide a computer program product.
[0086] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer program products, the one or more computer program products can include one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0087] Figure 6 Schematic diagram of a structure of an electronic device provided for an embodiment of the present invention. Refer to Figure 6 , Figure 6 The electronic device 12 shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present application. As Figure 6 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 can include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0088] The bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0089] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media accessible by the electronic device 12, including volatile and nonvolatile media, removable and non-removable media.
[0090] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 6 not shown, typically referred to as a "hard disk drive"). Although Figure 6 not shown in, a disk drive for reading and writing on removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data media interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present application.
[0091] A program / utilities 40 having a set (at least one) of program modules 46 can be stored, for example, in the memory 28. Such program modules 46 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples. The program modules 46 generally perform the functions and / or methods in the embodiments described in the present application.
[0092] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through the bus 18. It should be understood that although Figure 6 not shown in, other hardware and / or software modules can be used in conjunction with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0093] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, for example, implementing a user information processing method provided by an embodiment of the present invention: obtaining the data to be processed of the user to be processed, where the data to be processed includes the basic data and historical financial data of the user to be processed; obtaining the static prediction data and dynamic prediction data corresponding to the data to be processed based on the data to be processed, the pre-trained static model, and the pre-trained dynamic model; performing feature fusion and feature prediction on the static prediction data and the dynamic prediction data according to the pre-determined target meta-learner to obtain the output data corresponding to the data to be processed; determining the credit risk assessment information of the user to be processed based on the output data, and presenting the credit risk assessment information to the first user.
[0094] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements a user information processing method provided by all embodiments of the present invention: obtaining the data to be processed of the user to be processed, where the data to be processed includes the basic data and historical financial data of the user to be processed; obtaining the static prediction data and dynamic prediction data corresponding to the data to be processed based on the data to be processed, the pre-trained static model, and the pre-trained dynamic model; performing feature fusion and feature prediction on the static prediction data and the dynamic prediction data according to the pre-determined target meta-learner to obtain the output data corresponding to the data to be processed; determining the credit risk assessment information of the user to be processed based on the output data, and presenting the credit risk assessment information to the first user. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic device, apparatus, or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution electronic device, apparatus, or device.
[0095] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction-executing electronic device, apparatus, or device.
[0096] The program code contained on a computer-readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0097] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0098] Note that the above is only the preferred embodiment of the present invention and the applied technical principles. Those skilled in the art will understand that the present invention is not limited to the specific embodiments herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for processing user information, characterized in that, The method includes: Obtain the data to be processed of the user to be processed, where the data to be processed includes the basic data and historical financial data of the user to be processed; Based on the data to be processed, the pre-trained static model, and the pre-trained dynamic model, obtain the static prediction data and dynamic prediction data corresponding to the data to be processed; According to the pre-determined target meta-learner, perform feature fusion and feature prediction on the static prediction data and the dynamic prediction data to obtain the output data corresponding to the data to be processed; Based on the output data, determine the credit risk assessment information of the user to be processed, and display the credit risk assessment information to the first user.
2. The method according to claim 1, characterized in that, Based on the data to be processed, the pre-trained static model, and the pre-trained dynamic model, obtaining the static prediction data and dynamic prediction data corresponding to the data to be processed includes: Perform data preprocessing and static feature analysis on the data to be processed to obtain the static features of the data to be processed; Based on the time series information of the data to be processed, perform dynamic feature analysis on the data to be processed to obtain the dynamic features of the data to be processed; Based on the static features, the dynamic features, the pre-trained static model, and the pre-trained dynamic model, obtain the static prediction data and dynamic prediction data corresponding to the data to be processed.
3. The method according to claim 2, wherein Based on the static features, the dynamic features, the pre-trained static model, and the pre-trained dynamic model, obtaining the static prediction data and dynamic prediction data corresponding to the data to be processed includes: Perform feature fusion on the static features and the dynamic data to obtain fused features; Input the static features into the static model to obtain the static prediction data; where the static model is a model based on a lightweight gradient boosting machine; Input the dynamic features into the dynamic model to obtain the dynamic prediction data; where the dynamic model is a model based on a long short-term memory network.
4. The method according to claim 3, wherein According to the pre-determined target meta-learner, perform feature fusion and feature prediction on the static prediction data and the dynamic prediction data to obtain the output data corresponding to the data to be processed, including: Determine the static prediction data, the dynamic prediction data, the static features, the dynamic features, and the fused features as the input data of the target meta-learner; Perform feature fusion on the input data through the target meta-learner to obtain an input feature set; Perform weighted combination on the features in the input feature set to obtain the output data corresponding to the data to be processed.
5. The method according to claim 1, wherein Based on the output data, determining the credit risk assessment information of the user to be processed includes: According to the pre-set credit score mapping relationship, perform credit score conversion on the output data to obtain the credit score of the user to be processed; Generate the credit risk assessment information based on the credit score and the data to be processed of the user to be processed.
6. The method according to claim 1, wherein Before obtaining the data to be processed of the user to be processed, the method further includes: Obtain the historical training data of each user, and perform feature extraction on the historical training data to obtain training static features and training dynamic features; Input the training static features into the static model to obtain the static prediction result output by the static model; input the training dynamic features into the dynamic model to obtain the dynamic prediction result output by the dynamic model; Fuse the static prediction result and the dynamic prediction result to obtain a fusion result, and optimize the parameters of the pre-established initial meta-learner based on the fusion result, the static prediction result, and the dynamic prediction result to obtain the target meta-learner.
7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements a user information processing method according to any one of claims 1-6.
8. A user information processing device, characterized in that, It includes: A data acquisition module, configured to acquire the data to be processed of the user to be processed, where the data to be processed includes the basic data and historical financial data of the user to be processed; A first processing module, configured to obtain the static prediction data and dynamic prediction data corresponding to the data to be processed based on the data to be processed, the pre-trained static model, and the pre-trained dynamic model; A second processing module, configured to perform feature fusion and feature prediction on the static prediction data and the dynamic prediction data according to the pre-determined target meta-learner to obtain the output data corresponding to the data to be processed; A result determination module, configured to determine the credit risk assessment information of the user to be processed based on the output data, and display the credit risk assessment information to the first user.
9. An electronic device, the electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the user information processing method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the user information processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Credit risk assessment method and equipment based on network recommendation dynamic relationship, and medium
CN114418742A
Risk account prediction method and device and electronic equipment
CN115049484A
Method and device for determining overdue loan risk, equipment and medium
CN117635310A
Credit default risk prediction method and device, equipment and storage medium
CN118247041A
Suspicious transaction evaluation method and device, electronic equipment and storage medium
CN119444227A
Cited By
Cross-border electronic customs declaration inspection early warning method and system based on dynamic risk learning
CN120746425A