Risk assessment method, device, equipment, medium and program product
By screening and integrating multi-dimensional information and business information, and using risk prediction models to conduct loan risk assessment, the shortcomings of traditional models when facing complex risk patterns and multi-source data are resolved, and more efficient and accurate risk assessment is achieved.
Patent Information
- Application Number
- CN202510935805.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional loan risk assessment models perform poorly when faced with complex nonlinear risk patterns and multi-source data, have poor feature scalability, low computational efficiency, and weak interpretability, and find it difficult to dynamically integrate social information and real-time transaction data.
By obtaining multi-dimensional information of the target object and business information of the target business, screening mutually unrelated multi-dimensional risk assessment data, performing data fusion processing, and using risk prediction models for scoring, including calculating correlation coefficients, variance analysis and feature screening, combined with time window alignment, aggregation, cross-statistics and weighted fusion, the machine learning model with optimal performance is trained.
It improves the accuracy and efficiency of loan risk assessment, can capture complex nonlinear relationships, enhance the interpretability and computing speed of the model, and optimize large-scale data processing capabilities.
Smart Images

Figure CN120689134A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data, and specifically to a risk assessment method, apparatus, equipment, medium, and program product. Background Art
[0002] Traditional loan risk assessment relies primarily on credit scoring models, which conduct quantitative analysis based on fixed rules and historical data. These models typically examine a borrower's credit history (such as repayment record and credit card usage), income status, debt level, and macroeconomic indicators (such as unemployment and inflation rates). They then use statistical methods such as linear regression and logistic regression to generate a credit score, which ultimately determines loan approval and loan amount.
[0003] Although traditional models have long dominated the field of risk assessment, they have significant flaws: first, insufficient model capabilities. Linear models find it difficult to capture complex nonlinear risk patterns and have poor adaptability to emerging risk factors (such as behavioral data); second, poor feature scalability. They rely on a fixed feature set and find it difficult to dynamically integrate multi-source data (such as social information and real-time transactions); low computational efficiency: performance bottlenecks are prominent when faced with large-scale data, affecting the timeliness of assessments; third, weak interpretability and an opaque model decision-making process make it difficult for users and regulators to understand the scoring logic.
[0004] As financial data diversifies and risks become more complex, the limitations of traditional methods are becoming increasingly apparent. The industry urgently needs to introduce technologies such as machine learning and big data processing to enhance models' ability to identify nonlinear relationships, enhance the flexibility of feature fusion, and balance predictive accuracy and interpretability to meet the challenges of modern credit risk management. Summary of the Invention
[0005] In view of the above problems, the present application provides a risk assessment method, apparatus, device, medium and program product for improving prediction accuracy.
[0006] The first aspect of the present application provides a risk assessment method, which includes: obtaining multidimensional information of a target object and business information of a target business of the target object, wherein the multidimensional information is related to the risk level of the target object, and the business information is related to the risk level of the target business; screening mutually unrelated multidimensional risk assessment data from the multidimensional information and the business information; fusing the multidimensional information with the business information to obtain composite risk feature data; and evaluating the multidimensional risk assessment data and the composite risk feature data based on a risk prediction model to obtain a risk score of the target object in the target business.
[0007] According to an embodiment of the present application, the filtering of mutually unrelated multidimensional risk assessment data from the multidimensional information and the business information includes: calculating the correlation coefficient between each group of variable data in the multidimensional information and the business information, and when the correlation coefficient is greater than a preset threshold, filtering one of the two groups of variable data to add to the multidimensional risk assessment data; calculating the variance of each group of variable data in the multidimensional information and the business information to perform variance analysis, and filtering discrete features in the multidimensional information and the business information to add to the multidimensional risk assessment data.
[0008] According to an embodiment of the present application, the calculation of the correlation coefficient between each group of variable data in the multidimensional information and the business information includes: for the continuous variable data in the multidimensional information and the business information, if the continuous variable data conforms to the normal distribution, calculating the Pearson correlation coefficient between the continuous variable data of different groups; if the continuous variable data does not conform to the normal distribution, calculating the Spearman correlation coefficient between the continuous variable data of different groups; for the discrete variable data in the multidimensional information and the business information, performing a chi-square test on two groups of discrete variable data of different groups to calculate the chi-square value.
[0009] According to an embodiment of the present application, the method further includes: screening out data in the multidimensional information and the business information whose sample missing amount is greater than a preset ratio, is concentrated in a single value, and is continuous and whose variance tends to 0.
[0010] According to an embodiment of the present application, the fusing of the multidimensional information with the business information to obtain composite risk characteristic data includes: aligning the multidimensional information and the business information according to a time window; aggregating the multidimensional information and the business information to obtain first composite risk characteristic data; cross-statisticing the multidimensional information and the business information to obtain second composite risk characteristic data, wherein the cross-statistics include statistical target indicators based on the associated multidimensional information and the business information; and weighted fusing the multidimensional information and the business information based on business rules to obtain third composite risk characteristic data.
[0011] According to an embodiment of the present application, the method includes: training the risk prediction model, including: obtaining multidimensional information of historical target objects and business information of historical target businesses, screening out irrelevant historical multidimensional risk assessment data, and fusing the multidimensional information of historical target objects and business information of historical target businesses to obtain historical composite risk feature data, and forming a data feature set based on the historical multidimensional risk assessment data and the historical composite risk feature data; selecting at least two models from an alternative model library, training the at least two models based on the data feature set and expert knowledge, and using a cross-validation method to evaluate the performance of the at least two models, and screening the model with the best performance as the risk prediction model; regularly updating the data feature set and retraining the risk prediction model.
[0012] The second aspect of the present application provides a risk assessment device, characterized in that the device includes: an information acquisition module, used to obtain multidimensional information of a target object and business information of a target business of the target object, the multidimensional information is related to the risk level of the target object, and the business information is related to the risk level of the target business; a feature extraction module, used to filter out mutually unrelated multidimensional risk assessment data from the multidimensional information and the business information; a feature fusion module, used to fuse the multidimensional information with the business information to obtain composite risk feature data; a risk assessment module, used to evaluate the multidimensional risk assessment data and the composite risk feature data based on a risk prediction model to obtain a risk score of the target object in the target business.
[0013] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0014] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0015] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0017] Figure 1Schematically illustrates an application scenario diagram of the risk assessment method, apparatus, device, medium, and program product according to an embodiment of the present application;
[0018] Figure 2 The following schematically shows a flow chart of a risk assessment method according to an embodiment of the present application;
[0019] Figure 3 A schematic diagram of a risk assessment device according to an embodiment of the present application is shown; and
[0020] Figure 4 A block diagram of an electronic device suitable for implementing a risk assessment method according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0021] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0022] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0024] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0025] It should be noted that the risk assessment method and device provided in this application can be used in the field of financial technology for business risk assessment, and can also be used in any field other than the field of financial technology. The application field of the risk assessment method and device provided in this application is not limited.
[0026] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0027] An embodiment of the present application provides a risk assessment method, which includes: obtaining multidimensional information of a target object and business information of a target business of the target object, the multidimensional information being related to the risk level of the target object, and the business information being related to the risk level of the target business; screening mutually unrelated multidimensional risk assessment data from the multidimensional information and the business information; fusing the multidimensional information with the business information to obtain composite risk feature data; and evaluating the multidimensional risk assessment data and the composite risk feature data based on a risk prediction model to obtain a risk score of the target object in the target business.
[0028] Figure 1 The following schematically illustrates an application scenario diagram of the risk assessment method and device according to an embodiment of the present application.
[0029] like Figure 1 As shown, the application scenario 100 according to this embodiment may include an application scenario in the field of financial technology for loan risk assessment. A network 104 is used as a medium for providing a communication link between a first terminal device 101, a second terminal device 102, a third terminal device 103, and a server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0030] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0031] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0032] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0033] It should be noted that the risk assessment method provided in the embodiment of the present application can generally be executed by the server 105. Accordingly, the risk assessment device provided in the embodiment of the present application can generally be set in the server 105. The risk assessment method provided in the embodiment of the present application can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the risk assessment device provided in the embodiment of the present application can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0034] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0035] The following will be based on Figure 1 The scene described by Figure 2 The risk assessment method according to an embodiment of the present application is described in detail.
[0036] Figure 2 The flowchart of the risk assessment method according to an embodiment of the present application is schematically shown.
[0037] like Figure 2 As shown, the risk assessment method of this embodiment includes operations S210 to S240, and the transaction processing method can be executed sequentially.
[0038] In operation S210 , multidimensional information of a target object and business information of a target business are acquired, where the multidimensional information is related to the risk level of the target object, and the business information is related to the risk level of the target business.
[0039] In an embodiment of the present application, the user's consent or authorization may be obtained before obtaining the user's information. For example, before operation S210, a request to obtain multi-dimensional user information may be issued to the user. If the user agrees or authorizes the acquisition of multi-dimensional user information, operation S210 is performed.
[0040] In an embodiment of the present application, the target object may be a loan applicant or a lending enterprise. Multidimensional information refers to a data set that depicts the risk characteristics of the target object (such as an individual / enterprise) from multiple dimensions, typically including: credit dimension, financial dimension, behavioral dimension, and environmental dimension, etc. Among them, the data in the credit dimension include historical repayment records collected with the explicit authorization of the user, credit scores provided by legitimate credit reporting agencies, number of overdue payments, etc.; the data in the financial dimension include the debt-to-asset ratio proactively provided by the user, cash flow stability verified by bank statements, etc.; the data in the behavioral dimension include consumption frequency after desensitization processing, etc.; the data in the environmental dimension include industry risk levels, economic indicators of the place of enterprise registration, etc. Business information refers to risk influencing factors directly related to the target business, mainly including product attributes (such as business type, term, interest rate and other contract elements), business processing process characteristics and market environment, etc. The above data needs to be desensitized.
[0041] The embodiment of the present application uses the multi-dimensional information of the target object and the business information of the target business as the basic data for risk assessment. On the one hand, it can identify individual inherent risks from the user dimension through credit records, behavioral data, etc. On the other hand, it can quantify the risk volatility of the business itself from the business dimension by combining product characteristics and market environment. The comprehensive consideration of the two can avoid the blind spots of a single perspective and improve the accuracy of risk control.
[0042] In operation S220 , multi-dimensional risk assessment data that is not related to each other is filtered from the multi-dimensional information and the business information.
[0043] In risk assessment, filtering unrelated data from multidimensional information (user dimension) and business information (business dimension) can avoid feature redundancy, reduce the risk of model overfitting, and improve the interpretability of risk scores. For example, mathematically correlated features such as "monthly income" and "annual income" can be removed to reduce information redundancy. By filtering unrelated data, parameter estimation distortion caused by multicollinearity can be avoided, thereby enhancing model generalization. Data filtering reduces feature dimensionality, which can improve model calculation speed and optimize computational efficiency. It should be noted that when filtering data, only features with purely mathematical correlations are allowed to be removed, and features that may affect fairness are prohibited.
[0044] In the embodiment of the present application, meaningful features can be extracted from the original data, such as the frequency of credit card usage and loan repayment history explicitly authorized by the user, and data screening can be performed based on the extracted features. Before data screening, the data can be pre-processed as follows:
[0045] If the proportion of missing samples in multidimensional information and business information to the total number of samples is too large, the variable will be deleted;
[0046] If the variance of a continuous variable in the multidimensional information and business information approaches 0, it means that the variable does not change much and is not very helpful for model training or recognition, so the variable should be deleted;
[0047] If the frequency of a variable in the multidimensional information and business information is mostly concentrated on a single value, the variable is deleted.
[0048] Removing variables with high missingness rates can avoid the introduction of noise due to large numbers of missing values, reduce the model's reliance on imputation methods (such as mean filling), mitigate the risk of data bias, reduce the number of fields to be processed, and lower storage and computational costs. This is particularly crucial for large-scale datasets (e.g., tens of millions of samples). Fields with high missingness rates often indicate systemic collection issues (such as interface failures). Removing these fields can reduce model performance fluctuations caused by fluctuations in the data source.
[0049] Removing low-variance continuous variables can filter out invalid features and accelerate model convergence. For example, technical fields (such as device model) can be removed after they are confirmed to have no business significance.
[0050] Removing high-frequency variables can eliminate distribution bias. When a field has a high concentration of values, such as 95% of the values being the same, this feature has very low risk discrimination. Removing it can help focus on the most effective features.
[0051] It should be noted that before deleting a variable, it is necessary to verify that deleting the variable will not amplify bias in a specific group. If bias exists, data repair will be required for the affected user group. In addition, directly deleting some necessary or protected feature variables is prohibited.
[0052] After the pre-processing is completed, operation S220 may be performed. Operation S220 includes operations S221 and S222.
[0053] In operation S221 , the correlation coefficients between each set of variable data in the multidimensional information and the business information are calculated. When the correlation coefficient is greater than a preset threshold, one of the two sets of variable data is selected and added to the multidimensional risk assessment data.
[0054] In the embodiment of the present application, the following methods are included for calculating the correlation coefficients between each set of variable data in the multidimensional information and the business information.
[0055] First, for continuous variable data in multidimensional information and business information, if the continuous variable data conforms to the normal distribution, calculate the Pearson correlation coefficient between continuous variable data of different groups.
[0056] Assuming that the two continuous independent variables to be tested are X and Y, and assuming that the variables conform to the normal distribution, the Pearson correlation coefficient calculation formula is
[0057] in represents the covariance of the independent variables X and Y, E represents the mean, represents the mean of the independent variable X, represents the mean of the independent variable Y, represents the standard deviation of the independent variable X, Represents the standard deviation of the independent variable Y.
[0058] The Pearson correlation coefficient ranges from -1 to 1. When the Pearson correlation coefficient of two independent variables approaches 1, it indicates a more significant positive correlation. When the Pearson correlation coefficient of two independent variables approaches -1, it indicates a more significant negative correlation. When the Pearson correlation coefficient of two independent variables approaches 0, it indicates a greater lack of correlation. The higher the correlation between independent variables, the more likely it is to cause multicollinearity, which in turn leads to poor model stability. Even small perturbations in the sample can cause significant changes in the parameters. Therefore, only one characteristic of collinearity should be selected and the rest should be eliminated.
[0059] If the continuous variable data do not conform to the normal distribution, the Spearman correlation coefficient between the continuous variable data of different groups is calculated.
[0060] Assume that the two continuous independent variables to be tested are X and Y, and assume that the variables do not conform to the normal distribution
[0061] The calculation formula of Spearman correlation coefficient is:
[0062]
[0063] in, It represents the difference between the rank values of the i-th data pair, and n represents the total number of samples.
[0064] The Spearman correlation coefficient ranges from -1 to 1. When the Spearman correlation coefficient of two independent variables approaches 1, it indicates a more significant positive correlation. When the Spearman correlation coefficient of two independent variables approaches -1, it indicates a more significant negative correlation. When the Spearman correlation coefficient of two independent variables approaches 0, it indicates a greater lack of correlation. Select one of the collinearity features and eliminate the rest.
[0065] For discrete variable data in multidimensional and business information, perform a chi-square test on two different groups of discrete variable data to calculate the chi-square value. The chi-square test can be used to test the correlation between two discrete variables. Try to select independent variables with low correlation as the feature component.
[0066] By calculating correlation coefficients between multidimensional information and business information variables and removing highly correlated features (retaining only one), the number of features can be reduced, computing resources can be saved, and coefficient estimation distortion or sign contradictions caused by feature correlation in linear models (such as logistic regression) can be avoided. The retained features represent independent information dimensions, making model weights easier to interpret. For example, the impact of "credit score" on risk is no longer confounded by the strongly correlated "number of overdue payments." After optimizing the data, it is possible to focus on differentiated information, ensuring that each feature provides a unique risk signal and avoiding repeated calculations of similar information. Optimizing the data also helps combine weakly correlated features from user dimensions (such as consumption behavior) with business dimensions (such as product type), potentially revealing hidden risk patterns.
[0067] In operation S222, the variance of each group of variable data of the multidimensional information and the business information is calculated to perform variance analysis, and discrete features in the multidimensional information and the business information are screened and added to the multidimensional risk assessment data.
[0068] The purpose of analysis of variance (ANOVA) is to test whether there are significant differences in the means of different groups. ANOVA can be used to measure the degree of association between features and the target variable, thereby being used to screen discrete features.
[0069] It should be noted that if key fields in the multidimensional information and business information are removed during the above process, manual intervention is required to retain them.
[0070] In operation S230 , the multi-dimensional information is integrated with the business information to obtain composite risk feature data.
[0071] In the embodiments of this application, composite features can be created through data aggregation, interactive features, and other methods. By integrating multidimensional information with business information, it is possible to combine inherent user attributes (such as income stability) with dynamic business factors (such as loan term), avoiding blind spots in single-perspective assessments. For example, when high-income users apply for long-term loans, they need to consider economic cycle risks (such as the extended probability of default during a recession). Composite features (such as "debt-to-income ratio × product risk factor") have stronger risk prediction capabilities than raw features, potentially improving the KS value of the risk prediction model by 10%-30%. The interaction of user and business characteristics (such as "high-debt user × long-term product") often exhibits nonlinear risk jumps, and composite features can effectively capture these patterns. Based on composite risk feature data, it can be traced back to both user characteristics (such as credit score) and business parameters (such as product risk level), meeting regulatory transparency requirements.
[0072] Operation S230 includes S231 to S234.
[0073] In operation S231 , the multi-dimensional information and the business information are aligned according to a time window.
[0074] Align multi-dimensional information (such as user credit history and behavioral data) with business information (such as loan product parameters and market indicators) along a unified timeline to ensure temporal consistency. For example, a user's spending behavior over the past six months must be aligned with business activities during the same period (such as promotional periods and interest rate adjustments) to avoid feature distortion due to time skew. Alignment can be performed by using a key business point (such as the loan application date) as a benchmark and truncating the forward or backward period (e.g., ±30 days).
[0075] In operation S232, the multi-dimensional information and the business information are aggregated to obtain first composite risk feature data.
[0076] Aggregation operations are performed on the aligned data to extract cross-dimensional statistical features. For example, user-dimensional aggregation can calculate the average income and maximum debt ratio over the past three months; business-dimensional aggregation can calculate the average delinquency rate and channel approval rate of similar products; and spatiotemporal aggregation can summarize the correlation between user default rates and business indicators by region.
[0077] In operation S233, cross statistics are performed on the multi-dimensional information and the business information to obtain second composite risk feature data. The cross statistics include statistical target indicators based on the associated multi-dimensional information and business information.
[0078] Through the interactive calculation of multi-dimensional information and business information, potential risk patterns are discovered. Specific methods include conditional statistics, ratio intersection, and joint distribution. Conditional statistics analyzes user behavior within specific business scenarios (e.g., "prepayment rate of users of high-interest products"). Ratio intersection constructs cross-dimensional indicators (e.g., "user monthly repayment amount divided by average monthly payment for the product"). Joint distribution calculates the co-occurrence probability of user attributes and business attributes (e.g., "the proportion of young users choosing long-term loans"). These features can capture nonlinear correlations (e.g., "the default rate of users with low credit scores and high-risk products increases sharply"), improving the model's ability to identify complex risks.
[0079] In operation S234 , the multi-dimensional information and the business information are weighted and fused based on the business rules to obtain third composite risk feature data.
[0080] For example, you can define the sensitivity weights of business types (e.g., credit loans, mortgage loans) to user characteristics (e.g., income, debt), thereby calculating the user's risk profile for that business. Furthermore, you can adjust the timeliness of these weights based on macroeconomic indicators (e.g., GDP growth rate).
[0081] In an embodiment of the present application, the first composite risk feature data can provide basic risk trends, the second composite risk feature data can reveal the deep interaction effects between user information and business data, and the third composite risk feature data can inject business prior knowledge. The three together constitute hierarchical risk assessment data.
[0082] Before executing S240, the multidimensional risk assessment data and the composite risk feature data are standardized to improve the efficiency and stability of model training and to avoid certain features having too significant an impact on the model while certain features having too weak an impact on the model.
[0083] In operation S240 , the multi-dimensional risk assessment data and the composite risk feature data are evaluated based on the risk prediction model to obtain a risk score of the target object in the target business.
[0084] In the embodiment of the present application, the risk prediction model is trained based on historical data. Training the risk prediction model includes S241 to S243.
[0085] In operation S241, the multidimensional information of the historical target object and the business information of the historical target business are obtained, irrelevant historical multidimensional risk assessment data are filtered out, and the multidimensional information of the historical target object and the business information of the historical target business are integrated to obtain the historical composite risk feature data, and a data feature set is formed based on the historical multidimensional risk assessment data and the historical composite risk feature data.
[0086] In operation S242, at least two models are selected from the candidate model library, the at least two models are trained based on the data feature set and expert knowledge, and the performance of the at least two models is evaluated using a cross-validation method, and the model with the best performance is selected as the risk prediction model.
[0087] In the embodiments of the present application, a suitable machine learning algorithm is used for model training, and multiple models can be used to facilitate the selection of the optimal model in the model evaluation stage. Common models include random forest, xgboost, neural network, etc. Decision tree is a common method for machine learning, which establishes a model by maximizing information gain. In each candidate split during the learning process, a random subset of features is selected and multiple decision trees are trained to form a random forest, thereby improving the accuracy of the classifier. During xgboost training, the antecedent distribution algorithm is used for greedy learning. Each iteration learns a tree to fit the residual of the prediction results of the previous t-1 trees and the true value of the training sample. The forward propagation and backpropagation of the neural network are used to search for the minimum value of the objective function to establish a model.
[0088] In the disclosed embodiments, cross-validation, confusion matrix, and other methods are used to evaluate model performance and adjust model parameters to improve prediction accuracy. Model evaluation optimizes model parameters and selects the best performing model.
[0089] In operation S243, the data feature set is updated periodically and the risk prediction model is retrained.
[0090] In the disclosed embodiments, the prediction performance of the model can be regularly monitored to promptly identify any degradation in model performance. The model can be updated based on the latest data and business needs to ensure the long-term effectiveness of the model.
[0091] The risk assessment method provided in accordance with the embodiments of the present disclosure can optimize the ability to process large-scale data sets and significantly improve assessment efficiency and real-time performance. By integrating user information and business information, this method can expand data dimensions, mine hidden risk information, and improve the accuracy of risk assessment. By screening irrelevant feature data, this method enables the model to capture more complex nonlinear relationships, thereby improving the accuracy of predicting expected loan risks.
[0092] Based on the above risk assessment method, this application also provides a risk assessment device. Figure 3 The device is described in detail.
[0093] Figure 3 The following schematically shows a structural block diagram of a risk assessment device according to an embodiment of the present application.
[0094] like Figure 3 As shown, the risk assessment device 300 of this embodiment includes an information acquisition module 310 , a feature extraction module 320 , a feature fusion module 330 and a risk assessment module 340 .
[0095] The information acquisition module 310 is used to obtain multidimensional information of the target object and business information of the target business of the target object. The multidimensional information is related to the risk level of the target object, and the business information is related to the risk level of the target business. In one embodiment, the information acquisition module 310 can be used to perform the operation S210 described above, which will not be repeated here.
[0096] The feature extraction module 320 is used to filter out mutually unrelated multidimensional risk assessment data from the multidimensional information and the business information. In one embodiment, the feature extraction module 320 can be used to perform the operation S220 described above, which will not be repeated here.
[0097] The feature fusion module 330 is used to fuse the multi-dimensional information with the business information to obtain composite risk feature data. In one embodiment, the feature fusion module 330 can be used to perform the operation S230 described above, which will not be repeated here.
[0098] The risk assessment module 340 is used to evaluate the multi-dimensional risk assessment data and the composite risk feature data based on the risk prediction model to obtain the risk score of the target object in the target business. In one embodiment, the risk assessment module 340 can be used to perform the operation S240 described above, which will not be repeated here.
[0099] According to embodiments of the present application, any multiple modules among the information acquisition module 310, feature extraction module 320, feature fusion module 330, and risk assessment module 340 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the information acquisition module 310, feature extraction module 320, feature fusion module 330, and risk assessment module 340 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the information acquisition module 310 , the feature extraction module 320 , the feature fusion module 330 and the risk assessment module 340 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0100] Figure 4 A block diagram of an electronic device suitable for implementing a risk assessment method according to an embodiment of the present application is schematically shown.
[0101] like Figure 4 As shown, an electronic device 400 according to an embodiment of the present application includes a processor 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage unit 408 into a random access memory (RAM) 403. The processor 401 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 401 may also include onboard memory for caching purposes. The processor 401 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present application.
[0102] Various programs and data required for the operation of the electronic device 400 are stored in the RAM 403. The processor 401, ROM 402, and RAM 403 are connected to each other via a bus 404. The processor 401 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 402 and / or RAM 403. It should be noted that the programs may also be stored in one or more memories other than the ROM 402 and the RAM 403. The processor 401 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in the one or more memories.
[0103] According to an embodiment of the present application, electronic device 400 may further include an input / output (I / O) interface 405, which is also connected to bus 404. Electronic device 400 may also include one or more of the following components connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 408 including a hard disk; and a communication section 409 including a network interface card such as a LAN card or modem. Communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. Removable media 411, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 410 as needed, so that computer programs read from the removable media can be installed into storage section 408 as needed.
[0104] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.
[0105] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 402 and / or RAM 403 described above and / or one or more memories other than ROM 402 and RAM 403.
[0106] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to enable the computer system to implement the risk assessment method provided in the embodiments of the present application.
[0107] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the computer program is executed by the processor 401. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0108] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 409, and / or installed from a removable medium 411. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0109] In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the processor 401, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0110] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0112] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.
Claims
1. A risk assessment method, characterized in that: The method comprises: Acquire multidimensional information of a target object and business information of a target business, wherein the multidimensional information is related to the risk level of the target object, and the business information is related to the risk level of the target business; Filtering mutually unrelated multidimensional risk assessment data from the multidimensional information and the business information; Fusing the multi-dimensional information with the business information to obtain composite risk feature data; The multi-dimensional risk assessment data and the composite risk feature data are evaluated based on a risk prediction model to obtain a risk score of the target object in the target business.
2. The method according to claim 1, characterized in that The filtering of mutually unrelated multidimensional risk assessment data from the multidimensional information and the business information includes: Calculating the correlation coefficient between each set of variable data in the multidimensional information and the business information, and when the correlation coefficient is greater than a preset threshold, selecting one of the two sets of variable data to be added to the multidimensional risk assessment data; The variance of each group of variable data of the multidimensional information and the business information is calculated to perform variance analysis, and discrete features in the multidimensional information and the business information are screened and added to the multidimensional risk assessment data.
3. The method according to claim 2, characterized in that Calculating the correlation coefficients between each set of variable data in the multidimensional information and the business information comprises: For the continuous variable data in the multidimensional information and the business information, if the continuous variable data conforms to a normal distribution, calculating the Pearson correlation coefficient between the continuous variable data of different groups; If the continuous variable data do not conform to the normal distribution, calculate the Spearman correlation coefficient between the continuous variable data of different groups; For the discrete variable data in the multidimensional information and the business information, a chi-square test is performed on two groups of discrete variable data in different groups to calculate the chi-square value.
4. The method according to claim 2, characterized in that The method further comprises: Eliminate data in the multidimensional information and the business information whose sample missing amount is greater than a preset ratio, is concentrated in a single value, or is continuous and whose variance tends to 0.
5. The method according to claim 1, wherein The fusing of the multi-dimensional information with the business information to obtain composite risk feature data includes: Aligning the multidimensional information and the business information according to a time window; Aggregating the multi-dimensional information and the business information to obtain first composite risk feature data; Performing cross statistics on the multi-dimensional information and the business information to obtain second composite risk feature data, wherein the cross statistics include calculating target indicators based on the associated multi-dimensional information and the business information; The multi-dimensional information and the business information are weighted and fused based on business rules to obtain third composite risk feature data.
6. The method according to claim 1, wherein The method comprises: Training the risk prediction model comprises: Acquire multidimensional information of historical target objects and business information of historical target businesses, filter out irrelevant historical multidimensional risk assessment data, fuse the multidimensional information of historical target objects and business information of historical target businesses to obtain historical composite risk feature data, and form a data feature set based on the historical multidimensional risk assessment data and the historical composite risk feature data; selecting at least two models from a candidate model library, training the at least two models based on the data feature set and expert knowledge, and evaluating the performance of the at least two models using a cross-validation method, and selecting the model with the best performance as the risk prediction model; The data feature set is updated regularly, and the risk prediction model is retrained.
7. A risk assessment device, characterized in that: The device comprises: An information acquisition module, configured to acquire multi-dimensional information of a target object and business information of a target business, wherein the multi-dimensional information is related to the risk level of the target object, and the business information is related to the risk level of the target business; A feature extraction module, configured to filter mutually unrelated multidimensional risk assessment data from the multidimensional information and the business information; A feature fusion module, configured to fuse the multi-dimensional information with the business information to obtain composite risk feature data; The risk assessment module is used to evaluate the multi-dimensional risk assessment data and the composite risk feature data based on a risk prediction model to obtain a risk score of the target object in the target business.
8. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.