Risk assessment method and system based on artificial intelligence, and storage medium

By collecting user and external economic data, calculating risk experience and ability values, and selecting relevant feature training models, the problem of failure to fully consider user behavior and external environment changes in the existing technology is solved, and a more efficient and accurate risk assessment is achieved.

CN120278808APending Publication Date: 2025-07-08ICBC SANMENXIA BRANCH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510334837.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing technology fails to fully consider the user's behavioral characteristics, changes in the external economic environment, and users' cognition and response ability to risks in bank risk assessment, resulting in insufficient accuracy in the assessment.

Method used

By collecting user data and external economic data, calculating user behavioral characteristics and economic characteristics, standardized processing to generate combined characteristics, calculate risk experience values and ability values, and train the risk assessment model through the influencing value selection, combining risk experience values and ability values for comprehensive feature training, so as to improve the accuracy of the evaluation model.

Benefits of technology

It improves the training efficiency and evaluation accuracy of the risk assessment model, can more accurately predict the user's default probability and reduce the potential risks of the bank.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278808A_ABST
    Figure CN120278808A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a risk assessment method and system based on artificial intelligence and a storage medium. The method comprises the steps of collecting user data and external economic data, calculating a plurality of first behavior characteristics of a user based on the user data, calculating a plurality of second economic characteristics based on the external economic data, obtaining a combined characteristic based on the first behavior characteristics and the second economic characteristics, calculating a risk experience value and a risk capability value of each user based on the historical combination features, performing statistics on user data of each user to obtain a plurality of data features, calculating an influence value of each data feature, selecting related data features based on the influence values, combining all the related data features and the corresponding influence values to generate a comprehensive feature, and storing the comprehensive feature in a database; and taking the comprehensive features of the plurality of users as training data, taking the default probability as a target variable, and training a risk assessment model. The accuracy of risk assessment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a risk assessment method, system, and storage medium based on artificial intelligence. Background Art

[0002] In the financial service fields such as banks, the core of risk assessment is to predict the future default probability or risk level of users. Traditional methods mainly rely on data such as users' credit records, income levels, and asset statuses.

[0003] A similar prior art Chinese patent application with the publication number CN111882428A provides a commercial bank liquidity risk assessment method and device, including: obtaining historical business data and historical liquidity data of the commercial bank application service system; analyzing the obtained historical business data and historical liquidity data of the commercial bank by using pre-screened risk factors; evaluating the liquidity risk of the commercial bank through stress testing on the basis of risk factor analysis, and providing risk warnings in a timely manner according to the predicted liquidity risk of the commercial bank.

[0004] Another similar prior art is the Chinese patent application with the publication number CN118537008A, which provides a bank account transaction risk assessment method and device, including: obtaining account data and associated transaction data of the bank; constructing account nodes according to the account data, and constructing transaction edges according to the transaction data; constructing a transaction graph according to the account nodes and transaction edges, and using a GCN graph convolutional neural network model to extract features of the transaction graph to obtain account features and transaction features; according to the account features and transaction features, using a GRU gated recurrent unit to extract the temporal feature information of the account data and the temporal feature information of the transaction data in the transaction graph; performing transaction risk assessment according to the temporal feature information of the account data and the temporal feature information of the transaction data.

[0005] However, the behavioral characteristics of users, changes in the external economic environment, and users' awareness and response abilities to risks also have a greater impact on whether users will default in the future. Neither of the above two methods takes this issue into account. Therefore, the present invention provides a risk assessment method, system, and storage medium based on artificial intelligence. Summary of the Invention

[0006] This application provides a risk assessment method, system, and storage medium based on artificial intelligence for comprehensively and accurately assessing the risks of users.

[0007] In the first aspect, this application provides a risk assessment method based on artificial intelligence, and the method includes:

[0008] Step S1: Collect user data and external economic data. The user data includes internal data and behavioral data, and the external economic data includes economic data and industry market data;

[0009] Step S2: Calculate multiple first behavioral characteristics of the user based on the user data, calculate multiple second economic characteristics based on the external economic data, perform standardization processing on the first behavioral characteristics and the second economic characteristics to obtain combined characteristics, and calculate the risk experience value and risk ability value of each user based on the historical combined characteristics;

[0010] Step S3: Statistically obtain multiple data characteristics from the user data of each user, and also calculate the influence value of each data characteristic. Select relevant data characteristics based on the influence value, combine all relevant data characteristics and the corresponding influence values to generate comprehensive characteristics, use the comprehensive characteristics of several users as training data, and use the default probability as the target variable to train the risk assessment model;

[0011] Step S4: When a user handles a specific business, input the comprehensive characteristics of the user into the risk assessment model to obtain the corresponding default probability, and perform a risk assessment on the specific business handled by the user based on the default probability.

[0012] Combined with the first aspect, in the first implementation manner of the first aspect of the present application, calculating the risk experience value of each user based on the historical combined characteristics includes:

[0013] Obtain the first time period based on the second economic characteristic. For each first time period, obtain the first behavioral characteristics within the first time period, set different weight values for different first behavioral characteristics, calculate the asset growth rate brought by each first behavioral characteristic within the first time period, multiply the weight values and the asset growth rates of all first behavioral characteristics and then sum them up to obtain a result value as the first evaluation value, calculate the average value of all first evaluation values, and use the average value as the risk experience value.

[0014] Combined with the first aspect, in the second implementation manner of the first aspect of the present application, calculating the risk ability value of each user based on the historical combined characteristics includes:

[0015] Obtain all the combined features of the user, divide all the combined features into several groups of combined features in chronological order, and set different first weight values for each group of combined features. Calculate the corresponding first index, second index, and third index based on each group of combined features, and also set corresponding second weight values for different indexes. Multiply the corresponding first weight value by the corresponding three indexes and then sum them to obtain the first ability value of the combined features of the corresponding group. Calculate the weighted average of all the first ability values as the risk ability value, where the three indexes refer to the first index, the second index, and the third index.

[0016] Combined with the first aspect, in the third implementation manner of the first aspect of this application, statistical analysis is performed on the user data of each user to obtain multiple data features, including:

[0017] For each user, count the number of defaults of the user from the time of account opening to the current time from the user's user data, perform statistical analysis on the user data within a preset time period to obtain statistical data, and also obtain the latest credit score. Use the data obtained after standardizing the number of defaults, the statistical data, the credit score, the risk experience value, and the risk ability value of each user as the multiple data features of the user.

[0018] Combined with the first aspect, in the fourth implementation manner of the first aspect of this application, the influence value of each data feature is also calculated, including:

[0019] Step S31: Use the set composed of all data features as the data feature set, obtain the data feature sets of the first number of users, and select one data feature from the data feature set as the transformed data feature;

[0020] Step S32: Transform the original data feature set to obtain the first supplementary feature set and the second supplementary feature set. Use the original data feature set as the training data to generate the first model, use the first supplementary feature set as the training data to train and generate the second model, and train the second supplementary feature set to generate the third model. Compare the prediction results of the first model, the second model, and the third model to obtain the influence value of the transformed data feature;

[0021] Step S33: Determine whether the influence values of all data features have been calculated. If not, select the next data feature from the data feature set as the transformed data feature, and then return to the above step S32 to calculate the influence value of the corresponding transformed data feature. If so, end this step.

[0022] Combined with the first aspect, in the fifth implementation manner of the first aspect of this application, the original data feature set is transformed to obtain the first supplementary feature set and the second supplementary feature set, including:

[0023] Delete the transformed data features from the original first set of data features to obtain a first supplementary feature set. Use the remaining data features in the original data feature set except the transformed data features as fixed data features. Combine the transformed data features with the first supplementary feature set corresponding to each transformed data feature respectively to obtain a second set of new data features. The number of the obtained second set of new data features is used as the second supplementary feature set, and the second number is equal to (the first number minus one) multiplied by the first number.

[0024] Combined with the first aspect, in the sixth implementation manner of the first aspect of this application, obtaining the influence value of the transformed data features by comparing the prediction results of the first model, the second model, and the third model includes:

[0025] Select several data feature sets from the remaining data feature sets as the test data set. Input the test data set into the first model and the second model respectively to obtain the corresponding first accuracy rate and the second accuracy rate. Subtract the second accuracy rate from the first accuracy rate to obtain a first difference. If the first accuracy rate is greater than the second accuracy rate and the first difference is greater than a preset first threshold, obtain multiple groups of test data sets. Input the multiple groups of test data sets into the first model and the third model respectively to obtain multiple corresponding first prediction errors and multiple corresponding third prediction errors. Calculate the influence value of the transformed data features based on the multiple first prediction errors and the corresponding multiple third prediction errors. Otherwise, set the influence value of the transformed data features to zero.

[0026] Combined with the first aspect, in the seventh implementation manner of the first aspect of this application, inputting the test data set into the first model and the second model respectively to obtain the corresponding first prediction error and the second prediction error includes:

[0027] Obtain the corresponding third prediction error and the first prediction error, and subtract the first prediction error from the third prediction error to obtain multiple second differences. Calculate the first average value and the first standard deviation of the multiple second differences. Use the result value obtained by dividing the first average value by the first standard deviation as the influence value of the corresponding transformed data feature.

[0028] In the second aspect, this application provides a risk assessment system based on artificial intelligence. The system includes:

[0029] A collection unit for collecting user data and external economic data. The user data includes internal data and behavioral data, and the external economic data includes economic data and industry market data;

[0030] A calculation unit for calculating multiple first behavioral characteristics of a user based on the user data, calculating multiple second economic characteristics based on the external economic data, performing standardization processing on the first behavioral characteristics and the second economic characteristics to obtain combined characteristics, and calculating the risk experience value and the risk ability value of each user based on the historical combined characteristics;

[0031] A training unit is configured to statistically obtain multiple data features for the user data of each user, calculate the influence value of each data feature, select relevant data features based on the influence value, combine all the relevant data features and the corresponding influence values to generate a comprehensive feature, use the comprehensive features of several users as training data, and use the default probability as the target variable to train a risk assessment model.

[0032] An evaluation unit is configured to, when a user handles a specific service, input the comprehensive feature of the user into the risk assessment model, obtain the corresponding default probability, and perform a risk assessment on the specific service handled by the user based on the default probability.

[0033] A third aspect of the present application provides a computer-readable storage medium storing instructions, which, when running on a computer, cause the computer to execute the above-mentioned artificial intelligence-based risk assessment method.

[0034] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0035] In the technical solution provided by the present application, user data and external environment data are collected to provide sufficient data support for subsequent calculation of the risk experience value and the risk ability value. Then, multiple first behavior features and multiple second behavior features are obtained through statistics and calculation of the user data and the external environment data. The risk experience value and the risk ability value of the user are calculated based on the first behavior features and the second behavior features. Multiple data features of the user are obtained based on the user data, the corresponding influence value is calculated for each data feature, relevant data features related to the default probability are selected based on the influence value, and the relevant data features are obtained by calculating the influence value. This enables only the data features relevant to the default probability to be used for training when training the risk assessment model later, which can improve the training efficiency of the risk assessment model and also improve the assessment accuracy of the risk assessment model. The relevant data features and the corresponding influence values are combined with the risk experience value and the risk ability value to form a comprehensive feature, and the comprehensive feature is used to train the risk assessment model. Adding the risk experience value and the risk ability value to the training data can further improve the accuracy of the risk assessment model. Description of the Drawings

[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1Schematic diagram of an embodiment of the artificial intelligence-based risk assessment method in the embodiments of the present application;

[0038] Figure 2 Schematic diagram of an embodiment of calculating the influence value of data features in the embodiments of the present application;

[0039] Figure 3 Schematic diagram of an embodiment of calculating the influence value of transformed data features in the embodiments of the present application;

[0040] Figure 4 Schematic diagram of an embodiment of the artificial intelligence-based risk assessment system in the embodiments of the present application. Detailed implementation manners

[0041] The embodiments of the present application provide an artificial intelligence-based risk assessment method, system, and storage medium. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the term "comprising" or "having" and any deformation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0042] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 An embodiment of the artificial intelligence-based risk assessment method in the embodiments of the present application includes:

[0043] Step S1: Collect user data and external economic data. The user data includes internal data and behavioral data, and the external economic data includes economic data and industry market data.

[0044] Specifically, internal data refers to customer-related data that can be obtained from the bank, and behavioral data refers to the user's operation behaviors. Internal data includes, but is not limited to, age, occupation, credit record, income level, asset status, transaction history, consumption habits, and changes in account balance, etc. Behavioral data includes login frequency, transaction time, fund transfer pattern, and response to risk warnings. Economic data refers to macroeconomic data, including market interest rate, unemployment rate, inflation rate. Industry market data refers to relevant policy data and market data of various industries, including news, policy changes, market trends, stock market fluctuations, commodity prices, etc. of various industries. These data may affect the user's financial situation and may also affect the user's trading behavior.

[0045] Step S2: Calculate multiple first behavioral characteristics of the user based on the user data, calculate multiple second economic characteristics based on the external economic data, perform standardization processing on the first behavioral characteristics and the second economic characteristics to obtain combined characteristics, and calculate the risk experience value and risk ability value of each user based on the historical combined characteristics.

[0046] Specifically, the first behavioral characteristics include transaction characteristics, fund flow characteristics, and behavior pattern characteristics. Calculating the transaction characteristics, fund flow characteristics, and behavior pattern characteristics of the user based on the user data includes: calculating the average monthly consumption amount and consumption volatility of the user within a preset time period (such as the most recent three months). For example, if the consumption amounts of the user in the past three months are 5100, 4800, and 4500 respectively, then the average monthly consumption amount is 4800, the consumption amount of the month before the month with a consumption amount of 5100 is 4700, and the standard deviation of the consumption amounts in the past three months is 336.65. It also counts the average monthly large-amount transaction frequency and average monthly cross-border transaction frequency of the user within the preset time period, and takes the average monthly consumption amount, average monthly consumption volatility, average monthly large-amount transaction frequency, and average monthly cross-border transaction frequency as transaction characteristics. Calculate the average monthly ratio of fund inflow to outflow and average monthly change rate of account balance of the user, and take the average monthly ratio of fund inflow to outflow and average monthly balance change rate as fund flow characteristics. It also counts the average monthly login frequency of the user logging in to the bank application and the average risk response time to risk warnings, and takes the average monthly login frequency and average risk response time as behavior pattern characteristics.

[0047] The second economic feature includes the market volatility feature and the macroeconomic feature. The market volatility feature and the macroeconomic feature of the user's industry are calculated based on external economic data, including: calculating the monthly average investment value change rate, market interest rate change rate, and industry income change rate of the user within a preset time period (such as the last three months), calculating the monthly average investment value change rate of the user within a preset time period. For example, if the user invests in stocks and funds, the value three months ago was 100,000 yuan, and the current value is 90,000 yuan, then the investment portfolio change rate is (9 - 10) / 10 = -10%. Obtaining the inflation rate, unemployment rate, and exchange rate change rate within the last three months based on external economic data, taking the monthly average investment value change rate, market interest rate change rate, and industry income change rate as the industry market volatility feature, and taking the inflation rate, unemployment rate, and exchange rate change rate as the macroeconomic feature.

[0048] Normalize the first behavior feature and the second economic feature to obtain the combined feature, including: calculating the first behavior feature and the second economic feature for multiple preset time periods (such as a three - month cycle). Through calculation, multiple first behavior features and multiple second economic features can be obtained. Calculate the corresponding mean and standard deviation based on the feature values of all features. Use z - score to standardize each feature value based on the mean and standard deviation to obtain the corresponding standardized value. Combine the standardized values of all feature values as the combined feature of the corresponding user. Collectively call all the calculated combined features the historical combined features. Calculate the risk experience value and risk ability value of the user based on the historical combined features. The specific calculation method will be explained in detail later.

[0049] Step S3: Statistically analyze the user data of each user to obtain multiple data features, and also calculate the influence value of each data feature. Select relevant data features based on the influence value. Combine all relevant data features, the corresponding influence values, the risk experience value and risk ability value of the corresponding user to generate the comprehensive feature. Take the comprehensive features of several users as the training data, and take the default probability as the target variable to train the risk assessment model.

[0050] Specifically, statistical analysis is performed on the user data of each user to obtain multiple data features. Subsequently, the influence value of each data feature is calculated. The specific method for calculating the influence value will be explained in detail later. The influence value refers to the correlation between the data feature and the default probability. By calculating the influence value, all data features related to the default probability are obtained, and data features unrelated to the default probability are removed. By removing data features unrelated to the default probability, the training efficiency of the subsequent risk assessment model is effectively improved. All data features related to the default probability are used as relevant data features, and data features with an influence value greater than zero are also used as relevant data features. Additionally, all relevant data features are combined with the corresponding influence values, the risk experience values, and the risk capacity values of the corresponding users to generate comprehensive features. Subsequently, the comprehensive features of a large number of users are obtained and used as training data, with the default probability as the target variable, to train the risk assessment model. The risk assessment model can predict the default probability of users based on their comprehensive features. Since the user's behavioral characteristics, changes in the external economic environment, and the user's awareness and response ability to risks also have a significant impact on whether the user will default in the future, and the risk experience value and the risk capacity value are calculated based on behavioral characteristics and external environment data, adding the corresponding influence values, risk experience values, and risk capacity values to the comprehensive features can effectively improve the prediction accuracy of the risk assessment model.

[0051] Step S4: When a user handles a specific business, the comprehensive features of the user are input into the risk assessment model to obtain the corresponding default probability, and the risk assessment of the specific business handled by the user is performed based on the default probability.

[0052] Specifically, when a user handles a specific business, for example, when a user wants to apply for a loan at a bank, there is a possibility that the user may not be able to repay the loan. At this time, the comprehensive features of the corresponding user are input into the risk assessment model, and the default probability of the corresponding user is output based on the risk assessment model. For example, if the default probability exceeds a preset threshold, the bank can reject the loan application to reduce the potential risk of the bank.

[0053] The above method first collects user data and external environment data to provide sufficient data support for subsequent calculation of the risk experience value and the risk ability value. Then, through the statistics and calculation of the user data and the external environment data, multiple first behavioral characteristics and multiple second behavioral characteristics are obtained. Based on the first behavioral characteristics and the second behavioral characteristics, the risk experience value and the risk ability value of the user are calculated. Based on the user data, multiple data characteristics of the user are obtained, and the corresponding influence value is calculated for each data characteristic. Based on the influence value, the relevant data characteristics related to the default probability are selected, and the relevant data characteristics are obtained by calculating the influence value, so that only the data characteristics related to the default probability are used for training when training the risk assessment model subsequently, which can improve the efficiency of training the risk assessment model and also improve the evaluation accuracy of the risk assessment model. The relevant data characteristics and the corresponding influence values are combined with the risk experience value and the risk ability value to form comprehensive characteristics, and the comprehensive characteristics are used to train the risk assessment model. By adding the risk experience value and the risk ability value to the training data, the accuracy of the risk assessment model can be further improved.

[0054] In a specific embodiment, calculating the risk experience value of each user based on the historical combined characteristics specifically includes the following steps:

[0055] Based on the second economic characteristics, the first time period is obtained. For each first time period, the first behavioral characteristics within the first time period are obtained, different weight values are set for different first behavioral characteristics, the asset growth rate brought by each first behavioral characteristic within the first time period is calculated, and the result value obtained by multiplying the weight values of all the first behavioral characteristics by the asset growth rate and then adding them up is used as the first evaluation value. The average value of all the first evaluation values is calculated, and the average value is used as the risk experience value.

[0056] Specifically, the risk experience value refers to the maturity and coping ability of a user when facing economic risks (such as an economic downturn). Generally, experienced users will reduce risk investments, increase currency storage, and optimize the debt structure during an economic downturn. Such users are more likely to take effective coping measures when facing similar risks in the future, thereby reducing the possibility of default or loss. On the contrary, if a user increases risk investments and spending when facing economic risks, it is considered that these users may not have effective coping strategies when facing risks in the future, thereby increasing the possibility of default.

[0057] The first time period refers to the time period during an economic downturn. The second economic feature is obtained based on external economic data, including market interest rates, unemployment rates, etc. Therefore, the time period of the economic downturn can be obtained based on the second economic feature. The first behavioral feature refers to all behaviors that can increase the total assets of users during the first time period. There are various first behavioral features, such as increasing income, reducing consumption, selling stocks or funds before they decline, etc. Different weight values are set for different first behavioral features. For example, reducing consumption is the simplest and most direct behavior, so a lower weight of 2 is set; selling stocks or funds before they decline indicates that the user has relevant experience, so a higher weight of 5 is set; increasing income is set with a weight of 3. Calculate the asset growth rate brought by each first behavioral feature during the time period of the economic downturn. For example, if the user's average monthly salary is 5,000 and the user increases 500 in income through part-time work, the corresponding asset growth rate is 10%. The user's average monthly consumption is 2,000, and by reducing consumption by 200, the corresponding asset growth rate is 10%. The value of the fund bought is 10,000 before the decline and 9,500 after the decline, so the corresponding asset growth rate is 5%. Multiply the weight value of the first behavioral feature by the corresponding asset growth rate and then add them up. The resulting value is used as the first evaluation value. For example, 10% * 2 + 5% * 5 + 10% * 3 = 0.75. During the account opening period, the user may encounter multiple different economic downturn cycles. Therefore, calculate the corresponding first evaluation values for other first time periods. For example, during the account opening period, the user experiences a total of 3 first time periods, and the corresponding first evaluation values are 0.5, 0.75, and 0.85 respectively. Calculate the average value of the three first evaluation values, which is 0.7, and use the calculated average value as the risk experience value corresponding to the user.

[0058] In a specific embodiment, the risk ability value of each user is calculated based on historical portfolio features, and the specific steps are as follows:

[0059] Obtain all the portfolio features of the user. Divide all the portfolio features into several groups of portfolio features in chronological order, and set different first weight values for each group of portfolio features. Calculate the corresponding first index, second index, and third index based on each group of portfolio features. Also set corresponding second weight values for different indexes. Multiply the corresponding first weight value by the corresponding three indexes and then add them up to obtain the first ability value of the portfolio features of the corresponding group. Calculate the weighted average value of all the first ability values as the risk ability value, where the three indexes refer to the first index, the second index, and the third index.

[0060] Specifically, the risk ability value refers to the user's perception and sensitivity to future risks, reflecting whether the user can perceive risks in advance and take measures. In order to quantify the risk ability value, first, all combination features corresponding to the user are obtained. The combination features are divided into multiple groups of different combination features according to the chronological order, and different weight values are set for different combination features. The user may have different cognitions about the economic situation over time, and the cognitions generally become higher. Therefore, different first weight values are set for different groups of combination features according to the chronological order. For example, the combination features of the first group are collected based on the user's earliest user data, so the corresponding weight value is set relatively small. For example, if there are 5 groups of data in chronological order, the corresponding first weight values are set to 2, 4, 6, 8, and 10 respectively. The first indicator refers to the response rate of the user to risk prompts. For example, the click-through rate of risk prompts in the bank APP by the user. For example, if there are 5 prompts in total and 4 clicks are made, the response rate is 4 / 5 = 0.8. The second indicator refers to the user's attention to economic news. For example, it can be obtained through the search frequency of the user searching for relevant economic news in the bank APP. For example, if the total number of searches is 10 times and the number of searches for relevant economic news is 3 times, the value of the second indicator is 3 / 10 = 0.3. The third indicator refers to the response rate of the user to market dynamics. For example, the user has received 10 relevant risk prompts in total, among which 7 times have made corresponding adjustments to their own risk investments, and the other 3 times have not made adjustments. Then the response rate is 0.7. Corresponding second weight values are set for the first indicator, the second indicator, and the third indicator respectively, such as 0.2, 0.3, and 0.6. The result value calculated by multiplying each indicator by the corresponding weight value and then adding them up is used as the first ability value of the combination features of the corresponding group, and then the weighted average of all the first ability values is calculated as the corresponding risk ability value.

[0061] In a specific embodiment, the user data of each user is statistically analyzed to obtain a plurality of data features, and the specific steps are as follows:

[0062] For each user, the number of defaults of the user from the time of account opening to the current time is statistically analyzed from the user's user data. Statistical data is obtained by statistically analyzing the user data within a preset time period, and the latest credit score is also obtained. The data obtained after standardizing the number of defaults, statistical data, credit score, risk experience value, and risk ability value of each user is used as the plurality of data features of the user.

[0063] Specifically, the number of defaults refers to the total number of defaults of a user from the account opening to the current time. Defaults include overdue repayments of loans, overdue repayments of credit cards, etc. The statistical data includes average monthly income, average monthly consumption, average monthly transaction times, proportion of cross-border transactions, proportion of large-value transactions, investment revenue, investment losses, number of active repayments, and debt ratio. The credit score refers to the credit score calculated for each user in the bank system. In fact, the bank has a large number of users. For the sake of clear explanation here, it is assumed that there are a total of 5 users. One of the data, such as the number of defaults, is 2, 0, 3, 9, 1 respectively. Based on the number of defaults of these 5 users, standardization processing is performed. For example, the z-score method can be used to obtain the corresponding standardized data. For example, the standardized data are 0.32, 0.95, 0, 1.90, -0.63 respectively. Standardization processing is performed on other data to obtain the corresponding standardized data. Each standardized data is used as the data feature corresponding to the user. Multiple standardized data form multiple data features corresponding to the user.

[0064] In a specific embodiment, calculating the influence value of each data feature specifically includes the following steps:

[0065] Step S31: Use the set composed of all data features as the data feature set, obtain the data feature sets of the first number of users, and select one data feature from the data feature set as the transformed data feature;

[0066] Step S32: Transform the original data feature set to obtain the first supplementary feature set and the second supplementary feature set. Use the original data feature set as the training data to generate the first model, use the first supplementary feature set as the training data to train and generate the second model, and train the second supplementary feature set to generate the third model. Compare the prediction results of the first model, the second model, and the third model to obtain the influence value of the transformed data feature;

[0067] Step S33: Determine whether the influence values of all data features have been calculated. If not, select the next data feature from the data feature set as the transformed data feature, and then return to step S32 to calculate the influence value of the corresponding transformed data feature. If so, end this step.

[0068] Specifically, such as Figure 2The following is a flowchart for calculating the influence value of data features. Assume that there are a total of n data features that form a data feature set, and the data feature set includes {a1, a2, ……, an}. The actual first quantity may be very large, such as 100 or 200 or even more. Here, for the sake of clearly explaining this method, assume that the first quantity is 3. Then, the data feature sets of 3 users are obtained as A = {a1, a2, ……, an}, B = {b1, b2, ……, bn}, and C = {c1, c2, ……, cn} respectively. Select a data feature, such as the first data feature, as the transformation data feature. For example, select the first data feature. a1, b1, and c1 correspond to each other and are the first data features corresponding to each user. Transform the original data feature set to obtain a first supplementary feature set and a second supplementary feature set. The specific transformation method will be explained in detail later. Use the original data feature set as training data to generate a first model, use the first supplementary feature set as training data to train and generate a second model, and train the second supplementary feature set to generate a third model. Compare the prediction results of the first model, the second model, and the third model to obtain the influence value of the transformation data feature. The specific method for calculating the influence value will be explained in detail later. Then, determine whether the influence values of all data features have been calculated. If not, select the next data feature from the data feature set as the transformation data feature, such as a2, b2, and c2. Then, return to step S32 to calculate the influence value of the corresponding data feature. When the influence values of all data features have been calculated, it means that the influence values of all data features have been calculated, and this step ends.

[0069] The influence values of all data features can be calculated through the above method.

[0070] In a specific embodiment, transforming the original data feature set to obtain a first supplementary feature set and a second supplementary feature set specifically includes the following steps:

[0071] Delete the transformation data feature from the original first quantity of data feature sets to obtain a first supplementary feature set. Use the remaining data features in the original data feature set except the transformation data feature as fixed data features. Combine the transformation data feature with the first supplementary feature sets corresponding to the transformation data feature respectively to obtain a second quantity of new data feature sets. Use the obtained second quantity of new data feature sets as the second supplementary feature set. The second quantity is equal to the product of the first quantity minus one and the first quantity.

[0072] Specifically, assume that the original data feature sets of three users are A = {a1, a2, ……, an}, B = {b1, b2, ……, bn}, and C = {c1, c2, ……, cn} respectively. The transformed data features a1, b1, and c1 are respectively deleted from the original data feature sets to obtain A1 = {a2, ……, an}, B = {b2, ……, bn}, and C1 = {c2, ……, cn}. The obtained A1, B1, and C1 are the first supplementary feature sets. The remaining data features in the original data feature sets except the transformed data features are used as fixed data features. Taking {a2, ……, an} as the fixed data features, the corresponding {b2, ……, bn} and {c2, ……, cn} are also fixed data features. The transformed data feature a1 is respectively combined with B1 and C1, b1 is respectively combined with A1 and C1, and c1 is respectively combined with A1 and C1. The second number of new data feature sets obtained are B2 = {a1, b2, ……, bn}, C2 = {a1, c2, ……, cn}, A2 = {b1, a2, ……, an}, C3 = {b1, c2, ……, cn}, A3 = {c1, a2, ……, an}, B3 = {c1, b2, ……, bn}. These new data feature sets are used as the second supplementary feature sets. The second number (6) = (the first number - 1) * the first number = (3 - 1) * 3.

[0073] In a specific embodiment, the influence value of the transformed data feature is obtained by comparing the prediction results of the first model, the second model, and the third model, which specifically includes the following steps:

[0074] Select several data feature sets from the remaining data feature sets as the test data sets. The test data sets are respectively input into the first model and the second model to obtain the corresponding first accuracy rate and the second accuracy rate. Subtract the second accuracy rate from the first accuracy rate to obtain the first difference. If the first accuracy rate is greater than the second accuracy rate and the first difference is greater than the preset first threshold, obtain multiple groups of test data sets. The multiple groups of test data sets are respectively input into the first model and the third model to obtain multiple corresponding first prediction errors and multiple corresponding third prediction errors. Calculate the influence value of the transformed data feature based on the multiple first prediction errors and the corresponding multiple third prediction errors. Otherwise, set the influence value of the transformed data feature to zero.

[0075] Specifically, in the previous step, the first number of data feature sets are selected, such as Figure 3The figure shows a flowchart for calculating the influence value of transformed data features. Select several data feature sets that have not been selected before from the remaining data feature sets as the test data set. Input the test data set into the first model and the second model respectively to obtain the first accuracy rate of the first model and the second accuracy rate of the second model. The first model is the original data feature set, and the second model is the data feature set after deleting the transformed data feature. If the transformed data feature is related to the default probability, after deleting this transformed data feature, the prediction accuracy rate of the second model should decrease compared with that of the first model. If the transformed data feature is not related to the default probability, after deleting this transformed data feature, the prediction accuracy rates of the second model and the first model should be roughly the same. Therefore, by comparing the magnitudes of the first accuracy rate and the second accuracy rate, it can be determined whether the transformed data feature is related to the default probability. To quickly calculate the influence value, the first quantity is set to be small. Since the amount of training data is small, there may be some errors in the first model, the second model, and the third model trained. Therefore, when the first accuracy rate is greater than the second threshold, a first threshold is also set to improve the accuracy of the judgment. If the first accuracy rate is greater than the second accuracy rate and the first difference between the two is greater than the first threshold, it is determined that the transformed data feature is related to the default probability; otherwise, it is determined that the transformed data and the default probability are not related. Then, by comparing the prediction errors of the first model and the second model, the corresponding influence value is calculated, that is, the correlation between the transformed data feature and the default probability.

[0076] In a specific embodiment, the influence value of the transformed data feature is calculated based on the first prediction error and the third prediction error, and specifically includes the following steps:

[0077] Obtain the corresponding third prediction error and the first prediction error, subtract the first prediction error from the third prediction error to obtain multiple second differences, calculate the first average value and the first standard deviation of the multiple second differences, and use the result value obtained by dividing the first average value by the first standard deviation as the influence value of the corresponding transformed data feature.

[0078] Specifically, if the transformed data feature is related to the default probability, then the prediction error of the third model trained using the second supplementary data set should increase. The greater the increase in the error, the stronger the correlation between the transformed data feature and the default probability. Input multiple groups of test data sets into the first model and the third model to obtain multiple first prediction errors and the corresponding multiple third prediction errors, where the first prediction error and the third prediction error correspond to each other based on the same test data set. Subtract the first prediction error from the corresponding third prediction error to obtain multiple second differences, calculate the first average value and the first standard deviation of the multiple second differences, and use the result value obtained by dividing the first average value by the first standard deviation as the influence value of the corresponding transformed data feature. By using the above method, the influence value of the transformed data feature is quantified, which is convenient for training a more accurate risk assessment model based on the influence value in the future.

[0079] The above described the risk assessment method based on artificial intelligence in the embodiments of the present application. Next, the risk assessment system based on artificial intelligence in the embodiments of the present application will be described. Please refer to Figure 4 One embodiment of the risk assessment system based on artificial intelligence in the embodiments of the present application includes:

[0080] A collection unit, configured to collect user data and external economic data. The user data includes internal data and behavioral data, and the external economic data includes economic data and industry market data;

[0081] A calculation unit, configured to calculate multiple first behavioral characteristics of a user based on the user data, calculate multiple second economic characteristics based on the external economic data, perform normalization processing on the first behavioral characteristics and the second economic characteristics to obtain combined characteristics, and calculate the risk experience value and risk ability value of each user based on the historical combined characteristics;

[0082] A training unit, configured to perform statistics on the user data of each user to obtain multiple data characteristics, calculate the influence value of each data characteristic, select relevant data characteristics based on the influence value, combine all relevant data characteristics and the corresponding influence values to generate comprehensive characteristics, use the comprehensive characteristics of several users as training data, use the default probability as the target variable, and train a risk assessment model;

[0083] An evaluation unit, configured to, when a user handles a specific business, input the comprehensive characteristics of the user into the risk assessment model, obtain the corresponding default probability, and perform a risk assessment on the specific business handled by the user based on the default probability.

[0084] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the risk assessment method based on artificial intelligence.

[0085] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above described system, system and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0086] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0087] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. An artificial intelligence-based risk assessment method, characterized in that, The method includes: Step S1: Collect user data and external economic data. The user data includes internal data and behavioral data, and the external economic data includes economic data and industry market data. Step S2: Calculate multiple first behavioral characteristics of the user based on the user data, calculate multiple second economic characteristics based on the external economic data, perform normalization processing on the first behavioral characteristics and the second economic characteristics to obtain combined characteristics, and calculate the risk experience value and risk ability value of each user based on the historical combined characteristics. Step S3: Statistically obtain multiple data characteristics from the user data of each user, calculate the influence value of each data characteristic, select relevant data characteristics based on the influence value, combine all relevant data characteristics and the corresponding influence values to generate comprehensive characteristics, use the comprehensive characteristics of several users as training data, and use the default probability as the target variable to train the risk assessment model. Step S4: When a user handles a specific business, input the comprehensive characteristics of the user into the risk assessment model to obtain the corresponding default probability, and perform a risk assessment on the specific business handled by the user based on the default probability.

2. The method according to claim 1, characterized in that, Calculating the risk experience value of each user based on the historical combined characteristics includes: Obtain the first time period based on the second economic characteristic. For each of the first time periods, obtain the first behavioral characteristics within the first time period, set different weight values for different first behavioral characteristics, calculate the asset growth rate brought by each first behavioral characteristic within the first time period, use the result obtained by multiplying and then adding the weight values and the asset growth rates of all first behavioral characteristics as the first evaluation value, calculate the average value of all first evaluation values, and use the average value as the risk experience value.

3. The method according to claim 1, wherein Calculating the risk ability value of each user based on the historical combined characteristics includes: Obtain all combined characteristics of the user, divide all the combined characteristics into several groups of combined characteristics in chronological order, and set different first weight values for each group of combined characteristics. Calculate the corresponding first index, second index, and third index based on each group of combined characteristics, and set corresponding second weight values for different indexes. Use the result obtained by multiplying and then adding the corresponding first weight value and the three indexes as the first ability value of the combined characteristics of the corresponding group, and calculate the weighted average value of all first ability values as the risk ability value, where the three indexes refer to the first index, second index, and third index.

4. The method according to claim 1, characterized in that Statistically obtaining multiple data characteristics from the user data of each user includes: For each user, statistically count the number of defaults of the user from the account opening to the current time from the user data of the user, statistically obtain statistical data from the user data within a preset time period, and obtain the latest credit score. Use the data obtained by performing normalization processing on the number of defaults, the statistical data, the credit score, the risk experience value, and the risk ability value of each user as multiple data characteristics of the user.

5. The method according to claim 1, wherein Calculating the influence value of each data characteristic includes: Step S31: Use the set composed of all data features as the data feature set. Obtain the data feature sets of the first number of users, and select one data feature from the data feature set as the transformed data feature. Step S32: Transform the original data feature set to obtain a first supplementary feature set and a second supplementary feature set. Use the original data feature set as training data to generate a first model, use the first supplementary feature set as training data to train and generate a second model, and train the second supplementary feature set to generate a third model. Compare the prediction results of the first model, the second model, and the third model to obtain the influence value of the transformed data feature. Step S33: Determine whether the influence values of all data features have been calculated. If not, select the next data feature from the data feature set as the transformed data feature, and then return to Step S32 to calculate the influence value corresponding to the transformed data feature. If so, end this step.

6. The method according to claim 5, wherein Transforming the original data feature set to obtain a first supplementary feature set and a second supplementary feature set includes: Delete the transformed data feature from the original first number of data feature sets to obtain a first supplementary feature set. Use the remaining data features in the original data feature set except the transformed data feature as fixed data features. Combine the transformed data feature with the first supplementary feature sets corresponding to the transformed data feature respectively to obtain a second number of new data feature sets. Use the obtained second number of new data feature sets as the second supplementary feature set, and the second number is equal to (the first number minus one) multiplied by the first number.

7. The method according to claim 5, characterized in that, Comparing the prediction results of the first model, the second model, and the third model to obtain the influence value of the transformed data feature includes: Select several data feature sets from the remaining data feature sets as test data sets. Input the test data sets into the first model and the second model respectively to obtain the corresponding first accuracy rate and second accuracy rate. Subtract the second accuracy rate from the first accuracy rate to obtain a first difference. If the first accuracy rate is greater than the second accuracy rate and the first difference is greater than a preset first threshold, obtain multiple groups of test data sets. Input the multiple groups of test data sets into the first model and the third model respectively to obtain multiple corresponding first prediction errors and multiple corresponding third prediction errors. Calculate the influence value of the transformed data feature based on the multiple first prediction errors and the corresponding multiple third prediction errors. Otherwise, set the influence value of the transformed data feature to zero.

8. The method according to claim 7, characterized in that, Inputting the test data sets into the first model and the second model respectively to obtain the corresponding first prediction error and second prediction error includes: Obtain the corresponding third prediction error and first prediction error, and subtract the first prediction error from the third prediction error to obtain multiple second differences. Calculate the first average value and the first standard deviation of the multiple second differences. Use the result value obtained by dividing the first average value by the first standard deviation as the influence value of the corresponding transformed data feature.

9. An artificial intelligence-based risk assessment system for implementing the artificial intelligence-based risk assessment method according to any one of claims 1-8, characterized in that, The system includes: A collection unit, configured to collect user data and external economic data. The user data includes internal data and behavioral data, and the external economic data includes economic data and industry market data. A calculation unit, configured to calculate multiple first behavioral characteristics of a user based on user data, calculate multiple second economic characteristics based on the external economic data, perform normalization processing on the first behavioral characteristics and the second economic characteristics to obtain combined characteristics, and calculate a risk experience value and a risk ability value for each user based on historical combined characteristics; A training unit, configured to statistically obtain multiple data characteristics from the user data of each user, calculate the influence value of each data characteristic, select relevant data characteristics based on the influence value, combine all relevant data characteristics and the corresponding influence values to generate a comprehensive characteristic, use the comprehensive characteristics of several users as training data, use the default probability as the target variable, and train a risk assessment model; An evaluation unit, configured to, when a user handles a specific business, input the comprehensive characteristic of the user into the risk assessment model, obtain the corresponding default probability, and perform a risk assessment on the specific business handled by the user based on the default probability.

10. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instruction is executed by a processor, it implements the artificial intelligence-based risk assessment method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Commercial bank liquidity risk assessment method and device

    CN111882428A

  • Bank account transaction risk assessment method and device

    CN118537008A