A Credit Assessment Method and Device Based on Artificial Intelligence and Big Data Technology
By integrating multi-dimensional behavioral data and artificial intelligence technology in credit assessment, the problem of existing credit assessment methods relying on a single data source is solved, and more accurate credit risk assessment and personalized credit services are achieved.
Patent Information
- Application Number
- CN202510193758.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing credit assessment method relies on a single financial historical data, failing to make full use of customer data on other behavioral dimensions, resulting in insufficient credit risk assessment.
Using credit evaluation methods based on artificial intelligence and big data technology, we collect and analyze multi-dimensional behavioral data (including financial, social, consumption, geographical location and network behavior data), extract behavioral characteristics and combine pre-trained neural network models to generate credit risk assessment results.
Through the comprehensive assessment of multi-dimensional data, the accuracy of credit risk assessment is improved, and more accurate credit approval suggestions and personalized loan plans are provided.
Smart Images

Figure CN119671721B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a credit assessment method and device based on artificial intelligence and big data technology. Background Art
[0002] In the prior art, credit assessment usually relies on the historical credit records and financial data of the target customer for assessment, and mainly uses traditional financial data such as credit reports and bank transaction records to predict the credit risk of the customer. These methods often combine simple scoring models or rule-based analysis systems to generate the credit score of the customer, which is used as an important basis for credit approval.
[0003] However, the main problem faced in the prior art is that the data sources for credit assessment are relatively single, mainly relying on the financial history of the customer, and the data in other behavioral dimensions of the customer are not fully utilized.
[0004] Therefore, it is necessary to develop a new credit assessment method. Summary of the Invention
[0005] This application provides a credit assessment method and device based on artificial intelligence and big data technology to improve the accuracy of credit risk assessment.
[0006] This application provides a credit assessment method based on artificial intelligence and big data technology, including:
[0007] Collect multi-dimensional behavioral data of target customers, where the multi-dimensional behavioral data includes financial behavioral data, social behavioral data, consumption behavioral data, geographical location behavioral data, and network behavioral data; obtain multi-dimensional behavioral characteristics of target customers according to the multi-dimensional behavioral data, where the multi-dimensional behavioral characteristics include the behavioral stability index, abnormal behavior frequency, behavioral trend data, and multi-dimensional behavioral correlation data of target customers; the behavioral stability index is used to evaluate the consistency and volatility of target customers in each behavioral dimension; the abnormal behavior frequency refers to the frequency of abnormal events that deviate from the norm in the behavioral pattern of target customers; the behavioral trend data is used to reflect the change trend and development speed of the behavior of target customers over time; the multi-dimensional behavioral correlation data is used to reflect the mutual correlation between behavioral data in different dimensions; based on a pre-trained neural network model, generate a credit risk assessment result according to the credit report data, environmental risk assessment data, multi-dimensional behavioral data, and multi-dimensional behavioral characteristics of target customers; where the credit risk assessment result includes a credit risk score, the identification result of key risk factors, and future credit risk prediction; the environmental risk assessment data includes economic indicator data, and the economic indicator data includes GDP growth rate, unemployment rate, and industry risk data of the industry in which the target customer is engaged; generate a credit approval recommendation and a personalized loan plan based on the credit risk assessment result output by the neural network model.
[0008] In an alternative embodiment, the neural network model includes a feature extraction network, an anomaly detection network, and a multi-layer fusion network;
[0009] Among them, the input of the feature extraction network is the multi-dimensional behavioral data and multi-dimensional behavioral characteristics of target customers; it is implemented by a convolutional autoencoder and is used to extract features from the multi-dimensional behavioral data of target customers to obtain a non-linear behavioral feature vector;
[0010] The input of the anomaly detection network is the non-linear behavioral feature vector provided by the feature extraction network; the anomaly detection network is implemented by a recurrent neural network based on a multi-head self-attention mechanism and is used to obtain an anomaly behavior detection result and the prediction probability of anomaly behavior;
[0011] The input of the multi-layer fusion network is the non-linear behavioral feature vector provided by the feature extraction network, the anomaly behavior detection result and the prediction probability of anomaly behavior provided by the anomaly detection network, credit report data, and economic indicator data; the multi-layer fusion network is implemented by a capsule network and is used to obtain the credit risk assessment result of target customers.
[0012] In an alternative embodiment, the multi-layer convolutional autoencoder adopted by the feature extraction network includes:
[0013] The first input layer is used to receive multi-dimensional behavioral data and multi-dimensional behavioral characteristics of target customers;
[0014] The first convolutional layer is used to receive the multi-dimensional behavioral data provided by the first input layer. The first convolutional layer uses 64 convolutional kernels of size 3×3, performs a convolution operation to extract local features, and outputs 64 feature maps;
[0015] The first pooling layer is used to receive the output of the first convolutional layer, performs a max pooling operation using a pooling window of size 2×2, reduces the spatial dimension of the feature maps, and outputs 64 pooled feature maps;
[0016] The second convolutional layer is used to receive the output of the first pooling layer. The second convolutional layer uses 128 convolutional kernels of size 3×3 to extract deeper local features, capture temporal dependencies and non-linear relationships, and outputs 128 feature maps;
[0017] The second pooling layer is used to receive the output of the second convolutional layer. The second pooling layer performs a max pooling operation using a pooling window of size 2×2 to further reduce the spatial dimension of the features, and outputs 128 pooled feature maps;
[0018] The encoding layer is used to receive the output of the second pooling layer, compresses the feature vector to a 256-dimensional high-dimensional non-linear behavioral feature vector through a fully connected layer, which serves as the intermediate result of the feature extraction network and is used in the decoding stage;
[0019] The first deconvolutional layer is used to receive the 256-dimensional feature vector output by the encoding layer. The first deconvolutional layer uses 128 convolutional kernels of size 3×3 to perform a deconvolution operation, restore part of the spatial dimension, and outputs 128 feature maps;
[0020] The first upsampling layer is used to receive the output of the first deconvolutional layer, performs an upsampling operation of 2×2, restores the spatial resolution of the data, and generates 128 enlarged feature maps;
[0021] The second deconvolutional layer is used to receive the output of the first upsampling layer. The second deconvolutional layer uses 64 convolutional kernels of size 3×3 to further perform a deconvolution operation, restore a higher spatial resolution, and outputs 64 feature maps;
[0022] The second upsampling layer is used to receive the output of the second deconvolutional layer. The second upsampling layer performs an upsampling operation of 2×2, restores to a spatial resolution similar to that of the first convolutional layer, and outputs 64 feature maps;
[0023] The first output layer is used to further process the output of the second upsampling layer through a convolutional layer to generate reconstructed data similar to the first input layer; and generate a 256-dimensional high-dimensional non-linear behavior feature vector representing the abstract behavior characteristics of the target customer.
[0024] In an alternative embodiment, the anomaly detection network uses a recurrent neural network based on a multi-head self-attention mechanism, including:
[0025] The second input layer is used to receive the non-linear behavior feature vector output from the feature extraction network;
[0026] The spatio-temporal graph convolutional layer is used to receive the non-linear behavior feature vector provided by the second input layer. The spatio-temporal graph convolutional layer constructs a spatio-temporal graph based on the graph convolutional network and the time dimension, specifically for:
[0027] By arranging the multi-dimensional behavior data and multi-dimensional behavior characteristics of the target customer in chronological order, a spatio-temporal graph is formed, where the nodes represent the behavior states at different time points, and the edges represent the behavior relationships between different time points;
[0028] Through the graph convolutional network, convolutional operations are performed on the nodes in the spatio-temporal graph to capture the dependencies of user behavior in the time dimension and the spatial characteristics of the behavior data, and a 128-dimensional feature vector after spatio-temporal convolution is obtained;
[0029] The multi-head self-attention layer is used to receive the 128-dimensional feature vector output from the spatio-temporal graph convolutional layer, and uses the multi-head self-attention mechanism to weight the key behavior patterns to generate a 1024-dimensional feature vector;
[0030] The recurrent neural network layer is used to receive the 1024-dimensional feature vector from the multi-head self-attention layer, and combines the long short-term memory network to process the time series characteristics of the behavior data, and the output is a 512-dimensional time series feature vector representing the time characteristics of the abnormal behavior;
[0031] The second output layer is used to receive the 512-dimensional time series feature vector provided by the recurrent neural network layer, and generate the abnormal behavior prediction probability and abnormal behavior detection result of the target customer.
[0032] In an alternative embodiment, the multi-layer fusion network is implemented using a dynamic routing capsule network, and the dynamic routing capsule network includes:
[0033] The third input layer is used to receive the non-linear behavior feature vector provided by the feature extraction network, the abnormal behavior detection result and abnormal behavior prediction probability provided by the anomaly detection network, the credit report data, and the economic indicator data; the input data of the third input layer is integrated into a 919-dimensional feature vector as the input of the subsequent capsule layer;
[0034] The first capsule layer is used to receive the 919-dimensional feature vectors of the third input layer, and processes different types of input features through multiple capsule units. Each capsule unit contains 8 neurons, which are responsible for processing data of different dimensions. The output of each capsule unit is an 8-dimensional vector, representing the feature information of the input features in different dimensions. After passing through the first capsule layer, the output is a 64-dimensional feature vector. The dynamic routing mechanism of the first capsule layer includes: dynamically calculating the information transfer weights between capsule units according to the output of the capsule units. The vector length of the output of the capsule unit represents the feature intensity, and dynamically adjusting the information transfer paths between different capsule units to strengthen the capsule output related to key features.
[0035] The second capsule layer is used to receive the 64-dimensional vector from the first capsule layer, and further fuses and processes the feature information from different sources. The second capsule layer includes multiple first capsule units. Each first capsule unit includes 16 neurons, which are used to process multi-dimensional features. The second capsule layer divides and inputs the 64-dimensional feature vector into each capsule unit through the dynamic routing mechanism. The output of the second capsule layer is a 128-dimensional feature vector, reflecting the deep association between the target customer behavior data, economic data, and anomaly detection results.
[0036] The feature fusion layer is used to receive the 128-dimensional feature vector output by the second capsule layer and perform feature fusion through a fully connected layer. The output of the feature fusion layer is a 64-dimensional fused feature vector.
[0037] The third output layer is used to receive the 64-dimensional fused feature vector from the feature fusion layer and generate the comprehensive credit risk assessment result of the target customer. Specifically, the third output layer is used for:
[0038] Mapping the 64-dimensional feature vector to a credit risk score through the third fully connected layer, representing the credit risk level of the target customer.
[0039] Identifying and outputting the key factors affecting the customer's credit risk through the fourth fully connected layer, and the output is a set of classification labels or specific values. The key factors include abnormal behaviors, credit history, and economic environment changes.
[0040] Combining the 64-dimensional feature vector with historical behavior data through the time series analysis layer for prediction, generating the credit risk change trend in a specified future period, and the output is a future risk score, representing the potential credit risk of the customer in the future.
[0041] In an alternative embodiment, the calculation of the behavior stability index is achieved by the method of time-weighted average, specifically including:
[0042] Perform time slicing on multi-dimensional behavioral data and divide the behavioral data into different time windows;
[0043] Calculate the fluctuation range of behavioral characteristics in each dimension within each time window;
[0044] Perform weighted averaging on the behavioral fluctuations of each time window, where more recent time windows are given greater weights, to evaluate the stability of the target customer's current behavior and generate a behavioral stability index.
[0045] In an alternative embodiment, the detection of the abnormal behavior frequency is achieved through a threshold setting and an adaptive learning mechanism based on historical behavior patterns, including:
[0046] Establish a behavioral baseline pattern by analyzing the behavioral data of the target customer in the past 6 months to 1 year;
[0047] Set a threshold deviating from the historical behavior pattern, and behavioral events exceeding this threshold are marked as abnormal;
[0048] Apply an adaptive learning algorithm to update the threshold in real time based on the new behavioral data of the target customer, making the detection of abnormal behavior by the model more accurate.
[0049] In an alternative embodiment, the calculation of the multi-dimensional behavioral correlation data adopts a method based on the Pearson correlation coefficient, specifically including:
[0050] Normalize the behavioral characteristic data of the target customer in different dimensions;
[0051] Calculate the correlation between behavioral data in each dimension, including financial behavior and consumption behavior, social behavior and geographical location behavior, etc.;
[0052] Use the Pearson correlation coefficient to evaluate the linear correlation between behavioral data in different dimensions, generate a correlation matrix, and use it as the input for the subsequent credit risk assessment model.
[0053] In an alternative embodiment, the acquisition and update of the environmental risk assessment data are achieved through a dynamic macroeconomic monitoring system based on big data, specifically including:
[0054] Obtain the latest GDP growth rate, unemployment rate, and industry risk data from government statistical departments, industry reports, and third-party economic data providers;
[0055] Regularly update the economic indicators of the industry in which the target customer is engaged, and dynamically adjust the credit risk of the target customer in combination with macroeconomic environments such as economic recession and rising industry risks.
[0056] The present application also provides a credit assessment device based on artificial intelligence and big data technology, including:
[0057] A collection unit for collecting multi-dimensional behavior data of a target customer, where the multi-dimensional behavior data includes financial behavior data, social behavior data, consumption behavior data, geographical location behavior data, and network behavior data;
[0058] An obtaining unit for obtaining multi-dimensional behavior characteristics of the target customer according to the multi-dimensional behavior data, where the multi-dimensional behavior characteristics include a behavior stability index of the target customer, an abnormal behavior frequency, behavior trend data, and multi-dimensional behavior correlation data; the behavior stability index is used to evaluate the consistency and volatility of the target customer in each behavior dimension; the abnormal behavior frequency refers to the frequency of abnormal events deviating from the normal state in the behavior pattern of the target customer; the behavior trend data is used to reflect the change trend and development speed of the target customer's behavior over time; the multi-dimensional behavior correlation data is used to reflect the mutual correlation between behavior data in different dimensions;
[0059] A generating unit for generating a credit risk assessment result based on a pre-trained neural network model according to the credit report data, environmental risk assessment data, multi-dimensional behavior data, and multi-dimensional behavior characteristics of the target customer; where the credit risk assessment result includes a credit risk score, an identification result of key risk factors, and a future credit risk prediction; the environmental risk assessment data includes economic indicator data, and the economic indicator data includes GDP growth rate, unemployment rate, and industry risk data of the industry in which the target customer is engaged;
[0060] An approval unit for generating a credit approval recommendation and a personalized loan plan based on the credit risk assessment result output by the neural network model.
[0061] The technical solution proposed in the present application has the following beneficial technical effects:
[0062] (1) By collecting multi-dimensional behavior data of the target customer (including financial, social, consumption, geographical location, and network behavior data), the present invention can comprehensively evaluate the credit status of the customer from multiple aspects, providing a more accurate credit risk assessment result compared to relying only on traditional financial data. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a flowchart of a credit assessment method based on artificial intelligence and big data technology provided in the first embodiment of the present application. DETAILED DESCRIPTION
[0064] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.
[0065] The first embodiment of the present application provides a credit assessment method based on artificial intelligence and big data technologies. Please refer to Figure 1 , which is a schematic diagram of the first embodiment of the present application. The following will be combined with Figure 1 to detail a credit assessment method based on artificial intelligence and big data technologies provided by the first embodiment of the present application.
[0066] Step S101: Collect multi-dimensional behavior data of the target customer. The multi-dimensional behavior data includes financial behavior data, social behavior data, consumption behavior data, geographical location behavior data, and network behavior data.
[0067] It should be noted that the multi-dimensional behavior data here does not necessarily include all dimensions of behavior data. For example, financial behavior data, social behavior data, consumption behavior data, geographical location behavior data, and network behavior data can also be one or more of these dimension data.
[0068] First of all, financial behavior data can be collected by cooperating with financial institutions, third-party payment platforms, or data providers authorized by the customer to obtain information such as the target customer's bank account information, credit card transaction records, loan repayment history, investment behavior, etc.
[0069] Secondly, the collection of social behavior data requires the explicit authorization of the user and is carried out through legal third-party data service providers. This type of data mainly comes from social networking platforms, analyzing the user's interaction behavior, friendship relationship, changes in the social circle, and activity on the platform.
[0070] Consumption behavior data is mainly collected through the customer's shopping records, online consumption situations, etc. This can be obtained through the data interfaces of e-commerce platforms and payment platforms. The data types include information such as shopping categories, consumption frequencies, consumption amounts, payment methods, etc. These data can reflect the customer's consumption habits, expenditure stability, and financial health status. For large e-commerce platforms, the customer's consumption data can be directly obtained through the platform API and combined with the data of the payment platform authorized by the user to provide a more detailed consumption portrait.
[0071] The collection of geographical location behavior data is obtained through the GPS information of mobile devices, payment location records, and other location-related services. This data mainly reflects the daily travel patterns of target customers, changes in work and residence locations, and travel frequencies. Data sources include navigation applications used by customers, location information records of payment platforms, and authorized data provided by other mobile operators.
[0072] The collection of network behavior data is achieved by collaborating with third-party data providers authorized by users to obtain search records of target customers on the network, types of websites browsed, access frequencies, and online duration distributions.
[0073] Step S102: Obtain multi-dimensional behavior characteristics of the target customer based on the multi-dimensional behavior data. Among them, the multi-dimensional behavior characteristics include the behavior stability index, abnormal behavior frequency, behavior trend data, and multi-dimensional behavior correlation data of the target customer; the behavior stability index is used to evaluate the consistency and volatility of the target customer in each behavior dimension; the abnormal behavior frequency refers to the frequency of abnormal events that deviate from the normal state in the behavior pattern of the target customer; the behavior trend data is used to reflect the change trend and development speed of the target customer's behavior over time; the multi-dimensional behavior correlation data is used to reflect the mutual correlation between behavior data in different dimensions.
[0074] Step S102 involves extracting multi-dimensional behavior characteristics of the target customer based on multi-dimensional behavior data. These behavior characteristics include the behavior stability index, abnormal behavior frequency, behavior trend data, and multi-dimensional behavior correlation data. The generation process of each characteristic requires detailed calculation and analysis to ensure its accuracy and effectiveness in evaluating the credit risk of the target customer.
[0075] First, the calculation of the behavior stability index is achieved by analyzing the changes in behavior data of the target customer over different time periods. Specifically, time series analysis is performed on each type of behavior data (such as financial behavior, social behavior, consumption behavior, etc.), and divided into several fixed time windows (such as daily, weekly, or monthly). Within each time window, the volatility of each behavior data is calculated, and statistical methods (such as variance, standard deviation) are used to measure the amplitude of these fluctuations. Next, based on the volatility data within these time windows, the time-weighted average value is calculated, and higher weights are given to data closer to the current time to reflect the stability of the customer's recent behavior. In this way, the behavior stability index reflecting the consistency and volatility of the customer's behavior is obtained, and this index is used to evaluate the credit stability of the customer.
[0076] Secondly, the detection of abnormal behavior frequency is achieved by establishing a baseline of behavior patterns and continuously monitoring them. First, historical behavior data of target customers over a period of time (e.g., 6 months) is collected to establish a baseline behavior pattern. These baseline data are used to set the threshold for deviating from normal behavior. When the new behavior data of a customer exceeds the set threshold compared with the historical baseline, these behaviors are marked as abnormal behaviors. The system generates abnormal behavior frequency data by counting the frequency of these abnormal behaviors. To improve accuracy, machine learning algorithms can be used to dynamically adjust the threshold to ensure that over time, the system can adapt to changes in customer behavior and identify abnormal behavior patterns.
[0077] For the generation of behavior trend data, time series analysis methods are adopted to record the changes in the behavior of target customers over different time periods. By performing differencing on the behavior data of each time period, the change rate and direction of the behavior are calculated. This step can not only identify the overall trend of customer behavior but also capture the speed and magnitude of behavior changes. For example, consumption habits in financial behavior may increase or decrease sharply in the short term. The system generates corresponding behavior trend data by analyzing these change rates and uses them for further assessment of credit risk.
[0078] Finally, the calculation of multi-dimensional behavior correlation data is achieved through statistical methods and correlation analysis tools. The specific steps are as follows: normalize the behavior data of customers in different dimensions, and then calculate the correlation between these data. For example, the Pearson correlation coefficient or other common correlation calculation methods can be used to analyze the relationships between financial behavior and consumption behavior, geographical location behavior and social behavior, etc. These correlation data are used to evaluate the interactive effects between various behaviors, thereby providing more dimensional information support for subsequent credit risk assessment.
[0079] Through the above steps, the system can effectively extract the multi-dimensional behavior characteristics of customers, which will be used in the subsequent neural network model to generate more accurate credit risk assessment results.
[0080] Furthermore, the behavior stability index is calculated by the following formula:
[0081] ;
[0082] According to the above formula, the behavior stability index BSI is a complex calculation that combines multi-dimensional data and is used to quantify the behavior stability of target customers. The following is a detailed description of each parameter in Formula 1.
[0083] 1. BSI represents the behavior stability index:
[0084] BSI represents the comprehensive behavioral stability of customers across all behavioral dimensions. It evaluates the stability of customer behavior by combining factors such as behavioral data, volatility, and time-dependence in each dimension. A higher BSI value indicates more stable customer behavior; conversely, a lower value suggests greater volatility or abnormal changes in customer behavior.
[0085] 2. Represents the behavioral data of the dimension:
[0086] Represents the behavioral data of the target customer in the dimension, such as financial behavior, consumption behavior, social behavior, etc. The behavioral data for each dimension is obtained by monitoring the customer's daily activities or transaction behaviors. Specifically, financial behavior may include income and expenditure, and consumption behavior may include shopping frequency and amount, etc. The data for each dimension serves as an independent input.
[0087] 3. and represent the historical mean and historical standard deviation:
[0088] is the historical mean of the behavioral data in the dimension, obtained by long-term observation of the customer's behavioral data in this dimension, representing the long-term average behavioral pattern of the customer in this dimension.
[0089] is the standard deviation of the behavioral data in the dimension, representing the volatility or discreteness of the behavioral data, reflecting the amplitude of behavioral changes of the customer in this dimension. A higher standard deviation indicates unstable and more volatile behavior of the customer in this dimension, while a lower value indicates more stable behavior.
[0090] 4. and represent adjustment parameters:
[0091] and are parameters used to adjust the deviation of behavioral data and the impact of the standard deviation. These parameters are dynamically adjusted according to specific application requirements to enhance or weaken the weight of certain behavioral characteristics in the overall stability index.
[0092] Adjusts the weight of the behavioral data deviating from the historical mean , which reflects the degree of difference between the customer's current behavior and past behavior. If is larger, the impact on the deviation is greater.
[0093] It is the weight adjustment of the standard deviation , controlling the impact of historical fluctuations on the stability index. A higher means a greater impact of the standard deviation, that is, the behavioral volatility has a more significant impact on the stability of this dimension.
[0094] 5. represents the time weight:
[0095] is the time weight of the behavioral data of the dimension, used to reflect the time dependence of the customer's behavior in this dimension. This weight is usually based on the customer's behavior history and obtained by analyzing the trend of behavior over time. represents the customer's behavior pattern within a specific time. For example, a larger may indicate the long-term stability of the customer's behavior in this dimension, while a smaller may mean that the behavior in this dimension has relatively short-term volatility.
[0096] 6. represents the fluctuation correction factor:
[0097] is the fluctuation correction factor of the dimension, describing the fluctuation frequency or intensity of this behavioral dimension. By performing frequency analysis on the historical behavioral data of this dimension, is obtained.
[0098] For example, if the customer's behavior changes frequently in this dimension, the fluctuation correction factor will be larger, reflecting the instability of the behavior. If the behavioral fluctuations are small, will be smaller, indicating that the customer's behavior in this dimension is relatively stable.
[0099] 7. represents the time interval:
[0100] is the time interval between the latest behavioral data and the previous behavioral data in the dimension. It is used to measure the timeliness of the customer's behavior change. If the behavior interval is short (i.e., the customer's behavior changes frequently), then is smaller; on the contrary, a longer interval indicates that the behavior changes are relatively sparse.
[0101] 8. and represent the control parameters:
[0102] and It is a parameter that controls time-dependence and fluctuations, and is used to dynamically adjust the impact of time intervals on the stability of customer behavior.
[0103] Controls the sensitivity of time-dependence. A higher indicates that the time interval of behavioral data has a greater impact on stability, while is used to offset the time interval and adjust the non-linear impact of time changes on behavioral stability.
[0104] 9. Represents the weight:
[0105] Is the weight of the dimensional behavioral data. It represents the contribution degree of each dimension to the overall behavioral stability index. The setting of weights is usually based on business requirements or the importance of this dimension in the overall credit risk of customers. For example, financial behavior may be more important than social behavior for credit risk assessment, so the weight of financial behavior will be relatively high.
[0106] 10. Represents the total number of dimensions:
[0107] Is the total number of dimensions of customer behavioral data, indicating how many different types of behavioral data the system has analyzed. Usually, these dimensions include financial behavior, consumption behavior, social behavior, geographical location behavior, and network behavior, etc.
[0108] Through this formula, the system can comprehensively consider factors such as the volatility, time-dependence, and weight of each behavioral dimension, and dynamically calculate the customer's Behavioral Stability Index (BSI). This method ensures a comprehensive quantification of customer behavior. Especially in complex, multi-dimensional scenarios, it can capture the subtle changes in customer behavior and ensure the accuracy of the assessment.
[0109] Furthermore, the calculation of the behavioral stability index is achieved through a time-weighted average method, specifically including:
[0110] Perform time slicing on multi-dimensional behavioral data, dividing the behavioral data into different time windows;
[0111] Calculate the fluctuation range of the behavioral characteristics of each dimension within each time window;
[0112] Perform a weighted average on the behavioral fluctuations of each time window, where more recent time windows are assigned greater weights to evaluate the stability of the target customer's current behavior and generate the behavioral stability index.
[0113] First, the processing of behavioral data is carried out through time slicing technology, which divides multi-dimensional behavioral data (including financial behavior, social behavior, consumption behavior, etc.) at fixed time intervals. Specifically, these time windows can be set as daily, weekly or monthly according to actual needs. Each time window contains all the behavioral data of the target customer within that time period, thus forming time-series behavioral information. This time slicing processing ensures that subsequent analysis can identify the time dependence of behavioral data and lays the foundation for the calculation of the behavioral stability index.
[0114] Next, the system analyzes the behavioral data within each time window and calculates the fluctuation range of each behavioral feature. The fluctuation range reflects the change amplitude of the customer's behavioral data within a specific time window. Specifically, the calculation of the fluctuation range can be achieved through common statistical methods such as standard deviation or variance to measure the fluctuation degree of a certain behavior within that time window. For different behavioral dimensions (such as the amount fluctuation of consumption behavior, the frequency change of financial transactions, etc.), the system will calculate their respective fluctuation ranges to comprehensively measure the customer's behavioral changes.
[0115] After calculating the behavioral fluctuations of each time window, the system will perform a weighted average on these fluctuation data. The core of the weighted average lies in assigning different weights to the data of different time windows. Since the customer's recent behavior is more valuable for the current credit risk assessment, relatively recent time windows are usually assigned larger weights. Although relatively distant time windows also participate in the calculation, their weights are relatively low. The weight assignment can be adjusted according to business rules or automatically through algorithms, making the calculation result more flexible and able to adapt to the actual fluctuations of the customer's behavior. Such a weight design helps to improve the accuracy of the behavioral stability index, especially when there are significant fluctuations in the customer's recent behavior, it can reflect this change more promptly.
[0116] Finally, the system will generate a behavioral stability index based on the results of these weighted calculations. This index comprehensively considers the behavioral fluctuations of the customer within each time window and time factors, and obtains the overall stability of the customer's behavior through quantification. As an important credit assessment indicator, the behavioral stability index can help credit institutions better judge the credit risk of customers, especially playing a key role in evaluating the consistency and short-term volatility of customer behavior.
[0117] Furthermore, the detection of the abnormal behavior frequency is achieved through threshold setting based on historical behavior patterns and an adaptive learning mechanism, including:
[0118] By analyzing the behavioral data of the target customer in the past 6 months to 1 year, a behavioral baseline pattern is established;
[0119] Set a threshold that deviates from the historical behavior pattern, and behavior events exceeding this threshold are marked as anomalies;
[0120] Apply an adaptive learning algorithm to update the threshold in real time based on the new behavior data of the target customer, making the model's detection of abnormal behavior more accurate.
[0121] First, the system needs to collect and analyze the behavior data of the target customer, which usually involves multi-dimensional behavior data such as financial transactions, consumption behavior, social interactions, and geographical location changes. By deeply analyzing this behavior data over the past 6 months to 1 year, the system can identify the customer's regular behavior pattern. This baseline pattern not only reflects the characteristics of the customer's daily behavior but also includes the fluctuation range and frequency of these behaviors. For example, the customer's financial transactions may have a fixed periodicity and range, the consumption habits remain stable within a specific category and amount range, and the frequency and intensity of social activities also show a certain regularity. Through the analysis of this data, the system can form a comprehensive and quantitative behavior baseline.
[0122] After establishing the baseline pattern, the system sets a threshold for judging abnormal behavior. This threshold represents the allowable range for the customer's current behavior to deviate from their historical pattern. If the customer's behavior exceeds this preset threshold, the system will mark this behavior as abnormal. For example, if the customer's consumption amount suddenly increases significantly beyond the preset range of their normal consumption pattern, the system will mark this event as abnormal behavior. The setting of the threshold is usually achieved through statistical analysis, specifically by defining the degree of deviation based on the variance or standard deviation of the historical behavior data. A higher variance may indicate greater fluctuations in the customer's behavior, so the threshold will be relatively higher, while for customers with more stable behavior, the threshold will be relatively lower to capture subtle anomalies.
[0123] During the process of detecting abnormal behavior, an adaptive learning algorithm is applied to dynamically adjust this threshold. Over time, the customer's behavior may change, especially under the influence of the external environment or personal life changes. To ensure the accurate detection of abnormal behavior at all times, the system will update the threshold in real time through an adaptive learning mechanism. Specifically, when new behavior data is continuously input, the system will automatically adjust the deviation threshold based on the comparison between the new data and the historical baseline. If the customer's behavior gradually changes, such as a change in consumption habits, the system can learn this trend and gradually increase or decrease the threshold to maintain the accuracy of detecting abnormal behavior. This adaptive learning mechanism ensures that the model can flexibly respond to the dynamic changes in the customer's behavior and avoid overly frequent or insufficient abnormal behavior alerts.
[0124] The system ensures the accurate detection of abnormal behaviors through the establishment of historical behavior patterns, the dynamic setting of thresholds, and the application of an adaptive learning mechanism. This method can not only quickly identify abnormal events in customer behaviors but also continuously optimize the detection criteria as customer behaviors change, thus providing reliable data support for credit risk assessment.
[0125] Furthermore, the calculation of the multi-dimensional behavior correlation data adopts a method based on the Pearson correlation coefficient, specifically including:
[0126] Normalize the behavioral characteristic data of the target customer in different dimensions;
[0127] Calculate the correlation between the behavioral data of each dimension, including financial behavior and consumption behavior, social behavior and geographical location behavior, etc.;
[0128] Use the Pearson correlation coefficient to evaluate the linear correlation between the behavioral data of different dimensions, generate a correlation matrix, and use it as the input for the subsequent credit risk assessment model.
[0129] First, the system will normalize the multi-dimensional behavioral data of the target customer. This step is to ensure the comparability of data in different dimensions because there are usually differences in the order of magnitude among different types of behavioral data. For example, financial behavioral data may involve large transaction amounts, while social behavioral data may be related to interaction frequencies. Before normalization, these data may lead to biases in calculating correlations. Therefore, normalization can scale the data of each behavioral dimension to the same range, making them consistent in correlation calculations and ensuring that each behavioral characteristic can be reasonably considered.
[0130] Then, the system will calculate the correlation between the behavioral data of each dimension. Here, the different dimensions usually include financial behavior, consumption behavior, social behavior, geographical location behavior, and network behavior. By comparing and analyzing these behavioral data, the system can identify the potential connections between them. For example, the correlation between financial behavior and consumption behavior may reflect the relationship between the customer's financial situation and consumption habits; while the correlation between social behavior and geographical location behavior may reveal the relationship between the customer's social activities and changes in their geographical location. The relationships between these dimensions are of great significance for evaluating the overall credit risk of customers because they can provide additional context information to help identify hidden risk patterns.
[0131] In correlation calculation, the Pearson correlation coefficient is used to evaluate the linear relationship between two behavior dimensions. The Pearson correlation coefficient is a common statistical method for measuring the degree of linear correlation between two variables, with a value range between -1 and 1. A coefficient of 1 indicates a perfect positive correlation between the two variables, -1 indicates a perfect negative correlation, and 0 indicates no correlation. In this method, the system will calculate the Pearson correlation coefficient between different dimension behavior data pair by pair. For example, the system will calculate the correlation coefficient between financial behavior and consumption behavior, the correlation coefficient between social behavior and geographical location behavior, etc. Each calculation result reflects the strength of the linear relationship between different dimension behaviors.
[0132] Finally, the system generates a correlation matrix based on the calculated Pearson correlation coefficients. Each element in this matrix represents the correlation between two behavior data, visually presenting the relationships between all dimension behaviors in the form of a matrix. This correlation matrix not only provides important reference data for credit assessment but also serves as the input for the subsequent credit risk assessment model, helping the system comprehensively consider the multi-dimensional behavior characteristics of customers when conducting credit assessment.
[0133] This method combines data normalization processing, correlation calculation, and the application of the Pearson correlation coefficient to ensure the accurate analysis of the relevance of customers in each behavior dimension, providing a solid data foundation for credit risk assessment.
[0134] Furthermore, the acquisition and update of the environmental risk assessment data are realized through a big data-based dynamic macroeconomic monitoring system, specifically including:
[0135] Obtain the latest GDP growth rate, unemployment rate, and industry risk data from government statistical departments, industry reports, and third-party economic data providers;
[0136] Regularly update the economic indicators of the industries in which the target customers are engaged, and dynamically adjust the credit risks of the target customers in combination with macroeconomic environments such as economic recession and rising industry risks.
[0137] First, the system obtains the latest macroeconomic data from multiple sources. These data sources include government statistical departments, industry reports, and third-party economic data providers. Specifically, government statistical departments usually release authoritative economic indicators, such as the GDP growth rate and unemployment rate of a country or region. These data directly reflect the overall economic health. The GDP growth rate can be used as a sign of economic growth or contraction, while the unemployment rate reflects the supply and demand situation in the labor market. Industry reports provide detailed analyses of a specific industry, including current market demand, technological innovation, and competition status. Such information is particularly useful for assessing the risk level of the industry in which the target customers are engaged. Third-party economic data providers also offer a large amount of real-time updated economic data, covering key information such as global and regional economic trends, policy changes, etc. These data sources are integrated into the system through API interfaces or regular downloads to ensure the timeliness and accuracy of the data.
[0138] After the data collection is completed, the system processes these economic indicators and conducts an analysis on the specific industry in which the target customers are engaged. At this time, the system does not merely stay at the general macroeconomic level but delves into the specific economic environment of the customer's industry. For example, a certain industry may face the risk of economic recession, the uncertainty brought about by technological changes, or market shrinkage due to supply chain problems. These industry-specific economic factors will directly affect the credit risk of the customers. Therefore, the system regularly updates the relevant economic data according to the industry in which each customer is engaged and analyzes the possible impact of economic fluctuations inside and outside the industry on the future credit performance of the customers.
[0139] After obtaining and analyzing these macroeconomic and industry data, the system dynamically adjusts the credit risk assessment results of the customers. Specifically, the system automatically adjusts the credit risk scores of the customers according to changes in economic indicators, such as a slowdown in GDP growth rate, an increase in unemployment rate, or an intensification of industry risks. For example, if the industry in which the target customer is located is severely impacted due to an economic recession, the system will correspondingly increase its credit risk score to indicate the potential default possibility. On the contrary, in the case of economic recovery or an optimistic industry outlook, the system will appropriately lower the credit risk score of the customer to reflect its stable credit status in a good economic environment.
[0140] Step S103: Based on a pre-trained neural network model, generate a credit risk assessment result according to the credit investigation report data, environmental risk assessment data, multi-dimensional behavior data, and multi-dimensional behavior characteristics of the target customer; wherein, the credit risk assessment result includes a credit risk score, the identification result of key risk factors, and the prediction of future credit risk; the environmental risk assessment data includes economic indicator data, and the economic indicator data includes the GDP growth rate, the unemployment rate, and the industry risk data of the industry in which the target customer is engaged.
[0141] Step S103 involves generating a credit risk assessment result for the target customer based on a pre-trained neural network model. In this process, the system needs to use the multi-dimensional behavioral data, behavioral characteristics, credit report data, and environmental risk assessment data of the target customer as inputs. Through the calculation of the neural network model, the final credit risk assessment result is obtained. The specific implementation of this step requires multi-level data input and complex deep learning models to cooperate and process to ensure that the output results are accurate and comprehensive enough.
[0142] Furthermore, the neural network model includes a feature extraction network, an anomaly detection network, and a multi-layer fusion network;
[0143] First, the feature extraction network is responsible for processing the multi-dimensional behavioral data and multi-dimensional behavioral characteristics of the target customer. The multi-dimensional behavioral data usually includes various information such as finance, consumption, social, geographical location, and network behavior. To extract the key features from this complex data, the feature extraction network adopts the structure of a convolutional autoencoder. The convolutional autoencoder consists of two parts: an encoder and a decoder. The encoder part gradually compresses the input original behavioral data through multi-layer convolutional operations, extracts representative local features, and represents this information as a low-dimensional feature vector. This process helps to remove the noise in the data and retain important behavioral patterns. Next, the decoder reconstructs the compressed feature vector through transposed convolutional operations to evaluate whether the encoding accurately captures the key information in the input data. Finally, the output of the feature extraction network is a non-linear behavioral feature vector, which is a highly abstract feature representation that can reflect the complexity and internal connections of customer behavior.
[0144] Then, the task of the anomaly detection network is to identify abnormal behaviors based on the non-linear behavioral feature vector provided by the feature extraction network. The core of this network adopts a recurrent neural network based on the multi-head self-attention mechanism (RNN with Multi-head Self-Attention). In this network structure, the recurrent neural network is responsible for processing time series data and identifying abnormal behaviors in the behavioral patterns by capturing the temporal changes of the target customer's behavior. The output of this network includes two key results: the abnormal behavior detection result, which represents the behavior events found by the system that do not conform to the normal behavior pattern of the customer; the prediction probability of abnormal behavior, which represents the possibility that the customer will exhibit abnormal behavior again in the future. These outputs provide important reference bases for subsequent credit risk assessments.
[0145] Finally, the multi-layer fusion network integrates inputs from multiple data sources, including the non-linear behavior feature vectors of the feature extraction network, the abnormal behavior detection results and prediction probabilities of the anomaly detection network, credit report data, and economic indicator data. The multi-layer fusion network adopts the Capsule Network structure. By processing the multi-dimensional features and hierarchical relationships of the data, the Capsule Network can better maintain the interdependence and spatial relationships between different input features. Each capsule unit processes a different set of features and passes important information to the next layer through a dynamic routing mechanism to ensure that key information is not lost when processing complex multi-source data. Finally, the Capsule Network generates the credit risk assessment results of the target customers, which not only include the current credit status scores of the customers but also provide the possible future trends of credit risk changes by combining economic environment data and abnormal behavior predictions.
[0146] Furthermore, the multi-layer convolutional autoencoder adopted by the feature extraction network specifically includes:
[0147] The first input layer receives the multi-dimensional behavior data and multi-dimensional behavior features of the target customers. The multi-dimensional behavior data can include the customer's financial behavior, consumption behavior, social behavior, etc. Through this input layer, all data is uniformly input into the network as the basic data for subsequent convolutional operations.
[0148] In the first convolutional layer, the system uses 64 convolutional kernels of size 3×3 to perform convolutional operations on the input multi-dimensional behavior data. The role of convolution is to extract local features in the behavior data, such as local temporal patterns or correlations between behaviors. This convolutional operation generates 64 feature maps, which are the initial abstract representations of the original data and extract the basic patterns in the data.
[0149] Next, the first pooling layer receives the output of the first convolutional layer and performs max-pooling operations using a pooling window of size 2×2. The role of pooling is to reduce the spatial dimension of the feature maps, that is, to compress the data, making subsequent processing more efficient while retaining the most important features. The pooling operation outputs 64 pooled feature maps, which retain the main information of the local features while reducing the data redundancy.
[0150] Subsequently, the second convolutional layer receives the output of the first pooling layer and uses 128 convolutional kernels of size 3×3 to perform deeper convolutional operations. Different from the first convolutional layer, the purpose of the second convolutional layer is to capture more complex temporal dependencies and non-linear behavior features. The convolution of this layer can identify deeper feature patterns, thus generating 128 new feature maps.
[0151] The second pooling layer receives the output of the second convolutional layer and performs max pooling again using a 2×2 pooling window to further reduce the spatial dimension of the features and outputs 128 pooled feature maps. This step ensures that the feature maps are further compressed while retaining key information, laying the foundation for subsequent encoding processing.
[0152] The encoding layer receives the output of the second pooling layer and compresses the feature vector to 256 dimensions through a fully connected layer. This step further compresses the high-dimensional behavioral features into a high-dimensional non-linear vector. This 256-dimensional vector contains highly abstract features of the target customer behavior, representing complex patterns in multi-dimensional data. This vector serves as an intermediate result of the feature extraction network and will be used in the subsequent decoding stage.
[0153] During the decoding process, first, the first deconvolutional layer receives the 256-dimensional feature vector output by the encoding layer and performs deconvolution operations using 128 3×3 convolutional kernels. This step aims to gradually restore the spatial dimension of the data, generating 128 feature maps and reconstructing the structural information of part of the original behavioral data.
[0154] Subsequently, the first upsampling layer receives the output of the first deconvolutional layer and further restores the spatial resolution of the data by performing a 2×2 upsampling operation, generating 128 enlarged feature maps. This operation can expand the size of the feature maps, making them gradually approach the spatial structure of the original input data.
[0155] Next, the second deconvolutional layer receives the output of the first upsampling layer and performs further deconvolution operations using 64 3×3 convolutional kernels. This operation aims to restore a higher spatial resolution, generating 64 feature maps. These maps are similar to the output of the first convolutional layer but have undergone multiple layers of processing to extract deeper features.
[0156] The second upsampling layer receives the output of the second deconvolutional layer and performs a 2×2 upsampling operation to restore the feature maps to a spatial resolution similar to that of the first convolutional layer. At this point, the spatial structure of the data has been basically restored, and it is ready to enter the final output layer.
[0157] Finally, the first output layer receives the output of the second upsampling layer and further processes the data through an additional convolutional layer to generate reconstructed data similar to the first input layer. Compared with the original input data, these reconstructed data retain the key information while removing unnecessary noise. In addition, the first output layer also generates a 256-dimensional high-dimensional non-linear behavior feature vector, representing the abstract behavior characteristics of the target customer, and this vector will be used in the subsequent credit assessment process. Through this combination of convolution and deconvolution operations, the system can effectively extract and restore the multi-dimensional behavior characteristics of the target customer, providing a solid data foundation for the subsequent evaluation tasks.
[0158] Furthermore, the anomaly detection network adopts a recurrent neural network based on the multi-head self-attention mechanism, including:
[0159] A second input layer for receiving the non-linear behavior feature vector output from the feature extraction network; a spatio-temporal graph convolutional layer for receiving the non-linear behavior feature vector provided by the second input layer. The spatio-temporal graph convolutional layer constructs a spatio-temporal graph based on the graph convolutional network and the time dimension, specifically for:
[0160] By arranging the multi-dimensional behavior data and multi-dimensional behavior characteristics of the target customer in chronological order, a spatio-temporal graph is formed, where the nodes represent the behavior states at different time points, and the edges represent the behavior relationships between different time points;
[0161] Performing a convolution operation on the nodes in the spatio-temporal graph through the graph convolutional network to capture the dependence of the user's behavior in the time dimension and the spatial characteristics of the behavior data, and obtaining a 128-dimensional feature vector after spatio-temporal convolution;
[0162] A multi-head self-attention layer for receiving the 128-dimensional feature vector output from the spatio-temporal graph convolutional layer, and using the multi-head self-attention mechanism to weight the key behavior patterns to generate a 1024-dimensional feature vector;
[0163] A recurrent neural network layer for receiving the 1024-dimensional feature vector from the multi-head self-attention layer, and combining the long short-term memory network to process the time series characteristics of the behavior data, and the output is a 512-dimensional time series feature vector, representing the time characteristics of the abnormal behavior;
[0164] A second output layer for receiving the 512-dimensional time series feature vector provided by the recurrent neural network layer, and generating the abnormal behavior prediction probability and abnormal behavior detection result of the target customer.
[0165] The anomaly detection network adopts a recurrent neural network based on the multi-head self-attention mechanism to identify and predict the abnormal behaviors of target customers. This network consists of multiple layers, gradually processing the non-linear behavior feature vectors from the feature extraction network, and finally generating the prediction probability and detection result of abnormal behaviors.
[0166] First, the second input layer of the anomaly detection network receives the non-linear behavior feature vectors from the feature extraction network. These feature vectors are highly abstract behavior representations, extracted from the multi-dimensional behavior data of target customers, including financial behaviors, consumption behaviors, social behaviors, etc. These feature vectors capture the complexity and potential patterns of customer behaviors and serve as the input for subsequent layers.
[0167] Next, the spatio-temporal graph convolutional layer receives this non-linear behavior feature vector and further analyzes the dependencies of customer behaviors in the time and space dimensions by constructing a spatio-temporal graph. The nodes of the spatio-temporal graph represent the customer behavior states at different time points, and the edges represent the temporal correlations between these behaviors. By arranging the multi-dimensional behavior data of the target customer in chronological order, the spatio-temporal graph can effectively capture the temporal changes and spatial characteristics of behaviors. The graph convolutional network performs convolutional operations on the nodes in the spatio-temporal graph. The convolution process can extract the relationships and patterns between nodes and generate a 128-dimensional feature vector. These feature vectors not only contain the changes in behaviors at different time points but also reflect the temporal dependencies between behaviors, effectively describing the dynamic changes of customer behaviors.
[0168] After the spatio-temporal graph convolutional layer, the multi-head self-attention layer receives the 128-dimensional feature vector and weights the key behavior patterns through the self-attention mechanism. The core idea of the multi-head self-attention mechanism is to process data in parallel on multiple "heads" to enhance the network's attention ability to different feature dimensions. Each "head" independently calculates the attention weights to identify the key behavior patterns in the input feature vector. Through this mechanism, the network can more accurately identify the possible anomalies in customer behaviors. The output results of each head will be concatenated to generate a 1024-dimensional feature vector, which represents the important patterns and anomaly possibilities in customer behaviors.
[0169] Then, the recurrent neural network layer receives the 1024-dimensional feature vector from the multi-head self-attention layer and further processes the time series features of these data through a long short-term memory network. The LSTM can capture the changes in behaviors over a long time span and handle short-term and long-term behavior dependencies. The recurrent neural network will output a 512-dimensional time series feature vector, representing the characteristics of customer abnormal behaviors in the time dimension. These feature vectors not only consider the current behavior state but also combine the customer's past behavior sequences, thus improving the timeliness and accuracy of anomaly behavior detection.
[0170] Finally, the second output layer receives the 512-dimensional time series feature vector from the recurrent neural network layer and generates the prediction probability of the abnormal behavior of the target customer and the detection result. The prediction probability represents the likelihood of the customer having abnormal behavior within a certain future time period, and the detection result determines whether the customer has had abnormal behavior based on historical behavior data. These output data will be used for subsequent credit risk assessment and decision-making, providing an important reference for credit approval.
[0171] Furthermore, the multi-layer fusion network is implemented using a dynamic routing capsule network, and the dynamic routing capsule network includes:
[0172] A third input layer for receiving the non-linear behavior feature vector provided by the feature extraction network, the abnormal behavior detection result and prediction probability provided by the abnormal detection network, the credit report data, and the economic indicator data; the input data of the third input layer is integrated into a 919-dimensional feature vector as the input for the subsequent capsule layer;
[0173] A first capsule layer for receiving the 919-dimensional feature vector of the third input layer and processing different types of input features through multiple capsule units. Each capsule unit contains 8 neurons and is responsible for processing data in different dimensions; the output of each capsule unit is an 8-dimensional vector representing the feature information of the input feature in different dimensions; and after passing through the first capsule layer, the output is a 64-dimensional feature vector. The dynamic routing mechanism of the first capsule layer includes: dynamically calculating the information transfer weight between capsule units according to the output of the capsule units; the vector length of the output of the capsule unit represents the feature intensity, and dynamically adjusting the information transfer path between different capsule units to strengthen the output of the capsule related to the key feature;
[0174] A second capsule layer for receiving the 64-dimensional vector from the first capsule layer and further fusing and processing the feature information from different sources. The second capsule layer includes multiple first capsule units, where each first capsule unit includes 16 neurons for processing multi-dimensional features; the second capsule layer divides and inputs the 64-dimensional feature vector into each capsule unit through a dynamic routing mechanism; the output of the second capsule layer is a 128-dimensional feature vector, reflecting the deep association between the target customer's behavior data, economic data, and abnormal detection results;
[0175] A feature fusion layer for receiving the 128-dimensional feature vector output by the second capsule layer and performing feature fusion through a fully connected layer; the output of the feature fusion layer is a 64-dimensional fused feature vector;
[0176] A third output layer for receiving the 64-dimensional fused feature vector from the feature fusion layer and generating the comprehensive credit risk assessment result of the target customer. The third output layer is specifically used for:
[0177] Through the third fully connected layer, a 64-dimensional feature vector is mapped to a credit risk score, representing the credit risk level of the target customer;
[0178] Through the fourth fully connected layer, it is used to identify and output the key factors affecting the customer's credit risk, and the output is a set of classification labels or specific numerical values; among them, the key factors include abnormal behaviors, credit history, and changes in the economic environment;
[0179] Through the time series analysis layer, the 64-dimensional feature vector is combined with historical behavior data for prediction, generating the changing trend of credit risk in a specified future period, and the output is a future risk score, representing the potential credit risk of the customer in the future.
[0180] First, the third input layer receives feature data from multiple sources, including the non-linear behavior feature vector provided by the feature extraction network, the abnormal behavior detection results and prediction probabilities provided by the anomaly detection network, as well as the customer's credit report data and macroeconomic indicators. These data are integrated into a 919-dimensional feature vector, which combines the customer's behavior characteristics, historical credit information, and external economic environment, serving as the basis for subsequent processing. The input data undergoes preprocessing to ensure that data in different dimensions have a consistent format and comparability.
[0181] Next, the first capsule layer receives the 919-dimensional feature vector from the third input layer and uses multiple capsule units to process different types of input features in parallel. Each capsule unit contains 8 neurons and is responsible for processing specific feature dimensions. The output of the capsule unit is an 8-dimensional vector, representing the feature information of the input feature in different dimensions. This vectorized output can capture the key information in the input data and optimize the information transmission path between capsule units through a dynamic routing mechanism. The core of the dynamic routing mechanism is to dynamically adjust the weights and connection paths between different capsule units according to the length of the output vector of the capsule unit to strengthen the output related to key features. This mechanism ensures that the system can focus on the most important feature information while reducing the interference of noise and irrelevant information. After processing by this layer, the first capsule layer outputs a 64-dimensional feature vector, representing the initially extracted feature information.
[0182] Subsequently, the second capsule layer further receives the 64-dimensional vector from the first capsule layer and continues to process and fuse the feature information from multiple data sources. The second capsule layer contains multiple capsule units, each unit containing 16 neurons, which are used to process more complex multi-dimensional features. Through a further dynamic routing mechanism, this layer divides and inputs the 64-dimensional feature vector into each capsule unit. Through this process, the capsule network can capture the deep associations between different features, such as the relationship between a customer's behavior pattern and macroeconomic data. Finally, the second capsule layer outputs a 128-dimensional feature vector, which combines the customer's behavior data, economic data, and anomaly detection results, and reveals the complex interactions between dimensions.
[0183] The feature fusion layer receives the 128-dimensional feature vector from the second capsule layer and fuses these features through a fully connected layer. The fully connected layer unifies and fuses the feature information from different sources through weighted processing to ensure that the system can comprehensively consider the influencing factors of each dimension. After processing by this layer, the output is a 64-dimensional fused feature vector, representing the comprehensive credit risk characteristics of the customer.
[0184] Finally, the third output layer receives the 64-dimensional fused feature vector from the feature fusion layer and generates the comprehensive credit risk assessment result of the customer. This layer achieves the final assessment output through multiple fully connected layers. First, the third fully connected layer maps the 64-dimensional feature vector to a credit risk score, which represents the credit risk level of the customer and is a quantitative result of the customer's overall credit status. Second, the fourth fully connected layer is responsible for identifying and outputting the key factors that affect the customer's credit risk. These key factors can include abnormal behaviors, negative events in the historical credit record, and changes in the macroeconomic environment. The system outputs these key factors in the form of classification labels or numerical values to help credit assessment personnel identify the main sources of the customer's credit risk.
[0185] In addition, the system also predicts the future trend of the customer's credit risk through a time series analysis layer. The time series analysis layer combines the customer's historical behavior data and the current feature vector to predict the change trend of the customer's credit risk in a future period (such as 6 months or 1 year). The system finally outputs a future risk score, which represents the potential credit risk level of the customer in the future period. This prediction can help credit institutions prevent potential credit risks in advance and make more accurate credit decisions.
[0186] The following is the reference implementation code of the neural network model.
[0187] import torch;
[0188] import torch.nn as nn;
[0189] import torch.nn.functional as F;
[0190] # Feature extraction network: Multi-layer convolutional autoencoder
[0191] class FeatureExtractionNetwork(nn.Module):
[0192] def __init__(self):
[0193] super(FeatureExtractionNetwork, self).__init__();
[0194] # The first convolutional layer, receiving multi-dimensional behavioral data, 64 3x3 convolutional kernels;
[0195] self.conv1 = nn.Conv2d(in_channels=1, out_channels=64, kernel_size=3, padding=1);
[0196] # The first pooling layer, a 2x2 pooling window, reducing the spatial dimension;
[0197] self.pool1 = nn.MaxPool2d(kernel_size=2, stride=2);
[0198] # The second convolutional layer, 128 3x3 convolutional kernels, extracting deeper features
[0199] self.conv2 = nn.Conv2d(in_channels=64, out_channels=128, kernel_size=3, padding=1);
[0200] # The second pooling layer;
[0201] self.pool2 = nn.MaxPool2d(kernel_size=2, stride=2);
[0202] # Encoding layer, compressing features into 256 dimensions;
[0203] self.deconv1 = nn.ConvTranspose2d(in_channels=128, out_channels=64, kernel_size=3, padding=1);
[0204] self.up1 = nn.Upsample(scale_factor=2);
[0205] self.deconv2=nn.ConvTranspose2d(in_channels=64, out_channels=1,kernel_size=3, padding=1);
[0206] self.up2 = nn.Upsample(scale_factor=2);
[0207] def forward(self, x):
[0208] # First convolutional layer + ReLU activation;
[0209] x = F.relu(self.conv1(x));
[0210] # First pooling layer, reducing spatial resolution;
[0211] x = self.pool1(x);
[0212] # Second convolutional layer + ReLU activation;
[0213] x = F.relu(self.conv2(x));
[0214] # Second pooling layer;
[0215] x = self.pool2(x);
[0216] # Flatten the feature map and enter the encoding layer to generate a 256 - dimensional non - linear feature vector;
[0217] x = x.view(-1, 128, 8, 8);
[0218] x = F.relu(self.deconv1(x));
[0219] x = self.up1(x);
[0220] x = F.relu(self.deconv2(x));
[0221] x = self.up2(x);
[0222] return x;
[0223] # Anomaly Detection Network: Recurrent Neural Network (LSTM) Based on Multi-Head Self-Attention
[0224] class AnomalyDetectionNetwork(nn.Module):
[0225] def __init__(self, input_size, hidden_size, num_heads):
[0226] super(AnomalyDetectionNetwork, self).__init__();
[0227] # Multi-head self-attention mechanism to process spatio-temporal data
[0228] self.multihead_attention = nn.MultiheadAttention(embed_dim = input_size, num_heads = num_heads);
[0229] # LSTM layer to process time series features
[0230] self.lstm = nn.LSTM(input_size = input_size, hidden_size = hidden_size, num_layers = 2, batch_first = True);
[0231] # Fully connected layer to output detection results and prediction probabilities of abnormal behaviors
[0232] self.fc_out = nn.Linear(hidden_size, 2);
[0233] def forward(self, x):
[0234] # Multi-head self-attention to process feature vectors
[0235] attn_output, _ = self.multihead_attention(x, x, x);
[0236] # LSTM to process time series features
[0237] lstm_out, (hn, cn) = self.lstm(attn_output);
[0238] # Take the LSTM output at the last moment;
[0239] lstm_out = lstm_out[:, -1, :];
[0240] # Output the abnormal behavior detection result and prediction probability;
[0241] output = self.fc_out(lstm_out);
[0242] return output;
[0243] # Multi - layer fusion network: Capsule network implementation;
[0244] class CapsuleLayer(nn.Module):
[0245] def __init__(self, num_capsules, num_routes, in_dim, out_dim):
[0246] super(CapsuleLayer, self).__init__();
[0247] self.num_capsules = num_capsules;
[0248] self.num_routes = num_routes;
[0249] self.in_dim = in_dim;
[0250] self.out_dim = out_dim;
[0251] # Linear layer for each capsule unit;
[0252] self.capsules = nn.ModuleList([nn.Linear(in_dim, out_dim); for _ in range(num_capsules)]);
[0253] def forward(self, x):
[0254] # Calculate the output of each capsule;
[0255] u_hat = [capsule(x) for capsule in self.capsules]; It should be noted that there are some syntax errors in the original Chinese code snippet you provided. For example, in the `def__init__` function definition in line 22, there should be a space between `def` and `__init__`. The above translation is based on the corrected understanding of the code logic.
[0256] u_hat = torch.stack(u_hat, dim=1) # Combine all capsule units;
[0257] b_ij = torch.zeros(x.size(0), self.num_capsules, self.num_routes).to(x.device);
[0258] for i in range(3): # Number of dynamic routing iterations;
[0259] c_ij = F.softmax(b_ij, dim=2) # Calculate weights;
[0260] v_j = self.squash(s_j) # Compressive non - linear activation;
[0261] s_j_norm = torch.norm(s_j, dim=-1, keepdim=True);
[0262] class CapsuleNetwork(nn.Module):
[0263] def __init__(self):
[0264] super(CapsuleNetwork, self).__init__();
[0265] # The third input layer, integrating non - linear feature vectors, anomaly detection results, and economic data;
[0266] self.fc1 = nn.Linear(919, 128);
[0267] # The first capsule layer, 64 capsule units, each capsule with an 8 - dimensional output;
[0268] self.capsule1 = CapsuleLayer(num_capsules = 64, num_routes = 128, in_dim = 128, out_dim = 8);
[0269] # The second capsule layer, 128 capsule units, each capsule with a 16 - dimensional output;
[0270] self.capsule2 = CapsuleLayer(num_capsules=128, num_routes=64, in_dim=8, out_dim=16);
[0271] # Fully connected layer for feature fusion;
[0272] self.fc3 = nn.Linear(64, 1) # Credit risk score;
[0273] self.fc4 = nn.Linear(64, 3) # Key risk factors (3 factors);
[0274] def forward(self, x):
[0275] # Initial processing, non - linear feature vector integration;
[0276] x = F.relu(self.fc1(x));
[0277] # First capsule layer processing;
[0278] capsule1_out = self.capsule1(x);
[0279] # Second capsule layer processing;
[0280] capsule2_out = self.capsule2(capsule1_out);
[0281] # Feature fusion layer;
[0282] x = F.relu(self.fc2(capsule2_out.view(x.size(0), -1)));
[0283] # Output credit risk score;
[0284] credit_score = self.fc3(x);
[0285] # Output key risk factors;
[0286] risk_factors = self.fc4(x)
[0287] return credit_score, risk_factors;
[0288] # Comprehensive neural network model, combining each sub - network together;
[0289] class CreditRiskAssessmentModel(nn.Module):
[0290] def __init__(self):
[0291] super(CreditRiskAssessmentModel, self).__init__();
[0292] # Define each sub-network;
[0293] self.feature_extractor = FeatureExtractionNetwork();
[0294] self.anomaly_detector = AnomalyDetectionNetwork(input_size = 256, hidden_size = 512, num_heads = 8);
[0295] self.capsule_network = CapsuleNetwork();
[0296] def forward(self, behavior_data, credit_report_data, economic_data):
[0297] # Feature extraction network to extract behavior features;
[0298] behavior_features = self.feature_extractor(behavior_data);
[0299] # Anomaly detection network to detect abnormal behaviors and predict future abnormal behaviors
[0300] anomaly_detection_output = self.anomaly_detector(behavior_features);
[0301] # Integrate behavior features, anomaly detection results, credit report, and economic data;
[0302] combined_input = torch.cat([behavior_features, anomaly_detection_output, credit_report_data, economic_data], dim=1);
[0303] credit_score, risk_factors = self.capsule_network(combined_input);
[0304] return credit_score, risk_factors;
[0305] To train this neural network model, first, a training set containing multi-dimensional behavioral data, credit report data, and economic indicator data needs to be prepared. Each data sample should include the behavioral data, credit information, and the corresponding credit risk label of the target customer. The goal of the training process is to optimize the parameters of the network so that the model can accurately predict the credit risk score and effectively detect abnormal behaviors.
[0306] 1. Data Preparation: Divide the training data into input data and the corresponding labels (i.e., the credit risk scores of the customers). The input data includes the multi-dimensional behavioral data, abnormal behavior information, credit report data, and economic indicator data of the target customer. It is necessary to ensure that the data undergoes appropriate preprocessing, such as normalization or standardization, so that the network can learn more effectively.
[0307] 2. Initialize the Model: Create model instances, including a feature extraction network, an anomaly detection network, and a multi-layer fusion network. Combine these sub-networks into a complete model that can extract features from the input multi-dimensional data, detect anomalies, and generate credit risk assessment results.
[0308] 3. Define the Loss Function and Optimizer: Select suitable loss functions and optimizers. For example, the mean squared error (MSE) can be used as the loss function for credit score prediction, or the cross-entropy loss function can be used for the classification of abnormal behaviors. The Adam optimizer can be selected as the optimizer, which can quickly adjust the model parameters.
[0309] 4. Model Training: Input the input data and labels into the model for training. In each training iteration, the model will calculate the prediction results and compare them with the true labels. Through backpropagation, the optimizer will update the weights in the model so that the prediction results gradually approach the true values. During this process, the loss value should gradually decrease, indicating that the prediction ability of the model is gradually improving.
[0310] 5. Model Evaluation: During the training process, a validation dataset is used to periodically evaluate the model's performance. The data in the validation set is not involved in training and is only used to evaluate the model's performance on unseen data. Based on the results of the validation set, model hyperparameters, optimizer settings, etc. can be adjusted to improve the model's generalization ability.
[0311] Step S104: Generate a credit approval recommendation and a personalized loan plan based on the credit risk assessment result output by the neural network model.
[0312] Step S104 aims to provide a credit approval recommendation and a personalized loan plan for the target customer based on the credit risk assessment result generated by the neural network model. Specifically, this step will comprehensively analyze the customer's credit risk score, the identification results of key risk factors, and the future credit risk prediction to support the credit decision-making.
[0313] First, the system will use the credit risk score generated in Step S103 as the core basis. The score value usually ranges from 0 to 1. This score directly reflects the overall credit health of the customer. The higher the score, the lower the credit risk of the customer. For customers with a high credit score, the system will recommend a higher credit limit and more lenient loan terms. For customers with a low credit score, the system will suggest a more conservative credit decision, such as reducing the loan amount or tightening the loan conditions.
[0314] At the same time, the system will conduct an in-depth analysis of the identification results of key risk factors. By identifying the specific factors behind the customer's credit risk, such as a high frequency of abnormal behavior, an unstable economic environment, and overdue credit records, the system can formulate a more targeted loan plan. For example, if the main risk factor identified is related to the customer's fluctuating financial behavior, the system may suggest shortening the loan term or increasing the loan interest rate to reduce potential risks. In addition, the system can stratify customers according to different risk factors and formulate differentiated credit policies for customers in different risk levels to balance loan risks and returns.
[0315] The future credit risk prediction is also an important reference in the decision-making process. The system will combine the future credit risk prediction results within a certain period (such as 6 months or 1 year) to evaluate whether there is an upward trend in the customer's credit risk during the loan period. If the model predicts that the customer's future credit risk will increase significantly, the system may suggest a more conservative credit decision, such as shortening the loan cycle or setting up a re-evaluation mechanism in advance. For customers with relatively stable or decreasing future credit risks, the system will recommend more flexible loan terms, appropriately extending the repayment period or reducing the interest rate to attract high-quality customers.
[0316] When generating a personalized loan plan, the system not only bases on the credit risk assessment results, but also combines the personalized needs and preferences of customers, such as repayment ability, loan purpose, and the repayment method selected by the customer. By analyzing the historical behavior and preferences of customers, the system can recommend loan products that meet their needs and improve customer satisfaction. For example, for customers with stable repayment habits, the system may recommend a plan with a longer repayment period and a lower interest rate, while for customers who prefer to repay quickly, it will suggest a short-term loan product with a higher interest rate.
[0317] Finally, the system integrates all the analysis results to generate credit approval suggestions and personalized loan plans. This result will be submitted to the credit approval department as a decision-making basis, and at the same time, customers will also receive personalized loan plan suggestions to help them make the most suitable choice.
[0318] In the above embodiment, a credit assessment method based on artificial intelligence and big data technology is provided. Correspondingly, the present application also provides a credit assessment device based on artificial intelligence and big data technology. Since this embodiment, that is, the second embodiment, is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment. The embodiments described below are merely illustrative.
[0319] The second embodiment of the present application provides a credit assessment device based on artificial intelligence and big data technology, including:
[0320] A collection unit, configured to collect multi-dimensional behavior data of a target customer, where the multi-dimensional behavior data includes financial behavior data, social behavior data, consumption behavior data, geographical location behavior data, and network behavior data;
[0321] An obtaining unit, configured to obtain multi-dimensional behavior characteristics of the target customer according to the multi-dimensional behavior data, where the multi-dimensional behavior characteristics include a behavior stability index of the target customer, an abnormal behavior frequency, behavior trend data, and multi-dimensional behavior correlation data; the behavior stability index is used to evaluate the consistency and volatility of the target customer in each behavior dimension; the abnormal behavior frequency refers to the frequency of abnormal events in the behavior pattern of the target customer that deviate from the normal state; the behavior trend data is used to reflect the change trend of the target customer's behavior over time and its development speed; the multi-dimensional behavior correlation data is used to reflect the mutual correlation between different dimension behavior data;
[0322] A generation unit, configured to generate a credit risk assessment result based on a pre-trained neural network model according to the credit investigation report data, environmental risk assessment data, multi-dimensional behavior data, and multi-dimensional behavior characteristics of a target customer; wherein, the credit risk assessment result includes a credit risk score, an identification result of key risk factors, and a future credit risk prediction; the environmental risk assessment data includes economic indicator data, and the economic indicator data includes GDP growth rate, unemployment rate, and industry risk data of the industry in which the target customer is engaged;
[0323] An approval unit, configured to generate a credit approval recommendation and a personalized loan plan based on the credit risk assessment result output by the neural network model.
[0324] Although this application is disclosed above with preferred embodiments, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the protection scope of this application shall be subject to the scope defined by this application.
Claims
1. A credit assessment method based on artificial intelligence and big data technology, characterized in that: include: Collect multi-dimensional behavior data of target customers, including financial behavior data, social behavior data, consumption behavior data, geographic location behavior data, and network behavior data; According to the multi-dimensional behavior data, the multi-dimensional behavior characteristics of the target customer are obtained, wherein the multi-dimensional behavior characteristics include the behavior stability index, abnormal behavior frequency, behavior trend data and multi-dimensional behavior correlation data of the target customer; the behavior stability index is used to evaluate the consistency and volatility of the target customer in each behavior dimension; the abnormal behavior frequency refers to the frequency of abnormal events that deviate from the normal state in the behavior pattern of the target customer; the behavior trend data is used to reflect the change trend of the target customer's behavior over time and its development speed; the multi-dimensional behavior correlation data is used to reflect the mutual correlation between the behavior data of different dimensions; Generate credit risk assessment results based on the pre-trained neural network model and the target customer's credit report data, environmental risk assessment data, multi-dimensional behavior data and multi-dimensional behavior characteristics; wherein the credit risk assessment results include credit risk scores, identification results of key risk factors and future credit risk forecasts; the environmental risk assessment data include economic indicator data, and the economic indicator data include GDP growth rate, unemployment rate and industry risk data of the industry in which the target customer is engaged; Generate credit approval recommendations and personalized loan plans based on the credit risk assessment results output by the neural network model; The neural network model includes a feature extraction network, an anomaly detection network and a multi-layer fusion network; The input of the feature extraction network is the multi-dimensional behavior data and multi-dimensional behavior features of the target customer; a convolutional autoencoder is used to extract features from the multi-dimensional behavior data of the target customer to obtain a nonlinear behavior feature vector; The input of the anomaly detection network is the nonlinear behavior feature vector provided by the feature extraction network; the anomaly detection network is implemented by a recursive neural network based on a multi-head self-attention mechanism to obtain abnormal behavior detection results and prediction probabilities of abnormal behaviors; The input of the multi-layer fusion network is the nonlinear behavior feature vector provided by the feature extraction network, the abnormal behavior detection result and the predicted probability of abnormal behavior provided by the anomaly detection network, the credit report data and the economic indicator data; the multi-layer fusion network is implemented using a capsule network to obtain the credit risk assessment results of the target customers.
2. The credit assessment method according to claim 1, characterized in that: The multi-layer convolutional autoencoder used in the feature extraction network includes: The first input layer is used to receive multi-dimensional behavior data and multi-dimensional behavior characteristics of target customers; A first convolutional layer, used to receive the multi-dimensional behavior data provided by the first input layer, wherein the first convolutional layer uses 64 3×3 convolution kernels to perform a convolution operation to extract local features, and outputs 64 feature maps; The first pooling layer is used to receive the output of the first convolutional layer, use a 2×2 pooling window to perform a maximum pooling operation, reduce the spatial dimension of the feature map, and output 64 pooled feature maps; The second convolution layer is used to receive the output of the first pooling layer. The second convolution layer uses 128 3×3 convolution kernels to extract deeper local features, capture temporal dependencies and nonlinear relationships, and outputs 128 feature maps. The second pooling layer is used to receive the output of the second convolutional layer. The second pooling layer uses a 2×2 pooling window to perform a maximum pooling operation to further reduce the spatial dimension of the features, and the output is 128 pooled feature maps; The encoding layer receives the output of the second pooling layer and compresses the feature vector into a 256-dimensional high-dimensional nonlinear behavior feature vector through a fully connected layer as the intermediate result of the feature extraction network and used in the decoding stage. A first deconvolution layer is used to receive the 256-dimensional feature vector output by the encoding layer. The first deconvolution layer uses 128 3×3 convolution kernels to perform a deconvolution operation, restore part of the spatial dimensions, and output 128 feature maps; The first upsampling layer is used to receive the output of the first deconvolution layer, perform a 2×2 upsampling operation, restore the spatial resolution of the data, and generate 128 enlarged feature maps; A second deconvolution layer is used to receive the output of the first upsampling layer. The second deconvolution layer uses 64 3×3 convolution kernels to further perform a deconvolution operation to restore a higher spatial resolution and output 64 feature maps. A second upsampling layer, used to receive the output of the second deconvolution layer, the second upsampling layer performs a 2×2 upsampling operation to restore the spatial resolution to a similar spatial resolution as the first convolution layer, and outputs 64 feature maps; The first output layer is used to further process the output of the second upsampling layer through a convolutional layer to generate reconstructed data similar to the first input layer; and to generate a 256-dimensional high-dimensional nonlinear behavioral feature vector to represent the abstract behavioral characteristics of the target customers.
3. The credit assessment method according to claim 2, characterized in that: The anomaly detection network adopts a recursive neural network based on a multi-head self-attention mechanism, including: The second input layer is used to receive the nonlinear behavior feature vector output from the feature extraction network; The spatiotemporal graph convolution layer is used to receive the nonlinear behavior feature vector provided by the second input layer. The spatiotemporal graph convolution layer constructs a spatiotemporal graph based on the graph convolution network and the time dimension, and is specifically used for: By arranging the target customers' multi-dimensional behavior data and multi-dimensional behavior characteristics in chronological order, a spatiotemporal graph is formed, in which nodes represent the behavior states at different time points, and edges represent the behavior relationships between different time points; The graph convolutional network is used to convolve the nodes in the spatiotemporal graph to capture the dependency of user behavior in the time dimension and the spatial characteristics of behavior data, and obtain a 128-dimensional feature vector after spatiotemporal convolution. The multi-head self-attention layer is used to receive the 128-dimensional feature vector output from the spatiotemporal graph convolution layer, and use the multi-head self-attention mechanism to perform weighted processing on the key behavior patterns to generate a 1024-dimensional feature vector; The recursive neural network layer receives the 1024-dimensional feature vector from the multi-head self-attention layer, combines it with the long short-term memory network to process the time series characteristics of the behavior data, and outputs a 512-dimensional time series feature vector to represent the time characteristics of the abnormal behavior; The second output layer is used to receive the 512-dimensional time series feature vector provided by the recurrent neural network layer to generate the abnormal behavior prediction probability and abnormal behavior detection results of the target customers.
4. The credit assessment method according to claim 3, characterized in that: The multi-layer fusion network is implemented by a dynamic routing capsule network, and the dynamic routing capsule network includes: The third input layer is used to receive the nonlinear behavior feature vector provided by the feature extraction network, the abnormal behavior detection results and abnormal behavior prediction probability provided by the anomaly detection network, the credit report data, and the economic indicator data; the input data of the third input layer is integrated into a 919-dimensional feature vector as the input of the subsequent capsule layer; The first capsule layer is used to receive the 919-dimensional feature vector of the third input layer, and process different types of input features through multiple capsule units, wherein each capsule unit contains 8 neurons, which are responsible for processing data of different dimensions; the output of each capsule unit is an 8-dimensional vector, which represents the feature information of the input feature in different dimensions; and after passing through the first capsule layer, the output is a 64-dimensional feature vector; the dynamic routing mechanism of the first capsule layer includes: dynamically calculating the information transmission weight between capsule units according to the output of the capsule unit; the vector length of the capsule unit output represents the feature strength, and dynamically adjusting the information transmission path between different capsule units to strengthen the capsule output related to the key features; The second capsule layer is used to receive the 64-dimensional vector from the first capsule layer, and further fuse and process feature information from different sources; the second capsule layer includes multiple first capsule units, wherein each first capsule unit includes 16 neurons, and is used to process multi-dimensional features; the second capsule layer divides the 64-dimensional feature vector and inputs it into each capsule unit through a dynamic routing mechanism; the output of the second capsule layer is a 128-dimensional feature vector, which reflects the deep correlation between the target customer behavior data, economic data and anomaly detection results; A feature fusion layer is used to receive the 128-dimensional feature vector output by the second capsule layer and perform feature fusion through a fully connected layer; the output of the feature fusion layer is a 64-dimensional fused feature vector; The third output layer is used to receive the 64-dimensional fused feature vector from the feature fusion layer and generate a comprehensive credit risk assessment result of the target customer. The third output layer is specifically used to: Through the third fully connected layer, the 64-dimensional feature vector is mapped into a credit risk score, which represents the credit risk level of the target customer; The fourth fully connected layer is used to identify and output key factors that affect the customer's credit risk, and the output is a set of classification labels or specific values; wherein the key factors include abnormal behavior, credit history, and changes in the economic environment; Through the time series analysis layer, the 64-dimensional feature vector is combined with historical behavior data for prediction to generate the credit risk change trend for a specified period in the future, and the output is a future risk score that represents the customer's potential credit risk in the future.
5. The credit assessment method according to claim 1, characterized in that: The calculation of the behavior stability index is achieved by a time-weighted average method, specifically including: Time slice the multi-dimensional behavior data and divide the behavior data into different time windows; Calculate the fluctuation range of the behavioral characteristics of each dimension within each time window; The behavioral fluctuations of each time window are weighted averaged, where more recent time windows are given greater weights, to assess the stability of the target customer’s current behavior and generate a behavioral stability index.
6. The credit assessment method according to claim 1, characterized in that: The detection of abnormal behavior frequency is achieved through threshold setting and adaptive learning mechanism based on historical behavior patterns, including: Establish a behavioral baseline pattern by analyzing the target customers’ behavioral data over the past 6 months to 1 year; Setting a threshold for deviation from historical behavior patterns, and behavioral events exceeding this threshold are marked as abnormal; By applying an adaptive learning algorithm, the threshold is updated in real time based on the new behavior data of target customers, making the model more accurate in detecting abnormal behavior.
7. The credit assessment method according to claim 1, characterized in that: The calculation of the multi-dimensional behavior correlation data adopts a method based on the Pearson correlation coefficient, specifically including: Normalize the behavioral characteristic data of target customers in different dimensions; Calculate the correlation between behavioral data of various dimensions, including financial behavior and consumption behavior, social behavior and geographic location behavior; The Pearson correlation coefficient is used to evaluate the linear correlation between behavioral data of different dimensions, generate a correlation matrix, and use it as input for subsequent credit risk assessment models.
8. The credit assessment method according to claim 1, characterized in that: The acquisition and update of the environmental risk assessment data is achieved through a dynamic macroeconomic monitoring system based on big data, specifically including: Obtain the latest GDP growth rate, unemployment rate and industry risk data from government statistics departments, industry reports and third-party economic data providers; Regularly update the economic indicators of the industries in which target customers are engaged, and dynamically adjust the credit risks of target customers in light of the macroeconomic environment of economic recession and rising industry risks.
9. A credit assessment device based on artificial intelligence and big data technology, characterized in that: include: A collection unit, used to collect multi-dimensional behavior data of target customers, wherein the multi-dimensional behavior data includes financial behavior data, social behavior data, consumption behavior data, geographic location behavior data, and network behavior data; An obtaining unit is used to obtain the multi-dimensional behavior characteristics of the target customer according to the multi-dimensional behavior data, wherein the multi-dimensional behavior characteristics include the behavior stability index, abnormal behavior frequency, behavior trend data and multi-dimensional behavior correlation data of the target customer; the behavior stability index is used to evaluate the consistency and volatility of the target customer in each behavior dimension; the abnormal behavior frequency refers to the frequency of abnormal events that deviate from the normal state in the behavior pattern of the target customer; the behavior trend data is used to reflect the change trend of the target customer's behavior over time and its development speed; the multi-dimensional behavior correlation data is used to reflect the mutual correlation between the behavior data of different dimensions; A generating unit, for generating a credit risk assessment result based on a pre-trained neural network model and according to the target customer's credit report data, environmental risk assessment data, multi-dimensional behavior data and multi-dimensional behavior characteristics; wherein the credit risk assessment result includes a credit risk score, a key risk factor identification result and a future credit risk forecast; the environmental risk assessment data includes economic indicator data, and the economic indicator data includes GDP growth rate, unemployment rate and industry risk data of the industry in which the target customer is engaged; An approval unit, used to generate credit approval suggestions and personalized loan plans based on the credit risk assessment results output by the neural network model; The neural network model includes a feature extraction network, an anomaly detection network and a multi-layer fusion network; The input of the feature extraction network is the multi-dimensional behavior data and multi-dimensional behavior features of the target customer; a convolutional autoencoder is used to extract features from the multi-dimensional behavior data of the target customer to obtain a nonlinear behavior feature vector; The input of the anomaly detection network is the nonlinear behavior feature vector provided by the feature extraction network; the anomaly detection network is implemented by a recursive neural network based on a multi-head self-attention mechanism to obtain abnormal behavior detection results and prediction probabilities of abnormal behaviors; The input of the multi-layer fusion network is the nonlinear behavior feature vector provided by the feature extraction network, the abnormal behavior detection result and the predicted probability of abnormal behavior provided by the anomaly detection network, the credit report data and the economic indicator data; the multi-layer fusion network is implemented using a capsule network to obtain the credit risk assessment results of the target customers.
Citation Information
Patent Citations
Personal credit report query monitoring and early warning method based on artificial intelligence
CN117934159A
Credit granting data processing method and system of comprehensive credit system
CN118037440A