Credit risk prediction method based on multi-dimensional tense data
By collecting multi-dimensional temporal data and building a comprehensive risk prediction model, the problems of insufficient single dimensions of the existing credit risk prediction methods and insufficient generalization capabilities of the model are solved, and a more accurate credit risk assessment is achieved.
Patent Information
- Application Number
- CN202510176481.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing credit risk prediction methods lack single dimensions and insufficient generalization capabilities of the model, and cannot fully utilize multi-dimensional information and the ability to process temporal data.
The credit risk prediction method based on multi-dimensional temporal data is adopted to collect financial credit data, behavioral credit data and social network data, perform data preprocessing and feature extraction, build a comprehensive risk prediction model, and use dynamic weight adjustment mechanism and log-smooth processing to capture the time dependence of the data.
It provides a more comprehensive borrower portrait, improves the accuracy of credit risk prediction, and solves the problems of insufficient single dimensions and insufficient generalization capabilities of model.
Smart Images

Figure CN120146988A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of credit data processing, and particularly to a credit risk prediction method based on multi-dimensional temporal data. Background Art
[0002] Credit risk prediction is one of the core issues in the financial field, and its goal is to predict the likelihood of future default by evaluating the credit status of borrowers. Traditional credit risk prediction methods mainly rely on static financial data and historical credit records, and these methods have played an important role in the past few decades. However, with the increasing complexity of the financial environment and the rapid development of data technology, the limitations of traditional methods have gradually emerged, mainly reflected in the following aspects:
[0003] 1. Insufficiency in a single dimension:
[0004] Traditional methods usually only focus on financial data and credit records, ignoring other factors that may affect credit risk, such as behavioral data, social network data, etc.
[0005] Behavioral data (such as consumption habits, payment behaviors) and social network data (such as social activity, social influence) can provide a more comprehensive portrait of borrowers, but traditional methods fail to fully utilize this multi-dimensional information.
[0006] 2. Insufficiency in model generalization ability:
[0007] Traditional methods usually adopt linear models (such as logistic regression) or simple machine learning models (such as decision trees), and these models perform poorly in dealing with non-linear relationships and high-dimensional data.
[0008] With the increase in data dimension and the improvement of data complexity, the generalization ability of traditional models is insufficient, and problems such as overfitting or underfitting are likely to occur. Summary of the Invention
[0009] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title, and such simplifications or omissions shall not be used to limit the scope of the present invention.
[0010] In view of the problems existing in the above-mentioned existing credit risk prediction methods, the present invention is proposed.
[0011] Therefore, the technical problem solved by the present invention is to solve the problems of insufficient single dimension and insufficient model generalization ability of existing credit risk prediction methods.
[0012] To solve the above technical problems, the present invention provides the following technical solutions: A credit risk prediction method based on multi-dimensional temporal data, comprising the following steps: S1: Collect financial credit data at the current time point; S2: Statistically analyze behavioral credit data under a rated time series; S3: Statistically analyze social network data under a rated time series; S4: Construct a comprehensive risk prediction model, sequentially input the financial credit data, the behavioral credit data, and the social network data, and output a comprehensive risk prediction value; S5: Based on the comprehensive risk prediction value, conduct an object credit risk assessment at the current time point.
[0013] As a preferred embodiment of the credit risk prediction method based on multi-dimensional temporal data according to the present invention, wherein: The financial credit data specifically includes: asset data α and liability data β statistically obtained at the current time point; The behavioral credit data specifically includes: payment amount A, payment frequency B, and total income C under a rated time series; The social network data specifically includes: posting frequency a, interaction frequency b, online duration c, social circle size d, content diversity e, number of fans f, content dissemination range g, secondary interaction frequency h, social network centrality i, and KOL score j under a rated time series.
[0014] As a preferred embodiment of the credit risk prediction method based on multi-dimensional temporal data according to the present invention, wherein: After collecting the financial credit data, the behavioral credit data, and the social network data, it further includes sequentially performing data preprocessing on each data; wherein, the data preprocessing steps specifically include cleaning, normalization, and feature extraction.
[0015] As a preferred embodiment of the credit risk prediction method based on multi-dimensional temporal data according to the present invention, wherein: When performing feature extraction on the social network data, it specifically includes: Performing feature extraction on social activity information to obtain social activity reference parameters; Performing feature extraction on social influence information to obtain social influence reference parameters.
[0016] As a preferred embodiment of the credit risk prediction method based on multi-dimensional temporal data according to the present invention, wherein: The feature extraction of social activity information includes the following steps:
[0017] 1. Parameter definition:
[0018] Posting frequency:
[0019] Interaction frequency:
[0020] Online duration:
[0021] Social circle size: d = number of friends + number of fans
[0022] Content diversity:
[0023] 2. Introduce a dynamic weight adjustment mechanism, where the weight w of each parameter i changes over time, and the specific formula is as follows:
[0024]
[0025] where: t is the time variable, representing the difference between the current time and the reference time; α i is the weight adjustment coefficient of the i-th parameter, obtained through training with historical data;
[0026] 3. Based on the above parameters and dynamic weights, the calculation formula for social activity A(t) is:
[0027] A(t) = w a (t)·log(a + 1) + w b (t)·log(b + 1) + w c (t)·log(c + 1) + w d (t)·log(d + 1) + w e (t)·log(e + 1)
[0028] where: log(.) is used to smooth the data distribution and avoid the influence of extreme values; +1 is used to handle the zero value situation to ensure the legality of the logarithmic operation.
[0029] As a preferred solution of the credit risk prediction method based on multi-dimensional temporal data described in the present invention, where: the feature extraction of social influence information includes the following steps:
[0030] 1. Parameter definition:
[0031] Number of fans: f = total number of fans;
[0032] Content dissemination range:
[0033] Second interaction frequency:
[0034] Social network centrality: i = degree centrality (direct analysis of social graph)
[0035] KOL score: j = Klout direct score
[0036] 2. Introduce a dynamic weight adjustment mechanism, where the weight w of each parameter i changes over time, and the specific formula is as follows:
[0037]
[0038] where: t is a time variable, representing the difference between the current time and the reference time, and β i is the weight adjustment coefficient of the i-th parameter, obtained through training with historical data;
[0039] 3. Based on the above parameters and dynamic weights, the calculation formula for the social influence I(t) is:
[0040] I(t) = w f (t)·log(f + 1) + w g (t)·log(g + 1) + w h (t)·log(h + 1) + w i (t)·log(i + 1) + w j (t)·log(j + 1)
[0041] where: log(.) is used to smooth the data distribution and avoid the influence of extreme values; +1 is used to handle the zero value situation to ensure the legality of the logarithmic operation.
[0042] As a preferred solution of the credit risk prediction method based on multi-dimensional temporal data according to the present invention, where: the constructed comprehensive risk prediction model is specifically:
[0043]
[0044] where, δ is the comprehensive risk prediction value; t is a time variable, in days; α is the asset data statistically obtained at the current time point, in ten thousand; β is the liability data statistically obtained at the current time point, in ten thousand; A is the payment amount under the rated time series, in ten thousand; B is the number of payment times under the rated time series, in times; C is the total income under the rated time series, in ten thousand; A(t) is the social activity indication parameter; I(t) is the social influence indication parameter; 1.35, 1.06 and -1 / 2 are all adjustment constants; dx is the integral operation.
[0045] As a preferred solution of the credit risk prediction method based on multi-dimensional temporal data according to the present invention, where: when the comprehensive risk prediction value is higher than the preset threshold, it is defined that the credit risk at the current time point of the evaluation object is abnormal.
[0046] As a preferred solution of the credit risk prediction method based on multi-dimensional temporal data according to the present invention, where: the preset threshold is defined as 3.995 or 3.9953.
[0047] Advantages of the present invention: The present invention provides a credit risk prediction method based on multi-dimensional temporal data. It collects and statistically analyzes financial credit data, behavioral credit data, and social network data. After data preprocessing, it extracts features from various types of data, constructs a comprehensive risk prediction model after obtaining various feature parameters, and evaluates the credit risk of the object at the current time point based on the output comprehensive risk prediction value. The present invention not only considers traditional financial data but also introduces behavioral data and social network data, which can provide a more comprehensive borrower profile. At the same time, by constructing a model that can handle temporal data, the present invention captures the time dependence in the data, thereby improving the prediction accuracy and solving the problems of single-dimensional insufficiency and insufficient model generalization ability in existing credit risk prediction methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0049] Figure 1 It is the overall method flowchart of the credit risk prediction method based on multi-dimensional temporal data provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0051] With the increasing complexity of the financial environment and the rapid development of data technology, the limitations of traditional methods are gradually emerging, mainly reflected in the following aspects:
[0052] 1. Single-dimensional insufficiency:
[0053] Traditional methods usually only focus on financial data and credit records, ignoring other factors that may affect credit risk, such as behavioral data, social network data, etc.
[0054] Behavioral data (such as consumption habits, payment behavior) and social network data (such as social activity, social influence) can provide a more comprehensive borrower profile, but traditional methods fail to fully utilize this multi-dimensional information.
[0055] 2. Insufficient model generalization ability:
[0056] Traditional methods usually adopt linear models (such as logistic regression) or simple machine learning models (such as decision trees), which perform poorly in dealing with non - linear relationships and high - dimensional data.
[0057] As the data dimension increases and the data complexity improves, the generalization ability of traditional models is insufficient, and problems such as overfitting or underfitting are likely to occur.
[0058] Therefore, please refer to Figure 1 , the present invention provides a credit risk prediction method based on multi - dimensional temporal data, including the following steps:
[0059] S1: Collect financial credit data at the current time point;
[0060] S2: Statistically analyze behavioral credit data under a rated time series;
[0061] S3: Statistically analyze social network data under a rated time series;
[0062] S4: Construct a comprehensive risk prediction model, input financial credit data, behavioral credit data, and social network data in sequence, and output a comprehensive risk prediction value;
[0063] S5: Evaluate the credit risk of the object at the current time point based on the comprehensive risk prediction value.
[0064] Specifically, the financial credit data is specifically: asset data α and liability data β statistically obtained at the current time point;
[0065] Specifically, the behavioral credit data is specifically: payment amount A, payment times B, and total income C under a rated time series;
[0066] Specifically, the social network data is specifically: posting frequency a, interaction frequency b, online duration c, size of social circle d, content diversity e, number of fans f, content dissemination range g, second interaction frequency h, social network centrality i, and KOL score j under a rated time series.
[0067] It should be noted that: the present invention proposes that in credit risk prediction, social activity and social influence can be reflected by analyzing and quantifying social network data.
[0068] The following are the specific methods for expressing these two indicators with data:
[0069] 1. Data expression of social activity
[0070] Social activity refers to the degree of participation and interaction frequency of borrowers in the social network. It can be quantified by the following data:
[0071] ① Data source:
[0072] Social media platforms (such as Weibo, WeChat, Twitter, Facebook, etc.).
[0073] Social network analysis tools (such as social graph analysis).
[0074] ②Quantification method:
[0075] Posting frequency: Count the number of posts made by the borrower within a certain period of time (such as the number of posts per day, the number of posts per week).
[0076] Formula: Posting frequency = Total number of posts / Time period
[0077] Number of comments and likes: Count the number of comments and likes made by the borrower on others' content.
[0078] Formula: Interaction frequency = (Total number of comments + Total number of likes) / Time period
[0079] Online duration: Count the online duration of the borrower on the social platform.
[0080] Formula: Online duration = Total online time / Time period
[0081] Size of social circle: Count the number of friends or followers of the borrower.
[0082] Formula: Size of social circle = Number of friends + Number of followers
[0083] Content diversity: Count the diversity of types of content posted by the borrower (such as text, pictures, videos, etc.).
[0084] Formula: Content diversity = Number of different types of content / Total number of posts
[0085] Example:
[0086] If a borrower has posted 50 Weibo posts, commented 100 times, liked 200 times, had an online duration of 60 hours, and had 500 friends in the past month, then their social activity can be calculated as follows:
[0087] Posting frequency: 50 posts / 30 days ≈ 1.67 posts / day
[0088] Interaction frequency: (100 + 200) / 30 ≈ 10 times / day
[0089] Online duration: 60 hours / 30 days = 2 hours / day
[0090] Size of social circle: 500
[0091] Content diversity: Assuming there are 3 types of content posted (text, pictures, videos), then the diversity is 3 / 50 = 0.06
[0092] 2. Data Representation of Social Influence
[0093] Social influence refers to the ability of a borrower to influence the behavior or decisions of others in a social network. It can be quantified through the following data:
[0094] ① Data Sources:
[0095] Social media platforms (such as Weibo, WeChat, Twitter, Facebook, etc.).
[0096] Social network analysis tools (such as social graph analysis, influence scoring tools).
[0097] ② Quantification Methods:
[0098] Number of Fans: Count the number of fans or followers of the borrower.
[0099] Formula: Number of Fans = Total Number of Fans
[0100] Scope of Content Dissemination: Count the number of reposts, shares, or views of the content posted by the borrower.
[0101] Formula: Scope of Dissemination = (Total Number of Reposts + Total Number of Shares + Total Number of Views) / Time Period
[0102] Interaction Rate: Count the interaction rate of the content posted by the borrower (such as the ratio of the number of comments, likes to the number of fans).
[0103] Formula: Interaction Frequency = (Total Number of Comments + Total Number of Likes) / Number of Fans
[0104] Social Network Centrality: Calculate the centrality of the borrower in the social network through the social graph (such as degree centrality, closeness centrality, betweenness centrality, etc.).
[0105] Formula: Degree Centrality = Number of Direct Connections / (Total Number of Nodes - 1)
[0106] KOL (Key Opinion Leader) Score: Obtain the KOL score of the borrower through third-party tools (such as Weibo Influence Index, Klout Score).
[0107] Formula: KOL Score = Score Provided by Third-Party Tool
[0108] Example:
[0109] If a borrower has 100,000 fans on Weibo, and the content posted in the past month has been reposted 1,000 times, shared 500 times, with a view count of 500,000 times and an interaction rate of 5%, then their social influence can be calculated as follows:
[0110] Number of Fans: 100,000
[0111] Propagation range: (1000 + 500 + 500,000) / 30 ≈ 16,833 times per day
[0112] Interaction rate: 5%
[0113] Social network centrality: Assume the degree centrality of the borrower is 0.8 (obtained through social graph analysis)
[0114] KOL score: Assume the Klout score is 70
[0115] Additionally, in order to incorporate social activity and social influence into the credit risk prediction model, the above data can also be integrated and standardized (this step can also be omitted for more completeness)
[0116] Data integration: Integrate various indicators (such as posting frequency, interaction frequency, number of fans, propagation range, etc.) into a multi-dimensional feature vector - Standardization processing: Normalize each indicator so that they are comparable on the same scale.
[0117] Formula: Standardization = (original value - minimum value) / (maximum value - minimum value)
[0118] Example:
[0119] Assume the social activity and social influence data of a certain borrower are as follows: Posting frequency: 1.67 posts per day (standardized value: 0.5)
[0120] Interaction frequency: 10 times per day (standardized value: 0.6)
[0121] Number of fans: 100,000 (standardized value: 0.8)
[0122] Propagation range: 16,833 times per day (standardized value: 0.7)
[0123] Interaction rate: 5% (standardized value: 0.4)
[0124] Then it can be integrated into a feature vector: [0.5, 0.6, 0.8, 0.7, 0.4] for model input.
[0125] Furthermore, after collecting financial credit data, behavioral credit data, and social network data, it also includes performing data preprocessing on each data in sequence;
[0126] Among them, the data preprocessing steps specifically include cleaning, normalization, and feature extraction.
[0127] It should be noted that:
[0128] 1. Data cleaning
[0129] The purpose of data cleaning is to remove noise, outliers, and missing values in the data to ensure the accuracy and integrity of the data.
[0130] 1.1 Handling Missing Values
[0131] Deletion Method:
[0132] If the proportion of missing values in a certain record is too high (such as exceeding 50%), then directly delete that record.
[0133] Formula: Delete records that satisfy: the number of missing values / the total number of fields > 0.5.
[0134] Filling Method:
[0135] For a small number of missing values, use the following methods to fill them:
[0136] Mean Filling: Fill the missing values with the mean of this field.
[0137] Formula:
[0138] Median Filling: Fill the missing values with the median of this field.
[0139] Mode Filling: Fill the missing values with the mode of this field (applicable to categorical data).
[0140] Interpolation Method: For time series data, use linear interpolation or spline interpolation to fill the missing values.
[0141] Formula:
[0142] 1.2 Handling Outliers
[0143] 3σ Principle: For data that follows a normal distribution, consider the values that exceed the mean ± 3 times the standard deviation as outliers.
[0144] Formula:
[0145] Box Plot Method: Consider the values that exceed 1.5 times the interquartile range as outliers.
[0146] Formula:
[0147]
[0148] IQR = Q3 - Q1 (Interquartile Range).
[0149] 1.3 Handling Duplicate Values
[0150] Delete Duplicate Records: If all field values of two records are exactly the same, then delete one of them.
[0151] Formula: Delete duplicate records where record i = record j.
[0152] 2. Data normalization
[0153] The purpose of data normalization is to convert data with different dimensions to the same scale, avoiding certain features from dominating model training due to overly large numerical values.
[0154] 2.1 Min-Max normalization
[0155] Linearly map the data to the interval [0, 1].
[0156] 2.2 Z-Score standardization
[0157] Convert the data to a distribution with a mean of 0 and a standard deviation of 1.
[0158] 2.3 Decimal scaling normalization
[0159] Map the data to the interval [-1, 1] by dividing by the maximum value of the data.
[0160] 2.4 Log normalization
[0161] Compress the data range through logarithmic transformation.
[0162] Furthermore, when extracting features from social network data, it specifically includes:
[0163] Extract features from social activity information to obtain social activity reference parameters;
[0164] Extract features from social influence information to obtain social influence reference parameters.
[0165] Specifically, the steps for extracting features from social activity information are as follows:
[0166] 1. Parameter definition:
[0167] Posting frequency:
[0168] Interaction frequency:
[0169] Online duration:
[0170] Size of social circle: d = number of friends + number of fans
[0171] Content diversity:
[0172] 2. To reflect the contribution degrees of different parameters at different time periods, introduce a dynamic weight adjustment mechanism, and the weight w of each parameter i varies with time, and the specific formula is as follows:
[0173]
[0174] where: t is a time variable representing the difference between the current time and the reference time; α i is the weight adjustment coefficient of the i-th parameter, obtained by training with historical data;
[0175] 3. Based on the above parameters and dynamic weights, the calculation formula for social activity A(t) is:
[0176] A(t) = w a (t)·log(a + 1) + w b (t)·log(b + 1) + w c (t)·log(c + 1) + w d (t)·log(d + 1) + w e (t)·log(e + 1)
[0177] where: log(.) is used to smooth the data distribution and avoid the influence of extreme values; +1 is used to handle the zero value situation to ensure the legality of the logarithmic operation.
[0178] Dynamic weight adjustment: Existing technologies usually adopt fixed weights, while the present invention introduces a dynamic weight adjustment mechanism, which can automatically adjust the contribution degree of each parameter according to time changes, being more in line with the actual situation.
[0179] Logarithmic smoothing processing: Through logarithmic function smoothing processing, the influence of extreme values on the model is avoided, and the robustness of the model is improved.
[0180] Time-dependent modeling: Through the time variable t and the weight adjustment coefficient α i , the model can capture the time dependence of social activity, which is different from the static modeling method of existing technologies.
[0181] Example calculation
[0182] Suppose the data of a certain borrower in the past 30 days are as follows:
[0183] Posting frequency: a = 50 / 30 ≈ 1.67 = 50 / 30 ≈ 1.67
[0184] Interaction frequency: b = (100 + 200) / 30 ≈ 10 = (100 + 200) / 30 ≈ 10
[0185] Online duration: c = 60 / 30 = 2 = 60 / 30 = 2 hours per day
[0186] Size of social circle: d = 500 = 500
[0187] Content diversity: e = 3 / 50 = 0.06 = 3 / 50 = 0.06
[0188] Assume the current time t = 30 days, and the weight adjustment coefficient is:
[0189] α a == 0.1
[0190] α b = 0.2
[0191] α c = 0.15
[0192] α d = 0.25
[0193] α e = 0.3
[0194] Calculate the dynamic weight:
[0195]
[0196] Calculate the activity:
[0197] A(30) = w a (30)*log(1.67 + 1)+w b (30)*log(10 + 1)+w c (30)*log(2 + 1)+w d (30)*log(500 + 1)
[0198] +w e (30)*log(0.06 + 1);
[0199] Furthermore, the feature extraction of social influence information includes the following steps:
[0200] 1. Parameter definition:
[0201] Number of fans: f = total number of fans;
[0202] Content dissemination range:
[0203] Second interaction frequency:
[0204] Social network centrality: i = degree centrality (direct analysis of social graph)
[0205] KOL score: j = Klout direct score
[0206] 2. To reflect the contribution degree of different parameters in different time periods, a dynamic weight adjustment mechanism is introduced, and the weight w of each parameter iIt changes over time, and the specific formula is as follows:
[0207]
[0208] Where: t is the time variable, representing the difference between the current time and the reference time, and β i is the weight adjustment coefficient of the i-th parameter, obtained through training with historical data;
[0209] 3. Based on the above parameters and dynamic weights, the calculation formula for the social influence I(t) is:
[0210] I(t) = w f (t)·log(f + 1) + w g (t)·log(g + 1) + w h (t)·log(h + 1) + w i (t)·log(i + 1) + w j (t)·log(j + 1)
[0211] Where: log(.) is used to smooth the data distribution and avoid the influence of extreme values; +1 is used to handle the zero value situation to ensure the legality of the logarithmic operation.
[0212] Dynamic weight adjustment: Existing technologies usually adopt fixed weights, while the present invention introduces a dynamic weight adjustment mechanism, which can automatically adjust the contribution degree of each parameter according to time changes, being more in line with the actual situation.
[0213] Logarithmic smoothing processing: Through logarithmic function smoothing processing, the influence of extreme values on the model is avoided, and the robustness of the model is improved.
[0214] Time-dependent modeling: Through the time variable t and the weight adjustment coefficient β i , the model can capture the time dependence of social influence, different from the static modeling methods of existing technologies.
[0215] Example calculation
[0216] Suppose the data of a certain borrower in the past 30 days is as follows:
[0217] Number of fans: f = 100,000 f = 100,000
[0218] Content dissemination range: g = (1000 + 500 + 500,000) / 30 ≈ 16,833
[0219] Second interaction frequency: h = (100 + 200) / 100,000 = 0.003
[0220] Social network centrality: i = 0.8
[0221] KOL Score: j = 70
[0222] Assume the current time t = 30 days, and the weight adjustment coefficient is:
[0223] β f = 0.1
[0224] β g = 0.2
[0225] β h = 0.15
[0226] β i = 0.25
[0227] β j = 0.3
[0228] Calculate the dynamic weight:
[0229]
[0230] Calculate the influence:
[0231] I(30) = w f (30) * log(100,000 + 1) + w g (30) * log(16,833 + 1) + w h (30) * log(0.003 + 1) + w i (30) *
[0232] log(0.8 + 1) + w j (30) * log(70 + 1);
[0233] Furthermore, the constructed comprehensive risk prediction model is specifically:
[0234]
[0235] Among them, δ is the comprehensive risk prediction value; t is the time variable, in days; α is the asset data statistically obtained at the current time point, in ten thousand; β is the liability data statistically obtained at the current time point, in ten thousand; A is the payment amount under the rated time series, in ten thousand; B is the number of payment times under the rated time series, in times; C is the total income under the rated time series, in ten thousand; A(t) is the social activity reference parameter; I(t) is the social influence reference parameter; 1.35, 1.06 and -1 / 2 are all adjustment constants; dx is the integral operation.
[0236] Specifically, when the comprehensive risk prediction value is higher than the preset threshold, it is defined that the credit risk of the evaluation object at the current time point is abnormal.
[0237] Specifically, the preset threshold is defined as 3.995 or 3.9953.
[0238] To verify the technical effects of the present invention, the following simulation experiments are carried out:
[0239] Step 1: Data collection
[0240] Collect the financial credit data, behavioral credit data, and social network data of a group of borrowers. The following is an example table of the collected data:
[0241]
[0242] Step 2: Data preprocessing
[0243] Step 3: Model training and verification
[0244] Model type Accuracy Recall F1 score Traditional model 85% 80% 82.5% The model of the present invention 95% 92% 9%
[0245] The present invention provides a credit risk prediction method based on multi-dimensional temporal data, which collects and statistically analyzes financial credit data, behavioral credit data, and social network data. After data preprocessing, feature extraction is performed on various types of data, and after obtaining various feature parameters, a comprehensive risk prediction model is constructed. Based on the output comprehensive risk prediction value, the credit risk of the object at the current time point is evaluated. The present invention not only considers traditional financial data, but also introduces behavioral data and social network data, which can provide a more comprehensive borrower portrait. At the same time, the present invention constructs a model capable of processing temporal data to capture the time dependence in the data, thereby improving the prediction accuracy and solving the problems of single-dimensional deficiency and insufficient model generalization ability in existing credit risk prediction methods.
[0246] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A credit risk prediction method based on multi-dimensional temporal data, characterized in that: The steps include: S1: Collect financial credit data at the current time point; S2: Behavioral credit data under statistical rating time series; S3: Social network data under statistical rated time series; S4: constructing a comprehensive risk prediction model, inputting the financial credit data, the behavioral credit data and the social network data in sequence, and outputting a comprehensive risk prediction value; S5: Conduct credit risk assessment of the object at the current time point based on the comprehensive risk prediction value.
2. The credit risk prediction method based on multi-dimensional temporal data according to claim 1, characterized in that: The financial credit data specifically includes: asset data α and liability data β obtained by statistics at the current time point; The behavioral credit data specifically includes: payment amount A, payment times B and total income C under the rated time series; The social network data specifically include: posting frequency a, interaction frequency b, online time c, social circle size d, content diversity e, number of fans f, content dissemination range g, second interaction frequency h, social network centrality i and KOL score j under the rated time series.
3. The credit risk prediction method based on multi-dimensional temporal data according to claim 2, characterized in that: After collecting the financial credit data, the behavioral credit data and the social network data, it also includes preprocessing each data in turn; Among them, the data preprocessing steps specifically include cleaning, normalization and feature extraction.
4. The credit risk prediction method based on multi-dimensional temporal data according to claim 3 is characterized in that: When extracting features from the social network data, it specifically includes: Extract features from social activity information to obtain social activity indicators; Feature extraction is performed on social influence information to obtain social influence parameters.
5. The credit risk prediction method based on multi-dimensional temporal data according to claim 4 is characterized in that: Feature extraction of social activity information includes the following steps:
1. Parameter definition: Posting frequency: Interaction frequency: Online time: Size of social circle: d = number of friends + number of fans Content Diversity:
2. Introduce a dynamic weight adjustment mechanism. The weight w of each parameter i Changes over time, the specific formula is as follows: Among them: t is the time variable, which represents the difference between the current time and the reference time; α i is the weight adjustment coefficient of the i-th parameter, obtained through historical data training; 3. Based on the above parameters and dynamic weights, the calculation formula of social activity A(t) is: A(t)=w i (t).log(a+I)+w b (t).log(b+1)+w c (t).log(c+1) +w d (t)·log(d+1)+w e (t)·log(e+1) Among them: log(.) is used to smooth the data distribution and avoid the influence of extreme values; +1 is used to handle zero values and ensure the legitimacy of logarithmic operations.
6. The credit risk prediction method based on multi-dimensional temporal data according to claim 5, characterized in that: Feature extraction of social influence information includes the following steps:
1. Parameter definition: Number of fans: f = total number of fans; Content dissemination scope: Second interaction frequency: Social network centrality: i = degree centrality (direct analysis of social graph) KOL score: j = Klout direct score 2. Introduce a dynamic weight adjustment mechanism. The weight w of each parameter i Changes over time, the specific formula is as follows: Among them: t is the time variable, which means the difference between the current time and the reference time, β i is the weight adjustment coefficient of the i-th parameter, obtained through historical data training; 3. Based on the above parameters and dynamic weights, the calculation formula of social influence I(t) is: I(t)=w f (t)·log(f+1)+w g (t)·log(g+1)+w h (t)·log(h+1) +w i (t)·log(i+1)+w j (t)·log(j+1) Among them: log(.) is used to smooth the data distribution and avoid the influence of extreme values; +1 is used to handle zero values and ensure the legitimacy of logarithmic operations.
7. The credit risk prediction method based on multi-dimensional temporal data according to claim 6, characterized in that: The comprehensive risk prediction model constructed is specifically: Among them, δ is the comprehensive risk prediction value; t is the time variable, days; α is the asset data obtained statistically at the current time point, 10,000; β is the liability data obtained statistically at the current time point, 10,000; A is the payment amount under the rated time series, 10,000; B is the number of payments under the rated time series, times; C is the total income under the rated time series, 10,000; A(t) is the social activity parameter; I(t) is the social influence parameter; 1.35, 1.06 and -1 / 2 are all adjustment constants; dx is an integral operation.
8. The credit risk prediction method based on multi-dimensional temporal data according to claim 7, characterized in that: When the comprehensive risk prediction value is higher than a preset threshold, the credit risk abnormality of the assessment object at the current time point is defined.
9. The credit risk prediction method based on multi-dimensional temporal data according to claim 8, characterized in that: The preset threshold is defined as 3.995 or 3.9953.