Time cycle characteristic derivation method based on lawyer letter sample
The time-cycle feature derivation method generates high-dimensional feature variables, which solves the problem of inexplicability of the sample features of lawyers' letters, and improves the prediction accuracy and generalization ability of the model.
Patent Information
- Application Number
- CN202410198732.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-07-08
AI Technical Summary
现有的律师函件样本特征无可解释性,业务含义单一,导致机器学习模型无法挖掘到更多更有效的信息,导致模型预测准确度不高。
Through the time-cycle feature derivation method based on the lawyer's letter sample, the original credit line usage records are obtained, the basic data set of feature derivation is aggregated and calculated, high-dimensional feature variables are generated, and the time loop strategy is used to adjust the time window, low-contribution features are screened, and high-dimensional feature variables are retained.
High-dimensional feature variables with strong explanatory and high time business significance were generated, which improved the model's performance and generalization ability and improved the model's prediction accuracy.
Smart Images

Figure CN120277631A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing technology, and in particular relates to a time cycle feature derivation method based on lawyer letter samples. Background Art
[0002] In recent years, with the rapid development of big data technology at home and abroad, the trend of credit risk control shifting from traditional manual risk control to digital risk control has become increasingly strong. In digital risk control, the underlying machine learning algorithm has high requirements for the interpretability of variables, and it is best if the variables have practical business meanings (Yang Yemin; Zhang Huijun; Zhang Xiaolong. National Natural Science Foundation of China (No. 61572344), 1-2).
[0003] Existing letter sample variables are not highly interpretable, and their actual business meanings are relatively simple. Directly using machine learning algorithms for modeling often has poor results, does not contribute much to the model, and cannot mine more variable business information, which will cause serious overfitting in the model learning results and low model prediction accuracy (Guo Liang; Wang Chao; Song Liuyi; Mao Renxin; Liu Yang. Chinese Patent: CN114553395B, 2022-07-26, 1-2).
[0004] To address this phenomenon, we need a method for generating feature values that can reveal the complex relationships between the original features based on the original basic sample data, help the model more effectively utilize the information hidden in the data, and thus improve the performance and generalization ability of the model. Summary of the invention
[0005] The main purpose of the present invention is to provide a time cycle feature derivation method based on lawyer letter samples, aiming to solve the problems that the existing basic time dimension feature data has low data dimension, low interpretability, and poor effect of modeling directly using machine learning algorithms.
[0006] To achieve the above object, the present invention provides a method for deriving time cycle features based on lawyer's letter samples, the method for deriving time cycle features based on lawyer's letter samples comprising: Obtaining the original credit line usage records in the lawyer's letter sample, and obtaining a feature-derived basic data set based on the original credit line usage records; Based on the feature derivation basic data set, according to the data derivation algorithm, a feature derivation result is obtained; According to the time cycle strategy, the time window of the data derivation algorithm is adjusted, and the feature derivation result after the time parameter is updated is recalculated; A model contribution analysis is performed on the feature derivative results, feature derivative results with a contribution lower than a preset value are screened out, and the remaining feature derivative results are used as high-dimensional feature variables.
[0007] Optionally, obtaining the original credit limit usage record, and obtaining a feature-derived basic data set according to the original credit limit usage record, including: Based on the credit limit usage record database of the cooperation institution, obtaining the original credit limit usage record within a preset time range; Splitting the original credit limit usage record according to a preset unit time length to obtain sub-data of the credit limit usage record; Calculating the credit limit utilization rate corresponding to each sub-data of the credit limit usage record; Arranging the credit limit utilization rate according to the divided time to obtain a feature-derived basic data set.
[0008] Optionally, based on the feature-derived basic data set, obtaining a feature-derived result according to a data derivation algorithm, including: Performing data aggregation processing on the feature-derived basic data set according to a preset time window to obtain an aggregated feature value required for feature derivation; Generating a high-dimensional derived feature value according to the aggregated feature value according to a preset conversion algorithm; Using the aggregated feature value and the high-dimensional derived feature value as the feature-derived result.
[0009] Optionally, based on a data aggregation strategy, processing the feature-derived basic data set according to a preset unit time length to obtain an aggregated feature value required for feature derivation, including: Determining a sample segment participating in data aggregation in the feature-derived basic data set according to a preset time window; Calculating the aggregated feature value of the sample segment for the credit limit utilization rate data in the sample segment according to the data aggregation strategy; The aggregated feature value includes the mean value, trimmed mean value, maximum value, sum value, range, standard deviation, and coefficient of variation of the sample segment.
[0010] Optionally, generating a high-dimensional derived feature value according to a preset derivation algorithm according to the aggregated feature value, including: Determining the credit limit utilization rate within each unit time length in the sample segment according to the aggregated feature value and the preset unit time length, and calculating the change amount of the credit limit utilization rate between adjacent unit time lengths in the sample segment; Determining the overall change trend of the sample segment according to the change amount of the credit limit utilization rate; Calculating the relative proportion value of the credit limit utilization rate within each unit time length in the sample segment; Take the change amount and change trend of the quota utilization rate of the sample segment and the relative proportion value of the quota utilization rate within the unit time as high-dimensional derivative feature values.
[0011] Optionally, the model contribution degree analysis of the feature derivation result is performed, the feature derivation results lower than the preset contribution degree are screened out, and the remaining feature derivation results are used as high-dimensional feature variables, including: Divide the feature derivation results into multiple groups of derivative feature value data according to the data type; Divide each group of derivative feature value data into a training set and a validation set; Input the training set into the initial logistic regression model for training to obtain a trained logistic regression model; Perform performance evaluation on the trained logistic regression model through the corresponding validation set to obtain model performance index data; According to the model performance index data, determine the model contribution degree of each group of derivative feature value data, screen out the feature derivation results lower than the preset contribution degree, and use the remaining feature derivation results as high-dimensional feature variables.
[0012] Optionally, the performance evaluation of the trained logistic regression model through the corresponding validation set to obtain model performance index data includes: Input the feature data of the validation set into the trained logistic regression model to obtain a model prediction result; According to the model prediction result and the true result of the validation set, calculate the AUC value and KS value of the logistic regression model; Use the AUC value and KS value as the model performance index data of this model.
[0013] The beneficial effects of the present invention are as follows. Compared with the prior art, the present invention solves the problems that the characteristics of lawyer's letter samples are unexplainable, the business meaning is single, and more effective information cannot be mined for the machine learning model. By adopting a time loop method to aggregate the basic characteristics of lawyer's letter samples in dimensions such as coefficient of variation, range, variance, maximum value, mean value, trimmed mean, standard deviation, etc., hundreds of different high-dimensional features are generated. High-dimensional feature variables with strong interpretability and great time business significance are generated. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic diagram of the high-dimensional feature derivation process of the time loop feature derivation method based on lawyer's letter samples of the present invention; Figure 2 It is a schematic diagram of the process of the first embodiment of the time loop feature derivation method based on lawyer's letter samples of the present invention; Figure 3Schematic flowchart of the second embodiment of the time loop feature derivation method based on lawyer's letter samples according to the present invention; Figure 4 Schematic flowchart of the third embodiment of the time loop feature derivation method based on lawyer's letter samples according to the present invention; Figure 5 Schematic flowchart of the fourth embodiment of the time loop feature derivation method based on lawyer's letter samples according to the present invention.
[0015] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0016] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] Referring to Figure 1 , Figure 1 Schematic flowchart of the high-dimensional feature derivation process of the time loop feature derivation method based on lawyer's letter samples according to the present invention.
[0018] It can be understood that in the overall process of the present invention, first, relevant customer data information is exported from the lawyer's letter customer database and used as the original information, and then through a pre-set algorithm, the credit utilization rate under the algorithm parameters is obtained. Then, according to the time window and the time span of the original data, the derivative feature values of the basic dimension are obtained through multiple loop calculations. Based on this part of the data as the aggregation basis, the data is aggregated in multiple threads, and the contribution degree of the successfully aggregated high-dimensional feature value results is calculated. The high-dimensional feature values that do not meet the contribution degree threshold are screened out, and the finally retained high-dimensional feature values can meet the modeling requirements.
[0019] The embodiment of the present invention provides a feature derivation method based on the time loop of lawyer's letter samples. Referring to Figure 2 , Figure 2 Schematic flowchart of the first embodiment of the time loop feature derivation method based on lawyer's letter samples according to the present invention.
[0020] In this embodiment, the time loop feature derivation method based on lawyer's letters includes the following steps: Step S10: Obtain the original credit limit usage records in the lawyer's letter samples, and obtain the feature derivation basic data set according to the original credit limit usage records.
[0021] It should be noted that the original credit line usage record refers to a batch of data obtained from official institutions or cooperative institutions for machine learning training. This part of the data is obtained within the legal scope, including the consumption records generated by a large number of users during the use of bank credit cards or other types of credit products, as well as the credit line records given by their cooperative institutions when users use such credit products. Due to the changes in users' consumption habits and personal economic conditions, the consumption records, the credit lines given by the cooperative institutions at that time, and the changes in the credit lines are time-varying and have strong reference significance.
[0022] It can be understood that the data source containing the original credit line usage information is a database, an Excel file, or other forms of data sets. Ensure that the records include information about each credit account, such as account ID, transaction date, transaction amount, etc. Based on the time variable, some expanded and derived features can be generated. For example, the average daily transaction amount of each account can be calculated by dividing the total transaction amount of the account by the number of transaction days, which can provide information about the transaction activity of the account. In addition, in addition to time-related features, other statistical features can be supplemented. For example, the number of transactions of each account, the change curve of the transaction amount, and the credit line usage situation on a daily and monthly basis can be calculated. These statistical features can provide information about the transaction frequency, transaction scale, and transaction stability of the account.
[0023] It should be understood that in the above examples, only some ways to obtain the basic data set for feature derivation are pointed out. In actual situations, the obtained basic data set for feature derivation is data that can reflect some information on a lower data dimension after a certain preprocessing. As long as this part of the data can undergo derivative operations related to feature values or further dimensionality increase operations in the subsequent process, more meaningful feature functions can be obtained. On the other hand, taking the user credit line utilization rate as an example of the basic data set for feature derivation, based on this part of the basic data, the credit consumption situation of such users within a time range from the recorded time to the current time can be obtained, and it can be further refined to the credit line utilization rate arranged at a certain time interval until the current time. For example, when it is necessary to calculate with a month as the time truncation length, the credit consumption records within one month before the current time can be calculated, and then according to the credit line limit given at the current time, the credit line utilization rate within this time period can be determined. Then, the credit line utilization rates for the previous month, the previous two months, and so on until the original records are sampled and completed can be obtained. Of course, in the actual process, it is not limited to calculating the corresponding credit line usage records on a monthly basis. It is also feasible to calculate the corresponding credit line utilization rates on a half-monthly or weekly basis.
[0024] Step S20: Derive a basic dataset based on the features, and obtain a feature derivation result according to a data derivation algorithm.
[0025] It should be noted that the basic dataset for feature derivation is not the final analysis result, but the basis for providing feature inputs for subsequent data analysis and modeling. Further selection, transformation, or combination need to be performed according to specific analysis tasks and model requirements to finally form a feature matrix or feature vector for analysis and modeling.
[0026] It can be understood that, in order to further obtain a feature matrix or feature vector suitable for analysis and modeling, features with the highest correlation and predictive ability can be selected according to indicators such as feature importance, correlation, and collinearity. Statistical methods, machine learning models, or domain knowledge can also be used for feature selection, or the selected features can be further transformed, such as logarithmic transformation, standardization, normalization, etc., so that the features meet the requirements of model modeling while reducing the differences between features. Feature combination can also be performed, combining multiple features into new features, such as through addition, subtraction, multiplication, division, polynomial features, interaction features, etc., to enhance the expressive ability of features and capture more information.
[0027] It should be understood that the credit limit utilization rate can be aggregated in dimensions such as coefficient of variation, range, variance, maximum value, mean value, trimmed mean, standard deviation, etc. Since the specific credit limit utilization rate is related to the corresponding record range, for example, the derivative data generated by the credit limit usage within one month at the current time will have different explanatory powers from the derivative data generated within one week at the current time, and the applicable levels in model training are also different. Therefore, after determining the basic derivation algorithm, a series of data such as the mean value, trimmed mean, maximum value, sum value, range, standard deviation, and coefficient of variation are calculated for the obtained credit limit utilization rate according to the same derivation algorithm. The time range and time granularity corresponding to the credit limit utilization rate can be adjusted appropriately.
[0028] Step S30: According to the time loop strategy, adjust the time window of the data derivation algorithm, and recalculate to obtain the feature derivation result with updated time parameters.
[0029] It should be noted that according to the analysis task and the time range of the data, an appropriate initial time window size is selected, such as 3 months or 6 months. The specific selection can be determined according to the time distribution characteristics of the original credit limit usage records. If most of the single records in the original records are more than 12 months, the time window can be set to any time within 12 months at this time. For the reasonable utilization of data, a 30-day time window can be used, so that most records can be split into 12 groups of records. The obtained feature derivation result can be adjusted according to the same principle.
[0030] It can be understood that when the time factor changes, the meanings of the derivative eigenvalue obtained by the same set of derivative algorithms are different, and each newly generated set of derivative eigenvalues will be retained when the time parameter changes until the derivative eigenvalues generated in the combined results of all time windows are calculated. For example, when the time window is 30 days and the data in the original record has only valid records for the past year, 12 sets of data can be obtained according to the recent 30 days, the recent 30 days to 60 days, the recent 60 days to 90 days... and so on of the current time. When calculated according to twice the time window, that is, the recent 60 days of the current time, the recent 30 days to 90 days of the current time... and so on, 11 sets of data can be obtained. By repeating the above steps, according to the time cycle strategy, more derivative eigenvalues can be obtained by adjusting the actual time window size and the time sliding amount in the calculation based on the same basic data. Compared with directly inputting the original data into the model for training, the newly obtained derivative eigenvalue data has a larger data volume, and in terms of its mathematical meaning, the contribution ability of this type of data is better than the original data. After a certain data screening, it will be more suitable for modeling than the current original data.
[0031] Step S40: Perform model contribution degree analysis on the feature derivation result, screen out the feature derivation results with a contribution degree lower than the preset contribution degree, and use the remaining feature derivation results as high-dimensional feature variables.
[0032] It should be noted that performing model contribution degree analysis on the feature derivation result can help evaluate the importance of each feature and the contribution degree to the model prediction ability. By screening out the feature derivation results with a contribution degree lower than the preset contribution degree, the total amount of feature data can be reduced and the modeling efficiency can be improved. In addition, the actual model contribution degree analysis can evaluate the prediction effect of the established logistic regression model through commonly used model performance indicators in the industry (such as AUC and KS values). These indicators are used to measure the performance of the model on the training set and the validation set, and improve the accuracy of the logistic regression model in identifying lawyer letters.
[0033] It is understandable that the model performance metrics, the AUC (Area Under the ROC Curve) value and the KS (Kolmogorov-Smirnov) value, are used to evaluate the performance of a trained classification model. The AUC is the area under the ROC curve, representing the prediction ability of the classification model. The closer the AUC value is to 1, the better the prediction ability of the model; the closer the AUC value is to 0.5, the worse the prediction ability of the model. The KS value measures the model's prediction ability through the maximum vertical coordinate difference on the ROC curve. The larger the KS value, the greater the difference between positive and negative samples, and the better the discrimination ability of the model. Both the AUC value and the KS value are metrics for evaluating the performance of a classification model, used to measure the prediction accuracy and discrimination ability of the model.
[0034] It should be understood that the feature derivation results obtained in the previous steps are split into a training set and a validation set. The training set is used for model training and modeling, and the validation set is used for evaluating the model's capabilities. Specifically, it includes using the trained logistic regression model to predict the validation set and comparing it with the true labels of the validation set. The evaluation results can be obtained through the corresponding calculation methods of the model performance metrics, and further, the contribution degree of such feature values to model establishment can be obtained. Then, a part of the invalid data is filtered out according to methods such as the threshold or ranking of the contribution degree, and then a feature derivation result with a relatively high contribution degree can be obtained, and this part of the result is used as a high-dimensional feature variable.
[0035] Refer to Figure 3 , Figure 3 which is a schematic flow diagram of the second embodiment of the time-loop feature derivation method based on lawyer letter samples of the present invention.
[0036] Based on the above first embodiment, step S10 in the time-loop feature derivation method based on lawyer letter samples in this embodiment includes: Step S101: Based on the credit limit usage record database of the cooperation institution, obtain the original credit limit usage records within a preset time range.
[0037] It should be understood that, due to data privacy considerations, there are many legal issues in obtaining substantial consumption records. Therefore, during the process of obtaining actual basic data, simple desensitization markings are made on the original data. For example, for the credit consumption records of the same user, the user name is marked as a string of random data to hide the user's real information, and the credit usage records only record the rough time period and amount, without obtaining more in-depth privacy information. For example, the original usage records can be specific to minutes, the transaction object, and the transaction content, while during the desensitization process, only the time scale of days, the determination range of the amount, and the upper limit of the user's quota during this transaction period are recorded. Through the above steps, it can be ensured that the only function of such data is for machine learning purposes without disclosing the user's privacy.
[0038] It can be understood that due to the large amount of data and the uneven total credit transaction durations between users and banks, certain preliminary screening and processing need to be carried out in terms of the time range to make the subsequent feature generation smooth at the sample level.
[0039] Step S102: Split the original credit quota usage records according to a preset unit duration to obtain sub-data of credit quota usage records.
[0040] It can be understood that if the preset time range in the previous step is set as the original credit quota usage records in the past year near the current time, then the overall quota usage records can be split according to the unit duration. The preset unit duration can be set according to the actual situation. Generally, it is set according to a one-week duration (7 days) or a one-month duration (30 days). After splitting, the credit quota usage records within the unit duration are used as the splitting results of the original records.
[0041] Step S103: Calculate the sub-data of credit quota usage records to obtain the quota utilization rate corresponding to each sub-data of credit quota usage records.
[0042] It should be noted that there is a certain corresponding relationship between the quota usage records and the quota utilization rate. The quota utilization rates within consecutive multiple unit durations can, to a certain extent, reflect the user's credit consumption habits. In addition, due to the particularity of credit consumption, the quota utilization rates within the time periods close to the repayment date can also reflect the user's credit consumption habits at different time periods. In subsequent analysis and modeling, the quota utilization rate can be used as one of the features to assist in establishing a credit assessment model.
[0043] Step S104: Arrange the quota utilization rates according to the divided time to obtain a feature-derived basic data set.
[0044] It is understandable that the credit limit usage record refers to the amount of the credit limit actually used by a user within a specific time range. It records the amount of the credit limit used by the user at different time points or time periods. The limit utilization rate represents the proportion of the amount of the credit limit used by the user within a specific time range to the upper limit of the credit limit within that time range, usually expressed in percentage form. The limit utilization rate is an important indicator reflecting the usage degree of the credit limit and the management risk. A high limit utilization rate may indicate that the user has a relatively high dependence on credit and may not be able to bear more debts, while a low limit utilization rate can indicate that the user has better debt-bearing ability. Banks and financial institutions usually evaluate the credit risk of users based on the limit utilization rate and decide whether to adjust the credit limit or take other measures. Therefore, the limit utilization rate can be used as one of the features for constructing a credit assessment model to help predict the repayment ability and credit risk of users.
[0045] In this embodiment, by using the credit limit usage record database of the cooperative institution, the original credit limit usage records within a preset time range are obtained. The original credit limit usage records are split according to a preset unit time length to obtain sub-data of the credit limit usage records. The sub-data of the credit limit usage records are calculated to obtain the limit utilization rate corresponding to each sub-data of the credit limit usage records. The limit utilization rates are arranged according to the divided time to obtain a feature-derived basic data set. Through the above steps, in the process of obtaining the original credit limit usage records, the real information of the user is desensitized. Arranging according to the divided time can better utilize the information in the time dimension, better capture the credit usage trend and changes of the user, and provide a more accurate and reliable data basis for subsequent credit assessment and risk management.
[0046] Refer to Figure 4 , Figure 4 which is a schematic flowchart of the third embodiment of the time-loop feature derivation method based on lawyer's letter samples of the present invention.
[0047] Based on the above first embodiment, step S20 in the time-loop feature derivation method based on lawyer's letter samples in this embodiment includes: Step S201: Perform data aggregation processing on the feature-derived basic data set according to a preset time window to obtain the aggregated feature values required for feature derivation.
[0048] Further, the performing data aggregation processing on the feature-derived basic data set according to a preset time window to obtain the aggregated feature values required for feature derivation includes: Step S20101: Determine the sample segments participating in data aggregation in the feature-derived basic data set according to the preset time window.
[0049] It can be understood that the selection of the time window should be determined according to actual needs, and the selection of the window size can be determined according to the time distribution of the data and the analysis target. For example, a time window of three months, half a year, or one year can be selected. However, after the calculation of the eigenvalue starts, the size of the time window will remain unchanged. Only after the content of the sample segment is calculated can the time window parameters be changed accordingly.
[0050] Step S20102: Calculate the aggregated eigenvalue of the sample segment for the credit utilization rate data in the sample segment according to the data aggregation strategy. The aggregated eigenvalue includes the mean, trimmed mean, maximum value, sum value, range, standard deviation, and coefficient of variation of the sample segment.
[0051] It can be understood that for each sample segment, its credit utilization rate data can be aggregated. An appropriate aggregation strategy can be selected according to requirements. For example, statistical indicators such as the average value, maximum value, and minimum value of the credit utilization rate within the time window can be calculated, as well as other aggregated eigenvalues as needed.
[0052] It should be understood that the mean can be used as a measure of the central tendency of the data, reflecting the average level of the data. The trimmed mean is calculated by excluding a certain proportion of the maximum and minimum values from the data before calculating the mean. The role of the trimmed mean is to reduce the influence of outliers on the mean and more robustly describe the central tendency of the data. The maximum and minimum values can provide the range of the data to help understand the boundaries of the data. The sum value can show the cumulative amount of the data and is used to measure the overall magnitude of the data. The range is the difference between the maximum value and the minimum value. The range can show the dispersion of the data. A larger range indicates greater data fluctuations, and a smaller range indicates smaller data fluctuations. The standard deviation is a measure of the dispersion degree of a set of data, indicating the average deviation of the data from the mean. The larger the standard deviation, the higher the volatility of the data. The coefficient of variation is the ratio of the standard deviation to the mean and is used to compare the relative variation degrees of different data sets. A data set with a smaller coefficient of variation indicates relatively lower variability.
[0053] It should be noted that the sample segment refers to the data of a single user, and the corresponding credit utilization rate data is calculated based on this. The calculation of the aggregated eigenvalue can help extract the overall characteristics of the sample segment and reduce the data scale, facilitating subsequent feature derivation and analysis.
[0054] Step S202: Generate high-dimensional derivative eigenvalues according to the aggregated eigenvalues according to a preset conversion algorithm.
[0055] Furthermore, the generation of high-dimensional derivative eigenvalues according to the aggregated eigenvalues according to a preset conversion algorithm includes: Step S20201: Determine the quota utilization rate within each unit time period in the sample segment according to the aggregated eigenvalue and the preset unit time period, and calculate the change amount of the quota utilization rate between adjacent unit time periods in the sample segment.
[0056] It can be understood that the change in the quota usage amount within two adjacent unit time periods can, to a certain extent, reflect the changing trend of credit consumption habits. For example, when the quota utilization rate is relatively high but shows a downward trend in several consecutive unit time periods, it may still indicate that the consumer willingness of this user has declined during this period. On the other hand, when the overall quota utilization rate is relatively low but shows an upward trend in several consecutive unit time periods, it represents an upward trend in the consumer willingness of the user. This is more advanced and accurate than the traditional method of simply classifying based on the high or low level of the credit quota utilization rate.
[0057] It should be noted that since the time window is fixed before the calculation starts, the preset unit time period can be used to adjust the fineness of the trend change of the calculated quota utilization rate. For example, when the preset unit time periods are 30 days and 7 days respectively, they can reflect the monthly consumption trend and the weekly change trend of the quota utilization rate respectively.
[0058] Step S20202: Determine the overall change trend of the sample segment according to the change amount of the quota utilization rate.
[0059] It can be understood that determining the overall change trend of the sample segment according to the change amount of the quota utilization rate can be understood as judging the overall trend of the sample segment, such as rising, falling or fluctuating, based on the positive or negative and magnitude of the change amount of the quota utilization rate. In addition, the trend result reflected by the time period closer to the current time can better reflect the overall credit quota usage situation of the public under the current economic conditions.
[0060] Step S20203: Calculate the relative proportion value of the quota utilization rate within each unit time period in the sample segment.
[0061] It can be understood that the relative proportion value here is the ratio of two data. The denominator part can choose the maximum value, minimum value or average value of the quota utilization rate in this sample segment. Through this comparison process, the quota usage habit of this user and the change of the usage habit can be determined.
[0062] It should be understood that when there is a situation where the quota utilization rate is zero within a certain time period, and the number of such time periods will reflect a change in the concept of this user towards credit consumption or a potential change in consumption habits. The calculation result of this part of the data in terms of relative proportion is zero.
[0063] Step S20204: Use the change amount of the quota utilization rate, the change trend, and the relative proportion value of the quota utilization rate within the unit time of the sample segment as high-dimensional derivative feature values.
[0064] It can be understood that in the sample segment, high-dimensional derivative feature values are obtained by calculating the change amount and change trend of the quota utilization rate. The change amount of the quota utilization rate indicates what kind of changes have occurred in the quota usage situation within a period of time and the degree of change, which can be used to measure the stability of the user's credit behavior. The change trend indicates whether the quota utilization rate is gradually increasing or decreasing, reflecting the changing trend of the user's credit behavior. In addition, the relative proportion value of the quota utilization rate within the unit time can also be calculated to measure the user's usage habits in different time periods, such as the relative proportion of the quota utilization rate during the day and at night. It should be noted that these derivative feature values can enhance the model's ability to judge the user's credit status and provide more comprehensive information to support credit assessment and decision-making.
[0065] Step S203: Use the aggregated feature values and the high-dimensional derivative feature values as the feature derivation results.
[0066] It should be understood that various types of feature values in the above steps can all reflect some effective information. Integrate the feature values obtained by the aggregation method and the feature values obtained by the derivation method and use them together as the feature derivation results.
[0067] In this embodiment, the basic data set is aggregated according to a preset time window, and high-dimensional derivative feature values are generated based on the aggregated feature values and a preset conversion algorithm. These feature values can provide more comprehensive information to support credit assessment, such as the quota utilization rate, change trend, and relative proportion value, etc. By using the aggregated feature values and high-dimensional derivative feature values as the results of feature derivation, these feature values can provide more comprehensive information and help the model more accurately evaluate the user's credit status. By introducing these feature values, the credit assessment model can more accurately predict aspects such as the user's repayment ability, debt risk, and credit reliability.
[0068] Refer to Figure 5 , Figure 5 which is a schematic flowchart of the fourth embodiment of the time loop feature derivation method based on lawyer's letter samples of the present invention.
[0069] Based on the above first embodiment, in the time loop feature derivation method based on lawyer's letter samples in this embodiment, the step S40 includes: Step S401: Divide the feature derivation results into multiple groups of derivative feature value data according to the data type.
[0070] It should be noted that in the above steps, derivative eigenvalue data of various data types can be obtained through the aggregation algorithm and the derivative algorithm. Each type of feature derivative value data is related to both the calculation method and the time parameters during the calculation. Certain conditions are required for classification to distinguish the contribution ability of each type of derivative data to subsequent modeling.
[0071] Step S402: Divide each group of derivative eigenvalue data into a training set and a validation set.
[0072] It should be noted that in machine learning, the dataset is usually divided into a training set and a validation set to evaluate the performance of the model and perform model tuning. The training set is usually a dataset containing known input features and corresponding output labels. By training the model on the training set, the model can learn the relationship between the input features and the output labels, so as to perform accurate prediction or classification. The validation set is a dataset used to evaluate the performance of the model and perform model selection. The validation set is a part of the data separated from the original dataset. It contains data samples that have not been used in the training process. The validation set can be used to check the performance of the model on unseen data, and evaluate the accuracy, generalization ability and robustness of the model. By evaluating on the validation set, the model can be tuned, such as adjusting the hyperparameters of the model or optimizing the algorithm, to improve the performance of the model.
[0073] It can be understood that the division of the training set and the validation set should be random, and the distribution of the data and the representativeness of the features should be maintained. This can ensure the performance ability of the model in the real scenario. To better evaluate the generalization ability of the model, sometimes a third independent test set is also used to evaluate the final performance of the model.
[0074] Step S403: Input the training set into the initial logistic regression model for training to obtain a trained logistic regression model.
[0075] It should be noted that one type of derivative eigenvalue corresponds to a set of training set and validation set. That is to say, one type of derivative eigenvalue as the training set can first train the corresponding logistic regression model, and then verify the actual performance of the contribution degree of this type of derivative eigenvalue to the model through the validation set.
[0076] Step S404: Evaluate the performance of the trained logistic regression model through the corresponding validation set to obtain model performance index data.
[0077] Furthermore, input the feature data of the validation set into the trained logistic regression model to obtain the model prediction result; according to the model prediction result and the true result of the validation set, calculate the AUC value and KS value of the logistic regression model; use the AUC value and KS value as the model performance index data of this model.
[0078] It should be noted that before calculating the AUC value and the KS value, the prediction results of the model are compared with the true results of the validation set. The AUC value refers to "Area Under the Curve", that is, the area under the ROC curve. The ROC curve plots the variation between the true positive rate and the false positive rate of the model at different thresholds. Calculating the AUC value can evaluate the classification ability of the model for positive and negative samples. The higher the AUC value, the better the classification performance of the model. The KS value refers to the Kolmogorov-Smirnov statistic, which is used to evaluate the difference in the distribution of positive and negative samples of the model at different thresholds. The KS value represents the maximum difference between the cumulative distribution functions of positive and negative samples. The larger the KS value, the stronger the discrimination ability of the model. By calculating the AUC value and the KS value, the classification performance and discrimination ability of the model can be comprehensively evaluated.
[0079] It should be understood that recording the calculated AUC value and KS value as the model performance index data of this logistic regression model can be used to compare the performance of different models, select a model with better performance, and determine its corresponding derived eigenvalue.
[0080] Step S405: According to the model performance index data, determine the model contribution degree of each group of derived eigenvalue data, screen out the feature derivation results lower than the preset contribution degree, and use the remaining feature derivation results as high-dimensional feature variables.
[0081] It should be noted that although the directly calculated AUC value and KS value can both reflect the actual level of the model to a certain extent, the actual ability of the machine learning model needs to be considered comprehensively from multiple aspects. Since the AUC value and the KS value evaluate the classification performance and discrimination ability of the model respectively, it is necessary to perform a certain weight assignment to calculate the comprehensive level of the model in order to correspondingly determine the model contribution degree of each group of derived eigenvalue data.
[0082] It can be understood that after determining the model contribution degree of each group of derived eigenvalue data, a preset contribution degree can be set, and the feature derivation results lower than the preset contribution degree are screened out. The feature derivation results that are not screened out can be regarded as feature variables with high interpretability, rich business meaning, and high-dimensional value information.
[0083] In this embodiment, the feature derivation results are divided into multiple groups of derived feature value data according to the data type; each group of derived feature value data is divided into a training set and a validation set; the training set is input into the initial logistic regression model for training to obtain a trained logistic regression model; the performance of the trained logistic regression model is evaluated through the corresponding validation set to obtain model performance index data; according to the model performance index data, the model contribution degree of each group of derived feature value data is determined, the feature derivation results below the preset contribution degree are screened out, and the remaining feature derivation results are used as high-dimensional feature variables. Through the above steps, training and performance evaluation are respectively carried out on each group of derived feature value data, which can more accurately understand the contribution of each group of feature derivation results to the model performance, and screen out the feature derivation results below the preset contribution degree, which can help extract important feature derivation results, optimize the training and prediction processes of the model, thereby improving the model performance and prediction accuracy. In addition, it can also reduce the model complexity, improve the interpretability and application effect of the model.
[0084] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not limit this.
[0085] It should be understood that although the steps in the flowchart in the embodiments of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and they can be executed in other orders. Moreover, at least some of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0086] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and there is no limitation here.
[0087] In addition, it should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such a process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0088] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments.
[0089] The above are only the preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for deriving time-loop features based on lawyer's letter samples, characterized in that, The time loop feature derivation method based on lawyer's letter samples includes: Obtain the original credit line usage records in the lawyer's letter samples, and based on the original credit line usage records, obtain a feature derivation basic data set; Based on the feature derivation basic data set, according to the data derivation algorithm, obtain the feature derivation results; According to the time loop strategy, adjust the time window of the data derivation algorithm, and recalculate to obtain the feature derivation results with updated time parameters; Conduct a model contribution degree analysis on the feature derivation results, screen out the feature derivation results with a contribution degree lower than the preset contribution degree, and use the remaining feature derivation results as high-dimensional feature variables.
2. The method for deriving time loop features based on lawyer's letter samples according to claim 1, wherein, The obtaining of the original credit line usage records in the lawyer's letter samples and obtaining the feature derivation basic data set based on the original credit line usage records includes: Based on the credit line usage record database of the cooperative institution, obtain the original credit line usage records within a preset time range; Split the original credit line usage records according to a preset unit duration to obtain sub-data of credit line usage records; Calculate the usage rate of each sub-data of credit line usage records to obtain the usage rate corresponding to each sub-data of credit line usage records; Arrange the usage rates according to the divided time to obtain the feature derivation basic data set.
3. The method for deriving time loop features based on lawyer's letter samples according to claim 1, wherein, The obtaining of the feature derivation results based on the feature derivation basic data set according to the data derivation algorithm includes: Conduct data aggregation processing on the feature derivation basic data set according to a preset time window to obtain the aggregated feature values required for feature derivation; According to the aggregated feature values, generate high-dimensional derived feature values according to a preset conversion algorithm; Use the aggregated feature values and the high-dimensional derived feature values as the feature derivation results.
4. The method for deriving the time loop feature based on the lawyer's letter sample according to claim 3, wherein The processing of the feature derivation basic data set according to a preset unit duration based on the data aggregation strategy to obtain the aggregated feature values required for feature derivation includes: According to the preset time window, determine the sample segments participating in data aggregation in the feature derivation basic data set; For the usage rate data in the sample segments, calculate the aggregated feature values of the sample segments according to the data aggregation strategy; The aggregated feature values include the mean, trimmed mean, maximum value, sum value, range, standard deviation, and coefficient of variation of the sample segments.
5. The method for deriving time loop features based on lawyer's letter samples according to claim 3, wherein, The generating of high-dimensional derived feature values according to the aggregated feature values according to a preset derivation algorithm includes: According to the aggregated feature values and the preset unit duration, determine the usage rate of each unit duration in the sample segment, and calculate the change amount of the usage rate of adjacent unit durations in the sample segment; According to the change amount of the usage rate, determine the overall change trend of the sample segment; Calculate the relative proportion value of the usage rate of each unit duration in the sample segment; Use the change amount of the usage rate, change trend of the sample segment, and the relative proportion value of the usage rate within the unit duration as the high-dimensional derived feature values.
6. The method for deriving the time loop feature based on the lawyer's letter sample according to claim 1, wherein The conducting of a model contribution degree analysis on the feature derivation results, screening out the feature derivation results with a contribution degree lower than the preset contribution degree, and using the remaining feature derivation results as high-dimensional feature variables includes: Divide the feature derivation results into multiple groups of derived feature value data according to the data type; Divide each group of derived feature value data into a training set and a validation set; Input the training set into the initial logistic regression model for training to obtain a trained logistic regression model; Evaluate the performance of the trained logistic regression model through the corresponding validation set to obtain model performance index data; According to the model performance index data, determine the model contribution degree of each group of derived feature value data, screen out the feature derivation results lower than the preset contribution degree, and use the remaining feature derivation results as high-dimensional feature variables.
7. The method for deriving the time loop feature based on the lawyer's letter sample according to claim 6, characterized in that, The performance evaluation of the trained logistic regression model through the corresponding validation set to obtain model performance index data includes: Input the feature data of the validation set into the trained logistic regression model to obtain a model prediction result; Calculate the AUC value and KS value of the logistic regression model according to the model prediction result and the true result of the validation set; Use the AUC value and KS value as the model performance index data of the model.