Machine Learning-Based Intelligent Sorting Method, System and Device for Outbound Calls of Non-Performing Assets

Through data cleaning and standardization of non-performing asset out-of-call records, combined with the integration of multiple machine learning models and dynamic feature engineering, the problem of insufficient sorting accuracy and efficiency in the existing technology is solved, and efficient non-performing asset out-of-call sorting is achieved, and real-time optimization capabilities are provided.

CN119172476BActive Publication Date: 2025-07-18ZHEJIANG JIUMU HOLDING GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411273215.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-07-18
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

The existing non-performing asset out-of-call sorting methods cannot take into account accuracy and efficiency, the traditional methods cannot fully consider the borrower's real-time financial situation changes and behavioral patterns, and it is difficult to deal with data imbalance and feature engineering.

Method used

By cleaning and standardizing the data set of out-of-call records of non-performing assets, extracting time, amount, behavior and credit characteristics, selecting and combining features, using multiple machine learning models to build comprehensive sorting indicators, and dynamically adjusting weights and feature importance based on out-of-call feedback, and using incremental learning optimization model.

Benefits of technology

It improves the ordering accuracy and efficiency of non-performing asset outgoing calls, realizes continuous optimization based on real-time feedback, and has strong adaptability and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119172476B_ABST
    Figure CN119172476B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system and device for intelligent sorting of bad asset outbound calls based on machine learning. The method includes: performing data cleaning and data standardization on a data set containing bad asset outbound call records, performing feature selection on the extracted information, and creating composite features through feature combination; training one or more types of models using the composite features, and performing model fusion based on the Stacking integration method to obtain a fusion probability prediction; constructing a comprehensive sorting index based on the fusion probability prediction; sorting outbound customers in batches according to the comprehensive sorting index, and updating the weight coefficients in the comprehensive sorting index based on the outbound call feedback effect; dynamically adjusting the weights of training samples based on the outbound call feedback data to update the model in an incremental learning manner. Through the solution of the present application, the accuracy and efficiency of sorting can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electrical data processing, and particularly to an intelligent sorting method, system, and device for outbound calls of non-performing assets based on machine learning. Background Art

[0002] Outbound call sorting of non-performing assets has become an important link in the management and disposal of non-performing loans by financial institutions, which is closely related to improving asset recovery rates, optimizing resource allocation, and reducing operating costs. However, not all non-performing assets have the same recovery potential and urgency, which leads to the problem of distinguishing high-value targets from low-value targets. High-value targets are usually borrowers with higher repayment ability or willingness, while low-value targets may face more serious financial difficulties or lack of repayment willingness. Therefore, accurately identifying and prioritizing high-value targets is crucial for improving the efficiency and effectiveness of non-performing asset disposal. Outbound call sorting of non-performing assets involves analyzing and evaluating a large amount of borrower data to determine the optimal outbound call order. This process can not only improve the success rate of collection but also significantly reduce unnecessary human and time costs, thereby optimizing the overall resource allocation. Currently, the methods for outbound call sorting of non-performing assets are mainly based on traditional credit scoring models, historical repayment record analysis, and simple rule-based sorting. These methods determine the outbound call priority by evaluating factors such as the borrower's credit status, overdue time, and amount of arrears. Although these methods are effective to a certain extent, they have obvious limitations. For example, they cannot fully consider the changes in the borrower's real-time financial situation, ignore the impact of the external economic environment on the repayment ability, and it is difficult to capture the borrower's behavior patterns and psychological characteristics.

[0003] In recent years, with the rapid development of big data and artificial intelligence technologies, more advanced methods for outbound call sorting of non-performing assets have gradually attracted attention. Machine learning algorithms, especially deep learning models, have become the focus of research because they can process and analyze massive multi-dimensional data. These methods construct a more comprehensive borrower profile by integrating multi-source data, including the borrower's social media activities, consumption behavior, and changes in employment status. By learning the patterns in historical data through complex algorithm models, these methods can more accurately predict the borrower's repayment probability and the best contact time. In actual application scenarios, due to the limited availability of high-quality labeled data, model training often requires techniques such as data augmentation and transfer learning to improve the generalization ability of the model. However, the existing methods still face challenges in dealing with data imbalance, feature engineering, and model interpretability. For example, traditional random oversampling or undersampling techniques may lead to model overfitting or information loss.

[0004] Therefore, there is an urgent need for a technical solution to improve the accuracy and efficiency of sorting. Summary of the Invention

[0005] To address the deficiencies of the prior art, embodiments of the present application provide a method, system, and device for intelligent sorting of non-performing asset outbound calls based on machine learning. The present application solves technical problems such as the inability to balance accuracy and efficiency in the prior art.

[0006] Embodiments of the present application provide a method for intelligent sorting of non-performing asset outbound calls based on machine learning, including: performing data cleaning and data standardization on a dataset containing non-performing asset outbound call records to extract information including time features, amount features, behavior features, outbound call features, and / or credit features; performing feature selection on the extracted information and creating composite features through feature combination; training one or more types of models using the composite features and performing model fusion based on the Stacking integration method to obtain a fusion probability prediction; where the models include logistic regression, random forest, gradient boosting decision tree, and / or support vector machine; constructing a comprehensive sorting index based on the fusion probability prediction; sorting outbound customers in batches according to the comprehensive sorting index and updating the weight coefficients in the comprehensive sorting index based on the outbound feedback effect; dynamically adjusting the weights of training samples based on outbound feedback data to update the model through incremental learning.

[0007] In a possible implementation, performing data cleaning and data standardization on a dataset containing non-performing asset outbound call records includes: removing duplicate data and unifying the data format; filling numerical features with the mean, median, or mode; filling categorical features with the mode or creating unknown categories; using the principle or the interquartile range method to identify outliers for deletion or adjustment according to the outlier threshold; performing standardization processing on numerical features.

[0008] In a possible implementation, performing feature selection on the extracted information and creating composite features through feature combination includes: calculating the correlation coefficient matrix between features to remove highly correlated features; calculating the variance between features to remove features with variance below the variance threshold; using a random forest model to rank the feature importance to select a target number of features; performing polynomial combination on numerical features and / or cross combination on categorical features to construct composite features.

[0009] In a possible implementation, the method for obtaining the feature importance includes: , where, represents the importance of feature , represents a node in the decision tree, represents the child nodes on the node, represents summing over all nodes split on feature and performing summation, Indicates the number of samples in the node, Indicates the total number of samples, Indicates the entropy of the node Indicates the entropy of the child nodes on the node Indicates the left and right child nodes, Indicates the number of samples in the child nodes on the node

[0010] In one possible implementation, a comprehensive sorting index is constructed based on the fusion probability prediction, including: , where Indicates the comprehensive sorting index, Indicates the weight coefficient, Indicates the fusion probability prediction of repayment, Indicates the overdue amount, Indicates the customer value score.

[0011] In one possible implementation, the outbound customers are sorted in batches according to the comprehensive sorting index, and the weight coefficients in the comprehensive sorting index are updated based on the outbound feedback effect, including: evenly dividing the outbound customers into multiple batches based on the comprehensive sorting index; for each batch, the outbound customer with the highest comprehensive sorting index is selected with the first probability, and the outbound customers are randomly selected with the second probability; the method for determining the first probability includes: , where Indicates the first probability, Indicates the initial probability, Indicates the attenuation coefficient, Indicates the current batch number; a corresponding prior distribution is set for each weight coefficient in the comprehensive sorting index, and the distribution parameters are updated according to each outbound feedback effect to adjust the weight coefficients.

[0012] In one possible implementation, updating the distribution parameters according to each outbound feedback effect includes: , where and Indicates the prior Beta distribution parameter of the th weight coefficient, Indicates the success rate corresponding to the th weight coefficient.

[0013] In one possible implementation, dynamically adjusting the weights of the training samples based on the outbound feedback data includes: , where Indicates the adjusted sample weight,​​​ represents the sample weight before adjustment, represents the actual label, represents the predicted label, represents the learning rate.

[0014] The embodiment of the present application also provides an intelligent sorting system for outbound calls of non-performing assets based on machine learning, including: a data preprocessing unit, a prediction unit, and a sorting unit; wherein, the data preprocessing unit is used to perform data cleaning and data standardization on a data set containing outbound call records of non-performing assets, so as to extract information including time features, amount features, behavior features, outbound call features, and / or credit features; perform feature selection on the extracted information, and create composite features through feature combination; the prediction unit is used to train one or more types of models using the composite features, and perform model fusion based on the Stacking integration method to obtain a fusion probability prediction; the sorting unit is used to construct a comprehensive sorting index based on the fusion probability prediction; batch sort the outbound call customers according to the comprehensive sorting index, and update the weight coefficients in the comprehensive sorting index based on the outbound call feedback effect; the prediction unit is also used to dynamically adjust the weight of the training samples based on the outbound call feedback data, so as to update the model in an incremental learning manner.

[0015] The embodiment of the present application also provides an intelligent sorting device for outbound calls of non-performing assets based on machine learning, including: a processor, a memory, and a system bus; wherein, the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the method described in the above embodiment.

[0016] In the intelligent sorting method, system, and device for outbound calls of non-performing assets provided above, the embodiment of the present application can realize the intelligent sorting of non-performing assets through multi-dimensional data analysis, dynamic feature engineering, multi-model fusion, and real-time feedback optimization, which not only improves the outbound call efficiency and recovery rate, but also can be continuously optimized according to real-time feedback, and has strong adaptability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flow chart of an intelligent sorting method for outbound calls of non-performing assets based on machine learning provided by the embodiment of the present application;

[0019] Figure 2 This is a schematic block diagram of a non-performing asset outbound intelligent sorting system based on machine learning provided by an embodiment of the present application. Detailed implementation manners

[0020] Now, various exemplary embodiments of the present application will be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present application.

[0021] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present application are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them. It should also be understood that in the embodiments of the present application, "a plurality of" may refer to two or more, and "at least one" may refer to one, two, or more. It should also be understood that for any component, data, or structure mentioned in the embodiments of the present application, without clear limitation or contrary indication in the context, it can generally be understood as one or more. In addition, the term "and / or" in the present application is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects before and after. It should also be understood that the present application emphasizes the differences between the various embodiments, and the same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated one by one.

[0022] At the same time, it should be understood that for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present application and its application or use. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the specification. It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0024] Figure 1 It is a schematic flowchart of a method for intelligent sorting of bad asset outbound calls based on machine learning provided by an embodiment of this application. It should be noted that the disposal of bad assets is one of the important challenges faced by financial institutions. Traditional methods for sorting bad asset outbound calls often rely on manual experience, with low efficiency and prone to deviation. The solution of this application proposes a method for intelligent sorting of bad asset outbound calls based on machine learning, which realizes the intelligent sorting of bad assets through multi-dimensional data analysis and model training, improving the outbound call efficiency and recovery rate.

[0025] As Figure 1 shown, at step S01, the dataset containing bad asset outbound call records is subjected to data cleaning and data standardization to extract information including time features, amount features, behavior features, outbound call features, and / or credit features.

[0026] The dataset containing bad asset outbound call records may include, but is not limited to:

[0027] Basic customer information: age, gender, occupation, income level, etc.;

[0028] Loan information: loan amount, loan term, overdue duration, overdue amount, etc.;

[0029] Historical transaction data: repayment records, consumption habits, etc.;

[0030] Credit scores: internal scores, external credit investigation scores, etc.;

[0031] Contact information: phone number, address, etc.;

[0032] Outbound call records: historical outbound call times, connection rate, call duration, etc.

[0033] After obtaining the data with the customer's authorization and permission, the dataset containing bad asset outbound call records is subjected to data cleaning and data standardization. Among them, it includes: removing duplicate data and unifying the data format; filling numerical features with the mean, median, or mode; filling categorical features with the mode or creating unknown categories; using The principle or interquartile range method is used to identify outliers for deletion or adjustment according to the outlier threshold, which can be set by relevant business personnel according to their business experience; the numerical features are standardized. In order to eliminate the dimensional differences between different features, the Z-score standardization method can be used to standardize the numerical features: , where x is the original value, μ is the mean, and σ is the standard deviation.

[0034] It should be understood that for extracting information including time features, amount features, behavior features, outbound call features, and / or credit features, the following content can be extracted:

[0035] Time-related features: days overdue, time since last repayment, remaining loan term;

[0036] Amount-related features: ratio of overdue amount to total loan amount, ratio of average monthly repayment amount to income;

[0037] Behavior-related features: historical repayment times, historical overdue times, longest consecutive normal repayment periods;

[0038] Outbound call-related features: historical outbound call times, historical connection rate, average call duration;

[0039] Credit-related features: internal credit score, trend of external credit investigation score change.

[0040] At step S102, feature selection is performed on the extracted information, and composite features are created through feature combination. Specifically, it includes the following steps: calculating the correlation coefficient matrix between features to remove highly correlated features; calculating the variance between features to remove features with variance lower than the variance threshold; using the random forest model to rank the feature importance to select the target number of features; performing polynomial combination on numerical features and / or cross combination on categorical features to construct composite features.

[0041] Among them, the determination of highly correlated and variance threshold can be flexibly adjusted according to business requirements. The methods for obtaining the feature importance include: , where represents the importance of feature , represents the node in the decision tree, represents the child node on the node, represents the sum of all nodes split on feature , performing summation, represents the sample number in node , represents the total sample number, represents the node The entropy of represents the entropy of the child nodes on the node of which represents the left and right child nodes, represents the number of samples in the child nodes on the node in

[0042] The importance of a feature is measured by calculating the average reduction in impurity (such as Gini impurity or entropy) when the feature is used as a splitting node in all trees. For example, the calculation steps may include the following: For each tree in the random forest, traverse each internal node in the tree. If the node uses the feature for splitting, calculate the change in impurity before and after splitting, multiply the change by the weight of the node ( ), and accumulate it into the importance score of the feature. Take the average of the results for all trees to obtain the final feature importance score.

[0043] The specific implementation can be achieved by training a random forest model, traversing all internal nodes for each tree, calculating the change in impurity before and after splitting for each node, accumulating the importance score according to the splitting feature, normalizing the importance scores of all features, sorting the features according to the importance scores, and finally selecting the top-k features.

[0044] At step S103, one or more types of models are trained using composite features, and model fusion is performed based on the Stacking ensemble method to obtain a fused probability prediction.

[0045] In one embodiment, the following four different types of models can be selected as the base models: Logistic Regression (LR), Random Forest (RF), Gradient Boosting Decision Tree (GBDT), and Support Vector Machine (SVM). Each base model is trained, and the k-fold cross-validation method is used to evaluate the model performance. Taking Logistic Regression as an example, its objective function is: , where m is the number of samples, n is the number of features, is the prediction function, and λ is the regularization parameter.

[0046] Next, the Stacking ensemble method is used for model fusion: The prediction results of the base models are used as new features, and a meta-model (such as LightGBM) is trained as the final classifier. The calculation formula for the fused prediction probability is: , where is the prediction function of the meta-model, are the prediction probabilities of each base model respectively.

[0047] At step S104, a comprehensive ranking metric is constructed based on the fusion probability prediction. In one implementation scenario, it can be constructed as: , where represents the comprehensive ranking metric, represents the weight coefficient, represents the fusion probability prediction of repayment, represents the overdue amount, represents the customer value score.

[0048] In one implementation scenario, R can be further divided into several objectively quantifiable metrics. For example, it may include but is not limited to the following factors: historical repayment record, current economic situation, asset situation, income stability, and communication and cooperation with the company. The quantification scheme for the metrics can be set as:

[0049] Historical repayment record (full score 100 points), number of on-time repayments in the past 12 months: 5 points each time, maximum 60 points, time since the most recent overdue: within 3 months (0 points), 3 - 6 months (10 points), 6 - 12 months (20 points), 12 months and above (40 points).

[0050] Current economic situation (full score 100 points), employment status: full-time (50 points), part-time (30 points), unemployed (0 points). Monthly income level: < 3000 yuan (10 points), 3000 - 5000 yuan (20 points), 5000 - 10000 yuan (30 points), 10000 yuan and above (50 points).

[0051] Asset situation (full score 100 points), real estate: having own real estate (50 points), having no own real estate (0 points). Vehicle: having a car (30 points), having no car (0 points). Value of other liquidatable assets: < 10,000 yuan (5 points), 10,000 - 50,000 yuan (10 points), 50,000 yuan and above (20 points).

[0052] Income stability (full score 100 points), current working years: < 1 year (10), 1 - 3 years (20), 3 - 5 years (30), 5 years and above (40). Occupation type: civil servant / government agency employee (30), state-owned enterprise employee (25), private enterprise regular employee (20), self-employed (10), others (5). Whether having a side job: having (30), not having (0).

[0053] Communication and cooperation degree (full score 100 points), answering rate: 80% (40 points), 60 - 80% (30 points), 40 - 60% (20 points), < 40% (10 points). Attitude: actively cooperate (40), general (20), passive resistance (0). Whether actively providing a repayment plan: yes (20), no (0).

[0054] Calculation of customer value score R: R = (b1 * historical repayment record + b2 * current economic status + b3 * asset status + b4 * income stability + b5 * communication and cooperation degree) / (b1 + b2 + b3 + b4 + b5). Among them, the weights of b1 to b5 can be adjusted according to the actual situation. For example: b1 = 0.3, b2 = 0.2, b3 = 0.15, b4 = 0.2, b5 = 0.15. Finally, the score range of R is 0 - 100 points.

[0055] This quantification scheme can be adjusted according to the actual situation. For example, certain sub-items can be added or deleted, or the score distribution of each item can be adjusted. At the same time, the score can be evaluated and adjusted regularly (such as every quarter) according to the actual collection effect to ensure its continuous effectiveness.

[0056] At step S105, the outbound customers are sorted in batches according to the comprehensive sorting index, and the weight coefficients in the comprehensive sorting index are updated based on the outbound feedback effect. It includes: evenly dividing the outbound customers into multiple batches based on the comprehensive sorting index; for each batch, the outbound customer with the highest comprehensive sorting index is selected with the first probability, and the outbound customers are randomly selected with the second probability; the method for determining the first probability includes: , where represents the first probability, represents the initial probability, represents the attenuation coefficient, represents the current batch number; a corresponding prior distribution is set for each weight coefficient in the comprehensive sorting index, and the distribution parameters are updated according to the outbound feedback effect each time to adjust the weight coefficients.

[0057] Specifically, after batch outbound calls, a prior distribution (such as Beta distribution) is set for each weight coefficient in the comprehensive sorting index. After each outbound call, the distribution parameters are updated according to the actual effect, and new weight coefficients are sampled from the updated distribution. Among them, updating the distribution parameters according to the outbound feedback effect each time includes: , where and represents the th prior Beta distribution parameter of the weight coefficient, represents the success rate corresponding to the th weight coefficient. For example, in the embodiment described above, the three weight coefficients can be updated respectively.

[0058] At step S106, the weights of the training samples are dynamically adjusted based on the outbound call feedback data to update the model in an incremental learning manner. Among them, dynamically adjusting the weights of the training samples based on the outbound call feedback data includes: , where represents the adjusted sample weight, represents the sample weight before adjustment, and can be initialized within the numerical range of 0.75 to 1.25, represents the actual label, represents the predicted label, represents the learning rate. At the same time, an online learning algorithm (such as Online Gradient Descent) can be used for incremental update of the model: , where is the model parameter, η is the learning rate, L is the loss function, and are the features and labels of the new sample.

[0059] In addition, in order to be able to periodically re-evaluate the feature importance, the feature set can be dynamically adjusted: calculate the SHAP (SHapley Additive exPlanations) value of each feature, sort the features according to the SHAP value, retain the top k features with the highest importance, and eliminate or combine the features with low importance. The calculation formula of the SHAP value is: , where N is the set of all features, S is the subset that does not contain feature i, is the predicted value of the model on the feature subset S.

[0060] In summary, in the embodiment of the present application, through multi-dimensional data fusion, combining multi-dimensional data such as customer basic information, loan information, transaction data, and credit scores, the customer characteristics are comprehensively characterized. Through feature extraction, selection, and combination, a rich feature set is constructed, and the feature importance is dynamically adjusted according to the feedback. The Stacking integration method is used to fuse different types of machine learning models to improve the prediction accuracy. Combining the dynamic sorting strategy, the balance between exploration and exploitation is achieved, and the sorting weights are dynamically adjusted. Through sample weight update, incremental learning, and dynamic evaluation of feature importance, the continuous optimization of the model is realized. In short, the intelligent sorting method for bad assets outbound calls based on machine learning proposed in the solution of the present application realizes the intelligent sorting of bad assets through multi-dimensional data analysis, dynamic feature engineering, multi-model fusion, and real-time feedback optimization. This method not only improves the outbound call efficiency and recovery rate, but also can be continuously optimized according to real-time feedback, and has strong adaptability and scalability.

[0061] Figure 2A schematic block diagram of a non-performing asset outbound intelligent sorting system based on machine learning provided by an embodiment of the present application. It should be understood that the system shown in the figure is exemplary rather than restrictive. This means that the involved system architecture is not limited to a specific form or design, but is presented as an example. In other words, the architecture shown in the figure can be regarded as a way of expression to clearly describe relevant concepts and relationships, and does not exclude other forms of architecture. Therefore, when interpreting the architecture in the said picture, it should be understood that the model has flexibility and diversity, and its purpose is to provide an exemplary description rather than a restrictive regulation of a specific form.

[0062] The embodiment of the present application includes a data preprocessing unit 201, a prediction unit 202, and a sorting unit 203; wherein, the data preprocessing unit 201 is used to perform data cleaning and data standardization on a data set containing non-performing asset outbound records, so as to extract information including time features, amount features, behavior features, outbound features, and / or credit features; perform feature selection on the extracted information, and create composite features through feature combination; the prediction unit 202 is used to train one or more types of models using the composite features, and perform model fusion based on the Stacking integration method to obtain a fusion probability prediction; the sorting unit 203 is used to construct a comprehensive sorting index based on the fusion probability prediction; batch sort the outbound customers according to the comprehensive sorting index, and update the weight coefficients in the comprehensive sorting index based on the outbound feedback effect; the prediction unit 202 is further used to dynamically adjust the weights of the training samples based on the outbound feedback data, so as to update the model in an incremental learning manner.

[0063] Furthermore, the embodiment of the present application also provides a non-performing asset outbound intelligent sorting device based on machine learning, including: a processor, a memory, and a system bus; the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes any of the above methods.

[0064] Furthermore, the embodiment of the present application also provides a computer program product, which, when running on a terminal device, enables the terminal device to execute any of the above methods.

[0065] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.

[0066] It should be noted that the various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0067] It should also be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0068] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A bad asset outbound intelligent sorting method based on machine learning, characterized in that, Including: Performing data cleaning and data standardization on a dataset containing non-performing asset outbound call records to extract information including time features, amount features, behavior features, outbound call features, and / or credit features; Performing feature selection on the extracted information and creating composite features through feature combination, including: Calculating the correlation coefficient matrix between features to remove highly correlated features; Calculating the variance between features to remove features with variance lower than the variance threshold; Using a random forest model to rank feature importance to select a target number of features; Performing polynomial combination on numerical features and / or cross combination on categorical features to construct composite features; Wherein, the method for obtaining the feature importance includes: , Among them, represents the importance of the feature . represents a node in the decision tree represents the child nodes on the node represents the sum of all nodes split on the feature . Sum up represents the node The number of samples in represents the total number of samples represents the node Entropy of represents The entropy of the child nodes on the node . represents the left and right child nodes represents The child nodes on the node The number of samples in; Training one or more types of models using the composite features and performing model fusion based on the Stacking ensemble method to obtain a fusion probability prediction; wherein the models include logistic regression, random forest, gradient boosting decision tree, and / or support vector machine; Wherein, performing model fusion based on the Stacking ensemble method to obtain a fusion probability prediction includes: Use the prediction results of the base models as new features, train a meta-model as the final classifier, and the calculation formula for the fused prediction probability is: , where is the prediction function of the meta-model, are the prediction probabilities of each base model respectively; Constructing a comprehensive ranking index based on the fusion probability prediction; Sorting the outbound call customers in batches according to the comprehensive ranking index and updating the weight coefficients in the comprehensive ranking index based on the outbound call feedback effect, including: Dividing the outbound call customers into multiple batches evenly based on the comprehensive ranking index; For each batch, selecting the outbound call customer with the highest comprehensive ranking index with a first probability and randomly selecting outbound call customers with a second probability; Dynamically adjusting the weights of the training samples based on the outbound call feedback data to update the model in an incremental learning manner.

2. The intelligent sorting method according to claim 1, wherein, Wherein, Performing data cleaning and data standardization on a dataset containing non-performing asset outbound call records, including: Removing duplicate data and unifying the data format; Filling numerical features with mean, median, or mode; Filling categorical features with mode or creating unknown categories; Adopt Principle or interquartile range method to identify outliers for deletion or adjustment according to the outlier threshold; Performing standardization processing on numerical features.

3. The intelligent sorting method according to claim 1, wherein Wherein, Constructing a comprehensive ranking index based on the fusion probability prediction, including: , Among them, represents the comprehensive sorting index, represents the weight coefficient, represents the prediction of the fusion probability of repayment, represents the overdue amount, represents the customer value score.

4. The intelligent sorting method according to claim 1, wherein Wherein, The method for determining the first probability includes: , Among them, represents the first probability, represents the initial probability, represents the attenuation coefficient, represents the current batch number; Setting a corresponding prior distribution for each weight coefficient in the comprehensive ranking index and updating the distribution parameters according to each outbound call feedback effect to adjust the weight coefficients.

5. The intelligent sorting method according to claim 4, wherein, Wherein, Updating the distribution parameters according to each outbound call feedback effect, including: , Among them, and represent the prior Beta distribution parameters of the -th weight coefficient, represents the success rate corresponding to the -th weight coefficient.

6. The intelligent sorting method according to claim 1, wherein Wherein, Dynamically adjusting the weights of the training samples based on the outbound call feedback data, including: , Among them, represents the adjusted sample weight, represents the sample weight before adjustment, represents the actual label, represents the predicted label, represents the learning rate.

7. The intelligent sorting method according to claim 1, characterized in that Wherein, The method further includes: Regularly re-evaluating feature importance and dynamically adjusting the feature set: Calculate the SHAP values for each feature, sort the features according to the SHAP values, retain the top k features with the highest importance, and eliminate or combine the features with low importance. The formula for calculating the SHAP value is as follows: , where N is the set of all features, S is the subset that does not include feature i, is the predicted value of the model on the feature subset S.

8. A bad asset outbound intelligent sorting system based on machine learning for implementing the intelligent sorting method according to any one of claims 1-7, characterized in that Including: A data preprocessing unit, a prediction unit, and a sorting unit; wherein, The data preprocessing unit is used to perform data cleaning and data standardization on a dataset containing non-performing asset outbound call records to extract information including time features, amount features, behavior features, outbound call features, and / or credit features; perform feature selection on the extracted information and create composite features through feature combination; The prediction unit is used to train one or more types of models using the composite features and perform model fusion based on the Stacking ensemble method to obtain a fusion probability prediction; The sorting unit is used to construct a comprehensive sorting index based on the fusion probability prediction; batch sort the outbound customers according to the comprehensive sorting index, and update the weight coefficient in the comprehensive sorting index based on the outbound feedback effect; The prediction unit is further used to dynamically adjust the weights of the training samples based on the outbound feedback data, so as to update the model in an incremental learning manner.

9. Intelligent sorting device for bad assets outbound calls based on machine learning, characterized in that, Including: A processor, a memory, and a system bus; wherein, the processor and the memory are connected through the system bus; The memory is used to store one or more programs, and the one or more programs include instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Repayment probability prediction model building method and device

    CN108256691A

  • Loan collection number screening and sorting method and system, terminal and storage medium

    CN114358917A