User arrearage prediction processing method and device and readable storage medium

By combining the time series model and the arrears prediction model of the XGBoost classification module, the problem of low accuracy of arrears warning in the existing technology is solved, and accurate prediction and timely processing of user arrears are achieved, and the efficiency of arrears is improved.

CN120020836APending Publication Date: 2025-05-20CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311552825.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The prior art is not very accurate when warning due to arrears, and cannot effectively adapt to the consumption situation of each user, resulting in invalid warning and poor user experience.

Method used

Using arrears prediction model combining time series model and XGBoost classification module, sample feature data is extracted from the original sample data through Relief-RFE combined feature selection algorithm, and multi-dimensional model training is carried out to achieve more accurate arrears risk assessment.

Benefits of technology

It realizes a more accurate prediction of the user's arrears, can obtain and process a large amount of user data in real time, timely predict the user's arrears and determine the corresponding processing strategies, reduce manual intervention, and improve the efficiency of arrears processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020836A_ABST
    Figure CN120020836A_ABST
Patent Text Reader

Abstract

The invention provides a user arrearage prediction processing method and device and a readable storage medium, and the method specifically comprises the steps: inputting the user information data of a to-be-predicted target user into an arrearage prediction model; determining a final prediction result according to a first prediction result output by the time sequence model module based on the user information data and a second prediction result output by the XGBoost classification module based on the user information data; and determining an arrearage processing strategy of the target user according to the final prediction result. According to the method, the sample feature data is extracted from the original sample data through the Relief-RFE combined feature selection algorithm, model training is carried out based on the multi-dimensional sample data, the arrearage prediction model composed of the time sequence model module and the XGBoost classification module is obtained, the arrearage prediction model can provide a more accurate arrearage risk assessment result, and the arrearage risk assessment efficiency is improved. Manual intervention can be reduced in the processing process, and arrearage processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a method, apparatus, and readable storage medium for predicting and processing user arrears. Background Art

[0002] With the rapid development of mobile communication technologies, mobile phones have become an indispensable tool in people's lives. However, the subsequent problem is the frequent occurrence of mobile phone arrears, which brings a series of troubles to both operators and users. Among them, user arrears can cause the interruption of the business system, which may cause significant losses to users, is not conducive to maintaining user relationships and recovering arrears. If the user manager conducts collection after the arrears occur, this method has poor timeliness and is likely to cause the arrears amount to accumulate too much.

[0003] In the prior art, operators set a fixed warning threshold for the subscribed services of users. When the remaining amount of a user is lower than this threshold, a work order is assigned to the user manager to notify the corresponding user to renew the subscription. However, due to the different consumption situations of each user, setting a fixed warning threshold has low accuracy and cannot effectively adapt to each user. For example, for a user with a low consumption level, even if the remaining amount is lower than the fixed threshold, the user will still not be in arrears after deducting the payable fees for the current month. Then, warning the user when the remaining amount is lower than the fixed threshold is an ineffective warning, which brings a bad experience to the user.

[0004] Therefore, how to predict user arrears in advance is of great significance for taking measures in advance and optimizing the revenue management of operators. Summary of the Invention

[0005] The technical problem to be solved by this application is to provide a method, apparatus, and readable storage medium for predicting and processing user arrears in view of the above deficiencies of the prior art to solve the problems existing in the prior art.

[0006] In a first aspect, this application provides a method for predicting and processing user arrears, the method

[0007] comprises:

[0008] S1. Input the user information data of the target user to be predicted into the arrears prediction model;

[0009] wherein, the arrears prediction model includes a time series model module and an XGBoost classification module, the arrears prediction model is trained based on sample feature data, and the sample feature data is extracted from the original sample data through a filter-wrapper feature recursive elimination Relief-RFE combined feature selection algorithm;

[0010] S2. Determine the final prediction result based on the first prediction result output by the time series model module based on the user information data and the second prediction result output by the XGBoost classification module based on the user information data;

[0011] S3. Determine the arrears handling strategy for the target user according to the final prediction result.

[0012] In some embodiments, the original sample data and the user information data include at least one of user basic attribute data, user expense data, and extended feature data;

[0013] Among them, the user basic attribute data includes at least one of user identity ID, user phone number, network access time, and main product package ID;

[0014] The user expense data includes at least one of the user's historical arrears amount, consumption amount, fixed monthly rent, and functional monthly rent;

[0015] The extended feature data includes at least one of account opening duration, average monthly consumption, average monthly arrears, and the number of user arrears.

[0016] In some embodiments, the sample feature data is extracted from the original sample data by the Relief-RFE combined feature selection algorithm, specifically including:

[0017] S01. Construct a first feature set in the first dimension according to the original sample data;

[0018] S02. Calculate the direct correlation between each feature in the first feature set and the label category to obtain the first weight of each feature;

[0019] S03. Extract the features with the first weight greater than the first preset threshold as the second feature set;

[0020] S04. Obtain the second weight of each feature in the second feature set through the logical regression algorithm model training, and delete the features with the second weight less than the second preset threshold to obtain the sample feature data in the second dimension.

[0021] In some embodiments, S01 includes:

[0022] Perform missing value processing and normalization processing on the original sample data to obtain the first feature set in the first dimension.

[0023] In some embodiments, S2 includes:

[0024] S21. Obtain the first prediction result output by the time series model module based on the user information data, the first weight value corresponding to the time series model module, the second prediction result output by the XGBoost classification module based on the user information data, and the second weight value corresponding to the XGBoost classification module;

[0025] S22. Perform weighted summation according to the first prediction result, the first weight value, the second prediction result, and the second weight value to obtain the final prediction result.

[0026] In some embodiments, S3 includes:

[0027] S31. Determine whether the target user belongs to a user at high risk of overdue payment according to the final prediction result;

[0028] S32. If so, determine the corresponding overdue payment handling method based on the user information data of the user at high risk of overdue payment;

[0029] Wherein, the overdue payment handling method includes at least one of quickly shutting down the user's service, adjusting the user's credit limit, adjusting the user's credit score with the operator, and adjusting the user's delayed service shutdown duration.

[0030] In some embodiments, S31 includes:

[0031] Extract overdue users according to the final prediction result;

[0032] According to the overdue payment information corresponding to the overdue users and a preset determination rule for users at high risk of overdue payment, determine whether the overdue users belong to users at high risk of overdue payment.

[0033] In a second aspect, the present application provides a user overdue payment prediction processing device, and the device includes:

[0034] A data input module configured to input user information data of a target user to be predicted into an overdue payment prediction model; wherein, the overdue payment prediction model includes a time series model module and an XGBoost classification module, and the overdue payment prediction model is trained based on sample feature data, and the sample feature data is extracted from original sample data by a filter - wrapper feature recursive elimination Relief - RFE combined feature selection algorithm;

[0035] A result determination module configured to determine a final prediction result according to the first prediction result output by the time series model module based on the user information data and the second prediction result output by the XGBoost classification module based on the user information data;

[0036] An overdue payment processing module, which is configured to determine an overdue payment processing strategy for the target user according to the final prediction result.

[0037] In a third aspect, the present application provides a user overdue payment prediction processing device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to implement the user overdue payment prediction processing method described in the first aspect above.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the user overdue payment prediction processing method described in the first aspect above.

[0039] The user overdue payment prediction processing method, device and readable storage medium provided by the present application. Specifically, user information data of a target user to be predicted is input into an overdue payment prediction model. The overdue payment prediction model includes a time series model module and an XGBoost classification module. The overdue payment prediction model is trained based on sample feature data, and the sample feature data is extracted from original sample data through a filter-wrapper feature recursive elimination Relief-RFE combined feature selection algorithm. According to the first prediction result output by the time series model module based on the user information data and the second prediction result output by the XGBoost classification module based on the user information data, a final prediction result is determined. According to the final prediction result, an overdue payment processing strategy for the target user is determined. The present application provides a user overdue payment prediction processing method. By using the Relief-RFE combined feature selection algorithm to extract sample feature data from original sample data, model training is performed based on multi-dimensional sample data, and an overdue payment prediction model composed of a time series model module and an XGBoost classification module is obtained. This overdue payment prediction model can provide a more accurate overdue payment risk assessment result, can obtain and process a large amount of user data in real time, predict the overdue payment situation of users in a timely manner, and determine corresponding processing strategies. The processing process can reduce manual intervention and improve the efficiency of overdue payment processing. By predicting the overdue payment situation of users in advance, the present application is of great significance for taking measures in advance and optimizing the revenue management of operators. Description of the Drawings

[0040] The drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0041] Figure 1 It is a flowchart of a user overdue payment prediction processing method provided by an embodiment of the present application;

[0042] Figure 2A flowchart of another method for predicting user arrears provided by an embodiment of the present application;

[0043] Figure 3 A schematic structural diagram of a device for predicting user arrears provided by an embodiment of the present application;

[0044] Figure 4 A schematic structural diagram of another device for predicting user arrears provided by an embodiment of the present application.

[0045] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be provided later. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0046] To enable those skilled in the art to better understand the technical solutions of the present application, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0047] It can be understood that the specific embodiments and drawings described herein are only used to explain the present application, rather than limiting the present application.

[0048] It can be understood that, without conflict, the various embodiments and features in the embodiments of the present application can be combined with each other.

[0049] It can be understood that, for the convenience of description, only the parts related to the present application are shown in the drawings of the present application, and the parts unrelated to the present application are not shown in the drawings.

[0050] It can be understood that each unit and module involved in the embodiments of the present application may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple units and modules may also be integrated into one entity structure.

[0051] It can be understood that the terms "first", "second", etc. in the embodiments of the present application are used to distinguish different objects, or to distinguish different processes for the same object, rather than to describe the specific order of the objects.

[0052] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present application may occur in an order different from that marked in the drawings.

[0053] It can be understood that in the flowcharts and block diagrams of the present application, the possible system architectures, functions, and operations of the systems, devices, equipment, and methods according to the embodiments of the present application are shown. Among them, each block in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart can be implemented by a hardware-based system for implementing the specified function, or can be implemented by a combination of hardware and computer instructions.

[0054] It can be understood that the units and modules involved in the embodiments of the present application can be implemented in software or in hardware. For example, the units and modules can be located in the processor.

[0055] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0056] The present application provides a method for predicting and processing user arrears. The working process of this method can be implemented by an electronic device, such as a computer, a handheld intelligent terminal, etc. For the convenience of explanation, in the embodiments of the present application, the method execution subject is a computer for elaboration.

[0057] Figure 1 It is a schematic diagram of the method for predicting and processing user arrears provided by the embodiments of the present application. Figure 2 It is another schematic diagram of the method for predicting and processing user arrears provided by the embodiments of the present application. As Figure 1 and Figure 2 shown, the present application provides a method for predicting and processing user arrears, and the method includes:

[0058] S1. Input the user information data of the target user to be predicted into the arrears prediction model;

[0059] In this step, first, obtain the trained arrears prediction model. When it is necessary to predict user arrears, input the user information data of the target user into the arrears prediction model. The user information data includes at least one of user basic attribute data, user expense data, and extended feature data;

[0060] Among them, the user basic attribute data includes at least one of user identity ID, user phone number, network access time, and main product package ID; the user expense data includes at least one of user historical arrears amount, consumption amount, fixed monthly rent, and functional monthly rent; the extended feature data includes at least one of account opening duration, average monthly consumption, average monthly arrears, and user arrears times.

[0061] Specifically, the basic attribute data of users comes from the operator's product center, credit center, and user center. First, collect data on four dimensions including user ID, user phone number, network access time, and main product package ID from the user center and product center. The database uses the MYSQL database for storage. Based on these four dimensions of data, query and collect the data on the number of collection times and overdue duration of users in the credit control center for the past 12 months. The collection types include text messages for insufficient available credit, exhausted available credit, shutdown warning, semi-shutdown for overdue arrears, full-shutdown for overdue arrears, user-differentiated reminder text messages, and monthly rent pre-shutdown reminder text messages. The calculation method for the number of collection times is the sum of the number of text messages of each type received by the user in the current month. The calculation of the overdue duration data is the time of the user's monthly overdue arrears, in hours. Second, based on the user data stored in the database as mentioned above, query and collect the user star rating (1-5 stars), payment type (post-payment / pre-payment), credit limit, and customer ID data.

[0062] In addition, the user fee data comes from the operator's accounting center, including the overdue amount, consumption amount, fixed monthly rent, and functional monthly rent for each month of the user's past 12 months.

[0063] In addition, based on the data collected above, perform feature expansion through mathematical operations to obtain extended feature data. Specifically as follows: account opening duration, calculated as the difference in days between the account opening time and the current time; average monthly consumption, calculated as the sum of consumption in 12 months divided by 12; average monthly overdue amount, calculated as the sum of overdue amounts in 12 months divided by 12; number of user overdue times, calculated as the sum of the number of times of having overdue arrears in 12 months.

[0064] In this step, the overdue prediction model includes a time series model module and an XGBoost classification module. The overdue prediction model is trained based on sample feature data, and the sample feature data is extracted from the original sample data through the Relief-RFE combined feature selection algorithm of filter-encapsulation feature recursive elimination.

[0065] Optionally, the original sample data includes at least one of the user's basic attribute data, user fee data, and extended feature data;

[0066] Among them, the user's basic attribute data includes at least one of user identity ID, user phone number, network access time, and main product package ID; the user fee data includes at least one of the user's historical overdue amount, consumption amount, fixed monthly rent, and functional monthly rent; the extended feature data includes at least one of account opening duration, average monthly consumption, average monthly overdue amount, and number of user overdue times.

[0067] Optionally, the sample feature data is extracted from the original sample data by using the Relief-RFE combined feature selection algorithm, including S01-S04, specifically:

[0068] S01. Construct the first feature set of the first dimension based on the original sample data;

[0069] Optionally, this step includes: performing missing value processing and normalization processing on the original sample data to obtain the first feature set of the first dimension.

[0070] Specifically, missing values ​​are processed for the original sample data, including: missing arrears data in the last month are filled with the average arrears amount of the user in the previous 11 months, and the arrears amount of other previous months is filled with 0. Missing consumption data in the last month are filled with the average consumption amount of the user in the previous 11 months, and the consumption amount of other previous months is filled with 0. Missing values ​​of user star rating are uniformly filled with 1 star, and missing values ​​of credit limit are uniformly filled with 0 credit.

[0071] In addition, due to the large differences in the amount of arrears, consumption limit, and credit limit among different users, it is necessary to normalize the sample data to eliminate the dimensional influence between the features, solve the comparability between different sample features, and ensure the accuracy of the reimbursement data.

[0072] Based on multiple experimental verifications, this application uses the Z-score normalization processing method to standardize the mean and standard deviation of the original data so that the processed feature data conforms to the standard normal distribution with a mean of 0 and a standard deviation of 1. The conversion function is as follows:

[0073]

[0074] Wherein, x is the sample data before normalization, μ is the mean of all sample data, σ is the standard deviation of all sample data, and x′ is the sample data after normalization.

[0075] After the processing, the ratio of the samples in arrears to those in non-arrears is 1:1, that is, the data samples used in this application are balanced data sets. In addition, useless features are deleted to reduce the impact of model noise. The final first dimension is 26 dimensions. The arrears data of the most recent month is used as the label to mark the arrears and non-arrears data respectively, with the arrears data marked as 1 and the non-arrears data marked as 0.

[0076] S02. Calculate the direct correlation between each feature in the first feature set and the label category to obtain the first weight of each feature;

[0077] S03. Extract the features with the first weight greater than the first preset threshold as the second feature set;

[0078] Feature selection processing can effectively reduce the redundancy between sample features, delete non-critical features, reduce model noise, and improve the accuracy of model prediction. Starting from the influence of each feature weight of the sample on the sample label, this application combines the idea of feature subset search, combines the filtering algorithm Relief and the wrapper feature recursive elimination algorithm RFE, and constructs a Relief-RFE combined feature selection algorithm, which takes into account both reducing sample noise and effectively associating sample category correlations to the greatest extent.

[0079] Specifically, taking the original 26-dimensional sample data as the input sample set, then calculating the correlation between the features and the label category, calculating the average weight of each feature, that is, the first weight, and setting a first preset threshold in the calculation result, and returning the feature set T greater than this first preset threshold, that is, the second feature set;

[0080] S04. Train the second weight of each feature in the second feature set through a logistic regression algorithm model, and delete the features with the second weight less than the second preset threshold to obtain the sample feature data of the second dimension.

[0081] Specifically, taking the feature set T as the input, constructing a logistic regression algorithm model, using the above feature set as the input, training to obtain the second weight of each feature, and deleting the features with smaller weights, and repeating this until the optimal feature set is obtained, that is, all features are greater than the initially set second preset threshold. After feature selection processing, the finally selected sample set has a feature dimension of 21 dimensions, and this sample set is used as the sample feature data of the final prediction model.

[0082] In this application, after obtaining the sample feature data, a delinquent payment prediction model is constructed. The model construction is divided into two parts, one is a time series model module, and the other is an XGBoost classification module. Finally, the weights of the two models are combined.

[0083] First is the time series model module. This application uses the LSTM (Long Short-Term Memory) time series algorithm, which can also be called the LSTM neural network. The neural network includes an input gate, a forget gate, an output gate, and a memory unit. Each gate contains a sigmoid layer, and the function of this layer is to determine which information of the sample features should be retained, and the numerical output is a number between 0 and 1. In the calculation of each layer, if the calculation result is 0, it means that no information will be transmitted to the next gate unit, and when it is 1, it means that all information will be transmitted.

[0084] For the sample feature data obtained through step S04, five types of features, namely overdue data, consumption data, average overdue, average consumption, and number of overdue times, are taken as the input sample set of this sequence model. The labels are still the same as those marked in the original sample set, divided into two categories: overdue and not overdue. The time span of the entire overdue data and consumption data is 12 months (experiments have shown that the larger the time span, the better the final prediction effect). The output of the input gate will reach the forget gate unit, and through the calculation of sigmoid, the input information is calculated and transmitted to the memory unit. After each calculation of the input gate and forget gate, the memory unit will update its state. Finally, the output gate of the LSTM will determine the output of the entire sample information. Its output calculation is to multiply the value between 0 and 1 output by the sigmoid function with the tanh function, ensuring that the finally output value can be used as a value for classification judgment.

[0085] Secondly, it is the XGBoost classification module. In this application, the XGBoost package of sklearn is used for this module. For parameter selection, the depth of the XGBoost tree is set to 3, the learning rate is set to 0.1, the gbtree tree model is used as the base classifier, the number of iterative training times is set to 100, the ratio of the training set to the test set is set to 7:3, and the feature dimension of the sample set is all the remaining features of the entire feature set except for the time-related features used in the above time series module. The output of this module is also a binary classification result.

[0086] After both types of models are fully trained, it is necessary to perform a weighted comprehensive calculation. The calculation method is a linear function calculation, and the formula is as follows:

[0087] y = w1 × y1 + w2 × y2

[0088] Among them, y represents the final classification result, y1 represents the output classification result of the time series model LSTM, y2 represents the classification result of the XGBoost model, and w1 and w2 are weights. Preferably, after multiple experiments and effect evaluations, the weight parameter w1 is determined to be 0.6 and w2 is 0.4.

[0089] S2. Determine the final prediction result according to the first prediction result output by the time series model module based on the user information data and the second prediction result output by the XGBoost classification module based on the user information data;

[0090] This step includes S21 and S22. Specifically:

[0091] S21. Obtain the first prediction result output by the time series model module based on the user information data, the first weight value corresponding to the time series model module, the second prediction result output by the XGBoost classification module based on the user information data, and the second weight value corresponding to the XGBoost classification module;

[0092] S22. Perform weighted summation according to the first prediction result, the first weight value, the second prediction result, and the second weight value to obtain the final prediction result.

[0093] Specifically, according to the model training process, the final prediction result is obtained through the following formula:

[0094] y = w1 × y1 + w2 × y2

[0095] where y represents the final prediction result, y1 represents the first prediction result output by the time series model module based on the user information data, y2 represents the second prediction result output by the XGBoost classification module based on the user information data, w1 is the first weight value corresponding to the time series model module, w2 is the second weight value corresponding to the XGBoost classification module. Preferably, after multiple experiments and effect evaluations, the weight parameter w1 is determined to be 0.6 and w2 is determined to be 0.4.

[0096] S3. Determine the overdue payment handling strategy for the target user according to the final prediction result.

[0097] This step includes S31 and S32. Specifically:

[0098] S31. Determine whether the target user belongs to a high - risk overdue payment user according to the final prediction result;

[0099] Optionally, extract overdue payment users according to the final prediction result; according to the overdue payment information corresponding to the overdue payment users and the preset high - risk overdue payment user determination rule, determine whether the overdue payment users belong to high - risk overdue payment users.

[0100] Specifically, according to the overdue payment prediction situation of the users output by the model, further statistical analysis and strategy adjustment are carried out for the overdue payment users. Among them, the preset high - risk overdue payment user determination rule can be, for example:

[0101] 1. Statistically calculate the average monthly billing cost of the overdue payment users in the historical months, and calculate whether the current billing month cost is more than 2 times the average cost of previous months.

[0102] 2. Whether the item of the current month's billing cost includes an over - package item.

[0103] If an overdue user meets the above two conditions, the user is identified as a user with a high overdue risk.

[0104] It can be understood that the preset rules for determining high-overdue-risk users include, but are not limited to, the above two conditions, and can be adjusted according to the actual situation.

[0105] S32. If so, determine the corresponding overdue handling method based on the user information data of the high-overdue-risk user; wherein, the overdue handling method includes at least one of quickly suspending the user's service, adjusting the user's credit limit, adjusting the user's credit score with the operator, and adjusting the user's delayed service suspension duration.

[0106] Specifically, the corresponding overdue handling method can be selected according to the user's information data. For example, for users with a higher user star level and a longer network access time, the user's delayed service suspension duration can be appropriately increased; for users with a lower user star level and a shorter network access time, quick service suspension can be performed.

[0107] This application provides a user overdue prediction and processing method. Sample feature data is extracted from the original sample data through the Relief-RFE combined feature selection algorithm, and model training is performed based on multi-dimensional sample data to obtain an overdue prediction model composed of a time series model module and an XGBoost classification module. This overdue prediction model can provide a more accurate overdue risk assessment result, can obtain and process a large amount of user data in real time, predict the user's overdue situation in a timely manner, and determine corresponding processing strategies. The processing process can reduce manual intervention and improve the efficiency of overdue handling. By predicting the user's overdue situation in advance, this application is of great significance for taking measures in advance and optimizing the revenue management of operators.

[0108] It should be understood that although the steps in the flowcharts in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least some of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0109] Figure 3 For the schematic diagram of the user overdue prediction and processing device provided by the embodiments of this application, as Figure 3 shown, this application provides a user overdue prediction and processing device, and the device includes:

[0110] A data input module 11, which is configured to input user information data of a target user to be predicted into an overdue payment prediction model; wherein, the overdue payment prediction model includes a time series model module and an XGBoost classification module, the overdue payment prediction model is trained based on sample feature data, and the sample feature data is extracted from original sample data through a filtering - wrapper feature recursive elimination Relief - RFE combined feature selection algorithm;

[0111] A result determination module 12, which is configured to determine a final prediction result according to a first prediction result output by the time series model module based on the user information data and a second prediction result output by the XGBoost classification module based on the user information data;

[0112] An overdue payment processing module 13, which is configured to determine an overdue payment processing strategy for the target user according to the final prediction result.

[0113] Regarding the definition of the user overdue payment prediction processing device, reference can be made to the definition of the user overdue payment prediction processing method in the above - mentioned embodiments of the present application, and details are not described herein again.

[0114] Figure 4 This is another schematic diagram of the user overdue payment prediction processing device provided by the embodiments of the present application. As Figure 4 shown, in some embodiments, the present application provides a user overdue payment prediction processing device, including a memory 22 and a processor 21. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the user overdue payment prediction processing method in the above - mentioned embodiments of the present application.

[0115] Among them, the memory is connected to the processor. The memory can adopt flash memory, read - only memory or other memories, and the processor can adopt a central processing unit or a single - chip microcomputer.

[0116] In some embodiments, the present application provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the user overdue payment prediction processing method in the above - mentioned embodiments of the present application.

[0117] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, computer program modules, or other data. The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile discs (DVDs) or other optical disc storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0118] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present application. However, the present application is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present application, and these modifications and improvements are also regarded as the protection scope of the present application.

Claims

1. A method for predicting and processing user arrears, characterized in that: The method comprises: S1. Input the user information data of the target user to be predicted into the arrears prediction model; The arrears prediction model includes a time series model module and an XGBoost classification module. The arrears prediction model is trained based on sample feature data, and the sample feature data is extracted from the original sample data by a filtering-encapsulation feature recursive elimination Relief-RFE combined feature selection algorithm; S2. Determine a final prediction result according to the first prediction result output by the time series model module based on the user information data and the second prediction result output by the XGBoost classification module based on the user information data; S3. Determine the arrears handling strategy for the target user based on the final prediction result.

2. The method for predicting and processing user arrears according to claim 1, characterized in that: The original sample data and the user information data include at least one of user basic attribute data, user fee data and extended feature data; The basic attribute data of the user includes at least one of the user ID, user phone number, network access time and main product package ID; The user fee data includes at least one of the user's historical arrears amount, consumption amount, fixed monthly fee, and function monthly fee; The extended characteristic data includes at least one of the following: account opening duration, average monthly consumption, average monthly arrears, and number of arrears times of the user.

3. The method for predicting and processing user arrears according to claim 1, characterized in that: The sample feature data is extracted from the original sample data by using the Relief-RFE combined feature selection algorithm, specifically including: S01. Construct a first feature set of the first dimension based on the original sample data; S02. Calculate the direct correlation between each feature in the first feature set and the label category to obtain a first weight for each feature; S03. Extracting features with a first weight greater than a first preset threshold as a second feature set; S04. Obtain the second weight of each feature in the second feature set through logistic regression algorithm model training, and delete the features whose second weight is less than a second preset threshold to obtain the sample feature data of the second dimension.

4. The method for predicting and processing user arrears according to claim 3, characterized in that: S01, including: The original sample data is processed for missing values ​​and normalized to obtain a first feature set of the first dimension.

5. The method for predicting and processing user arrears according to claim 1, characterized in that: S2, including: S21. Obtain the first prediction result output by the time series model module based on the user information data, the first weight value corresponding to the time series model module, and the second prediction result output by the XGBoost classification module based on the user information data, and the second weight value corresponding to the XGBoost classification module; S22. Perform weighted summation according to the first prediction result, the first weight value, the second prediction result and the second weight value to obtain the final prediction result.

6. The method for predicting and processing user arrears according to claim 1, characterized in that: S3, including: S31. Determine whether the target user is a high-amount arrears risk user according to the final prediction result; S32. If so, the corresponding arrears processing method is determined based on the user information data of the user with high arrears risk; The arrears processing method includes at least one of quickly shutting down the user, adjusting the user's credit limit, adjusting the user's credit points with the operator, and adjusting the user's delayed shutdown duration.

7. The method for predicting and processing user arrears according to claim 6, characterized in that: S31, including: Extracting users in arrears according to the final prediction result; According to the arrears information corresponding to the arrears user and the preset high-amount arrears risk user determination rule, it is determined whether the arrears user is a high-amount arrears risk user.

8. A user arrears prediction and processing device, characterized in that: The device comprises: A data input module, which is configured to input user information data of the target user to be predicted into the arrears prediction model; wherein the arrears prediction model includes a time series model module and an XGBoost classification module, and the arrears prediction model is trained based on sample feature data, and the sample feature data is extracted from the original sample data by a filtering-encapsulation feature recursive elimination Relief-RFE combined feature selection algorithm; A result determination module, which is configured to determine a final prediction result based on a first prediction result output by the time series model module based on the user information data and a second prediction result output by the XGBoost classification module based on the user information data; The arrears processing module is configured to determine an arrears processing strategy for the target user based on the final prediction result.

9. A user arrears prediction and processing device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the user arrears prediction and processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for predicting and processing user arrears according to any one of claims 1 to 7 is implemented.