User data processing method and system

By cleaning, transforming and standardizing user behavior data, and generating churn indexes with calibrated communication indexes, the problems of low churn prediction accuracy and coverage in the existing technology are solved, and higher churn prediction accuracy and coverage are achieved.

CN119762142BActive Publication Date: 2025-05-13CHINA UNICOM (JIANGXI) IND INTERNET CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510259418.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-13
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The prior art makes churn prediction by analyzing the changes in single behavior data of users, with low accuracy and low coverage for users to be churn.

Method used

By obtaining the user's unique identification data, extracting user behavior data from the report library, performing data cleaning and transformation, standardizing user behavior data, combining calibration of communication indexes, generating churn indexes, and determining the user's churn risk level.

Benefits of technology

The accuracy and coverage of churn prediction are improved. Through data streamlining and multi-dimensional analysis, the impact of Kangyu data or low-correlation data on prediction is reduced, and the ability to predict user churn risk is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762142B_ABST
    Figure CN119762142B_ABST
Patent Text Reader

Abstract

The present invention provides a user data processing method and system, the method comprising: obtaining unique identification data to extract a number of user behavior data, performing data cleaning on the user behavior data to form first standby behavior data; performing data transformation on the first standby behavior data, and then selecting a number of final behavior data; obtaining key communication data corresponding to the unique identification data, obtaining a single communication index between the two, and generating a calibration communication index based on a number of single communication indexes; obtaining a first behavior index and a second behavior index, generating a churn index based on the calibrated communication index, the first behavior index, and the second behavior index, and determining the churn risk level of non-churned users based on the churn index. By standardizing different user behavior data to the same dimension, and then judging whether a user has a churn risk by considering a variety of user behavior data, the accuracy of churn prediction is improved, and the corresponding prediction coverage is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data prediction, and in particular to a user data processing method and system. Background Art

[0002] As the population grows, the number of network accesses has skyrocketed, greatly expanding the market size of operators. Operators have entered an era of stock operation.

[0003] Stock operation refers to a series of business policies and strategies that operators use for their existing customers through refined management and differentiated services to improve customer loyalty and release customer value. When the personalized needs of users do not match the services provided by operators, it is easy to cause user loss. Therefore, if the user's loss intention can be predicted in advance, the user's intention can be pulled back in advance.

[0004] Currently, user churn prediction is generally achieved by subjectively selecting a single user's behavior data and analyzing the behavior data. For example, if a user's traffic usage drops suddenly or the call duration drops suddenly, it is determined that the user is about to churn. However, this prediction method has low accuracy and low coverage of users who are about to churn. Summary of the invention

[0005] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a user data processing method and system, aiming to solve the technical problems in the prior art of predicting churn by analyzing changes in a single user's behavioral data, which has low accuracy and low coverage of users who are about to churn.

[0006] In order to achieve the above objectives, in a first aspect, an embodiment of the present application provides a user data processing method, comprising the following steps:

[0007] Acquire unique identification data of the user, the attribute types of the user include non-churned users and churned users, extract a number of user behavior data from a report library based on the unique identification data, the user behavior data include sub-behavior data for a number of consecutive months, perform data cleaning on the user behavior data to form first standby behavior data, the first standby behavior data include first standby sub-data for a number of consecutive months;

[0008] Performing data transformation on the first standby behavior data to obtain second standby behavior data, the second standby behavior data including second standby sub-data for a number of consecutive months, and determining the relevance of the second standby behavior data to a user based on the second standby sub-data to select a number of final behavior data from the number of second standby behavior data, the final behavior data including final sub-data for a number of consecutive months;

[0009] Acquire a number of key interaction data corresponding to the unique identification data, acquire a single interaction index between the unique identification data and the key interaction data, and generate a calibrated interaction index corresponding to the unique identification data based on the number of single interaction indexes;

[0010] The first behavior index and the second behavior index corresponding to the final behavior data are obtained through the final sub-data, and a churn index corresponding to the unique identification data is generated based on the calibrated interaction index, the first behavior index and the second behavior index. The churn risk level of non-churned users is determined based on the churn index.

[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: by extracting a number of the user behavior data from the report library, after performing the data cleaning and the data changes, the different user behavior data are standardized to the same dimension, and then by considering a plurality of the user behavior data to determine whether the user has a churn risk, the accuracy of churn prediction is improved, and the corresponding prediction coverage is improved; by selecting the final behavior data, that is, by eliminating part of the second standby behavior data through the correlation between the second standby behavior data and the user, the data is streamlined, and the influence of redundant data or low-correlation data on the accuracy of churn prediction is avoided; on the basis of the user behavior data, the calibrated interaction index is introduced, and the prediction of churn risk is assisted through the stability of the interaction circle formed between the unique identification data and the key interaction data, thereby further improving the accuracy and coverage of churn prediction.

[0012] Furthermore, the step of cleaning the user behavior data includes:

[0013] Obtaining a maximum number of the sub-behavior data in the plurality of user behavior data, and determining whether there are missing months in the user behavior data based on the maximum number;

[0014] If there are missing months in the user behavior data, obtain the missing number of the missing months, and compare the missing number with the missing threshold;

[0015] If the missing number is less than the missing threshold, the sub-behavior data corresponding to the two months adjacent to the missing month are obtained to generate filling behavior data, and the filling behavior data is filled into the missing month;

[0016] If the missing number is greater than the missing threshold, the preset behavior data is added to fill in the missing month;

[0017] A threshold range is constructed based on all of the sub-behavior data in the user behavior data, and abnormal data is replaced on the user behavior data based on the threshold range.

[0018] Furthermore, the step of performing data transformation on the first standby behavior data to obtain second standby behavior data, wherein the second standby behavior data includes second standby sub-data of a plurality of consecutive months, comprises:

[0019] Acquire the maximum value and the minimum value of the first standby sub-data in the first standby behavior data;

[0020] converting each of the first unused sub-data into second unused sub-data based on the maximum value and the minimum value;

[0021] The second standby sub-data of several consecutive months are combined into the second standby behavior data.

[0022] Furthermore, the calculation formula of the second standby sub-data is:

[0023] ,

[0024] in, represents the second standby sub-data of the i-th month in the second standby behavior data corresponding to the j-th first standby behavior data, represents the first standby sub-data of the i-th month in the j-th first standby behavior data, represents the minimum value of the first standby sub-data in the jth first standby behavior data, Represents the maximum value of the first standby sub-data in the jth first standby behavior data.

[0025] Furthermore, the step of determining the relevance between the second standby behavior data and the user based on the second standby sub-data to select a plurality of final behavior data from a plurality of the second standby behavior data includes:

[0026] generating a behavior change feature based on the adjacent second standby sub-data, generating a change speed feature based on the adjacent behavior change feature, and generating a change trend feature based on the adjacent change speed feature;

[0027] Fitting a plurality of the change trend features into a trend line, obtaining the slope of the trend line, and comparing the slope with the attribute type;

[0028] If the slope is a positive value and the attribute type is a non-churned user, determining that the second standby behavior data corresponding to the slope is the final behavior data;

[0029] If the slope is a negative value and the attribute type is lost users, the second standby behavior data corresponding to the slope is determined to be final behavior data.

[0030] Furthermore, the step of obtaining a plurality of key communication data corresponding to the unique identification data includes:

[0031] Acquire a plurality of interaction identification data associated with the unique identification data, and acquire the cumulative number of interactions between the unique identification data and the interaction identification data;

[0032] A plurality of key interaction data are selected from a plurality of interaction identification data based on the accumulated number of interactions.

[0033] Furthermore, the step of obtaining a single interaction index between the unique identification data and the key interaction data includes:

[0034] Within a preset time period, obtaining a daily interaction value, a three-day interaction value, a weekly interaction value, a ten-day interaction value, and a monthly interaction value between the unique identification data and the key interaction data;

[0035] The single interaction index is generated based on the daily interaction value, the three-day interaction value, the weekly interaction value, the ten-day interaction value, and the monthly interaction value.

[0036] Furthermore, the calculation formula of the daily exchange value is:

[0037] ,

[0038] in, Represents the daily interaction value, Indicates the number of interactions between the unique identification data and the key interaction data based on days. Indicates the total number of days in the preset time period;

[0039] The calculation formula of the three-day interaction value is:

[0040] ,

[0041] in, Indicates the three-day interaction value. The number of interactions between the unique identification data and the key interaction data based on a 3-day period;

[0042] The calculation formula of the single interaction index is:

[0043] ,

[0044] in, Represents a single interaction index between the unique identification data and the ath key interaction data, Represents the weekly exchange value, Indicates the ten-day exchange value, Indicates the monthly transaction value.

[0045] Furthermore, the step of determining the churn risk level of non-churned users based on the churn index includes:

[0046] Setting a number of preset risk levels and index ranges corresponding to the preset risk levels;

[0047] The churn index is matched with a plurality of the index ranges to select a churn risk level corresponding to non-churned users from among the plurality of the preset risk levels.

[0048] In a second aspect, an embodiment of the present application provides a user data processing system, which is applied to the user data processing method as described in the first aspect above, and the system includes:

[0049] A processing module is used to obtain unique identification data of a user, wherein the attribute types of the user include non-lost users and lost users, extract a plurality of user behavior data from a report library based on the unique identification data, wherein the user behavior data includes sub-behavior data for a plurality of consecutive months, and perform data cleaning on the user behavior data to form first standby behavior data, wherein the first standby behavior data includes first standby sub-data for a plurality of consecutive months;

[0050] a screening module, configured to perform data transformation on the first standby behavior data to obtain second standby behavior data, wherein the second standby behavior data includes second standby sub-data for a number of consecutive months, and to determine the relevance of the second standby behavior data to the user based on the second standby sub-data, so as to select a number of final behavior data from the number of second standby behavior data, wherein the final behavior data includes final sub-data for a number of consecutive months;

[0051] A first analysis module is used to obtain a number of key communication data corresponding to the unique identification data, obtain a single communication index between the unique identification data and the key communication data, and generate a calibrated communication index corresponding to the unique identification data based on the single communication indexes;

[0052] The second analysis module is used to obtain the first behavior index and the second behavior index corresponding to the final behavior data through the final sub-data, generate a churn index corresponding to the unique identification data based on the calibrated interaction index, the first behavior index and the second behavior index, and determine the churn risk level of non-churned users based on the churn index.

[0053] In a third aspect, an embodiment of the present application provides a computer, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the user data processing method as described in the first aspect above is implemented.

[0054] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the user data processing method described in the first aspect above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a flowchart of a method for processing user data in a first embodiment of the present invention;

[0056] Figure 2 is a structural block diagram of a user data processing system in a second embodiment of the present invention;

[0057] The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0058] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0059] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0061] See also Figure 1 The user data processing method provided by the first embodiment of the present invention comprises the following steps:

[0062] S10: Acquire unique identification data of the user, the attribute types of the user include non-lost users and lost users, extract a number of user behavior data from a report library based on the unique identification data, the user behavior data includes sub-behavior data of a number of consecutive months, and perform data cleaning on the user behavior data to form first standby behavior data, the first standby behavior data includes first standby sub-data of a number of consecutive months;

[0063] In this embodiment, the unique identification data is the user number. There are several sub-report libraries in the report library, and each sub-report library has the corresponding user behavior data. In this embodiment, the user behavior data includes monthly traffic usage, monthly call duration, monthly call times, monthly SMS sending volume and monthly usage fees. Due to system errors and other reasons, after extracting the user behavior data, the sub-behavior data may be abnormal. Therefore, in order to ensure the accuracy of the subsequent loss risk level, the user behavior data needs to be processed accordingly.

[0064] The step S10 comprises:

[0065] S110: Obtaining a maximum number of the sub-behavior data in the plurality of user behavior data, and determining whether there is a missing month in the user behavior data based on the maximum number;

[0066] It can be understood that the number of the sub-behavior data is the same as the number of months, and each sub-behavior data corresponds to the month, such as a user used 1000M of traffic in January.

[0067] S120: If there are missing months in the user behavior data, obtain the missing number of the missing months, and compare the missing number with a missing threshold;

[0068] The missing threshold value may be changed according to the number of the sub-behavior data in the user behavior data. In this embodiment, the missing threshold value is 30% of the number of the sub-behavior data in the user behavior data.

[0069] S130: If the missing number is less than the missing threshold, obtaining the sub-behavior data corresponding to two months adjacent to the missing month to generate filling behavior data, and filling the missing month with the filling behavior data;

[0070] S140: If the missing number is greater than the missing threshold, fill the missing month with preset behavior data;

[0071] In this embodiment, the preset behavior data is 0.

[0072] S150: constructing a threshold range based on all the sub-behavior data in the user behavior data, and replacing abnormal data in the user behavior data based on the threshold range;

[0073] Specifically, a plurality of the sub-behavior data are sorted to select an upper quartile and a lower quartile from the plurality of the sub-behavior data; a standard value is determined based on the upper quartile and the lower quartile, the sum of the upper quartile and the standard value is used as the high threshold, and the difference between the lower quartile and the standard value is used as the low threshold; the threshold range is formed by the low threshold and the high threshold. The sub-behavior data is compared with the threshold range. If the sub-behavior data is not within the threshold range, the sub-behavior data is determined as abnormal data and replaced. The replacement of the abnormal data can be carried out by replacing it with the average of two adjacent sub-behavior data, or by replacing it with the data of the same month in different years after averaging.

[0074] S20: performing data transformation on the first standby behavior data to obtain second standby behavior data, wherein the second standby behavior data includes second standby sub-data for a number of consecutive months, and determining the relevance of the second standby behavior data to the user based on the second standby sub-data, so as to select a number of final behavior data from the number of second standby behavior data, wherein the final behavior data includes final sub-data for a number of consecutive months;

[0075] The step S20 comprises:

[0076] S210: Acquire the maximum value and the minimum value of the first standby sub-data in the first standby behavior data;

[0077] S220: Convert each of the first unused sub-data into second unused sub-data based on the maximum value and the minimum value;

[0078] The calculation formula of the second standby sub-data is:

[0079] ,

[0080] in, represents the second standby sub-data of the i-th month in the second standby behavior data corresponding to the j-th first standby behavior data, represents the first standby sub-data of the i-th month in the j-th first standby behavior data, represents the minimum value of the first standby sub-data in the jth first standby behavior data, Represents the maximum value of the first standby sub-data in the jth first standby behavior data.

[0081] S230: combining the second standby sub-data of a plurality of consecutive months into the second standby behavior data;

[0082] Since different first standby behavior data have different data dimensions, it is difficult to achieve unified integration of data if the first standby behavior data are directly used for subsequent analysis and prediction. Therefore, it is necessary to convert them into the second standby behavior data.

[0083] S240: generating a behavior change feature based on the adjacent second stand-by sub-data, generating a change speed feature based on the adjacent behavior change feature, and generating a change trend feature based on the adjacent change speed feature;

[0084] The formula for obtaining the behavior change characteristics is:

[0085] ,

[0086] in, represents the behavior change characteristics between the second standby sub-data of the i-th month and the second standby sub-data of the i-1-th month, represents the second standby sub-data of the i-th month in the second standby behavior data corresponding to the j-th first standby behavior data, represents the second standby sub-data of the i-1th month in the second standby behavior data corresponding to the jth first standby behavior data, The calculation method of the change speed feature and the change trend feature is consistent with the calculation method of the behavior change feature, and will not be repeated here.

[0087] S250: fitting a plurality of the change trend features into a trend line, obtaining the slope of the trend line, and comparing the slope with the attribute type;

[0088] S260: If the slope is a positive value and the attribute type is a non-churned user, determining that the second standby behavior data corresponding to the slope is final behavior data;

[0089] S270: If the slope is a negative value and the attribute type is lost users, determining that the second standby behavior data corresponding to the slope is final behavior data;

[0090] It can be understood that if the slope is a positive value and the attribute type is a lost user, or the slope is a negative value and the attribute type is a non-lost user, it is determined that the second standby behavior data corresponding to the slope has a poor correlation with the user, and the second standby behavior data corresponding to the slope is eliminated.

[0091] S30: Acquire a plurality of key communication data corresponding to the unique identification data, acquire a single communication index between the unique identification data and the key communication data, and generate a calibrated communication index corresponding to the unique identification data based on the plurality of single communication indexes;

[0092] The step S30 comprises:

[0093] S310: Acquire a plurality of communication identification data associated with the unique identification data, and acquire the cumulative number of interactions between the unique identification data and the communication identification data;

[0094] S320: Selecting a plurality of key interaction data from a plurality of interaction identification data based on the accumulated number of interactions;

[0095] In this embodiment, the communication identification data and the key communication data are both communication numbers. The accuracy of data analysis is improved by screening and eliminating low-frequency communication numbers.

[0096] S330: obtaining a daily interaction value, a three-day interaction value, a weekly interaction value, a ten-day interaction value, and a monthly interaction value between the unique identification data and the key interaction data within a preset time period;

[0097] In this embodiment, the preset time period is 3 months.

[0098] The calculation formula of the daily exchange value is:

[0099] ,

[0100] in, Represents the daily interaction value, Indicates the number of interactions between the unique identification data and the key interaction data based on days. Indicates the total number of days in the preset time period. It should be noted that, taking 3 months as an example, the total number of days is 90. If there is an interaction on a certain day, only one interaction will be recorded.

[0101] The calculation formula of the three-day interaction value is:

[0102] ,

[0103] in, Indicates the three-day interaction value. When 3 days is used as the standard, the number of interactions between the unique identification data and the key interaction data; it can be understood that, taking 3 months as an example, a total of 90 days, if there is interaction within 3 days, then 1 interaction is recorded. The acquisition method of the weekly interaction value, the ten-day interaction value and the monthly interaction value is similar, and will not be repeated here.

[0104] S340: generating the single interaction index based on the daily interaction value, the three-day interaction value, the weekly interaction value, the ten-day interaction value, and the monthly interaction value;

[0105] The calculation formula of the single interaction index is:

[0106] ,

[0107] in, Represents a single interaction index between the unique identification data and the ath key interaction data, Represents the weekly exchange value, Indicates the ten-day exchange value, Indicates the monthly transaction value.

[0108] After obtaining the plurality of single interaction indexes, the plurality of single interaction indexes are averaged to obtain the calibrated interaction index.

[0109] S40: obtaining a first behavior index and a second behavior index corresponding to the final behavior data through the final sub-data, generating a churn index corresponding to the unique identification data based on the calibrated interaction index, the first behavior index and the second behavior index, and determining a churn risk level of non-churned users based on the churn index;

[0110] The formula for obtaining the first behavior index is:

[0111] ,

[0112] in, represents the first behavior index of the mth final behavior data, Indicates the nth final sub-data in the mth final behavior data, Indicates the number of final sub-data in the mth final behavior data;

[0113] The formula for obtaining the second behavior index is:

[0114] ,

[0115] in, represents the second behavior index of the mth final behavior data, Represents the final sub-data of the last month in the mth final behavioral data.

[0116] The formula for obtaining the churn index is:

[0117] ,

[0118] in, represents the churn index corresponding to user b, represents the number of final behavior data corresponding to user b, , , All represent weights, and the sum of the three is 1. Represents the calibrated interaction index corresponding to user b.

[0119] The step S40 comprises:

[0120] S410: Setting a number of preset risk levels and index ranges corresponding to the preset risk levels;

[0121] In this embodiment, four preset risk levels are set, namely low risk, medium risk, medium-high risk and high risk. The index range for high risk is: 0~0.25, the index range for medium-high risk is: 0.25~0.5, the index range for medium risk is: 0.5~0.75, and the index range for low risk is: 0.75~1.

[0122] S420: Matching the churn index with a plurality of the index ranges to select a churn risk level corresponding to non-churned users from a plurality of the preset risk levels.

[0123] It can be understood that after determining the churn risk level, the prediction of user churn is completed, and then the corresponding user retention strategy can be determined according to the churn risk level to achieve the purpose of stock retention.

[0124] By extracting a number of the user behavior data from the report library, and after performing the data cleaning and data changes, the different user behavior data are standardized to the same dimension, and then by considering a plurality of the user behavior data to determine whether the user has a churn risk, the accuracy of churn prediction is improved, and the corresponding prediction coverage is improved; by selecting the final behavior data, that is, by eliminating part of the second standby behavior data through the correlation between the second standby behavior data and the user, the data is streamlined, and the influence of redundant data or low-correlation data on the accuracy of churn prediction is avoided; on the basis of the user behavior data, the calibrated interaction index is introduced, and the prediction of churn risk is assisted through the stability of the interaction circle formed between the unique identification data and the key interaction data, thereby further improving the accuracy and coverage of churn prediction.

[0125] See also Figure 2The second embodiment of the present invention provides a user data processing system, which is applied to the user data processing method described in the above embodiment, and will not be repeated here. As used below, the terms "module", "unit", "subunit" and the like can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0126] The system comprises:

[0127] The processing module 10 is used to obtain the unique identification data of the user, the attribute types of the user include non-lost users and lost users, extract a number of user behavior data from the report library based on the unique identification data, the user behavior data includes sub-behavior data of a number of consecutive months, and perform data cleaning on the user behavior data to form first standby behavior data, the first standby behavior data includes first standby sub-data of a number of consecutive months;

[0128] The processing module 10 comprises:

[0129] The first unit is used to obtain the maximum number of the sub-behavior data in the plurality of user behavior data, and determine whether there is a missing month in the user behavior data based on the maximum number;

[0130] The second unit is configured to obtain the missing number of the missing month if there is a missing month in the user behavior data, and compare the missing number with a missing threshold;

[0131] The third unit is configured to obtain the sub-behavior data corresponding to two months adjacent to the missing month to generate filling behavior data if the missing number is less than the missing threshold, and fill the filling behavior data into the missing month;

[0132] The fourth unit is used to fill the missing month with preset behavior data if the missing number is greater than the missing threshold;

[0133] A fifth unit, configured to construct a threshold range based on all the sub-behavior data in the user behavior data, and replace abnormal data of the user behavior data based on the threshold range;

[0134] A screening module 20 is used to perform data transformation on the first standby behavior data to obtain second standby behavior data, wherein the second standby behavior data includes second standby sub-data for a number of consecutive months, and to determine the relevance of the second standby behavior data to the user based on the second standby sub-data, so as to select a number of final behavior data from the number of second standby behavior data, wherein the final behavior data includes final sub-data for a number of consecutive months;

[0135] The screening module 20 comprises:

[0136] A sixth unit is used to obtain a maximum value and a minimum value of the first standby sub-data in the first standby behavior data;

[0137] A seventh unit, configured to convert each of the first unused sub-data into second unused sub-data based on the maximum value and the minimum value;

[0138] An eighth unit is used to combine the second standby sub-data of a plurality of consecutive months into the second standby behavior data;

[0139] A ninth unit, configured to generate a behavior change feature based on the adjacent second stand-by sub-data, generate a change speed feature based on the adjacent behavior change feature, and generate a change trend feature based on the adjacent change speed feature;

[0140] A tenth unit is used to fit a plurality of the change trend features into a trend line, obtain the slope of the trend line, and compare the slope with the attribute type;

[0141] An eleventh unit is configured to determine that the second standby behavior data corresponding to the slope is final behavior data if the slope is a positive value and the attribute type is a non-churned user;

[0142] A twelfth unit, configured to determine that the second standby behavior data corresponding to the slope is final behavior data if the slope is a negative value and the attribute type is lost users;

[0143] The first analysis module 30 is used to obtain a number of key communication data corresponding to the unique identification data, obtain a single communication index between the unique identification data and the key communication data, and generate a calibrated communication index corresponding to the unique identification data based on the single communication indexes;

[0144] The first analysis module 30 includes:

[0145] The thirteenth unit is used to obtain a plurality of interaction identification data associated with the unique identification data, and obtain the cumulative number of interactions between the unique identification data and the interaction identification data;

[0146] A fourteenth unit is used to select a plurality of key interaction data from a plurality of interaction identification data based on the accumulated number of interactions;

[0147] The fifteenth unit is used to obtain the daily interaction value, three-day interaction value, weekly interaction value, ten-day interaction value and monthly interaction value between the unique identification data and the key interaction data within a preset time period;

[0148] A sixteenth unit is used to generate the single interaction index based on the daily interaction value, the three-day interaction value, the weekly interaction value, the ten-day interaction value and the monthly interaction value;

[0149] A second analysis module 40 is configured to obtain a first behavior index and a second behavior index corresponding to the final behavior data through the final sub-data, generate a churn index corresponding to the unique identification data based on the calibrated interaction index, the first behavior index and the second behavior index, and determine a churn risk level of non-churned users based on the churn index;

[0150] The second analysis module 40 includes:

[0151] Unit 17 is used to set a number of preset risk levels and index ranges corresponding to the preset risk levels;

[0152] The eighteenth unit is used to match the churn index with a plurality of the index ranges to select a churn risk level corresponding to non-churned users from a plurality of the preset risk levels.

[0153] The present invention also provides a computer, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the user data processing method as described in the above technical solution when executing the computer program.

[0154] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the user data processing method described in the above technical solution is implemented.

[0155] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0156] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A user data processing method, characterized in that: The following steps are involved: Acquire unique identification data of the user, the attribute types of the user include non-churned users and churned users, extract a number of user behavior data from a report library based on the unique identification data, the user behavior data include sub-behavior data for a number of consecutive months, perform data cleaning on the user behavior data to form first standby behavior data, the first standby behavior data include first standby sub-data for a number of consecutive months; Performing data transformation on the first standby behavior data to obtain second standby behavior data, the second standby behavior data including second standby sub-data for a number of consecutive months, and determining the relevance of the second standby behavior data to a user based on the second standby sub-data to select a number of final behavior data from the number of second standby behavior data, the final behavior data including final sub-data for a number of consecutive months; The step of performing data transformation on the first standby behavior data to obtain second standby behavior data, wherein the second standby behavior data includes second standby sub-data of a plurality of consecutive months, comprises: Acquire the maximum value and the minimum value of the first standby sub-data in the first standby behavior data; converting each of the first unused sub-data into second unused sub-data based on the maximum value and the minimum value; combining the second standby sub-data of a plurality of consecutive months into the second standby behavior data; Acquire a number of key interaction data corresponding to the unique identification data, acquire a single interaction index between the unique identification data and the key interaction data, and generate a calibrated interaction index corresponding to the unique identification data based on the number of single interaction indexes; The first behavior index and the second behavior index corresponding to the final behavior data are obtained through the final sub-data, and a churn index corresponding to the unique identification data is generated based on the calibrated interaction index, the first behavior index and the second behavior index. The churn risk level of non-churned users is determined based on the churn index.

2. The user data processing method according to claim 1, characterized in that: The step of cleaning the user behavior data comprises: Obtaining a maximum number of the sub-behavior data in the plurality of user behavior data, and determining whether there are missing months in the user behavior data based on the maximum number; If there are missing months in the user behavior data, obtain the missing number of the missing months, and compare the missing number with the missing threshold; If the missing number is less than the missing threshold, the sub-behavior data corresponding to the two months adjacent to the missing month are obtained to generate filling behavior data, and the filling behavior data is filled into the missing month; If the missing number is greater than the missing threshold, the preset behavior data is added to fill in the missing month; A threshold range is constructed based on all of the sub-behavior data in the user behavior data, and abnormal data is replaced on the user behavior data based on the threshold range.

3. The user data processing method according to claim 1, characterized in that: The calculation formula of the second standby sub-data is: , in, represents the second standby sub-data of the i-th month in the second standby behavior data corresponding to the j-th first standby behavior data, represents the first standby sub-data of the i-th month in the j-th first standby behavior data, represents the minimum value of the first standby sub-data in the jth first standby behavior data, Represents the maximum value of the first standby sub-data in the jth first standby behavior data.

4. The user data processing method according to claim 1, characterized in that: The step of determining the relevance between the second standby behavior data and the user based on the second standby sub-data to select a plurality of final behavior data from a plurality of the second standby behavior data comprises: generating a behavior change feature based on the adjacent second standby sub-data, generating a change speed feature based on the adjacent behavior change feature, and generating a change trend feature based on the adjacent change speed feature; Fitting a plurality of the change trend features into a trend line, obtaining the slope of the trend line, and comparing the slope with the attribute type; If the slope is a positive value and the attribute type is a non-churned user, determining that the second standby behavior data corresponding to the slope is the final behavior data; If the slope is a negative value and the attribute type is lost users, the second standby behavior data corresponding to the slope is determined to be final behavior data.

5. The user data processing method according to claim 1, characterized in that: The step of obtaining a plurality of key communication data corresponding to the unique identification data comprises: Acquire a plurality of interaction identification data associated with the unique identification data, and acquire the cumulative number of interactions between the unique identification data and the interaction identification data; A plurality of key interaction data are selected from a plurality of interaction identification data based on the accumulated number of interactions.

6. The user data processing method according to claim 1, characterized in that: The step of obtaining a single interaction index between the unique identification data and the key interaction data comprises: Within a preset time period, obtaining a daily interaction value, a three-day interaction value, a weekly interaction value, a ten-day interaction value, and a monthly interaction value between the unique identification data and the key interaction data; The single interaction index is generated based on the daily interaction value, the three-day interaction value, the weekly interaction value, the ten-day interaction value, and the monthly interaction value.

7. The user data processing method according to claim 6, characterized in that: The calculation formula of the daily exchange value is: , in, Represents the daily interaction value, Indicates the number of interactions between the unique identification data and the key interaction data based on days. Indicates the total number of days in the preset time period; The calculation formula of the three-day interaction value is: , in, Indicates the three-day interaction value. The number of interactions between the unique identification data and the key interaction data based on a 3-day period; The calculation formula of the single interaction index is: , in, Represents a single interaction index between the unique identification data and the ath key interaction data, Represents the weekly exchange value, Indicates the ten-day exchange value, Indicates the monthly transaction value.

8. The user data processing method according to claim 1, characterized in that: The step of determining the churn risk level of non-churned users based on the churn index comprises: Setting a number of preset risk levels and index ranges corresponding to the preset risk levels; The churn index is matched with a plurality of the index ranges to select a churn risk level corresponding to non-churned users from among the plurality of the preset risk levels.

9. A user data processing system, applied to the user data processing method according to any one of claims 1 to 8, characterized in that: The system comprises: A processing module is used to obtain unique identification data of a user, wherein the attribute types of the user include non-lost users and lost users, extract a plurality of user behavior data from a report library based on the unique identification data, wherein the user behavior data includes sub-behavior data for a plurality of consecutive months, and perform data cleaning on the user behavior data to form first standby behavior data, wherein the first standby behavior data includes first standby sub-data for a plurality of consecutive months; a screening module, configured to perform data transformation on the first standby behavior data to obtain second standby behavior data, wherein the second standby behavior data includes second standby sub-data for a number of consecutive months, and to determine the relevance of the second standby behavior data to the user based on the second standby sub-data, so as to select a number of final behavior data from the number of second standby behavior data, wherein the final behavior data includes final sub-data for a number of consecutive months; The screening module comprises: A sixth unit is used to obtain a maximum value and a minimum value of the first standby sub-data in the first standby behavior data; A seventh unit, configured to convert each of the first unused sub-data into second unused sub-data based on the maximum value and the minimum value; An eighth unit is used to combine the second standby sub-data of a plurality of consecutive months into the second standby behavior data; A first analysis module is used to obtain a number of key communication data corresponding to the unique identification data, obtain a single communication index between the unique identification data and the key communication data, and generate a calibrated communication index corresponding to the unique identification data based on the single communication indexes; The second analysis module is used to obtain the first behavior index and the second behavior index corresponding to the final behavior data through the final sub-data, generate a churn index corresponding to the unique identification data based on the calibrated interaction index, the first behavior index and the second behavior index, and determine the churn risk level of non-churned users based on the churn index.

Citation Information

Patent Citations

  • Telecom customer loss probability prediction method and system based on end-to-end model

    CN111538873A

  • Risk prediction processing method and apparatus, computer device and medium

    WO2020037942A1