Data processing method and device, computer device and storage medium

By binning and cross-processing bank customer characteristics, combined with customer overlap rate analysis, the problem of low accuracy in determining customer characteristic value ranges in existing technologies has been solved, achieving more efficient and accurate determination of customer characteristic value ranges.

CN115187353BActive Publication Date: 2025-12-16INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210814320.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-12
Publication Date
2025-12-16
Estimated Expiration
2042-07-12

AI Technical Summary

Technical Problem

In the banking and finance sector, existing technologies rely on expert experience to determine key ranges of customer characteristic values, resulting in low accuracy.

Method used

By binning customer characteristics and combining feature cross-processing and customer overlap rate analysis, the characteristic value interval combination of the target customer group is determined. The feature cross-processing round is used to screen the characteristic value interval combination that meets the customer overlap rate requirements.

Benefits of technology

It improves the accuracy and speed of determining the key intervals of customer feature values ​​and reduces the computational load of feature cross processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187353B_ABST
    Figure CN115187353B_ABST
Patent Text Reader

Abstract

The application relates to a data processing method and device, computer equipment and a storage medium, and relates to the field of big data. The method comprises: for any customer feature, performing binning processing on the customer feature to obtain a plurality of feature value intervals corresponding to the customer feature; in the i-th round of feature cross processing, performing feature cross processing on a first feature value interval combination determined in the (i-1)-th round of feature cross processing and the plurality of feature value intervals corresponding to each customer feature to obtain a plurality of second feature value interval combinations; for any second feature value interval combination, determining a customer overlap rate of a target customer group and an overall customer group on the second feature value interval combination; according to the customer overlap rate, determining a third feature value interval combination from the second feature value interval combination; when i is equal to N, taking the third feature value interval combination as a target feature value interval combination corresponding to the target customer group. The method can improve the determination accuracy of a customer feature value key interval.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of big data, and in particular, to a data processing method and device, computer equipment and a storage medium. BACKGROUND

[0002] In the field of bank finance, it is often necessary to analyze the behavior patterns of customers, find the key interval of customer characteristic values that have a greater impact on the customer type to which the customer belongs, so as to provide a basis for the future business development of the bank.

[0003] In related technologies, a traditional data mining mode is often used to determine the rules between customer characteristics and customer behavior. This method requires experts to analyze data based on experience and relies too much on the subjective judgment of experts, resulting in low accuracy in determining the key interval of customer characteristic values. SUMMARY

[0004] Therefore, it is necessary to provide a data processing method, device, computer equipment and storage medium to solve the above technical problems.

[0005] In a first aspect, the present application provides a data processing method. The method comprises:

[0006] For any customer characteristic, the customer characteristic is subjected to binning processing to obtain a plurality of characteristic value intervals corresponding to the customer characteristic;

[0007] In the i-th round of feature cross processing, the first characteristic value interval combination determined in the (i-1)-th round of feature cross processing and the plurality of characteristic value intervals corresponding to each customer characteristic are subjected to feature cross processing to obtain a plurality of second characteristic value interval combinations, wherein the characteristic value interval combination includes at least one characteristic value interval;

[0008] For any second characteristic value interval combination, the customer overlap rate of a target customer group and an all customer group on the second characteristic value interval combination is determined, the customer overlap rate is used to represent the overlap degree of a target customer in the all customer group and a target customer in the target customer group, the target customer is a user whose customer characteristics match the second characteristic value interval combination, and at least one customer belonging to the same customer group corresponds to the same customer category;

[0009] According to the customer overlap rate, a third characteristic value interval combination is determined from the second characteristic value interval combination;

[0010] In the case where i is equal to N, the third characteristic value interval combination is taken as a target characteristic value interval combination corresponding to the target customer group, wherein i and N are positive integers, and N is the total number of feature cross processing rounds.

[0011] In one of the embodiments, the determining of the customer overlap rate of the target customer group and the whole customer group on the second characteristic value interval combination comprises:

[0012] determining a first number of first target customers in the target customer group whose customer characteristics match the second characteristic value interval combination, and determining a second number of second target customers in the whole customer group whose customer characteristics match the second characteristic value interval combination;

[0013] obtaining the customer overlap rate of the target customer group and the whole customer group on the second characteristic value interval combination according to the first number and the second number.

[0014] In one of the embodiments, the obtaining of the customer overlap rate of the target customer group and the whole customer group on the second characteristic value interval combination according to the first number and the second number comprises:

[0015] determining a ratio of the first number and the second number, and taking the ratio as the customer overlap rate of the target customer group and the whole customer group on the second characteristic value interval combination.

[0016] In one of the embodiments, the binning processing of the customer characteristics to obtain the multiple characteristic value intervals corresponding to the customer characteristics comprises:

[0017] for any of the customer characteristics, determining an influence index of the customer characteristic on the target customer group, the influence index being used to represent the importance of the customer characteristic on the target customer group;

[0018] determining target customer characteristics from the customer characteristics according to the influence index of each of the customer characteristics on the target customer group;

[0019] performing binning processing on each of the target customer characteristics to obtain multiple characteristic value intervals corresponding to the target customer characteristics.

[0020] In one of the embodiments, the determining of the target customer characteristics from the customer characteristics according to the influence index of each of the customer characteristics on the target customer group comprises:

[0021] sequencing each of the customer characteristics from high to low according to the influence index of each of the customer characteristics to obtain an influence index ranking sequence, and determining the first preset number of customer characteristics in the influence index ranking sequence as the target customer characteristics;

[0022] or, determining the customer characteristics whose influence index is greater than an influence index threshold as the target customer characteristics.

[0023] In one of the embodiments, the method further comprises:

[0024] determining a customer category corresponding to each of the customers according to product purchase data of the customers;

[0025] dividing all the customers into a plurality of customer groups according to the customer categories, at least one customer in the customer groups corresponding to a same customer category.

[0026] In one of the embodiments, the method further comprises:

[0027] determining a customer overlap rate of the target customer group and the all customer group in the feature value interval for any feature value interval in the target feature value interval combination;

[0028] sorting and displaying each of the feature value intervals in the target feature value interval combination according to the customer overlap rate corresponding to the feature value interval.

[0029] In a second aspect, the application further provides a data processing device. The device comprises:

[0030] a binning module configured to perform binning processing on any customer feature to obtain a plurality of feature value intervals corresponding to the customer feature;

[0031] a feature intersection module configured to perform feature intersection processing on a first feature value interval combination determined in an i-1th round of feature intersection processing and a plurality of feature value intervals corresponding to each of the customer features to obtain a plurality of second feature value interval combinations in an ith round of feature intersection processing, wherein the feature value interval combination comprises at least one of the feature value intervals;

[0032] a first determination module configured to determine a customer overlap rate of a target customer group and an all customer group in the second feature value interval combination for any of the second feature value interval combinations, the customer overlap rate representing an overlap degree of a target customer in the all customer group and a target customer in the target customer group, the target customer being a user whose customer feature matches the second feature value interval combination, wherein at least one customer belonging to a same customer group corresponds to a same customer category;

[0033] a second determination module configured to determine a third feature value interval combination from the second feature value interval combination according to the customer overlap rate;

[0034] The third determining module is configured to, in a case where i is equal to N, combine the third characteristic value interval group as a target characteristic value interval group corresponding to the target customer group, where i and N are positive integers, and N is a total number of characteristic cross processing rounds.

[0035] In one of the embodiments, the first determining module is further configured to:

[0036] determine a first number of first target customers in the target customer group whose customer characteristics match the second characteristic value interval group, and determine a second number of second target customers in the total customer group whose customer characteristics match the second characteristic value interval group;

[0037] obtain a customer overlap rate of the target customer group and the total customer group on the second characteristic value interval group according to the first number and the second number.

[0038] In one of the embodiments, the first determining module is further configured to:

[0039] determine a ratio of the first number and the second number, and take the ratio as the customer overlap rate of the target customer group and the total customer group on the second characteristic value interval group.

[0040] In one of the embodiments, the binning module is further configured to:

[0041] for any of the customer characteristics, determine an influence index of the customer characteristic for the target customer group, the influence index being used to represent an importance degree of the customer characteristic for the target customer group;

[0042] determine target customer characteristics from the customer characteristics according to the influence index of each of the customer characteristics for the target customer group;

[0043] perform binning processing on each of the target customer characteristics to obtain a plurality of characteristic value intervals corresponding to the target customer characteristics.

[0044] In one of the embodiments, the binning module is further configured to:

[0045] sort each of the customer characteristics from high to low according to the influence index of each of the customer characteristics to obtain an influence index ranking sequence, and determine a preset number of customer characteristics in the influence index ranking sequence as target customer characteristics;

[0046] or, determine the customer characteristics with an influence index greater than an influence index threshold as the target customer characteristics.

[0047] In one of the embodiments, the apparatus further comprises:

[0048] a fourth determining module, configured to determine a customer category corresponding to each of the customers according to product purchase data of the customers;

[0049] a grouping module, configured to group all the customers into a plurality of customer groups according to the customer categories, wherein at least one customer group includes at least one customer, and at least one customer in the customer group corresponds to the same customer category.

[0050] In one of the embodiments, the apparatus further includes:

[0051] a fifth determining module, configured to determine, for any characteristic value interval in the target characteristic value interval combination, a customer overlap rate of the target customer group and the all customer groups in the characteristic value interval;

[0052] a sorting module, configured to sort and display each of the characteristic value intervals in the target characteristic value interval combination according to the customer overlap rate corresponding to the characteristic value interval.

[0053] In a third aspect, the present application provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements any of the above methods when executing the computer program.

[0054] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement any of the above methods.

[0055] In a fifth aspect, the present application provides a computer program product. The computer program product includes a computer program, and the computer program is executed by a processor to implement any of the above methods.

[0056] The aforementioned data processing method, apparatus, computer equipment, and storage medium can repeatedly perform feature cross-processing on combinations of feature value intervals that meet customer overlap requirements and multiple feature value intervals corresponding to each customer feature, until the number of rounds of feature cross-processing reaches a preset number. The feature value interval combinations that meet customer overlap requirements determined in the last round are then identified as the key customer feature value intervals for the target customer group. In this embodiment, feature value interval combinations are screened in each round of feature cross-processing, ensuring that only those combinations that meet customer overlap requirements enter the next round. This ensures that each round yields feature value interval combinations with high customer overlap rates, reducing the computational load of feature cross-processing, accelerating the determination of key customer feature value intervals, and improving the accuracy of determining key customer feature value intervals. Attached Figure Description

[0057] Figure 1 This is a flowchart illustrating a data processing method in one embodiment;

[0058] Figure 2 This is a flowchart illustrating step 106 in one embodiment;

[0059] Figure 3 This is a flowchart illustrating step 102 in one embodiment;

[0060] Figure 4 This is a flowchart illustrating step 304 in one embodiment;

[0061] Figure 5 This is a flowchart illustrating a data processing method in one embodiment;

[0062] Figure 6 This is a flowchart illustrating a data processing method in one embodiment;

[0063] Figure 7 This is a schematic diagram of a data processing method in one embodiment;

[0064] Figure 8 This is a schematic diagram of a data processing method in one embodiment;

[0065] Figure 9 This is a structural block diagram of a data processing device in one embodiment;

[0066] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0067] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, and are not intended to limit the present application.

[0068] In one embodiment, as shown in Figure 1 A data processing method is provided, and the embodiment is exemplified by the method applied to a terminal. It should be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and can be realized through the interaction of the terminal and the server. In the embodiment, the method includes the following steps.

[0069] In step 102, for any customer feature, the customer feature is binned to obtain a plurality of feature value intervals corresponding to the customer feature.

[0070] In the embodiment of the present application, for any customer feature, the customer feature can be divided into a plurality of feature value intervals. The customer feature can include features such as basic information, asset information, and in-row product information, for example, the customer feature can include age, income, and in-row deposit. Binning the customer feature means dividing the feature values corresponding to all customers on the customer feature into a plurality of feature value intervals, for example, if the customer feature is a continuous numerical value such as age, binning the age means dividing the age into numerical intervals, for example, dividing the age into the following numerical intervals: under 20 years old, 21-30 years old, 31-40 years old, etc. If the customer feature is a categorical numerical value such as gender, the feature value intervals can be divided according to the categories, for example, binning the gender means dividing the customers into two categories according to the gender categories: male and female.

[0071] In this way, after performing binning processing for each customer feature, a plurality of feature value intervals corresponding to each customer feature can be obtained.

[0072] It should be noted that the method of binning the customer feature is not specifically limited in the embodiment of the present application, and any method capable of binning the customer feature is applicable to the embodiment of the present application, for example, equal frequency binning method, optimal binning method, etc.

[0073] In step 104, in the i-th round of feature cross processing, the first feature value interval combination determined in the (i-1)-th round of feature cross processing and the plurality of feature value intervals corresponding to each customer feature are subjected to feature cross processing to obtain a plurality of second feature value interval combinations, wherein the feature value interval combination includes at least one feature value interval.

[0074] Step 106, determining, for any second feature value interval combination, a customer overlap rate of the target customer group and the whole customer group on the second feature value interval combination, the customer overlap rate being used to represent the overlapping degree of the target customer in the whole customer group and the target customer in the target customer group, the target customer being a user whose customer feature matches the second feature value interval combination, wherein at least one customer belonging to the same customer group corresponds to the same customer category.

[0075] Step 108, determining, according to the customer overlap rate, a third feature value interval combination from the second feature value interval combination.

[0076] Step 110, in the case that i is equal to N, taking the third feature value interval combination as the target feature value interval combination corresponding to the target customer group, wherein i and N are positive integers, and N is the total number of feature intersection processing rounds.

[0077] In the embodiments of the present application, N rounds of feature intersection processing can be performed. In any round of feature intersection processing, a corresponding feature value interval combination can be obtained based on the result of the previous round of feature intersection, and the feature value interval combination includes at least one feature value interval. The second feature value interval combination obtained in the i th round of feature intersection processing includes i feature value intervals, and the i feature value intervals belong to different customer features respectively.

[0078] For example, the customers can be grouped in advance to obtain a plurality of customer groups, wherein the customers belonging to the same customer group correspond to the same customer category. The target customer group can be any customer group in the plurality of customer groups. For the target customer group, when the i th round of feature intersection processing is performed, the first feature value interval combination whose customer overlap rate meets the requirement and obtained in the i-1 th round of feature intersection processing can be subjected to feature intersection processing with the plurality of feature value intervals corresponding to each customer feature to obtain the second feature value interval combination, and the customer overlap rate corresponding to each second feature value interval combination can be determined to screen the third feature value interval combination from the second feature value interval combination according to the customer overlap rate, so that the third feature value interval combination participates in the i+1 th round of feature intersection processing, until the feature intersection round reaches N. The value of N can be set by the person skilled in the art according to the actual demand, for example, when the feature value interval included in the target feature value interval combination needs to be more, N can be set larger; when the feature value interval included in the target feature value interval combination needs to be less, N can be set smaller. The customer overlap rate meeting the requirement can mean that the customer overlap rate is greater than a certain threshold, or the ranking of the customer overlap rate is higher than a certain threshold, which is not limited in the embodiments of the present application.

[0079] The customer overlap rate is used to represent the overlap degree of the customers corresponding to the feature value interval combination and the customers in the target customer group. The customer overlap rate can be determined by the number of customers corresponding to the feature value interval combination in the target customer group and the number of customers corresponding to the feature value interval combination in the whole customer group. For example, if the feature value interval combination is (A<2000, 1<B<6), the customer overlap rate of the target customer group corresponding to the feature value interval combination can be obtained according to the number of customers in the target customer group whose feature A and feature B satisfy (A<2000, 1<B<6) and the number of customers in the whole customer group whose feature A and feature B satisfy (A<2000, 1<B<6).

[0080] If the customer overlap rate is high, it indicates that most of the customers corresponding to the feature value interval combination belong to the target customer group, and therefore the feature value interval combination is more likely to be a key interval of the customer feature value of the target customer group. If the customer overlap rate is low, it indicates that most of the customers corresponding to the feature value interval combination do not belong to the target customer group, and therefore the feature value interval combination is less likely to be a key interval of the customer feature value of the target customer group.

[0081] When i is equal to 1, that is, when the first round of feature cross processing is performed, there is no first feature value interval combination determined in the (i-1)th round of feature cross processing. At this time, all feature value intervals corresponding to each customer feature can be taken as second feature value interval combinations, that is, any second feature value interval combination includes one feature value interval at this time. Then, a third feature value interval combination satisfying the customer overlap rate requirement is determined from the second feature value interval combinations. When the second round of feature cross processing is performed, the third feature value interval combination can be taken as the first feature value interval combination in the second round of feature cross processing.

[0082] When i is greater than 1, there is a first feature value interval combination determined in the (i-1)th round of feature cross processing. For any first feature value interval combination, the customer feature corresponding to the feature value interval included in the first feature value interval combination can be excluded from all customer features, and the feature value intervals corresponding to the remaining customer features are cross-processed with the first feature value interval combination to obtain second feature value interval combinations. Then, a third feature value interval combination satisfying the customer overlap rate requirement is determined from the second feature value interval combinations. When the (i+1)th round of feature cross processing is performed, the third feature value interval combination can be taken as the first feature value interval combination in the (i+1)th round of feature cross processing.

[0083] The customer overlap rate meeting the requirement can refer to the customer overlap rate being greater than a certain threshold value, or the ranking of the customer overlap rate being higher than a certain threshold value. For example, the second feature value interval combinations can be ranked from high to low according to the customer overlap rates of the second feature value interval combinations, and the top preset number of second feature value interval combinations are selected as the third feature value interval combinations. For example, when the preset number is 3, the second feature value interval combinations with the top 3 customer overlap rates are selected as the third feature value interval combinations. The preset number can be set by a person skilled in the art according to requirements. For example, when the third feature value interval combinations to be obtained are less, and the speed of determining the target feature value interval combination is to be accelerated, the preset number can be set to be less; when the third feature value interval combinations to be obtained are more, and the accuracy of determining the target feature value interval combination is to be improved, the preset number can be set to be more.

[0084] Alternatively, the customer feature with the customer overlap rate greater than the customer overlap rate threshold value can also be determined as the target customer feature. The customer overlap rate threshold value can be set by a person skilled in the art according to requirements. For example, when the third feature value interval combinations to be obtained are less, and the speed of determining the target feature value interval combination is to be accelerated, the customer overlap rate threshold value can be set to be higher; when the third feature value interval combinations to be obtained are more, and the accuracy of determining the target feature value interval combination is to be improved, the customer overlap rate threshold value can be set to be lower.

[0085] It should be noted that the preset number and the customer overlap rate threshold value can be different in different feature intersection rounds. For example, the preset number can be set to 3 in the first round of feature intersection processing, and the preset number can be set to 10 in the second round of feature intersection processing. The present application does not make a specific limitation on this.

[0086] When i=N, the third feature value interval combination determined in the i-th round can be taken as the target feature value interval combination corresponding to the target customer group, and the target feature value interval combination is the customer feature value key interval corresponding to the target customer group.

[0087] For example, taking the example of i being equal to 3, the customer features being A, B, C and D, and the two third feature value interval combinations determined in the second round of feature cross processing being feature value interval combination one (A < 2000, 1 < B < 6) and feature value interval combination two (7 < B < 17, 10000 < C), in the third round of feature cross processing, since the feature value intervals included in feature value interval combination one correspond to customer feature A and customer feature B, at this time, the second feature value interval combination can be obtained by performing second-order feature cross processing on all feature value intervals corresponding to customer feature C and customer feature D in feature value interval combination one, and the second feature value interval combination includes any one of feature value interval combination one, customer feature C or customer feature D. Since the feature value intervals included in feature value interval combination two correspond to customer feature B and customer feature C, at this time, the second feature value interval combination can be obtained by performing second-order feature cross processing on all feature value intervals corresponding to customer feature A and customer feature D in feature value interval combination two, and the second feature value interval combination includes any one of feature value interval combination two, customer feature A or customer feature D.

[0088] Suppose that customer feature A corresponds to 3 feature value intervals in total, customer feature C corresponds to 2 feature value intervals in total, and customer feature D corresponds to 4 feature value intervals in total, then 13 second feature value interval combinations can be finally obtained. The embodiments of the present application can further calculate the customer overlap rates corresponding to the 13 second feature value interval combinations, and filter out third feature value interval combinations whose customer overlap rates meet the requirements, so as to take the third feature value interval combinations obtained in the third round of feature cross processing as the first feature value interval combinations in the fourth round of feature cross processing, and perform feature cross processing on all feature value intervals corresponding to each customer feature except for the customer features corresponding to the feature value intervals included in the first feature value interval combinations.

[0089] The data processing method provided by the embodiments of the present application can repeatedly perform feature cross processing on feature value interval combinations whose customer overlap rates meet the requirements and multiple feature value intervals corresponding to each customer feature until the round of feature cross processing reaches the pre-set round, and determine the feature value interval combination whose customer overlap rate meets the requirements obtained in the last round as the customer feature value key interval of the target customer group. The embodiments of the present application filter feature value interval combinations in the process of each round of feature cross processing, and only enable feature value interval combinations whose customer overlap rates meet the requirements to enter the next round of feature cross processing, so as to enable the feature value interval combinations obtained in each round to be feature value interval combinations with higher customer overlap rates, which not only can reduce the calculation amount of feature cross processing, accelerate the speed of determining the customer feature value key interval, but also can improve the accuracy of determining the customer feature value key interval.

[0090] In one embodiment, as shown in FIG. 1, the method comprises the following steps: Figure 2 In step 106, the method further comprises the following steps:

[0091] In step 202, the method further comprises the following steps:

[0092] In step 204, the method further comprises the following steps:

[0093] In the embodiments of the present application, the customer overlap rate of the target customer group and the whole customer group in the second characteristic value interval combination can be determined by the first number of the first target customers in the target customer group and the second number of the second target customers in the whole customer group.

[0094] For example, if the customer characteristics are A, B, C and D, and the second characteristic value interval combination is (A<2000, 1<B<6), the customers in the target customer group whose characteristic data of customer characteristic A is less than 2000 and whose characteristic data of customer characteristic B is greater than 1 and less than 6 are all the first target customers whose customer characteristics match the second characteristic value interval combination; the customers in the whole customer group whose characteristic data of customer characteristic A is less than 2000 and whose characteristic data of customer characteristic B is greater than 1 and less than 6 are all the second target customers whose customer characteristics match the second characteristic value interval combination.

[0095] After the first number of the first target customers and the second number of the second target customers are determined respectively, the customer overlap rate of the target customer group and the whole customer group in the second characteristic value interval combination can be determined further. For example, the method for determining the customer overlap rate by the first number and the second number can be taking the ratio of the first number and the second number, and the embodiments of the present application are not limited in this regard.

[0096] The data processing method provided in the embodiments of the present application can determine the customer overlap rate of the target customer group and the total customer group on the second characteristic value interval combination according to the first quantity of the first target customer in the target customer group matching the second characteristic value interval combination and the second quantity of the second target customer in the total customer group matching the second characteristic value interval combination, and then filter the characteristic value interval combination by using the customer overlap rate in the process of each round of feature cross processing, so that only the characteristic value interval combination with the customer overlap rate meeting the requirement can enter the next round of feature cross processing. Therefore, the characteristic value interval combination obtained in each round is the characteristic value interval combination with a higher customer overlap rate, which can not only reduce the calculation amount of feature cross processing and speed up the determination of the customer characteristic value key interval, but also improve the accuracy of the determination of the customer characteristic value key interval.

[0097] In one embodiment, in step 204, the customer overlap rate of the target customer group and the total customer group on the second characteristic value interval combination is obtained according to the first quantity and the second quantity, including:

[0098] The ratio of the first quantity and the second quantity is determined, and the ratio is taken as the customer overlap rate of the target customer group and the total customer group on the second characteristic value interval combination.

[0099] In the embodiments of the present application, the ratio of the first quantity and the second quantity can be determined, and the ratio is taken as the customer overlap rate of the target customer group and the total customer group on the second characteristic value interval combination. For example, if the first quantity is 967 and the second quantity is 1000, the ratio of the first quantity and the second quantity is 967 / 1000=0.967, and then 0.967 can be taken as the customer overlap rate of the target customer group and the total customer group on the second characteristic value interval combination.

[0100] The data processing method provided in the embodiments of the present application can determine the customer overlap rate of the target customer group and the total customer group on the second characteristic value interval combination by using the ratio of the first quantity and the second quantity, and then filter the characteristic value interval combination by using the customer overlap rate in the process of each round of feature cross processing, so that only the characteristic value interval combination with the customer overlap rate meeting the requirement can enter the next round of feature cross processing. Therefore, the characteristic value interval combination obtained in each round is the characteristic value interval combination with a higher customer overlap rate, which can not only reduce the calculation amount of feature cross processing and speed up the determination of the customer characteristic value key interval, but also improve the accuracy of the determination of the customer characteristic value key interval.

[0101] In one embodiment, as shown in Figure 3 In step 102, the customer characteristics are subjected to binning processing to obtain a plurality of characteristic value intervals corresponding to the customer characteristics, including:

[0102] At step 302, for any customer feature, an influence index of the customer feature on the target customer group is determined, the influence index being used to represent an importance of the customer feature on the target customer group.

[0103] At step 304, according to the influence index of each customer feature on the target customer group, a target customer feature is determined from the customer features.

[0104] At step 306, each target customer feature is binned to obtain a plurality of feature value intervals corresponding to the target customer feature.

[0105] In the embodiments of the present application, the target customer feature corresponding to the target customer group can be determined from the customer features according to the influence index of each customer feature on the target customer group, and the target customer feature is binned to obtain a plurality of feature value intervals corresponding to the target customer feature, so as to perform feature cross processing on the plurality of feature value intervals corresponding to the target customer feature to obtain a target feature value interval combination corresponding to the target customer group.

[0106] The influence index is used to represent an importance of the customer feature on the target customer group, that is, an influence size of the customer feature on whether a certain customer belongs to the target customer group. For example, the influence index can be information gain, information value, etc. of the customer feature. The way of determining the target customer feature according to the influence index can be to select customer features with top threshold number of influence indexes as the target customer features, or to select customer features with influence indexes greater than a certain threshold as the target customer features, which are not limited in the embodiments of the present application.

[0107] Taking the information gain of the influence index as an example, the information gain of each customer feature can be determined by the lightGBM algorithm (Light Gradient Boosting Machine) in the embodiments of the present application. The information gain is used to represent a contribution of the customer feature to the division of customers, and the higher the information gain, the higher the probability that only one type of customers is divided out from each customer group after the customers are divided according to the customer feature. The lightGBM algorithm is a gradient boosting algorithm based on decision trees. When splitting the decision tree, the lightGBM algorithm calculates the information gain of each customer feature, and first splits the decision tree based on the customer feature with the maximum information gain. Therefore, in the decision tree constructed by the lightGBM algorithm, the closer the splitting node based on a certain customer feature to the root node of the decision tree, the higher the information gain of the customer feature. Therefore, the order of the splitting nodes from the root node to the leaf node of the decision tree is the information gain order of each customer feature.

[0108] In this embodiment of the application, when determining the target customer characteristics corresponding to the target customer group using the lightGBM algorithm, data labels can be set for each customer data first. For example, the data label of the customer data corresponding to the customer belonging to the target customer group can be set to 1, and the data label of the customer data corresponding to the customer not belonging to the target customer group can be set to 0, so that the lightGBM algorithm can only calculate the information gain of each customer characteristic based on the customer data with data label 1, that is, only calculate the ability of each customer characteristic to divide the customers belonging to the target customer group from all customers.

[0109] Furthermore, customer data can be preprocessed. In this embodiment, customer features whose corresponding customer data are continuous values ​​can be set as numerical variables in the lightGBM algorithm, and customer features whose corresponding customer data are discrete values ​​can be set as categorical variables in the lightGBM algorithm. When processing outliers in customer data, this embodiment can remove customer features that correspond to only one customer data point; when customer data values ​​are abnormal, for example, if the value of a customer data point that should be a number is a character, the value of that customer data can be set to -1. If customer data corresponding to a certain customer feature is missing, when the missing data is more than 85%, the customer feature can be removed; when the missing data is less than 85%, if the customer data value is a number, the value of the missing customer data can be set to 0, and if the customer data value is a category, the value of the missing customer data can be set to N.

[0110] The embodiments of this application can further construct a decision tree using the lightGBM algorithm, and determine the information gain ranking of each customer feature according to the order of each split node.

[0111] The data processing method provided in this application can determine target customer characteristics from customer characteristics based on the influence index of each customer characteristic on the target customer group. This application filter customer characteristics, selecting only those characteristics that significantly influence whether a customer belongs to the target customer group as target customer characteristics. It also performs binning and feature cross-processing only on the target customer characteristics, thus improving the speed of determining the combination of target characteristic value ranges.

[0112] In one embodiment, such as Figure 4 As shown, in step 304, target customer characteristics are determined from the customer characteristics based on the impact index of each customer characteristic on the target customer group, including:

[0113] Step 402: Based on the influence index of each customer feature, sort each customer feature from high to low to obtain the influence index ranking sequence, and determine the first number of customer features in the influence index ranking sequence as target customer features.

[0114] In the embodiments of the present application, the customer features can be ranked from high to low according to the influence indexes of the customer features, and the top pre-set number of customer features are selected as the target customer features. For example, when the pre-set number is 3, the top 3 customer features in terms of the influence indexes are selected as the target customer features. The pre-set number can be set by those skilled in the art according to requirements. For example, when the number of feature value interval combinations to be obtained is small and the speed of determining the target feature value interval combination is to be accelerated, the pre-set number can be set to be small; when the number of feature value interval combinations to be obtained is large and the accuracy of determining the target feature value interval combination is to be improved, the pre-set number can be set to be large.

[0115] In step 404, or the customer features with the influence indexes greater than the influence index threshold value are determined as the target customer features.

[0116] In the embodiments of the present application, the customer features with the influence indexes greater than the influence index threshold value are determined as the target customer features. The influence index threshold value can be set by those skilled in the art according to requirements. For example, when the number of feature value interval combinations to be obtained is small and the speed of determining the target feature value interval combination is to be accelerated, the influence index threshold value can be set to be high; when the number of feature value interval combinations to be obtained is large and the accuracy of determining the target feature value interval combination is to be improved, the influence index threshold value can be set to be low.

[0117] The data processing method provided in the embodiments of the present application can select the top pre-set number of customer features in terms of the influence indexes as the target customer features, or select the customer features with the influence indexes greater than the influence index threshold value as the target customer features. The embodiments of the present application screen the customer features, and only select the customer features that have a greater impact on whether a customer belongs to a target customer group as the target customer features, and only perform binning and feature intersection processing on the target customer features, so that the speed of determining the target feature value interval combination can be improved.

[0118] In one embodiment, as shown in FIG. 5, the above method further includes: Figure 5

[0119] In step 502, the customer categories corresponding to the customers are determined according to the product purchase data of the customers.

[0120] In step 504, all the customers are divided into a plurality of customer groups according to the customer categories, and at least one customer is included in each customer group, and at least one customer in each customer group corresponds to the same customer category.

[0121] ​In the embodiments of the present application, the product purchase data of each customer can be used to determine the customer category corresponding to each customer, and then all customers can be divided into multiple customer groups according to the customer categories. The product purchase data can include the current product purchase data and the historical same-period product purchase data. The length of a period can be set by those skilled in the art according to actual needs, for example, one week, one month, one year, etc.

[0122] For example, the current product purchase data can be the current product purchase amount data, and the historical same-period product purchase data can be the historical same-period product purchase amount data. The customer category corresponding to each customer can be determined by subtracting the historical same-period product purchase amount data of the customer from the current product purchase amount data of the customer, that is, the customer category corresponding to each customer is the customer with a decreased product purchase amount or the customer with an increased product purchase amount, and all customers can be divided into a customer group with a decreased product purchase amount and a customer group with an increased product purchase amount.

[0123] Further, the customers can be classified more carefully. For example, the current product purchase data of the customer can include the current product purchase price daily average data (hereinafter referred to as the current price) and the current product purchase quantity daily average data (hereinafter referred to as the current daily average), and the historical same-period product purchase data of the customer can include the historical same-period product purchase price daily average data (hereinafter referred to as the same-period price) and the historical same-period product purchase quantity daily average data (hereinafter referred to as the same-period daily average). According to the difference between the current price and the historical price of the customer and the difference between the current daily average and the historical daily average of the customer, multiple customer categories can be obtained, and then all customers can be divided into multiple customer groups (see Formula 1, Formula 2, Formula 3, and Table 1).

[0124] Product business volume contribution difference=(current daily average-same-period daily average)×current price (Formula 1)

[0125] The product business volume contribution difference is used to represent the difference between the actual purchase amount of the customer in the current period and the product amount that the customer will purchase in the current period, assuming that the customer maintains the historical same-period product purchase daily average quantity.

[0126] Product price contribution difference=(current price-same-period price)×current daily average (Formula 2)

[0127] The product price contribution difference is used to represent the difference between the actual purchase amount of the customer in the current period and the product amount that the customer will purchase in the current period, assuming that the customer maintains the historical same-period product purchase daily average price.

[0128] The sum of the product business volume contribution difference and the product price contribution difference is the product total contribution difference, which is used to represent the increase trend of the actual purchase amount of the customer in the current period compared with the actual purchase amount of the customer in the same period (see Formula Three):

[0129] Product total contribution difference = product business volume contribution difference + product price contribution difference (Formula Three)

[0130] According to the product business volume contribution difference, the product price contribution difference, and the product total contribution difference, six customer categories can be set: (1) both quantity and price increase, that is, the product purchase quantity and the product purchase price of the customer in the current period increase, and the actual purchase amount of the customer in the current period increases compared with the actual purchase amount of the customer in the same period (product business volume contribution difference > 0, product price contribution difference > 0, product total contribution difference > 0); (2) quantity compensates for price, that is, the product purchase price of the customer in the current period decreases, but the product purchase quantity of the customer in the current period increases, and the actual purchase amount of the customer in the current period increases compared with the actual purchase amount of the customer in the same period (product business volume contribution difference > 0, product price contribution difference < 0, product total contribution difference > 0); (3) price compensates for quantity, that is, the product purchase quantity of the customer in the current period decreases, but the product purchase price of the customer in the current period increases, and the actual purchase amount of the customer in the current period increases compared with the actual purchase amount of the customer in the same period (product business volume contribution difference < 0, product price contribution difference > 0, product total contribution difference > 0); (4) quantity does not compensate for price, that is, the product purchase price of the customer in the current period decreases, although the product purchase quantity of the customer in the current period increases, but the actual purchase amount of the customer in the current period decreases compared with the actual purchase amount of the customer in the same period (product business volume contribution difference > 0, product price contribution difference < 0, product total contribution difference < 0); (5) price does not compensate for quantity, that is, the product purchase quantity of the customer in the current period decreases, although the product purchase price of the customer in the current period increases, but the actual purchase amount of the customer in the current period decreases compared with the actual purchase amount of the customer in the same period (product business volume contribution difference > 0, product price contribution difference < 0, product total contribution difference < 0); (6) both quantity and price decrease, that is, the product purchase quantity and the product purchase price of the customer in the current period decrease, and the actual purchase amount of the customer in the current period decreases compared with the actual purchase amount of the customer in the same period (product business volume contribution difference < 0, product price contribution difference < 0, product total contribution difference < 0) (see Table 1):

[0131] Table 1

[0132]

[0133] The embodiments of the present application can further divide all customers into six customer groups according to the customer categories, and determine the target customer characteristics corresponding to each customer group.

[0134] The data processing method provided in the embodiments of the present application can divide all customers into a plurality of customer groups according to the product purchase data of the customers, so as to determine the target characteristic value interval combination corresponding to each customer group. The embodiments of the present application can determine the target characteristic value interval combination corresponding to each customer group respectively, so that the target characteristic value interval combination which has a greater impact on whether a customer belongs to the customer group can be displayed for different customer groups, and thus the reference value of the determined target characteristic value interval combination can be improved.

[0135] In one embodiment, as shown in FIG. 6, the method further includes: Figure 6

[0136] Step 602: For any characteristic value interval in the target characteristic value interval combination, determining a customer overlap rate of the target customer group and the all customer group on the characteristic value interval.

[0137] Step 604: According to the customer overlap rates corresponding to each characteristic value interval in the target characteristic value interval combination, sorting and displaying each characteristic value interval in the target characteristic value interval combination.

[0138] In the embodiments of the present application, the customer overlap rate corresponding to each characteristic value interval in the target characteristic value interval combination can be determined, and each characteristic value interval in each target characteristic value interval combination can be sorted and displayed according to the customer overlap rate, so as to preferentially display the characteristic value interval with the highest customer overlap rate.

[0139] For example, according to the customer overlap rates corresponding to each characteristic value interval, each characteristic value interval can be sorted from high to low to obtain a ranking sequence, and each characteristic value interval can be displayed according to the ranking of each characteristic value interval in the ranking sequence. The higher the ranking of a characteristic value interval in the ranking sequence, the higher the customer overlap rate corresponding to the characteristic value interval, that is, the customers corresponding to the characteristic value interval are more likely to belong to the target customer group, and therefore the characteristic value interval is more important than other characteristic value intervals in the target characteristic value interval combination for judging whether a customer belongs to the target customer group.

[0140] For example, if the determined target characteristic value interval combination is (A, B, C), and the customer overlap rate corresponding to the A characteristic value interval is 0.8, the customer overlap rate corresponding to the B characteristic value interval is 0.6, and the customer overlap rate corresponding to the C characteristic value interval is 0.7, then each characteristic value interval can be sorted from high to low according to the customer overlap rate corresponding to each characteristic value interval to obtain a ranking sequence: A, C, B. When displaying each characteristic value interval, the A characteristic value interval can be preferentially displayed, then the C characteristic value interval, and finally the B characteristic value interval. For example, one example of finally displaying the target characteristic value interval combination and each characteristic value interval can be as shown in Table 2.

[0141] ​Table 2

[0142]

[0143] The further left the characteristic value interval is in the table, the higher the customer overlap rate of the characteristic value interval.

[0144] The data processing method provided in the embodiments of the present application can determine the importance ranking of each characteristic value interval according to the customer overlap rate of each characteristic value interval in the target characteristic value interval combination, and preferentially display the characteristic value interval with a higher importance ranking according to the importance ranking. When the target characteristic value interval combination is displayed, the embodiments of the present application can also display the importance of each characteristic value interval in the target characteristic value interval combination to the target customer group, so as to improve the reference value of the determined target characteristic value interval combination.

[0145] In order for those skilled in the art to better understand the embodiments of the present application, the embodiments of the present application are described below through specific examples.

[0146] As shown in the examples of FIGS. 1 and 2, a flowchart of a data processing method is shown. Figure 7 Figure 8 As shown in the examples of FIGS. 1 and 2, a flowchart of a data processing method is shown.

[0147] In the embodiments of the present application, customer information such as customer assets, credit risk, product signing identifier, etc. can be selected, and each customer feature can be constructed according to the customer information. The customer features can include customer age, income, in-row deposit, etc.

[0148] Further, all customers can be grouped according to customer categories to determine the target characteristic value interval combination corresponding to each customer group. The grouping method of all customers can refer to the related description of the foregoing embodiments, which will not be described herein again.

[0149] For any customer group, the embodiments of the present application can further determine the information gain ordering of each customer feature through the lightGBM algorithm, and determine the preset number of customer features in the ordering as the target customer features of the customer group. The method of determining the information gain ordering of each customer feature through the lightGBM algorithm can refer to the related description of the foregoing embodiments, which will not be described herein again. For example, the preset number can be 5, that is, in the embodiments of the present application, each customer group corresponds to 5 target customer features.

[0150] ​After determining the target customer characteristics corresponding to the customer category, the embodiments of the present application can perform binning processing on each target customer characteristic based on the optimal binning method of the decision tree to obtain a plurality of characteristic value intervals. For example, the embodiments of the present application can set the maximum number of bins to 6, and the IV value (Information Value) of each bin is at least 0.05. That is, for any customer characteristic, the corresponding characteristic value of the customer characteristic is at most divided into 6 characteristic value intervals, and the IV value corresponding to each characteristic value interval is greater than 0.05. In this way, while reducing the number of characteristic value intervals, it is ensured that each characteristic value interval contributes to determining the customer type to which the customer belongs.

[0151] After obtaining the plurality of characteristic value intervals corresponding to each target customer characteristic, the embodiments of the present application can perform multi-round feature intersection processing on each characteristic value interval to obtain a target characteristic value interval combination corresponding to the customer group. The specific method of performing multi-round feature intersection processing on each characteristic value interval can refer to the related description of the foregoing embodiments, which will not be described here again. The feature intersection includes at least first-order intersection, second-order intersection, and third-order intersection, and each round of feature intersection processing is based on the feature intersection processing result of the previous round.

[0152] For example, taking the Cartesian product of the characteristic intersection method as an example, when there are two target customer characteristics, the first customer characteristic corresponds to M characteristic value intervals, and the second customer characteristic corresponds to N characteristic value intervals, in the second-order intersection of each characteristic value interval, a first characteristic value interval can be selected from the plurality of characteristic value intervals corresponding to the first customer characteristic, a second characteristic value interval can be selected from the characteristic value intervals corresponding to the second customer characteristic, and the first characteristic value interval and the second characteristic value interval can be processed by feature intersection processing, and finally M×N characteristic value interval combinations are obtained.

[0153] The embodiments of the present application can further sort each characteristic value interval according to the customer overlap rate of each characteristic value interval in the target characteristic value interval combination, so as to preferentially display the characteristic value intervals with high customer overlap rate. In addition, the embodiments of the present application can also display the first target customer proportion (i.e., the ratio of the number of first target customers to the number of all customers in the customer group) of the target characteristic value interval combination, so as to provide more information about the target characteristic value interval combination for those skilled in the art (see Table 3 below):

[0154] Table 3

[0155]

[0156] In the table, the more left the characteristic value interval is, the higher the customer overlap rate of the characteristic value interval is, which can be considered in the financial budget management, evaluation, and resource allocation of the bank.

[0157] The data processing method provided by the embodiments of the present application can divide all customers into multiple customer groups and determine a target characteristic value interval combination corresponding to each customer group. When applied to the field of bank finance, according to the target characteristic value interval combination, the performance evaluation of institutions at various levels can be more intuitively solved, the main characteristic interval values affecting the performance of institutions and products can be more intuitively given, the more important characteristic value intervals in the characteristics of customers can be better displayed, and stronger basis can be provided for budget management, evaluation, resource allocation and other work, and the decision support role can be better played.

[0158] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless explicitly stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.

[0159] Based on the same inventive concept, the embodiments of the present application also provide a data processing device for implementing the above-mentioned data processing method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more data processing device embodiments provided below can refer to the limitations of the data processing method described above, and will not be repeated here.

[0160] In one embodiment, as shown in FIG. 9, Figure 9 a data processing device is provided, comprising: a binning module 902, a feature intersection module 904, a first determination module 906, a second determination module 908, and a third determination module 910, wherein:

[0161] The binning module 902 is configured to perform binning processing on any customer characteristic to obtain a plurality of characteristic value intervals corresponding to the customer characteristic.

[0162] The feature intersection module 904 is configured to perform feature intersection processing on the first characteristic value interval combination determined in the i-1th round of feature intersection processing and the plurality of characteristic value intervals corresponding to each customer characteristic to obtain a plurality of second characteristic value interval combinations, wherein the characteristic value interval combination includes at least one characteristic value interval.

[0163] The first determining module 906 is configured to determine, for any second feature value interval combination, a customer overlap rate of a target customer group and an overall customer group on the second feature value interval combination, where the customer overlap rate is used to represent an overlapping degree of target customers in the overall customer group and target customers in the target customer group, and the target customers are users whose customer features match the second feature value interval combination, and at least one customer belonging to the same customer group corresponds to the same customer category.

[0164] The second determining module 908 is configured to determine, according to the customer overlap rate, a third feature value interval combination from the second feature value interval combination.

[0165] The third determining module 910 is configured to, in a case where i is equal to N, take the third feature value interval combination as a target feature value interval combination corresponding to the target customer group, where i and N are positive integers, and N is a total number of feature cross processing rounds.

[0166] The data processing apparatus provided by the embodiments of the present application can repeatedly perform feature cross processing on feature value interval combinations whose customer overlap rates meet requirements and a plurality of feature value intervals corresponding to each customer feature until a round of feature cross processing reaches a pre-set round, and determine a feature value interval combination whose customer overlap rate meets requirements in the last round as a customer feature value key interval of a target customer group. In each round of feature cross processing, the embodiments of the present application perform screening on feature value interval combinations, and only enable feature value interval combinations whose customer overlap rates meet requirements to enter the next round of feature cross processing, so that the feature value interval combinations obtained in each round are feature value interval combinations with higher customer overlap rates, which not only reduces the calculation amount of feature cross processing, speeds up the determination of a customer feature value key interval, but also improves the accuracy of the determination of a customer feature value key interval.

[0167] In one of the embodiments, the first determining module 906 is further configured to:

[0168] determine a first number of first target customers in the target customer group whose customer features match the second feature value interval combination, and determine a second number of second target customers in the overall customer group whose customer features match the second feature value interval combination;

[0169] obtain the customer overlap rate of the target customer group and the overall customer group on the second feature value interval combination according to the first number and the second number.

[0170] In one of the embodiments, the first determining module 906 is further configured to:

[0171] determining a ratio of the first quantity and the second quantity, and taking the ratio as a customer overlap rate of the target customer group and the whole customer group on the second feature value interval combination.

[0172] In one of the embodiments, the binning module 902 is further configured to:

[0173] For any of the customer features, determining an influence index of the customer feature on the target customer group, the influence index being used to represent an importance of the customer feature on the target customer group;

[0174] According to the influence index of each of the customer features on the target customer group, determining target customer features from the customer features;

[0175] Binning each of the target customer features to obtain a plurality of feature value intervals corresponding to the target customer features.

[0176] In one of the embodiments, the binning module 902 is further configured to:

[0177] According to the influence index of each of the customer features, ranking each of the customer features from high to low to obtain an influence index ranking sequence, and determining a preset number of customer features in the influence index ranking sequence as the target customer features;

[0178] Or, determining the customer features with the influence index greater than an influence index threshold as the target customer features.

[0179] In one of the embodiments, the apparatus further comprises:

[0180] A fourth determination module configured to determine a customer category corresponding to each of the customers according to product purchase data of the customers;

[0181] A grouping module configured to group all of the customers into a plurality of customer groups according to the customer categories, at least one customer in the customer groups corresponding to a same customer category.

[0182] In one of the embodiments, the apparatus further comprises:

[0183] A fifth determination module configured to determine, for any of the feature value intervals in the target feature value interval combination, a customer overlap rate of the target customer group and the whole customer group on the feature value interval;

[0184] A sorting module configured to sort and display each of the feature value intervals in the target feature value interval combination according to the customer overlap rate corresponding to the feature value interval.

[0185] Each module in the data processing apparatus can be implemented by software, hardware and a combination thereof in whole or in part. Each module can be embedded in or independent of a processor in the computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by the processor to perform operations corresponding to each module.

[0186] In an embodiment, a computer device, which can be a server, has an internal structure as shown in Figure 10 The computer device includes a processor, a memory and a network interface connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement a data processing method.

[0187] Those skilled in the art can understand that Figure 10 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0188] In an embodiment, a computer device includes a memory and a processor. The memory stores a computer program. The processor executes the computer program to implement the steps in each method embodiment.

[0189] In an embodiment, a computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the steps in each method embodiment.

[0190] In an embodiment, a computer program product includes a computer program. The computer program is executed by a processor to implement the steps in each method embodiment.

[0191] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.

[0192] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0193] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0194] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A data processing method, characterized in that, The method includes: For any customer feature, the customer feature is binned to obtain multiple feature value ranges corresponding to the customer feature; In the i-th round of feature cross processing, the first feature value interval combination determined in the (i-1)-th round of feature cross processing is subjected to feature cross processing with the multiple feature value intervals corresponding to each customer feature to obtain multiple second feature value interval combinations, wherein the feature value interval combination includes at least one of the feature value intervals. For any combination of the second feature value intervals, determine the customer overlap rate between the target customer group and all customer groups in the combination of the second feature value intervals. The customer overlap rate is used to characterize the degree of overlap between the target customers in the all customer groups and the target customers in the target customer group. The target customers are users whose customer features match the combination of the second feature value intervals, wherein at least one customer belonging to the same customer group corresponds to the same customer category. Based on the customer overlap rate, a third feature value interval combination is determined from the second feature value interval combination; When i equals N, the combination of the third feature value intervals is taken as the target feature value interval combination corresponding to the target customer group, where i and N are positive integers, and N is the total number of feature cross-processing rounds; The binning process for the customer characteristics yields multiple feature value ranges corresponding to the customer characteristics, including: For any of the aforementioned customer characteristics, an impact index of the customer characteristic on the target customer group is determined. The impact index is used to characterize the importance of the customer characteristic to the target customer group. The impact index is the information gain and information value of the customer characteristic. Target customer characteristics are determined from the customer characteristics based on the impact index of each of the customer characteristics on the target customer group; The target customer features are binned to obtain multiple feature value ranges corresponding to the target customer features.

2. The method according to claim 1, characterized in that, Determining the customer overlap rate between the target customer group and all customer groups within the second feature value interval includes: Determine a first number of first target customers in the target customer group whose customer characteristics match the second feature value range combination, and determine a second number of second target customers in all customer groups whose customer characteristics match the second feature value range combination; Based on the first quantity and the second quantity, the customer overlap rate between the target customer group and the entire customer group in the second feature value interval combination is obtained.

3. The method according to claim 2, characterized in that, The step of obtaining the customer overlap rate between the target customer group and the entire customer group in the second feature value interval combination based on the first quantity and the second quantity includes: Determine the ratio of the first quantity to the second quantity, and use the ratio as the customer overlap rate between the target customer group and the entire customer group in the second feature value interval combination.

4. The method according to claim 1, characterized in that, The step of determining the target customer characteristics from the customer characteristics based on the influence index of each of the customer characteristics on the target customer group includes: Based on the influence index of each customer feature, the customer features are sorted from high to low to obtain an influence index ranking sequence, and the first number of customer features in the influence index ranking sequence are determined as target customer features. Alternatively, customer characteristics whose influence index is greater than the influence index threshold can be identified as target customer characteristics.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Based on each customer's product purchase data, determine the customer category corresponding to each customer; Based on each of the customer categories, all the customers are divided into multiple customer groups, each customer group including at least one customer, and at least one customer in each customer group corresponds to the same customer category.

6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: For any feature value interval in the combination of target feature value intervals, determine the customer overlap rate between the target customer group and the all customer groups in the feature value interval; Based on the customer overlap rate corresponding to each of the feature value intervals in the target feature value interval combination, the feature value intervals in the target feature value interval combination are sorted and displayed.

7. A data processing apparatus, characterized in that, The device includes: The binning module is used to bin any customer feature to obtain multiple feature value ranges corresponding to the customer feature. The feature crossover module is used to perform feature crossover processing on the first feature value interval combination determined in the (i-1)th round of feature crossover processing and the multiple feature value intervals corresponding to each customer feature during the i-th round of feature crossover processing, so as to obtain multiple second feature value interval combinations, wherein the feature value interval combination includes at least one of the feature value intervals. The first determining module is used to determine the customer overlap rate between the target customer group and all customer groups in the second feature value interval combination for any second feature value interval combination. The customer overlap rate is used to characterize the degree of overlap between the target customers in the all customer groups and the target customers in the target customer group. The target customers are users whose customer features match the second feature value interval combination, wherein at least one customer belonging to the same customer group corresponds to the same customer category. The second determining module is used to determine a third feature value interval combination from the second feature value interval combination based on the customer overlap rate; The third determining module is used to take the third feature value interval combination as the target feature value interval combination corresponding to the target customer group when i equals N, where i and N are positive integers and N is the total number of feature cross-processing rounds; The binning module is specifically used to determine the impact index of any customer feature on the target customer group, wherein the impact index is used to characterize the importance of the customer feature to the target customer group; the impact index is the information gain and information value of the customer feature; based on the impact index of each customer feature on the target customer group, target customer features are determined from the customer features; and binning is performed on each target customer feature to obtain multiple feature value intervals corresponding to the target customer features.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Feature selection optimization method and device and readable storage medium

    CN112633414A

  • Feature crossing method and device, computer readable storage medium and program product

    CN112668046A