Information screening method and device, equipment and storage medium

By combining the LightGBM model and the K-Means model, the similarity and category scores of customer feature data are calculated, and the target customer list is screened out, which solves the problem of low customer screening efficiency in bank retail credit business and improves the work efficiency and effectiveness of marketers.

CN120563154APending Publication Date: 2025-08-29CHONGQING YUYIN FINANCIAL TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510711560.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In bank retail credit business, it is difficult to efficiently screen out existing customers suitable for secondary marketing, resulting in low work efficiency and poor marketing results for marketers.

Method used

Using a combination of LightGBM model and K-Means model, the similarity score and category score of customer feature data are calculated by training and adjusting the model, target customers are selected, and priority customers are generated by sorting according to the scores.

Benefits of technology

It improves the efficiency of customer information screening, reduces invalid lists, optimizes the work arrangements of marketers, and improves marketing effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563154A_ABST
    Figure CN120563154A_ABST
Patent Text Reader

Abstract

The invention discloses an information screening method and device, equipment and a storage medium, and relates to the technical field of computers, and the method comprises the steps: obtaining a target LightGBM model based on a preset hyper-parameter, historical customer feature data, a preset performance index and an initial LightGBM model; calculating based on the first customer feature data of the stock customer and the second customer feature data of the customer adapted to the target product to obtain a similarity score; processing the first customer feature data by using a preset K-Means model to obtain a customer type, and setting a quantity proportion of customers adapted to the target product in the stock customers as a category score of the customer type; whether the similarity score and the category score are smaller than a target similarity score threshold value and a target category score threshold value or not is judged, if not, the stock customer is set as a target customer, the target customers are sorted, and a target customer list is obtained. Therefore, the efficiency of screening the customer information can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an information screening method, device, equipment and storage medium. Background Art

[0002] As it becomes increasingly difficult for banks to attract new customers in their retail credit business, tapping into existing customers across different products, such as identifying potential customers for other mortgage loan products among existing mortgage customers and conducting secondary marketing for them, is an idea for increasing the scale of retail credit.

[0003] Currently, the process of identifying suitable customers for a product is primarily based on business rules. This involves selecting existing customers who meet the basic requirements for new retail credit products. This initial screening list is then handed over to relevant marketing personnel for targeted outreach. However, banks typically have limited marketing staff, and the resulting list still includes a large number of customers. This large number of ineffective lists consumes marketing staff's time and undermines their confidence, thus reducing actual marketing effectiveness. Furthermore, the resulting list lacks a clear order of priority, forcing marketers to follow a default sequence of marketing efforts, preventing them from effectively planning their marketing plans.

[0004] From the above, it can be seen that how to improve the efficiency of screening customer information during the information screening process is an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an information screening method, device, equipment and storage medium, which can improve the efficiency of screening customer information during the information screening process. The specific scheme is as follows:

[0006] In a first aspect, the present application provides an information screening method, comprising:

[0007] The initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and the LightGBM model to be adjusted is adjusted using preset performance indicators to obtain a target LightGBM model;

[0008] Utilizing the target LightGBM model and calculating a similarity score based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product-adapted customers, a similarity score is obtained;

[0009] Using a preset K-Means model, each of the first customer feature data is processed separately to obtain a customer type corresponding to each of the existing customers, and the proportion of the number of customers suitable for the target product among each of the existing customers corresponding to each customer type is set as a category score corresponding to the customer type;

[0010] Determine whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and target category score threshold; if both are not less than, set the corresponding existing customer as the target customer, and sort the target customers in descending order according to the similarity score to obtain a target customer list.

[0011] Optionally, before the initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and the LightGBM model to be adjusted is adjusted using preset performance indicators to obtain a target LightGBM model, the method further includes:

[0012] Obtaining initial historical customer characteristic data of existing customers, including basic information data and behavioral information data; the basic information data includes customer age, customer occupation, and customer education; the behavioral information data includes card activation behavior, mobile banking operation behavior, and banking business processing behavior;

[0013] The initial historical customer feature data is processed in sequence using a preset feature elimination rule, a WOE encoding algorithm, and a preset correlation analysis rule to obtain feature data to be standardized;

[0014] The preset centralization processing rule is used to perform a data center movement operation on the feature data to be standardized to obtain feature data to be scaled, and then the preset scaling processing rule is used to scale the feature data to be scaled to a preset scale to obtain historical customer feature data.

[0015] Optionally, the initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and the LightGBM model to be adjusted is adjusted using preset performance indicators to obtain a target LightGBM model, including:

[0016] Using a preset structured method and setting hyperparameters based on development requirements to obtain the preset hyperparameters, the target historical customer feature data is divided into a training set and a validation set for consecutive time periods based on business cycle characteristics; the preset hyperparameters include the number of leaf nodes of a tree configured based on feature dimensions, the depth of the tree configured based on preset rules for balancing model complexity and generalization ability, and a learning rate set using a preset staged decay mechanism;

[0017] The initial LightGBM model is trained using the preset hyperparameters, the training set, and the validation set to obtain a LightGBM model to be adjusted; the LightGBM model to be adjusted is adjusted using a preset performance indicator and a preset dynamic early stopping mechanism to obtain a target LightGBM model; the preset performance indicator includes AUC.

[0018] Optionally, the similarity score is calculated using the target LightGBM model based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product-adapted customers to obtain the similarity score, including:

[0019] The target LightGBM model is used to extract common factors of the first customer feature data corresponding to each existing customer and the second customer feature data corresponding to the target product adaptation customer;

[0020] Based on the common factors, the potential correlation between each of the existing customers and the target product-compatible customers is determined, and a standardized similarity score is output; the similarity score is a score determined by using a preset nonlinear mapping function and based on a distance metric in a model implicit space; the numerical value of the similarity score is positively correlated with the numerical value of the similarity between the first customer feature data and the second customer feature data.

[0021] Optionally, the using of a preset K-Means model to process each of the first customer feature data separately to obtain a customer type corresponding to each of the existing customers, and setting the proportion of the number of target product-compatible customers among each of the existing customers corresponding to each customer type as a category score corresponding to the customer type, includes:

[0022] Using a preset K-Means model and based on several customer dimensions, each of the first customer characteristic data is processed separately to obtain a customer type corresponding to each of the existing customers; the customer dimensions include consumption behavior, demographic attributes, and interaction records;

[0023] Determining target product-compatible customers from the existing customers corresponding to each customer type based on a combination of preset conditions; the combination of preset conditions includes historical purchase records of related categories, conformity to target customer profile characteristics, and responsiveness demonstrated in pilot marketing;

[0024] The ratio of the first number of customers who are suitable for the target product in each customer type to the second number of existing customers in each customer type is counted, and the ratio is set as the category score corresponding to the customer type.

[0025] Optionally, the determining whether the similarity score and the category score are respectively less than corresponding target similarity score thresholds and target category score thresholds, and if both are not less than, setting the corresponding existing customer as a target customer includes:

[0026] Processing each of the similarity scores and each of the category scores using a preset probability density function and a preset distribution curve determination function to obtain a first score distribution corresponding to the category score and a second score distribution corresponding to the similarity score;

[0027] Setting corresponding initial category score thresholds and initial similarity score thresholds based on the first score distribution and the second score distribution;

[0028] The initial category score threshold and the initial similarity score threshold are adjusted based on the marketing capabilities of the marketing personnel to obtain a target category score threshold and a target similarity score threshold corresponding to the marketing personnel; the marketing capabilities include customer conversion rate, depth of demand insight, and product matching accuracy;

[0029] Determine whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and the target category score threshold; if the similarity score and the category score are respectively less than the corresponding target similarity score threshold and the target category score threshold, set the corresponding existing customer as the target customer.

[0030] Optionally, the target customers are sorted in descending order according to the similarity scores to obtain a target customer list, including:

[0031] Using a preset asymmetric weighted algorithm and a preset sliding window technology, and arranging the target customers in descending order based on the similarity score, to obtain an initial customer list;

[0032] Determine whether there are any associated customers among the target customers in the initial customer list; if so, process the initial customer list using a preset graph neural network to identify potential communication paths among the target customers;

[0033] Determine whether there are any customers with cross-channel interaction behaviors among the target customers in the initial customer list. If there are any customers with cross-channel interaction behaviors among the target customers in the initial customer list, use a preset time decay model and a preset industry knowledge graph, and process the initial customer list based on customer needs to obtain a ranking index;

[0034] The initial customer list is flexibly adjusted based on the potential propagation path and the ranking index to obtain a target customer list.

[0035] In a second aspect, the present application provides an information screening device, comprising:

[0036] A LightGBM model determination module is used to train the initial LightGBM model based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and to adjust the LightGBM model to be adjusted using preset performance indicators to obtain a target LightGBM model;

[0037] A similarity score determination module is used to calculate a similarity score based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product adaptation customers using the target LightGBM model to obtain a similarity score;

[0038] a category score determination module, configured to process each of the first customer feature data using a preset K-Means model to obtain a customer type corresponding to each of the existing customers, and set the proportion of the number of target product-compatible customers among each of the existing customers corresponding to each customer type as a category score corresponding to the customer type;

[0039] The customer list determination module is used to determine whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and target category score threshold. If both are not less than, the corresponding existing customers are set as target customers, and the target customers are sorted in descending order according to the similarity score to obtain a target customer list.

[0040] In a third aspect, the present application provides an electronic device, comprising:

[0041] Memory, used to store computer programs;

[0042] The processor is used to execute the computer program to implement the aforementioned information screening method.

[0043] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned information screening method when executed by a processor.

[0044] As can be seen from the above, before information screening, this application needs to train the initial LightGBM model based on the preset hyperparameters and the target historical customer characteristic data to obtain the LightGBM model to be adjusted, and use the preset performance indicators to adjust the LightGBM model to obtain the target LightGBM model; use the target LightGBM model and calculate the similarity score based on the existing customer characteristic data and the customer characteristic data corresponding to the target product adaptation customer to obtain the similarity score; use the preset K-Means model to process each existing customer characteristic data separately to obtain the customer type corresponding to each existing customer characteristic data, and set the proportion of the number of target product adaptation customers in each existing customer corresponding to each customer type as the category score corresponding to the customer type; determine whether the similarity score and the category score are less than the corresponding target similarity score threshold and the target category score threshold respectively. If both are not less than, the corresponding existing customer is set as the target customer, and the target customer is sorted in descending order according to the similarity score to obtain a list of target customers.

[0045] It can be seen that this application first needs to train the initial LightGBM model based on the preset hyperparameters, target historical customer feature data and preset performance indicators to obtain the target LightGBM model; then, the target LightGBM model is used to calculate the similarity score based on the existing customer feature data and the customer feature data corresponding to the target product adaptation customer to obtain the similarity score; then, the preset K-Means model is used to process each existing customer feature data separately to obtain the customer type corresponding to each existing customer feature data, and the proportion of the number of target product adaptation customers in each existing customer corresponding to each customer type is set as the category score corresponding to the customer type; finally, it is determined whether the similarity score and the category score are less than the corresponding target similarity score threshold and the target category score threshold respectively. If both are not less than, the corresponding existing customer is set as the target customer, and each target customer is sorted in descending order according to the similarity score to obtain a target customer list. In this way, while increasing the speed of screening out invalid customer lists, the remaining customer lists can be prioritized, thereby improving the efficiency of screening customer information. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0047] Figure 1A flow chart of an information screening method disclosed in this application;

[0048] Figure 2 This is a schematic structural diagram of an information screening device disclosed in this application;

[0049] Figure 3 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0051] Currently, in the process of determining customers who are suitable for marketing products, customers are primarily screened based on business rules to select customers who meet the basic loan requirements of new retail credit products from existing customers, and then the preliminary screening list is handed over to relevant marketing personnel for targeted outreach. However, banks usually have limited marketing personnel, and there are still many customers on the list obtained after the preliminary screening. To this end, this application provides an information screening method that can improve the efficiency of screening customer information during the information screening process.

[0052] See also Figure 1 As shown, an embodiment of the present invention discloses an information screening method, comprising:

[0053] Step S11: Train the initial LightGBM model based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and adjust the LightGBM model to be adjusted using preset performance indicators to obtain a target LightGBM model.

[0054] In this embodiment, it is necessary to initialize a LightGBM classifier and set relevant hyperparameters, such as the number of leaf nodes of the tree, the depth of the tree, the learning rate, etc. The preprocessed feature data is used as input and input into the initial LightGBM model for training. In addition, the data set is divided into a training set and a validation set so that the performance of the model can be monitored during the model training process. That is, in the process of training the LightGBM model, the changes in performance indicators such as AUC (Area Under Curve, that is, the area under the ROC curve and the coordinate axis) are observed to monitor the performance of the training set and the validation set. In a specific embodiment, the embodiment of the present application uses an early stopping method to prevent overfitting, thereby stopping the model training in advance.

[0055] Specifically, the initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain the LightGBM model to be adjusted, and the LightGBM model to be adjusted is adjusted using preset performance indicators to obtain the target LightGBM model, which may include: using a preset structured method and setting hyperparameters based on development requirements to obtain preset hyperparameters, and dividing the target historical customer feature data into training sets and validation sets for continuous time periods based on business cycle characteristics; the preset hyperparameters include the number of leaf nodes of the tree configured based on the feature dimension, the depth of the tree configured based on the preset balance model complexity and generalization ability rules, and the learning rate set using a preset staged attenuation mechanism; the initial LightGBM model is trained using the preset hyperparameters, training set and validation set to obtain the LightGBM model to be adjusted; the LightGBM model to be adjusted is adjusted using preset performance indicators and a preset dynamic early stopping mechanism to obtain the target LightGBM model; the preset performance indicators include AUC.

[0056] In this embodiment, feature data corresponding to existing customers needs to be obtained, where the feature data corresponding to existing customers includes basic information data and behavioral information data. In one specific embodiment, the basic information data includes the customer's age, occupation, and educational background, and the behavioral information data includes the customer's card activation behavior, mobile banking operation behavior, and bank business transaction behavior. Feature engineering for customers can fully utilize the customer's basic feature data, such as age, occupation, and educational background, as well as their behavioral data, such as card activation behavior, mobile banking operation behavior, and bank business transaction behavior, thereby more accurately assessing the similarity between existing customers and customers for the target product. Specifically, an initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain a target LightGBM model, and the target LightGBM model is adjusted using preset performance indicators. Before obtaining the target LightGBM model, the following steps may also be performed: obtaining initial historical customer feature data for existing customers, including basic information data and behavioral information data; the basic information data includes the customer's age, occupation, and educational background; and the behavioral information data includes the customer's card activation behavior, mobile banking operation behavior, and bank business transaction behavior.

[0057] Furthermore, after obtaining the feature data corresponding to the existing customers, the embodiment of the present application needs to perform feature engineering and standardization on the feature data of the existing customers. First, the features with unique values ​​and missing values ​​greater than the preset missing value threshold are deleted from the feature data corresponding to the existing customers. Then, outliers in the feature variables are removed, and missing values ​​in the feature variables are removed or filled. Then, WOE encoding (Weight of Evidence Encoding, a coding method for binary classification problems) is performed on character-type and numeric feature variables with a low number of values. Then, correlation analysis is performed on the feature variables to screen variables with low inter-variable correlation, thereby avoiding multicollinearity problems. Finally, the feature variables are standardized. Among them, standardization includes centering and scaling. Centering is to subtract the mean of the data set from each data point and move the center of the data to zero to eliminate the influence of the original mean of the data on the analysis results, so that data of different dimensions can be compared fairly. Scaling is to scale the data to the same scale by dividing by the standard deviation of the data set. The purpose is to eliminate the dimensional differences between different features and make the data scale-invariant.

[0058] Specifically, the initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain the LightGBM model to be adjusted, and the LightGBM model to be adjusted is adjusted using preset performance indicators. Before obtaining the target LightGBM model, it may also include: using preset feature elimination rules, WOE encoding algorithm and preset correlation analysis rules to process the initial historical customer feature data in sequence to obtain feature data to be standardized; using preset centralization processing rules to perform data center movement operations on the standardized feature data to obtain feature data to be scaled, and then using preset scaling processing rules to scale the feature data to a preset scale to obtain historical customer feature data.

[0059] Step S12: Utilize the target LightGBM model and calculate a similarity score based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product adaptation customers to obtain a similarity score.

[0060] In this embodiment, the LightGBM algorithm with supervised machine learning algorithm and the K-Means clustering algorithm with unsupervised machine learning algorithm are used to provide a double guarantee for the process of screening the list, thereby improving the accuracy and reliability of the model prediction. After obtaining the target LightGBM model of the training number, the implementation of this application can use the target LightGBM model to score the existing customers, where the model score range is [0,1]. The higher the score, the higher the similarity between the existing customers and the target product adaptation customers. Specifically, the target LightGBM model is used to calculate the similarity score based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product adaptation customers to obtain the similarity score, which can include: using the target LightGBM model to extract the common factors of the first customer feature data corresponding to each existing customer and the second customer feature data corresponding to the target product adaptation customer; determining the potential correlation between each existing customer and the target product adaptation customer based on the common factors, and outputting a standardized similarity score; the similarity score is a score determined by using a preset nonlinear mapping function and based on a distance metric in the model implicit space; the numerical value of the similarity score is positively correlated with the numerical value of the similarity between the first customer feature data and the second customer feature data.

[0061] Step S13: Use the preset K-Means model to process the characteristic data of each of the first customers separately to obtain the customer type corresponding to each of the existing customers, and set the proportion of the number of customers suitable for the target product among each of the existing customers corresponding to each customer type as the category score corresponding to the customer type.

[0062] In this embodiment, the K-Means clustering model is used to process the first customer feature data to obtain the categories and category scores corresponding to the existing customers. Among them, when clustering the first customer feature data using the K-Means clustering model, it is first necessary to select a suitable K value, that is, the number of samples to be classified. In addition, the embodiment of the present application can increase the K value in sequence during the model training process until sample islands appear, and then stop the process of training the model to obtain the target model. Subsequently, the embodiment of the present application needs to process the first customer feature data with the trained model to obtain the category corresponding to each existing customer, and use the proportion of target product-adapted customers in each category sample as the category score of the category customers, and the score range is [0,1]. The higher the score, the higher the similarity between the category customers and the target product-adapted customers.

[0063] Specifically, the preset K-Means model is used to process each first customer characteristic data separately to obtain the customer type corresponding to each existing customer, and the proportion of the number of target product-compatible customers in each existing customer corresponding to each customer type is set as the category score corresponding to the customer type, which may include: using the preset K-Means model and processing each first customer characteristic data separately based on several customer dimensions to obtain the customer type corresponding to each existing customer; the customer dimensions include consumption behavior, demographic attributes and interaction records; based on a preset condition combination, the target product-compatible customers are determined from each existing customer corresponding to each customer type; the preset condition combination includes historical purchase records of related categories, compliance with target customer profile characteristics and response tendencies in pilot marketing; counting the proportion of the number of first customers of target product-compatible customers in each customer type to the number of second customers of existing customers in each customer type, and setting the proportion as the category score corresponding to the customer type.

[0064] Step S14: determine whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and target category score threshold; if both are not less than, set the corresponding existing customer as the target customer, and sort the target customers in descending order according to the similarity score to obtain a target customer list.

[0065] In this embodiment, after obtaining the similarity score and category score, the embodiment of the present application needs to screen the existing customers to obtain customers that meet both the category score threshold and the model score threshold requirements, and then sort the retained customers according to the model score. In other words, the category score distribution and similarity model score distribution of the existing customers are obtained, and the corresponding category score threshold and similarity model score threshold are determined based on the specific marketing capability limit of the marketer. Ultimately, the customers who meet both the category score threshold and the model score threshold requirements are retained and sorted according to the model score. Specifically, judging whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and the target category score threshold, and if both are not less than, setting the corresponding existing customer as the target customer, can include: using a preset probability density function and a preset distribution curve to determine the function to process each similarity score and each category score to obtain a first score distribution corresponding to the category score and a second score distribution corresponding to the similarity score; setting a corresponding initial category score threshold and an initial similarity score threshold based on the first score distribution and the second score distribution; adjusting the initial category score threshold and the initial similarity score threshold based on the marketing ability of the marketer to obtain a target category score threshold and a target similarity score threshold corresponding to the marketer; marketing ability includes customer conversion rate, depth of demand insight and product matching accuracy; judging whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and target category score threshold, and if both the similarity score and the category score are respectively less than the corresponding target similarity score threshold and target category score threshold, setting the corresponding existing customer as the target customer.

[0066] It is worth mentioning that in the process of sorting each target customer in descending order according to the similarity score to obtain the target customer list, the embodiment of the present application needs to use a preset asymmetric weighted algorithm and a preset sliding window technology, and arrange each target customer in descending order based on the similarity score to obtain an initial customer list; then, determine whether there are customers with associated relationships between the target customers in the initial customer list. If there are customers with associated relationships between the target customers, the initial customer list is processed using a preset graph neural network to identify potential communication paths among each target customer; furthermore, determine whether there are customers with cross-channel interaction behaviors between the target customers in the initial customer list. If there are customers with cross-channel interaction behaviors between the target customers in the initial customer list, the preset time decay model and the preset industry knowledge graph are used, and the initial customer list is processed based on customer needs to obtain a sorting index; finally, the initial customer list is flexibly adjusted based on the potential communication path and the sorting index to obtain a target customer list.

[0067] It can be seen that this application first needs to train the initial LightGBM model based on the preset hyperparameters, target historical customer feature data and preset performance indicators to obtain the target LightGBM model; then, use the target LightGBM model and calculate the similarity score based on the existing customer feature data and the customer feature data corresponding to the target product adaptation customer to obtain the similarity score; then, use the preset K-Means model to process each existing customer feature data separately to obtain the customer type corresponding to each existing customer feature data, and set the proportion of the number of target product adaptation customers in each existing customer corresponding to each customer type as the category score corresponding to the customer type; finally, determine whether the similarity score and category score are less than the corresponding target similarity score threshold and target category score threshold respectively. If both are not less than, the corresponding existing customer is set as the target customer, and each target customer is sorted in descending order according to the similarity score to obtain a list of target customers. In this way, the efficiency of screening customer information is improved.

[0068] Accordingly, see Figure 2 As shown, the present application also provides an information screening device, comprising:

[0069] A LightGBM model determination module 11 is configured to train an initial LightGBM model based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and to adjust the LightGBM model to be adjusted using preset performance indicators to obtain a target LightGBM model;

[0070] A similarity score determination module 12 is configured to calculate a similarity score using the target LightGBM model based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product-adapted customers, thereby obtaining a similarity score;

[0071] a category score determination module 13 for processing each of the first customer feature data using a preset K-Means model to obtain a customer type corresponding to each of the existing customers, and setting the proportion of the number of target product-compatible customers among each of the existing customers corresponding to each customer type as a category score corresponding to the customer type;

[0072] The customer list determination module 14 is used to determine whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and target category score threshold. If both are not less than, the corresponding existing customers are set as target customers, and the target customers are sorted in descending order according to the similarity score to obtain a target customer list.

[0073] As can be seen from the above, before performing information screening, the embodiment of the present application first needs to train the initial LightGBM model based on preset hyperparameters, target historical customer feature data and preset performance indicators to obtain a target LightGBM model; then, the target LightGBM model is used to calculate the similarity score based on the existing customer feature data and the customer feature data corresponding to the target product adaptation customer to obtain a similarity score; then, the preset K-Means model is used to process each existing customer feature data separately to obtain the customer type corresponding to each existing customer feature data, and the proportion of the number of target product adaptation customers in each existing customer corresponding to each customer type is set as the category score corresponding to the customer type; finally, it is determined whether the similarity score and the category score are less than the corresponding target similarity score threshold and the target category score threshold respectively. If both are not less than, the corresponding existing customer is set as the target customer, and each target customer is sorted in descending order according to the similarity score to obtain a target customer list. In this way, the efficiency of screening customer information is improved.

[0074] In some specific implementations, the information screening device may further include:

[0075] The first characteristic data acquisition unit is used to acquire initial historical customer characteristic data of existing customers, including basic information data and behavioral information data; the basic information data includes customer age, customer occupation, and customer education; the behavioral information data includes card activation behavior, mobile banking operation behavior, and banking business processing behavior;

[0076] A second feature data acquisition unit is used to process the initial historical customer feature data in sequence using a preset feature elimination rule, a WOE encoding algorithm, and a preset correlation analysis rule to obtain feature data to be standardized;

[0077] The third feature data acquisition unit is used to perform a data center movement operation on the feature data to be standardized using a preset centralization processing rule to obtain feature data to be scaled, and then scale the feature data to be scaled to a preset scale using a preset scaling processing rule to obtain historical customer feature data.

[0078] In some specific implementations, the LightGBM model determination module 11 may specifically include:

[0079] A feature data partitioning unit is configured to use a preset structured method and set hyperparameters based on development requirements to obtain the preset hyperparameters, and to partition the target historical customer feature data into a training set and a validation set for consecutive time periods based on business cycle characteristics; the preset hyperparameters include the number of leaf nodes of a tree configured based on feature dimensions, the depth of the tree configured based on preset rules for balancing model complexity and generalization ability, and a learning rate set using a preset staged decay mechanism;

[0080] A model parameter adjustment unit is used to train the initial LightGBM model using the preset hyperparameters, the training set and the validation set to obtain the LightGBM model to be adjusted; the LightGBM model to be adjusted is adjusted using the preset performance indicators and the preset dynamic early stopping mechanism to obtain the target LightGBM model; the preset performance indicators include AUC.

[0081] In some specific implementations, the similarity score determination module 12 may specifically include:

[0082] A common factor extraction unit is used to extract common factors between the first customer feature data corresponding to each existing customer and the second customer feature data corresponding to the target product adaptation customer using the target LightGBM model;

[0083] A similarity score determination subunit is used to determine the potential correlation between each of the existing customers and the target product adaptation customers based on the common factors, and output a standardized similarity score; the similarity score is a score determined by using a preset nonlinear mapping function and based on a distance metric in the model implicit space; the numerical value of the similarity score is positively correlated with the numerical value of the similarity between the first customer feature data and the second customer feature data.

[0084] In some specific implementations, the category score determination module 13 may specifically include:

[0085] a customer type determination unit, configured to process each of the first customer feature data using a preset K-Means model and based on a plurality of customer dimensions to obtain a customer type corresponding to each of the existing customers; the customer dimensions including consumption behavior, demographic attributes, and interaction records;

[0086] a target product suitable customer determination unit, configured to determine target product suitable customers from the existing customers corresponding to each customer type based on a combination of preset conditions; the preset condition combination including historical purchase records of related categories, conformity with target customer profile characteristics, and response tendency demonstrated in pilot marketing;

[0087] The category score determination subunit is used to count the ratio of the first number of customers who are suitable for the target product in each customer type to the second number of existing customers in each customer type, and set the ratio as the category score corresponding to the customer type.

[0088] In some specific implementations, the customer list determination module 14 may specifically include:

[0089] a score distribution determining unit, configured to process each of the similarity scores and each of the category scores using a preset probability density function and a preset distribution curve determination function to obtain a first score distribution corresponding to the category score and a second score distribution corresponding to the similarity score;

[0090] a score threshold determination unit, configured to set a corresponding initial category score threshold and an initial similarity score threshold based on the first score distribution and the second score distribution;

[0091] a score threshold adjustment unit, configured to adjust the initial category score threshold and the initial similarity score threshold based on the marketing capability of the marketing personnel to obtain a target category score threshold and a target similarity score threshold corresponding to the marketing personnel; the marketing capability includes customer conversion rate, depth of demand insight, and product matching accuracy;

[0092] The score judgment unit is used to judge whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and the target category score threshold; if the similarity score and the category score are respectively less than the corresponding target similarity score threshold and the target category score threshold, the corresponding existing customer is set as the target customer.

[0093] In some specific implementations, the customer list determination module 14 may specifically include:

[0094] A customer ranking unit, configured to use a preset asymmetric weighting algorithm and a preset sliding window technology to rank the target customers in descending order based on the similarity score to obtain an initial customer list;

[0095] a propagation path determination unit, configured to determine whether there are any associated customers among the target customers in the initial customer list; and if so, to process the initial customer list using a preset graph neural network to identify potential propagation paths among the target customers;

[0096] a ranking index determination unit, configured to determine whether there are any customers with cross-channel interaction behaviors among the target customers in the initial customer list; if there are any customers with cross-channel interaction behaviors among the target customers in the initial customer list, the initial customer list is processed based on customer needs using a preset time decay model and a preset industry knowledge graph to obtain a ranking index;

[0097] The customer list determination subunit is used to flexibly adjust the initial customer list based on the potential propagation path and the ranking index to obtain a target customer list.

[0098] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 3 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the information screening method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0099] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0100] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0101] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the information screening method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program capable of implementing other specific tasks.

[0102] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned disclosed information screening method. The specific steps of this method can be referred to the corresponding contents disclosed in the aforementioned embodiments and will not be repeated here.

[0103] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0104] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0105] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0106] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0107] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An information screening method, characterized in that: include: The initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and the LightGBM model to be adjusted is adjusted using preset performance indicators to obtain a target LightGBM model; Utilizing the target LightGBM model and calculating a similarity score based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product-adapted customers, a similarity score is obtained; Using a preset K-Means model, each of the first customer feature data is processed separately to obtain a customer type corresponding to each of the existing customers, and the proportion of the number of customers suitable for the target product among each of the existing customers corresponding to each customer type is set as a category score corresponding to the customer type; Determine whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and target category score threshold; if both are not less than, set the corresponding existing customer as the target customer, and sort the target customers in descending order according to the similarity score to obtain a target customer list.

2. The information screening method according to claim 1, characterized in that: The initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and the LightGBM model to be adjusted is adjusted using preset performance indicators to obtain a target LightGBM model, further comprising: Obtaining initial historical customer characteristic data of existing customers, including basic information data and behavioral information data; the basic information data includes customer age, customer occupation, and customer education; the behavioral information data includes card activation behavior, mobile banking operation behavior, and banking business processing behavior; The initial historical customer feature data is processed in sequence using a preset feature elimination rule, a WOE encoding algorithm, and a preset correlation analysis rule to obtain feature data to be standardized; The preset centralization processing rule is used to perform a data center movement operation on the feature data to be standardized to obtain feature data to be scaled, and then the preset scaling processing rule is used to scale the feature data to be scaled to a preset scale to obtain historical customer feature data.

3. The information screening method according to claim 1, wherein: The initial LightGBM model is trained based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and the LightGBM model to be adjusted is adjusted using preset performance indicators to obtain a target LightGBM model, including: Using a preset structured method and setting hyperparameters based on development requirements to obtain the preset hyperparameters, the target historical customer feature data is divided into a training set and a validation set for consecutive time periods based on business cycle characteristics; the preset hyperparameters include the number of leaf nodes of a tree configured based on feature dimensions, the depth of the tree configured based on preset rules for balancing model complexity and generalization ability, and a learning rate set using a preset staged decay mechanism; The initial LightGBM model is trained using the preset hyperparameters, the training set, and the validation set to obtain a LightGBM model to be adjusted; the LightGBM model to be adjusted is adjusted using a preset performance indicator and a preset dynamic early stopping mechanism to obtain a target LightGBM model; the preset performance indicator includes AUC.

4. The information screening method according to claim 1, wherein: The similarity score is calculated by using the target LightGBM model and based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product adaptation customers to obtain the similarity score, including: The target LightGBM model is used to extract common factors of the first customer feature data corresponding to each existing customer and the second customer feature data corresponding to the target product adaptation customer; Based on the common factors, the potential correlation between each of the existing customers and the target product-compatible customers is determined, and a standardized similarity score is output; the similarity score is a score determined by using a preset nonlinear mapping function and based on a distance metric in a model implicit space; the numerical value of the similarity score is positively correlated with the numerical value of the similarity between the first customer feature data and the second customer feature data.

5. The information screening method according to claim 1, characterized in that: The method of using a preset K-Means model to process each of the first customer feature data to obtain a customer type corresponding to each of the existing customers, and setting the proportion of the number of target product-compatible customers among each of the existing customers corresponding to each customer type as a category score corresponding to the customer type includes: Using a preset K-Means model and based on several customer dimensions, each of the first customer characteristic data is processed separately to obtain a customer type corresponding to each of the existing customers; the customer dimensions include consumption behavior, demographic attributes, and interaction records; Determining target product-compatible customers from the existing customers corresponding to each customer type based on a combination of preset conditions; the combination of preset conditions includes historical purchase records of related categories, conformity to target customer profile characteristics, and responsiveness demonstrated in pilot marketing; The ratio of the first number of customers who are suitable for the target product in each customer type to the second number of existing customers in each customer type is counted, and the ratio is set as the category score corresponding to the customer type.

6. The information screening method according to claim 1, characterized in that: The determining whether the similarity score and the category score are respectively less than corresponding target similarity score thresholds and target category score thresholds, and if both are not less than, setting the corresponding existing customer as a target customer includes: Processing each of the similarity scores and each of the category scores using a preset probability density function and a preset distribution curve determination function to obtain a first score distribution corresponding to the category score and a second score distribution corresponding to the similarity score; Setting corresponding initial category score thresholds and initial similarity score thresholds based on the first score distribution and the second score distribution; The initial category score threshold and the initial similarity score threshold are adjusted based on the marketing capabilities of the marketing personnel to obtain a target category score threshold and a target similarity score threshold corresponding to the marketing personnel; the marketing capabilities include customer conversion rate, depth of demand insight, and product matching accuracy; Determine whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and the target category score threshold; if the similarity score and the category score are respectively less than the corresponding target similarity score threshold and the target category score threshold, set the corresponding existing customer as the target customer.

7. The information screening method according to any one of claims 1 to 6, characterized in that: The target customers are sorted in descending order according to the similarity scores to obtain a target customer list, including: Using a preset asymmetric weighted algorithm and a preset sliding window technology, and arranging the target customers in descending order based on the similarity score, to obtain an initial customer list; Determine whether there are any associated customers among the target customers in the initial customer list; if so, process the initial customer list using a preset graph neural network to identify potential communication paths among the target customers; Determine whether there are any customers with cross-channel interaction behaviors among the target customers in the initial customer list. If there are any customers with cross-channel interaction behaviors among the target customers in the initial customer list, use a preset time decay model and a preset industry knowledge graph, and process the initial customer list based on customer needs to obtain a ranking index; The initial customer list is flexibly adjusted based on the potential propagation path and the ranking index to obtain a target customer list.

8. An information screening device, characterized in that: include: A LightGBM model determination module is used to train the initial LightGBM model based on preset hyperparameters and historical customer feature data to obtain a LightGBM model to be adjusted, and to adjust the LightGBM model to be adjusted using preset performance indicators to obtain a target LightGBM model; A similarity score determination module is used to calculate a similarity score based on the first customer feature data corresponding to the existing customers and the second customer feature data corresponding to the target product adaptation customers using the target LightGBM model to obtain a similarity score; a category score determination module, configured to process each of the first customer feature data using a preset K-Means model to obtain a customer type corresponding to each of the existing customers, and set the proportion of the number of target product-compatible customers among each of the existing customers corresponding to each customer type as a category score corresponding to the customer type; The customer list determination module is used to determine whether the similarity score and the category score are respectively less than the corresponding target similarity score threshold and target category score threshold. If both are not less than, the corresponding existing customers are set as target customers, and the target customers are sorted in descending order according to the similarity score to obtain a target customer list.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the information screening method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, the information screening method according to any one of claims 1 to 7 is implemented.