Method, device and storage medium for generating distinctive tags
The continuous eigenvalues are binned by the feature binning clustering model and the K-means algorithm, which solves the problem of difficult to determine the importance of labels in the risk control model, and achieves efficient and accurate risk monitoring and control.
Patent Information
- Application Number
- CN202111679757.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-12-31
AI Technical Summary
It is difficult for existing risk control models to determine the importance of labels before going online in online business, resulting in low degree of automation of risk control analysis and inability to effectively monitor and control risks.
The feature binning clustering model is used to bin multiple continuous eigenvalues in the sample, and the optimal binning strategy is determined through the K-means algorithm and simulated binning strategy, the importance of the tag customer base is quantified, and significant tags are extracted.
It realizes automatic determination of label importance without data labeling, improves the efficiency and accuracy of the risk control model, and supports risk monitoring and control.
Smart Images

Figure CN114429178B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of risk identification and monitoring control, and specifically relates to a method, device and storage medium for generating a significant label. Background Art
[0002] With the rapid development of information technology and the internet, online businesses, such as payment, registration, marketing, and credit lending, have rapidly grown and become widely used. However, with this rapid growth, some criminals often exploit various means to commit fraud within these businesses. Consequently, increasing attention and emphasis are being placed on improving the security and risk management of online businesses. Against this backdrop, risk control models (or risk models) specific to online businesses (or business scenarios) have emerged.
[0003] After a model is launched in the risk control field, it often takes a full observation period before the model's features can be ranked by importance. This is because only after a full observation period will the online data generate corresponding data labels. Only then can a model be built based on the online data and data labels, and the results of the feature importance analysis of the online data can be output to conduct risk control. In addition, the analysis object in the risk control field is generally a feature, not a labeled customer group. For example, the risk control field will generally divide the feature of the previous month's loan balance into several labeled customer groups for analysis and comparison. For example, the feature of the previous month's loan balance can be divided into four labels: less than 5,000, 5,000-10,000, 10,000-20,000, and more than 20,000. According to existing technology, even if the feature importance is ranked at this time, the importance of the label cannot be determined. This is a major bottleneck in the realization of data analysis automation in the risk control field. Summary of the Invention
[0004] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a method, device and storage medium for generating significant labels. The method is based on a feature binning clustering model, bins multiple continuous feature values in a sample, and uses cluster analysis to quantify the importance of the label customer group corresponding to the target business scenario, thereby extracting the significant label customer group and achieving accurate risk monitoring and control.
[0005] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:
[0006] A method for generating a salient tag, the method comprising:
[0007] Obtain sample data for the target business scenario;
[0008] extracting a plurality of continuous feature values from the sample data,
[0009] Using a feature binning clustering model to bin the plurality of continuous feature values, obtaining an optimal binning strategy corresponding to the plurality of continuous feature values, and importance ranking results of each tag customer group obtained by binning the continuous feature values under the optimal binning strategy;
[0010] Extract salient tags based on the importance ranking result.
[0011] According to a specific embodiment, in the above-mentioned risk label processing method, the use of a feature binning clustering model to bin the plurality of continuous feature values includes:
[0012] Traversing each of the continuous eigenvalues, performing N simulated binning for each of the continuous eigenvalues, and calculating the difference distribution value corresponding to the current simulated binning strategy in each simulated binning; determining the optimal binning strategy corresponding to the continuous eigenvalue based on the difference distribution value of the N simulated binnings, and the importance ranking results of each label customer group obtained by binning the continuous eigenvalue under the optimal binning strategy;
[0013] After the traversal is completed, the optimal binning strategy corresponding to each continuous feature value and the importance ranking result of the label customer group are obtained.
[0014] According to a specific embodiment, in the above-mentioned risk label processing method, before extracting multiple continuous feature values from the sample data,
[0015] The method further includes: obtaining a risk model corresponding to a target business scenario, and using the risk model to predict the sample data to obtain a first prediction result;
[0016] In each simulated binning, the difference distribution value corresponding to the current simulated binning strategy is calculated, including:
[0017] Assume that the current simulation binning strategy divides the continuous feature value A into M label customer groups, and performs masking processing on the M label customer groups to obtain M masked feature values corresponding to the continuous feature value A; use the M masked feature values to replace the continuous feature value A in the sample data to obtain M masked samples;
[0018] The risk model is used to predict M masked samples to obtain M masked prediction results, and the difference values between the M masked prediction results and the first prediction result are calculated respectively. The difference distribution value corresponding to the current simulated binning strategy is generated according to the difference values of the M masked prediction results.
[0019] According to a specific embodiment, in the above-mentioned risk label processing method, the importance ranking results of each label customer group obtained by binning the continuous feature value under the optimal binning strategy are determined by the following method, including:
[0020] Obtain the difference values of each label customer group obtained by binning the continuous feature value under the optimal binning strategy, and sort the customer groups from small to large based on the difference values to obtain the importance ranking results of each label customer group.
[0021] According to a specific implementation, in the above-mentioned risk label processing method, the feature binning clustering model is established based on the K-means algorithm.
[0022] According to a specific embodiment, in the above-mentioned risk label processing method, obtaining sample data in the target business scenario includes:
[0023] Create user profiles for the target business scenario, and generate sample data for the target business scenario based on the original user profile data.
[0024] According to a specific implementation, in the above-mentioned risk label processing method, the target business scenarios include: violation and prohibition monitoring business scenarios, investment and financial management business scenarios, lending business scenarios, and insurance business scenarios.
[0025] Another aspect of the present invention provides a device for generating a salient tag, the device comprising:
[0026] The acquisition module is used to obtain sample data under the target business scenario;
[0027] A parsing module, used to extract multiple continuous feature values from the sample data,
[0028] a calculation module, configured to perform binning processing on the plurality of continuous eigenvalues using a feature binning clustering model, obtain an optimal binning strategy corresponding to the plurality of continuous eigenvalues, and obtain an importance ranking result of each tag customer group obtained by binning the continuous eigenvalues under the optimal binning strategy;
[0029] A generation module extracts significant tags according to the importance ranking result.
[0030] Another aspect of the present invention provides an electronic device, comprising:
[0031] one or more processors;
[0032] A memory is used to store one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the above-mentioned risk label processing method, or the above-mentioned significant label generation method.
[0033] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-mentioned risk label processing method, or implements the above-mentioned significant label generation method.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] The method provided by the embodiment of the present invention obtains sample data in a target business scenario; extracts multiple continuous eigenvalues from the sample data, and uses a feature binning and clustering model to bin the multiple continuous eigenvalues to obtain the optimal binning strategy corresponding to the multiple continuous eigenvalues, as well as the importance ranking results of each labeled customer group obtained by binning the continuous eigenvalues under the optimal binning strategy; extracts significant labels according to the importance ranking results; accordingly, the present invention uses a feature binning and clustering model to iteratively bin the continuous eigenvalues to obtain the optimal binning strategy for the continuous eigenvalues, and at the same time objectively quantifies the importance of the labeled customer groups obtained by binning, and extracts significant labels, which can provide a basis for the generation of risk control strategies and improve the efficiency and accuracy of risk control. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Schematic diagram of a method for generating salient tags according to an embodiment of the present invention;
[0037] Figure 2 Schematic diagram of the binning process of the feature binning clustering model according to an embodiment of the present invention.
[0038] Figure 3 Schematic diagram of the architecture of a salient tag generating device according to an embodiment of the present invention;
[0039] Figure 4 Schematic diagram of an electronic device architecture according to an embodiment of the present invention DETAILED DESCRIPTION
[0040] The following describes the embodiments of the present invention through specific examples. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention.
[0041] Example 1
[0042] See also Figure 1 , Figure 1 The risk label processing method provided by an embodiment of the present invention includes the following steps:
[0043] S1: Obtain sample data under the target business scenario;
[0044] S2: extracting multiple continuous feature values from the sample data;
[0045] S3: Use a feature binning clustering model to bin the multiple continuous feature values (binning means dividing the features into multiple label customer groups), obtain the optimal binning strategy corresponding to the multiple continuous feature values, and the importance ranking results of each label customer group obtained by binning the continuous feature values under the optimal binning strategy.
[0046] In this embodiment, the continuous eigenvalues are iteratively binned through the feature binning clustering model to obtain the optimal binning strategy for the continuous eigenvalues. At the same time, the importance of the labeled customer groups obtained by binning is objectively quantified, and significant labels are extracted, which can provide a basis for the generation of risk control strategies and improve the efficiency and accuracy of risk control.
[0047] In one possible implementation, in S1 above, in a preferred embodiment of the present invention, the business sample data primarily comes from user basic profiles, user behavior information, and user information provided by third parties. The user basic profile primarily includes age, gender, education level, marital status, and region. Accordingly, target business scenarios include: illegal and prohibited business scenarios, anti-fraud business scenarios, investment and financial management business scenarios, fraud business scenarios, and credit lending business scenarios. Taking the credit lending business scenario as an example, the user's credit transaction flow over the past period, user basic profiles, and other features are captured to construct the business sample data: X = [x1, x2, ..., xi].
[0048] In one possible implementation, labeling is generally performed on continuous-value features, such as height and weight, and is generally not performed on enumerated-value features, such as gender and education level. Therefore, in S2, multiple continuous feature values are extracted from the sample data by filtering the enumerated values.
[0049] It is understood that the features in the feature matrix of the training samples described in the embodiment of the present invention refer to a column of data in the feature matrix other than the label column; the features are labeled and encoded, and the corresponding labels are different from the training sample labels described in the embodiment of the present invention (the data labels are also the labels of the samples). The labels in the labeled encoding can be each specific value of the enumerated feature, such as the enumerated feature of gender, which has two labels: male and female. It can also be each specific part of a continuous feature after binning. For example: in a risky lending scenario, the user's loan balance last month is a feature. After the user's loan balance last month is binned by label, four labels can be obtained: less than 5000, 5000-10000, 10000-20000, and greater than 20000; based on this, a new feature matrix X_new is generated.
[0050] In one possible implementation, in S3 above, the K-means algorithm is first modified, and its number of clusters is extracted as a hyperparameter in the subsequent training process. At the same time, its internal details are improved so that it meets the output style requirements under the significant labels required to be solved in this embodiment of the present invention. Then, based on this improved K-means algorithm, features are labeled. When using the K-means algorithm to label features, the number of clusters of the K-means algorithm is designed as a hyperparameter to participate in the training of the subsequent feature relative importance ranking algorithm. After the training is completed, the optimal cluster number hyperparameter can be obtained without manual adjustment.
[0051] In a possible implementation, in the above S3, the feature binning clustering model is used to perform binning processing on the plurality of continuous feature values, including: traversing each of the continuous feature values, performing N simulated binning for each of the continuous feature values, and calculating the difference distribution value corresponding to the current simulated binning strategy in each simulated binning; determining the optimal binning strategy corresponding to the continuous feature value based on the difference distribution value of the N simulated binnings, and the importance ranking result of each label customer group obtained by binning the continuous feature value under the optimal binning strategy. After the traversal is completed, the optimal binning strategy corresponding to each of the continuous feature values and the importance ranking result of the label customer group are obtained. Wherein, for each of the continuous feature values, N simulated binning is performed, including: setting the maximum number of binning N, from binning 1 time, binning 2 times... to binning N times, and performing N simulated binning for each of the continuous feature values.
[0052] In one possible implementation, in S1 above, while acquiring sample data for the target business scenario, the corresponding risk model is also acquired. For example, a logistic regression risk model model for a credit loan business scenario is acquired. The acquired risk model is then used to perform predictive analysis on the sample data, and the resulting prediction result is used as the baseline result, result_0. Using the risk model to predict multiple sample data sets, prediction results for multiple raw data samples are obtained, which serve as the basis for calculating the differences in each subsequent sample clustering bin.
[0053] In a possible implementation, in each simulated binning, calculating the difference distribution value corresponding to the current simulated binning strategy includes:
[0054] Assume that the current simulated binning strategy divides the continuous eigenvalue A into M labeled customer groups, and performs masking processing on the M labeled customer groups respectively to obtain M masked eigenvalues corresponding to the continuous eigenvalue A; use the M masked eigenvalues to replace the continuous eigenvalue A in the sample data respectively to obtain M masked samples; use the risk model to predict the M masked samples to obtain M masked prediction results, and calculate the difference values between the M masked prediction results and the first prediction result respectively, and generate the difference distribution value corresponding to the current simulated binning strategy according to the difference values of the M masked prediction results.
[0055] Among them, after the current simulation binning strategy divides the continuous feature value A into M label customer groups, A becomes: {A1, A2…A m}, then perform masking on the M labeled customer groups respectively, specifically: perform masking on A1 to transform A into: {A2…A m}、Mask A2 to change A into: {A1, A3…A m}、…to A m Perform masking to transform A into: {A1, ...A m-1}; thereby obtaining M masked feature values corresponding to A. The M masked feature values are then used to replace the continuous feature values A in the sample data, thereby obtaining M masked samples. Optionally, a Cohen's D value function or a Pearson coefficient function is used to calculate the difference between the prediction results corresponding to the M masked samples and the first prediction result.
[0056] Figure 2The specific process of binning continuous features using the feature binning clustering model described in an embodiment of the present invention is shown. When using the feature binning clustering model for binning, the number of clusters for each feature to be binned is used as a hyperparameter nj, representing the optimal number of bins for the nth feature. nj participates in the subsequent relative importance ranking algorithm, which ultimately returns the optimal number of bins and the relative importance ranking of the labels under the optimal binning strategy:
[0057] 1. Assume that the online model at this time is model, the feature matrix of the online data is X = [x1, x2, ..., xi], and i is the number of features in the feature matrix.
[0058] 2. When traversing to the feature xi to be binned, traverse different numbers of bins, with the maximum number of bins being 10.
[0059] 3. When the number of bins is nj, a custom K-means algorithm is called. This K-means algorithm improves the output of the original K-means algorithm. The output format is:
[0060] [[part1_indexs],[part2_indexs],...,[partnj_indexs]] represents the sample index of each bin segment.
[0061] 4. Then traverse each bin segment at this time. When traversing to the bin segment xi_part, randomly mask the xi_part part of the original feature xi to generate a new xi, update the X feature matrix, and then use the model model for prediction to obtain the number of bins nj and the prediction result result_i_index when the bin segment is xi_part.
[0062] 5. Calculate the statistical difference between result_i_index and the original data prediction result result_0. The difference function can be implemented using Cohen's D value function or Pearson coefficient function.
[0063] 6. Finally, after traversing the possible bins of each feature under each binning method, if the number of bins is key, and the weighted average difference of all labels is the smallest at this time, then this number of bins is the optimal number of bins.
[0064] 7. The binning result at this time is the feature labeling result, and the label difference score is the basis for evaluating the importance of the label. The smaller the difference score, the higher the label importance.
[0065] The key pseudo code of the above feature binning clustering model is as follows:
[0066]
[0067] 8. After executing the above algorithm on the feature matrix, we can obtain the optimal binning strategy and the relative importance ranking of the labels. Output the most important label customer groups as the significant label customer groups for subsequent analysis.
[0068] Another aspect of the present invention further provides a risk control method, comprising:
[0069] S201: Calculate and obtain a salient label corresponding to the target business scenario using the above-mentioned salient label generation method;
[0070] S202. Set a corresponding monitoring dimension for each high-risk tag, and generate an SQL monitoring description corresponding to each high-risk tag based on the monitoring dimension of each high-risk tag to monitor the high-risk tags in the target business scenario.
[0071] In one possible implementation, the monitoring dimensions set in S202 above include but are not limited to: the ratio of the labeled customer group to all customer groups, the rate of change of the labeled customer group in the current cycle compared to the previous cycle, and the instantaneous pulse volume of the labeled customer group (pulse volume = the ratio of the labeled customer group to all customer groups * (the number of current labeled customer groups / the number of previous labeled customer groups)^2, which is used to measure the instantaneous impulse of this customer group. Based on the monitoring dimensions, a monitoring SQL code is automatically generated through a regular expression or abstract syntax tree algorithm to continuously monitor the relevant statistical indicators of the risk customer group. Implementation examples of the two methods are as follows:
[0072]
[0073] In this embodiment, all continuous-valued features can be automatically binned using the K-means algorithm without the need for data labels. The number of clusters for each feature is then used as a hyperparameter to model the relative importance ranking algorithm for the labels. Upon completion of the ranking, the optimal number of clusters for each feature is determined. After outputting the risk labels, SQL code for label statistical monitoring can be automatically generated, enabling automated risk label monitoring deployment and accurate risk control.
[0074] Example 2
[0075] Another aspect of the present invention is Figure 3 The present invention provides a device for generating a significant tag, including:
[0076] The acquisition module is used to obtain sample data under the target business scenario;
[0077] A parsing module, used to extract multiple continuous feature values from the sample data,
[0078] a calculation module, configured to perform binning processing on the plurality of continuous eigenvalues using a feature binning clustering model, obtain an optimal binning strategy corresponding to the plurality of continuous eigenvalues, and obtain an importance ranking result of each tag customer group obtained by binning the continuous eigenvalues under the optimal binning strategy;
[0079] A generation module extracts significant tags according to the importance ranking result.
[0080] Another aspect of the present invention is Figure 4 The risk control system provided by an embodiment of the present application is shown, comprising: the aforementioned salient tag generation device, an SQL component generation module, and an execution module. The SQL component generation module is configured to set corresponding monitoring dimensions for each salient tag, generate an SQL monitoring description corresponding to each salient tag based on the monitoring dimensions of each high-risk tag, and the execution module is configured to monitor high-risk tags in the target business scenario based on the generated SQL monitoring description.
[0081] In a possible implementation, the SQL component generation module is configured to generate an SQL monitoring description corresponding to each high-risk tag through a regular expression or abstract syntax tree algorithm.
[0082] In another aspect of the present invention, an electronic device is also provided. Figure 4 As shown, the electronic device (which can be a terminal, server, computer, etc.) includes a processor, a network interface and a memory, and the processor, the network interface and the memory are interconnected, wherein the memory is used to store a computer program, and the computer program includes program instructions, and the processor is configured to call the program instructions to execute the above-mentioned significant label generation method.
[0083] In embodiments of the present invention, the processor may be an integrated circuit chip having signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0084] The methods, steps, and logic diagrams disclosed in the embodiments of the present invention can be implemented or executed. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software modules can be located in a storage medium well-established in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The processor reads the information from the storage medium and, in conjunction with its hardware, completes the steps of the aforementioned methods.
[0085] In another aspect of the present invention, a computer storage medium is provided, wherein program instructions are stored in the computer storage medium. When the program instructions are executed by at least one processor, the program instructions are used to implement the above-mentioned method for generating a salient tag.
[0086] In one possible implementation, the storage medium may be a memory, for example, a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories.
[0087] Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
[0088] Volatile memory may be random access memory (RAM), which is used as an external cache memory. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM).
[0089] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.
[0090] It should be understood that the system disclosed herein can be implemented in other ways. For example, the module division described above is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, the communication connections between modules can be through interfaces, indirect coupling or communication connections between servers or units, and can be electrical or otherwise.
[0091] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each module may exist physically separately, or two or more modules may be integrated into a single processing unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
Claims
1. A method for generating a salient tag, characterized in that: The method comprises: Obtain sample data for the target business scenario; extracting a plurality of continuous feature values from the sample data; Using a feature binning clustering model to bin the plurality of continuous feature values, obtaining an optimal binning strategy corresponding to the plurality of continuous feature values, and importance ranking results of each tag customer group obtained by binning the continuous feature values under the optimal binning strategy; Extracting significant tags according to the importance ranking result; The adopting the feature binning clustering model to perform binning processing on the plurality of continuous feature values includes: Traversing each of the continuous eigenvalues, performing N simulated binning for each of the continuous eigenvalues, and calculating the difference distribution value corresponding to the current simulated binning strategy in each simulated binning; determining the optimal binning strategy corresponding to the continuous eigenvalue based on the difference distribution value of the N simulated binnings, and the importance ranking results of each label customer group obtained by binning the continuous eigenvalue under the optimal binning strategy; After the traversal is completed, the optimal binning strategy corresponding to each continuous feature value and the importance ranking result of the label customer group are obtained; Before extracting the plurality of continuous feature values from the sample data, the method further comprises: obtaining a risk model corresponding to a target business scenario, and using the risk model to predict the sample data to obtain a first prediction result; In each simulated binning, the difference distribution value corresponding to the current simulated binning strategy is calculated, including: Assume that the current simulation binning strategy divides the continuous feature value A into M label customer groups, and performs masking processing on the M label customer groups to obtain M masked feature values corresponding to the continuous feature value A; use the M masked feature values to replace the continuous feature value A in the sample data to obtain M masked samples; The risk model is used to predict M masked samples to obtain M masked prediction results, and the difference values between the M masked prediction results and the first prediction result are calculated respectively. The difference distribution value corresponding to the current simulated binning strategy is generated according to the difference values of the M masked prediction results.
2. The method for generating a salient tag according to claim 1, wherein: The importance ranking results of each tag customer group obtained by binning the continuous feature value under the optimal binning strategy are determined by the following method, including: Obtain the difference values of each label customer group obtained by binning the continuous feature value under the optimal binning strategy, and sort the customer groups from small to large based on the difference values to obtain the importance ranking results of each label customer group.
3. The method for generating a salient tag according to claim 1 or 2, wherein: The feature binning clustering model is established based on the K-means algorithm.
4. The method for generating a salient tag according to claim 1 or 2, wherein: The obtaining of sample data in the target business scenario includes: Create user profiles for the target business scenario, and generate sample data for the target business scenario based on the original user profile data.
5. The method for generating a salient tag according to claim 4, wherein: The target business scenarios include: violation and prohibition monitoring business scenarios, investment and financial management business scenarios, lending business scenarios, and insurance business scenarios.
6. A distinctive tag generation device, configured to execute the distinctive tag generation method according to any one of claims 1 to 5, characterized in that: The device comprises: The acquisition module is used to obtain sample data under the target business scenario; A parsing module, configured to extract a plurality of continuous feature values from the sample data; a calculation module, configured to perform binning processing on the plurality of continuous eigenvalues using a feature binning clustering model, obtain an optimal binning strategy corresponding to the plurality of continuous eigenvalues, and obtain an importance ranking result of each tag customer group obtained by binning the continuous eigenvalues under the optimal binning strategy; A generation module extracts significant tags according to the importance ranking result.
7. An electronic device, characterized in that: The electronic device comprises: one or more processors; The memory is configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a significant label according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a distinctive label according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Data processing method, device and equipment and computer readable storage medium
CN112632045A
Method and device for determining model feature binning scheme
CN113792205A