Target enterprise user screening method and device, computer device and storage medium
By constructing a predictive scoring model using the LightGBM algorithm, the problem of accurately and quickly identifying the signatory in enterprise user agreement signing is solved, achieving efficient screening of target enterprise users.
Patent Information
- Application Number
- CN202211086182.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2042-09-06
AI Technical Summary
In existing technologies, it is difficult to accurately and quickly identify the intended contracting party during the enterprise user agreement signing process, resulting in low sales efficiency.
A predictive scoring model is constructed using the LightGBM algorithm. By dividing enterprise users into positive and negative samples, feature elements are obtained and filtered to construct comprehensive features. The LightGBM algorithm is then used to predict scores and identify target enterprise users.
It enables precise and rapid screening of enterprise users, helping the parties establishing the agreement to efficiently find the intended parties to sign the agreement.
Smart Images

Figure CN115409559B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and user screening technology, and in particular to a method, apparatus, computer equipment and storage medium for screening target enterprise users. Background Technology
[0002] In the era of big data, group user profiling based on big data and tagging has become an important tool to help improve user screening.
[0003] Currently, when signing agreements, it is common for a particular enterprise user to sign such an agreement, and other enterprises with similar characteristics also need to sign similar agreements. Now, for agreements between similar enterprises, sales personnel often promote them, or the parties to the agreement actively seek out other parties to sign the agreement. These methods are not conducive to the party setting up the agreement accurately and quickly finding the intended parties to sign it. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, computer equipment, and storage medium for screening target enterprise users, so as to help the parties establishing the agreement to accurately and quickly find the intended parties to the agreement.
[0005] To address the aforementioned technical problems, this application provides a method for screening target enterprise users, employing the following technical solution:
[0006] A method for screening target enterprise users includes the following steps:
[0007] Step 201: Obtain basic information of several enterprise users, and divide the several enterprise users into positive and negative samples based on the basic information. The basic information includes the enterprise name and the name of the group agreement signed by the enterprise user.
[0008] Step 202: Based on the preset form, obtain the first feature set and the second feature set corresponding to each enterprise user in the same positive sample;
[0009] Step 203: Filter the feature elements in the first feature set and the second feature set respectively to obtain feature elements that meet the preset filtering conditions, and construct the comprehensive feature corresponding to the current sample.
[0010] Step 204: Take the comprehensive features corresponding to each positive sample as input parameters and input them into the prediction scoring model based on the LightGBM algorithm to obtain the prediction score corresponding to each positive sample.
[0011] Step 205: Obtain the first feature set and the second feature set corresponding to each enterprise user in the negative sample, and execute step 203 to obtain the comprehensive features corresponding to each enterprise user in the negative sample. Input the comprehensive features into the prediction scoring model to obtain the prediction score.
[0012] Step 206: Based on the predicted score and the predicted score corresponding to each positive sample, determine the positive samples corresponding to each enterprise user in the negative samples, and complete the target enterprise user screening.
[0013] Furthermore, the step of dividing the plurality of enterprise users into positive and negative samples based on the basic information, with preset positive sample screening conditions, specifically includes:
[0014] Based on the positive sample screening criteria and the company name, select the company users that meet the positive sample screening criteria from among the plurality of company users, wherein the positive sample screening criteria include the names of the group agreements signed by the plurality of company users;
[0015] Selected enterprise users will be used as positive samples;
[0016] Enterprise users that were not selected were used as negative samples.
[0017] Furthermore, the step of selecting enterprise users from the plurality of enterprise users that meet the positive sample screening criteria based on the positive sample screening criteria and the enterprise name specifically includes:
[0018] Based on the group agreement name and the enterprise name, identify the enterprise users among the plurality of enterprise users who have signed the same group agreement;
[0019] Enterprise users who have signed the same group agreement are divided into a positive sample, and N positive samples are generated, where N is a positive integer and N represents the type of group agreement signed by the enterprise users.
[0020] Enterprises that have not signed any group agreements among the aforementioned enterprise users are considered negative samples.
[0021] Furthermore, the step of obtaining the first feature set and the second feature set corresponding to each enterprise user in the same positive sample based on a preset form specifically includes:
[0022] Based on the form, the profile attributes of the enterprise user are obtained, and a first tag set is constructed. The profile attributes of the enterprise user include enterprise operation information, scale information, product information, group agreement information and non-group agreement information.
[0023] Based on the form, the profile attributes of the decision-makers of the enterprise users are obtained, and a second tag set is constructed. The profile attributes of the decision-makers of the enterprise users include the basic information, wealth information, and personal agreement information of the decision-makers.
[0024] Furthermore, the step of filtering the feature elements in the first feature set and the second feature set respectively to obtain feature elements that meet the preset filtering conditions and constructing the comprehensive feature corresponding to the current sample specifically includes:
[0025] Feature binning is performed on the feature elements in the first feature set and the second feature set respectively to obtain the feature factor IV value corresponding to each feature element;
[0026] Based on the preset feature factor IV value threshold and the feature factor IV value corresponding to each feature element, data cleaning is performed on the feature elements in the first feature set and the second feature set to obtain feature elements that meet the initial selection conditions.
[0027] Correlation analysis and principal component analysis are used to perform feature analysis on the feature elements that meet the initial selection criteria, obtain the feature elements that meet the final selection criteria, and construct the comprehensive features corresponding to the current sample.
[0028] Furthermore, before the step of performing feature binning on the feature elements of the first feature set and the second feature set respectively to obtain the feature factor IV value corresponding to each feature element, the method further includes:
[0029] According to the preset assignment rules, the feature elements in the first feature set and the second feature set are numerically assigned values.
[0030] The extreme values in the numerical assignment results are smoothly replaced according to a preset replacement rule, wherein the extreme values include maximum values and minimum values.
[0031] Furthermore, the step of cleaning the feature elements in the first and second feature sets based on a preset feature factor IV value threshold and the feature factor IV value corresponding to each feature element to obtain feature elements that meet the initial selection conditions specifically includes:
[0032] Determine whether the feature factor IV values corresponding to the feature elements in the first feature set and the second feature set meet the corresponding feature factor IV value thresholds.
[0033] If the conditions are met, the feature element meets the initial selection criteria and is retained.
[0034] If the conditions are not met, the feature element does not meet the initial selection criteria and is therefore eliminated.
[0035] Furthermore, the step of using correlation analysis and principal component analysis to perform feature analysis on the feature elements that meet the initial selection criteria, obtaining feature elements that meet the final selection criteria, and constructing the comprehensive features corresponding to the current positive sample specifically includes:
[0036] Using the aforementioned correlation analysis method, correlation analysis is performed on the feature elements that meet the initial selection criteria to obtain the correlation characterization value of each feature element.
[0037] Based on the correlation characterization value and the preset correlation threshold, feature elements that are greater than the correlation threshold are filtered and removed to obtain feature elements after secondary filtering.
[0038] Principal component analysis was used to perform principal component analysis on the feature elements after the secondary screening to obtain the weight coefficients corresponding to the feature elements.
[0039] Based on the weight coefficients corresponding to the feature elements and the preset weight coefficient threshold, feature elements with weight coefficients greater than the weight coefficient threshold are selected, and the feature elements are combined according to the preset combination rules to construct a comprehensive feature.
[0040] Furthermore, the step of determining the positive samples corresponding to each enterprise user in the negative samples based on the predicted score and the predicted score corresponding to each positive sample, thereby completing the target enterprise user screening, specifically includes:
[0041] The predicted scores for each enterprise user in the negative samples are compared sequentially with the predicted scores for each positive sample.
[0042] Based on the comparison results, the target positive samples corresponding to each enterprise user in the negative samples were selected.
[0043] The group-type protocol name corresponding to the target positive sample is recommended to the enterprise users corresponding to the target positive sample, thus completing the target enterprise user screening.
[0044] To address the aforementioned technical problems, this application also provides a target enterprise user screening device, which employs the following technical solution:
[0045] A target enterprise user screening device, comprising:
[0046] The positive and negative sample segmentation module is used to obtain basic information of several enterprise users and perform positive and negative sample segmentation on the several enterprise users based on the basic information. The basic information includes the enterprise name and the name of the group agreement signed by the enterprise user.
[0047] The feature set acquisition module is used to acquire the first feature set and the second feature set corresponding to each enterprise user in the same positive sample based on a preset form.
[0048] The data cleaning module is used to filter the feature elements in the first feature set and the second feature set respectively, obtain the feature elements that meet the preset filtering conditions, and construct the comprehensive feature corresponding to the current sample.
[0049] The positive sample prediction module is used to take the comprehensive features corresponding to each positive sample as input parameters and input them into the prediction scoring model based on the LightGBM algorithm to obtain the prediction score corresponding to each positive sample.
[0050] The negative sample prediction module is used to obtain the first feature set and the second feature set corresponding to each enterprise user in the negative sample, and to execute the data cleaning module to obtain the comprehensive features corresponding to each enterprise user in the negative sample. The comprehensive features are then input into the prediction scoring model to obtain the prediction score.
[0051] The target enterprise user screening module is used to determine the positive samples corresponding to each enterprise user in the negative samples based on the predicted score and the predicted score corresponding to each positive sample, thereby completing the target enterprise user screening.
[0052] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution:
[0053] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the target enterprise user screening method described above.
[0054] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below:
[0055] A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the target enterprise user screening method described above.
[0056] Compared with the prior art, the embodiments of this application have the following main advantages:
[0057] The target enterprise user screening method described in this application involves: dividing a number of enterprise users into positive and negative samples; obtaining the feature sets corresponding to each enterprise user in the same positive sample; obtaining feature elements that meet the initial selection criteria; obtaining feature elements that meet the final selection criteria to construct the comprehensive features corresponding to the current sample; sequentially using the comprehensive features corresponding to each positive sample as input parameters to a prediction scoring model based on the LightGBM algorithm to obtain the predicted score for each positive sample; obtaining the comprehensive features of the negative samples, inputting the comprehensive features into the prediction scoring model to obtain the predicted score; comparing the results to determine the positive samples corresponding to each enterprise user in the negative samples, thus completing the target enterprise user screening. This facilitates the agreement initiator in accurately and quickly finding the intended agreement signatories. Attached Figure Description
[0058] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0059] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;
[0060] Figure 2 A flowchart of an embodiment of the target enterprise user screening method according to this application;
[0061] Figure 3 yes Figure 2 A flowchart of a specific implementation of step 203 shown;
[0062] Figure 4 A schematic diagram of a structural diagram of an embodiment of the target enterprise user screening device according to this application;
[0063] Figure 5 yes Figure 4 A schematic diagram of a specific embodiment of 403 is shown;
[0064] Figure 6 A schematic diagram of the structure of an embodiment of the computer device according to this application. Detailed Implementation
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0066] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0067] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0068] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0069] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0070] Terminal devices 101, 102, and 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc.
[0071] Server 105 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal devices 101, 102, and 103.
[0072] It should be noted that the target enterprise user screening method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the target enterprise user screening device is generally set in the server / terminal device.
[0073] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0074] Continue to refer to Figure 2 A flowchart illustrating an embodiment of the target enterprise user screening method according to this application is shown. The target enterprise user screening method includes the following steps:
[0075] Step 201: Obtain basic information of several enterprise users, and divide the several enterprise users into positive and negative samples based on the basic information. The basic information includes the enterprise name and the name of the group agreement signed by the enterprise user.
[0076] In this embodiment, a positive sample screening condition is preset. The step of dividing the plurality of enterprise users into positive and negative samples based on the basic information specifically includes: screening out enterprise users that meet the positive sample screening condition based on the positive sample screening condition and the enterprise name, wherein the positive sample screening condition includes the name of the group agreement signed by the plurality of enterprise users; taking the screened enterprise users as positive samples; and taking the enterprise users that are not screened as negative samples.
[0077] Taking group insurance policies as an example, suppose that among the company's corporate clients, some clients have signed multiple unspecified group insurance policies as a company, while some clients have not yet signed group insurance policies. In this case, the clients who have signed group insurance policies and the clients who have not signed group insurance policies are divided into positive samples and negative samples.
[0078] By dividing samples into positive and negative samples, it is easy to predict the appropriate group-type agreements for negative samples based on positive samples, helping the agreement setter to accurately and quickly find target enterprise users.
[0079] In this embodiment, the step of selecting enterprise users from the plurality of enterprise users that meet the positive sample screening criteria based on the positive sample screening criteria and the enterprise name specifically includes: identifying enterprise users from the plurality of enterprise users who have signed the same group agreement based on the group agreement name and the enterprise name; dividing the enterprise users who have signed the same group agreement into a positive sample, generating N positive samples, where N is a positive integer and N represents the type of group agreement signed by the plurality of enterprise users; and treating enterprise users from the plurality of enterprise users who have not signed any group agreement as negative samples.
[0080] Taking group insurance policies as an example, suppose that among the company's corporate clients, there are 100 clients who have signed group insurance policies as a company. Among them, 10 clients signed Group Insurance Policy A, 20 clients signed Group Insurance Policy B, 30 clients signed Group Insurance Policy C, and 40 clients signed Group Insurance Policy D. In this case, N is 4, representing 4 different types of group insurance policies. Therefore, 4 positive samples need to be constructed based on the different types of insurance policies.
[0081] By constructing different types of positive samples based on the names of group-type protocols, it is easy to predict the group-type protocols suitable for negative samples based on different types of positive samples, helping protocol setters to accurately and quickly find target enterprise users.
[0082] Step 202: Based on the preset form, obtain the first feature set and the second feature set corresponding to each enterprise user in the same positive sample.
[0083] In this embodiment, the step of obtaining the first feature set and the second feature set corresponding to each enterprise user in the same positive sample based on a preset form specifically includes: obtaining the profile attributes of the enterprise users based on the form and constructing a first tag set, wherein the profile attributes of the enterprise users include enterprise operation information, scale information, product information, group agreement information and non-group agreement information; obtaining the profile attributes of the decision-makers of the enterprise users based on the form and constructing a second tag set, wherein the profile attributes of the decision-makers of the enterprise users include the basic information, wealth information and personal agreement information of the decision-makers.
[0084] In this embodiment, the preset form includes enterprise operation information, scale information, product information, group agreement information and non-group agreement information, as well as the decision-maker's basic information, wealth information and personal agreement information. The source can be extracted through a big data platform or historical relevant information reserved by enterprise users.
[0085] In this embodiment, the decision-maker refers to the entity that can legally represent the enterprise in signing agreements with external parties, including the enterprise's legal representative and agents with corresponding civil capacity who are authorized by the enterprise.
[0086] Taking Company E as an example, assume that Company E's operating information includes operating duration and industry; its size information includes registered capital and whether it is a Fortune 500 company; its product information includes product demand and customer value; its group agreement information includes auto insurance underwriting information signed by enterprises; and its non-group agreement information includes insurance policies not signed by enterprises. The basic information of Company E's decision-makers includes age and position; its wealth information includes assets and value; and its personal agreement information includes personal auto insurance underwriting information. In this case, we can construct the company feature set (the first feature set) and the decision-maker feature set (the second feature set).
[0087] By constructing a first feature set and a second feature set, enterprises are associated with their owners, i.e., decision-makers. Through joint predictive analysis by enterprises and relevant individuals, the reliability of the predictive analysis results is improved.
[0088] Step 203: Filter the feature elements in the first feature set and the second feature set respectively to obtain feature elements that meet the preset filtering conditions, and construct the comprehensive feature corresponding to the current sample.
[0089] In this embodiment, the step of filtering the feature elements in the first feature set and the second feature set respectively to obtain feature elements that meet the preset filtering conditions and constructing the comprehensive features corresponding to the current sample specifically includes: performing feature binning on the feature elements in the first feature set and the second feature set respectively to obtain the feature factor IV value corresponding to each feature element; performing data cleaning on the feature elements in the first feature set and the second feature set based on the preset feature factor IV value threshold and the feature factor IV value corresponding to each feature element to obtain feature elements that meet the initial selection conditions; and using correlation analysis and principal component analysis to perform feature analysis on the feature elements that meet the initial selection conditions to obtain feature elements that meet the final selection conditions and construct the comprehensive features corresponding to the current sample.
[0090] Continue to refer to Figure 3 , Figure 3 yes Figure 2 A flowchart of a specific implementation of step 203 shown includes the following steps:
[0091] Step 301: Perform feature binning on the feature elements in the first feature set and the second feature set respectively, and obtain the feature factor IV value corresponding to each feature element.
[0092] In this embodiment, the feature binning is essentially a data discretization process performed on the feature elements of the first feature set and the second feature set, converting continuous data into non-continuous data.
[0093] Feature binning avoids having too many original categories, which could result in some categories having insufficient sample sizes. For numerical variables, this could lead to extreme values, affecting the robustness of the analysis. It improves stability, making the analysis results or model predictions more robust.
[0094] In this embodiment, the IV value of the feature factor is the contribution of the feature factor to the model prediction. The larger the IV value, the greater the contribution of the corresponding feature factor to the model prediction, and vice versa.
[0095] In this embodiment, before the step of performing feature binning on the feature elements of the first feature set and the second feature set respectively to obtain the feature factor IV value corresponding to each feature element, the method further includes: performing numerical assignment processing on the feature elements of the first feature set and the second feature set according to a preset assignment rule; and performing smooth replacement on the extreme values in the numerical assignment processing result according to a preset replacement rule, wherein the extreme values include maximum values and minimum values.
[0096] In this embodiment, numerical values are assigned and numerical smoothing is performed on each feature element through assignment rules and replacement rules to avoid overfitting of the prediction results of the prediction model.
[0097] Step 302: Based on the preset feature factor IV value threshold and the feature factor IV value corresponding to each feature element, perform data cleaning on the feature elements in the first feature set and the second feature set to obtain feature elements that meet the initial selection conditions.
[0098] In this embodiment, the step of cleaning the feature elements in the first feature set and the second feature set based on the preset feature factor IV value threshold and the feature factor IV value corresponding to each feature element to obtain feature elements that meet the initial selection conditions specifically includes: determining whether the feature factor IV value corresponding to the feature element in the first feature set and the second feature set meets the corresponding feature factor IV value threshold; if it does, the feature element meets the initial selection conditions and is retained; if it does not, the feature element does not meet the initial selection conditions and is removed.
[0099] Through initial screening, features that contribute little to the model's predictions are removed to avoid having too many model parameters.
[0100] Step 303: Use correlation analysis and principal component analysis to perform feature analysis on the feature elements that meet the initial selection criteria, obtain the feature elements that meet the final selection criteria, and construct the comprehensive features corresponding to the current sample.
[0101] In this embodiment, the step of using correlation analysis and principal component analysis to perform feature analysis on the feature elements that meet the initial selection criteria, obtain feature elements that meet the final selection criteria, and construct the comprehensive feature corresponding to the current positive sample specifically includes: using the correlation analysis method to perform correlation analysis on the feature elements that meet the initial selection criteria, and obtaining the correlation characterization value of each feature element; based on the correlation characterization value and a preset correlation threshold, filtering and removing feature elements that are greater than the correlation threshold to obtain feature elements after secondary filtering; using principal component analysis to perform principal component analysis on the feature elements after secondary filtering to obtain the weight coefficients corresponding to the feature elements; based on the weight coefficients corresponding to the feature elements and a preset weight coefficient threshold, filtering out feature elements that are greater than the weight coefficient threshold, and combining the feature elements according to a preset combination rule to construct the comprehensive feature.
[0102] In this embodiment, the preset combination rules are customized by the programmer in combination with the business scenario. For example, the feature elements selected by principal component analysis include historical car insurance and car insurance quotes in the past three years, which can be combined into a comprehensive feature: the ratio of car insurance quotes in the past three years to historical car insurance.
[0103] The purpose of using correlation analysis is to eliminate features corresponding to strongly correlated indicators, thereby further ensuring the reliability of the predicted score. Principal component analysis is then used to further screen out the final features, and the final features are weighted and combined to determine the final parameters of the model, thus ensuring the reliability of the predicted score once again.
[0104] Step 204: Take the comprehensive features corresponding to each positive sample as input parameters and input them into the prediction scoring model based on the LightGBM algorithm to obtain the prediction score corresponding to each positive sample.
[0105] The prediction scoring model based on the LightGBM algorithm supports highly efficient parallel training, with faster iteration speed, lower memory consumption, and better accuracy. It supports distributed processing and can quickly handle massive amounts of data. It can train multiple positive samples in parallel to obtain prediction scores without waiting for each positive sample to be predicted and scored sequentially, thus ensuring prediction efficiency.
[0106] Step 205: Obtain the first feature set and the second feature set corresponding to each enterprise user in the negative sample, and execute step 203 to obtain the comprehensive features corresponding to each enterprise user in the negative sample. Input the comprehensive features into the prediction scoring model to obtain the prediction score.
[0107] The prediction scoring model based on the LightGBM algorithm supports highly efficient parallel training, with faster iteration speed, lower memory consumption, and better accuracy. It supports distributed processing and can quickly handle massive amounts of data. It can train multiple enterprise users in the negative sample in parallel to obtain the predicted score, without having to wait for each individual enterprise user to make a prediction score in turn, thus ensuring the prediction efficiency for negative samples.
[0108] Step 206: Based on the predicted score and the predicted score corresponding to each positive sample, determine the positive samples corresponding to each enterprise user in the negative samples, and complete the target enterprise user screening.
[0109] In this embodiment, the step of determining the positive samples corresponding to each enterprise user in the negative samples based on the predicted scores and the predicted scores corresponding to each positive sample, and completing the target enterprise user screening, specifically includes: comparing the predicted scores corresponding to each enterprise user in the negative samples with the predicted scores corresponding to each positive sample in turn; screening out the target positive samples corresponding to each enterprise user in the negative samples according to the comparison results; and recommending the group-type protocol names corresponding to the target positive samples to the enterprise users corresponding to the target positive samples, thereby completing the target enterprise user screening.
[0110] This application involves dividing a number of enterprise users into positive and negative samples; obtaining the feature sets corresponding to each enterprise user in the same positive sample; obtaining feature elements that meet the initial selection criteria; obtaining feature elements that meet the final selection criteria, and constructing the comprehensive features corresponding to the current sample; sequentially using the comprehensive features corresponding to each positive sample as input parameters to a prediction scoring model based on the LightGBM algorithm to obtain the predicted score for each positive sample; obtaining the comprehensive features of the negative samples, inputting the comprehensive features into the prediction scoring model to obtain the predicted score; comparing these, and determining the positive samples corresponding to each enterprise user in the negative samples, thus completing the target enterprise user screening. This facilitates the agreement initiator in accurately and quickly finding the intended agreement signatories.
[0111] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0112] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0113] In this embodiment of the application, a predictive scoring model is constructed to obtain scores for positive and negative samples respectively. Then, by comparison, the intention group type agreement that matches each enterprise user in the negative sample is determined, which helps the agreement initiator to accurately and quickly predict the parties who intend to sign the agreement.
[0114] Further reference Figure 4 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a target enterprise user screening device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0115] like Figure 4 As shown, the target enterprise user screening device 400 described in this embodiment includes: a positive and negative sample partitioning module 401, a feature set acquisition module 402, a data cleaning module 403, a positive sample prediction module 404, a negative sample prediction module 405, and a target enterprise user screening module 406. Wherein:
[0116] The positive and negative sample segmentation module 401 is used to obtain basic information of several enterprise users and perform positive and negative sample segmentation on the several enterprise users based on the basic information. The basic information includes the enterprise name and the name of the group agreement signed by the enterprise user.
[0117] The feature set acquisition module 402 is used to acquire the first feature set and the second feature set corresponding to each enterprise user in the same positive sample based on a preset form.
[0118] The data cleaning module 403 is used to filter the feature elements in the first feature set and the second feature set respectively to obtain feature elements that meet the preset filtering conditions and construct the comprehensive features corresponding to the current sample.
[0119] The positive sample prediction module 404 is used to sequentially input the comprehensive features corresponding to each positive sample as input parameters into the prediction scoring model based on the LightGBM algorithm to obtain the prediction score corresponding to each positive sample.
[0120] The negative sample prediction module 405 is used to obtain the first feature set and the second feature set corresponding to each enterprise user in the negative sample, and to execute the data cleaning module to obtain the comprehensive features corresponding to each enterprise user in the negative sample, and input the comprehensive features into the prediction scoring model to obtain the prediction score.
[0121] The target enterprise user screening module 406 is used to determine the positive samples corresponding to each enterprise user in the negative samples based on the predicted score and the predicted score corresponding to each positive sample, thereby completing the target enterprise user screening.
[0122] This application involves dividing a number of enterprise users into positive and negative samples; obtaining the feature sets corresponding to each enterprise user in the same positive sample; obtaining feature elements that meet the initial selection criteria; obtaining feature elements that meet the final selection criteria, and constructing the comprehensive features corresponding to the current sample; sequentially using the comprehensive features corresponding to each positive sample as input parameters to a prediction scoring model based on the LightGBM algorithm to obtain the predicted score for each positive sample; obtaining the comprehensive features of the negative samples, inputting the comprehensive features into the prediction scoring model to obtain the predicted score; comparing these, and determining the positive samples corresponding to each enterprise user in the negative samples, thus completing the target enterprise user screening. This facilitates the agreement initiator in accurately and quickly finding the intended agreement signatories.
[0123] Continue to refer to Figure 5 , Figure 5 yes Figure 4 The diagram shows a specific embodiment of the data cleaning module 403. The data cleaning module 403 includes a feature factor IV value acquisition submodule 4031, a preliminary screening submodule 4032, and a comprehensive feature acquisition submodule 4033.
[0124] The feature factor IV value acquisition submodule 4031 is used to perform feature binning processing on the feature elements in the first feature set and the second feature set respectively, and obtain the feature factor IV value corresponding to each feature element.
[0125] The initial screening submodule 4032 is used to perform data cleaning on the feature elements in the first feature set and the second feature set based on a preset feature factor IV value threshold and the feature factor IV value corresponding to each feature element, so as to obtain feature elements that meet the initial screening conditions.
[0126] The comprehensive feature acquisition submodule 4033 is used to perform feature analysis on the feature elements that meet the initial selection conditions using correlation analysis and principal component analysis, to obtain the feature elements that meet the final selection conditions, and to construct the comprehensive features corresponding to the current sample.
[0127] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0128] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0129] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 6 , Figure 6 This is a basic structural block diagram of the computer device in this embodiment.
[0130] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are interconnected via a system bus. It should be noted that only the computer device 6 with components 61-63 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0131] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0132] The memory 61 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 6. Of course, the memory 61 may include both the internal storage unit and its external storage device of the computer device 6. In this embodiment, the memory 61 is typically used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions for target enterprise user screening methods. In addition, the memory 61 can also be used to temporarily store various types of data that have been output or will be output.
[0133] In some embodiments, the processor 62 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 62 is typically used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to execute computer-readable instructions stored in the memory 61 or to process data, for example, to execute computer-readable instructions for the target enterprise user screening method.
[0134] The network interface 63 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 6 and other electronic devices.
[0135] The computer device proposed in this embodiment belongs to the field of customer screening technology. This application divides several enterprise users into positive and negative samples; obtains the feature sets corresponding to each enterprise user in the same positive sample; obtains feature elements that meet the initial selection criteria; obtains feature elements that meet the final selection criteria, and constructs the comprehensive features corresponding to the current sample; sequentially, it uses the comprehensive features corresponding to each positive sample as input parameters and inputs them into a prediction scoring model based on the LightGBM algorithm to obtain the predicted score for each positive sample; it obtains the comprehensive features of the negative samples, inputs the comprehensive features into the prediction scoring model to obtain the predicted score; and compares these to determine the positive samples corresponding to each enterprise user in the negative samples, thus completing the target enterprise user screening. This facilitates the agreement initiator in accurately and quickly finding the intended agreement signatories.
[0136] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to perform the steps of the target enterprise user screening method described above.
[0137] The computer-readable storage medium proposed in this embodiment belongs to the field of customer screening technology. This application divides several enterprise users into positive and negative samples; obtains the feature sets corresponding to each enterprise user in the same positive sample; obtains feature elements that meet the initial selection criteria; obtains feature elements that meet the final selection criteria, constructing a comprehensive feature corresponding to the current sample; sequentially, it inputs the comprehensive feature corresponding to each positive sample as input parameters into a prediction scoring model based on the LightGBM algorithm to obtain the predicted score for each positive sample; it obtains the comprehensive feature of the negative samples, inputs the comprehensive feature into the prediction scoring model to obtain the predicted score; and compares these to determine the positive samples corresponding to each enterprise user in the negative samples, thus completing the target enterprise user screening. This facilitates the agreement initiator in accurately and quickly finding the intended agreement signatories.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0139] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A method for screening target enterprise users, characterized in that, The method comprises the following steps: Step 201, obtaining basic information of a plurality of enterprise users, and performing positive and negative sample division on the plurality of enterprise users according to the basic information, wherein the basic information comprises an enterprise name and a group agreement name signed by the enterprise users; Step 202, obtaining a first feature set and a second feature set corresponding to each enterprise user in the same positive sample based on a preset form; Step 203, performing screening processing on feature elements in the first feature set and the second feature set respectively to obtain feature elements meeting a preset screening condition, and constructing comprehensive features corresponding to the current sample, wherein the step of performing screening processing on the feature elements in the first feature set and the second feature set respectively to obtain the feature elements meeting the preset screening condition and constructing the comprehensive features corresponding to the current sample comprises: performing feature binning processing on the feature elements in the first feature set and the second feature set respectively to obtain feature factor IV values corresponding to the feature elements, wherein the feature factor IV value is a contribution degree of a feature factor to model prediction, the greater the IV value is, the greater the contribution degree of the corresponding feature factor to model prediction is, and vice versa; performing data cleaning on the feature elements in the first feature set and the second feature set based on a preset feature factor IV value threshold and the feature factor IV values corresponding to the feature elements to obtain feature elements meeting a preliminary selection condition; performing feature analysis on the feature elements meeting the preliminary selection condition by using a correlation analysis method and a principal component analysis method to obtain feature elements meeting a final selection condition, and constructing comprehensive features corresponding to the current sample; Step 204, inputting the comprehensive features corresponding to each positive sample as input parameters into a prediction score model based on a LightGBM algorithm in sequence to obtain prediction scores corresponding to each positive sample; Step 205, obtaining a first feature set and a second feature set corresponding to each enterprise user in a negative sample, performing step 203 to obtain comprehensive features corresponding to each enterprise user in the negative sample, inputting the comprehensive features into the prediction score model, and obtaining prediction scores; Step 206, determining positive samples corresponding to each enterprise user in the negative sample based on the prediction scores and the prediction scores corresponding to each positive sample, and completing target enterprise user screening.
2. The targeted enterprise user screening method of claim 1, wherein, A preset positive sample screening condition, and the step of performing positive and negative sample division on the plurality of enterprise users according to the basic information comprises: screening enterprise users meeting the positive sample screening condition from the plurality of enterprise users based on the positive sample screening condition and the enterprise name, wherein the positive sample screening condition comprises a group agreement name signed by the plurality of enterprise users; screening out the enterprise users as positive samples; screening out the enterprise users as negative samples.
3. The targeted enterprise user screening method of claim 2, wherein, The step of screening enterprise users meeting the positive sample screening condition from the plurality of enterprise users based on the positive sample screening condition and the enterprise name comprises: Identify, based on the group-type agreement name and the enterprise name, enterprise users in the plurality of enterprise users who have signed the same group-type agreement; Divide the enterprise users who have signed the same group-type agreement into one positive sample, and generate N positive samples, where N is a positive integer, and N represents the number of types of group-type agreements signed by the plurality of enterprise users; Take enterprise users in the plurality of enterprise users who have not signed any group-type agreement as negative samples.
4. The targeted enterprise user screening method of claim 1, wherein, The step of obtaining, based on the preset form, a first feature set and a second feature set corresponding to each enterprise user in the same positive sample, specifically includes: Obtain, based on the form, portrait attributes of the enterprise user, and construct a first label set, where the portrait attributes of the enterprise user include enterprise operation information, scale information, product information, group-type agreement information, and non-group-type agreement information; Obtain, based on the form, portrait attributes of a decision maker of the enterprise user, and construct a second label set, where the portrait attributes of the decision maker of the enterprise user include basic information, wealth information, and personal agreement information of the decision maker.
5. The targeted enterprise user screening method of claim 1, wherein, Before the step of performing feature binning processing on feature elements in the first feature set and the second feature set respectively, and obtaining feature factor IV values corresponding to each feature element, the method further includes: According to a preset assignment rule, perform numerical assignment processing on the feature elements in the first feature set and the second feature set; Smoothly replace extreme values in the numerical assignment processing result according to a preset replacement rule, where the extreme values include maximum values and minimum values.
6. The targeted enterprise user screening method of claim 1, wherein, The step of performing data cleaning on the feature elements in the first feature set and the second feature set based on a preset feature factor IV value threshold and the feature factor IV values corresponding to each feature element, and obtaining feature elements meeting a preliminary selection condition, specifically includes: Determine whether the feature factor IV values corresponding to the feature elements in the first feature set and the second feature set meet corresponding feature factor IV value thresholds; If yes, the feature elements meet the preliminary selection condition and are retained; If no, the feature elements do not meet the preliminary selection condition and are removed.
7. The targeted enterprise user screening method of claim 1, wherein, The step of performing feature analysis on the feature elements meeting the preliminary selection condition using a correlation analysis method and a principal component analysis method, and obtaining feature elements meeting a final selection condition and constructing a comprehensive feature corresponding to the current positive sample, specifically includes: Perform correlation analysis on the feature elements meeting the preliminary selection condition using the correlation analysis method, and obtain correlation representation values of each feature element; Based on the correlation representation values and a preset correlation threshold, filter and remove feature elements greater than the correlation threshold, and obtain feature elements after secondary filtering; Perform principal component analysis on the feature elements after secondary filtering using the principal component analysis method, and obtain weight coefficients corresponding to the feature elements; Based on the weight coefficients corresponding to the feature elements and a preset weight coefficient threshold, filter out feature elements greater than the weight coefficient threshold, combine the feature elements according to a preset combination rule, and construct a comprehensive feature.
8. The targeted enterprise user screening method of any one of claims 1 to 3, wherein, The step of determining the positive samples corresponding to each enterprise user in the negative samples respectively based on the prediction score and the prediction score corresponding to each positive sample, and completing the target enterprise user screening, specifically comprises: Comparing the prediction score corresponding to each enterprise user in the negative samples with the prediction score corresponding to each positive sample in turn; According to the comparison result, the target positive sample corresponding to each enterprise user in the negative samples is screened out; The group agreement name corresponding to the target positive sample is recommended to the enterprise user corresponding to the target positive sample, and the target enterprise user screening is completed.
9. A target enterprise user screening apparatus characterized by comprising: It comprises: The positive and negative sample division module is used for obtaining the basic information of a plurality of enterprise users, and dividing the plurality of enterprise users into positive and negative samples according to the basic information, wherein the basic information comprises enterprise name and group agreement name signed by the enterprise user; The feature set acquisition module is used for acquiring the first feature set and the second feature set corresponding to each enterprise user in the same positive sample based on a preset form; The data cleaning module is used for respectively screening the feature elements in the first feature set and the second feature set, obtaining the feature elements meeting the preset screening condition, and constructing the comprehensive feature corresponding to the current sample, wherein the step of respectively screening the feature elements in the first feature set and the second feature set, obtaining the feature elements meeting the preset screening condition, and constructing the comprehensive feature corresponding to the current sample, specifically comprises: The feature elements in the first feature set and the second feature set are respectively subjected to feature binning processing, and the feature factor IV value corresponding to each feature element is acquired, the feature factor IV value being the contribution degree of the corresponding feature factor to the model prediction, the greater the IV value, the greater the contribution degree of the corresponding feature factor to the model prediction, and vice versa, the smaller the IV value, the smaller the contribution degree of the corresponding feature factor to the model prediction; Based on the preset feature factor IV value threshold and the feature factor IV value corresponding to each feature element, the feature elements in the first feature set and the second feature set are subjected to data cleaning, and the feature elements meeting the preliminary selection condition are obtained; The correlation analysis method and the principal component analysis method are used to analyze the feature elements meeting the preliminary selection condition, and the feature elements meeting the final selection condition are obtained, and the comprehensive feature corresponding to the current sample is constructed; The positive sample prediction module is used for sequentially inputting the comprehensive feature corresponding to each positive sample as an input parameter into a prediction scoring model based on a LightGBM algorithm, and acquiring the prediction score corresponding to each positive sample; The negative sample prediction module is used for acquiring the first feature set and the second feature set corresponding to each enterprise user in the negative sample, executing the data cleaning module, acquiring the comprehensive feature corresponding to each enterprise user in the negative sample, and inputting the comprehensive feature into the prediction scoring model to acquire the prediction score; The target enterprise user screening module is used for determining the positive samples corresponding to each enterprise user in the negative samples based on the prediction score and the prediction score corresponding to each positive sample, and completing the target enterprise user screening. 10.A computer device, comprising a memory and a processor, wherein the memory stores computer readable instructions, and the processor executes the computer readable instructions to implement the steps of the target enterprise user screening method according to any one of claims 1 to 8.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer readable instructions, and the processor executes the computer readable instructions to implement the steps of the target enterprise user screening method according to any one of claims 1 to 8.
Citation Information
Patent Citations
A data processing method and system based on a similarity model
CN109636482A
Game resource delivery method and device
CN111973996A