A method and apparatus for determining a target user

By performing feature analysis and model training on historical user information in the seed package, a target model is established, which solves the problems of low flexibility and accuracy in target user determination in existing technologies and realizes a more efficient user push strategy.

CN114119094BActive Publication Date: 2026-05-29WEBANK (CHINA)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WEBANK (CHINA)
Filing Date
2021-11-30
Publication Date
2026-05-29

Smart Images

  • Figure CN114119094B_ABST
    Figure CN114119094B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a kind of method and device for determining target user.The method comprises: the feature analysis is carried out to the user information of each historical user in selected seed package, and the significant feature of the seed package is determined;From the corresponding relationship between significant feature and algorithm tool, the first algorithm tool of the seed package is determined;With the seed package as sample, the model training is carried out to the first algorithm tool, and target model is obtained;For any existing user, whether the existing user is the target user in conformity with expectation is determined by the target model.Finally determined target user is not limited to the historical user in seed package.If target user is determined for different activities, it only needs to adjust selected seed package, so the method for determining target user is more flexible.Through the modeling of historical user in seed package by preset first algorithm tool, the participation of algorithm personnel is not needed, guaranteeing that prediction accuracy is high while saving manpower.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of financial technology, and in particular to a method, apparatus, computing device, and computer-readable storage medium for identifying target users. Background Technology

[0002] With the development of computer technology, more and more technologies are being applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech). However, due to the security and real-time requirements of the financial industry, higher demands are being placed on technology.

[0003] When promoting activities, operations staff often don't send out activity messages to all registered users, resulting in a lack of targeting and poor effectiveness. Instead, they use experience to filter target user groups for more targeted activity messages.

[0004] The existing push notification strategy involves identifying candidate user groups based on experience, sending messages to these groups, and recording users whose performance meets expectations as the target user group. In subsequent similar promotional campaigns, these target user groups can be selected for related message pushes. For example, if the operations team identifies users with over 100,000 RMB in deposits on the platform as candidate users, they send message A to these users. Message A could be related to product benefits or discounts. User feedback is recorded, and users who claim the product benefits or place orders through the discount information are selected as the target user group. In subsequent campaigns related to promoting product benefits or discounts, these target user groups will continue to receive campaign information.

[0005] It can be observed that the target audience identified using the above method is limited to users with deposits exceeding 100,000 yuan, and the same target audience is identified for different activities. This means that these target audiences are reached multiple times, while other users are reached with a low probability. The method for identifying the target audience lacks flexibility.

[0006] In summary, the embodiments of the present invention provide a method for determining target users, which can flexibly determine target users in different activities and improve the accuracy of target user determination. Summary of the Invention

[0007] This invention provides a method for determining target users, which can flexibly identify target users in different activities and improve the accuracy of target user identification.

[0008] In a first aspect, embodiments of the present invention provide a method for determining a target user, comprising:

[0009] Feature analysis is performed on the user information of each historical user in the selected seed package to determine the significant features of the seed package; each historical user is a user with a result label of a set target after a historical event; the set target is whether it meets or does not meet expectations;

[0010] The first algorithm tool for the seed package is determined from the correspondence between salient features and algorithm tools;

[0011] Using the seed package as a sample, the first algorithm tool is trained to obtain a target model; the target model is used to determine the probability that a user meets a set target based on the user information.

[0012] For any existing user, the target model is used to determine whether the existing user is a target user as expected.

[0013] Business personnel select seed packages of various historical events based on their needs. These packages contain historical users who have experienced the events and have set target labels. Since each historical user in the seed package has a set target, the salient features of each historical user in the seed package can be determined based on their user information. Based on the correspondence between salient features and algorithm tools, a first algorithm tool is easily determined, allowing for rapid feature extraction from the user information in the seed package. Thus, for any existing user, it is possible to determine whether they are a target user. The final target users are not limited to the historical users in the seed package. If target users are determined for different activities, only the selected seed package needs to be adjusted, making this method of determining target users more flexible. By modeling the historical users in the seed package using a pre-set first algorithm tool, no algorithm personnel are required, ensuring high prediction accuracy while saving manpower. Because the correspondence between each salient feature and the algorithm tool is pre-set, the operation is simple, fast, and efficient.

[0014] Optionally, the seed package includes at least a first seed package and a second seed package; the historical users in the first seed package are those whose results after historical events are as expected; the historical users in the second seed package are those whose results after historical events are not as expected.

[0015] For any existing user, determining whether the existing user is a target user according to the target model includes:

[0016] For any existing user, a first probability that the existing user meets the expectation is determined using the target model of the first seed package; a second probability that the existing user does not meet the expectation is determined using the target model of the second seed package.

[0017] Based on the first probability and the second probability, determine whether the existing user is the expected target user.

[0018] By selecting a first seed package containing users whose results are labeled as "meeting expectations" and a second seed package containing users whose results are labeled as "not meeting expectations," and by incorporating more historical user information, the probability of existing users being target users is determined from two dimensions. This approach more accurately reflects the potential intentions of existing users, thereby improving the accuracy of target user identification.

[0019] Optionally, the seed package includes at least a first seed package and a second seed package;

[0020] Each historical user in the first seed package and each historical user in the second seed package has experienced different historical events; or

[0021] The result labels for each historical user in the first seed package are as expected, while the result labels for each historical user in the second seed package are not as expected.

[0022] Selecting various types of seed packages based on needs, and integrating more information, improves the accuracy of identifying target users.

[0023] Optionally, determining the salient features of the seed package includes:

[0024] Based on any selected feature mining tool, feature analysis is performed on the user information of each historical user in the seed package to determine the salient features of the seed package.

[0025] This gives business personnel more options, not just limiting them to the system's preset feature mining tools. Business personnel can choose feature mining tools according to their needs. The feature mining process is more flexible and efficient.

[0026] Optionally, it also includes:

[0027] If the correspondence between the salient features and the algorithm tool does not include the salient features of the seed package, then return to the step of determining the salient features of the seed package, and select features that meet the set conditions from the remaining features as salient features, until the correspondence between the salient features and the algorithm tool includes the salient features of the seed package.

[0028] If the correspondence between salient features and algorithm tools does not include the salient features of the seed package, then the salient features are updated until the target algorithm tool is determined. This eliminates the need for manual selection, saving manpower and making the process convenient and fast.

[0029] Optionally, using the seed package as a sample, the first algorithm tool is trained to obtain a target model, including:

[0030] Using the seed package as a sample, the first algorithm tool is trained to obtain a first selectable model;

[0031] Using the seed package as a sample, one or more second algorithm tools selected from the algorithm tool library are used for model training to obtain one or more second optional models;

[0032] The user information and result labels of each test user in the test set are input into the first optional model and the second optional model to determine the accuracy of the first optional model and the second optional model; based on the accuracy of the first optional model and the second optional model, the target model is selected.

[0033] Not only can modeling be performed using a pre-set first algorithm tool, but users can also freely choose a second algorithm tool. The accuracy of each model is then compared to determine the target model. This provides business personnel with more options, allowing them to quickly select an algorithm tool, rapidly verify accuracy, and find the target model based on accuracy comparisons. This improves both the accuracy and flexibility in determining the target model.

[0034] Optionally, after determining whether the existing users are the expected target users, the method further includes:

[0035] For the target user, based on the historical behavior of historical users whose result tags in the seed package match the expected behavior, push information is determined for the target user and pushed to the target user.

[0036] This increases the likelihood that the existing user will respond to the historical behavior.

[0037] Secondly, embodiments of the present invention also provide an apparatus for determining a target user, comprising:

[0038] Determine the unit, used for:

[0039] Feature analysis is performed on the user information of each historical user in the selected seed package to determine the significant features of the seed package; each historical user is a user with a result label of a set target after a historical event; the set target is whether it meets or does not meet expectations;

[0040] The first algorithm tool for the seed package is determined from the correspondence between salient features and algorithm tools;

[0041] Processing unit, used for:

[0042] Using the seed package as a sample, the first algorithm tool is trained to obtain a target model; the target model is used to determine the probability that a user meets a set target based on the user information.

[0043] For any existing user, the target model is used to determine whether the existing user is a target user as expected.

[0044] Optionally, the seed package includes at least a first seed package and a second seed package; the historical users in the first seed package are those whose results after historical events are as expected; the historical users in the second seed package are those whose results after historical events are not as expected.

[0045] The processing unit is specifically used for:

[0046] For any existing user, a first probability that the existing user meets the expectation is determined using the target model of the first seed package; a second probability that the existing user does not meet the expectation is determined using the target model of the second seed package.

[0047] Based on the first probability and the second probability, determine whether the existing user is the expected target user.

[0048] Optionally, the seed package includes at least a first seed package and a second seed package;

[0049] Each historical user in the first seed package and each historical user in the second seed package has experienced different historical events; or

[0050] The result labels for each historical user in the first seed package are as expected, while the result labels for each historical user in the second seed package are not as expected.

[0051] Optionally, the determining unit is specifically used for:

[0052] Based on any selected feature mining tool, feature analysis is performed on the user information of each historical user in the seed package to determine the salient features of the seed package.

[0053] Optionally, the determining unit is further configured to:

[0054] If the correspondence between the salient features and the algorithm tool does not include the salient features of the seed package, then return to the step of determining the salient features of the seed package, and select features that meet the set conditions from the remaining features as salient features, until the correspondence between the salient features and the algorithm tool includes the salient features of the seed package.

[0055] Optionally, the processing unit is specifically used for:

[0056] Using the seed package as a sample, the first algorithm tool is trained to obtain a first selectable model;

[0057] Using the seed package as a sample, one or more second algorithm tools selected from the algorithm tool library are used for model training to obtain one or more second optional models;

[0058] The user information and result labels of each test user in the test set are input into the first optional model and the second optional model to determine the accuracy of the first optional model and the second optional model; based on the accuracy of the first optional model and the second optional model, the target model is selected.

[0059] Optionally, the processing unit is further configured to:

[0060] For the target user, based on the historical behavior of historical users whose result tags in the seed package match the expected behavior, push information is determined for the target user and pushed to the target user.

[0061] Thirdly, embodiments of the present invention also provide a computing device, comprising:

[0062] Memory, used to store computer programs;

[0063] The processor is configured to invoke a computer program stored in the memory and execute the method for determining the target user as listed above, according to the obtained program.

[0064] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer-executable program, the computer-executable program being used to cause a computer to perform the method for determining a target user listed in any of the above methods. Attached Figure Description

[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0066] Figure 1 A schematic diagram of a system architecture provided for an embodiment of the present invention;

[0067] Figure 2 A flowchart illustrating a possible method for determining a target user, as provided in an embodiment of the present invention;

[0068] Figure 3 A schematic diagram of the interface of a possible intelligent modeling system 300 provided in an embodiment of the present invention;

[0069] Figure 4A schematic diagram illustrating a specific application scenario provided by an embodiment of the present invention;

[0070] Figure 5 A flowchart illustrating a method for determining a target user according to an embodiment of the present invention;

[0071] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0072] To make the objectives, implementation methods and advantages of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.

[0073] Based on the exemplary embodiments described in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the appended claims. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete implementation on its own.

[0074] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0075] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities and do not necessarily imply a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be used interchangeably where appropriate, for example, to implement the application in a sequence other than those given in the embodiments illustrated or described herein.

[0076] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclusively include, for example, a product or device that includes a series of components is not necessarily limited to those that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.

[0077] Figure 1 An exemplary system architecture applicable to an embodiment of the present invention is shown. This system architecture can be a server 100, including a processor 110, a communication interface 120, and a memory 130.

[0078] The communication interface 120 is used to communicate with the terminal device, send and receive information transmitted by the terminal device, and realize communication.

[0079] The processor 110 is the control center of the server 100, connecting various parts of the server 100 through various interfaces and routes. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 130, and by calling data stored in the memory 130. Optionally, the processor 110 may include one or more processing units.

[0080] The memory 130 can be used to store software programs and modules. The processor 110 executes various functional applications and data processing by running the software programs and modules stored in the memory 130. The memory 130 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to business processing, etc. In addition, the memory 130 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0081] It should be noted that the above Figure 1 The structure shown is merely an example, and the embodiments of the present invention are not limited thereto.

[0082] The above Figure 1 The server shown can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0083] This invention provides a possible method for determining a target user, such as... Figure 2 As shown, it includes:

[0084] Step 201: Perform feature analysis on the user information of each historical user in the selected seed package to determine the significant features of the seed package; each historical user is a user with a set target result label after a historical event; the set target is whether it meets expectations or does not meet expectations.

[0085] Step 202: Determine the first algorithm tool of the seed package from the correspondence between salient features and algorithm tools.

[0086] Step 203: Using the seed package as a sample, train the first algorithm tool to obtain a target model; the target model is used to determine the probability that a user meets the set target based on the user information.

[0087] Step 204: For any existing user, determine whether the existing user is a target user as expected using the target model.

[0088] Figure 3 A schematic diagram of a possible intelligent modeling system 300 is shown, including a seed package 310, a feature mining tool library 320, and an algorithm tool library 330. When business personnel need to select target users for message push for a certain activity, they can select seed packages, feature mining tools, and algorithm tools according to their needs in the intelligent modeling system. Finally, the system models the data in the seed package selected by the business personnel based on the algorithm tools selected or the system's preset algorithm tools. The resulting model can be used to determine the target users.

[0089] The intelligent system can store multiple seed packages containing users with the same historical event but different set goals, different historical events with the same set goal, or different historical events and different set goals. Each seed package contains users who share the same historical event and the same set goal. For example, seed package 1 contains users successfully activated during the May account opening referral campaign (the historical event was the May account opening referral campaign, and the set goal met expectations); seed package 2 contains users successfully activated during the April dormant customer reactivation campaign (the historical event was the April dormant customer reactivation campaign, and the set goal met expectations); seed package 3 contains users who were not successfully activated during the May order-for-phone-credit campaign (the historical event was the May order-for-phone-credit campaign, and the set goal did not meet expectations); and seed package 4 contains users successfully activated during the May order-for-phone-credit campaign (the historical event was the May order-for-phone-credit campaign, and the set goal met expectations). Grouping historical users who meet the same historical event and the same set goal into seed packages allows business personnel to select seed packages based on their needs and extract features from the user information within each seed package, thus determining the salient features of the historical users in that seed package.

[0090] When selecting seed packages, business personnel can choose one or multiple seed packages. For example, seed packages 1 and 3 are selected. Seed package 1 is the seed package with the target being met (positive seed package), and seed package 3 is the seed package with the target being not met (negative seed package). Feature extraction is performed on the positive and negative seed packages to analyze the significant characteristics of historical users who meet the expectations and those who do not, thereby guiding the subsequent modeling process.

[0091] Each seed package contains user information for each historical user who participated in the historical event, such as user ID, gender, age, province, and education level, as shown in Table 1.

[0092] Table 1

[0093] User ID gender age province Education A female 23 Shanghai Undergraduate B male 25 Guangdong Undergraduate C female 23 Shanghai master D female 25 Guangdong PhD

[0094] The system performs feature analysis on the user information of historical users in the selected sub-packages to determine the significant features of each sub-package.

[0095] This paper takes the feature analysis of historical user information in seed package 1 as an example to introduce the method for determining significant features. The feature mining tool library in the interface of the intelligent modeling system includes various feature mining tools, such as LR (Linear Regression), FM (Factorization Machine), GBDT (Gradient Boosting Decision Tree), and DT (Decision Tree). Business personnel can select feature mining tools themselves in the interface of the intelligent modeling system, or the intelligent modeling system can pre-set feature mining tools for each seed package. This embodiment of the invention does not impose any limitations on this.

[0096] For example, business personnel might choose a decision tree as their feature mining tool, which offers strong interpretability. This tool is used to determine the information value of any feature for each historical user in the seed package. Then, based on the information values ​​of multiple features, one or more features are selected as salient features. Specifically, the information value of any feature is obtained as follows: for any feature value, calculate the first distribution ratio of that feature value in the seed package and the second distribution ratio of that feature value among all users; and determine the information value of the feature based on the first and second distribution ratios of the multiple feature values.

[0097] For example, if the user information in Seed Package 1 is as shown in Table 1, and the "age" feature includes two values: 23 years old and 25 years old, then first calculate the distribution ratio of 23-year-old users in the Seed Package and the distribution ratio of 23-year-old users in the total user population (the total user population can be registered users of the platform or selected users of a certain type, such as users with deposits of over 100,000; this embodiment of the invention does not impose any restrictions on this). Then calculate the distribution ratio of 25-year-old users in the Seed Package and the distribution ratio of 25-year-old users in the total user population (since the users in Seed Package 1 only have the two feature values ​​of 23 years old and 25 years old, only these two ages are calculated).

[0098] The IV (Information Value) of the feature "age" can be obtained using formula (1). Here, i is the feature value.

[0099] IV = ∑(seed packet distribution i - total user distribution i) * ln(seed packet distribution i / total user distribution i) (1)

[0100] Similarly, the information values ​​for features such as "province" and "age" can be obtained. The larger the information value, the more significant the feature is compared to the overall user population in seed package 1, reflecting the uniqueness of users in this seed package. The information values ​​are sorted, and one or more features with the largest information values ​​are selected as significant features.

[0101] If the information values ​​of each feature are ordered as follows: age, province, education level, gender, then the first one can be selected as the significant feature, or the first three can be selected as the significant features; there is no restriction here.

[0102] Next, the algorithm tool is determined based on the selected salient features. This invention provides three methods for determining the algorithm tool.

[0103] Option 1: Business personnel can select a second algorithm tool from the algorithm tool library on the interface of the intelligent modeling system according to their needs.

[0104] The intelligent modeling system's interface includes a library of algorithms such as LR (Linear Regression), FM (Factorization Machine), GBDT (Gradient Boosting Decision Tree), and DT (Decision Tree). Business personnel can freely choose from these options.

[0105] For example, if "age" was identified as a significant feature in the previous step, business personnel will select FM as the second algorithm tool in this step.

[0106] Method 2: The system pre-determines the correspondence between significant features and algorithm tools. Once the significant features are determined, the algorithm tool can be directly determined.

[0107] For example, in the previous step, "age" was identified as a salient feature. In this step, in the correspondence between salient features and algorithm tools, the first algorithm tool corresponding to the salient features of seed package 1 is identified as LR.

[0108] If the salient feature identified in this step is not found in the mapping between salient features and algorithm tools, the process returns to the step of determining the salient features of the seed package. From the remaining features, features that meet the set conditions are selected as salient features until the mapping between salient features and algorithm tools includes the salient features of the seed package. For example, if it is determined that the salient feature "age" is not found in the mapping between salient features and algorithm tools, the process returns to redetermining salient features: the information values ​​of the features determined in the previous step are sorted as follows: age, province, education level, gender. The second-ranked feature, "province," is selected as the salient feature. This continues until the salient feature is included in the mapping between salient features and algorithm tools.

[0109] Method 3: Use both the system's preset algorithm tools and the algorithm tools selected by business personnel to train the seed package, and select the optimal algorithm tool.

[0110] Before introducing Method 3, we need to first introduce how to use algorithmic tools to perform modeling in order to obtain the target model.

[0111] The modeling method involves using a selected seed package as a sample, inputting the user information and result labels of each historical user in the seed package into the algorithm tool, and training the model until the target model is obtained.

[0112] For example, for historical users in seed package 1, the LR algorithm tool is used to train the model. The model formula is formula (2). Where xi is the value of each feature, ωi is the weight corresponding to the value of each feature, that is, the variable to be calculated in the model, and y' is the calculated predicted value.

[0113] h(x)=ω1x1+ω2x2+…+ωnxn+b (2)

[0114] y'=1 / (1+eh(x)) (3)

[0115] Accordingly, the loss function is:

[0116]

[0117] y i That is, the true value of the i-th sample, y′ i The predicted value is obtained by substituting the feature value of the i-th sample into formulas (2) and (3). Substituting the user information of all historical users in seed package 1 into formulas (2), (3), and (4), we calculate ωi to minimize the loss function L. The update of ωi can be achieved using gradient descent, i.e.:

[0118]

[0119] Where, ω tLet ω be the value corresponding to the t-th iteration, α be the learning rate (the speed at which samples are learned), set empirically to 0.01, and ω0 be randomly assigned. First, ω0 is randomly assigned. Then, all samples (steps (1) and (2) above) are substituted into formulas (1), (2), and (3) to calculate the L value. Next, the L value is substituted into formula (4) to update ω. This process is repeated iteratively until the L value is no longer updated. At this point, the model training is complete. The ω obtained at this time represents the weights corresponding to the values ​​of each feature in formula (1), reflecting the importance of each feature value.

[0120] For example, x1 represents whether the person is 18 years old; if yes, x1 = 1, otherwise x1 = 0; x2 represents whether the person is 19 years old; if yes, x2 = 1, otherwise x2 = 0; ... x6 represents whether the person is 23 years old; if yes, x6 = 1, otherwise x6 = 0; ... x8 represents whether the person is 25 years old; if yes, x8 = 1, otherwise x8 = 0; ... x9 represents whether the person is female; if yes, x9 = 1, otherwise x9 = 0; x10 represents whether the person is from Guangdong; if yes, x10 = 1, otherwise x10 = 0... Thus, the value of each feature can be represented by 1 or 0. The above is just an example; xi can also represent a range, such as whether the person is 18-30 years old, etc., and the embodiments of the present invention do not limit this.

[0121] First, presuppose ω in formula (2). t For each historical user in seed package 1 as shown in Table 1, for A, substitute the value of A into formula (2) to obtain the predicted value y′1 in (3); substitute the value of B into formula (2) to obtain the predicted value y′2 in (3); substitute the value of C into formula (2) to obtain the predicted value y′3 in (3); substitute the value of D into formula (2) to obtain the predicted value y′4 in (3), substitute the obtained predicted value into formula (4) to obtain the loss function L1; update ω according to formula (5). t The value of ω is obtained by repeating the above steps to obtain the loss function L2... until the loss function value is no longer updated. t The value of is the parameter of the trained target model. Table 2 shows one possible set of determined model parameters.

[0122] Table 2

[0123]

[0124] Next, we will introduce method three mentioned above.

[0125] Using the seed package as a sample, the first algorithm tool preset by the system is used for modeling to obtain the first optional model. The modeling method is as described above.

[0126] Using the seed package as a sample, business personnel select one or more second algorithm tools from the algorithm tool library to train the model and obtain one or more second optional models.

[0127] The user information and result labels of each test user in the test set are input into the first optional model and the second optional model to determine the accuracy of the first optional model and the second optional model. The model with an accuracy that meets a preset threshold or has the highest accuracy is determined as the final target model.

[0128] Figure 4 The document illustrates a specific application scenario. It includes the following steps:

[0129] Step 401: The system identifies salient features.

[0130] Step 402: Search for the algorithm tool corresponding to the salient feature in the correspondence between salient features and algorithm tools.

[0131] Step 403: Does the significant feature exist? If it exists, proceed to step 404; if it does not exist, return to step 401.

[0132] Step 404: Determine the first algorithm tool;

[0133] Step 405: The system pops up a request to the business personnel: whether to use the first algorithm tool;

[0134] Step 406: Use the first algorithm tool to train the model and obtain the first selectable model.

[0135] Step 407: Business personnel select one or more second algorithm tools from the algorithm tool library.

[0136] Step 408: The system pops up a request to the business personnel: whether to use the first optional model.

[0137] Step 409: Obtain one or more second optional models.

[0138] Step 410: Based on the accuracy, determine the target model from the first optional model and / or the second optional model.

[0139] Step 411, Model Application.

[0140] Once the system identifies a salient feature, it searches for the corresponding algorithm tool in the feature-algorithm mapping. If the salient feature does not exist, it returns to re-identify the salient feature until it appears in the mapping. Then, it determines the first algorithm tool corresponding to that salient feature. At this point, the system prompts the business user with a request: "Do you want to use this first algorithm tool?" If the user selects "Yes," the system uses this tool to train the model, obtaining a first selectable model. If the user selects "No," an algorithm tool selection interface appears, allowing the user to choose an algorithm tool. The user can select one or more second algorithm tools. The system then obtains one or more second selectable models based on these second algorithm tools. The target model is determined based on the accuracy of the second selectable models. Finally, the model is applied.

[0141] After obtaining the first optional model, the system can also prompt the business personnel with a request: whether to use the first optional model. If the business personnel select yes, the model is applied; if the business personnel select no, the process proceeds to step 407, where the business personnel select one or more second algorithm tools. The system obtains one or more second optional models based on one or more second algorithm tools. The target model is determined based on the accuracy of the first and second optional models. Finally, the model is applied.

[0142] Optionally, multiple salient features can be identified, and multiple first algorithm tools can be determined from the correspondence between salient features and algorithm tools. In this way, multiple target models can be obtained, the accuracy of multiple target models can be tested, and the final target model can be determined.

[0143] Optionally, multiple salient features can be identified, and the algorithm tools corresponding to combinations of multiple salient features can be stored in the correspondence between salient features and algorithm tools, so that one or more corresponding algorithm tools can be identified.

[0144] In this way, business personnel can quickly verify the effectiveness of multiple models, update the models or adjust parameters in a timely manner, and find the optimal target model by comparing the effects of multiple models, thereby improving the accuracy of identifying target users.

[0145] The above describes the process of determining the target model for seed package 1. The same process can be performed for seed package 3 selected by business personnel to obtain their respective target models.

[0146] For any existing user, input the user information (e.g., female, 23 years old) into the target model of seed package 1 to obtain the first probability that the existing user meets the expectation. For example, the first probability is output as 0.6.

[0147] Input the existing user's information (e.g., female, 23 years old) into the target model of seed package 2 to obtain the second probability that the existing user does not meet expectations. For example, output the second probability as 0.2.

[0148] Based on the first probability and the second probability, it is determined whether the existing user is the expected target user. Subtracting the second probability from the first probability yields a probability of 0.4 that the user meets the expectation.

[0149] If the probability of meeting the expectation is greater than the preset threshold, then the existing user is the target user.

[0150] Optionally, one or more seed packages that meet the expectations can be selected for the above method process; one or more seed packages that do not meet the expectations can also be selected for the above method process; alternatively, one or more seed packages that meet the expectations can be selected to obtain one or more target models, thereby obtaining one or more first probabilities that existing users meet the expectations, and one or more seed packages that do not meet the expectations can be selected to obtain one or more target models, thereby obtaining one or more second probabilities that existing users do not meet the expectations. Then, the sum of multiple first probabilities minus the sum of multiple second probabilities is used to obtain the probability that existing users meet the expectations. This embodiment of the invention does not limit this. For example, if the business personnel select seed package 3, the system can extract significant features from the users in seed package 3, determine the first algorithm tool, and use the first algorithm tool to model the historical users in seed package 3. Since the historical users in seed package 3 are all users who do not meet the expectations, the obtained target model can be used to determine the probability that any existing user does not meet the expectations, and thus the probability that they meet the expectations can be known, that is, whether they are target users who can receive message pushes. For example, if the probability that any existing user does not meet the expectations is determined to be 0.2, then the probability that they meet the expectations is determined to be 1-0.2=0.8.

[0151] For example, after selecting two seed packages that meet the expectations and two seed packages that do not meet the expectations to train the model and obtain four target models, the first probability of a certain existing user is calculated to be 0.8 and 0.5, and the second probability is calculated to be 0.4 and 0.3. Then the probability of the existing user meeting the expectations is (0.8+0.5)-(0.4+0.3)=0.6.

[0152] Optionally, the next push notification action can also be directed to the identified target users. That is, after determining whether existing users are the expected target users, the process also includes:

[0153] For the target user, based on the historical behavior of historical users whose result tags in the seed package match the expected behavior, push information is determined for the target user and pushed to the target user.

[0154] For example, seed package 1 stores the historical action taken against a historical user during this event, which was sending a promotional SMS message. Seed package 2 stores the historical action taken against a historical user during this event, which was giving away three months of video streaming privileges. For a certain existing user, the first probability determined by the target model of seed package 1 is 0.6, and the first probability determined by the target model of seed package 2 is 0.2. Therefore, it is recommended that this existing user adopt the historical action from seed package 1: sending a promotional SMS message.

[0155] Since the existing user is more consistent with the user information of the historical user in Seed Package 1, implementing the historical behavior of Seed Package 1 on the existing user can increase the likelihood that the existing user will respond to the historical behavior.

[0156] Business personnel select seed packages of various historical events based on their needs. These packages contain historical users who have experienced the events and have set target labels. Since each historical user in the seed package has a set target, the salient features of each historical user in the seed package can be determined based on their user information. Based on the correspondence between salient features and algorithm tools, a first algorithm tool is easily determined, allowing for rapid feature extraction from the user information in the seed package. Thus, for any existing user, it is possible to determine whether they are a target user. The final target users are not limited to the historical users in the seed package. If target users are determined for different activities, only the selected seed package needs to be adjusted, making this method of determining target users more flexible. By modeling the historical users in the seed package using a pre-set first algorithm tool, no algorithm personnel are required, ensuring high prediction accuracy while saving manpower. Because the correspondence between each salient feature and the algorithm tool is pre-set, the operation is simple, fast, and efficient.

[0157] Based on the same technological concept Figure 5 An exemplary embodiment of the present invention provides a structure for a target user determination device, which can execute a process for determining a target user.

[0158] like Figure 5 As shown, the device specifically includes:

[0159] Determine unit 501, used for:

[0160] Feature analysis is performed on the user information of each historical user in the selected seed package to determine the significant features of the seed package; each historical user is a user with a result label of a set target after a historical event; the set target is whether it meets or does not meet expectations;

[0161] The first algorithm tool for the seed package is determined from the correspondence between salient features and algorithm tools;

[0162] Processing unit 502 is used for:

[0163] Using the seed package as a sample, the first algorithm tool is trained to obtain a target model; the target model is used to determine the probability that a user meets a set target based on the user information.

[0164] For any existing user, the target model is used to determine whether the existing user is a target user as expected.

[0165] Optionally, the seed package includes at least a first seed package and a second seed package; the historical users in the first seed package are those whose results after historical events are as expected; the historical users in the second seed package are those whose results after historical events are not as expected.

[0166] The processing unit 502 is specifically used for:

[0167] For any existing user, a first probability that the existing user meets the expectation is determined using the target model of the first seed package; a second probability that the existing user does not meet the expectation is determined using the target model of the second seed package.

[0168] Based on the first probability and the second probability, determine whether the existing user is the expected target user.

[0169] Optionally, the seed package includes at least a first seed package and a second seed package;

[0170] Each historical user in the first seed package and each historical user in the second seed package has experienced different historical events; or

[0171] The result labels for each historical user in the first seed package are as expected, while the result labels for each historical user in the second seed package are not as expected.

[0172] Optionally, the determining unit 501 is specifically used for:

[0173] Based on any selected feature mining tool, feature analysis is performed on the user information of each historical user in the seed package to determine the salient features of the seed package.

[0174] Optionally, the determining unit 501 is further configured to:

[0175] If the correspondence between the salient features and the algorithm tool does not include the salient features of the seed package, then return to the step of determining the salient features of the seed package, and select features that meet the set conditions from the remaining features as salient features, until the correspondence between the salient features and the algorithm tool includes the salient features of the seed package.

[0176] Optionally, the processing unit 502 is specifically used for:

[0177] Using the seed package as a sample, the first algorithm tool is trained to obtain a first selectable model;

[0178] Using the seed package as a sample, one or more second algorithm tools selected from the algorithm tool library are used for model training to obtain one or more second optional models;

[0179] The user information and result labels of each test user in the test set are input into the first optional model and the second optional model to determine the accuracy of the first optional model and the second optional model; based on the accuracy of the first optional model and the second optional model, the target model is selected.

[0180] Optionally, the processing unit 502 is further configured to:

[0181] For the target user, based on the historical behavior of historical users whose result tags in the seed package match the expected behavior, push information is determined for the target user and pushed to the target user.

[0182] Based on the same technical concept, embodiments of this application provide a computer device, such as... Figure 6 As shown, it includes at least one processor 601 and a memory 602 connected to at least one processor. In this embodiment, the specific connection medium between the processor 601 and the memory 602 is not limited. Figure 6 Taking the connection between the processor 601 and the memory 602 via a bus as an example, the bus can be divided into address bus, data bus, control bus, etc.

[0183] In this embodiment of the application, the memory 602 stores instructions that can be executed by at least one processor 601. By executing the instructions stored in the memory 602, at least one processor 601 can perform the steps of the above-described method for determining the target user.

[0184] The processor 601 is the control center of the computer device. It can connect to various parts of the computer device using various interfaces and lines. It identifies the target user by running or executing instructions stored in the memory 602 and calling data stored in the memory 602. Optionally, the processor 601 may include one or more processing units. The processor 601 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 601. In some embodiments, the processor 601 and the memory 602 can be implemented on the same chip; in some embodiments, they can also be implemented on separate chips.

[0185] Processor 601 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0186] Memory 602, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 602 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 602 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 602 may also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0187] Based on the same technical concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer-executable program, the computer-executable program being used to cause a computer to perform the method for determining a target user listed in any of the above methods.

[0188] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0189] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0190] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0191] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0192] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for identifying target users, characterized in that, include: Feature analysis is performed on the user information of each historical user in the selected seed package to determine the significant features of the seed package; Each historical user is a user whose result label after a historical event has a set target; the set target is whether it meets or does not meet expectations; the seed package includes N first seed packages and N second seed packages; the historical users in the first seed package are users whose result label after a historical event meets expectations; the historical users in the second seed package are users whose result label after a historical event does not meet expectations; N is greater than or equal to 1; The first algorithm tool for the seed package is determined from the correspondence between salient features and algorithm tools; Using the seed package as a sample, the first algorithm tool is trained to obtain a target model; the target model is used to determine the probability that a user meets a set target based on the user information. For any existing user, N first probabilities that the existing user meets expectations are determined using the N target models corresponding to the N first seed packages; N second probabilities that the existing user does not meet expectations are determined using the N target models corresponding to the N second seed packages. The existing user is determined to be the expected target user by subtracting the sum of the N first probabilities from the sum of the N second probabilities.

2. The method as described in claim 1, characterized in that, The seed package includes at least a first seed package and a second seed package; Each historical user in the first seed package and each historical user in the second seed package has experienced different historical events; or The result labels for each historical user in the first seed package are as expected, while the result labels for each historical user in the second seed package are not as expected.

3. The method as described in claim 1, characterized in that, Feature analysis is performed on the user information of each historical user in the selected seed package to determine the significant features of the seed package, including: Based on the selected feature mining tool, determine the information value of any feature of each historical user in the seed package; Based on the information values ​​of multiple features, select one or more features as salient features; The information value of any feature is obtained in the following ways: For any feature value among the features, calculate the first distribution ratio of the feature value in the seed package and the second distribution ratio of the feature value in the total number of users; The information value of the feature is determined based on the first distribution ratio and the second distribution ratio of multiple feature values ​​in the feature.

4. The method as described in claim 1, characterized in that, Also includes: If the correspondence between the salient features and the algorithm tool does not include the salient features of the seed package, then return to the step of determining the salient features of the seed package, and select features that meet the set conditions from the remaining features as salient features, until the correspondence between the salient features and the algorithm tool includes the salient features of the seed package.

5. The method as described in claim 1, characterized in that, Using the seed package as a sample, the first algorithm tool is trained to obtain the target model, including: Using the seed package as a sample, the first algorithm tool is trained to obtain a first selectable model; Using the seed package as a sample, one or more second algorithm tools selected from the algorithm tool library are used for model training to obtain one or more second optional models; The user information and result labels of each test user in the test set are input into the first optional model and the second optional model to determine the accuracy of the first optional model and the second optional model; based on the accuracy of the first optional model and the second optional model, the target model is selected.

6. The method according to any one of claims 1-5, characterized in that, After determining whether the existing users are the expected target users, the process also includes: For the target user, based on the historical behavior of historical users whose result tags in the seed package match the expected behavior, push information is determined for the target user and pushed to the target user.

7. An apparatus for identifying a target user, characterized in that, include: Determine the unit, used for: Feature analysis is performed on the user information of each historical user in the selected seed package to determine the significant features of the seed package; Each historical user is a user whose result label after a historical event has a set target; the set target is whether it meets or does not meet expectations; the seed package includes N first seed packages and N second seed packages; the historical users in the first seed package are users whose result label after a historical event meets expectations; the historical users in the second seed package are users whose result label after a historical event does not meet expectations; N is greater than or equal to 1; The first algorithm tool for the seed package is determined from the correspondence between salient features and algorithm tools; Processing unit, used for: Using the seed package as a sample, the first algorithm tool is trained to obtain a target model; the target model is used to determine the probability that a user meets a set target based on the user information. For any existing user, N first probabilities that the existing user meets expectations are determined using the N target models corresponding to the N first seed packages; N second probabilities that the existing user does not meet expectations are determined using the N target models corresponding to the N second seed packages. The existing user is determined to be the expected target user by subtracting the sum of the N first probabilities from the sum of the N second probabilities.

8. A computing device, characterized in that, include: Memory, used to store computer programs; A processor is configured to invoke a computer program stored in the memory and execute the method according to any one of claims 1 to 6 in accordance with the obtained program.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer-executable program for causing a computer to perform the method according to any one of claims 1 to 6.