Method, device, and storage medium for determining user type information

By entering the user's characteristic data into the prediction model to obtain conversion rate, the user type problem in the prior art that cannot reflect different marketing conditions is solved, and the effective user type division under preset execution conditions is realized.

CN114298232BActive Publication Date: 2025-06-10WEBANK (CHINA)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111655734.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-06-10
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The existing user type classification method cannot reflect the user types under different marketing conditions, resulting in the low validity of the divided user types.

Method used

By obtaining the user's characteristic data, inputting it into the first prediction model and the second prediction model respectively, the conversion rate output by the two prediction models is obtained, and the target data generation probability is respectively used to predict the preset execution conditions and under the preset execution conditions without applying the preset execution conditions, and the user's type information is determined based on the two conversion rates.

Benefits of technology

The user type is effectively divided under preset execution conditions, which improves the effectiveness of user type division.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298232B_ABST
    Figure CN114298232B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method, device, and storage medium for determining the type information of a user, belonging to the field of technology finance. The method includes: obtaining the feature data of the user; respectively inputting the feature data of the user into a first prediction model and a second prediction model, and obtaining a first conversion rate output by the first prediction model and a second conversion rate output by the second prediction model. The first prediction model is used to predict the probability of generating the target data of the user under preset execution conditions, and the second prediction model is used to predict the probability of generating the target data of the user without applying the preset execution conditions; determining the type information of the user according to the first conversion rate and the second conversion rate. In this way, the types of users can be effectively classified under preset execution conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of fintech, and in particular, to a method, device, and storage medium for determining type information of users. Background Art

[0002] With the development of computer technology, more and more technologies are applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech), and the technology for determining the type information of users is no exception. However, due to the security and real-time requirements of the financial industry, higher requirements are also put forward for the technology.

[0003] In related technologies, by obtaining a sample of the cumulative credit utilization rate of a customer, using the basic information of the customer (identity, transactions, assets, credit, purchases, etc.) as features, and inputting them into a machine learning model, the predicted cumulative credit utilization rate of the customer's loan is output. If the cumulative credit utilization rate of the customer's loan exceeds a preset threshold, it indicates that the customer is a potential customer, thereby classifying the type of the user.

[0004] However, in the existing method for classifying the type of users, it is impossible to reflect the type of users under different marketing conditions, resulting in low effectiveness of the classified type of users. Summary of the Invention

[0005] Embodiments of the present disclosure provide a method, device, and storage medium for determining type information of users to solve the problem of low effectiveness of the classified type of users in the prior art.

[0006] In a first aspect, embodiments of the present disclosure provide a method for determining type information of users, the method including:

[0007] Obtain feature data of a user;

[0008] Input the feature data of the user into a first prediction model and a second prediction model respectively, and obtain a first conversion rate output by the first prediction model and a second conversion rate output by the second prediction model. The first prediction model is used to predict the probability of generating target data of the user under a preset execution condition, and the second prediction model is used to predict the probability of generating target data of the user without applying the preset execution condition;

[0009] Determine the type information of the user according to the first conversion rate and the second conversion rate.

[0010] In an optional implementation manner, the determining the type information of the user according to the first conversion rate and the second conversion rate includes:

[0011] Use the difference between the first conversion rate and the second conversion rate as the contribution value of the probability of generating the target data of the user

[0012] ;

[0013] Determine the type information of the user according to the first conversion rate, the second conversion rate and the contribution value.

[0014] In an optional implementation manner, the first prediction model is generated after being trained by a first sample set, and the first sample set includes feature data of historical users and result data of generating the target data of the historical users under the preset execution condition;

[0015] The second prediction model is generated after being trained by a second sample set, and the second sample set includes feature data of historical users and result data of generating the target data of the historical users without applying the preset execution condition.

[0016] In an optional implementation manner, after determining the type information of the user according to the first conversion rate and the second conversion rate, the method further includes:

[0017] Query the transfer users of the users of the target type in the database according to the multi-dimensional vector of the user, where the multi-dimensional vector is used to represent the association relationship between users in multiple dimensions, and the transfer users are the users to whom the preset execution condition is applied for magnitude expansion.

[0018] In an optional implementation manner, the querying the transfer users of the users of the target type in the database according to the multi-dimensional vector of the user includes:

[0019] Determine whether the user to be queried is the transfer user according to the cosine similarity between the multi-dimensional vector of the user to be queried in the database and the multi-dimensional vector of the user of the target type.

[0020] In an optional implementation manner, the querying the transfer users of the users of the target type in the database according to the multi-dimensional vector of the user includes:

[0021] Sample the multi-dimensional vectors of the users of the target type and the multi-dimensional vectors of the users of non-target types according to a preset sampling ratio to generate a third sample set, where the multi-dimensional vectors of the users of the target type are positive samples of the third sample set, and the multi-dimensional vectors of the users of non-target types are negative samples of the third sample set;

[0022] Use the third sample to train the similar population expansion model;

[0023] Input the multi-dimensional vector of the user to be queried in the database into the trained similar population expansion model, and obtain the population conversion probability output by the trained similar population expansion model;

[0024] Determine whether the user to be queried is the transferred user according to the population conversion probability.

[0025] In an optional implementation manner, before querying for the transferred users of the target type in the database according to the multi-dimensional vector of the user, the method further includes:

[0026] Determine the multi-dimensional vector of the users in the database according to the association information between the users.

[0027] In an optional implementation manner, the determining the multi-dimensional vector of the users in the database according to the association information between the users includes:

[0028] Select a target user from the database as the target node in the user relationship network;

[0029] Sequentially determine the next user node of the current end node in the associated user node sequence of the target node until the length of the node sequence of the target user reaches a preset sequence length;

[0030] Generate an associated node array of the target node according to the associated user node sequence of the target node;

[0031] Determine the multi-dimensional vector of the users in the database according to the associated node array of the target node.

[0032] In an optional implementation manner, the determining the next user node of the current end node in the associated user node sequence of the target node includes:

[0033] Determine the normalized transfer probability of the current end node and perform weighted sampling on the associated nodes of the current end node to determine the next user node.

[0034] In an optional implementation manner, the determining the normalized transfer probability of the current end node includes:

[0035] Determine the transfer probability between the current end node and any associated user node according to the association information between the users;

[0036] Normalize the transfer probability between the current end node and any associated user node to determine the normalized transfer probability of the current end node.

[0037] In an alternative embodiment, determining the transition probability between the current end node and any associated user node according to the association information between the users includes:

[0038] Generating weight data between the current end node and any associated user node according to the association information between the users and the identifiers of the users;

[0039] Determining a weight correction coefficient between the current end node and any associated user node according to the value of the shortest path distance between the current end node and any associated user node in the user relationship network;

[0040] Determining the transition probability between the current end node and any associated user node according to the weight data between the current end node and any associated user node and the weight correction coefficient between the current end node and any associated user node.

[0041] In an alternative embodiment, the value of the shortest path distance includes a first value, a second value, and a third value; there is a mapping relationship between the value of the shortest path distance and the weight correction coefficient;

[0042] If the associated user node is the previous node of the current end node, the value of the shortest path distance is the first value; if the associated user node and the current end node are adjacent nodes, the value of the shortest path distance is the second value; if the associated user node is not the previous node or the adjacent node of the current end node, the value of the shortest path distance is the third value.

[0043] In a second aspect, an embodiment of the present disclosure provides a device for determining the type information of a user, including:

[0044] An acquisition module, configured to acquire the feature data of the user.

[0045] A prediction module, configured to respectively input the feature data of the user into a first prediction model and a second prediction model, and acquire a first conversion rate output by the first prediction model and a second conversion rate output by the second prediction model. The first prediction model is used to predict the probability of generating the target data of the user under a preset execution condition, and the second prediction model is used to predict the probability of generating the target data of the user without applying the preset execution condition.

[0046] A determination module, configured to determine the type information of the user according to the first conversion rate and the second conversion rate.

[0047] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0048] Among them, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to determine the type information of the user according to any one of the first aspect and its optional modes.

[0049] In a fourth aspect, an embodiment of the present disclosure provides a computer storage medium, which stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor to determine the type information of the user according to any one of the first aspect.

[0050] The method, device, and storage medium for determining the type information of the user provided by the embodiments of the present disclosure first obtain the feature data of the user. Subsequently, the feature data of the user are respectively input into the first prediction model and the second prediction model to obtain the first conversion rate output by the first prediction model and the second conversion rate output by the second prediction model. The first prediction model is used to predict the probability of generating the target data of the user under the preset execution condition, and the second prediction model is used to predict the probability of generating the target data of the user without applying the preset execution condition. Finally, according to the first conversion rate and the second conversion rate, the type information of the user is determined. In this way, the type of the user can be effectively divided under the preset execution condition. Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions in the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a schematic diagram of a scenario of an operating environment provided by an embodiment of the present disclosure;

[0053] Figure 2 It is a schematic flowchart of a method for determining the type information of a user provided by an embodiment of the present disclosure;

[0054] Figure 3 It is a schematic flowchart of a method for mining transferred users provided by an embodiment of the present disclosure;

[0055] Figure 4 It is a schematic flowchart of another method for mining transferred users provided by an embodiment of the present disclosure;

[0056] Figure 5 It is a schematic diagram of the association relationship between user nodes provided by an embodiment of the present disclosure;

[0057] Figure 6 It is a schematic diagram of determining the multi-dimensional vector of a user provided by an embodiment of the present disclosure;

[0058] Figure 7 A schematic diagram of the transition probability between nodes provided by an embodiment of the present disclosure;

[0059] Figure 8 A schematic structural diagram of a determining device for user type information provided by an embodiment of the present disclosure;

[0060] Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0062] Inclusive finance is one of the important tones for the transformation and development of the Chinese financial industry at present. Many commercial banks are creating new marketing models for online customer acquisition and operation. How to accurately locate potential customers of their own banks among a large number of people, recommend the most suitable products to them, and improve the active stickiness of existing customers has become an important issue for many banks to consider.

[0063] In the related art, by obtaining the cumulative quota utilization rate samples of customers, taking the basic information of customers (identity, transactions, assets, credit, purchases, etc.) as features, and inputting them into a machine learning model, the predicted cumulative quota utilization rate of customers' loans is output. If the cumulative quota utilization rate of a customer's loan exceeds a preset threshold, it indicates that the customer is a potential customer, thereby classifying the type of user. However, in the existing method for classifying the type of user, the type of user under different marketing conditions cannot be reflected, resulting in low effectiveness of the classified type of user.

[0064] To solve the above problems, the embodiments of the present disclosure provide a method, device, and storage medium for determining user type information, predicting the first conversion rate of generating target data of a user under preset execution conditions and the second conversion rate of generating target data of a user without applying the preset execution conditions, and then determining the user type information based on the first conversion rate and the second conversion rate, so that the type of user can be effectively classified under the preset execution conditions.

[0065] Before describing the method for determining user type information of the present disclosure, first according to Figure 1 to understand the exemplary operating environment of the present disclosure.

[0066] Figure 1 A schematic diagram of a scenario of an operating environment provided by an embodiment of the present disclosure. As Figure 1 shown, it shows entities that want to obtain the type information of users, such as enterprise 101, banking institution 102, etc. These entities can request to query the type information of users from the system platform 103 as needed. Of course, the above entities are only for illustrative purposes. In fact, there are other entities that can initiate queries, such as the system platform 103 automatically initiating queries, which will not be exemplified one by one here. Query requests from various entities are provided to the system platform 103 through the network. The system platform 103 is used to perform the task of determining the type information of users. The system platform 103 may not only include a query module for querying the type information of users, but also include a marketing module and a mining module. The marketing module is used to provide different marketing plans for different types of users, and the mining module is used to mine potential transfer users of target type users. In addition, during the process of querying transfer users, the system platform 103 can also provide a database 104, which contains users to be queried, such as an enterprise information database. It should be understood that the enterprise information database in the exemplary environment is only exemplary, and other types of information databases all fall within the protection scope of the present disclosure. And, in the above exemplary operating scenario, the entities that obtain the type information of users can access the network using various devices, such as personal computers, servers, tablets, mobile phones, PDAs, notebooks, or any other computing device with networking capabilities. The system platform 103 can be implemented by a server or a server group with more powerful processing capabilities and higher security. And the network used between them can include various types of wired and wireless networks, such as but not limited to: the Internet, local area networks, WIFI, WLAN, cellular communication networks (GPRS, CDMA, 2G / 3G / 4G / 5G cellular networks), satellite communication networks, etc.

[0067] It can be understood that the method for determining the type information of the above users can be implemented by the device for determining the type information of users provided by an embodiment of the present disclosure. The device for determining the type information of users can be part or all of a certain device, such as a server or a chip of a server.

[0068] Taking a server integrated or installed with relevant execution code as an example below, the technical solution of an embodiment of the present disclosure will be described in detail with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0069] Figure 2A flowchart of a method for determining the type information of a user provided by an embodiment of the present disclosure. This embodiment relates to the process of how a server determines the type information of a user. Different from the existing methods for determining the type information of a user, the present disclosure respectively predicts a first conversion rate of generating target data of a user under a preset execution condition and a second conversion rate of generating target data of a user without applying the preset execution condition, and then determines the type information of the user according to the first conversion rate and the second conversion rate. Therefore, the method for determining the type information of a user provided by the present disclosure can effectively classify the type of a user under the influence of the preset execution condition on the user type.

[0070] Specifically, as Figure 2 shown, the method includes:

[0071] S201. Obtain the feature data of the user.

[0072] In the present disclosure, when it is necessary to classify the type of a user, the server can obtain the feature data of the user.

[0073] It should be understood that the present disclosure embodiment does not limit how to obtain the feature data of the user. In some embodiments, the unique identification (ID) of the user can be obtained, and then according to the unique identification of the user, the feature information related to the user can be determined.

[0074] It should be noted that the users involved in the embodiments of the present disclosure can be enterprise users or individual users, and the present disclosure embodiments do not limit this.

[0075] It should be understood that the present disclosure embodiment also does not limit the feature data of the user, which can be specifically determined according to the actual situation. Exemplarily, if it is an enterprise user, the feature data of the user can include but is not limited to financial data, industrial and commercial data, regional data, equity data, and business exception data.

[0076] Among them, the financial data is used to represent the total assets, owner's equity, investment income, operating income, non-operating income, total profit, main business income, net profit, total liabilities, total tax payment, operating cost, sales expense, asset impairment loss, etc. The industrial and commercial data is used to represent the company type, establishment years, industry, enterprise status, registered capital, number of industrial and commercial changes, etc. The regional data is used to represent the province, city, etc. The equity data is used to represent the number of direct shareholders, the number of natural person direct shareholders, the shareholding ratio of natural person direct shareholders, the number of non-natural person direct shareholders, the shareholding ratio of non-natural person direct shareholders, etc. The business exception is used to represent the number of administrative penalties, the number of business exceptions, the number of tax violations, the number of defendants, the number of times of being dishonestly executed, etc.

[0077] S202. Input the user's feature data into the first prediction model and the second prediction model respectively, and obtain the first conversion rate output by the first prediction model and the second conversion rate output by the second prediction model. The first prediction model is used to predict the probability of generating the user's target data under preset execution conditions, and the second prediction model is used to predict the probability of generating the user's target data without applying the preset execution conditions.

[0078] In this step, after the server obtains the user's feature data, it can input the user's feature data into the first prediction model and the second prediction model respectively, so as to obtain the first conversion rate and the second conversion rate.

[0079] It should be understood that the present disclosure embodiments do not limit the preset execution conditions and the target data. In some embodiments, the preset execution conditions may be to conduct marketing on the user. Correspondingly, the target data may be the transaction data generated by the user. Through the first conversion rate and the second conversion rate, it can be reflected the probability that the user is converted into a transaction behavior when being marketed and not being marketed.

[0080] Next, how to construct the first prediction model and the second prediction model will be described.

[0081] In some embodiments, the first prediction model is generated after being trained by a first sample set, and the first sample set contains the feature data of historical users and the result data of generating the target data of historical users under preset execution conditions. Correspondingly, the second prediction model is generated after being trained by a second sample set, and the second sample set contains the feature data of historical users and the result data of generating the target data of historical users without applying the preset execution conditions.

[0082] Exemplarily, the server can divide all customers into two groups according to whether they are marketed (treated): the treated group and the non-marketed group (control). For the two groups of customers, the result data of whether they are converted (responded) is used as the target label (label), and the training set and the validation set are divided.

[0083] For example, the marketed population can be taken out separately, and the converted population among them is used as label1, and the non-converted population is used as label0. Randomly sample 80% as the training set and 20% as the validation set, and train the first prediction model based on the XgBoost algorithm or other classification methods. For the non-marketed customers, use the same processing method to train the second prediction model.

[0084] It should be understood that the present disclosure embodiments do not limit the types of the first prediction model and the second prediction model. Exemplarily, the first prediction model and the second prediction model can be binary classification models.

[0085] S203. Determine the user type information based on the first conversion rate and the second conversion rate.

[0086] In this step, after the server obtains the first conversion rate and the second conversion rate, it can determine the user type information based on the first conversion rate and the second conversion rate.

[0087] It should be understood that the embodiments of the present disclosure do not limit how to determine the user type information. In some embodiments, the server can use the difference between the first conversion rate and the second conversion rate as the contribution value of the preset execution condition to the probability of generating the user's target data, and thus determine the user type information based on the first conversion rate, the second conversion rate, and the contribution value.

[0088] Exemplarily, the first prediction model and the second prediction model can be used to predict the conversion rate p of the user in the marketing and non-marketing scenarios, so as to obtain the first conversion rate p in the marketing scenario treated and the second conversion rate p in the non-marketing scenario control . By taking the difference between the first conversion rate and the second conversion rate, the contribution value lift of marketing to the probability of generating the user's transaction data can be obtained. Where lift = p treated - p control .

[0089] Subsequently, in some embodiments, if the first conversion rate is greater than the first threshold, the second conversion rate is less than or equal to the first threshold, and the contribution value is greater than the second threshold, the server can determine that the user belongs to the first user type.

[0090] If the first conversion rate is greater than the first threshold, the second conversion rate is greater than the first threshold, and the contribution value is greater than or equal to the third threshold and less than or equal to the second threshold, the server can determine that the user belongs to the second user type.

[0091] If the first conversion rate is less than or equal to the first threshold, the second conversion rate is less than or equal to the first threshold, and the contribution value is greater than or equal to the third threshold and less than or equal to the second threshold, the server can determine that the user belongs to the third user type.

[0092] If the first conversion rate is less than or equal to the first threshold, the second conversion rate is greater than the first threshold, and the contribution value is less than the third threshold, the server can determine that the user belongs to the fourth user type;

[0093] Wherein, the absolute value of the second threshold is equal to the absolute value of the third threshold.

[0094] It should be understood that the embodiments of the present disclosure do not limit the value of the first threshold. Exemplarily, the first threshold can be 0.5.

[0095] It should be understood that the embodiments of the present disclosure do not limit the value of the second threshold either. Exemplarily, the second threshold thres is the threshold at which lift changes significantly. It can be a positive decimal less than 1 and relatively close to 0, and can be determined by the quantile of the overall statistical distribution. Correspondingly, the third threshold can be -thres.

[0096] Exemplarily, if p_treated > 0.5, p_control ≤ 0.5, and lift > thres, then the user is a user of the first type. If p_treated > 0.5, p_control > 0.5, and -thres ≤ lift ≤ thres, then the user is a user of the second type. If p_treated ≤ 0.5, p_control ≤ 0.5, and -thres ≤ lift ≤ thres, then the user is a user of the third type. If p_treated ≤ 0.5, p_control > 0.5, and lift < -thres, then the user is a user of the fourth type.

[0097] Among them, the probability of users of the first user type generating target data increases when the preset execution condition is applied; the probability of users of the second user type generating target data is higher than the target upper limit value whether the preset execution condition is applied or not; the probability of users of the third user type generating target data is lower than the target lower limit value whether the preset execution condition is applied or not; the probability of users of the fourth user type generating target data decreases when the preset execution condition is applied.

[0098] Exemplarily, taking the preset execution condition as marketing to users as an example, the four types of users can correspondingly include marketing-sensitive people, natural conversion people, indifferent people, and counteractive people.

[0099] Among them, the active proportion of marketing-sensitive people is relatively low, but they are easily affected by marketing activities and generate active behaviors. For this part of users, they can be further stratified and managed according to their sensitivity to price, discount, profit concession, etc.

[0100] Natural conversion people are spontaneous active users. Even if the bank does not invest marketing resources in them, users will be spontaneously active, which is relatively high-quality. For this part of users, a similar user expansion model can be used to find more similar users in the enterprise information database and guide them to become bank users through online marketing, telephone marketing and other means.

[0101] Indifferent people are users who have been determined to have lost and cannot be recovered through marketing, or users who rarely read marketing messages, and no more marketing resources need to be invested.

[0102] The reaction group will be spontaneously active, but will be more averse to marketing interruptions. It is necessary to avoid marketing interruptions to this part of users and there is no need to invest marketing resources.

[0103] In some embodiments, the marketing-sensitive population can be further divided, so as to adopt different marketing methods based on the further divided types.

[0104] Exemplarily, the marketing-sensitive population can be further divided into two types. The first type is price-sensitive. When there are marketing activities related to subsidies and discounts, there will be corresponding active behaviors. The second type is price-insensitive. Whether there is a subsidy has little impact on it. For this type of users, only regular marketing reminders need to be carried out. After constructing price-sensitive and price-insensitive samples, classic classification methods can be used for user classification, which will not be elaborated here.

[0105] Exemplarily, when a user has one of the following behaviors, it can be classified as a price-sensitive label = 1 sample (here N is an adjustable threshold, and different values can be set based on different data distributions, usually not greater than 30)):

[0106] 1. The average loan interest rate in the past 3 months is within the lowest N% of the group;

[0107] 2. The subsidy utilization rate of loan documents in the past 3 months is above top N%;

[0108] 3. Actively share channel information to receive coupons.

[0109] Exemplarily, except for the above-mentioned price-sensitive customers, other customers are price-insensitive customers (label = 0).

[0110] In some embodiments, for price-sensitive customers, when conducting marketing, it is possible to focus on recommending activity rights and interests with greater preferential efforts to encourage them to have better active performances. For price-insensitive customers, through regular marketing reminders, multi-product cross-recommendations within the bank can be made for them.

[0111] The method for determining the type information of users provided by the embodiments of the present disclosure divides users in a relatively detailed manner, which can avoid waste of marketing resources. Combining models such as a marketing response model and a price sensitivity model, the existing and incremental users can be carefully divided, which helps the bank to tilt marketing resources to the most marketing-sensitive users in need.

[0112] The method for determining the type information of a user provided by the embodiments of the present disclosure first obtains the feature data of the user. Subsequently, the feature data of the user are respectively input into a first prediction model and a second prediction model to obtain a first conversion rate output by the first prediction model and a second conversion rate output by the second prediction model. The first prediction model is used to predict the probability of generating the target data of the user under preset execution conditions, and the second prediction model is used to predict the probability of generating the target data of the user without applying the preset execution conditions. Finally, the type information of the user is determined according to the first conversion rate and the second conversion rate. In this way, the types of users can be effectively divided under preset execution conditions.

[0113] On the basis of the above embodiments, after determining the type information of the user, the server can also query the transfer users of the users of the target type in the database according to the multi-dimensional vector of the user. The multi-dimensional vector is used to represent the association relationship between users in multiple dimensions, and the transfer user is the user to whom the preset execution conditions are applied for magnitude expansion. Figure 3 A flowchart of a method for mining transfer users provided by the embodiments of the present disclosure is shown in Figure 3 As shown, the method includes:

[0114] S301. According to the cosine similarity between the multi-dimensional vector of the user to be queried in the database and the multi-dimensional vector of the user of the target type.

[0115] It should be understood that the embodiments of the present disclosure do not limit the users of the target type. In some embodiments, the users of the target type may be the users of the second type described above, that is, the natural conversion population.

[0116] Exemplarily, let the multi-dimensional vector (embedding) vectors of two users ui and uj be (x_i1, x_i2,..., x_in) and (x_j1, x_j2,..., x_jn) respectively. Then the cosine similarity cos(u i ,u j ) between the users ui and uj can be determined by formula (1).

[0117]

[0118] It should be understood that the embodiments of the present disclosure do not limit how to determine the multi-dimensional vector of the user. In some embodiments, the server can determine the multi-dimensional vector of the user in the database according to the association information between users.

[0119] Exemplarily, the server may first select a target user from the database as the target node in the user relationship network. Secondly, the server may sequentially determine the next user node of the current end node in the associated user node sequence of the target node until the length of the node sequence of the target user reaches a preset sequence length. Thirdly, the server may generate an associated node array of the target node according to the associated user node sequence of the target node. Finally, the server may determine the multi-dimensional vectors of the users in the database according to the associated node array of the target node.

[0120] It should be understood that to determine the normalized transition probability of the current end node, the transition probability between the current end node and any associated user node may first be determined according to the association information between users. Subsequently, the transition probability between the current end node and any associated user node is normalized to determine the normalized transition probability of the current end node.

[0121] It should be understood that the embodiments of the present disclosure do not limit how to determine the transition probability between user nodes. Exemplarily, the server may first generate weight data between the current end node and any associated user node according to the association information between users and the identifiers of the users. Subsequently, the server determines a weight correction coefficient between the current end node and any associated user node according to the value of the shortest path distance between the current end node and any associated user node in the user relationship network. Finally, the server determines the transition probability between the current end node and any associated user node according to the weight data between the current end node and any associated user node and the weight correction coefficient between the current end node and any associated user node.

[0122] Among them, the value of the shortest path distance includes a first value, a second value, and a third value; there is a mapping relationship between the value of the shortest path distance and the weight correction coefficient;

[0123] If the associated user node is the previous node of the current end node, the value of the shortest path distance is the first value; if the associated user node and the current end node are adjacent nodes, the value of the shortest path distance is the second value; if the associated user node is not the previous node or the adjacent node of the current end node, the value of the shortest path distance is the third value.

[0124] S302. Determine whether the user to be queried is a transfer user according to the cosine similarity.

[0125] Among them, a transfer user may be understood as a user who can undergo marketing conversion.

[0126] It should be understood that the embodiments of the present disclosure do not limit how to determine whether the user to be queried is a transferred user according to the cosine similarity. In some embodiments, the greater the cosine similarity distance between two users, the more similar they are. Correspondingly, the marketing target customers can be found by finding the customers with the smallest cosine similarity distance from the users of the target type.

[0127] It should be noted that Figure 3 The method for mining transferred users can be applicable to the case where the quantity of users of the target type is small, for example, less than or equal to 200 people. When the quantity of users of the target type is large (for example, greater than 200 people), the method shown in Figure 4 can be adopted.

[0128] Figure 4 is a schematic flowchart of another method for mining transferred users provided by the embodiments of the present disclosure. As shown in Figure 4 , this method for mining transferred users includes:

[0129] S401. Sample the multi-dimensional vectors of the users of the target type and the multi-dimensional vectors of the users of the non-target type according to a preset sampling ratio to generate a third sample set. The multi-dimensional vectors of the users of the target type are the positive samples of the third sample set, and the multi-dimensional vectors of the users of the non-target type are the negative samples of the third sample set.

[0130] Exemplarily, the natural conversion population can be used as the seed customers (label = 1), and the non-natural conversion customers can be used as the negative samples (label = 0). Appropriate sampling is performed on the sample quantity so that label1:label0 is between 1:1 and 1:3. Subsequently, the samples are randomly divided, and 80% is taken as the training set, and the remaining 20% is taken as the validation set.

[0131] S402. Use the third sample to train the similar population expansion model.

[0132] Exemplarily, for the sample users in the training set and the validation set, their equity holding embedding vectors can be used as features for model training (XgBoost / LR feature combination, etc. can be used), and the binary classification model lookalike.model is saved.

[0133] S403. Input the multi-dimensional vector of the user to be queried in the database into the trained similar population expansion model, and obtain the population conversion probability output by the trained similar population expansion model.

[0134] Exemplarily, a large number of users with the same embedding vector features can be used as the users to be queried, so as to use the trained lookalike.model model for prediction to obtain the probability score of the conversion population of the user to be queried.

[0135] S404. Determine whether the user to be queried is a transferred user according to the population conversion probability.

[0136] It should be understood that the embodiments of the present disclosure do not limit how to determine whether the user to be queried is a transferred user according to the population conversion probability. In some embodiments, the population conversion probability can be compared with a threshold.

[0137] Exemplarily, if score≥thres, it can be determined that the user to be queried is a transferred user; if score<thres, it can be determined that the user to be queried is not a transferred user.

[0138] It should be understood that Figure 3 and Figure 4 The method for mining transferred users provided uses the multi-dimensional vectors of users. Since the multi-dimensional vectors of users contain the holding homogeneity similarity and holding structure similarity between users, compared with the traditional method, it can mine the friendship or acquaintance relationship between users, thereby improving the transfer success rate of transferred users.

[0139] In Figure 3 and Figure 4 Based on the method for mining transferred users provided, the server can determine the multi-dimensional vectors of the users in the database according to the association information between users. The following will describe how to determine the multi-dimensional vectors of users.

[0140] Figure 5 FIG. is a schematic diagram of the association relationship between user nodes provided by the embodiments of the present disclosure. As Figure 5 shown, all the nodes in the figure represent an enterprise, the edges between the nodes represent the holding relationship, pointing from the investing enterprise to the controlled enterprise, and the weight of the edge represents the holding ratio.

[0141] As Figure 5 can be seen, the enterprises corresponding to the two nodes can include two similarity relationships.

[0142] The first similarity relationship. Considering that enterprise u has a neighbor relationship with s1, s2, s3, and s4, it can be considered that there is a certain similarity between enterprise u and enterprises s1, s2, s3, and s4, which is called homogeneity.

[0143] The second similarity relationship. Both u and s6 are the central nodes of the corresponding subgraphs and have the largest degrees in the corresponding subgraphs, and there is also a certain similarity, which can be called structural similarity.

[0144] It should be noted that to simultaneously discover homogeneity and structural similarity and reflect them in the embedding results, both depth-first traversal (DFS) and breadth-first traversal (BFS) are required. To better integrate the advantages of the two traversal methods, the node2vec algorithm can be used. This algorithm uses a random walk approach, which can balance depth-first traversal and breadth-first traversal to generate a traversal node queue composed of nodes. Then, using the traversal node queue as the context, the skip-gram method is used to obtain the embedding vector representation of each node.

[0145] Figure 6 Schematic diagram for determining the multi-dimensional vector of a user provided by an embodiment of the present disclosure, as Figure 6 shown, the method includes:

[0146] S501. Perform unique table encoding on the user's name.

[0147] It should be understood that the embodiments of the present disclosure do not limit how to encode the user's name, and it can be encoded according to a preset encoding order. Exemplarily, "Shenzhen Qianhai Weizhong Bank Co., Ltd." can be encoded as s5.

[0148] S502. Generate weight data between user nodes corresponding to each directed edge in the user relationship network.

[0149] In some embodiments, the weight data between user nodes may include a start node, an end node, and a weight system. Exemplarily, as Figure 5 shown, the weight data can be, for example, "u s1 0.7", "u s2 0.35", "u s3 0.65", etc.

[0150] S503. Determine the weight correction coefficient between nodes according to the value of the shortest path distance between two nodes in the user relationship network.

[0151] Among them, the value of the shortest path distance includes a first value, a second value, and a third value; there is a mapping relationship between the value of the shortest path distance and the weight correction coefficient;

[0152] If the associated user node is the previous node of the current end node, the value of the shortest path distance is the first value; if the associated user node and the current end node are adjacent nodes, the value of the shortest path distance is the second value; if the associated user node is neither the previous node of the current end node nor an adjacent node of the current end node, the value of the shortest path distance is the third value.

[0153] Exemplarily, if the currently located node is v and the previous node of v is t (there is a directed edge from t to v), then for the adjacent node x of the current node v, the weight correction coefficient can be defined as shown in formula (2):

[0154]

[0155] Where dtx represents the shortest path distance between x and the vertex t. There are only three cases for this shortest path distance: if it returns to node t again (regardless of the directionality of the edge), then dtx = 0; if x and t are directly adjacent, then dtx = 1; in other cases, dtx = 2.

[0156] It should be understood that the specific values of p and q can be specified in advance and are not limited here.

[0157] S504. Determine the transition probability between nodes according to the weight data between user nodes and the weight correction coefficient between nodes.

[0158] Exemplarily, the transition probability between nodes can be determined by formula (3).

[0159] π(v,x) = α(t,x)·w vx (3)

[0160] Where wvx is the weight of the edge between nodes v and x in the user relationship network, and π(v,x) is the transition probability from node v to node x.

[0161] Figure 7 This is a schematic diagram of the transition probability between nodes provided by an embodiment of the present disclosure. As Figure 7 shown, s7 is the current node and s6 is the previous node of s7. Then for the two adjacent nodes of s7: s8 and s5, their transition probabilities are respectively: π(s_7,s_6) = 1 / p * 60%, π(s_7,s_8) = 1 * 15%, π(s_7,s_5) = 1 / q * 25%.

[0162] S505. Determine the normalized probability between nodes according to the transition probability between nodes.

[0163] Exemplarily, for each adjacent node xi of node v, obtain the transition probability πi and perform normalization of the transition probability

[0164] S506. Select a target user from the database as the target node in the user relationship network.

[0165] Exemplarily, from Figure 5 the user relationship network shown, a node t can be randomly selected, and an adjacent node v of t can be selected as the target node, where there is a directed edge from t to v.

[0166] S507. Sequentially determine the next user node of the current last node in the associated user node sequence of the target node until the length of the node sequence of the target user reaches the preset sequence length.

[0167] The embodiments of the present disclosure do not limit how to determine the next user node. In some embodiments, the next user node can be determined by determining the normalized transition probability of the current last node and performing weighted sampling on the associated nodes of the current last node.

[0168] Among them, the weighted sampling can specifically be alias sampling.

[0169] Exemplarily, for all adjacent nodes xi of the target node v, calculate its normalized transition probability through pi, and perform node sampling based on alias sampling to obtain the next node xi. At this time, the sequence is (v, x1). Subsequently, repeat the above process for the last node of the node sequence of the target user to obtain the next user node, resulting in (v, x 1 , x 2 ). By predefining the sequence length to be obtained as m + 1 and repeating the above process m times, the sequence result of node v can be obtained: (v, x1, x2,... xm).

[0170] S508. Generate an associated node array of the target node according to the associated user node sequence of the target node.

[0171] Exemplarily, by predefining the required number of sequences M, M adjacent nodes of node t can be selected, and the user node sequences of each of the M adjacent nodes of t can be calculated accordingly. By combining the user node sequences of each of the M adjacent nodes of t, the following associated node array of the target node can be obtained:

[0172] v1, x11, x12,... x1m

[0173] v2, x21, x22,... x2m

[0174] ...

[0175] vM, xM1, xM2,... xMm

[0176] S509. Determine the multi-dimensional vector of the user in the database according to the associated node array of the target node.

[0177] Exemplarily, the associated node array of the target node can be input into word2vec to obtain the n-dimensional (n can be manually set) embedding vector representation of each node, and the format is as follows:

[0178] “id1 0.13716 0.05973 -0.05692 0.34796…

[0179] id2 0.55362 -0.24561 0.67832 0.89571…

[0180] …”

[0181] In the embodiments of the present disclosure, a better application method is proposed for relationship data between enterprises (such as equity holding relationships, supply chain relationships, etc.), thereby improving the accuracy of determining transferred users, and to a great extent, it can solve the problems of difficult customer acquisition and difficult activation of existing customers in the current banking industry.

[0182] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.

[0183] Figure 8 FIG. is a schematic structural diagram of a device for determining the type information of a user provided by an embodiment of the present disclosure. The device for determining the type information of the user can be implemented by software, hardware, or a combination of both to execute the method for determining the type information of the user in the above embodiment. As Figure 8 shown, the device 600 for determining the type information of the user includes: an acquisition module 601, a prediction module 602, and a determination module 603.

[0184] The acquisition module 601 is configured to acquire feature data of a user.

[0185] The prediction module 602 is configured to input the feature data of the user into a first prediction model and a second prediction model respectively, and obtain a first conversion rate output by the first prediction model and a second conversion rate output by the second prediction model. The first prediction model is used to predict the probability of generating target data of the user under preset execution conditions, and the second prediction model is used to predict the probability of generating target data of the user without applying the preset execution conditions.

[0186] The determination module 603 is configured to determine the type information of the user according to the first conversion rate and the second conversion rate.

[0187] The device for determining the type information of the user provided by the embodiments of the present disclosure can perform the actions of the method for determining the type information of the user in the above embodiments. The implementation principle and technical effects are similar and will not be elaborated herein.

[0188] Figure 9The structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 9 shown, the electronic device may include: at least one processor 701 and a memory 702. Figure 9 The electronic device shown takes one processor as an example.

[0189] The memory 702 is used to store a program. Specifically, the program may include program code, and the program code includes computer operation instructions.

[0190] The memory 702 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0191] The processor 701 is used to execute the computer execution instructions stored in the memory 702 to implement the method for determining the type information of the above user;

[0192] Among them, the processor 701 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present disclosure.

[0193] Optionally, in specific implementation, if the communication interface, the memory 702 and the processor 701 are independently implemented, the communication interface, the memory 702 and the processor 701 may be interconnected through a bus and complete communication with each other. The bus may be an industry standard architecture (ISA for short) bus, a peripheral component interconnect (PCI for short) bus, or an extended industry standard architecture (EISA for short) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.

[0194] Optionally, in specific implementation, if the communication interface, the memory 702 and the processor 701 are integrated on a chip, the communication interface, the memory 702 and the processor 701 may complete communication through an internal interface.

[0195] An embodiment of the present disclosure also provides a chip, including a processor and an interface. The interface is used to input and output the data or instructions processed by the processor. The processor is used to execute the method provided in the above method embodiments.

[0196] The present disclosure also provides a computer-readable storage medium, which may include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs. Specifically, program information is stored in the computer-readable storage medium, and the program information is used for the method for determining the type information of the above-mentioned user.

[0197] The embodiments of the present disclosure also provide a program, which is used to execute the method for determining the type information of the user provided by the above method embodiments when executed by a processor.

[0198] The embodiments of the present disclosure also provide a program product, such as a computer-readable storage medium. Instructions are stored in the program product. When it runs on a computer, the computer is caused to execute the method for determining the type information of the user provided by the above method embodiments.

[0199] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. A method for determining the type information of a user, Characterized in that, The method includes: Obtaining the characteristic data of the user; Inputting the characteristic data of the user into a first prediction model and a second prediction model respectively, obtaining a first conversion rate output by the first prediction model and a second conversion rate output by the second prediction model. The first prediction model is used to predict the probability of generating the target data of the user under a preset execution condition, and the second prediction model is used to predict the probability of generating the target data of the user without applying the preset execution condition; Taking the difference between the first conversion rate and the second conversion rate as the contribution value of the preset execution condition to the probability of generating the target data of the user; If the first conversion rate is greater than a first threshold, the second conversion rate is less than or equal to the first threshold, and the contribution value is greater than a second threshold, the server can determine that the user belongs to the first user type; If the first conversion rate is greater than the first threshold, the second conversion rate is greater than the first threshold, the contribution value is greater than or equal to a third threshold and less than or equal to the second threshold, the server can determine that the user belongs to the second user type; If the first conversion rate is less than or equal to the first threshold, the second conversion rate is less than or equal to the first threshold, the contribution value is greater than or equal to the third threshold and less than or equal to the second threshold, the server can determine that the user belongs to the third user type; If the first conversion rate is less than or equal to the first threshold, the second conversion rate is greater than the first threshold, and the contribution value is less than the third threshold, the server can determine that the user belongs to the fourth user type; wherein, the absolute values of the second threshold and the third threshold are equal; the probability of the user of the first user type generating the target data increases when the preset execution condition is applied; the probability of the user of the second user type generating the target data is higher than the target upper limit value whether the preset execution condition is applied or not; the probability of the user of the third user type generating the target data is lower than the target lower limit value whether the preset execution condition is applied or not; the probability of the user of the fourth user type generating the target data decreases when the preset execution condition is applied.

2. The method according to claim 1, Characterized in that, The first prediction model is generated after being trained by a first sample set, and the first sample set contains the characteristic data of historical users and the result data of generating the target data of the historical users under the preset execution condition; The second prediction model is generated after being trained by a second sample set, and the second sample set contains the characteristic data of historical users and the result data of generating the target data of the historical users without applying the preset execution condition.

3. The method according to claim 1 or 2, Characterized in that, After determining the type information of the user according to the first conversion rate and the second conversion rate, the method further includes: Query for transferred users of users of the target type in the database according to the multi-dimensional vector of the user. The multi-dimensional vector is used to represent the association relationship between users in multiple dimensions. The transferred user is the user to whom the preset execution condition is applied for magnitude expansion.

4. The method according to claim 3, wherein, the querying for transferred users of users of the target type in the database according to the multi-dimensional vector of the user includes: Determine whether the user to be queried is the transferred user according to the cosine similarity between the multi-dimensional vector of the user to be queried in the database and the multi-dimensional vector of the user of the target type.

5. The method according to claim 3, wherein, the querying for transferred users of users of the target type in the database according to the multi-dimensional vector of the user includes: Sample the multi-dimensional vectors of the users of the target type and the multi-dimensional vectors of the non-target type users according to a preset sampling ratio to generate a third sample set. The multi-dimensional vectors of the users of the target type are the positive samples of the third sample set, and the multi-dimensional vectors of the non-target type users are the negative samples of the third sample set; Use the third sample set to train the similar population expansion model; Input the multi-dimensional vector of the user to be queried in the database into the trained similar population expansion model and obtain the population conversion probability output by the trained similar population expansion model; Determine whether the user to be queried is the transferred user according to the population conversion probability.

6. The method according to claim 3, wherein, Before the querying for transferred users of users of the target type in the database according to the multi-dimensional vector of the user, the method further includes: Determine the multi-dimensional vectors of the users in the database according to the association information between the users.

7. The method according to claim 6, wherein, the determining the multi-dimensional vectors of the users in the database according to the association information between the users includes: Select a target user from the database as the target node in the user relationship network; Sequentially determine the next user node of the current end node in the associated user node sequence of the target node until the length of the node sequence of the target user reaches the preset sequence length; Generate the associated node array of the target node according to the associated user node sequence of the target node; Determine the multi-dimensional vectors of the users in the database according to the associated node array of the target node.

8. The method according to claim 7, wherein, the determining the next user node of the current end node in the associated user node sequence of the target node includes: Determine the normalized transition probability of the current end node and perform weighted sampling on the associated nodes of the current end node to determine the next user node.

9. The method according to claim 8, wherein, the determining the normalized transition probability of the current end node includes: Determine the transition probability between the current end node and any associated user node according to the association information between the users. Normalize the transition probability between the current end node and any associated user node to determine the normalized transition probability of the current end node.

10. The method according to claim 9, wherein, the determining the transition probability between the current end node and any associated user node according to the association information between users includes: generating weight data between the current end node and any associated user node according to the association information between users and the identifiers of the users; determining a weight correction coefficient between the current end node and any associated user node according to the value of the shortest path distance between the current end node and any associated user node in the user relationship network; determining the transition probability between the current end node and any associated user node according to the weight data between the current end node and any associated user node and the weight correction coefficient between the current end node and any associated user node.

11. The method according to claim 10, wherein, the value of the shortest path distance includes a first value, a second value and a third value; there is a mapping relationship between the value of the shortest path distance and the weight correction coefficient; if the associated user node is the previous node of the current end node, the value of the shortest path distance is the first value, if the associated user node and the current end node are adjacent nodes, the value of the shortest path distance is the second value, if the associated user node is not the previous node or an adjacent node of the current end node, the value of the shortest path distance is the third value.

12. An electronic device, wherein, comprising: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the method according to any one of claims 1-11.

13. A computer storage medium, wherein, the computer storage medium stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor to perform the method steps according to any one of claims 1-11.

Citation Information

Patent Citations

  • Game resource delivery method and device

    CN111973996A

  • Target object identification method and device

    CN112199600A

  • User information management method and device

    CN113763019A