Training of User Information Classification Model, User Information Classification Method and Device

By obtaining the correct and wrong mark sets of user information, calculating the interval and weight probability of mark pairs, and training the user information classification model, the problem of low accuracy in the multi-marking classification of user information is solved, and higher classification accuracy and customer image quality are achieved.

CN114897054BActive Publication Date: 2025-07-04AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210385791.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-07-04
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

The existing decision tree algorithm has low accuracy when classifying user information with multiple tags, and fails to effectively consider the weight problem between tags, which affects the accuracy of model prediction.

Method used

By obtaining the correct and wrong mark sets of user information, calculating the interval and weight probability of mark pairs, training the user information classification model, introducing calibration marks to sort the mark sets, considering the order relationship and weight size between marks, and optimizing linear function parameters.

Benefits of technology

It improves the accuracy of user information classification, outputs more accurate marking sequences, and improves the quality of customer portraits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114897054B_ABST
    Figure CN114897054B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for a user information classification model, a user information classification method and an apparatus. The training method for the user information classification model obtains a correct tag set and an incorrect tag set corresponding to user information, forms a first tag pair by combining one tag in the correct tag set of a user information with one tag in the incorrect tag set, trains the user information classification model by using the relationship between the first tag pairs, and introduces calibration tags, so that when the user information classification model inputs the user information to be classified, it can distinguish the correct tags from the incorrect tags of the user information to be classified and obtain the correct tag set of the user information to be classified. In addition, the user information classification model can be trained according to the weight size relationship between the tags, so that the user information classification model outputs a tag sequence arranged according to the importance degree of the tags to the user, thereby improving the accuracy of classifying the user information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly relates to a training method for a user information classification model, a user information classification method, and a device. Background Art

[0002] In recent years, with the rise of fintech, "smart banks" are gradually playing an increasingly important role in various banking operations and product R & D. Among them, customer portraits can endow banks with the function of providing intelligent decisions based on customers' user information and behavioral data. Therefore, the performance of the classification model for tokenizing customers in customer portraits plays a crucial role in the quality of customer portraits.

[0003] Currently, the decision tree algorithm can be used to implement multi-label classification of user information, but this method is not accurate enough when implementing multi-label classification of user information. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a training method for a user information classification model, a user information classification method, and a device to improve the accuracy when performing multi-label classification of user information.

[0005] To solve the above problems, the technical solutions provided by the embodiments of the present application are as follows:

[0006] A training method for a user information classification model, the method comprising:

[0007] Obtain a user data set, where the user data set includes a plurality of user information and a correct label set corresponding to each user information;

[0008] Determine an incorrect label set corresponding to each user information, where the incorrect label set corresponding to each user information is obtained by removing the correct label set corresponding to the user information from all label sets;

[0009] Obtain a first label pair corresponding to a target user information, where the first label pair is composed of any one label in the correct label set corresponding to the target user information and any one label in the incorrect label set corresponding to the target user information; the target user information is any one of the user information;

[0010] Calculate the interval of the first label pair corresponding to each user information;

[0011] With the goal of maximizing the interval and minimizing the first loss function, train a user information classification model and determine a calibration label; the user information classification model is used to sort the labels in the entire label set when inputting the user information to be classified, and determine the correct label set corresponding to the user information to be classified for the labels sorted before the calibration label;

[0012] Obtain a second label pair corresponding to the target user information, where the second label pair is composed of any two labels in the correct label set corresponding to the target user information;

[0013] With the goal of minimizing the second loss function, retrain the user information classification model, where the second loss function is determined according to the weight probabilities of the second label pairs corresponding to each user information and the corresponding actual weight probabilities; the user information classification model is also used to sort the labels in the correct label set corresponding to the user information to be classified according to the weight probabilities.

[0014] In a possible implementation, the user information classification model is composed of Q linear functions, where Q is the number of labels in the entire label set, Q is a positive integer, and the linear functions correspond one-to-one with the labels in the entire label set;

[0015] The training of the user information classification model includes:

[0016] Solve the optimal solutions of the linear parameters of each linear function.

[0017] In a possible implementation, the calculation of the interval of the first label pair corresponding to each user information includes:

[0018] Determine the decision boundary;

[0019] According to the distance from each user information and the corresponding correct label set to the decision boundary, calculate the interval of the first label pair corresponding to each user information.

[0020] A user information classification method, the method includes:

[0021] Obtain the user information to be classified;

[0022] Input the user information to be classified into the user information classification model to obtain the correct tag set corresponding to the user information to be classified; the user information classification model is used to sort the tags in all tag sets when the user information to be classified is input, determine the correct tag set corresponding to the user information to be classified as the tags before the calibration tag in the sorting, and sort the tags in the correct tag set corresponding to the user information to be classified according to the weight probability; the user information classification model is trained according to the above training method of the user information classification model; the calibration tag is obtained during the training of the user information classification model.

[0023] A training device for a user information classification model, the device includes:

[0024] A first acquisition unit, configured to acquire a user data set, where the user data set includes a plurality of user information and the correct tag set corresponding to each piece of user information;

[0025] A determination unit, configured to determine the incorrect tag set corresponding to each piece of user information, where the incorrect tag set corresponding to each piece of user information is obtained by removing the correct tag set corresponding to the user information from all tag sets;

[0026] A second acquisition unit, configured to acquire a first tag pair corresponding to the target user information, where the first tag pair is composed of any one tag in the correct tag set corresponding to the target user information and any one tag in the incorrect tag set corresponding to the target user information; the target user information is any one of the user information;

[0027] A calculation unit, configured to calculate the interval of the first tag pair corresponding to each piece of user information;

[0028] A first training unit, configured to train the user information classification model with the goal of maximizing the interval and minimizing the first loss function, and determine the calibration tag; the user information classification model is used to sort the tags in all tag sets when the user information to be classified is input, and determine the correct tag set corresponding to the user information to be classified as the tags before the calibration tag in the sorting;

[0029] A third acquisition unit, configured to acquire a second tag pair corresponding to the target user information, where the second tag pair is composed of any two tags in the correct tag set corresponding to the target user information;

[0030] A second training unit is configured to retrain the user information classification model with the objective of minimizing a second loss function, where the second loss function is determined based on the weight probabilities of the second label pairs corresponding to the respective user information and the corresponding actual weight probabilities; the user information classification model is further configured to rank the labels in the correct label set corresponding to the user information to be classified according to the weight probabilities.

[0031] In a possible implementation, the user information classification model consists of Q linear functions, where Q is the number of labels in the entire label set, Q is a positive integer, and the linear functions correspond one-to-one with the labels in the entire label set;

[0032] Training the user information classification model includes:

[0033] Solving for the optimal solutions of the linear parameters of each of the linear functions.

[0034] In a possible implementation, the calculation unit is specifically configured to:

[0035] Determine the decision boundary;

[0036] Calculate the margin of the first label pair corresponding to each user information based on the distance from each user information and the corresponding correct label set to the decision boundary.

[0037] A user information classification device, the device includes:

[0038] A fourth acquisition unit is configured to acquire user information to be classified;

[0039] A classification unit is configured to input the user information to be classified into the user information classification model to obtain the correct label set corresponding to the user information to be classified; the user information classification model is configured to rank the labels in the entire label set when the user information to be classified is input, determine the labels ranked before the calibration label as the correct label set corresponding to the user information to be classified, and rank the labels in the correct label set corresponding to the user information to be classified according to the weight probabilities; the user information classification model is trained according to the above training method of the user information classification model; the calibration label is obtained during the process of training the user information classification model.

[0040] An electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the above training method of the user information classification model, or the above user information classification method.

[0041] A computer-readable storage medium stores instructions which, when running on a terminal device, cause the terminal device to execute the training method of the user information classification model as described above, or the user information classification method as described above.

[0042] Therefore, the embodiments of the present application have the following beneficial effects:

[0043] The embodiments of the present application obtain a correct tag set and an incorrect tag set corresponding to user information, form a first tag pair by combining one tag in the correct tag set of a user information with one tag in the incorrect tag set, utilize the relationship between the first tag pairs to train the user information classification model, and introduce calibration tags, so that when inputting the user information to be classified by using the user information classification model, the correct tag and the incorrect tag of the user information to be classified can be distinguished, and the correct tag set of the user information to be classified can be obtained. In addition, the user information classification model can be trained according to the weight size relationship between the tags, so that the user information classification model outputs a tag sequence arranged according to the importance degree of the tags to the user, thereby improving the accuracy of classifying the user information. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flowchart of a training method of a user information classification model provided by an embodiment of the present application;

[0045] Figure 2 It is a schematic diagram of a training method of a user information classification model provided by an embodiment of the present application;

[0046] Figure 3 It is a flowchart of a user information classification method provided by an embodiment of the present application;

[0047] Figure 4 It is a schematic diagram of a training device of a user information classification model provided by an embodiment of the present application;

[0048] Figure 5 It is a schematic diagram of a user information classification device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the above objects, features and advantages of the present application more obvious and understandable, the embodiments of the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0050] To facilitate the understanding and interpretation of the technical solutions provided by the embodiments of the present application, the background technology of the embodiments of the present application will be described first.

[0051] The bank's customer portrait is a systematic description of user information in a specific business scenario, and models all aspects of user information. The full picture of all types of information of a customer can be abstracted to build a labeling system. The user information classification model of the customer portrait in the construction of the labeling system is a type of algorithm that labels user information, namely multi-label learning. The goal of multi-label learning is to learn a set of real-valued function mappings from feature space to label space, that is, input a data set, each sample in the data set includes an example and a label set corresponding to the example, and output a classification model after algorithm training. When an example of an unknown label is input into the classification model, the classification model can output a predicted label set. Combined with the application scenario of the embodiment of the present application, multi-label learning is that each sample in the data set includes user information and the corresponding correct label set, and a user information classification model is trained and generated. When the user information to be classified is input into the user information classification model, the predicted correct label set corresponding to the user information to be classified can be output.

[0052] At present, the user tagging system uses the technical solution of decision tree algorithm. When the occurrence probability of each tag corresponding to the user information is known, a decision tree is constructed to find the probability that the expected value of the net present value is greater than or equal to 0, and the tag set corresponding to the user information is obtained by using this tree branch structure. However, the accuracy of this method is low.

[0053] In addition, the inventors have found through research that the existing algorithm for building a user tagging system does not consider the weight of the relevant tags to the user, and the weight relationship between tags is one of the factors that affect the accuracy of the model's prediction of user tags. User information, as an example, corresponds to a tag set, in which each tag has a different "degree of description" for this user information, that is, each tag has a different degree of importance to the user information, which means that each tag should have a predetermined weight value. When the algorithm considers the weight value of the tag, it can output a more accurate prediction result.

[0054] Based on this, the embodiments of the present application provide a method for training a user information classification model, a method and device for classifying user information, and an apparatus. The user information classification model is trained using the correct tag set and the error tag set corresponding to the user information, and a calibration tag is introduced, so that the user information classification model can distinguish the correct tags and error tags of the user information to be classified when the user information to be classified is input, thereby improving the accuracy of user information classification. On the other hand, when training the user information classification model, the weight probability of the tag is considered, and the user information classification model can output a tag sequence arranged according to the importance of the tag to the customer (i.e., the weight probability), thereby showing better user tag mining performance.

[0055] To facilitate the understanding of the embodiments of the present application, a method for training a user information classification model provided by the embodiments of the present application will be described below with reference to the accompanying drawings.

[0056] See Figure 1 As shown, this figure is a flowchart of a method for training a user information classification model provided by the embodiments of the present application. As Figure 1 shown, the method may include S101 - S105:

[0057] S101: Obtain a user data set, where the user data set includes multiple user information and the correct label set corresponding to each user information.

[0058] The user data set may include multiple user information, and the user information includes basic information such as name, gender, and age. At the same time, each user information corresponds to a correct label set, and the labels in the correct label set are the labels corresponding to the user information. The label can be understood as a classification attribute corresponding to the user information. In actual applications, the label system, that is, which labels the user information can correspond to, can be constructed according to actual business requirements.

[0059] It should be noted that in the embodiments of the present application, the user information, the correct label set and the wrong label set corresponding to the user information do not involve the sensitive information of the user and can be obtained and used after the user's authorization. In one example, before obtaining the user information and the correct label set corresponding to the user information, a prompt message related to obtaining data usage authorization is displayed on the corresponding interface, and the user determines whether to agree to the authorization based on the prompt message.

[0060] S102: Determine the wrong label set corresponding to each user information. The wrong label set corresponding to each user information is obtained by removing the correct label set corresponding to the user information from all the label sets.

[0061] According to whether a certain label belongs to the correct label set corresponding to the current user information, the labels in all the label sets can be divided into a correct label set and a wrong label set. The correct label set is the label set composed of the labels corresponding to the user information, and the wrong label set is the label set composed of the labels that are not the labels corresponding to the user information. Then, for each user information, the wrong label set corresponding to the user information can be obtained by removing the correct label set corresponding to the user information from all the label sets.

[0062] S103: Obtain the first label pair corresponding to the target user information. The first label pair is composed of any one label in the correct label set corresponding to the target user information and any one label in the wrong label set corresponding to the target user information; the target user information is any one of the user information.

[0063] Embodiments of the present application can utilize binary labels when training a user information classification model based on a second-order strategy multi-label algorithm, taking into account the correlation between pairwise labels, and using the correlation between labels in each label pair to train the user information classification model.

[0064] Taking each user information as the target user information respectively, select any one label from the correct label set corresponding to the target user information and any one label from the wrong label set corresponding to the target user information to form the first label pair corresponding to the target user information, that is, the "correct - wrong" label pair. There is an order relationship between the two labels in the first label pair, and this order relationship can represent a sorting related or unrelated to the user information.

[0065] S104: Calculate the interval of the first label pair corresponding to each user information.

[0066] For the first label pair, the order relationship between the two labels in the first label pair can be determined by the distance between the sample point and the decision boundary, and then the interval of the first label pair corresponding to each user information is calculated according to this distance, so as to be used to guide the training of the user information classification model.

[0067] For the specific method of calculating the interval of the first label pair corresponding to each user information, please refer to the candidate embodiments and will not be elaborated here.

[0068] S105: Taking maximizing the interval and minimizing the first loss function as the goal, train the user information classification model and determine the calibration label; the user information classification model is used to sort the labels in the entire label set when inputting the user information to be classified, and determine the correct label set corresponding to the user information to be classified for the labels sorted before the calibration label.

[0069] Taking maximizing the interval and minimizing the first loss function as the goal, a user information classification model can be trained. At the same time, introducing a calibration label, after the user information to be classified is input into the user information classification model, the sorting of the labels in the entire label set can be obtained. Based on the calibration label, the labels before the calibration label are the relevant labels (i.e., correct labels) of the user information to be classified, and the labels after the calibration label are the irrelevant labels (i.e., wrong labels) of the user information to be classified, so that the correct label set corresponding to the user information to be classified can be obtained.

[0070] In a possible implementation, the user information classification model is composed of Q linear functions, where Q is the number of labels in the entire label set, Q is a positive integer, and the linear functions correspond one-to-one with the labels in the entire label set. Then training the user information classification model is specifically to solve the optimal solutions of the linear parameters of each linear function. After the user information to be classified is input into the user information classification model, the sorting of the labels in the entire label set can be obtained, which is sorted according to the values of the linear functions corresponding to each label.

[0071] In the embodiment of the present application, a correct tag set and an incorrect tag set corresponding to user information are obtained, a first tag pair is formed by one tag in the correct tag set of a user information and one tag in the incorrect tag set, the relationship between the first tag pairs is utilized to train a user information classification model, and a calibration tag is introduced. Thus, when inputting the user information to be classified by using the user information classification model, the correct tags and incorrect tags of the user information to be classified can be distinguished, and the correct tag set of the user information to be classified is obtained.

[0072] S106: Obtain a second tag pair corresponding to the target user information, where the second tag pair is composed of any two tags in the correct tag set corresponding to the target user information.

[0073] Since the weight size relationship between tags is also one of the factors affecting the prediction accuracy of user tags by the user information classification model, the embodiment of the present application can also use the second tag pair, that is, the "correct-correct" tag pair, to continue training the above-trained user information classification model.

[0074] S107: Retrain the user information classification model with the goal of minimizing the second loss function, where the second loss function is determined according to the weight probabilities of the second tag pairs corresponding to each user information and the corresponding actual weight probabilities; the user information classification model is also used to sort the tags in the correct tag set corresponding to the user information to be classified according to the weight probabilities.

[0075] For the second tag pair, that is, the "correct-correct" tag pair, the relative weight probability between the two tags in the "correct-correct" tag pair is defined based on the principle of the change of information entropy, so as to represent the relative weight size relationship between the two tags. Then, by calculating the order relationship between the tags and the relative weight size relationship, the user information classification model can output a tag sequence with higher accuracy.

[0076] The following Figure 2 is used to illustrate the training method of the user information classification model provided by the embodiment of the present application.

[0077] Let the input space be a d-dimensional example space: X = R d , the entire tag set: y = {1,..., Q}, where Q is the number of tags in the entire tag set.

[0078] The data set D = {(x1, Y1), (x2, Y2),...,(x m , Y m ), where |D| = m, x i ∈X, the i-th example x i represents the i-th user information and can be represented by a d-dimensional feature vector: (x i1 , xi2 ,..., x id ). Y i is the tag sequence corresponding to the i-th example (i.e., the correct tag set): (y1,..., y k ,...), and any component element y k ∈ y, and it is agreed that if a certain tag y t is an element of the tag sequence Y i , then it is represented by the "∈" symbol in the set, that is: y t ∈ Y i .

[0079] The trained user information classification model h: X → 2 y , that is: h = (f1, f2,..., f Q ). Among them, the linear function f k = <w k , x> + b k corresponds to a certain tag k in the entire tag set y, is the weight vector of the linear function corresponding to the tag k, is the bias of the linear function corresponding to the tag k, and training the user information classification model is to obtain the optimal solutions of the parameters w k and b k through quadratic optimization.

[0080] For a sample (x i , Y i ) in the data set D, that is, a user information and the corresponding correct tag set. If there is a tag R ∈ Y i , Among them, then the tag pair composed of the tag R and the tag U is the first tag pair, that is, the "correct - wrong" tag pair. When calculating the order relationship between tags, the optimal parameters can be obtained while maximizing the margin and minimizing the first loss function, and a calibration tag K0 is introduced as the boundary value of the obtained tag set. When using the user information classification model, the tags ranked before K0 are used as the correct tag set corresponding to the user information to be classified. Then, in practical applications, using the first tag pair to train the user information classification model can include the following steps:

[0081] Step 1, determine the decision boundary of the first tag pair. Take a certain example x i in the data set D. In the case where each correct tag R and wrong tag U of x i form the first tag pair, the corresponding decision boundary g(x i ) is:

[0082] <w R - w U, x i > +b R -b U = 0

[0083] Step 2, determine the margin of the first label pair. Based on the defined decision boundary, calculate the distance from the sample point (x i , Y i ) to the decision boundary. Then, the minimum value is the margin m(x i ) of the first label pair:

[0084]

[0085] From the margin of the sample point (x i , Y i ) defined above, the margin of the dataset D can be obtained:

[0086]

[0087] That is, the above S104 calculates the margin of the first label pair corresponding to each user information to determine the decision boundary, which may include: determining the decision boundary; calculating the margin of the first label pair corresponding to each user information according to the distance from each user information and the corresponding correct label set to the decision boundary.

[0088] Step 3, maximize the margin. In an ideal situation, the correct label set corresponding to the example x i should be the correct label selected according to the order relationship calculated for each label pair, that is, each label in Y i has a relatively higher ranking compared to each label in . Then, the margin formula of the dataset D should take a positive value, that is, <w R -w U , x i > +b R -b U > 0; by appropriately scaling the parameters, we can get <w R -w U , x i > +b R -b U ≥ 1. Then, the problem of maximizing the margin can be calculated as:

[0089]

[0090] In the case where the dataset is large enough, the above formula can be transformed into:

[0091]

[0092] To reduce the impact of the maximization operator on subsequent formula calculations and further simplify the formula, the summation operator is used to approximately replace the maximization operator. Then, the maximum margin formula can be transformed into:

[0093]

[0094] Satisfying the condition: <w R -w U ,x i >+b R -b U ≥1, where x i ∈X,

[0095] For this optimization objective with constraints, since not all constraints in the actual situation can be satisfied, slack variables ξ iRU : Another objective function is the first loss function R(f) to be minimized, and its calculation formula is:

[0096]

[0097] Satisfying the condition:

[0098] Parameter I is introduced to adjust these two objective functions. The above two objective functions (the maximum margin formula and the first loss function) are transformed into a quadratic programming problem for solving the optimal parameters w and b in the form of addition. The formula is:

[0099]

[0100] Satisfying the condition:

[0101] Thus, after solving the optimal parameters, the training of the user information classification model is completed.

[0102] Step 5, the calibration marker K0 determines the correct marker set. Take Using the correct marker and the boundary between the correct and incorrect markers with marker K0 as an example, and using the markers before K0 as the correct marker set in the example, the correct and incorrect markers of the customer can be distinguished.

[0103] Continue to combine Figure 2 , and further illustrate the training method of the user information classification model provided in the embodiments of the present application.

[0104] Calculate the relative weight magnitude relationship between the tags on the second tag pair, i.e., the "correct - correct" (R - R') tag pair. In fact, it is to further optimize the parameter w output during the above training process. Define the weight probability of the tag pair. The weight probability of the tag pair defined here is not used to directly determine whether the weight of tag R is greater than the weight of tag R', but is used to represent the probability value of the event that the weight of tag R is greater than the weight of tag R'. And use the difference between the predicted probability value and the actual weight relationship probability value to define the second loss function. Minimizing the second loss function can obtain the optimal solution of parameter w. In practical applications, using the second tag pair to train the user information classification model may include the following steps:

[0105] Step 1, determine the weight probability of the second tag pair, i.e., the (R - R') tag pair. It represents the probability that the weight of tag R is greater than the weight of tag R', and is denoted by the symbol P R,R' (f):

[0106]

[0107] Among them, The parameter is used to adjust the weight distribution ratio between the tags.

[0108] If the weight value of tag R is greater than the weight value of tag R', the weight probability P R,R' (f) should approach 1 infinitely; if the weight value of tag R is less than the weight value of tag R', then P R,R' (f) approaches 0 infinitely; if the weight value of tag R is equal to the weight value of tag R', then P R,R' (f) is equal to 0.5.

[0109] Assume that the actual weight probability between these two tags (R - R') is Then The formula of is:[[]]

[0110]

[0111] Among them,

[0112] Step 2, the objective function for calculating the weight magnitude relationship between the second tag pairs should be to compare the predicted weight probability value with the actual weight probability value and minimize the error between the two. Define the second loss function C R,R' (f),

[0113] Minimize C R,R' (f) to solve for the parameter w R :[[]]

[0114]

[0115] Define the loss function according to the weight probability value constructed by the model and the true weight probability value. Solving the weight size relationship between tags is to solve the parameter w in minimizing the second loss function. R The optimal solution.

[0116] Therefore, the user information classification model proposed in the embodiments of the present application can consider the order relationship and weight size relationship between tags when constructing the tokenization system of customers, so that the user information classification model outputs a more accurate tag sequence for a to-be-classified user information, that is, a tag sequence arranged according to the importance of tags to users can be output.

[0117] After training and generating the user information classification model, the user information classification model can be used for user information classification.

[0118] The following describes a user information classification method provided in the embodiments of the present application with reference to the accompanying drawings.

[0119] Refer to Figure 3 As shown, this figure is a flowchart of a user information classification method provided in the embodiments of the present application. As Figure 3 shown, the method may include S301 - S302:

[0120] S301: Obtain the to-be-classified user information.

[0121] The to-be-classified user information includes basic information such as name, gender, and age. In the embodiments of the present application, the to-be-classified user information does not involve sensitive information of users and can be obtained and used after user authorization. In one example, before obtaining the to-be-classified user information, a prompt message related to obtaining data usage authorization is displayed on the corresponding interface, and the user determines whether to agree to the authorization based on the prompt message.

[0122] S302: Input the to-be-classified user information into the user information classification model to obtain the correct tag set corresponding to the to-be-classified user information; the user information classification model is used to sort the tags in the entire tag set when inputting the to-be-classified user information, and determine the tags before the calibration tag in the sorting as the correct tag set corresponding to the to-be-classified user information; the user information classification model is trained according to the above training method of the user information classification model; the calibration tag is obtained during the process of training the user information classification model.

[0123] Input the to-be-classified user information into the user information classification model, and the user information classification model sorts the tags in the entire tag set, and determines the tags before the calibration tag in the sorting as the correct tag set corresponding to the to-be-classified user information. When the user information classification model is composed of Q linear functions, the user information classification model can sort the tags in the entire tag set according to the values of the linear functions corresponding to each tag.

[0124] For the training of the user information classification model and the obtaining of the calibration labels, reference may be made to the descriptions in the above embodiments, which will not be elaborated here.

[0125] In a possible implementation manner, the user information classification model is further configured to sort the labels in the correct label set corresponding to the user information to be classified according to the weight probability.

[0126] After the user information classification model is trained according to the second label pair, the user information classification model can also sort the labels in the correct label set corresponding to the user information to be classified according to the weight probability, so as to obtain a more accurate label sequence and effectively improve the accuracy of the customer portrait.

[0127] Based on the training method of a user information classification model provided in the above method embodiment, an embodiment of the present application further provides a training device for a user information classification model, which will be described below with reference to the accompanying drawings.

[0128] See Figure 4 As shown, this figure is a schematic structural diagram of a training device for a user information classification model provided in an embodiment of the present application. As Figure 4 shown, the training device for the user information classification model includes:

[0129] A first obtaining unit 401, configured to obtain a user data set, where the user data set includes a plurality of user information and a correct label set corresponding to each piece of the user information;

[0130] A determining unit 402, configured to determine an incorrect label set corresponding to each piece of the user information, where the incorrect label set corresponding to each piece of the user information is obtained by removing the correct label set corresponding to the user information from all the label sets;

[0131] A second obtaining unit 403, configured to obtain a first label pair corresponding to the target user information, where the first label pair is composed of any one label in the correct label set corresponding to the target user information and any one label in the incorrect label set corresponding to the target user information; the target user information is any one of the user information;

[0132] A calculating unit 404, configured to calculate the interval of the first label pair corresponding to each piece of the user information;

[0133] A first training unit 405, configured to train a user information classification model with the goal of maximizing the interval and minimizing a first loss function, and determine a calibration label; the user information classification model is configured to sort the labels in all the label sets when inputting the user information to be classified, and determine the labels sorted before the calibration label as the correct label set corresponding to the user information to be classified;

[0134] A third acquisition unit 406, configured to acquire a second tag pair corresponding to the target user information, where the second tag pair is composed of any two tags in the correct tag set corresponding to the target user information;

[0135] A second training unit 407, configured to retrain the user information classification model with the goal of minimizing a second loss function, where the second loss function is determined according to the weight probabilities of the second tag pairs corresponding to each user information and the corresponding actual weight probabilities; the user information classification model is further configured to sort the tags in the correct tag set corresponding to the user information to be classified according to the weight probabilities.

[0136] In a possible implementation manner, the user information classification model is composed of Q linear functions, where Q is the number of tags in the entire tag set, Q is a positive integer, and the linear functions correspond to the tags in the entire tag set one by one;

[0137] Training the user information classification model includes: solving the optimal solutions of the linear parameters of each linear function.

[0138] In a possible implementation manner, the calculation unit is specifically configured to:

[0139] Determine a decision boundary;

[0140] According to the distance from each user information and the corresponding correct tag set to the decision boundary, calculate the interval of the first tag pair corresponding to each user information.

[0141] Based on the user information classification method provided in the above method embodiments, an embodiment of the present application further provides a user information classification device, which will be described below with reference to the accompanying drawings.

[0142] See Figure 5 As shown in the figure, the figure is a schematic structural diagram of a user information classification device provided by an embodiment of the present application. As Figure 5 shown, the user information classification device includes:

[0143] A fourth acquisition unit 501, configured to acquire user information to be classified;

[0144] A classification unit 502 is configured to input the user information to be classified into the user information classification model to obtain a correct tag set corresponding to the user information to be classified; the user information classification model is configured to sort tags in all tag sets when the user information to be classified is input, determine the tags sorted before the calibration tag as the correct tag set corresponding to the user information to be classified, and sort the tags in the correct tag set corresponding to the user information to be classified according to weight probabilities; the user information classification model is obtained by training according to the training method of the user information classification model as described above; the calibration tag is obtained during the training of the user information classification model.

[0145] In addition, an embodiment of the present application further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the training method of the user information classification model as described above, or the user information classification method as described above is implemented.

[0146] An embodiment of the present application further provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a terminal device, the terminal device is caused to execute the training method of the user information classification model as described above, or the user information classification method as described above.

[0147] The embodiment of the present application obtains a correct tag set and an incorrect tag set corresponding to user information, forms a first tag pair by combining one tag in the correct tag set of a user information with one tag in the incorrect tag set, trains the user information classification model by using the relationship between the first tag pairs, and introduces a calibration tag, so that the user information classification model can distinguish the correct tags and incorrect tags of the user information to be classified when the user information to be classified is input, and obtain the correct tag set of the user information to be classified. In addition, the user information classification model can be trained according to the weight magnitude relationship between tags, so that the user information classification model outputs a tag sequence arranged according to the importance degree of tags to the user, thereby improving the accuracy of classifying user information.

[0148] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method part.

[0149] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0150] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0151] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0152] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A training method for a user information classification model, characterized in that The method includes: Obtaining a user data set, where the user data set includes multiple user information and a correct tag set corresponding to each piece of the user information; Determining an incorrect tag set corresponding to each piece of the user information, where the incorrect tag set corresponding to each piece of the user information is obtained by removing the correct tag set corresponding to the user information from all tag sets; Obtaining a first tag pair corresponding to the target user information, where the first tag pair is composed of any one tag in the correct tag set corresponding to the target user information and any one tag in the incorrect tag set corresponding to the target user information; the target user information is any one of the user information; Calculating the interval of the first tag pair corresponding to each piece of the user information; Training a user information classification model with the goal of maximizing the interval and minimizing a first loss function, and determining a calibration tag; the user information classification model is used to sort the tags in all tag sets when inputting the user information to be classified, and determine the correct tag set corresponding to the user information to be classified as the tags sorted before the calibration tag; where the user information classification model is composed of Q linear functions, Q is the number of tags in all tag sets, Q is a positive integer, and the linear functions correspond one-to-one with the tags in all tag sets; training the user information classification model includes: solving the optimal solutions of the linear parameters of each linear function; the user information classification model is specifically used to: sort the tags in all tag sets according to the values of the linear functions corresponding to each tag; Obtaining a second tag pair corresponding to the target user information, where the second tag pair is composed of any two tags in the correct tag set corresponding to the target user information; Retraining the user information classification model with the goal of minimizing a second loss function, where the second loss function is determined according to the weight probabilities of the second tag pairs corresponding to each piece of the user information and the corresponding actual weight probabilities; the user information classification model is also used to sort the tags in the correct tag set corresponding to the user information to be classified according to the weight probabilities.

2. The method according to claim 1, characterized in that The calculating the interval of the first tag pair corresponding to each piece of the user information includes: Determining a decision boundary; Calculating the interval of the first tag pair corresponding to each piece of the user information according to the distance from each piece of the user information and the corresponding correct tag set to the decision boundary.

3. A method for classifying user information, characterized in that, The method includes: Obtaining the user information to be classified; Input the user information to be classified into the user information classification model to obtain the correct tag set corresponding to the user information to be classified; the user information classification model is used to sort the tags in the entire tag set when inputting the user information to be classified, determine the correct tag set corresponding to the user information to be classified for the tags sorted before the calibration tag, and sort the tags in the correct tag set corresponding to the user information to be classified according to the weight probability; wherein, when the user information classification model is composed of Q linear functions, Q is the number of tags in the entire tag set, Q is a positive integer, the linear functions correspond to the tags in the entire tag set one by one, and the user information classification model is specifically used to: sort the tags in the entire tag set according to the values of the linear functions corresponding to each tag; the user information classification model is trained according to the training method of the user information classification model according to any one of claims 1-2; the calibration tag is obtained during the process of training the user information classification model.

4. A training device for a user information classification model, characterized in that, The device includes: A first acquisition unit, configured to acquire a user data set, where the user data set includes a plurality of user information and the correct tag set corresponding to each piece of user information; A determination unit, configured to determine the incorrect tag set corresponding to each piece of user information, where the incorrect tag set corresponding to each piece of user information is obtained by removing the correct tag set corresponding to the user information from the entire tag set; A second acquisition unit, configured to acquire a first tag pair corresponding to the target user information, where the first tag pair is composed of any one tag in the correct tag set corresponding to the target user information and any one tag in the incorrect tag set corresponding to the target user information; the target user information is any one of the user information; A calculation unit, configured to calculate the interval of the first tag pair corresponding to each piece of user information; A first training unit, configured to train the user information classification model with the goal of maximizing the interval and minimizing the first loss function, and determine the calibration tag; the user information classification model is used to sort the tags in the entire tag set when inputting the user information to be classified, and determine the correct tag set corresponding to the user information to be classified for the tags sorted before the calibration tag; wherein, the user information classification model is composed of Q linear functions, Q is the number of tags in the entire tag set, Q is a positive integer, the linear functions correspond to the tags in the entire tag set one by one; training the user information classification model includes: solving the optimal solution of the linear parameters of each linear function; the first training unit is specifically used to: sort the tags in the entire tag set according to the values of the linear functions corresponding to each tag; A third acquisition unit, configured to acquire a second tag pair corresponding to the target user information, where the second tag pair is composed of any two tags in the correct tag set corresponding to the target user information; A second training unit for retraining the user information classification model with the goal of minimizing a second loss function, where the second loss function is determined based on the weight probabilities of the second label pairs corresponding to each piece of the user information and the corresponding actual weight probabilities; the user information classification model is further used to sort the labels in the correct label set corresponding to the user information to be classified according to the weight probabilities.

5. The device according to claim 4, characterized in that, Specifically, the calculation unit is configured to: Determine a decision boundary; Calculate the margin of the first label pair corresponding to each piece of the user information according to the distance from each piece of the user information and the corresponding correct label set to the decision boundary.

6. A user information classification device, characterized in that, The apparatus includes: A fourth acquisition unit for acquiring the user information to be classified; A classification unit for inputting the user information to be classified into the user information classification model to obtain the correct label set corresponding to the user information to be classified; the user information classification model is used to sort the labels in the entire label set when the user information to be classified is input, determine the labels sorted before the calibration label as the correct label set corresponding to the user information to be classified, and sort the labels in the correct label set corresponding to the user information to be classified according to the weight probabilities; wherein, when the user information classification model is composed of Q linear functions, Q is the number of labels in the entire label set, Q is a positive integer, the linear functions correspond to the labels in the entire label set one by one, and the user information classification model is specifically configured to: sort the labels in the entire label set according to the values of the linear functions corresponding to each label; the user information classification model is trained according to the training method of the user information classification model according to any one of claims 1-2; the calibration label is obtained during the training of the user information classification model.

7. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the training method of the user information classification model according to any one of claims 1-2, or the user information classification method according to claim 3.

8. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions run on the terminal device, the terminal device is caused to execute the training method of the user information classification model according to any one of claims 1-2, or the user information classification method according to claim 3.