Classification model training method, user classification method, device and electronic equipment
By training and validating multiple initial models and obtaining positive and negative samples from multiple training sets, the system identifies the time minors spend using apps, solving the problem of limiting game time for minors and improving identification accuracy and user experience.
Patent Information
- Application Number
- CN202210435625.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-04-24
AI Technical Summary
In the current technology, it is impossible to effectively limit the time minors spend using game apps, and some users bypass control measures by falsifying their age, making it impossible to accurately identify and limit the game time of minors, which affects their physical and mental health and studies.
By acquiring positive and negative training samples from multiple training sets, a classification model is obtained by training multiple initial models. Based on the validation set, the parameters are validated to determine the target classification model, which is used to identify user types, thus avoiding the need to perform complex real-name authentication operations for each user.
It improves the accuracy of identifying the time minors spend using the app, reduces user annoyance, enhances the user experience, and ensures that the model still has good classification performance even when there is a large difference between positive and negative samples.
Smart Images

Figure CN117009797B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and more particularly, to a classification model training method, a user classification method, an apparatus and an electronic device. BACKGROUND
[0002] With the development of electronic devices, various application programs (APPs) are emerging, such as shopping APPs, video APPs, live broadcast APPs, and game APPs. More and more people use such APPs and even become addicted to such APPs. Currently, such users who are addicted to APPs usually need to limit the time length of using the APPs according to their own conditions. For example, the time length of using the APPs is limited according to age, physical condition, consumption ability, work condition, and the like. Taking the limitation according to age as an example, since the self-control ability of minors is weak, with the increase of game APPs, minors have no self-control for playing games, and the phenomenon of minors playing games for several hours in succession occurs, thereby causing certain influence on the physical and mental health and studies of the minors.
[0003] In the related art, in order to avoid the login time length in the game application being limited, some users may perform age fraud when registering the game application, which leads to inconsistency with the actual age stage of the user. Correspondingly, if the user is controlled in the game application according to the age registered by the user in the game application, effective control may not be performed. SUMMARY
[0004] In view of the above problems, the embodiments of the present application propose a classification model training method, a user classification method, an apparatus and an electronic device to improve the above problems.
[0005] In a first aspect, the embodiments of the present application provide a classification model training method, which comprises: obtaining at least two initial models and a plurality of training sets obtained based on training samples, wherein the training samples comprise positive and negative training samples obtained in a registration stage and positive and negative training samples obtained in a use stage, each training set comprises positive and negative training samples corresponding to at least one stage, the positive training samples are feature data of a first type of user, and the negative training samples are feature data of a second type of user; for each initial model, training the initial model by using the positive and negative training samples in each training set to obtain a plurality of classification models corresponding to the initial model and training parameters of each classification model, and each classification model corresponds to one training set; determining a target initial model based on the training parameters of the classification models corresponding to each initial model; obtaining a verification set corresponding to each training set, wherein each verification sample in the verification set is obtained based on a user type identification operation of a user in the registration stage or the use stage, and the feature types of the training samples in the training set and the verification samples in the corresponding verification set are the same; for each classification model corresponding to the target initial model, verifying the classification model based on the verification set corresponding to the training set used to train the classification model to obtain verification parameters of the classification model; and determining a target classification model based on the verification parameters of the classification models.
[0006] In a second aspect, the embodiments of the present application provide a classification method, which comprises: obtaining feature data of a user to be classified; processing the feature data by using a target classification model to obtain a user type corresponding to the feature data of the user to be classified, wherein the user type is a first type of user or a second type of user.
[0007] In a third aspect, an embodiment of the present application provides a classification model training apparatus, comprising a first obtaining module, a training module, an initial model determining module, a second obtaining module, a verifying module, and a classification model determining module. The first obtaining module is configured to obtain at least two initial models and a plurality of training sets obtained based on training samples, wherein the training samples comprise positive and negative training samples obtained in a registration stage and positive and negative training samples obtained in a use stage, each training set comprises positive and negative training samples corresponding to at least one stage, the positive training samples are feature data of a first type of user, and the negative training samples are feature data of a second type of user; the training module is configured to train each initial model using the positive and negative training samples in each training set, to obtain a plurality of classification models corresponding to the initial model and training parameters of each classification model, and each classification model corresponds to one training set; the initial model determining module is configured to determine a target initial model based on the training parameters of the classification models corresponding to each initial model; the second obtaining module is configured to obtain a verification set corresponding to each training set, wherein each verification sample in the verification set is obtained based on a user type identification operation of a user in the registration stage or the use stage, and the feature types of the training samples in the training set and the verification samples in the corresponding verification set are the same; the verifying module is configured to verify each classification model corresponding to the target initial model based on the verification set corresponding to the training set used to train the classification model, to obtain verification parameters of the classification model; and the classification model determining module is configured to determine a target classification model based on the verification parameters of each classification model.
[0008] In an implementation, the training module is further configured to train each initial model using positive and negative training samples selected from each training set according to a target preset ratio, to obtain a first classification model corresponding to each preset classification threshold and training parameters of each first classification model, wherein one preset classification threshold corresponds to one first classification model, and the target preset ratio is used to represent the ratio of positive training samples to negative training samples in the training set. The initial model determining module is further configured to determine a target preset classification threshold and a target initial model according to the training parameters of the first classification models corresponding to each initial model under the plurality of preset classification thresholds, wherein the first classification model corresponding to the target preset classification threshold and trained based on the target initial model is the classification model corresponding to the target initial model.
[0009] In an implementation, the classification model training apparatus further comprises a preset ratio determining module, and the training module is further configured to train each initial model by selecting a plurality of sets of positive and negative training samples from the training set according to a plurality of preset ratios, to obtain a second classification model corresponding to a set of positive and negative training samples when a classification threshold is a set value and a training parameter of the second classification model, wherein each preset ratio corresponds to a set of positive and negative training samples. The preset ratio determining module is configured to determine the target preset ratio according to the training parameters of the second classification models corresponding to the plurality of preset ratios.
[0010] In an implementation, the training parameters of the second classification model are at least two, and the preset ratio determining module is further configured to perform weighted summation on the training parameters of the second classification models corresponding to each preset ratio to obtain a training parameter calculation value corresponding to each preset ratio, wherein the greater the value of the same training parameter, the greater the corresponding weight value; and select the preset ratio corresponding to the maximum training parameter calculation value as the target preset ratio.
[0011] In an implementation, the initial model determining module comprises a calculation submodule, a classification threshold determining submodule, and an initial model determining submodule. The calculation submodule is configured to perform weighted summation on the training parameters of the first classification models corresponding to each initial model under each preset classification threshold to obtain a training parameter calculation value of each initial model under the preset classification threshold, and determine a target parameter calculation value based on the training parameter values of each initial model, wherein the greater the value of the same training parameter, the greater the corresponding weight value. The classification threshold determining submodule is configured to determine the preset classification threshold corresponding to the maximum target parameter calculation value as a target preset classification threshold. The initial model determining submodule is configured to determine the initial model corresponding to the maximum training parameter calculation value in the training parameter calculation values of each initial model corresponding to the target classification threshold as a target initial model.
[0012] In an implementation, the verification parameters of the first classification model are at least two, and the classification model determining module comprises a calculation submodule and a classification model determining submodule. The calculation submodule is configured to perform weighted summation on the verification parameters of each classification model to obtain a verification parameter calculation value corresponding to each classification model, wherein the greater the value of the same verification parameter, the greater the corresponding weight value. The classification model determining submodule is configured to select the classification model corresponding to the maximum verification parameter calculation value as a target classification model.
[0013] In an implementation, the first type of user includes a first age group, the second type of user includes a second age group, and the second obtaining module includes a threshold determining submodule and a second obtaining submodule. The threshold determining submodule is configured to determine a delivery threshold of each classification model corresponding to the target initial model according to a ratio of positive training samples to negative training samples in the training samples. The second obtaining submodule is configured to, for each classification model corresponding to the target initial model, deploy the classification model on a server, and use the classification model to perform classification prediction on a plurality of to-be-predicted feature data, to obtain a probability that each to-be-predicted feature data classification prediction result is a first type of user, sort the probabilities of the to-be-predicted feature data from large to small, and obtain each target to-be-predicted feature data and a user type corresponding to each target registered user that are arranged before the delivery threshold of the classification model, as a verification sample in a verification set corresponding to a training set used to train the classification model. The user type corresponding to each target registered user is obtained based on performing a user type identification operation on a user in a registration stage or a use stage.
[0014] In an implementation, the first obtaining module includes a first obtaining submodule and an expansion submodule. The first obtaining submodule is configured to obtain training samples, the training samples including positive training samples and negative training samples, the positive training samples including positive and negative training samples obtained in a registration stage and positive training samples obtained in a use stage, and the negative training samples including negative training samples obtained in the registration stage and negative training samples obtained in the use stage. The expansion submodule is configured to, when a ratio of the positive training samples to the negative training samples in the training samples is less than a preset threshold, expand the positive training samples in the training samples to obtain expanded training samples, and obtain a plurality of training sets based on the expanded training samples.
[0015] In an implementation, the expansion submodule is further configured to expand the positive training samples in the training samples based on an oversampling algorithm to obtain the expanded training samples.
[0016] In an implementation, the training parameters include at least one of a precision rate, a recall rate, an AUC value, and an F1 score, and the model parameters of the verification parameters include at least one of the precision rate, the recall rate, the AUC value, and the F1 score.
[0017] In a fourth aspect, an embodiment of the present application provides a user classification device. The device includes a feature data obtaining module and a user classification module. The feature data obtaining module is configured to obtain feature data of a to-be-classified user. The user classification module is configured to process the feature data using a target classification model to obtain a user type corresponding to the feature data of the to-be-classified user, the user type being a first type of user or a second type of user.
[0018] In a fifth aspect, the embodiments of the present application further provide an electronic device, comprising: a processor; a memory, wherein the memory stores computer readable instructions, and the computer readable instructions are executed by the processor to implement the classification model training method or the user classification method as described above.
[0019] In a sixth aspect, the embodiments of the present application provide a computer readable storage medium, which stores computer readable instructions, and the computer readable instructions are executed by a processor to implement the classification model training method or the user classification method as described above.
[0020] In a seventh aspect, the embodiments of the present application provide a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device acquires the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the method as described above.
[0021] The classification model training method, the user classification method, the device and the electronic device provided by the embodiments of the present application can obtain at least two initial models and a plurality of training sets obtained based on training samples, so that in the model training process, for each initial model, the positive and negative training samples in each training set are used to train the initial model respectively, a plurality of classification models corresponding to the initial model and the training parameters of each classification model are obtained, a target initial model is determined based on the training parameters of the classification model corresponding to each initial model, so that the type of the selected initial model can be effectively ensured to be the optimal type for classification. In addition, since each verification sample in the verification set is obtained based on the user type identification operation of the user in the registration stage or the use stage, the sample proportion of the obtained positive and negative verification samples usually has a large difference, so that in the verification stage, for each classification model corresponding to the target initial model, the verification set corresponding to the training set used to train the classification model is used to verify the classification model to obtain the verification parameter of the classification model, and the target classification model is determined based on the verification parameters of the classification models, so that the model can also have good classification effect under the condition that the positive and negative verification samples have a large difference, thereby improving the accuracy of the user type obtained when the target classification model is used to identify the feature data subsequently. BRIEF DESCRIPTION OF DRAWINGS
[0022] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles behind the present application. It is apparent that the accompanying drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0023] Figure 1 An authentication interface schematic diagram when a user logs in an application is shown;
[0024] Figure 2 A face recognition authentication interface schematic diagram when a user logs in an application is shown;
[0025] Figure 3 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied is shown;
[0026] Figure 4 A flowchart of a classification model training method according to an embodiment of the present application is shown;
[0027] Figure 5 A schematic diagram of a training sample according to an embodiment of the present application is shown;
[0028] Figure 6 A schematic diagram of selecting K neighbor samples of a positive training sample according to an embodiment of the present application is shown;
[0029] Figure 7 A schematic diagram of selecting a target neighbor sample of a positive training sample according to an embodiment of the present application is shown;
[0030] Figure 8 A flowchart of a classification model training method according to another embodiment of the present application is shown;
[0031] Figure 9 A flowchart of a user classification method according to an embodiment of the present application is shown;
[0032] Figure 10 A flowchart of an application scenario of a classification model training method according to an embodiment of the present application is shown;
[0033] Figure 11 A connection block diagram of a classification model training apparatus according to an embodiment of the present application is shown;
[0034] Figure 12 A connection block diagram of a user classification apparatus according to an embodiment of the present application is shown;
[0035] Figure 13 A structural schematic diagram of an electronic device suitable for implementing the embodiments of the present application is shown. DETAILED DESCRIPTION
[0036] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations can be implemented in any
[0037] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the
[0038] The block diagrams in the drawings show only the functionality of the embodiments and do not imply any particular physical or architectural arrangement of the devices. No inference should be made regarding the architecture (i.e., configuration, design, etc.) of devices implementing functionality in accordance with the embodiments as shown in the drawings. Functionality described in this disclosure can be implemented in software or hardware, or a combination of both, as desired.
[0039] The flow diagrams depicted in the drawings show the functionality of example embodiments and do not imply any particular physical or architectural arrangement of the devices. No inference should be made regarding the architecture (i.e., configuration, design, etc.) of devices implementing functionality in accordance with the examples as shown in the drawings. Different arrangements can be used and still be within the scope of the examples.
[0040] It should be noted that "a plurality" refers to two or more.
[0041] With the research and progress of artificial intelligence technology, artificial intelligence technology is researched and applied in many fields, and plays an increasingly important value.
[0042] Artificial intelligence (AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Take artificial intelligence applied in machine learning as an example:
[0043] Machine Learning (ML) is a multi-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory, etc. It is a specialized study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure, and continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. It is applied in various fields of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, etc. The scheme of the present application mainly learns the feature data of different types of users by machine learning to obtain a classification model. The classification model is used to classify the feature data of a user to be classified to obtain the type corresponding to the user to be classified.
[0044] Before making specific descriptions, the terms involved in the present application are explained as follows:
[0045] The initial model refers to an algorithm model used to make decisions or assign items to categories. In the present application, at least two initial models can be at least two algorithm models selected from a random forest algorithm model, a machine learning model, a logistic regression model, a decision tree model, a support vector machine model, and a naive Bayes model.
[0046] The training sample refers to the feature data of different types of users authenticated by the user during the login and use of the APP. The training sample includes positive training samples and negative training samples. The positive training sample is the feature data of the first type of user, and the negative training sample is the feature data of the second type of user. The different types of users can be users of different ages (e.g., the different types of users include first age group users and second age group users), users of different physical health states (e.g., the different types of users include healthy users and sub-healthy users), users of different consumption abilities (e.g., the different types of users can include users with good consumption ability and users with poor consumption ability), or users of different working conditions (e.g., the different types of users include working users and unemployed users), etc. If the first type of user refers to the first age group user and the second type of user refers to the second age group user, the first age group can be the age group of minors, i.e., the age group less than 18 years old, and the second age group can be the age group of adults, i.e., the age group greater than or equal to 18 years old. The first age group can also be an age group less than a preset age, and the second age group can be an age group greater than or equal to a preset age. The preset age can be 12 years old, 14 years old, or 15 years old, etc.
[0047] The training set includes positive and negative training samples corresponding to at least one stage of the training samples. It should be understood that the types of feature data of the training samples corresponding to different stages are different. Therefore, if the training set includes positive and negative training samples corresponding to two stages, all feature types corresponding to the two stages can be used as the feature types of the training samples in the training set, and if a training sample does not have a certain feature type, the feature type is filled with a specified value (for example, zero). For example, if the types of feature data of the training samples of the first stage include A, B, C, and D, a total of 4 types, and the types of feature data of the training samples of the second stage include B, C, D, E, and F, a total of 5 types, if the training set includes training samples of the first stage and training samples of the second stage, the types of feature data of the training samples included in the training set are A, B, C, D, E, and F, a total of 6 types.
[0048] The validation set includes each validation sample obtained based on the user type identification operation of the user in the registration stage or the use stage, and the feature types of the training samples in the training set and the validation samples in the corresponding validation set are the same. The above-mentioned validation set can include the feature data of the registration stage or the use stage obtained when the trained model is deployed on the server, and the user type corresponding to the feature data obtained by identifying the user.
[0049] The training parameters include at least one of precision, recall, AUC value, and F1 score. The recall is the ratio of the number of users of the first type found (the number of training samples predicted as positive) to the number of actual users of the first type (the number of actual positive training samples). The higher the recall, the higher the probability that the actual first-class user is predicted. The precision, in terms of prediction results, means the probability that all training samples predicted as positive are actually positive training samples. AUC (Area Under Curve) is defined as the area surrounded by the coordinate axis under the ROC curve (receiver operating characteristic curve). Obviously, the value of this area cannot be greater than 1. Since the ROC curve is generally above the line y=x, the value of AUC ranges between 0.5 and 1. The closer the AUC is to 1.0, the better the model; when it is equal to 0.5, it is the worst and has no application value. F1 score: F1 score is used to represent the relationship between precision and recall. F1 score considers both precision and recall, and allows both to reach the highest, taking a balance, and the formula of F1 score is F1=2*precision*recall / (precision+recall).
[0050] The verification parameters include at least one of precision, recall, AUC value, and F1 score. The specific meanings of the verification parameters can refer to the foregoing specific description of the training parameters.
[0051] At present, for various application programs (APPs), a control method for preventing users (part of users) from being addicted to the APPs is started. Taking a game application program as an example, since all network game companies can only provide a period of service to minors in part of time periods, and cannot provide network game services to minors (here, the minors can refer to users under the age of 18) in any form in other time periods, and need to strictly implement the real-name registration and login requirements of network game user accounts, and cannot provide game services to users who have not registered and logged in, therefore, in the related art, when a user uses a game APP, it is usually necessary to register and perform real-name authentication to authenticate whether the user is a minor. For example, as shown in Figure 1 After the user logs in to the game APP, a forced real-name pop-up window is implemented for the game real-name user or the newly registered user, the user needs to output authentication information such as name and ID card number in the pop-up window for real-name authentication, and after the authentication information is input, a submission control (control A) can be clicked to send the authentication information to the game background for authentication. If the game background confirms that the user is an adult according to the ID card, further authentication is needed, for example, face pop-up window authentication, before the authentication is performed, as shown in Figure 2 After the user inputs the name and ID card number again, a control (control B) for starting face recognition is selected to perform face recognition to verify whether the face information and the identity information corresponding to the ID card number are consistent. If they are not consistent or the final authentication is a minor, the user is limited as a minor. For the user, if each user performs the foregoing login and face recognition verification operation, the user may be disturbed, and even the game retention may be affected.
[0052] Therefore, the application provides a classification model training method. The classification model obtained by using the method can accurately identify whether a user is a minor, thereby effectively alleviating the above problems. The method comprises the following steps: obtaining at least two initial models and a plurality of training sets obtained based on training samples, wherein the training samples comprise positive and negative training samples obtained in a registration stage and positive and negative training samples obtained in a use stage, each training set comprises positive and negative training samples corresponding to at least one stage, the positive training samples are feature data of a first type of user, and the negative training samples are feature data of a second type of user; for each initial model, the positive and negative training samples in each training set are used to train the initial model respectively, thereby obtaining a plurality of classification models corresponding to the initial model and training parameters of each classification model, and each classification model corresponds to a training set; determining a target initial model based on the training parameters of the classification models corresponding to each initial model; obtaining a verification set corresponding to each training set, wherein each verification sample in the verification set is obtained based on a user type identification operation of a user in the registration stage or the use stage, and the feature types of the training samples in the training set and the verification samples in the corresponding verification set are the same; for each classification model corresponding to the target initial model, the classification model is verified based on the verification set corresponding to the training set used to train the classification model, thereby obtaining verification parameters of the classification model; and determining a target classification model based on the verification parameters of the classification models. By using the above target classification model, the feature data of a to-be-identified user is classified to obtain the user type (such as the age range of the user), thereby confirming whether to restrict the use of the game APP by the user according to the identified user type. The above identification operation (such as a face recognition verification operation) does not need to be performed for each user, thereby improving the user experience.
[0053] The implementation details of the technical solutions of the embodiments of the application are described in detail as follows:
[0054] Figure 3 is a schematic diagram of an application scenario according to an embodiment of the application, as shown in Figure 3 The application scenario comprises a terminal device 10 and a server 20 in communication connection with the terminal device 10 through a network. The network can be a wide area network or a local area network, or a combination of the two. The terminal device 10 can be a smart phone, a tablet computer, a computer or the like. Figure 3 In the embodiment, only a schematic diagram of the terminal device 10 as a smart phone is shown.
[0055] A user can log in or use an APP through the terminal device 10, so that the terminal device 10 can obtain feature data in the user login or use APP stage, and identify the user to obtain the user category (for example, Figure 3 In the embodiment, a face recognition operation is adopted as Figure 2The interface shown identifies the user to obtain the age of the user, so as to associate the feature data with the user type to obtain the sample. For example, the terminal device 10 can upload the sample to the server 20 after obtaining the sample, or can be directly stored so as to be called by the server 20.
[0056] The server 20 can obtain at least two initial models, a plurality of training sets based on the training samples, the training samples including positive and negative training samples obtained in the registration stage and positive and negative training samples obtained in the use stage, each training set including positive and negative training samples corresponding to at least one stage, the positive training samples being feature data of the first type of users, and the negative training samples being feature data of the second type of users; for each initial model, the positive and negative training samples in each training set are used to train the initial model respectively, to obtain a plurality of classification models corresponding to the initial model and training parameters of each classification model, each classification model corresponding to a training set; a target initial model is determined based on the training parameters of the classification models corresponding to each initial model; a verification set corresponding to each training set is obtained, each verification sample in the verification set being obtained based on the user type identification operation of the user in the registration stage or the use stage, and the feature types of the training samples in the training set and the verification samples in the corresponding verification set being the same; for each classification model corresponding to the target initial model, the verification set corresponding to the training set used to train the classification model is used to verify the classification model to obtain verification parameters of the classification model; and a target classification model is determined based on the verification parameters of the classification models.
[0057] By using the above method, the training of the target classification model can be completed, so that the feature data of the user who logs in or uses the APP subsequently can be predicted by using the target classification model to obtain the user type.
[0058] It should be understood that the prediction process described above can be performed on the server 20, or can be performed on the terminal device 10. When the prediction process is performed on the server 20, the target classification model is deployed on the server 20; when the prediction process is performed on the terminal device 10, the target classification model is deployed on the terminal device 10. Similarly, the model training process described above can be performed on the server 20, or can be performed on the terminal device 10.
[0059] Figure 4 is a flowchart of a classification model training method according to an embodiment of the present application. The method can be performed by an electronic device with processing capability, for example, by a server, a terminal, or by a server and a terminal interacting to implement the scheme, and the like, which is not specifically limited here. Referring to Figure 4 As shown, the method includes at least steps S110 to S140, which are described in detail as follows:
[0060] Step S110, obtaining at least two initial models, and a plurality of training sets based on training samples.
[0061] The training samples include positive and negative training samples obtained in the registration stage and positive and negative training samples obtained in the use stage, each training set includes positive and negative training samples corresponding to at least one stage, the positive training samples are feature data of the first type of users, and the negative training samples are feature data of the second type of users.
[0062] The first type of users and the second type of users can be classified according to age, can be classified according to physical health, or can be classified according to consumption ability or work situation, and are not limited here.
[0063] In an implementation manner, the first age group users can be users with an age less than a preset age, and the second age group users can be users with an age greater than or equal to the preset age. The preset age can be 12 years old, 14 years old, 15 years old, or 18 years old, etc. In an implementation manner of the present application, the preset age is 18 years old, that is, the first age group users are underage users, and the second age group users are adult users.
[0064] The registration stage can be feature data obtained by the user on the day of registering to use the APP, and the use stage can be feature data obtained after using the APP for a period of time, such as feature data obtained after using the APP for three days, one week, two weeks, or one month, etc.
[0065] The feature data in the registration stage can include one or more of the following: the number of attempts to register within a specified time period (such as one week, two weeks, or one month, etc.) before the registration time, the number of attempts to register since the APP is put into use with a real-name strategy, the age filled in on the registration day, the attributes on the registration day (such as whether it is a holiday, a weekday, or a specified date, etc.), the earliest time and the latest time on the registration day, the active time length in different time periods on the registration day, the data of the devices used for registration on the registration day (such as the number of all accounts, the number of accounts of the first age group users, the number of natural persons, and the number of persons of the first age group, etc.), and the data of the registration devices used on the registration day within a period of time (such as within one week or one month, etc.) (such as the number of login accounts, the number of accounts of the first age group users, the number of accounts triggering the real-name authentication pop-up window, etc.).
[0066] The feature data in the use stage can include one or more of the above-mentioned feature data in the registration stage, and can also include one or more of the following: the number of active days in the use stage after registration, the proportion of active days on weekdays, the proportion of active days on holidays, the active time period on weekdays, the active time period on holidays, and the use time length in different active time periods, etc.
[0067] The at least two initial models obtained can be at least two of a random forest algorithm model, a machine learning model, a logistic regression model, a decision tree model, a support vector machine model, and a naive Bayes model, without specific limitation here, and can be set according to actual needs.
[0068] The at least two initial models can be obtained from a database or from the memory of the electronic device. That is, a plurality of algorithm models can be stored in the database or the electronic device.
[0069] The plurality of training sets obtained based on the training samples can be obtained in the following manner:
[0070] The training samples can be obtained by associating feature data of a plurality of terminal devices during user login or after the user logs in and uses an APP and a user type (such as a user age group) obtained by identifying the user.
[0071] In general, the proportions of different types of users using an APP are different. For example, for a game APP, a video playing APP, and a content interaction APP, the number of young users using such an APP is generally much smaller than the number of adult users. For an education APP or a classroom live streaming APP, the number of young users using such an APP is generally much smaller than the number of adult users. Therefore, the proportions of positive and negative samples in the training samples obtained above generally differ greatly depending on the type of the APP.
[0072] For example, for a game APP, a video playing APP, and a content interaction APP, the first age group is minors, and the second age group is adults. In this case, the ratio between positive training samples and negative training samples obtained from the training samples of users during registration and after registration is generally small. For an education APP or a classroom live streaming APP, the first age group is minors, and the second age group is adults. In this case, the ratio between positive training samples and negative training samples obtained from the training samples of users during registration and after registration is generally large.
[0073] To avoid the imbalance between positive and negative training samples leading to a larger amount of information about adults than about minors in the machine learning process, if all prediction results are biased towards adults, the model can achieve a high accuracy rate, which affects the final prediction effect. In the embodiments of the present application, to make the classification model trained using the positive and negative training samples in the subsequent training phase better, obtaining the plurality of training sets based on the training samples can specifically include the following steps:
[0074] Step S112: Obtain training samples, the training samples including positive training samples and negative training samples, the positive training samples including positive training samples obtained in the registration stage and positive training samples obtained in the use stage, and the negative training samples including negative training samples obtained in the registration stage and negative training samples obtained in the use stage.
[0075] The specific description of obtaining the training samples can refer to the foregoing specific description, which will not be repeated here.
[0076] Step S114: If the ratio of the positive training samples to the negative training samples in the training samples is less than a preset threshold, the positive training samples in the training samples are expanded to obtain expanded training samples, and a plurality of training sets are obtained based on the expanded training samples.
[0077] The preset threshold can be any value such as 0.5, 0.8, or 0.9, as long as the quantity difference between the positive training samples and the negative training samples is small.
[0078] The method of expanding the positive training samples in the training samples can be to expand the positive training samples in the training samples by using an oversampling algorithm, where the oversampling algorithm can be a random oversampling algorithm, an SMOTE algorithm, or a DOPING algorithm. The positive samples in the training samples can also be expanded by using a data enhancement algorithm. It should be understood that there can be many ways to expand the positive samples in the training samples, and the above are only exemplary and should not be considered as a limitation of the present solution.
[0079] In an implementation manner of the present application, step S114 can be to expand the positive training samples in the training samples based on the SMOTE algorithm to obtain the expanded training samples.
[0080] If the ratio of the two types of samples of minors and adults in the training set is n:m*n (m is a positive integer), the SMOTE algorithm will expand the minor type to generate a*n samples, where 0
[0081] Please refer to Figure 5 , Figure 6 and Figure 7 for an example. For example, the positive training samples include feature data corresponding to minors in the registration stage and the training stage, and the negative training samples include feature data corresponding to adults in the registration stage and the training stage. Figure 5The obtained training set is represented by triangles, where positive training samples are represented by triangles and negative training samples by circles. As can be seen from the diagram, the number of negative training samples is much larger than the number of positive training samples; therefore, it is necessary to expand the positive training samples. When expanding the positive training samples using the SMOTE algorithm, the selected positive training samples (e.g., ...) can be calculated using the Euclidean distance formula. Figure 6 The distance between the positive training samples (selected by the solid line in the middle) and other positive training samples is used to select K nearest neighbor samples (e.g., ...). Figure 6 The positive training samples are shown in the dashed lines. Then, a ratio i (where i is a positive integer) is set, and new sample points are randomly selected from the selected positive training samples and their k nearest neighbors. Finally, the number of training samples for the minor class is increased to i*n, narrowing the gap between the number of minors and the number of adults. For example, as shown... Figure 7 As shown, a new sample can be calculated between the selected positive training sample and any target nearest neighbor sample using the sample expansion calculation formula. If the above expansion calculation formula is used to calculate N times, the number of positive training samples can be expanded from the original M to M+N.
[0082] It should be understood that if the preset ratio is greater than a specified threshold, the negative training samples in the training samples can be expanded to obtain expanded training samples, and multiple training sets can be obtained based on the expanded training samples. The specified threshold mentioned above can be 1.5, 2, 3, or 5, etc., and can be set according to actual needs.
[0083] Regarding the methods for expanding negative training samples in the training samples, the process of expanding positive training samples in the training samples was described above, and will not be repeated here.
[0084] Step S120: For each initial model, train the initial model using the positive and negative training samples in each training set to obtain multiple classification models corresponding to the initial model and the training parameters of each classification model. Each classification model corresponds to a training set.
[0085] Specifically, when training the initial model using positive and negative training samples from each training set, the training samples from each training set can be input into the initial model for training. During the training process, the model parameters are continuously adjusted until the model converges, resulting in multiple classification models corresponding to the initial model, with each classification model corresponding to a training set.
[0086] The training parameter of each classification model can be obtained after obtaining the classification model, and a part of positive and negative training samples can be selected from the corresponding training set to input into the classification model to obtain a prediction result (predicted age range), and the training parameter of the classification model can be calculated according to the user type (for example, user age range) and the prediction type (for example, predicted age range) corresponding to each positive and negative training sample.
[0087] The training parameter of the classification model can be a parameter for evaluating the pros and cons of the classification effect of the classification model. Specifically, the training parameter of the classification model can include at least one of precision, recall, AUC value, and F1 score.
[0088] In step S130, a target initial model is determined based on the training parameter of the classification model corresponding to each initial model.
[0089] If the training parameter of each classification model is multiple, the above-mentioned method of determining the target initial model can be multiple.
[0090] In an implementation manner, one training parameter can be selected from the training parameter of each classification model. Since the training parameter is used to represent the pros and cons of the classification effect of the classification model, the classification model can be determined from the multiple classification models based on the value size of the selected training parameter of each classification model, and the initial model corresponding to the determined classification model is taken as the target initial model.
[0091] In another implementation manner, the training parameters of each classification model can be weighted and summed to obtain the training parameter calculation value of each classification model, and the classification model can be determined based on the training parameter calculation value of each classification model, and the initial model corresponding to the determined classification model is taken as the target initial model. The greater the value of the same training parameter, the greater the corresponding weight value.
[0092] In step S140, a verification set corresponding to each training set is obtained. Each verification sample in the verification set is obtained based on the user type identification operation of the user in the registration stage or the use stage, and the feature types of the training samples in the training set and the verification samples in the corresponding verification set are the same.
[0093] The method of obtaining the verification set corresponding to each training set can be: obtaining the verification sample obtained by associating the feature data collected in the registration stage with the identification result obtained by classifying and identifying (face recognition) the feature data; and obtaining the verification sample obtained by associating the feature data collected in the use stage after the user registration with the identification result obtained by classifying and identifying (face recognition) the feature data.
[0094] It should be understood that the proportion of positive and negative verification samples in each of the above obtained verification sets is determined according to the users actually using the APP.
[0095] The manner of obtaining the verification set corresponding to each training set can also be that, for each classification model corresponding to the target initial model, deploying the classification model on a server, and using the classification model to perform classification prediction on a plurality of to-be-predicted feature data to obtain a probability that each to-be-predicted feature data classification prediction result is a first age group user, sorting the probabilities of the to-be-predicted feature data from large to small, obtaining the user types of each target to-be-predicted feature data and each target registered user corresponding to the target to-be-predicted feature data that are ranked before the delivery threshold value of the classification model corresponding to the target to-be-predicted feature data, as the verification samples in the verification set corresponding to the training set for training the classification model, and the user types of each target registered user are obtained based on performing a user type identification operation (such as a face recognition operation) on the users in the registration stage or the use stage.
[0096] Specifically, since the above verification set is usually obtained in the APP registration stage and the APP use stage after registration, when obtaining the user types (such as user age groups) of each target to-be-predicted feature data and each target registered user that are ranked before the delivery threshold value of the classification model corresponding to the target to-be-predicted feature data, specifically, a face recognition verification window for verifying the user age can be displayed in the form of a pop-up window for the users corresponding to the target to-be-predicted feature data, to verify whether the user age group corresponds to the first age group or the second age group.
[0097] The delivery threshold value can be determined according to the proportion of positive and negative samples in the training samples. It can also be determined according to the proportion of positive and negative samples in the training samples and a preset confidence. It can also be determined according to the utility rate of the APP, the proportion of positive and negative samples in the training samples, and a preset confidence, which can be set according to actual needs.
[0098] In an implementation manner of the present application, the delivery threshold value can be calculated by the following formula to obtain the delivery threshold value of each classification model corresponding to the target initial model, wherein X is the delivery threshold value, m is an absolute difference value, P * is the proportion of the number of the first type of users in the training samples (such as the proportion of the number of the first age group users), and Z * is a statistical quantity, and is a value obtained when the confidence is 95% and the normal distribution is followed.
[0099] Step S150, for each classification model corresponding to the target initial model, verifying the classification model based on the verification set corresponding to the training set for training the classification model to obtain the verification parameter of the classification model.
[0100] The manner of verifying the classification model based on the verification set corresponding to the training set used to train the classification model to obtain the verification parameter of the classification model can be: inputting each verification sample in the verification set into the corresponding classification model to identify the verification sample in the corresponding verification set by using the classification model to obtain the user type prediction result of each verification sample, and obtaining the model parameter of the classification model based on the user type label corresponding to each verification sample and the user type prediction result of each verification sample.
[0101] The verification parameter of the classification model can be at least one of precision, recall, AUC value, and F1 score.
[0102] Step S160: determining the target classification model based on the verification parameter of each classification model.
[0103] If the verification parameter of each classification model is multiple, the manner of determining the target classification model can be multiple. The verification parameter can be at least one of precision, recall, AUC value, and F1 score.
[0104] In one implementation manner, one verification parameter can be selected from the verification parameter of each classification model. Since the verification parameter is used to represent the advantages and disadvantages of the classification effect of the classification model, the target classification model can be determined from the multiple classification models based on the value of the selected verification parameter of each classification model. For example, when determining the target classification model, the classification model with the maximum value of the verification parameter of each classification model can be determined as the target classification model.
[0105] In another implementation manner, the verification parameter of each classification model can be weighted and summed to obtain the verification parameter calculation value of each classification model, and the target classification model can be determined based on the verification parameter calculation value of each classification model. The greater the value of the same verification parameter, the greater the corresponding weight value. For example, when determining the target classification model, the classification model with the maximum calculation value of the verification parameter calculation value of each classification model can be determined as the target classification model.
[0106] The application provides a classification model training method. The method comprises the following steps: obtaining at least two initial models and a plurality of training sets obtained based on training samples; in a model training process, for each initial model, training the initial model by using positive and negative training samples in each training set to obtain a plurality of classification models corresponding to the initial model and training parameters of each classification model; determining a target initial model based on the training parameters of the classification models corresponding to each initial model, so as to effectively ensure that the type of the selected initial model is the optimal type for user classification. In addition, since each verification sample in the verification set is obtained based on the user type identification operation of the user in the registration stage or the use stage, the sample proportion of the obtained positive and negative verification samples usually has a large difference. Therefore, in the verification stage, for each classification model corresponding to the target initial model, the classification model is verified based on the verification set corresponding to the training set used for training the classification model to obtain verification parameters of the classification model; the target classification model is determined based on the verification parameters of the classification models, so as to effectively ensure that the model also has a good classification effect in the case that the positive and negative verification samples have a large difference. Therefore, the accuracy of the user age obtained by using the target classification model to identify the feature data in the subsequent use can be improved.
[0107] Please refer to Figure 8 In another embodiment of the application, a classification model training method is provided. The method comprises the following steps:
[0108] In step S210, at least two initial models and a plurality of training sets obtained based on training samples are obtained.
[0109] The training samples comprise positive and negative training samples obtained in the registration stage and positive and negative training samples obtained in the use stage. Each training set comprises positive and negative training samples corresponding to at least one stage. The positive training samples are feature data of the first type of users, and the negative training samples are feature data of the second type of users.
[0110] For specific description of step S210, please refer to the specific description of step S110 in the foregoing embodiment, which will not be repeated here.
[0111] In step S220, for each initial model, positive and negative training samples are selected from each training set according to a target preset ratio to train the initial model, so as to obtain a plurality of first classification models corresponding to a plurality of preset classification thresholds and training parameters of each first classification model.
[0112] Each preset classification threshold corresponds to a first classification model, and the preset ratio is used to represent the ratio of the positive training samples to the negative training samples in the training set.
[0113] The preset ratio can be preset or determined from multiple ratios, and can be set according to actual needs.
[0114] If the preset ratio is preset, the preset ratio can be 5:1, 2:1, 3:2, 1:1, 1:2, 2:3, or 1:5, etc.
[0115] If the preset ratio is determined from multiple ratios, the specific determination method can be: multiple groups of training samples with different ratios are selected from a training set and input into at least one initial model respectively, to obtain a second classification model obtained by each initial model after training with each group of training samples. For each second classification model, a prediction result obtained by using the second classification model to predict a plurality of specified feature data with user type labels (such as age labels) is obtained, and a prediction parameter of the second classification model is obtained based on the prediction result of each specified feature data and the label of each feature data. The prediction parameter can be at least one of accuracy, recall rate, or AUC value, etc. A target ratio is determined according to the parameters of each second classification model. The specified feature data can be selected from the training set.
[0116] That is, before step S220 is performed, the method further includes:
[0117] For each initial model, multiple groups of positive and negative training samples are selected from a training set according to multiple preset ratios to train the initial model respectively, to obtain a second classification model corresponding to each group of positive and negative training samples when the classification threshold is set to a value, and a training parameter of the second classification model, wherein each preset ratio corresponds to a group of positive and negative training samples; and a target preset ratio is determined according to the training parameters of the second classification models corresponding to the multiple preset ratios.
[0118] The multiple preset ratios can be preset, and the multiple preset ratios can specifically include at least two of 10:1, 5:1, 2:1, 1:1, 1:2, 1:5, 1:10, etc., which can be set according to actual needs, and the embodiment is not limited specifically.
[0119] The training set used to determine the target preset ratio can include positive and negative training samples obtained in the registration stage, can include positive and negative training samples obtained in the use stage, or can include positive and negative training samples obtained in the registration stage and the use stage, which can be set according to actual needs.
[0120] In an implementable manner of the present application, the training set used to determine the target preset ratio is a training set composed of positive and negative training samples obtained in the registration stage.
[0121] If the verification parameter of each second classification model is multiple, the target preset ratio can be determined according to the model parameter of the second classification model corresponding to each preset ratio.
[0122] In an implementation manner, one training parameter can be selected from the training parameters of each second classification model. Since the training parameter is used to represent the pros and cons of the classification effect of the classification model, the target preset ratio can be determined from the multiple classification models based on the value of the selected training parameter of each classification model. For example, when determining the target preset ratio, the preset ratio corresponding to the second classification model with the maximum value of the training parameter can be determined as the target preset ratio.
[0123] In another implementation manner, the training parameters of the second classification model corresponding to each preset ratio can be weighted and summed to obtain the training parameter calculation value of the second classification model corresponding to each preset ratio, and a target second classification model can be determined based on the training parameter calculation value of the second classification model corresponding to each preset ratio, and the preset ratio corresponding to the target second classification model is taken as the target preset ratio. For example, when determining the target preset ratio, the preset ratio corresponding to the second classification model with the maximum calculation value of the training parameter calculation value can be determined as the target preset ratio.
[0124] For the process of selecting positive and negative samples from each training set according to the target preset ratio to train each initial model, refer to the specific description of step S120 in the foregoing, which will not be repeated here.
[0125] It should be noted that the classification threshold refers to a classification critical point. When the classification calculation value obtained by classifying a certain feature data is greater than or equal to the classification threshold, it can be determined that the user type corresponding to the feature data is the first type of user (i.e., the age range is the first age range). Correspondingly, if it is less than the classification threshold, it can be determined that the user type corresponding to the feature data is the second type of user (i.e., the age range is the second age range). During the training process of each initial model, different classification thresholds are selected, and the model accuracy of the classification model corresponding to each classification threshold is also different, and the parameters of the classification model corresponding to each classification threshold are also different.
[0126] The multiple classification thresholds can be preselected or set, and can include at least two of 0.5, 0.6, 0.65, 0.7, 0.75, 0.8, and 0.9. The actual demand can be set, and the embodiments of the present application are not limited.
[0127] In step S230, a target preset classification threshold and a target initial model are determined according to the training parameters of the first classification model corresponding to each initial model at the plurality of preset classification thresholds.
[0128] The first classification model corresponding to the target preset classification threshold and trained based on the target initial model is the classification model corresponding to the target initial model.
[0129] The training parameters of the first classification model are parameters for evaluating the advantages and disadvantages of the classification effect of the classification model, and can specifically include at least one of precision, recall, AUC value, and F1 score.
[0130] If the training parameters of each first classification model are multiple, the above-mentioned manner of determining the target initial model can be multiple.
[0131] In an implementation manner, for each first classification model, a training parameter is selected from the training parameters of the first classification model. Since the training parameter is used to represent the advantages and disadvantages of the classification effect of the classification model, a target classification threshold and a target first classification model can be determined based on the value of the selected parameter of the training parameters of the first classification model corresponding to the plurality of classification thresholds, and the initial model corresponding to the target first classification model is taken as the target initial model. For example, the preset ratio corresponding to the maximum training parameter can be selected as the target preset ratio.
[0132] In another implementation manner, the training parameters of each first classification model are weighted and summed to obtain a training parameter calculation value of each first classification model, and a target classification model and a target classification threshold are determined based on the training parameter calculation value of each first classification model, and the initial model corresponding to the target classification model is taken as the target initial model. For example, the preset ratio corresponding to the maximum training parameter calculation value can be selected as the target preset ratio.
[0133] In yet another implementation manner, for each initial model, the plurality of training parameters of the first classification model corresponding to the plurality of classification thresholds obtained based on the initial model are weighted and summed to obtain a training parameter calculation value of each first classification model, and a candidate classification model and a candidate classification threshold are determined based on the training parameter calculation value of each first classification model. The target classification model and the target classification threshold are determined according to the training parameter calculation value of the candidate classification model corresponding to each initial model, and the initial model corresponding to the target classification model is taken as the target initial model.
[0134] In still another implementation, for each preset classification threshold, the training parameters of the first classification model corresponding to each initial model at the preset classification threshold can be weighted and summed to obtain a training parameter calculation value of each initial model at the preset classification threshold, and a target parameter calculation value can be determined based on the training parameter value of each initial model, wherein the greater the value of the same training parameter, the greater the corresponding weight value; the preset classification threshold corresponding to the maximum target parameter calculation value is determined as the target preset classification threshold; and the initial model corresponding to the maximum training parameter calculation value in the training parameter calculation value of each initial model corresponding to the target classification threshold is determined as the target initial model.
[0135] In this way, for each preset classification threshold, the target parameter calculation value can be determined based on the training parameter value of each initial model in the following way: for each preset classification threshold, the maximum value in the training parameter values of the multiple initial models corresponding to the preset classification threshold is taken as the target training parameter value; or for each preset classification threshold, the mean value of the training parameter values of the multiple initial models corresponding to the preset classification threshold is taken as the target training parameter value.
[0136] It should be understood that the above-described way of determining the target initial model and the target classification threshold can also have multiple ways, and the above examples are only illustrative and should not be considered as limiting.
[0137] In step S240, a verification set corresponding to each training set is obtained, each verification sample in the verification set is obtained based on the user type identification operation of the user in the registration stage or the use stage, and the feature categories of the training samples in the training set and the verification samples in the corresponding verification set are the same.
[0138] In step S250, for each classification model corresponding to the target initial model, the classification model is verified based on the verification set corresponding to the training set used to train the classification model to obtain the verification parameter of the classification model.
[0139] In step S260, the target classification model is determined based on the verification parameters of the classification models.
[0140] For the specific description of steps S240-S160, refer to the specific description of steps S140-S160 in the foregoing embodiments, which will not be repeated here.
[0141] The application provides a classification model training method. The method comprises the following steps: obtaining at least two initial models and a plurality of training sets obtained based on training samples; in a model training process, for each initial model, selecting positive and negative training samples from each training set according to a target preset ratio to train the initial model, obtaining a plurality of first classification models corresponding to a plurality of preset classification thresholds respectively and training parameters of each first classification model, wherein the target preset ratio is determined based on a plurality of preset ratios, and a target preset classification threshold and a target initial model are determined according to the training parameters of the first classification model corresponding to each initial model at a plurality of preset classification thresholds. Therefore, the proportion of the selected positive and negative samples can be effectively ensured to be optimal, the classification threshold is optimal, and the type of the selected initial model is the optimal type for user classification. In addition, since each verification sample in the verification set is obtained based on the user type identification operation of the user in the registration stage or the use stage, the sample proportion of the obtained positive and negative verification samples usually has a large difference. Therefore, in the verification stage, for each classification model corresponding to the target initial model, the classification model is verified based on the verification set corresponding to the training set used for training the classification model to obtain the verification parameter of the classification model, and the target classification model is determined based on the verification parameters of the classification models, so that the model also has good classification effect under the condition that the positive and negative verification samples have a large difference. Therefore, the accuracy of the subsequent identification of the feature data by using the target classification model to determine the user type is ensured.
[0142] Please refer to Figure 9 The application also provides a user classification method, which can comprise the following steps:
[0143] In step S310, the feature data of a user to be classified is obtained.
[0144] The feature data of the user to be classified can be obtained in the use stage of the APP or in the registration stage, which is not limited here.
[0145] In step S320, the target classification model is used to process the feature data to obtain the user type corresponding to the feature data of the user to be classified, and the user type is a first type of user or a second type of user.
[0146] It should be noted that the feature types included in the feature data of the user to be classified should be the same as the feature types included in the training samples used for training the target classification model. Therefore, the specific description of the feature data can refer to the specific description of the feature data in the training samples described above, which is not repeated here.
[0147] The step S320 can be specifically that the feature data is classified and calculated by using the target classification model to obtain a classification calculation result, and the classification calculation result is compared with a target classification threshold in the target classification model to determine the category corresponding to the feature data of the user to be classified. When the classification calculation result is greater than or equal to the target classification threshold, it can be determined that the user type corresponding to the feature data of the user to be classified is the first user type. When the classification calculation result is less than the target classification threshold, it can be determined that the user type corresponding to the feature data of the user to be classified is the second user type.
[0148] For example, if the positive training sample in the training sample of the target classification model is the feature data of the user in the first age group, and the negative training sample is the feature data of the user in the second age group, then when the classification calculation result is greater than or equal to the target classification threshold, it can be determined that the age group corresponding to the feature data of the user to be classified is the first age group. When the classification calculation result is less than the target classification threshold, it can be determined that the user type corresponding to the feature data of the user to be classified is the second age group.
[0149] The obtaining process of the target classification model can refer to the specific description of the classification model training method in the foregoing, and will not be described here.
[0150] By using the classification method of the present application, accurate classification of the user to be classified can be realized, so as to facilitate determining the restriction mode of using the APP by the user to be classified according to the age classification result of the user to be classified.
[0151] Please refer to Figure 10 The training and use scene of the classification model is a game scene, and therefore the present application proposes a classification model training method for identifying whether the user used in the game registration and use process is an adult user or a minor user. The method comprises the following steps:
[0152] In step S410, at least two initial models and training samples obtained based on a pop-up window authentication mode are obtained.
[0153] The way of obtaining the training samples based on the pop-up window authentication mode can refer to the description of the method for obtaining the training samples based on the pop-up window authentication mode in the foregoing. Figure 1
[0154] The at least two models obtained in the method specifically include a random forest algorithm model (RF model) and a machine learning model (GBDT model). Each training sample is a sample obtained by using the method shown in Figure 3 after performing the face recognition operation.
[0155] Step S420, if the ratio of positive training samples to negative training samples in the training samples is less than a preset threshold, the positive training samples in the training samples are expanded based on an oversampling algorithm, to obtain expanded training samples, and a plurality of training sets are obtained based on the expanded training samples.
[0156] The plurality of training sets obtained based on the training samples are three, specifically a first training set, a second training set, and a third training set. The first training set includes feature data of users on the registration day (registration stage), the second training set includes feature data of users in a period after registration (use stage), and the third training set includes feature data of users on the registration day and feature data of users in a period after registration. Each training set includes positive and negative training samples, and the positive training samples are feature data of minor users, and the negative training samples are feature data of adult users. The categories of the feature data of different training sets can be specifically referred to the foregoing specific description of step S110.
[0157] Step S430: For each initial model, a plurality of groups of positive and negative training samples are selected from the first training set according to a plurality of preset ratios, and the initial model is trained respectively to obtain a second classification model corresponding to each group of positive and negative training samples when the classification threshold is set, and a training parameter of the second classification model.
[0158] Each preset ratio corresponds to a group of positive and negative training samples, and the parameters of the second classification model include precision, recall, AUC value, and F1 score. As shown in Table 1, for each initial model (RF model or GBDT model), a plurality of training parameters of the second classification model are obtained by training with positive and negative sample ratios of 1:1, 1:5, and 1:10, respectively. Table 1 is as follows:
[0159]
[0160]
[0161] Step S440: The training parameters of the second classification model corresponding to each preset ratio are weighted and summed to obtain a training parameter calculation value corresponding to each preset ratio, and the preset ratio corresponding to the maximum training parameter calculation value is selected as a target preset ratio.
[0162] The greater the value of the same training parameter, the greater the corresponding weight value. The target preset ratio obtained by using step S440 with respect to Table 1 is 1:1.
[0163] Step S450, for each initial model, selecting positive and negative training samples from each training set according to the target preset ratio to train the initial model, obtaining a plurality of first classification models corresponding to a plurality of preset classification thresholds and training parameters of each first classification model.
[0164] Wherein, one preset classification threshold corresponds to one first classification model. The training parameters of the first classification model can include precision, recall, AUC value and F1 score, as shown in Table 2, Table 3 and Table 4, which are the training parameters of a plurality of first classification models obtained by training each initial model (RF model or GBDT model) with classification thresholds of 0.5, 0.6, 0.7, 0.8 and 0.9 respectively.
[0165] Table 2 is the training parameters of each first classification model when the classification threshold is 0.5, 0.6, 0.7, 0.8 and 0.9, which is obtained by selecting a plurality of training samples from the first training set according to the target preset ratio of 1:1 for each model. Table 2 is as follows:
[0166]
[0167] Table 3 is the training parameters of each first classification model when the classification threshold is 0.5, 0.6, 0.7, 0.8 and 0.9, which is obtained by selecting a plurality of training samples from the second training set according to the target preset ratio of 1:1 for each model. Table 3 is as follows:
[0168]
[0169]
[0170] Table 4 is the training parameters of each first classification model when the classification threshold is 0.5, 0.6, 0.7, 0.8 and 0.9, which is obtained by selecting a plurality of training samples from the third training set according to the target preset ratio of 1:1 for each model. Table 4 is as follows:
[0171]
[0172] Step S460, according to the training parameters of the first classification model corresponding to each initial model under a plurality of preset classification thresholds, determining a target preset classification threshold and a target initial model.
[0173] Wherein, the first classification model corresponding to the target preset classification threshold trained based on the target initial model is the classification model corresponding to the target initial model.
[0174] Specifically, for each preset classification threshold, the training parameters of the first classification model corresponding to each initial model under the preset classification threshold are weighted and summed to obtain the training parameter calculation value of each initial model under the preset classification threshold, and the target parameter calculation value is determined based on the training parameter value of each initial model, wherein the greater the value of the same training parameter, the greater the corresponding weight value; the preset classification threshold corresponding to the maximum target parameter calculation value is determined as the target preset classification threshold; and the initial model corresponding to the maximum training parameter calculation value in the training parameter calculation value of each initial model corresponding to the target classification threshold is determined as the target initial model. That is, as shown in Tables 2-4, the target preset classification threshold is 0.7, the classification effect is best, and under the same classification threshold, the classification effect of the GBDT model is obviously better than that of the RF model. That is, the target initial model determined is the GBDT model.
[0175] In step S470, the delivery threshold of each classification model corresponding to the target initial model is determined according to the proportion of positive and negative training samples in the training samples.
[0176] Specifically, the delivery threshold of each classification model corresponding to the target initial model can be calculated by using the calculation formula , wherein X is the delivery threshold, m is an absolute difference value, P is obtained according to the classification model for a specified number of feature data with age stage labels, and Z is the proportion of the number of users in the first age stage in the training samples. * is a statistical value obtained when the confidence level is 95% and the normal distribution is followed. *
[0177] In step S480, the verification set corresponding to each classification model is obtained based on the delivery threshold of each classification model.
[0178] Specifically, for each classification model corresponding to the target initial model, the classification model is deployed on the server, and the classification model is used to classify and predict a plurality of to-be-predicted feature data to obtain the probability that each to-be-predicted feature data classification prediction result is a user in the first age stage. The probabilities of the to-be-predicted feature data are sorted from large to small, and the age stages of the users corresponding to each target to-be-predicted feature data and each target registered user are obtained as the verification samples in the verification set corresponding to the training set of the classification model, and the age stages of the users corresponding to each target registered user are obtained based on the face recognition operation performed on the users in the registration stage or the use stage.
[0179] Step S3490: For each classification model corresponding to the target initial model, verifying the classification model based on the verification set corresponding to the training set used to train the classification model to obtain the verification parameter of the classification model, and determining the target classification model based on the verification parameters of each classification model.
[0180] After the training of the target classification model is completed, the target classification model described above can be deployed on a server to use the target classification model to predict whether a newly online or newly registered user is an adult user. Thus, the inconvenience caused by the need to perform face authentication for each user when logging into a game APP is avoided, greatly improving the efficiency and experience of user age classification.
[0181] As shown in Table 5, the results are obtained by analyzing the login situation of underage users in the game server in a time period (such as the period from December 23 to December 26 in XX) when each classification model is deployed on the game server. Specifically, Table 5 shows the underage authentication proportion confirmed based on the first training set and the second training set when using each classification model corresponding to the target initial model for online prediction, and the proportion of newly registered users who pass the authentication and the proportion of minors determined when authenticating the proportion of minors. The proportion of newly registered users who pass the authentication and the proportion of minors determined when using the classification model corresponding to the first training set and the classification model corresponding to the second training set to authenticate the proportion of minors. The proportion of newly registered users who pass the authentication and the proportion of minors determined when using the classification model corresponding to the third training set to authenticate the proportion of minors. The proportion of newly registered users who pass the authentication and the proportion of minors determined when the game server performs game login authentication using the pop-up window to confirm the proportion of minors. The proportion of newly registered users who pass the authentication and the proportion of minors determined when authenticating the proportion of minors. And the high-risk rules set in advance, wherein the high-risk rules set the proportion of minors authentication, and the proportion of minors determined when authenticating the proportion of minors. Table 5 is as follows:
[0182]
[0183] From Table 5, it can be seen that the proportion of minor authentication confirmed based on the first training set and the second training set is 10.09%, which is relatively increased by 601% compared with the pop-up authentication result in the same period, and is relatively increased by 345% compared with the high-risk rule in the same period; the proportion of being determined as a minor is 72%, which is relatively increased by 6% compared with the pop-up authentication result in the same period, and is relatively decreased by 4% compared with the high-risk rule in the same period. The proportion of minor authentication confirmed based on the third training set is 10.76%, which is relatively increased by 647% compared with the pop-up authentication result in the same period, and is relatively increased by 374% compared with the high-risk rule in the same period; the proportion of being determined as a minor is 77%, which is relatively increased by 13% compared with the pop-up authentication result in the same period, and is relatively increased by 2% compared with the high-risk rule in the same period. It can be seen that the classification model trained by the third training set has the best effect. That is, the target classification model described above is the classification model obtained based on the third training set.
[0184] After obtaining the target classification model described above, the target classification model can be deployed on a game server to predict whether a user registered online and using the game is a minor user. The user predicted as a minor can also be subjected to corresponding game restrictions, such as allowing the user to perform game operations only for a certain duration in a certain time period.
[0185] The device embodiments of the present application are described below, which can be used to execute the methods in the embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments described above.
[0186] Please refer to Figure 11 The embodiments of the present application also provide a classification model training device 500 applicable to an electronic device, which comprises a first acquisition module 510, a training module 520, an initial model determination module 530, a second acquisition module 540, a verification module 550, and a classification model determination module 560.
[0187] The first obtaining module 510 is configured to obtain at least two initial models and a plurality of training sets obtained based on training samples, wherein the training samples include positive and negative training samples obtained in a registration stage and positive and negative training samples obtained in a use stage, each training set includes positive and negative training samples corresponding to at least one stage, the positive training samples are feature data of a first type of user, and the negative training samples are feature data of a second type of user; the training module 520 is configured to, for each initial model, train the initial model by using the positive and negative training samples in each training set, to obtain a plurality of classification models corresponding to the initial model and training parameters of each classification model, and each classification model corresponds to one training set; the initial model determination module 530 is configured to determine a target initial model based on the training parameters of the classification models corresponding to each initial model; the second obtaining module 540 is configured to obtain a verification set corresponding to each training set, wherein each verification sample in the verification set is obtained based on a user type identification operation of a user in the registration stage or the use stage, and the feature types of the training samples in the training set and the verification samples in the corresponding verification set are the same; the verification module 550 is configured to, for each classification model corresponding to the target initial model, verify the classification model based on the verification set corresponding to the training set used to train the classification model, to obtain verification parameters of the classification model; and the classification model determination module 560 is configured to determine a target classification model based on the verification parameters of the classification models.
[0188] In an implementation manner, the first obtaining module 510 includes a first obtaining submodule and an expansion submodule. The first obtaining submodule is configured to obtain training samples, wherein the training samples include positive training samples and negative training samples, the positive training samples include positive and negative training samples obtained in a registration stage and positive training samples obtained in a use stage, and the negative training samples include negative training samples obtained in the registration stage and negative training samples obtained in the use stage. The expansion submodule is configured to, when a ratio of the positive training samples to the negative training samples in the training samples is less than a preset threshold, expand the positive training samples in the training samples to obtain expanded training samples, and obtain a plurality of training sets based on the expanded training samples.
[0189] In this implementation manner, the expansion submodule is further configured to expand the positive training samples in the training samples based on an oversampling algorithm to obtain the expanded training samples.
[0190] In an implementation, the training module 520 is further configured to train each initial model by selecting a pair of positive and negative training samples from each training set according to a preset ratio of target, to obtain a first classification model corresponding to each preset classification threshold and training parameters of each first classification model, wherein one preset classification threshold corresponds to one first classification model, and the preset ratio of target is used to represent a ratio of positive training samples to negative training samples in the training set. The initial model determination module is further configured to determine a target preset classification threshold and a target initial model according to the training parameters of the first classification model corresponding to each preset classification threshold of each initial model, wherein the first classification model corresponding to the target preset classification threshold and trained based on the target initial model is the classification model corresponding to the target initial model.
[0191] In this way, the classification model training apparatus 500 further includes a preset ratio determination module, and the training module 520 is further configured to train each initial model by selecting a plurality of pairs of positive and negative training samples from one training set according to a plurality of preset ratios, to obtain a second classification model when a preset classification threshold corresponding to each pair of positive and negative training samples is set to a preset value and training parameters of the second classification model, wherein each preset ratio corresponds to one pair of positive and negative training samples. The preset ratio determination module is configured to determine a target preset ratio according to the training parameters of the second classification model corresponding to each preset ratio.
[0192] In an implementation, the training parameters of the second classification model are at least two, and the preset ratio determination module is further configured to perform weighted summation on the training parameters of the second classification model corresponding to each preset ratio to obtain a training parameter calculation value corresponding to each preset ratio, wherein the greater the value of the same training parameter, the greater the corresponding weight value; and the preset ratio corresponding to the maximum training parameter calculation value is selected as the target preset ratio.
[0193] In an implementation, the initial model determination module 530 includes a calculation sub-module, a classification threshold determination sub-module, and an initial model determination sub-module. The calculation sub-module is configured to, for each preset classification threshold, perform weighted summation on the training parameters of the first classification model corresponding to each initial model at the preset classification threshold to obtain a training parameter calculation value of each initial model at the preset classification threshold, and determine a target parameter calculation value based on the training parameter value of each initial model, wherein the greater the value of the same training parameter, the greater the corresponding weight value. The classification threshold determination sub-module is configured to determine that the preset classification threshold corresponding to the maximum target parameter calculation value is the target preset classification threshold. The initial model determination sub-module is configured to determine that the initial model corresponding to the maximum training parameter calculation value in the training parameter calculation value of each initial model corresponding to the target classification threshold is the target initial model.
[0194] In an implementation, the first type of user includes a first age group, the second type of user includes a second age group, and the second obtaining module 540 includes a threshold determining submodule and a second obtaining submodule. The threshold determining submodule is configured to determine a delivery threshold of each classification model corresponding to the target initial model according to a proportion of positive training samples to negative training samples in the training samples. The second obtaining submodule is configured to, for each classification model corresponding to the target initial model, deploy the classification model on a server, and use the classification model to perform classification prediction on the plurality of to-be-predicted feature data to obtain a probability that each to-be-predicted feature data classification prediction result is a first type of user, sort the probabilities of the to-be-predicted feature data from large to small, and obtain each target to-be-predicted feature data and a user type corresponding to each target registered user that are located before the delivery threshold corresponding to the classification model as verification samples in a verification set corresponding to a training set used for training the classification model. The user type corresponding to each target registered user is obtained based on performing a user type identification operation on a user in a registration stage or a use stage.
[0195] In an implementation, the verification parameters of the first classification model are at least two, and the classification model determining module 560 includes a calculation submodule and a classification model determining submodule. The calculation submodule is configured to perform weighted summation on the verification parameters of each classification model to obtain a verification parameter calculation value corresponding to each classification model, where the greater the value of the same verification parameter, the greater the corresponding weight value. The classification model determining submodule is configured to select a classification model corresponding to the greatest verification parameter calculation value as a target classification model.
[0196] In an implementation, the training parameters include at least one of precision, recall, AUC value, and F1 score, and the model parameters of the verification parameters include at least one of precision, recall, AUC value, and F1 score.
[0197] Please refer to Figure 12 The embodiments of the present application also provide a user classification device 600 applicable to an electronic device, which includes a feature data obtaining module 610 and a user classification module 620.
[0198] The feature data obtaining module 610 is configured to obtain feature data of a to-be-classified user. The user classification module 620 is configured to process the feature data by a target classification model to obtain a user type corresponding to the feature data of the to-be-classified user, where the user type is a first type of user or a second type of user.
[0199] It should be noted that the device embodiments in the present application correspond to the foregoing method embodiments, and the specific principles in the device embodiments can be referred to the content in the foregoing method embodiments, which will not be described herein again.
[0200] The following will be described in combination with Figure 13An electronic device 100 provided by the present application is described.
[0201] Referring to Figure 13 Based on the classification model training method and the user classification method provided in the above embodiments, the present application further provides another electronic device 100 including a processor 102 that can execute the foregoing methods. The electronic device 100 can be a server 10 or a terminal device. The terminal device can be a smartphone, a tablet computer, a computer, a portable computer, or the like.
[0202] The electronic device 100 further includes a memory 104. The memory 104 stores programs that can execute the content of the foregoing embodiments, and the processor 102 can execute the programs stored in the memory 104.
[0203] The processor 102 can include one or more cores for processing data and a message matrix unit. The processor 102 connects various parts of the entire electronic device 100 through various interfaces and lines, executes various functions and processes data of the electronic device 100 by running or executing instructions, programs, code sets or instruction sets stored in the memory 104, and calling data stored in the memory 104. Alternatively, the processor 102 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 102 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 102, but can be realized by a separate communication chip.
[0204] The memory 104 can include a random access memory (RAM) and can also include a read-only memory (ROM). The memory 104 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 104 can include a program storage area and a data storage area, where the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing each of the methods described below, and the like. The data storage area can also store data (such as training samples and validation samples) acquired by the electronic device 100 in use, and the like.
[0205] The electronic device 100 can also include a network module and a screen, where the network module is used to receive and send electromagnetic waves, to achieve mutual conversion between electromagnetic waves and electrical signals, and to communicate with a communication network or other devices, such as an audio playback device. The network module can include various existing circuit elements for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, and the like. The network module can communicate with various networks, such as the Internet, an intranet, a wireless network, or other devices through a wireless network. The wireless network described above can include a cellular telephone network, a wireless local area network, or a metropolitan area network. The screen can display interface content and perform data interaction.
[0206] In some embodiments, the electronic device 100 can further include a peripheral interface 106 and at least one peripheral device. The processor 102, the memory 104, and the peripheral interface 106 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral interface through a bus, a signal line, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency component 108, a camera 114, an audio component 116, a display screen 118, a power supply 122, and the like.
[0207] The peripheral interface 106 can be used to connect at least one peripheral device related to I / O (input / output) to the processor 102 and the memory 104. In some embodiments, the processor 102, the memory 104, and the peripheral interface 106 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 102, the memory 104, and the peripheral interface 106 can be implemented on a separate chip or circuit board, and the embodiments of the present application do not limit this.
[0208] The radio frequency component 108 is configured to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency component 108 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency component 108 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency component 108 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency component 108 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency component 108 can also include NFC (Near Field Communication) related circuitry, which is not limited in the present application.
[0209] The camera 114 is configured to capture images or videos (such as the to-be-detected image in the present solution). Optionally, the camera 114 includes a front camera and a rear camera. Generally, the front camera is arranged on the front panel of the electronic device 100, and the rear camera is arranged on the back of the electronic device 100. In some embodiments, the rear camera is at least two, which is any one of a main camera, a depth-of-field camera, a wide-angle camera, and a long-focus camera, to realize the background virtualization function of the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function of the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera 114 can also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0210] The audio component 116 can include a microphone and a speaker. The microphone is used to collect sound waves of a user and an environment, and convert the sound waves into an electrical signal input to the processor 102 for processing, or input to the radio frequency component 108 to realize voice communication. The microphone can be multiple for the purpose of stereo sound collection or noise reduction, and arranged at different parts of the electronic device 100 respectively. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert an electrical signal from the processor 102 or the radio frequency component 108 into sound waves. The speaker can be a traditional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert an electrical signal into a sound wave audible to humans, but also convert an electrical signal into a sound wave inaudible to humans for ranging purposes. In some embodiments, the audio component 114 can also include a headphone jack.
[0211] The display screen 118 is used to display a UI (User Interface). The UI can include graphics, text, icons, videos, and any combination thereof. When the display screen 118 is a touch display screen, the display screen 118 also has the ability to collect a touch signal on or above the surface of the display screen 118. The touch signal can be input to the processor 102 as a control signal for processing. At this time, the display screen 118 can also be used to provide a virtual button and / or a virtual keyboard, also known as a soft button and / or a soft keyboard. In some embodiments, the display screen 118 can be one, arranged on the front panel of the electronic device 100; in other embodiments, the display screen 118 can be at least two, arranged on different surfaces of the electronic device 100 or in a folding design; in yet other embodiments, the display screen 118 can be a flexible display screen, arranged on a curved surface or a folding surface of the electronic device 100. Even, the display screen 118 can also be arranged in an irregular shape other than a rectangle, i.e. a special-shaped screen. The display screen 118 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0212] The power supply 122 is used to supply power to each component in the electronic device 100. The power supply 122 can be alternating current, direct current, disposable batteries, or rechargeable batteries. When the power supply 122 includes rechargeable batteries, the rechargeable batteries can be wired rechargeable batteries or wireless rechargeable batteries. The wired rechargeable batteries are batteries charged through wired lines, and the wireless rechargeable batteries are batteries charged through wireless coils. The rechargeable batteries can also be used to support fast charging technology.
[0213] The embodiment of the present application further provides a computer readable storage medium. The computer readable medium stores program codes, and the program codes can be invoked by a processor to execute the method described in the above method embodiment.
[0214] The computer readable storage medium can be an electronic storage such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer readable storage medium comprises a non-transitory computer readable medium. The computer readable storage medium has a storage space for storing program codes for executing any of the above methods. The program codes can be read from or written into one or more computer program products. The program codes can be compressed in a suitable form, for example.
[0215] The embodiment of the present application further provides a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described in the above various optional implementation manners.
[0216] It should be noted that the information (such as one or more of device information of a user, identity information (information for real-name authentication) of the user, and the like) and feature data (including but not limited to one or more of login data, registration time, active time length, and the like for analysis) involved in the above-described embodiments of the present application are all information and data authorized by a user or authorized by all parties, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0217] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units.
[0218] Those skilled in the art can easily understand, through the above description of the embodiments, that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to perform the methods according to the embodiments of the present application.
[0219] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the embodiments disclosed herein. It is intended that the present application cover any and all variations of the present application that come within the scope of the claims and that the claims be prevailing over any prior art. It is intended that the specification and examples be considered exemplary only, with the true scope of the present application being indicated by the following claims.
[0220] It should be understood that the application is not limited to the precise construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application. The scope of the application should only be limited by the appended claims.
Claims
1. A classification model training method, characterized in that, The method comprises: obtaining at least two initial models, a plurality of training sets obtained based on training samples, the training samples comprising positive and negative training samples obtained in a registration stage and positive and negative training samples obtained in a use stage, each training set comprising positive and negative training samples corresponding to at least one stage, the positive training samples being feature data of a first type of user, and the negative training samples being feature data of a second type of user; for each initial model, training the initial model by using the positive and negative training samples in each training set to obtain a plurality of classification models corresponding to the initial model and training parameters of each classification model, each classification model corresponding to a training set; determining a target initial model based on the training parameters of the classification models corresponding to each initial model; obtaining a verification set corresponding to each training set, each verification sample in the verification set being obtained based on a user type identification operation of a user in the registration stage or the use stage, and the feature types of the training samples in the training set and the verification samples in the corresponding verification set being the same; for each classification model corresponding to the target initial model, verifying the classification model based on the verification set corresponding to the training set used to train the classification model to obtain verification parameters of the classification model; determining a target classification model based on the verification parameters of the classification models. 2.The method of claim 1, wherein, The method further comprises: for each initial model, selecting positive and negative training samples from each training set according to a target preset ratio to train the initial model, to obtain a first classification model corresponding to each preset classification threshold and training parameters of each first classification model, wherein one preset classification threshold corresponds to one first classification model, and the target preset ratio is used to represent the ratio of positive training samples to negative training samples in the training set. The method further comprises: determining a target preset classification threshold and a target initial model according to the training parameters of the first classification models corresponding to each initial model under the plurality of preset classification thresholds, wherein the first classification model corresponding to the target preset classification threshold and trained based on the target initial model is the classification model corresponding to the target initial model. 3.The method of claim 2, wherein, The method further comprises: for each initial model, selecting a plurality of groups of positive and negative training samples from one training set according to a plurality of preset ratios to train the initial model, to obtain a second classification model corresponding to each group of positive and negative training samples when a set value of a classification threshold and training parameters of the second classification model, wherein each preset ratio corresponds to one group of positive and negative training samples. The target preset ratio is determined according to training parameters of the second classification model corresponding to a plurality of preset ratios.
4. The classification model training method of claim 3, wherein, The training parameters of the second classification model are at least two, and the target preset ratio is determined according to the training parameters of the second classification model corresponding to a plurality of preset ratios, including: The training parameters of the second classification model corresponding to each preset ratio are weighted and summed to obtain a training parameter calculation value corresponding to each preset ratio, wherein the greater the value of the same training parameter, the greater the corresponding weight. The preset ratio corresponding to the maximum training parameter calculation value is selected as the target preset ratio. 5.The method of claim 2, wherein, A target preset classification threshold and a target initial model are determined according to training parameters of the first classification model corresponding to each initial model under a plurality of preset classification thresholds, including: For each preset classification threshold, the training parameters of the first classification model corresponding to each initial model under the preset classification threshold are weighted and summed to obtain a training parameter calculation value of each initial model under the preset classification threshold, and a target parameter calculation value is determined based on the training parameter value of each initial model, wherein the greater the value of the same training parameter, the greater the corresponding weight. The preset classification threshold corresponding to the maximum target parameter calculation value is determined as the target preset classification threshold. The initial model corresponding to the maximum training parameter calculation value in the training parameter calculation value of each initial model corresponding to the target classification threshold is determined as the target initial model. 6.The method of claim 2, wherein, The verification parameters of the first classification model are at least two, and the target classification model is determined based on the verification parameters of each classification model, including: The verification parameters of each classification model are weighted and summed to obtain a verification parameter calculation value corresponding to each classification model, wherein the greater the value of the same verification parameter, the greater the corresponding weight. The classification model corresponding to the maximum verification parameter calculation value is selected as the target classification model. 7.The method of claim 1, wherein, The first type of user includes a first age group user, and the second type of user includes a second age group user, and the verification set corresponding to each training set is obtained, including: The delivery threshold of each classification model corresponding to the target initial model is determined according to the proportion of positive and negative training samples in the training sample. For each classification model corresponding to the target initial model, the classification model is deployed on a server, and the classification model is used to classify and predict a plurality of to-be-predicted feature data to obtain a probability that each to-be-predicted feature data classification prediction result is a first type of user, the probabilities of each to-be-predicted feature data are sorted from large to small, and the user types corresponding to each target to-be-predicted feature data and each target registered user are obtained as verification samples in a verification set corresponding to a training set for training the classification model, and the user type corresponding to each target registered user is obtained based on performing a user type identification operation on a user in a registration stage or a use stage. 8.The method of claim 7, wherein, The delivery threshold corresponding to each classification model is determined according to the proportion of positive and negative training samples in the training sample, including: The delivery threshold corresponding to each classification model is determined according to the preset confidence and the proportion of positive and negative training samples obtained in the registration stage. 9.The method of claim 1, wherein, The method comprises: Obtaining a plurality of training sets based on training samples, comprising: Obtaining training samples, the training samples comprising positive training samples and negative training samples, the positive training samples comprising positive and negative training samples obtained in a registration stage and positive training samples obtained in a use stage, and the negative training samples comprising negative training samples obtained in the registration stage and negative training samples obtained in the use stage; 10.The method of claim 9, wherein, If the ratio of the positive training samples to the negative training samples in the training samples is less than a preset threshold, the positive training samples in the training samples are expanded to obtain expanded training samples, and a plurality of training sets are obtained based on the expanded training samples. The method comprises: 11.The method of any one of claims 1-10, wherein, The positive training samples in the training samples are expanded based on an oversampling algorithm to obtain expanded training samples.
12. A method of classifying users, characterized by, The training parameters comprise at least one of precision, recall, AUC value and F1 score, and the model parameters of the verification parameters comprise at least one of precision, recall, AUC value and F1 score. The method comprises: Obtaining feature data of a user to be classified; 13. A classification model training apparatus characterized by comprising: Processing the feature data by using a target classification model determined by the classification model training method of any one of claims 1-11 to obtain a user type corresponding to the feature data of the user to be classified, the user type being a first type of user or a second type of user. The method comprises: A first obtaining module is configured to obtain at least two initial models and a plurality of training sets based on training samples, the training samples comprising positive and negative training samples obtained in a registration stage and positive and negative training samples obtained in a use stage, and each training set comprising positive and negative training samples corresponding to at least one stage, the positive training samples being feature data of a first type of user, and the negative training samples being feature data of a second type of user; A training module is configured to train each initial model by using the positive and negative training samples in each training set to obtain a plurality of classification models corresponding to the initial model and training parameters of each classification model; An initial model determining module is configured to determine a target initial model based on the training parameters of the classification models corresponding to each initial model; A second obtaining module is configured to obtain a verification set corresponding to each training set, each verification sample in the verification set being obtained based on user type identification operations in the registration stage or the use stage, and the feature types of the training samples in the training set and the verification samples in the corresponding verification set being the same; A verification module is configured to verify each classification model corresponding to the target initial model based on the verification set corresponding to the training set used to train the classification model to obtain verification parameters of the classification model; 14. A user classification apparatus characterized by comprising: A classification model determining module is configured to determine a target classification model based on the verification parameters of the classification models. The device comprises: A feature data obtaining module is configured to obtain feature data of a user to be classified; The user classification module is configured to determine a user type corresponding to the feature data of the user to be classified by processing the feature data using the target classification model determined by the classification model training apparatus in claim 13, wherein the user type is a first type of user or a second type of user.
15. An electronic device, comprising: The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or 12. The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or 12. The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or 12. The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or 12.
16. A computer readable storage medium, characterized in that, The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or 12. The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or 12. The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or 12. The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or 12. The computer readable storage medium stores program codes, which can be invoked by a processor to execute the method according to any one of claims 1-11 or
Citation Information
Patent Citations
Model training method and device, storage medium and electronic equipment
CN111191590A
Training method of user classification network, and user classification method and device
CN113457167A