Machine Learning-Based Real User Identification Method, Apparatus, Device, and Medium
Through the real user identification method based on machine learning, combining basic information and behavioral data, and using logistic regression algorithms to predict real user probability, the problem of low accuracy of real user identification in the existing technology is solved and data security is improved.
Patent Information
- Application Number
- CN202210507727.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-05-10
AI Technical Summary
The prior art uses basic information for identity identification, and it is impossible to effectively determine whether the user is a real user, resulting in a reduction in the accuracy of real user identification and the data security is threatened.
A real user identification method based on machine learning is proposed. By obtaining the basic information of the target identity, verifying data and behavioral data sets, rating and classification prediction, and using a classification prediction model trained by logistic regression algorithm to predict real user probability.
The accuracy of real user identification is improved, and by combining basic information and behavioral data, the judgment of user identity is enhanced and data security risks are reduced.
Smart Images

Figure CN114817878B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a real user identification method, device, equipment and medium based on machine learning. Background Art
[0002] Existing identity recognition technologies use basic information for identity recognition. For example, face brushing, iris recognition, document recognition, fingerprint recognition, etc. Since the user operates on another terminal, there will inevitably be "loopholes" that can be exploited, resulting in vulnerability to being exploited. Merely using basic information for identity recognition cannot determine whether the person operating on the terminal is a real user, reducing the accuracy of real user recognition and data security. Summary of the Invention
[0003] The main purpose of this application is to provide a real user identification method, device, equipment and medium based on machine learning, aiming to solve the problem that existing technologies use basic information for identity recognition, which is vulnerable to being exploited due to the user operating on another terminal, reducing the accuracy of real user recognition.
[0004] To achieve the above invention purpose, this application proposes a real user identification method based on machine learning, and the method includes:
[0005] Obtain target basic information verification data and a target behavior data set corresponding to a target identity identifier;
[0006] Score each verification category of the target basic information verification data to obtain a single verification category score set;
[0007] Score each behavior category of the target behavior data set to obtain a single behavior category score set;
[0008] Input the single verification category score set and the single behavior category score set into a preset classification prediction model for real user probability prediction to obtain a real user probability prediction result, where the classification prediction model is a model trained based on a logistic regression algorithm;
[0009] Determine a real user identification result according to the real user probability prediction result and a preset probability threshold.
[0010] Further, the step of scoring each verification category of the target basic information verification data to obtain a single verification category score set includes:
[0011] Divide the target basic information verification data according to the verification category to obtain multiple single verification category verification data sets;
[0012] Score each validation data in the specified single validation category validation dataset to obtain a single validation data score, where the specified single validation category validation dataset is any one of the single validation category validation datasets;
[0013] Filter the scores of each single validation data to obtain the single validation category score corresponding to the specified single validation category validation dataset;
[0014] Use the scores of each single validation category as the single validation category score set.
[0015] Further, the step of filtering the scores of each single validation data to obtain the single validation category score corresponding to the specified single validation category validation dataset includes:
[0016] Obtain the first comprehensive type corresponding to the specified single validation category validation dataset;
[0017] When the first comprehensive type is the maximum value, find the single validation data score with the maximum value among the scores of each single validation data as the single validation category score corresponding to the specified single validation category validation dataset;
[0018] When the first comprehensive type is weighted summation, perform weighted summation on the scores of each single validation data according to the weight data corresponding to the specified single validation category validation dataset to obtain the single validation category score corresponding to the specified single validation category validation dataset.
[0019] Further, the step of scoring each behavior category of the target behavior dataset to obtain a single behavior category score set includes:
[0020] Divide the target behavior dataset according to the behavior category to obtain multiple single behavior category behavior datasets;
[0021] Score each behavior data in the specified single behavior category behavior dataset to obtain a single behavior data score, where the specified single behavior category behavior dataset is any one of the single behavior category behavior datasets;
[0022] Filter the scores of each single behavior data to obtain the single behavior category score corresponding to the specified single behavior category behavior dataset;
[0023] Use the scores of each single behavior category as the single behavior category score set.
[0024] Further, the step of filtering the scores of each single behavior data to obtain the single behavior category score corresponding to the specified single behavior category behavior dataset includes:
[0025] Obtain a second comprehensive type corresponding to the specified single - behavior category behavior data set;
[0026] When the second comprehensive type is the maximum value, find the single - behavior data score with the maximum value from each of the single - behavior data scores as the single - behavior category score corresponding to the specified single - behavior category behavior data set;
[0027] When the second comprehensive type is weighted summation, perform weighted summation on each of the single - behavior data scores according to the weight data corresponding to the specified single - behavior category behavior data set to obtain the single - behavior category score corresponding to the specified single - behavior category behavior data set.
[0028] Further, before the step of inputting the single - verification category score set and the single - behavior category score set into a preset classification prediction model to perform real - user probability prediction to obtain a real - user probability prediction result, it further includes:
[0029] Obtain a plurality of training samples and an initial model, where the initial model is a model obtained based on a logistic regression algorithm;
[0030] Use a preset division ratio to divide each of the training samples to obtain a first sample set, a second sample set, and a third sample set;
[0031] Use the first sample set to train the initial model to obtain a first model;
[0032] Use the second sample set to train the initial model to obtain a second model;
[0033] Use the third sample set to train the initial model to obtain a third model;
[0034] Calculate the average value of the same parameter for each parameter of the first model, each parameter of the second model, and each parameter of the third model to obtain a target parameter set;
[0035] Use the target parameter set to update each parameter of the initial model to obtain the classification prediction model.
[0036] Further, the step of determining a real - user recognition result according to the real - user probability prediction result and a preset probability threshold includes:
[0037] When the real - user probability prediction result is greater than the probability threshold, determine that the real - user recognition result is a real user;
[0038] When the true user probability prediction result is less than or equal to the probability threshold, determine that the true user identification result is a non-true user.
[0039] This application also proposes a true user identification device based on machine learning. The device includes:
[0040] A data acquisition module for acquiring target basic information verification data and a target behavior data set corresponding to a target identity identifier;
[0041] A first scoring module for scoring the target basic information verification data for each verification category to obtain a single verification category score set;
[0042] A second scoring module for scoring the target behavior data set for each behavior category to obtain a single behavior category score set;
[0043] A true user probability prediction result determination module for inputting the single verification category score set and the single behavior category score set into a preset classification prediction model for true user probability prediction to obtain a true user probability prediction result, where the classification prediction model is a model trained based on a logistic regression algorithm;
[0044] A true user identification result determination module for determining a true user identification result according to the true user probability prediction result and a preset probability threshold.
[0045] This application also proposes a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0046] This application also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0047] The true user identification method, device, equipment and medium based on machine learning of this application. The method realizes true user probability prediction based on target basic information verification data and a target behavior data set, increases user features, and improves the accuracy of true user identification; the classification prediction model is a model trained based on a logistic regression algorithm, which improves the accuracy of true user probability prediction and further improves the accuracy of true user identification; for the scoring of each verification category and each behavior category, the influence of a single category on the scoring result is reduced, and the accuracy of true user identification is further improved. Description of the Drawings
[0048] Figure 1Schematic flowchart of a real user identification method based on machine learning according to an embodiment of the present application;
[0049] Figure 2 Schematic block diagram of the structure of a real user identification device based on machine learning according to an embodiment of the present application;
[0050] Figure 3 Schematic block diagram of the structure of a computer device according to an embodiment of the present application.
[0051] The realization, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.
[0053] Refer to Figure 1 , an embodiment of the present application provides a real user identification method based on machine learning, and the method includes:
[0054] S1: Obtain target basic information verification data and a target behavior data set corresponding to a target identity identifier;
[0055] S2: Score each verification category for the target basic information verification data to obtain a single verification category score set;
[0056] S3: Score each behavior category for the target behavior data set to obtain a single behavior category score set;
[0057] S4: Input the single verification category score set and the single behavior category score set into a preset classification prediction model for real user probability prediction to obtain a real user probability prediction result, where the classification prediction model is a model trained based on a logistic regression algorithm;
[0058] S5: Determine a real user identification result according to the real user probability prediction result and a preset probability threshold.
[0059] This embodiment realizes real user probability prediction based on target basic information verification data and a target behavior data set, increases user features, and improves the accuracy of real user identification; the classification prediction model is a model trained based on a logistic regression algorithm, which improves the accuracy of real user probability prediction and further improves the accuracy of real user identification; for the scoring of each verification category and each behavior category, the influence of a single category on the scoring result is reduced, and the accuracy of real user identification is further improved.
[0060] For S1, a real user identification request input by a user can be obtained, or a real user identification request sent by a third-party application system can be obtained, or a real user identification request triggered by a program implementing this application according to preset conditions can be obtained. For example, the preset condition is to send a real user identification request when applying to use function A.
[0061] A real user identification request is a request to identify whether the user performing the operation is a real user.
[0062] Among them, the identity identifier carried in the real user identification request is used as the target identity identifier. The identity identifier can be data that uniquely identifies a user, such as a user account, an ID number, etc.
[0063] A real user is the actual owner of the identity identifier. Since the user is operating on another terminal, it is possible that a non-real user steals the account corresponding to the identity identifier for operation.
[0064] Among them, the basic information verification data corresponding to the target identity identifier can be obtained from the local database as the target basic information verification data, or the basic information verification data corresponding to the target identity identifier can be obtained from a third-party application system as the target basic information verification data, or the basic information verification data corresponding to the target identity identifier input by the user can be obtained as the target basic information verification data.
[0065] Among them, the behavior data set corresponding to the target identity identifier can be obtained from the local database as the target behavior data set, or the behavior data set corresponding to the target identity identifier can be obtained from a third-party application system as the target behavior data set, or the behavior data set corresponding to the target identity identifier input by the user can be obtained as the target behavior data set.
[0066] It can be understood that each data in the basic information verification data can be distributed and stored in one or more storage spaces. Each behavior data in the behavior data set can be distributed and stored in one or more storage spaces.
[0067] The basic information verification data includes but is not limited to: mobile communication operator authentication records, bank card binding records, face recognition records, Ministry of Public Security real-name authentication records, and manual return visit records.
[0068] The behavior data set includes but is not limited to: querying identity information records, querying order information records, opening an in-app wallet account records, querying customer rights and interests records, receiving customer rights and interests records, using customer rights and interests records, binding vehicle information records, purchasing product records, and signing in continuously for more than N days.
[0069] For S2, each verification category may correspond to multiple verification methods. For the accuracy of subsequent real user identification, it is necessary to score each verification category to remove noisy data. Therefore, score the target basic information verification data for each verification category, use the result of scoring for one verification category as the single verification category score, and use each single verification category score as the single verification category score set.
[0070] For S3, each behavior category may correspond to multiple behavior methods. For the accuracy of subsequent real user identification, it is necessary to score each behavior category to remove noisy data. Therefore, score the target behavior data set for each behavior category, use the result of scoring for one behavior category as the single behavior category score, and use each single behavior category score as the single behavior category score set.
[0071] For S4, input the single verification category score set and the single behavior category score set into a preset classification prediction model to predict the real user probability, and use the predicted data as the real user probability prediction result.
[0072] It can be understood that the real user probability prediction result is a value between 0 and 1, which can be equal to 0 or equal to 1.
[0073] Among them, use multiple training samples to train the model obtained based on the logistic regression algorithm, and use the trained model as the classification prediction model. The training samples include: single verification category score sample set, single behavior category score sample set, and real user probability calibration value. The single verification category score sample set includes scores for multiple verification categories. The single behavior category score sample set includes scores for multiple behavior categories. The real user probability calibration value is the accurate calibration result of whether it is a real user for the single verification category score sample set and the single behavior category score sample set with the same identity identifier. When the real user probability calibration value is 0, it means it is not a real user operation; when the real user probability calibration value is 1, it means it is a real user operation.
[0074] For S5, if the real user probability prediction result is greater than the probability threshold, determine that the real user identification result is a real user; if the real user probability prediction result is less than or equal to the probability threshold, determine that the real user identification result is a non-real user.
[0075] Optionally, after the step of determining the real user identification result according to the real user probability prediction result and the preset probability threshold, it further includes: if the real user identification result is a non-real user, log out the account corresponding to the target identity identifier, and mark the account status corresponding to the target identity identifier as abnormal. Thus, the expansion of losses is avoided.
[0076] In one embodiment, the step of scoring each verification category for the target basic information verification data to obtain a single verification category score set includes:
[0077] S21: Divide the target basic information verification data according to the verification category to obtain multiple single verification category verification data sets;
[0078] S22: Score each verification data in the specified single verification category verification data set to obtain a single verification data score, where the specified single verification category verification data set is any one of the single verification category verification data sets;
[0079] S23: Screen each of the single verification data scores to obtain a single verification category score corresponding to the specified single verification category verification data set;
[0080] S24: Use each of the single verification category scores as the single verification category score set.
[0081] In this embodiment, by dividing the target basic information verification data according to the verification category and then scoring each divided single verification category verification data set, noise data is removed, the influence of a single category on the scoring result is reduced, and the accuracy of real user identification is further improved.
[0082] For S21, divide the target basic information verification data according to the verification category, and use each divided set as a single verification category verification data set. That is to say, the verification data in the single verification category verification data set belongs to the same verification category.
[0083] For S22, use any one of the single verification category verification data sets as the specified single verification category verification data set; for each verification data in the specified single verification category verification data set, search for the verification data in a preset verification score list, and use the verification score corresponding to the found verification data in the verification score list as a single verification data score.
[0084] The verification score list includes: verification data and verification scores.
[0085] For S23, use a preset first screening condition to screen each of the single verification data scores, and use the screened scores as the single verification category scores corresponding to the specified single verification category verification data set.
[0086] Optionally, the preset first screening condition is the maximum value. That is, the single verification data score with the maximum value among each of the single verification data scores is used as the single verification category score corresponding to the specified single verification category verification dataset.
[0087] In one embodiment, the step of screening each of the single verification data scores to obtain the single verification category score corresponding to the specified single verification category verification dataset includes:
[0088] S231: Obtain a first comprehensive type corresponding to the specified single verification category verification dataset;
[0089] S232: When the first comprehensive type is the maximum value, find the single verification data score with the maximum value from each of the single verification data scores as the single verification category score corresponding to the specified single verification category verification dataset;
[0090] S233: When the first comprehensive type is weighted summation, perform weighted summation on each of the single verification data scores according to the weight data corresponding to the specified single verification category verification dataset to obtain the single verification category score corresponding to the specified single verification category verification dataset.
[0091] In this embodiment, different screening methods are adopted according to the first comprehensive type, meeting the personalized screening requirements, improving the accuracy of the single verification category score, and further improving the accuracy of real user identification.
[0092] For S231, the verification category corresponding to the specified single verification category verification dataset is searched for in the preset verification category mapping table, and the comprehensive type corresponding to the found verification category in the verification category mapping table is used as the first comprehensive type.
[0093] The verification category mapping table includes: verification category and comprehensive type. The value range of the comprehensive type includes: maximum value and weighted summation.
[0094] For S232, when the first comprehensive type is the maximum value, it means screening the maximum value as the single verification category score. Therefore, the single verification data score with the maximum value is found from each of the single verification data scores, and the found single verification data score is used as the single verification category score corresponding to the specified single verification category verification dataset.
[0095] For S233, when the first comprehensive type is weighted summation, it means that the result of the weighted summation is used as the single verification category score. Therefore, the verification category corresponding to the specified single verification category verification dataset is looked up from the verification category mapping table, and the weight data corresponding to the looked-up verification category in the verification category mapping table is used to perform a weighted summation on each of the single verification data scores, and the data obtained from the weighted summation is used as the single verification category score corresponding to the specified single verification category verification dataset.
[0096] In one embodiment, the step of scoring each behavior category of the target behavior dataset to obtain a single behavior category score set includes:
[0097] S31: Divide the target behavior dataset according to the behavior category to obtain a plurality of single behavior category behavior datasets;
[0098] S32: Score each behavior data in the specified single behavior category behavior dataset to obtain a single behavior data score, where the specified single behavior category behavior dataset is any one of the single behavior category behavior datasets;
[0099] S33: Screen each of the single behavior data scores to obtain the single behavior category score corresponding to the specified single behavior category behavior dataset;
[0100] S34: Use each of the single behavior category scores as the single behavior category score set.
[0101] In this embodiment, by dividing the target behavior dataset according to the behavior category and then scoring each divided single behavior category behavior dataset, noise data is removed, the influence of a single category on the scoring result is reduced, and the accuracy of real user identification is further improved.
[0102] For S31, divide the target behavior dataset according to the behavior category, and use each divided set as a single behavior category behavior dataset. That is to say, the behavior data in the single behavior category behavior dataset belongs to the same behavior category.
[0103] For S32, use any one of the single behavior category behavior datasets as the specified single behavior category behavior dataset; look up each behavior data in the specified single behavior category behavior dataset from the preset behavior score list, and use the behavior score corresponding to the looked-up behavior data in the behavior score list as a single behavior data score.
[0104] The behavior score list includes: behavior data and behavior scores.
[0105] Optionally, input the specified behavior data into the behavior data scoring model corresponding to the specified behavior data to obtain the single behavior data score corresponding to the specified behavior data, where the specified behavior data is any behavior data in the specified single behavior category behavior dataset.
[0106] Optionally, the behavior data scoring model is a scoring model trained based on a neural network.
[0107] Optionally, the behavior data scoring model can also adopt a scoring card model.
[0108] For S33, use a preset second screening condition to screen each of the single behavior data scores, and use the screened scores as the single behavior category scores corresponding to the specified single behavior category behavior dataset.
[0109] Optionally, the preset second screening condition is the maximum value. That is, use the single behavior data score with the maximum value among each of the single behavior data scores as the single behavior category score corresponding to the specified single behavior category behavior dataset.
[0110] In one embodiment, the step of screening each of the single behavior data scores to obtain the single behavior category score corresponding to the specified single behavior category behavior dataset includes:
[0111] S331: Obtain the second comprehensive type corresponding to the specified single behavior category behavior dataset;
[0112] S332: When the second comprehensive type is the maximum value, find the single behavior data score with the maximum value from each of the single behavior data scores as the single behavior category score corresponding to the specified single behavior category behavior dataset;
[0113] S333: When the second comprehensive type is weighted summation, perform weighted summation on each of the single behavior data scores according to the weight data corresponding to the specified single behavior category behavior dataset to obtain the single behavior category score corresponding to the specified single behavior category behavior dataset.
[0114] This embodiment adopts different screening methods according to the second comprehensive type, meets the personalized screening requirements, improves the accuracy of the single behavior category score, and further improves the accuracy of real user identification.
[0115] For S331, look up the behavior category corresponding to the specified single behavior category behavior dataset from the preset behavior category mapping table, and use the comprehensive type corresponding to the found behavior category in the behavior category mapping table as the second comprehensive type.
[0116] The behavior category mapping table includes: behavior categories and comprehensive types.
[0117] For S332, when the second comprehensive type is the maximum value, it means screening the maximum value as the single behavior category score. Therefore, find the single behavior data score with the largest value from each of the single behavior data scores, and use the found single behavior data score as the single behavior category score corresponding to the specified single behavior category behavior dataset.
[0118] For S333, when the second comprehensive type is weighted summation, it means using the result of the weighted summation as the single behavior category score. Therefore, find the behavior category corresponding to the specified single behavior category behavior dataset from the behavior category mapping table, and use the weight data corresponding to the found behavior category in the behavior category mapping table to perform weighted summation on each of the single behavior data scores, and use the data obtained from the weighted summation as the single behavior category score corresponding to the specified single behavior category behavior dataset.
[0119] In one embodiment, before the step of inputting the single verification category score set and the single behavior category score set into a preset classification prediction model to perform real user probability prediction and obtaining the real user probability prediction result, the following steps are further included:
[0120] S41: Obtain a plurality of training samples and an initial model, where the initial model is a model obtained based on a logistic regression algorithm;
[0121] S42: Divide each of the training samples according to a preset division ratio to obtain a first sample set, a second sample set, and a third sample set;
[0122] S43: Train the initial model using the first sample set to obtain a first model;
[0123] S44: Train the initial model using the second sample set to obtain a second model;
[0124] S45: Train the initial model using the third sample set to obtain a third model;
[0125] S46: Calculate the average value of the same parameter for each parameter of the first model, each parameter of the second model, and each parameter of the third model to obtain a target parameter set;
[0126] S47: Update each parameter of the initial model using the target parameter set to obtain the classification prediction model.
[0127] In this embodiment, the first sample set, the second sample set, and the third sample set are used to train models respectively. Then, the average value of the same parameter of each of the three trained models is calculated, and the calculated parameters are used to update the parameters of the initial model. The updated initial model is used as the classification prediction model, thus avoiding the low accuracy of using multiple training samples for the same model, improving the robustness of the classification prediction model, and further improving the accuracy of real user probability prediction.
[0128] For S41, multiple training samples and an initial model can be obtained from the local database, or from a third-party application system, or the multiple training samples and the initial model input by the user can be obtained.
[0129] For S43, each of the training samples in the first sample set is used to train the initial model in sequence, and the initial model after the training is completed is used as the first model.
[0130] Optionally, each of the training samples in the first sample set is used to train the initial model in sequence and in a loop, and the initial model after the training is completed is used as the first model. Sequentially and in a loop means that, using the traversal method, the training samples are sequentially obtained from the first sample set to train the initial model. If the traversal is completed and the initial model after being trained with the first sample set has not converged, then the first sample set is traversed again to continue training the initial model.
[0131] For S44, each of the training samples in the second sample set is used to train the initial model in sequence, and the initial model after the training is completed is used as the second model.
[0132] Optionally, each of the training samples in the second sample set is used to train the initial model in sequence and in a loop, and the initial model after the training is completed is used as the second model. Sequentially and in a loop means that, using the traversal method, the training samples are sequentially obtained from the second sample set to train the initial model. If the traversal is completed and the initial model after being trained with the second sample set has not converged, then the second sample set is traversed again to continue training the initial model.
[0133] For S45, each of the training samples in the third sample set is used to train the initial model in sequence, and the initial model after the training is completed is used as the third model.
[0134] Optionally, each of the training samples in the third sample set is sequentially and cyclically used to train the initial model, and the initial model after the training is completed is used as the third model. Sequentially and cyclically means, that is, using a traversal method, sequentially obtaining the training samples from the third sample set to train the initial model. If the traversal is completed and the initial model after being trained with the third sample set has not converged, then perform the next traversal of the third sample set to continue training the initial model.
[0135] For S46, calculate the average value of the same parameter for each parameter of the first model, each parameter of the second model, and each parameter of the third model, and use the calculated data as the target parameter set. The number of parameters in the target parameter set is the same as the number of parameters of the first model, the number of parameters in the target parameter set is the same as the number of parameters of the second model, and the number of parameters in the target parameter set is the same as the number of parameters of the third model.
[0136] For S47, use the target parameter set to perform replacement and update of the same parameter for each parameter of the initial model, and use the initial model after the replacement and update is completed as the classification prediction model.
[0137] In one embodiment, the step of determining the real user identification result according to the real user probability prediction result and the preset probability threshold includes:
[0138] S51: When the real user probability prediction result is greater than the probability threshold, determine that the real user identification result is a real user;
[0139] S52: When the real user probability prediction result is less than or equal to the probability threshold, determine that the real user identification result is a non-real user.
[0140] This embodiment determines the real user identification result according to the real user probability prediction result and the preset probability threshold, improving the accuracy of the real user identification result.
[0141] For S51, when the real user probability prediction result is greater than the probability threshold, it means that the target basic information verification data and the target behavior data set corresponding to the target identity identifier meet the identification requirements of a real user. Therefore, determine that the real user identification result is a real user.
[0142] For S52, when the real user probability prediction result is less than or equal to the probability threshold, it means that the target basic information verification data and the target behavior data set corresponding to the target identity identifier do not meet the identification requirements of a real user. Therefore, determine that the real user identification result is a non-real user.
[0143] Referring to Figure 2 , this application also provides a real user identification device based on machine learning. The device includes:
[0144] A data acquisition module 100, configured to acquire target basic information verification data and a target behavior data set corresponding to a target identity identifier;
[0145] A first scoring module 200, configured to score the target basic information verification data for each verification category to obtain a single verification category score set;
[0146] A second scoring module 300, configured to score the target behavior data set for each behavior category to obtain a single behavior category score set;
[0147] A real user probability prediction result determination module 400, configured to input the single verification category score set and the single behavior category score set into a preset classification prediction model for real user probability prediction to obtain a real user probability prediction result, where the classification prediction model is a model trained based on a logistic regression algorithm;
[0148] A real user identification result determination module 500, configured to determine a real user identification result according to the real user probability prediction result and a preset probability threshold.
[0149] This embodiment realizes real user probability prediction based on target basic information verification data and a target behavior data set, increases user features, and improves the accuracy of real user identification; the classification prediction model is a model trained based on a logistic regression algorithm, which improves the accuracy of real user probability prediction and further improves the accuracy of real user identification; for the scoring of each verification category and each behavior category, the influence of a single category on the scoring result is reduced, and the accuracy of real user identification is further improved.
[0150] Referring to Figure 3 , this application embodiment also provides a computer device, which may be a server, and its internal structure may be as Figure 3As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as a real user identification method based on machine learning. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a real user identification method based on machine learning. The real user identification method based on machine learning includes: obtaining target basic information verification data and a target behavior data set corresponding to a target identity identifier; scoring the target basic information verification data for each verification category to obtain a single verification category score set; scoring the target behavior data set for each behavior category to obtain a single behavior category score set; inputting the single verification category score set and the single behavior category score set into a preset classification prediction model for real user probability prediction to obtain a real user probability prediction result, where the classification prediction model is a model trained based on a logistic regression algorithm; determining a real user identification result according to the real user probability prediction result and a preset probability threshold.
[0151] This embodiment realizes real user probability prediction based on target basic information verification data and a target behavior data set, increases user features, and improves the accuracy of real user identification; the classification prediction model is a model trained based on a logistic regression algorithm, which improves the accuracy of real user probability prediction and further improves the accuracy of real user identification; for the scoring of each verification category and each behavior category, it reduces the influence of a single category on the scoring result and further improves the accuracy of real user identification.
[0152] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a real user identification method based on machine learning, including the steps of: obtaining target basic information verification data and a target behavior data set corresponding to a target identity identifier; scoring the target basic information verification data for each verification category to obtain a single verification category score set; scoring the target behavior data set for each behavior category to obtain a single behavior category score set; inputting the single verification category score set and the single behavior category score set into a preset classification prediction model for real user probability prediction to obtain a real user probability prediction result, where the classification prediction model is a model trained based on a logistic regression algorithm; determining a real user identification result according to the real user probability prediction result and a preset probability threshold.
[0153] The above-mentioned machine learning-based real user identification method realizes the prediction of real user probability based on the target basic information verification data and the target behavior data set, adds user features, and improves the accuracy of real user identification; the classification prediction model is a model trained based on the logistic regression algorithm, which improves the accuracy of real user probability prediction and further improves the accuracy of real user identification; for the scores of each verification category and each behavior category, the influence of a single category on the scoring result is reduced, and the accuracy of real user identification is further improved.
[0154] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0155] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article, or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, device, article, or method including that element.
[0156] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. A method for identifying real users based on machine learning, characterized in that, The method includes: Obtaining target basic information verification data and a target behavior data set corresponding to a target identity identifier, where the target behavior data set includes a query identity information record, a query order information record, and an in-app wallet account opening record; Scoring the target basic information verification data for each verification category to obtain a single verification category score set; Scoring the target behavior data set for each behavior category to obtain a single behavior category score set; Inputting the single verification category score set and the single behavior category score set into a preset classification prediction model for real user probability prediction to obtain a real user probability prediction result, where the classification prediction model is a model trained based on a logistic regression algorithm; Determining a real user identification result according to the real user probability prediction result and a preset probability threshold; The step of scoring the target basic information verification data for each verification category to obtain a single verification category score set includes: Dividing the target basic information verification data according to the verification category to obtain multiple single verification category verification data sets; Scoring each verification data in a specified single verification category verification data set to obtain a single verification data score, where the specified single verification category verification data set is any one of the single verification category verification data sets; Screening each of the single verification data scores to obtain a single verification category score corresponding to the specified single verification category verification data set; Using each of the single verification category scores as the single verification category score set.
2. The method for identifying real users based on machine learning according to claim 1, characterized in that, The step of screening each of the single verification data scores to obtain a single verification category score corresponding to the specified single verification category verification data set includes: Obtaining a first comprehensive type corresponding to the specified single verification category verification data set; When the first comprehensive type is the maximum value, finding the single verification data score with the maximum value from each of the single verification data scores as the single verification category score corresponding to the specified single verification category verification data set; When the first comprehensive type is weighted summation, performing weighted summation on each of the single verification data scores according to the weight data corresponding to the specified single verification category verification data set to obtain the single verification category score corresponding to the specified single verification category verification data set.
3. The method for identifying real users based on machine learning according to claim 1, characterized in that, The step of scoring the target behavior data set for each behavior category to obtain a single behavior category score set includes: Dividing the target behavior data set according to the behavior category to obtain multiple single behavior category behavior data sets; Scoring each behavior data in a specified single behavior category behavior data set to obtain a single behavior data score, where the specified single behavior category behavior data set is any one of the single behavior category behavior data sets; Screening each of the single behavior data scores to obtain a single behavior category score corresponding to the specified single behavior category behavior data set; Using each of the single behavior category scores as the single behavior category score set.
4. The method for identifying real users based on machine learning according to claim 3, characterized in that, The step of screening the scores of each single-behavior data to obtain the single-behavior category score corresponding to the specified single-behavior category behavior dataset includes: Obtaining a second comprehensive type corresponding to the specified single-behavior category behavior dataset; When the second comprehensive type is the maximum value, finding the single-behavior data score with the maximum value from each of the single-behavior data scores as the single-behavior category score corresponding to the specified single-behavior category behavior dataset; When the second comprehensive type is weighted summation, performing weighted summation on each of the single-behavior data scores according to the weight data corresponding to the specified single-behavior category behavior dataset to obtain the single-behavior category score corresponding to the specified single-behavior category behavior dataset.
5. The method for identifying real users based on machine learning according to claim 1, characterized in that, Before the step of inputting the single-verification category score set and the single-behavior category score set into a preset classification prediction model to perform real-user probability prediction and obtain a real-user probability prediction result, it further includes: Obtaining a plurality of training samples and an initial model, where the initial model is a model obtained based on a logistic regression algorithm; Dividing each of the training samples according to a preset division ratio to obtain a first sample set, a second sample set, and a third sample set; Training the initial model using the first sample set to obtain a first model; Training the initial model using the second sample set to obtain a second model; Training the initial model using the third sample set to obtain a third model; Calculating the average value of the same parameter for each parameter of the first model, each parameter of the second model, and each parameter of the third model to obtain a target parameter set; Updating each parameter of the initial model using the target parameter set to obtain the classification prediction model.
6. The method for identifying real users based on machine learning according to claim 1, wherein, The step of determining the real-user identification result according to the real-user probability prediction result and a preset probability threshold includes: When the real-user probability prediction result is greater than the probability threshold, determining the real-user identification result as a real user; When the real-user probability prediction result is less than or equal to the probability threshold, determining the real-user identification result as a non-real user.
7. A device for identifying real users based on machine learning, which is used to implement the method according to any one of claims 1-6, wherein, The device includes: A data acquisition module, configured to acquire target basic information verification data and a target behavior dataset corresponding to a target identity identifier; A first scoring module, configured to score each verification category of the target basic information verification data to obtain a single-verification category score set; A second scoring module, configured to score each behavior category of the target behavior dataset to obtain a single-behavior category score set; A real-user probability prediction result determination module, configured to input the single-verification category score set and the single-behavior category score set into a preset classification prediction model to perform real-user probability prediction and obtain a real-user probability prediction result, where the classification prediction model is a model trained based on a logistic regression algorithm; A real-user identification result determination module, configured to determine the real-user identification result according to the real-user probability prediction result and a preset probability threshold.
8. A computer device, including a memory and a processor, the memory stores a computer program, wherein, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, on which a computer program is stored, wherein, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Utilizing behavioral features to identify robot
CN108604272A
Identity verification method and device
CN112613005A