A login behavior risk scoring method and system

Through the collaborative collection of multi-dimensional user behavior data and the construction of a fusion model, the shortcomings of judging the risk of logging based on a single indicator in the existing technology are solved, and a more accurate and efficient login behavior risk assessment is achieved.

CN119766578BActive Publication Date: 2025-05-13FUZHOU PUBLIC SECURITY BUREAU +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510253526.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-13
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

When judging the risk of online banking login, the existing technology is mostly based on a single indicator, and cannot fully capture the diversity and complexity of user login behavior, resulting in low accuracy, poor adaptability and lag in prevention.

Method used

Through the three parties on the public security side, operator and bank side, multi-dimensional user behavior data is collected, combined with the black and gray IP library and the improved weighted DBSCAN algorithm, abnormal IP is identified, and a joint fusion risk scoring model of the XGBoost model and the rule judgment model is constructed to perform login behavior risk scoring.

Benefits of technology

It improves the ability to identify the robbery behavior, enhances the generalization ability of the risk scoring model, significantly improves the accuracy and efficiency of the risk assessment of login behavior, and at the same time, model training is carried out on the premise of ensuring user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119766578B_ABST
    Figure CN119766578B_ABST
Patent Text Reader

Abstract

The present invention discloses a login behavior risk scoring method and system, the method includes: collecting multi-dimensional user behavior data through the public security side, the operator and the bank side to establish a multi-dimensional user behavior data set; according to the fraud data provided by the public security side, the banks on the bank side screen and extract the corresponding login IP to add to the black and gray IP library; extracting the spatiotemporal feature vector of the data sample, clustering the data sample based on the improved weighted DBSCAN algorithm, performing IP anomaly recognition detection according to the clustering result, identifying the black and gray IP and adding it to the black and gray IP library to improve the black and gray IP library; constructing an XGBoost model and a rule judgment model, fusing the trained XGBoost model with the rule judgment model to obtain a joint risk scoring model, and performing risk scoring on the login behavior to be predicted. The present invention solves the problem that the prior art cannot fully capture the diversity and complexity of user login behavior.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information security technology, and in particular relates to a login behavior risk scoring method and system. Background Art

[0002] With the rapid development of Internet technology, online banking has become an important channel for people to conduct financial transactions. However, it is accompanied by the increasing rampant cybercrime, especially the frequent occurrence of online banking fraud and fraudulent card transactions, which seriously threaten the property safety of financial institutions and users. In order to deal with these security risks, the financial and network security industries have adopted a variety of security protection measures, including identity authentication, device identification, etc.

[0003] Although these security measures have improved the security of online banking to a certain extent, due to the continuous innovation of cyber attack methods, the existing security measures still have many limitations and cannot fully cope with the increasingly complex security threats. Specifically, the existing security technical measures still have the following deficiencies:

[0004] Low accuracy: Existing technologies mostly use a single indicator to determine whether there is a risk of unauthorized login, such as identity verification or fixed device binding only through verification of login passwords, SMS verification codes, etc. These methods are either easy to crack, or lack flexibility and convenience, and cannot accurately identify complex unauthorized login behaviors.

[0005] Poor adaptability: With the continuous changes in the network environment and the continuous upgrading of attack methods, existing security measures are difficult to cope with new types of illegal login threats and lack sufficient adaptability.

[0006] Delayed prevention: Due to the lack of real-time monitoring and effective early warning mechanisms, existing security protection systems can often only detect problems after the fraudulent use of credit cards occurs, and are unable to block the fraudulent use of credit cards in a timely manner, resulting in greater losses for financial institutions and users.

[0007] The Chinese patent with the publication number CN110933080A discloses a method and device for identifying IP groups with abnormal user logins, the method comprising: obtaining login logs, counting login logs within each preset period, and obtaining login frequency sequences of each IP; using the login frequency sequence as a sample set to train an isolation forest algorithm to obtain scores for each IP address; for each score, obtaining the mode of the score, and obtaining a login log set corresponding to the mode; filtering out the frequency sequence of the login log corresponding to the mode from the login frequency sequence, and binarizing the filtered frequency sequence to obtain the mark of each IP in each period; according to the mark of each IP in each period, using the kappa algorithm to obtain the kappa coefficient between the data of the login log set, and the login log set with a kappa coefficient greater than a preset threshold is used as a login abnormal group. Although this method can effectively use the login frequency data of the IP address to identify abnormal behavior, this method only relies on the single indicator of login frequency, and cannot fully capture the diversity and complexity of user login behavior, and may have certain limitations when dealing with complex scenarios. Summary of the invention

[0008] The purpose of the present invention is to provide a login behavior risk scoring method and system to solve the problem that the existing technology is mostly based on a single indicator to judge whether there is a risk of unauthorized login, and cannot fully capture the diversity and complexity of user login behavior.

[0009] The technical solution of the present invention is as follows:

[0010] In one aspect, the present invention provides a login behavior risk scoring method, comprising the following steps:

[0011] Through the collaboration of the public security side, operators and banks, multi-dimensional user behavior data including the identity and transaction behavior data of defrauded users, account security and transaction behavior data, user communication identity and communication behavior data are collected, and a multi-dimensional user behavior data set is established for the collected multi-dimensional user behavior data.

[0012] Based on the identity and transaction behavior data of the defrauded users provided by the public security department, the banks on the banking side screen the transaction behaviors and extract the corresponding login IPs as preliminary black and gray IPs to be added to the black and gray IP database.

[0013] Based on the collected multi-dimensional user behavior data set, spatiotemporal feature vectors are extracted from data samples, and data samples are clustered based on the improved weighted DBSCAN algorithm. IP anomaly identification and detection are performed according to the clustering results. By analyzing the clustering center and combining the black-gray IP library, the black-gray IPs are identified and added to the black-gray IP library to improve the black-gray IP library. The black-gray IP library is used to construct a rule judgment model.

[0014] The multidimensional user behavior dataset is preprocessed and divided into training set and test set, the XGBoost model and rule judgment model are constructed, the XGBoost model is trained by the multidimensional user behavior dataset, the trained XGBoost model is integrated with the rule judgment model to obtain the risk joint scoring model, and the risk scoring is performed on the predicted login behavior.

[0015] Preferably, the multi-dimensional user behavior data is specifically:

[0016] The public security department provides the identity and transaction behavior data of the defrauded user, including but not limited to the user's name, ID number, household registration address, permanent address, mobile phone number, bank card number, transaction time, and transaction amount.

[0017] The bank provides account security and transaction behavior data, including but not limited to user name, ID number, bank card number, commonly used device fingerprint, transaction amount, transaction time, consumption location, login time, login IP, login device fingerprint, number of logins, and verification method change records.

[0018] Operators provide user communication identity and communication behavior data, including but not limited to user name, ID number, mobile phone number, real-name status of mobile phone number, network connection duration, abnormal status changes, base station location address, abnormal traffic, overseas roaming, SIM card replacement, and whether SMS forwarding is set up.

[0019] Preferably, the spatiotemporal feature vector is defined as:

[0020]

[0021] In the formula, is the space-time feature vector; is the geographic location offset; is the entropy value of the active time period; is the request frequency gradient; The device association degree.

[0022] The geographic location offset is solved as follows:

[0023]

[0024] In the formula, For the The spherical distance between the secondary login location and the user's permanent residence; is the historical average offset distance; The number of logins.

[0025] The entropy value of the active time period is calculated as follows:

[0026]

[0027] In the formula, is the total number of time periods; For the The probability of activity in a time period.

[0028] The request frequency gradient is solved as follows:

[0029]

[0030] In the formula, The number of login requests within the current time window; The interval between adjacent windows.

[0031] The device association degree is solved as follows:

[0032]

[0033] In the formula, The number of logged-in devices within a specific time window; The total number of logins within a specific time window.

[0034] Preferably, the improved weighted DBSCAN algorithm clusters the spatiotemporal feature vectors as follows:

[0035] Define the weighted distance measurement function, set the neighborhood radius and the adaptive adjustment strategy of the minimum number of neighborhood points.

[0036] Traverse all unvisited spatiotemporal feature vector samples to calculate the number of samples in their weighted distance neighborhood. If the number of samples in the neighborhood of the traversed sample is greater than or equal to the minimum number of neighborhood points, mark the sample as a core point and expand the cluster.

[0037] The boundary points in the neighborhood of the core point are recursively merged until there are no sample points that can be added to the cluster. The samples that are not included in any cluster are marked as noise points.

[0038] All samples are visited and classified to obtain the final clustering results.

[0039] Preferably, the weighted distance metric function is expressed as:

[0040]

[0041] In the formula, is the weighted distance metric function; For the feature weights; For sample No. eigenvalues; For sample No. eigenvalues; is the range of each feature value.

[0042] The neighborhood radius dynamic adaptive adjustment strategy is expressed as:

[0043]

[0044] In the formula, is the neighborhood radius; is the sample distance mean; is the sample distance standard deviation; is the adjustment factor.

[0045] The minimum number of neighborhood points adaptive adjustment strategy is expressed as:

[0046]

[0047] In the formula, is the minimum number of neighborhood points in the current window; is the number of samples in the current window; is the total sample number; is the scaling factor; It is the minimum neighborhood point number benchmark value.

[0048] Preferably, IP anomaly identification and detection is performed based on the clustering results. By analyzing the cluster center and combining the black and gray IP library, identifying the black and gray IP specifically includes:

[0049] Discrimination based on cluster centers: For each cluster, calculate its cluster center and compare the similarity between the cluster center and the black-gray IP spatiotemporal feature vector. If the similarity between the cluster center and the black-gray IP exceeds the preset threshold, the cluster is determined to be a black-gray IP cluster.

[0050] Discrimination based on the black-gray IP library: Calculate the overlap between each cluster and the black-gray IP in the black-gray IP library. If the overlap of the cluster exceeds the set threshold, the cluster is determined to be a black-gray IP cluster.

[0051] Preferably, the rule judgment model is specifically:

[0052] The rule base is designed based on expert experience and historical illegal login patterns, and a rule judgment model is constructed based on the black and gray IP library and the rule base. The rule judgment function of the rule judgment model is expressed as:

[0053]

[0054] In the formula, For login behavior Rule judgment function; is the black-gray IP weight factor; The number of rules triggered for login behavior, ; For rules The influence weight of .

[0055] Preferably, the XGBoost model is trained by a multi-dimensional user behavior data set, and the trained XGBoost model is integrated with the rule judgment model to obtain a joint risk scoring model. The risk scoring of the predicted login behavior is specifically as follows:

[0056] The XGBoost model is trained based on federated learning using the training set. The XGBoost model optimizes the loss function by iteratively constructing decision trees. Each decision tree is modified based on the output of the previous decision tree until the XGBoost model converges.

[0057] After training is completed, the XGBoost model outputs the user's login behavior risk assessment value through the login behavior risk assessment function, which is expressed as:

[0058]

[0059] In the formula, It is the login behavior risk assessment function; for A decision tree for login behavior The cumulative forecast results.

[0060] The trained XGBoost model is integrated with the rule judgment model to obtain a risk joint scoring model. The risk joint scoring function of the risk joint scoring model is expressed as:

[0061]

[0062] In the formula, For login behavior The rule judgment function.

[0063] If the risk combined score of the login behavior exceeds the set threshold, an early warning and transfer restriction action will be triggered.

[0064] On the other hand, the present invention provides a login behavior risk scoring system, including: a data collection module, a black-gray IP library preliminary construction module, a black-gray IP library improvement module and a risk scoring module.

[0065] The data collection module is used to collect multi-dimensional user behavior data, including the identity and transaction behavior data of the defrauded user, account security and transaction behavior data, user communication identity and communication behavior data, through the collaboration of the public security side, operators and banks, and to establish a multi-dimensional user behavior data set for the collected multi-dimensional user behavior data.

[0066] The initial construction module of the black and gray IP database is used to screen the transaction behaviors of various banks on the bank side based on the identities of the defrauded users and transaction behavior data provided by the public security side, and extract the corresponding login IP as the preliminary identified black and gray IP to be added to the black and gray IP database.

[0067] The black-gray IP library improvement module is used to extract spatiotemporal feature vectors from data samples based on the collected multi-dimensional user behavior data set, cluster the data samples based on the improved weighted DBSCAN algorithm, perform IP anomaly recognition and detection according to the clustering results, and identify black-gray IPs by analyzing the cluster centers and combining the black-gray IP library and adding them to the black-gray IP library to improve the black-gray IP library, which is used to build a rule judgment model.

[0068] The risk scoring module is used to preprocess the multidimensional user behavior data set and divide it into training set and test set, build XGBoost model and rule judgment model, train the XGBoost model through the multidimensional user behavior data set, integrate the trained XGBoost model with the rule judgment model to obtain the risk joint scoring model, and perform risk scoring on the predicted login behavior.

[0069] On the other hand, the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the login behavior risk scoring method as described in any embodiment of the present invention is implemented.

[0070] Compared with the prior art, the present invention has the following technical effects:

[0071] 1. The present invention collects multi-dimensional user behavior data and establishes a multi-dimensional user behavior data set through cooperation between the public security side, operators and banks. Combined with the black and gray IP library and the improved weighted DBSCAN algorithm, abnormal IPs are identified and the black and gray IP library is improved. A risk joint scoring model of "XGBoost+rule judgment" is constructed, which can more accurately determine whether the user login is a stolen login behavior. At the same time, the present invention uses privacy computing technology to realize model training under the premise of protecting user privacy.

[0072] 2. The present invention improves the ability to identify illegal login behaviors and enhances the generalization ability of the risk joint scoring model by integrating the XGBoost model and the rule model. The rule model can be dynamically expanded and easily adjusted, which greatly improves the accuracy and efficiency of login behavior risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 It is an overall flow chart of the login behavior risk scoring method described in the present invention. DETAILED DESCRIPTION

[0074] In order to make the objectives, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in combination with specific embodiments of the present application and with reference to the accompanying drawings.

[0075] Embodiment 1

[0076] This embodiment provides a login behavior risk scoring method, which combines privacy computing technology to collaboratively use the data from the public security side, the operator side, and the bank side to perform login behavior risk scoring without leaving the domain. Figure 1 As shown, the following steps are included:

[0077] Through the collaboration of the public security side, operators and banks, multi-dimensional user behavior data including the identity and transaction behavior data of defrauded users, account security and transaction behavior data, user communication identity and communication behavior data are collected, and a multi-dimensional user behavior data set is established for the collected multi-dimensional user behavior data.

[0078] As a preferred implementation of this embodiment, the data provided by the three parties in the multi-dimensional user behavior data all contain the user's name and ID number for data alignment, specifically:

[0079] The public security department provides the identity and transaction behavior data of the defrauded user, including but not limited to the user's name, ID number, household registration address, permanent address, mobile phone number, bank card number, transaction time, and transaction amount.

[0080] The bank provides account security and transaction behavior data, including but not limited to user name, ID number, bank card number, commonly used device fingerprint, transaction amount, transaction time, consumption location, login time, login IP, login device fingerprint, number of logins, and verification method change records.

[0081] Operators provide user communication identity and communication behavior data, including but not limited to user name, ID number, mobile phone number, real-name status of mobile phone number, network connection duration, abnormal status changes, base station location address, abnormal traffic, overseas roaming, SIM card replacement, and whether SMS forwarding is set up.

[0082] Based on the identity and transaction behavior data of the victim provided by the public security department, the banks on the banking side screen the transaction behavior and extract the corresponding login IP as the preliminary black and gray IP to be added to the black and gray IP database. Furthermore, the public security and banking sides build the Private Set Intersection (PSI) field based on the user's ID number, bank card number and transaction time. Through the PSI, the banks on the banking side screen the transaction records, extract the login IP and mark it as a black and gray IP.

[0083] As a preferred implementation of this embodiment, in order to ensure the effectiveness and timeliness of the black-gray IP library, an attenuation mechanism can be introduced to optimize the black-gray IP library. Specifically, a black-gray IP attenuation function is introduced to perform weight evaluation on IPs in the black-gray IP library that are no longer judged as black-gray IPs. The black-gray IP attenuation function is expressed as:

[0084]

[0085] In the formula, is the black-gray IP attenuation function; is the decay rate factor, which is used to control the decay speed; is the current time; is the start time, which can indicate the time when the black and gray IP enters the black and gray IP database; is the maximum time window.

[0086] The black-gray IP decay function represents the importance of the black-gray IP. Over time, the weight of the black-gray IP will gradually decrease. When it is lower than the preset threshold, the black-gray IP can be removed from the black-gray IP database. The decay mechanism introduces nonlinear decay and adjusts the aging speed of the black-gray IP according to actual conditions, ensuring the effectiveness and timeliness of the black-gray IP database.

[0087] In addition, when we receive feedback from users that they cannot make normal transfers due to IP restrictions, they need to be removed from the black and gray IP database.

[0088] Based on the collected multi-dimensional user behavior data set, spatiotemporal feature vectors are extracted from data samples, and data samples are clustered based on the improved weighted DBSCAN algorithm. IP anomaly identification and detection are performed according to the clustering results. By analyzing the clustering center and combining the black-gray IP library, the black-gray IPs are identified and added to the black-gray IP library to improve the black-gray IP library. The black-gray IP library is used to construct a rule judgment model.

[0089] As a preferred implementation of this embodiment, the spatiotemporal feature vector is defined as:

[0090]

[0091] In the formula, is the space-time feature vector; is the geographic location offset; is the entropy value of the active time period; is the request frequency gradient; The device association degree.

[0092] The geographic location offset is used to calculate the offset between the user's login location and the historical common location under the same IP address. The solution is as follows:

[0093]

[0094] In the formula, For the The spherical distance between the second login location and the user's permanent address (calculated using the Haversine formula); is the historical average offset distance; The number of logins.

[0095] The entropy value of the active time period is used to measure the randomness of the distribution of IP active time, and the solution is as follows:

[0096]

[0097] In the formula, is the total number of time periods; For the The activity probability of a time period is calculated by The percentage of login times in the total number of times during the time period.

[0098] The request frequency gradient is used to quantify the mutation characteristics of the IP login request frequency, and the solution is as follows:

[0099]

[0100] In the formula, The number of login requests within the current time window; The interval between adjacent windows.

[0101] Device association is used to evaluate the diversity of IP-associated devices and is solved as follows:

[0102]

[0103] In the formula, The number of logged-in devices within a specific time window; The total number of logins within a specific time window.

[0104] As a preferred implementation of this embodiment, the improved weighted DBSCAN algorithm clusters the spatiotemporal feature vectors as follows:

[0105] Define a weighted distance metric function, set the neighborhood radius and the adaptive adjustment strategy of the minimum number of neighborhood points. The adaptive adjustment strategy of the neighborhood radius can adaptively adjust the cluster neighborhood range to deal with the problem of uneven data density distribution. The adaptive adjustment strategy of the minimum number of neighborhood points dynamically sets the minimum density threshold of the cluster according to the change of the data volume in the time window, avoiding the misjudgment of the fixed threshold when the data volume fluctuates.

[0106] Traverse all unvisited spatiotemporal feature vector samples to calculate the number of samples in their weighted distance neighborhood. If the number of samples in the neighborhood of the traversed sample is greater than or equal to the minimum number of neighborhood points, mark the sample as a core point and expand the cluster.

[0107] The boundary points in the neighborhood of the core point are recursively merged until there are no sample points that can be added to the cluster. The samples that are not included in any cluster are marked as noise points.

[0108] The final clustering result is obtained when all samples are visited and classified, there are no more core points to expand the neighborhood, or all density-reachable points have been assigned to cluster clusters.

[0109] As a preferred implementation of this embodiment, the weighted distance metric function is expressed as:

[0110]

[0111] In the formula, is the weighted distance metric function; For the The feature weights are set according to the actual scenario requirements or determined by the optimization algorithm, and are not limited here; For sample No. eigenvalues; For sample No. eigenvalues; is the range of each characteristic value, (such as ).

[0112] The neighborhood radius dynamic adaptive adjustment strategy is adaptively calculated based on the feature space density, which is expressed as:

[0113]

[0114] In the formula, is the neighborhood radius; is the sample distance mean; is the sample distance standard deviation; It is an adjustment factor, which can be set according to actual scene requirements or experience. There is no restriction here. For example, the preferred setting interval of the experience value is .

[0115] The minimum number of neighborhood points adaptive adjustment strategy is expressed as:

[0116]

[0117] In the formula, is the minimum number of neighborhood points in the current window; is the number of samples in the current window; is the total sample number; is the scaling factor; It is the minimum neighborhood point number benchmark value.

[0118] As a preferred implementation of this embodiment, IP anomaly identification and detection is performed according to the clustering results. By analyzing the cluster center and combining the black and gray IP library, identifying the black and gray IP specifically includes:

[0119] Discrimination based on cluster centers: Calculate the cluster center for each cluster, and compare the similarity between the cluster center and the black-gray IP spatiotemporal feature vector. If the similarity between the cluster center and the black-gray IP exceeds a preset threshold, the cluster is determined to be a black-gray IP cluster. Further, for each cluster, calculate the mean of each feature to obtain the feature vector of the cluster center. The similarity between the cluster center and the black-gray IP is not limited here, and can be selected according to actual needs, such as cosine similarity, etc.

[0120] Discrimination based on the black-gray IP library: Calculate the overlap between each cluster and the black-gray IP in the black-gray IP library. If the overlap of the cluster exceeds the set threshold, the cluster is determined to be a black-gray IP cluster. Further, the overlap is expressed as the ratio of the intersection of the IP in the cluster to the IP in the black-gray IP library to the total number of IPs in the cluster.

[0121] The multidimensional user behavior data set is preprocessed and divided into a training set and a test set, an XGBoost model and a rule judgment model are constructed, the XGBoost model is trained through the multidimensional user behavior data set, and the trained XGBoost model is integrated with the rule judgment model to obtain a joint risk scoring model, and the risk scoring of the predicted login behavior is performed. Furthermore, the preprocessing includes but is not limited to: missing value processing (filling or deleting missing values ​​in the data), outlier processing (detecting and processing outliers in the data), feature conversion (converting character type data into numerical data) and feature engineering (extracting and generating new features according to requirements). In addition, before model training, multiple PSI techniques are used to intersect multiple data to align the data between multiple parties for the next step of model training.

[0122] As a preferred implementation of this embodiment, the rule judgment model is specifically:

[0123] The rule base is designed based on expert experience and historical fraudulent login patterns. The rule settings of the rule base are not limited here and can be established according to actual scenario requirements. For example, it can be designed based on categories such as geographical location anomalies, time behavior anomalies, device fingerprint risks, identity real-name risks, transaction model mutations, etc. A rule judgment model is constructed based on the black and gray IP library and the rule base.

[0124] The rule judgment function of the rule judgment model is expressed as:

[0125]

[0126] In the formula, For login behavior Rule judgment function; is the black-gray IP weight factor; The number of rules triggered for login behavior, ; For rules The influence weight of .

[0127] As a preferred implementation of this embodiment, the rule influence weight of the rule judgment function can be updated through a self-optimization mechanism, specifically:

[0128]

[0129] In the formula, For rules Updated impact weights; For rules The original impact weight; It is a dynamic learning rate, which controls the weight update amplitude; For rules The number of correct interceptions triggered; For rules The number of false interceptions; For rules The number of missed detections where fraudulent logins actually occurred but were not triggered.

[0130] As a preferred implementation of this embodiment, the XGBoost model is trained through a multi-dimensional user behavior data set, and the trained XGBoost model is integrated with the rule judgment model to obtain a risk joint scoring model. The risk scoring of the predicted login behavior is specifically as follows:

[0131] The XGBoost model is trained based on federated learning using the training set. The XGBoost model optimizes the loss function by iteratively constructing decision trees, and each decision tree is modified based on the output of the previous decision tree until the XGBoost model converges. Furthermore, the federated learning is a distributed learning method in which each participant trains the model locally and only shares model parameters or gradients. Through federated learning, the public security side, the operator side, and the bank side can jointly train an XGBoost global model while protecting their respective data privacy. The federated learning training process of the XGBoost model includes: the public security side, the operator side, and the bank side perform local XGBoost model training; after the local training is completed, the model parameters or gradients of the local training are sent to the central party, which aggregates them to generate a global XGBoost model; the global XGBoost model is returned to each party as initialization, and iterative training is performed until the XGBoost model converges.

[0132] After training is completed, the XGBoost model outputs the user's login behavior risk assessment value through the login behavior risk assessment function, which is expressed as:

[0133]

[0134] In the formula, It is the login behavior risk assessment function; for A decision tree for login behavior The cumulative forecast results.

[0135] The trained XGBoost model is integrated with the rule judgment model to obtain a risk joint scoring model. The risk joint scoring function of the risk joint scoring model is expressed as:

[0136]

[0137] In the formula, For login behavior The rule judgment function.

[0138] If the risk combined score of the login behavior exceeds the set threshold, an early warning and transfer restriction action will be triggered.

[0139] Embodiment 2

[0140] Accordingly, this embodiment provides a login behavior risk scoring system, which is used to implement the login behavior risk scoring method as described in any embodiment of the present invention, including: a data collection module, a black-gray IP library preliminary construction module, a black-gray IP library improvement module and a risk scoring module.

[0141] The data collection module is used to collect multi-dimensional user behavior data, including the identity and transaction behavior data of the defrauded user, account security and transaction behavior data, user communication identity and communication behavior data, through the collaboration of the public security side, operators and banks, and to establish a multi-dimensional user behavior data set for the collected multi-dimensional user behavior data.

[0142] The initial construction module of the black and gray IP database is used to screen the transaction behaviors of various banks on the bank side based on the identities of the defrauded users and transaction behavior data provided by the public security side, and extract the corresponding login IP as the preliminary identified black and gray IP to be added to the black and gray IP database.

[0143] The black-gray IP library improvement module is used to extract spatiotemporal feature vectors from data samples based on the collected multi-dimensional user behavior data set, cluster the data samples based on the improved weighted DBSCAN algorithm, perform IP anomaly recognition and detection according to the clustering results, and identify black-gray IPs by analyzing the cluster centers and combining the black-gray IP library and adding them to the black-gray IP library to improve the black-gray IP library, which is used to build a rule judgment model.

[0144] The risk scoring module is used to preprocess the multidimensional user behavior data set and divide it into training set and test set, build XGBoost model and rule judgment model, train the XGBoost model through the multidimensional user behavior data set, integrate the trained XGBoost model with the rule judgment model to obtain the risk joint scoring model, and perform risk scoring on the predicted login behavior.

[0145] Embodiment 3

[0146] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the login behavior risk scoring method as described in any embodiment of the present invention is implemented.

[0147] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.

[0148] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0149] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0150] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program codes.

[0151] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A login behavior risk scoring method, characterized in that: The following steps are involved: Through the collaboration of the public security, operators and banks, multi-dimensional user behavior data including the identity and transaction behavior data of the defrauded user, account security and transaction behavior data, user communication identity and communication behavior data are collected, and a multi-dimensional user behavior data set is established for the collected multi-dimensional user behavior data; Based on the identity and transaction behavior data of the victim provided by the public security department, the banks screen the transaction behaviors and extract the corresponding login IPs as the preliminary black and gray IPs to be added to the black and gray IP database; Extract spatiotemporal feature vectors from data samples based on the collected multi-dimensional user behavior data set, cluster the data samples based on the improved weighted DBSCAN algorithm, perform IP anomaly recognition and detection based on the clustering results, identify black and gray IPs by analyzing the cluster centers and combining the black and gray IP library, and add them to the black and gray IP library to improve the black and gray IP library, which is used to build a rule judgment model; Preprocess the multidimensional user behavior dataset and divide it into training set and test set, build XGBoost model and rule judgment model, train the XGBoost model with the multidimensional user behavior dataset, fuse the trained XGBoost model with the rule judgment model to obtain the joint risk scoring model, and perform risk scoring on the predicted login behavior; The spatiotemporal feature vector is defined as: In the formula, is the space-time feature vector; is the geographic location offset; is the entropy value of the active time period; is the request frequency gradient; is the device association degree; The geographic location offset is solved as follows: In the formula, For the The spherical distance between the secondary login location and the user's permanent residence; is the historical average offset distance; is the number of logins; The entropy value of the active time period is calculated as follows: In the formula, is the total number of time periods; For the Activity probability of time period; The request frequency gradient is solved as follows: In the formula, The number of login requests within the current time window; is the interval between adjacent windows; The device association degree is solved as follows: In the formula, The number of logged-in devices within a specific time window; is the total number of logins within a specific time window; Clustering the spatiotemporal feature vectors based on the improved weighted DBSCAN algorithm is specifically as follows: Define the weighted distance measurement function, set the neighborhood radius and the adaptive adjustment strategy of the minimum number of neighborhood points; Traverse all unvisited spatiotemporal feature vector samples to calculate the number of samples in their weighted distance neighborhood. If the number of samples in the neighborhood of the traversed sample is greater than or equal to the minimum number of neighborhood points, mark the sample as a core point and expand the cluster; Recursively merge the boundary points in the neighborhood of the core point until there are no sample points that can be added to the cluster. Mark the samples that are not included in any cluster as noise points. All samples were visited and classified to obtain the final clustering results; The weighted distance metric function is expressed as: In the formula, is the weighted distance metric function; For the feature weights; For sample No. eigenvalues; For sample No. eigenvalues; is the range of each feature value; The neighborhood radius dynamic adaptive adjustment strategy is expressed as: In the formula, is the neighborhood radius; is the sample distance mean; is the sample distance standard deviation; is the regulating factor; The minimum number of neighborhood points adaptive adjustment strategy is expressed as: In the formula, is the minimum number of neighborhood points in the current window; is the number of samples in the current window; is the total sample number; is the scaling factor; It is the minimum neighborhood point number benchmark value.

2. The login behavior risk scoring method according to claim 1, characterized in that: The multi-dimensional user behavior data specifically includes: The public security department provides the victim's identity and transaction behavior data, including but not limited to the user's name, ID number, household registration address, permanent address, mobile phone number, bank card number, transaction time, and transaction amount; The bank provides account security and transaction behavior data, including but not limited to user name, ID number, bank card number, commonly used device fingerprint, transaction amount, transaction time, consumption location, login time, login IP, login device fingerprint, login times, and verification method change records; Operators provide user communication identity and communication behavior data, including but not limited to user name, ID number, mobile phone number, real-name status of mobile phone number, network connection duration, abnormal status changes, base station location address, abnormal traffic, overseas roaming, SIM card replacement, and whether SMS forwarding is set up.

3. The login behavior risk scoring method according to claim 1, characterized in that: According to the clustering results, IP anomaly identification and detection are performed. By analyzing the cluster center and combining the black and gray IP database, the identification of black and gray IPs includes: Discrimination based on cluster centers: For each cluster, calculate its cluster center and compare the similarity between the cluster center and the black-gray IP spatiotemporal feature vector. If the similarity between the cluster center and the black-gray IP exceeds the preset threshold, the cluster is determined to be a black-gray IP cluster. Discrimination based on the black-gray IP library: Calculate the overlap between each cluster and the black-gray IP in the black-gray IP library. If the overlap of the cluster exceeds the set threshold, the cluster is determined to be a black-gray IP cluster.

4. The login behavior risk scoring method according to claim 1, characterized in that: The rule judgment model is specifically: The rule base is designed based on expert experience and historical illegal login patterns, and a rule judgment model is constructed based on the black and gray IP library and the rule base. The rule judgment function of the rule judgment model is expressed as: In the formula, For login behavior Rule judgment function; is the black-gray IP weight factor; The number of rules triggered for login behavior, ; For rules The influence weight of .

5. The login behavior risk scoring method according to claim 4, characterized in that: The XGBoost model is trained through a multi-dimensional user behavior dataset, and the trained XGBoost model is integrated with the rule judgment model to obtain a joint risk scoring model. The risk scoring of the predicted login behavior is as follows: The XGBoost model is trained based on federated learning using the training set. The XGBoost model optimizes the loss function by iteratively building decision trees. Each decision tree is modified based on the output of the previous decision tree until the XGBoost model converges. After training is completed, the XGBoost model outputs the user's login behavior risk assessment value through the login behavior risk assessment function, which is expressed as: In the formula, It is the login behavior risk assessment function; for A decision tree for login behavior The cumulative forecast results of The trained XGBoost model is integrated with the rule judgment model to obtain a risk joint scoring model. The risk joint scoring function of the risk joint scoring model is expressed as: In the formula, For login behavior Rule judgment function; If the risk combined score of the login behavior exceeds the set threshold, an early warning and transfer restriction action will be triggered.

6. A login behavior risk scoring system, characterized in that: The system is used to implement the login behavior risk scoring method according to any one of claims 1 to 5, comprising: a data collection module, a black-gray IP library preliminary construction module, a black-gray IP library improvement module and a risk scoring module; The data collection module is used to collect multi-dimensional user behavior data including the identity and transaction behavior data of the defrauded user, account security and transaction behavior data, user communication identity and communication behavior data through the collaboration of the public security side, the operator and the bank side, and to establish a multi-dimensional user behavior data set for the collected multi-dimensional user behavior data; The initial construction module of the black and gray IP database is used to screen the transaction behaviors of the fraudulent users and transaction behavior data provided by the public security side, and extract the corresponding login IP as the preliminary black and gray IP to be added to the black and gray IP database; The black-gray IP library improvement module is used to extract spatiotemporal feature vectors from data samples based on the collected multi-dimensional user behavior data set, cluster the data samples based on the improved weighted DBSCAN algorithm, perform IP anomaly identification and detection based on the clustering results, and identify black-gray IPs by analyzing the cluster center and combining the black-gray IP library and adding them to the black-gray IP library to improve the black-gray IP library, which is used to build a rule judgment model; The risk scoring module is used to preprocess the multidimensional user behavior data set and divide it into training set and test set, build XGBoost model and rule judgment model, train the XGBoost model through the multidimensional user behavior data set, integrate the trained XGBoost model with the rule judgment model to obtain the risk joint scoring model, and perform risk scoring on the predicted login behavior.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the login behavior risk scoring method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • IP group identification method and device for abnormal user login

    CN110933080A

  • Network behavior security early warning method and system

    CN115987615A

  • Method and device for judging abnormal association between users and electronic equipment

    CN116561608A