User Behavior Abnormal Detection Method, System and Terminal Device
By using the suffix probability tree model for supervised training and unsupervised updates in user behavior abnormality detection, the problems of poor adaptability and low detection accuracy for new business behaviors in the existing technology are solved, and higher abnormal prediction accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202011091429.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-10-13
AI Technical Summary
In the detection of user behavior abnormalities, the prior art has problems such as poor adaptability to new business behaviors and low detection accuracy, especially in complex and multi-branched behavior sequences.
Supervised training and unsupervised updates are used to use the probability distribution model based on the suffix probability tree. Anomaly detection is performed on the set of user behavior sequences to be tested, and online updates are performed in the sub-sequence containing new behavioral points.
It improves the adaptability of abnormal prediction scenarios in the user behavior sequence set to be tested, which contains new behavioral subsequences with added behavioral points, and enhances the accuracy of behavioral abnormality detection, especially in complex and multi-branched behavioral sequences.
Smart Images

Figure CN114357849B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminal devices, and particularly to a method, system, and terminal device for detecting abnormal user behavior.
Background Art
[0002] With the rapid development of the global information electronics industry, various smartphones, e-commerce platforms, and application software emerge in an endless stream. How to accurately evaluate the true feelings of users when using the above products is one of the keys to improving product quality and competitiveness. Currently, the commonly adopted solution is to collect non-privacy behavior data of users during the operation process under the premise of user authorization, and analyze abnormal problems in the user operation behavior sequence from the dimension of big data. For example, a user browses a lot on a certain e-commerce platform, but the purchase behavior is relatively small, indicating that the price or variety of the product cannot meet the user's needs. Regarding how to mine abnormal behavior sequences from user behavior big data, the existing solutions in the industry include hidden Markov models, clustering models, deep learning models, etc. based on conditional transition probabilities. These models have high requirements for the quality of samples and are difficult to adapt to the current business scenarios with rapid version iteration and update.
[0003] As Figure 1 shown, there is a commonly used abnormal detection method based on a hidden Markov model in the prior art, which constructs a user behavior abnormal detection model based on the hidden Markov model. The state set of this model corresponds to the predefined types of user behavior patterns, the observation value set is the original behavior data of the user, and the probability parameters of the hidden Markov model are obtained by training with normal behavior data. Finally, it is judged whether the user behavior sequence is normal according to the predicted probability value output by the model. However, the disadvantage of this technical solution is that it is necessary to classify the original behavior data of the user according to the behavior pattern. In practice, the user behavior patterns corresponding to different business products are not the same, it is difficult to describe them uniformly, and the adaptability to newly added behavior data is poor, and it is necessary to retrain the model parameters.
[0004] As Figure 2 shown, there is another commonly used clustering-based abnormal detection method in the prior art. According to a certain distance measurement rule, it uses the clustering method to divide user behavior data into several categories, and determines the outliers far from each group as abnormal user behaviors. However, this clustering method only judges the similarity between behavior sequences based on the distance, and the clustering result has great unpredictability. In addition, when the user behavior sequence is long, the clustering method cannot find local abnormal problems in the sequence, resulting in a high missed detection rate.
[0005] As Figure 3Shown is another commonly used anomaly detection method based on feature extraction in the prior art. Its principle is to extract features from the user behavior sequence and form a feature vector, that is, to quantify the user behavior sequence, such as the number of uses, the usage duration, etc. Then, a classification model of deep learning or neural network is constructed, and the model is trained in a supervised manner to obtain a classification model that can distinguish normal behavior sequences from abnormal behavior sequences. However, this solution requires manual determination of the feature extraction method for behaviors, which involves a large amount of work, and there is a risk of missing detection of unknown abnormal behaviors. In addition, the training of the model requires a large number of labeled samples, and for the anomaly detection scenario, the number of abnormal samples is usually much smaller than the number of normal samples, so the training of the model is more difficult.
Summary of the Invention
[0006] In view of this, embodiments of the present application provide a user behavior anomaly detection method, system, and terminal device to solve the technical problems of poor adaptability to new behavior marking in the business and low accuracy of behavior anomaly detection existing in the prior art.
[0007] In the first aspect, embodiments of the present application provide a user behavior anomaly detection method, which includes the following steps: establishing a probability distribution model based on a user behavior sequence set; training the probability distribution model based on the user behavior sequence set to obtain an initial probability distribution model; based on the initial probability distribution model, performing anomaly detection on the to-be-tested user behavior sequence set Seq s and updating the subsequences containing new behavior marking in the to-be-tested user behavior sequence set Seq s in an unsupervised manner.
[0008] Through the solution provided in this embodiment, through the supervised training of the probability distribution model, the probability distribution model can perform anomaly detection and unsupervised update on the to-be-tested user behavior sequence set, so as to improve the adaptability to the anomaly prediction scenario of the subsequences containing new behavior marking in the to-be-tested user behavior sequence set, and improve the accuracy of behavior anomaly detection, especially beneficial for improving the accuracy of anomaly detection of complex and multi-branched behavior sequences in the to-be-tested user behavior sequence set.
[0009] In a preferred implementation, in the step of training the probability distribution model based on the user behavior sequence set to obtain an initial probability distribution model, the probability distribution model has the structure of a suffix probability tree, and this step includes: training the probability distribution model with the normal behavior sequences in the user behavior sequence set to obtain an initial suffix probability tree; for the case of subsequences containing new behavior marking in the to-be-tested user behavior sequence, performing online update on the initial suffix probability tree in an unsupervised manner to obtain an updated suffix probability tree; where the to-be-tested user behavior sequence set Seqs ={seq 1 ,seq 2 ,…,seq n}, the subsequence seq containing the newly added behavior mark i ={…newAct p ,…,newAct q ,…}, newAct p and newAct q Indicates the addition of a new behavior.
[0010] Through the solution provided in this embodiment, a suffix probability tree structure is used to represent the probability distribution model, and supervised training and unsupervised updating are performed based on the suffix probability tree, thereby ensuring the basic accuracy of the probability distribution model in detecting anomalies of the user behavior sequence set to be tested, as well as the dynamic adaptability to the addition of new business behavior points.
[0011] In a preferred embodiment, the user behavior sequence set Seq to be tested based on the initial probability distribution model s Perform anomaly detection on the user behavior sequence set Seq s The step of updating the subsequence containing the newly added behavior mark in an unsupervised manner includes: detecting each behavior act in the user behavior sequence set to be tested p-1 suffix; if the behavior is act p-1 The suffix is newAct p , then the last node in the suffix probability tree is act p-1 For all nodes, add newAct to the suffix probability vector of each node p The unsupervised suffix probability vector P(newAct p |[act p-k ,…,act p-1 ]), and finally obtain the updated suffix probability tree; if the action p-1 NewAct is a new behavior management p-1 , then add the last digit newAct from the root node of the suffix probability tree in sequence p-1 The prefix node and its unsupervised suffix probability vector [P(act ’ p |[act p-k ,…,newAct p-1 ]),P(act ” p |[act p-k ,…,newAct p-1 ]),…], and finally obtain the updated suffix probability tree.
[0012] Through the solution provided by this embodiment, the set of user behavior sequences to be measured Seq containing new behavior marking is s directly input into the probability distribution model. The probability distribution model online updates the unsupervised suffix probability vectors of the nodes in the suffix probability tree, and determines whether the subsequence containing the new behavior marking is abnormal according to the characteristics of the distribution of the unsupervised suffix probability vectors, thereby avoiding re-labeling samples and training the model.
[0013] In a preferred implementation, after obtaining the updated suffix probability tree, the following steps are included: traversing each user behavior act in the updated suffix probability tree p 's set of prefix nodes {[act p-1 ,…,[act p-k ,…,act p-1}, and respectively obtaining the suffix probability vector V m =[…,P(act p |prefix m ),…]; where, prefix m represents the prefix subsequence of the mth suffix containing act p ; judging the type of the suffix probability vector V m of the prefix subsequence prefix m ; if the suffix probability vector V m of the prefix subsequence prefix m is a supervised probability vector, then judge the size relationship between P(act p |prefix m ) and the preset first threshold P thresh ; if it is less than, output the result that the subsequence is abnormal; if it is greater than, output the result that the subsequence is normal; if the suffix probability vector V m of the prefix subsequence prefix m is an unsupervised probability vector, then judge the size relationship between P(act p |prefix m ) and the preset second threshold P ’ thresh ; if it is less than, analyze the concentration of the probability distribution of the suffix probability vector V m ; if it is greater than, output the result that the subsequence is normal.
[0014] Through the solution provided by this embodiment, by traversing the suffix probability tree, the set of user behavior sequences to be measured Seq is matched sThe subsequence of the upper part, and according to whether the corresponding suffix probability vector is supervised or unsupervised, different transfer probability thresholds are used to determine whether the subsequence is abnormal, thereby realizing the anomaly detection of the entire sequence set of the user behavior to be measured, without being limited by the sequence length, and avoiding the problem of missing detection of local anomalies caused by detecting anomalies according to the sequence distance similarity.
[0015] In a preferred embodiment, the step of analyzing the concentration of the probability distribution of the suffix probability vector V m includes: judging the standard deviation Std(V m ) of the unsupervised probability vector V m and the size relationship with the preset concentration threshold S thresh ; if the standard deviation Std(V m ) is greater than the concentration threshold S thresh , then judge the size relationship between P(act p |prefix m ) and the second threshold P ’ thresh ; if it is less, then output the result that the subsequence is abnormal; if it is greater, then output the result that the subsequence is normal; if the standard deviation Std(V m ) is less than the concentration threshold S thresh , then reverse the sequence containing act s in the sequence set Seq p of the user behavior to be measured to construct an unsupervised reverse suffix probability tree, and then traverse the unsupervised reverse suffix probability tree to search for the node with act p at the end in the reverse suffix probability tree, and obtain the suffix probability vector preV m of the node; judge the size relationship between the standard deviation Std(preVm) of the suffix probability vector preV m of the node and the concentration threshold S thresh ; if the standard deviation Std(preV m ) of the suffix probability vector preV m of the node is greater than the concentration threshold S thresh , then judge the size relationship between P(prefix m |act p ) and the second threshold P ’ thresh ; if it is less, then output the result that the subsequence is abnormal; if it is greater, then output the result that the subsequence is normal; if the standard deviation Std(preV m ) of the suffix probability vector preV m of the node is less than the concentration threshold S thresh , then output the result that the subsequence is normal; where, the dot sequence seqi ={…newAct p ,…,newAct q ,…}, newAct p and newAct q Indicates the addition of a new behavior.
[0016] Through the solution provided in this embodiment, for subsequences containing newly added behavior dots, especially complex subsequences with multiple branches, a judgment method combining probability centrality analysis with reverse suffix probability tree is adopted to judge whether the behavior transfer in the subsequence is abnormal from two dimensions: the distribution of suffix probability vectors and the distribution of prefix probability vectors, thereby reducing a large number of misjudgment problems caused by judging only based on the size of the transfer probability.
[0017] In a preferred embodiment, the first threshold value P thresh With the second threshold P ’ thresh The sizes of the second threshold P are different, and ’ thresh Not zero.
[0018] Through the solution provided in this embodiment, for the supervised situation, the probability threshold P is set thresh The purpose is to avoid the situation where manual annotation has bias, such as manual annotation of act p-1 to act p The subsequence of is normal, but P(act p |act p-1 ) value did not reach the expected P thresh , which still indicates that there is a problem with this path. For the unsupervised case, set the probability threshold P' thresh The purpose of P is to identify low-probability subsequences as abnormal subsequences without manual labeling. thresh Can be 0, but P' thresh Cannot be 0.
[0019] In a preferred embodiment, the suffix probability tree includes a plurality of nodes, wherein the nodes include a local prefix subsequence of the user behavior sequence and a corresponding suffix probability vector, wherein the suffix probability vector includes an initial supervised suffix probability vector and an unsupervised updated suffix probability vector.
[0020] In a second aspect, an embodiment of the present application provides a user behavior anomaly detection system, characterized in that the system includes: a creation module for establishing a probability distribution model based on a user behavior sequence set; a training module for training the probability distribution model based on the user behavior sequence set to obtain an initial probability distribution model; a detection module for detecting the user behavior sequence set Seq to be tested based on the initial probability distribution model.s Perform anomaly detection. For the set of user behavior sequences Seq to be measured s For the subsequences containing newly added behavior markings in it, update them in an unsupervised manner.
[0021] Through the solution provided in this embodiment, through the supervised training of the probability distribution model by the training module, the detection module can perform anomaly detection and unsupervised update on the set of user behavior sequences to be measured through this probability distribution model, so as to improve the adaptability to the anomaly prediction scenario of the subsequences containing newly added behavior markings in the set of user behavior sequences to be measured, and improve the accuracy of behavior anomaly detection. In particular, it is beneficial to improve the accuracy of anomaly detection for complex and multi-branched behavior sequences in the set of user behavior sequences to be measured.
[0022] In a third aspect, an embodiment of the present application provides a terminal device, including: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the terminal device executes the method described in the first aspect.
[0023] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including a program or an instruction. When the program or the instruction runs on a computer, the method described in the first aspect is executed.
[0024] Compared with the prior art, the technical solution of the present application has at least the following beneficial effects:
[0025] The user behavior anomaly detection method, system and terminal device disclosed in the embodiments of the present application support supervised training and unsupervised update, have both high sequence anomaly prediction accuracy, and can adapt to the anomaly prediction scenario of subsequences containing newly added behavior markings in the set of user behavior sequences to be measured. This solution effectively solves the problem that traditional methods cannot quickly adapt to the needs of newly added behavior markings in the business, can greatly reduce the workload of manual re-labeling in the process of model application, and in addition, effectively reduces the false positive rate of the model detecting abnormal behaviors, and can be applied to scenarios without manual labeling and with many path branches.
Description of the Drawings
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0027] Figure 1 It is a schematic diagram of an anomaly detection method using a hidden Markov model in the prior art;
[0028] Figure 2 It is a schematic diagram of an anomaly detection method based on clustering in the prior art;
[0029] Figure 3 It is a schematic diagram of an anomaly detection method based on feature extraction in the prior art;
[0030] Figure 4 It is a schematic diagram of the overall steps of the user behavior anomaly detection method provided in Embodiment 2 of the present application;
[0031] Figure 5 It is a schematic diagram of the steps of Step200 in the user behavior anomaly detection method provided in Embodiment 2 of the present application;
[0032] Figure 6 It is a schematic diagram of the steps of Step300 in the user behavior anomaly detection method provided in Embodiment 2 of the present application;
[0033] Figure 7 It is the training process of the probability distribution model in the user behavior anomaly detection method provided in Embodiment 2 of the present application;
[0034] Figure 8 It is a schematic diagram of the steps after Step304 is executed in the user behavior anomaly detection method provided in Embodiment 2 of the present application;
[0035] Figure 9 It is the specific flowchart of Step311 in the user behavior anomaly detection method provided in Embodiment 2 of the present application;
[0036] Figure 10 It is a comparison chart of the detection results between the pure supervised method and the user behavior anomaly detection method provided in Embodiment 2 of the present application for a sequence containing new behavior logging;
[0037] Figure 11 It is a schematic diagram of the user behavior anomaly detection system provided in Embodiment 3 of the present application.
[0038] Reference numerals:
[0039] 10 - Creation module; 20 - Training module; 30 - Detection module.
Detailed implementation manners
[0040] For a better understanding of the technical solution of the present application, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0041] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.
[0042] Example 1
[0043] Embodiment 1 of the present application discloses a terminal device, comprising: a memory and a processor: the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the terminal device executes the method disclosed in Embodiment 2.
[0044] The terminal device can be a mobile phone (also known as a smart terminal device), a tablet personal computer, a personal digital assistant, an e-book reader or a virtual reality interactive device, etc. The terminal device can be accessed into various types of communication systems, such as: long term evolution (LTE) system, the future fifth generation (5G) system, the new generation of wireless access technology (NR), and future communication systems such as 6G system; it can also be a wireless local area network (WLAN), etc.
[0045] Example 2
[0046] like Figure 4 What is shown is the user behavior anomaly detection method provided in Example 2 of the present application. This method is a user behavior anomaly detection method based on a variable-order Markov probability distribution model. This method has the ability of supervised training and unsupervised updating, and is used to solve two technical problems existing in the prior art: 1) The adaptability to new business behavior is poor, and the existing method usually needs to re-label sequence samples and retrain the model; 2) The accuracy of behavior anomaly detection is low, especially for complex multi-branch behavior operation sequences.
[0047] Therefore, the user behavior abnormality detection method of this embodiment specifically includes the following steps:
[0048] Step 100: Establish a probability distribution model based on the user behavior sequence set.
[0049] In step 100, a variable-order Markov probability distribution model of the user behavior sequence is established. The model is represented by a suffix probability tree structure and is used to describe the forward and backward transition probability distribution of subsequences in the user behavior sequence set to be tested as a sample set in practical applications.
[0050] In the method of this embodiment, the probability distribution model is represented by a suffix probability tree. The suffix probability tree includes multiple nodes. A node includes a local prefix subsequence of the user behavior sequence and a corresponding suffix probability vector. The suffix probability vector includes an initial supervised suffix probability vector and an unsupervised updated suffix probability vector. The path from the root node to the leaf node represents the process of gradually expanding the prefix forward. The depth of the suffix probability tree represents the maximum order of the variable-order Markov model.
[0051] Step200: Train the probability distribution model based on the user behavior sequence set to obtain an initial probability distribution model.
[0052] In this step Step200, the probability distribution model undergoes a supervised initialization process, that is, the probability distribution parameters of the model are initialized using an artificially labeled set of normal user behavior sequences to form a probability expression for the normal user behavior sequences.
[0053] Step300: Perform anomaly detection on the set of user behavior sequences to be measured Seq s and update the subsequences containing new behavior markings in the set of user behavior sequences to be measured Seq s in an unsupervised manner.
[0054] In this step Step300, the probability distribution model undergoes an unsupervised update process. In this unsupervised update process, the subsequences containing new behavior markings are not artificially labeled. For the set of user behavior sequences to be measured Seq s containing subsequences with new business behavior markings, update the node information at the corresponding positions in the probability distribution model so that the model can represent the forward and backward transition probability distributions of the new behavior markings in the behavior sequence.
[0055] In the user behavior anomaly detection method of this embodiment, to ensure the accuracy of behavior anomaly detection, for the subsequences in the set of user behavior sequences to be measured Seq s that do not contain new behavior markings, directly use the supervised initialization to obtain the suffix transition probability vector to determine whether there is an anomaly; for the subsequences in the set of user behavior sequences to be measured Seq s that contain new behavior markings, use the suffix transition probability vector obtained by unsupervised update and combine it with probability centrality analysis and the reverse suffix probability tree to determine whether the subsequence is abnormal.
[0056] According to the steps of the above user behavior anomaly detection method, the training process of the probability distribution model based on the structure represented by the suffix probability tree is divided into two stages: one is supervised initialization, that is, using the user normal behavior sequence samples to train the probability distribution model to obtain an initial suffix probability tree containing the suffix probability distribution of local subsequences. The other is unsupervised update, that is, for the newly added user behavior points in the business, input the behavior sequence set containing the newly added points into the above initial suffix probability tree to perform online update on the suffix probability tree.
[0057] Through the supervised training and unsupervised update of the probability distribution model, the probability distribution model can perform anomaly detection on the to-be-tested user behavior sequence set, thereby improving the adaptability to the anomaly prediction scenario of the subsequences containing newly added behavior points in the to-be-tested user behavior sequence set, improving the accuracy of user behavior anomaly detection, especially beneficial for improving the accuracy of anomaly detection of complex and multi-branched behavior sequences in the to-be-tested user behavior sequence set.
[0058] As Figure 5 shown, in the user behavior anomaly detection method of this embodiment, step Step200 specifically includes the following steps:
[0059] Step201: Use the normal behavior sequences in the user behavior sequence set to train the probability distribution model to obtain an initial suffix probability tree.
[0060] Step202: For the case of the subsequence seq s containing newly added behavior points in the to-be-tested user behavior sequence Seq i adopt an unsupervised method to perform online update on the initial suffix probability tree to obtain an updated suffix probability tree.
[0061] Establish a probability distribution model including supervised probability distribution and unsupervised probability distribution, and use the structure of the suffix probability tree to represent the probability distribution model, where the supervised probability distribution is generated from the user normal sample set to ensure the basic accuracy of the model for behavior sequence anomaly detection; the unsupervised probability distribution is updated and generated from the sample set containing newly added behavior points to achieve dynamic adaptation to the newly added behavior points in the business. Based on the suffix probability tree for supervised training and unsupervised update, ensure the basic accuracy of the anomaly detection of the probability distribution model for the to-be-tested user behavior sequence set and the dynamic adaptability to the newly added behavior points in the business.
[0062] As Figure 6 and Figure 7 shown, in the user behavior anomaly detection method of this embodiment, step Step300 specifically includes the following steps:
[0063] Step301: Detect the to-be-tested user behavior sequence set Seqs Each action in p-1 suffix; if the behavior is act p-1 The suffix is newAct p , then execute Step 302; if the action is p-1 NewAct is a new behavior management p-1 , then execute Step 303.
[0064] In step 301, for the newly added user behavior mark of the service, the behavior sequence set containing the newly added mark is input into the above initial suffix probability tree, and the suffix probability tree is updated online. For the convenience of description, let the user behavior sequence set to be tested be Seqs = {seq 1 ,seq 2 ,…,seq n}, containing the subsequence seq of the newly added behavior mark i ={…newAct p ,…,newAct q ,…}, newAct p and newAct q Indicates the newly added behavior mark. Traverse the subsequence seq containing the newly added behavior mark in sequence i Dot the behavior in to update the suffix probability tree.
[0065] Step 302: Traverse the suffix probability tree and the last node is act p-1 For all nodes, add newAct to the suffix probability vector of each node p The unsupervised suffix probability vector P(newAct p |[act p-k ,…,act p-1 ]); and execute step 304. Among them, the unsupervised suffix probability P(newAct p |[act p-k ,…,act p-1 ])=Count([act p-k ,…,act p-1 ,newAct p ]) / Count([act p-k ,…,act p-1 ]), k represents the order of the Markov chain, [act p-k ,…,act p-1 ] indicates newAct p The prefix subsequence of .
[0066] Step 303: Add the last digit of the suffix probability tree to the newActp-1 and its prefix nodes and unsupervised suffix probability vectors [P(act ’ p |[act p-k ,…,newAct p-1 ), P(act ” p |[act p-k ,…,newAct p-1 ),…]; and perform step Step304.
[0067] Step304: Obtain the updated suffix probability tree.
[0068] In the user behavior anomaly detection method of this embodiment, the set of user behavior sequences to be tested Seq containing new behavior logging s is directly input into the probability distribution model, and the probability distribution model online updates the unsupervised suffix probability vectors [P(act ’ p |[act p-k ,…,newAct p-1 ), P(act ” p |[act p-k ,…,newAct p-1 ),…] of the nodes in the suffix probability tree, and based on the characteristics of the unsupervised suffix probability vectors [P(act ’ p |[act p-k ,…,newAct p-1 ), P(act ” p |[act p-k ,…,newAct p-1 ),…] distribution, determine whether the subsequence seq containing new behavior logging i is abnormal, thus avoiding re-labeling samples and training the model.
[0069] As Figure 8 shown, in the user behavior anomaly detection method of this embodiment, after step Step304, the following steps are further included:
[0070] Step305: Traverse each set of prefix nodes {[act p ,…,[act p-1 ,…,act p-k ,…,act p-1} of each user behavior act in the updated suffix probability tree, and respectively obtain the suffix probability vector V m =[…, P(act p|prefix m ), …]; where, prefix m represents the prefix subsequence contained in the m-th suffix that contains act p .
[0071] For a certain user operation act p in the sequence, traverse the set of prefix nodes of each user behavior act p in accordance with the depth-first principle.
[0072] Step306: Determine the type of the suffix probability vector V m of the prefix subsequence prefix m ; if the suffix probability vector V m of the prefix subsequence prefix m is a supervised probability vector, then execute Step307; if the suffix probability vector V m of the prefix subsequence prefix m is an unsupervised probability vector, then execute Step308.
[0073] Step307: Determine the magnitude relationship between P(act p |prefix m ) and the preset first threshold P thresh ; if it is less than, then execute Step309; if it is greater than, then execute Step310.
[0074] Step308: Determine the magnitude relationship between P(act p |prefix m ) and the preset second threshold P ’ thresh ; if it is less than, then execute Step311; if it is greater than, then execute Step310.
[0075] Step309: Output the result of subsequence anomaly.
[0076] Step310: Output the result of normal subsequence.
[0077] Step311: Analyze the concentration of the probability distribution of the suffix probability vector V m .
[0078] In the user behavior anomaly detection method of this embodiment, after obtaining the updated suffix probability tree, by traversing the updated suffix probability tree, match the set of user behavior sequences to be measured Seq sThe subsequence of the upper part, and according to whether the corresponding suffix probability vector is supervised or unsupervised, different transition probability thresholds are used to determine whether the subsequence is abnormal, thereby realizing the anomaly detection of the entire user behavior sequence set to be measured, which is not limited by the sequence length and avoids the problem of missing local anomalies detected according to the sequence distance similarity.
[0079] The first threshold P thresh and the second threshold P ’ thresh are of different sizes, and the second threshold P ’ thresh is not zero. For the supervised case, setting the probability threshold P thresh is to avoid the situation of deviation in manual annotation. For example, the subsequence from manual annotation act p-1 to act p is normal, but the value of P(act p |act p-1 ) does not reach the expected P thresh , which still indicates that there is a problem with this path. For the unsupervised case, setting the probability threshold P’ thresh is to judge low-probability subsequences as abnormal subsequences without manual annotation. Therefore, P thresh can be 0, but P’ thresh cannot be 0.
[0080] As Figure 9 shown, in the user behavior anomaly detection method of this embodiment, for business scenarios where user behavior logging is relatively complicated, there are often multiple possibilities for the user's next operation behavior. At this time, for the case where the suffix probability vector V m is an unsupervised probability vector, if P(act p |prefix m ) < P’ thresh , it cannot be determined that the subsequence must be abnormal, and it is necessary to further analyze the concentration of the probability distribution of V m . Therefore, it should be further determined whether the subsequence is abnormal according to the concentration of the probability distribution. Here, the standard deviation Std(V m ) is used to represent the concentration of the probability distribution of V m .
[0081] Therefore, step Step311 includes the following steps:
[0082] Step312: Determine the size relationship between the standard deviation Std(V m ) of the unsupervised probability vector V m and the preset concentration threshold S thresh ; that is, Std(V m ) > Sthresh Whether it holds. If the standard deviation Std(V m ) is greater than the concentration threshold S thresh , then execute Step 313; if the standard deviation Std(V m ) is less than the concentration threshold S thresh , then execute Step 314.
[0083] Step 313: Determine the magnitude relationship between P(act p |prefix m ) and the second threshold P ’ thresh ; that is, whether P(act p |prefix m ) < P ’ thresh holds. If it is less, then execute Step 309; if it is greater, then execute Step 310.
[0084] If Std(V m ) > S thresh (S thresh represents the given concentration threshold), it means that the probability distribution of the unsupervised probability vector V m is concentrated, that is, for a normal user behavior sequence, the next behavior of the prefix subsequence prefix m is known and determined. At this time, if P(act p |prefix m ) < P’ thresh , it indicates that the subsequence is abnormal (see Figure 9 path ①→②→③); otherwise, it is a normal subsequence (see Figure 9 path ①→②→④).
[0085] Step 314: Reverse the sequences in the set of user behavior sequences to be tested Seq s that contain act p to construct an unsupervised reverse suffix probability tree, and then traverse the unsupervised reverse suffix probability tree to search for the node at the end of the reverse suffix probability tree that is act p to obtain the suffix probability vector preV m of the node.
[0086] Step 315: Determine the magnitude relationship between the standard deviation Std(preV m ) of the suffix probability vector preV of the node and the concentration threshold S m ; that is, Std(preV thresh ) > S m ) > S threshIs it true if the suffix probability vector preV of the node m The standard deviation Std(preV m ) is greater than the concentration threshold S thresh , then execute step Step316; if the suffix probability vector preV of the node m The standard deviation Std(preV m ) is less than the concentration threshold S thresh , then execute step Step317.
[0087] Step 316: Determine P(prefix m |act p ) and the second threshold value P ’ thresh The size relationship; that is, P(prefix m |act p )<P ’ thresh Is it true? If it is less than, execute step 309; if it is greater than, execute step 310.
[0088] Step317: The output subsequence is random behavior.
[0089] If Std(V m ) thresh , then it represents the unsupervised probability vector V m The probability distribution of is dispersed, that is, for a normal user behavior sequence, the prefix subsequence prefix m The next behavior of is undefined. Then the prefix subsequence prefix m to act p The subsequences can be divided into the following three scenarios:
[0090] (1) is an abnormal subsequence;
[0091] (2) is a normal subsequence, except that the prefix m to act p The sample size is small;
[0092] (3)act p is a random behavior, that is, act p Can appear anywhere in the user behavior sequence.
[0093] To accurately identify the above three scenarios, this embodiment proposes a discrimination method based on the reverse suffix probability tree. The reverse suffix probability tree is obtained by reversing the sequence containing the new behavior markers and then constructing an unsupervised suffix probability tree. The characteristics of this unsupervised suffix probability tree are as follows: the suffix probability vector of each node represents the prefix probability distribution of the current node in the original behavior sequence. Traverse the reverse suffix probability tree and search for the node whose last position is act p and assume that the suffix probability vector of this node is preV m .
[0094] Then, the steps to identify the above three scenarios are as follows:
[0095] (1) If Std(preV m ) > S thresh , it means that the prefix probability distribution of act p is concentrated. Combining with Std(V m ) < S thresh indicating that the probability distribution of the unsupervised probability vector V m is dispersed, it can be known that the next behavior of prefix m is uncertain, but the previous behavior of act p is known and determined. Therefore, if P(prefix m |act p ) < P’ thresh , it indicates that the subsequence is abnormal (see Figure 9 path ①→⑤→⑥→⑦→⑨); otherwise, it is a normal subsequence (see Figure 9 path ①→⑤→⑥→⑦→④).
[0096] (2) If Std(preV m ) < S thresh , it means that the prefix probability distribution of act p is dispersed. Combining with Std(V m ) < S thresh indicating that the probability distribution of the unsupervised probability vector V m is dispersed, it can be known that it is uncertain whether act p appears after prefix m , and conversely, the previous behavior of act p is also uncertain. This indicates that act p is a random behavior (see Figure 9 path ①→⑤→⑥→⑧). At this time, the model determines the subsequence from prefix m to act p as a normal sequence.
[0097] In step Step311, for subsequences containing new behavior logging, especially complex subsequences with multiple branches, a discrimination method combining probability concentration analysis and reverse suffix probability tree is adopted to judge whether the behavior transfer in the subsequence is abnormal from two dimensions: the distribution of the suffix probability vector and the distribution of the prefix probability vector, reducing a large number of misjudgment problems caused by judging only based on the magnitude of the transfer probability.
[0098] See Figure 10 , Figure 10 which represents the proportion of user abnormal behavior sequences obtained by using a pure supervised method for new behavior logging, and the proportion of user abnormal behavior sequences obtained by using the combined supervised and unsupervised method proposed in this application. As can be seen from the appendix Figure 10 , the results of the two methods are very close, indicating that the abnormal sequence detection result of the technical solution of this application is credible, and it avoids the need to re-label new behavior logging and train the model in the traditional pure supervised manner.
[0099] Compared with Figures 4 to 8 the user behavior anomaly detection method shown, Figure 9 the user behavior anomaly detection method shown proposes a method combining probability concentration analysis and reverse suffix probability tree for scenarios containing new behavior logging and having many path branches, judging whether the current subsequence is abnormal from two directions: suffix and prefix, reducing misjudgment problems caused by relying only on the transfer probability value to judge whether it is abnormal.
[0100] Embodiment 3
[0101] As Figure 11 shown is the user behavior anomaly detection system provided by Embodiment 3 of this application. The system includes a creation module 10, a training module 20, and a detection module 30.
[0102] Among them, the creation module 10 is used to establish a probability distribution model based on the user behavior sequence set; the training module 20 is used to train the probability distribution model based on the user behavior sequence set to obtain an initial probability distribution model; the detection module 30 is used to perform anomaly detection on the to-be-tested user behavior sequence set Seq s and update the subsequences containing new behavior logging in the to-be-tested user behavior sequence set Seq s in an unsupervised manner.
[0103] Through the supervised training of the probability distribution model by the training module 20 of the system, the detection module 30 can perform anomaly detection and unsupervised update on the set of user behavior sequences to be measured through the probability distribution model, so as to improve the adaptability to the anomaly prediction scenario of the subsequence containing new behavior points in the set of user behavior sequences to be measured, and improve the accuracy of behavior anomaly detection. In particular, it is beneficial to improve the accuracy of anomaly detection of complex and multi-branched behavior sequences in the set of user behavior sequences to be measured.
[0104] Embodiment 4
[0105] Embodiment 4 of the present application provides a terminal device, including: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the terminal device executes the method disclosed in Embodiment 2 of the present application.
[0106] The terminal device of this embodiment includes, but is not limited to, products such as websites, application software, and smart phones with the function of marking user operation behaviors, and is particularly suitable for scenarios where software version updates are fast (new functions may be incorporated in different versions) and user behavior marking is complicated.
[0107] Embodiment 5
[0108] Embodiment 5 of the present application provides a computer-readable storage medium, including a program or instruction, when the program or instruction runs on a computer, the method as described in Embodiment 2 of the present application is executed.
[0109] The user behavior anomaly detection method, system and terminal device disclosed in the embodiments of the present application support supervised training and unsupervised update, have both high sequence anomaly prediction accuracy and can adapt to the anomaly prediction scenario of the subsequence containing new behavior points in the set of user behavior sequences to be measured. This solution effectively solves the problem that traditional methods cannot quickly adapt to the need for new behavior points in the business, can greatly reduce the workload of manual re-labeling in the process of model application, and in addition, effectively reduces the false positive rate of the model in detecting abnormal behaviors, and can be applied to scenarios without manual labeling and with many path branches.
[0110] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Video Disc (DVD) with high density), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.
[0111] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0112] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for detecting abnormal user behavior, characterized in that, the method comprises the following steps: establishing a probability distribution model based on a set of user behavior sequences; training the probability distribution model based on the set of user behavior sequences to obtain an initial probability distribution model; Based on the initial probability distribution model, the user behavior sequence set Seq to be tested s Perform anomaly detection on the user behavior sequence set Seq s The subsequence containing the newly added behavior marks is updated in an unsupervised manner; in the step of training the probability distribution model based on the set of user behavior sequences to obtain an initial probability distribution model, the probability distribution model has a structure of a suffix probability tree, and this step comprises: training the probability distribution model with normal behavior sequences in the set of user behavior sequences to obtain an initial suffix probability tree; for the case where the subsequence of the to-be-detected user behavior sequence contains new behavior markers, performing online update on the initial suffix probability tree in an unsupervised manner to obtain an updated suffix probability tree; Among them, the set of user behavior sequences to be measured Seq s ={seq 1 , seq 2 , …, seq n}, and the subsequence seq i containing new behavior logging points = {… newAct p , …, newAct q , …}, where newAct p and newAct q represent new behavior logging points; The user behavior sequence set Seq to be tested based on the initial probability distribution model s Perform anomaly detection on the user behavior sequence set Seq s The steps of updating the subsequence containing the newly added behavior marks in an unsupervised manner include: Detect the suffix of each behavior act in the set of user behavior sequences to be measured p-1 of; If the suffix of the behavior act p-1 is newAct p , traverse all the nodes in the suffix probability tree whose node end is act p-1 , and add the unsupervised suffix probability vector P(newAct p |[act p ,…,act p-k ,…,act p-1 ) to the suffix probability vector of each node, and finally obtain the updated suffix probability tree; If the behavior p-1 NewAct is a new behavior management p-1 , then add the last digit of the suffix probability tree to the newAct p-1 The prefix node and its unsupervised suffix probability vector [P(act' p |[act p-k ,…,newAct p-1 ]),P(act” p |[act p-k ,…,newAct p-1 ]),…], and finally obtain the updated suffix probability tree; after obtaining the updated suffix probability tree, the following steps are included: Traverse each user behavior act in the updated suffix probability tree p of the set of prefix nodes {[act p-1 , …, [act p-k , …, act p-1}, and respectively obtain the suffix probability vector V m = […, P(act p |prefix m ), …]; where, prefix m represents the prefix subsequence of the m-th suffix containing act p ; Determine the suffix probability vector V m of the prefix subsequence prefix m and its type; If the suffix probability vector V m of the prefix subsequence prefix m is a supervised probability vector, then judge the magnitude relationship between P(act p |prefix m ) and the preset first threshold P thresh ; if it is less, output the result of abnormal subsequence; if it is greater, output the result of normal subsequence; If the suffix probability vector V m of the prefix subsequence prefix m is an unsupervised probability vector, then judge the magnitude relationship between P(act p |prefix m ) and the preset second threshold P’ thresh ; if it is less, then analyze the probability distribution concentration of the suffix probability vector V m ; if it is greater, then output the result that the subsequence is normal; The step of analyzing the concentration of the probability distribution of the suffix probability vector V m includes: Determine the standard deviation Std(V m ) of the unsupervised probability vector V m and the magnitude relationship with a preset centrality threshold S thresh ; If the standard deviation Std(V m ) is greater than the centrality threshold S thresh , then determine the magnitude relationship between P(act p |prefix m ) and the second threshold P’ thresh ; if it is less, then output the result of abnormal subsequence; if it is greater, then output the result of normal subsequence; If the standard deviation Std(V m ) is less than the centrality threshold S thresh , then reverse the sequences containing act s in the sequence set Seq p to be tested, construct an unsupervised reverse suffix probability tree, then traverse the unsupervised reverse suffix probability tree, search for the node with act p at the end in the reverse suffix probability tree, and obtain the suffix probability vector preV m of the node; Judge the magnitude relationship between the standard deviation Std(preVm) of the suffix probability vector preV of the node and the centrality threshold S m ; thresh If the standard deviation Std(preV m ) of the suffix probability vector preV of the node m is greater than the centrality threshold S thresh , then judge the magnitude relationship between P(prefix m |act p ) and the second threshold P' thresh ; if it is less than, output the result of abnormal subsequence; if it is greater than, output the result of normal subsequence; If the standard deviation Std(preV m ) of the suffix probability vector preV of the node m is less than the centrality threshold S thresh , then output the result that the subsequence is normal; Among them, the subsequence seq i ={…newAct p ,…,newAct q ,…}, newAct p and newAct q represent new behavior markers; the suffix probability tree includes a plurality of nodes, the nodes include local prefix subsequences of the user behavior sequence and corresponding suffix probability vectors, and the suffix probability vectors include initial supervised suffix probability vectors and unsupervised updated suffix probability vectors.
2. The method for detecting abnormal user behavior according to claim 1, characterized in that, The first threshold P thresh is different in magnitude from the second threshold P' thresh , and the second threshold P' thresh is non-zero.
3. A system for detecting abnormal user behavior, characterized in that, the system comprises: a creation module for establishing a probability distribution model based on a set of user behavior sequences; a training module for training the probability distribution model based on the set of user behavior sequences to obtain an initial probability distribution model; A detection module, which is used to perform anomaly detection on the sequence set Seq of the user behavior to be measured based on the initial probability distribution model. s For the sequence set Seq of the user behavior to be measured s the subsequences containing newly added behavior markers are updated in an unsupervised manner. the training module is further configured to train the probability distribution model with normal behavior sequences in the set of user behavior sequences to obtain an initial suffix probability tree; for the case where the subsequence of the to-be-detected user behavior sequence contains new behavior markers, performing online update on the initial suffix probability tree in an unsupervised manner to obtain an updated suffix probability tree; Among them, the set of user behavior sequences to be measured Seq s ={seq 1 , seq 2 , …, seq n}, the subsequence seq i containing new behavior logging = {… newAct p , …, newAct q , …}, newAct p and newAct q represent new behavior logging, and the probability distribution model has the structure of a suffix probability tree; The detection module is further configured to detect the suffix of each behavior act in the to-be-detected user behavior sequence set p-1 ; If the suffix of the behavior act p-1 is newAct p , then traverse all the nodes in the suffix probability tree whose last digit is act p-1 , and add the unsupervised suffix probability vector P(newAct p |[act p ,…,act p-k ,…,act p-1 ) to the suffix probability vector of each node, and finally obtain the updated suffix probability tree; If the behavior act p-1 is the new behavior marking newAct p-1 , then starting from the root node of the suffix probability tree, prefix nodes with the end being newAct p-1 and their unsupervised suffix probability vectors [P(act’ p |[act p-k ,…,newAct p-1 ), P(act” p |[act p-k ,…,newAct p-1 ),…] are added in sequence, and finally an updated suffix probability tree is obtained; The detection module is further configured to traverse each user behavior act in the updated suffix probability tree p 's set of prefix nodes {[act p-1 ,…,[act p-k ,…,act p-1} to obtain a suffix probability vector V m =[…,P(act p |prefix m ),…]; where prefix m represents the prefix subsequence of the m-th suffix containing act p ; Determine the type of the suffix probability vector V m of the prefix subsequence prefix m ; If the suffix probability vector V m of the prefix subsequence prefix m is a supervised probability vector, then determine the magnitude relationship between P(act p |prefix m ) and a preset first threshold P thresh ; if it is less, then output the result of abnormal subsequence; if it is greater, then output the result of normal subsequence; If the suffix probability vector V m of the prefix subsequence prefix m is an unsupervised probability vector, then determine the magnitude relationship between P(act p |prefix m ) and the preset second threshold P’ thresh ; if it is less, then analyze the probability distribution concentration of the suffix probability vector V m ; if it is greater, then output the result that the subsequence is normal; The detection module is further configured to determine the standard deviation Std(V m ) of the unsupervised probability vector V m and the magnitude relationship with a preset centrality threshold S thresh ; If the standard deviation Std(V m ) is greater than the centrality threshold S thresh , then judge the magnitude relationship between P(act p |prefix m ) and the second threshold P' thresh ; if it is less, output the result of abnormal subsequence; if it is greater, output the result of normal subsequence; If the standard deviation Std(V m ) is less than the centrality threshold S thresh , then reverse the sequences containing act s in the user behavior sequence set Seq p to be tested, construct an unsupervised reverse suffix probability tree, then traverse the unsupervised reverse suffix probability tree, search for the node with act p at the end in the reverse suffix probability tree, and obtain the suffix probability vector preV m of the node; Judge the magnitude relationship between the standard deviation Std(preVm) of the suffix probability vector preV of the node and the centrality threshold S m ; thresh If the standard deviation Std(preV m ) of the suffix probability vector preV of the node m is greater than the centrality threshold S thresh , then judge the magnitude relationship between P(prefix m |act p ) and the second threshold P' thresh ; if it is less than, output the result of abnormal subsequence; if it is greater than, output the result of normal subsequence; If the standard deviation Std(preV m ) of the suffix probability vector preV of the node m is less than the centrality threshold S thresh , then output the result that the subsequence is normal; Among them, the subsequence seq i ={…newAct p ,…,newAct q ,…}, newAct p and newAct q represent new behavior logging; the suffix probability tree includes a plurality of nodes, the nodes include local prefix subsequences of the user behavior sequence and corresponding suffix probability vectors, and the suffix probability vectors include initial supervised suffix probability vectors and unsupervised updated suffix probability vectors.
4. A terminal device, characterized in that, comprises: a memory and a processor: the memory is used for storing a computer program; the processor is configured to execute the computer program stored in the memory so that the terminal device executes the method according to claim 1 or 2.
5. A computer-readable storage medium, characterized in that, comprises a program or an instruction, and when the program or the instruction runs on a computer, the method according to claim 1 or 2 is executed.
Citation Information
Patent Citations
A method for detecting outliers in time series
CN109542952A
User abnormal behavior detection method and system
CN109889538A