Abnormal behavior recognition method and device, equipment and storage medium

Through the encoding, clustering, frequent sub-sequence mining and probability prediction model analysis of user behavior sequences of live broadcast platforms, the behavior patterns of abnormal users in live broadcast platforms are identified, and the problems of low recognition accuracy and high misidentification rate in the prior art are solved, and more efficient abnormal behavior recognition is achieved.

CN119942630APending Publication Date: 2025-05-06GUANGZHOU HUYA INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411670156.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In live broadcast platforms, it is difficult for the prior art to accurately identify abnormal users in gang fraud, resulting in low recognition accuracy and high false recognition rate.

Method used

By obtaining the behavior sequence of live broadcast platform users, coding and clustering, identifying user groups with similar behavior patterns, further conducting frequent sub-sequence mining and probability prediction model analysis to identify whether there are abnormal behaviors.

Benefits of technology

It improves the accuracy of user abnormal behavior recognition, reduces the rate of false recognition, and effectively resists fraud in live broadcast platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942630A_ABST
    Figure CN119942630A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of Internet, and discloses an abnormal behavior recognition method, device and equipment and a storage medium. The abnormal behavior identification method comprises the following steps: acquiring a behavior sequence of each user in a live broadcast platform, encoding the behavior sequence, screening and clustering each user based on the behavior code of each user, and further obtaining a plurality of user groups with similar behavior patterns; mining a frequent subsequence of the behavior sequence of each user group, and estimating a generation probability of the frequent behavior subsequence of each user group through a probability estimation model; and according to the generation probability of the frequent behavior subsequence of each user group, identifying whether the corresponding user group has an abnormal behavior or not. According to the invention, the identification precision of the abnormal behavior of the user is improved, and the error identification rate is reduced, so that fraudulent behaviors in a live broadcast platform are effectively resisted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technology, and in particular to an abnormal behavior recognition method, device, equipment and storage medium. Background Art

[0002] With the rise of live streaming platforms, crowdsourcing gang fraud is becoming more and more common. In this type of fraud, organizers often share cheating strategies through forums and social groups to guide users to participate in fraudulent activities, such as selling accounts after completing registration and real-name authentication, or completing new user acquisition tasks issued by organizers to obtain rewards. The operating media of these users are very scattered, and it is difficult to accurately identify them only by relying on device and IP information aggregation, which poses a challenge to the construction of graph models based on media association. Despite this, the behavioral links of these gang members often show certain similarities and can be identified by behavioral sequence clustering. However, behavioral clustering alone may lead to misidentification or mistaken killing. For example, during the live broadcast of a large-scale event, many normal users may also show similar behavioral aggregation patterns in a short period of time. Summary of the invention

[0003] The main purpose of the present invention is to provide a method, device, equipment and storage medium for identifying abnormal behavior, aiming to solve the technical problem of how to improve the recognition accuracy of user abnormal behavior and reduce the false recognition rate.

[0004] A first aspect of the present invention provides an abnormal behavior recognition method, the abnormal behavior recognition method comprising:

[0005] Obtaining and encoding the behavior sequences of each user in the live broadcast platform to obtain a number of user behavior codes corresponding to each of the users;

[0006] Based on the behavior codes of each user, each user is screened and clustered to obtain several user groups with similar behavior patterns;

[0007] Obtaining user behavior codes of each of the user groups and performing frequent subsequence mining to obtain frequent behavior subsequence codes of the user groups;

[0008] Inputting the frequent behavior subsequence encoding of each user group into a pre-trained probability estimation model, and outputting the generation probability of the frequent behavior subsequence of each user group;

[0009] Based on the generation probability of the frequent behavior subsequences of each of the user groups, it is identified whether the corresponding user group has abnormal behavior.

[0010] Optionally, in a first implementation of the first aspect of the present invention, the acquiring and encoding the behavior sequence of each user in the live broadcast platform to obtain a plurality of user behavior codes corresponding to each of the users includes:

[0011] Collect user behavior events of each user in the live broadcast platform and combine them into behavior sequences;

[0012] Based on the preset identifier, each of the user behavior events in the behavior sequence is mapped and encoded to obtain a plurality of user behavior codes.

[0013] Optionally, in a second implementation of the first aspect of the present invention, the screening and clustering of the users based on the user behavior codes to obtain several user groups with similar behavior patterns include:

[0014] Calculate the behavior sequence weight corresponding to each of the user behavior codes respectively;

[0015] Based on a preset weight threshold and a weight of each behavior sequence, each user behavior code is screened to obtain the screened user behavior code;

[0016] Using a preset similarity measurement algorithm, performing similarity measurement between each of the screened user behavior codes to obtain a similarity measurement result;

[0017] The users having similar behavior patterns in the similarity measurement results are identified to obtain a plurality of user groups.

[0018] Optionally, in a third implementation of the first aspect of the present invention, respectively calculating the behavior sequence weights corresponding to the user behavior codes includes:

[0019] Merging the behavior sequences of all the users to obtain a long event list containing all user behavior events;

[0020] Counting the number of occurrences of each user behavior event in the long event list and normalizing the number of occurrences to obtain the occurrence frequency of each user behavior event;

[0021] Set the highest weight threshold and the lowest weight threshold, and use the preset event weight calculation formula to calculate the weight of each user behavior event based on the frequency of occurrence of each user behavior event;

[0022] Based on the weight of each user behavior event, the behavior sequence weight corresponding to each user behavior code is calculated.

[0023] Optionally, in a fourth implementation of the first aspect of the present invention, the similarity measurement result obtained by using a preset similarity measurement algorithm to perform similarity measurement between each of the filtered user behavior codes includes:

[0024] Selecting two user behavior codes from the screened user behavior codes in sequence;

[0025] Determine whether the two currently selected user behavior codes are completely matched or contain each other;

[0026] If there is a complete match or a containment relationship, the edit distance between the two currently selected user behavior codes is determined to be zero and saved in a preset sparse matrix;

[0027] If it is not a complete match or does not belong to a containment relationship, it is determined whether the length difference between the two currently selected user behavior codes is less than the set edit distance filtering threshold;

[0028] If it is less than the set edit distance filtering threshold, the edit distance between the two currently selected user behavior codes is calculated and saved in the sparse matrix;

[0029] When there is no user behavior code whose edit distance has not been calculated in the screened user behavior codes, the edit distances in the sparse matrix that are lower than a preset edit distance threshold are retained to obtain a distance sparse matrix as a similarity measurement result, wherein each element in the distance sparse matrix represents the distance between two behavior sequences.

[0030] Optionally, in a fifth implementation of the first aspect of the present invention, the identifying the users having similar behavior patterns in the similarity measurement results to obtain a number of user groups includes:

[0031] Initializing an empty first undirected graph according to the dimension of the distance sparse matrix;

[0032] Traversing the distance sparse matrix to add edges to the empty first undirected graph to obtain a second undirected graph, wherein for each non-zero element in the distance sparse matrix that is less than a preset distance threshold, an edge is added between two corresponding nodes in the first undirected graph;

[0033] A preset connected component algorithm is used to identify all connected components in the graph, and a number of user groups corresponding to each connected component are obtained, wherein each connected component includes a group of interconnected nodes, which are used to represent a group of user groups with similar behavior patterns.

[0034] Optionally, in a sixth implementation of the first aspect of the present invention, the probability estimation model adopts a hidden Markov model and is obtained by the following training method:

[0035] Obtaining and encoding the behavior sequences of normal users on the live broadcast platform to obtain training data;

[0036] Based on the training data, a hidden Markov model is constructed, and the hidden Markov model is trained to obtain a trained probability estimation model.

[0037] A second aspect of the present invention provides an abnormal behavior recognition device, the abnormal behavior recognition device comprising:

[0038] The encoding module is used to obtain and encode the behavior sequence of each user in the live broadcast platform to obtain a number of user behavior codes corresponding to each of the users;

[0039] A clustering module, for screening and clustering each of the users based on the behavior codes of each of the users, to obtain a number of user groups with similar behavior patterns;

[0040] A mining module, used to obtain the user behavior code of each user group and perform frequent subsequence mining to obtain the frequent behavior subsequence code of the user group;

[0041] An estimation module, used for encoding the frequent behavior subsequences of each user group into a pre-trained probability estimation model, and outputting the generation probability of the frequent behavior subsequences of each user group;

[0042] The identification module is used to identify whether there is abnormal behavior corresponding to the user group based on the generation probability of the frequent behavior subsequence of each user group.

[0043] Optionally, in a first implementation manner of the second aspect of the present invention, the encoding module is specifically used to:

[0044] Collect user behavior events of each user in the live broadcast platform and combine them into behavior sequences;

[0045] Based on the preset identifier, each of the user behavior events in the behavior sequence is mapped and encoded to obtain a plurality of user behavior codes.

[0046] Optionally, in a second implementation of the second aspect of the present invention, the clustering module includes:

[0047] A calculation unit, used to respectively calculate the behavior sequence weight corresponding to each of the user behavior codes;

[0048] A screening unit, configured to screen each of the user behavior codes based on a preset weight threshold and a weight of each of the behavior sequences to obtain the screened user behavior codes;

[0049] A measurement unit, used to use a preset similarity measurement algorithm to perform similarity measurement between each of the screened user behavior codes to obtain a similarity measurement result;

[0050] The identification unit is used to identify the users with similar behavior patterns in the similarity measurement results to obtain a plurality of user groups.

[0051] Optionally, in a third implementation manner of the second aspect of the present invention, the computing unit is specifically configured to:

[0052] Merging the behavior sequences of all the users to obtain a long event list containing all user behavior events;

[0053] Counting the number of occurrences of each user behavior event in the long event list and normalizing the number of occurrences to obtain the occurrence frequency of each user behavior event;

[0054] Set the highest weight threshold and the lowest weight threshold, and use the preset event weight calculation formula to calculate the weight of each user behavior event based on the frequency of occurrence of each user behavior event;

[0055] Based on the weight of each user behavior event, the behavior sequence weight corresponding to each user behavior code is calculated.

[0056] Optionally, in a fourth implementation manner of the second aspect of the present invention, the measurement unit is specifically used to:

[0057] Selecting two user behavior codes from the screened user behavior codes in sequence;

[0058] Determine whether the two currently selected user behavior codes are completely matched or contain each other;

[0059] If there is a complete match or a containment relationship, the edit distance between the two currently selected user behavior codes is determined to be zero and saved in a preset sparse matrix;

[0060] If it is not a complete match or does not belong to a containment relationship, it is determined whether the length difference between the two currently selected user behavior codes is less than the set edit distance filtering threshold;

[0061] If it is less than the set edit distance filtering threshold, the edit distance between the two currently selected user behavior codes is calculated and saved in the sparse matrix;

[0062] When there is no user behavior code whose edit distance has not been calculated in the screened user behavior codes, the edit distances in the sparse matrix that are lower than a preset edit distance threshold are retained to obtain a distance sparse matrix as a similarity measurement result, wherein each element in the distance sparse matrix represents the distance between two behavior sequences.

[0063] Optionally, in a fifth implementation of the second aspect of the present invention, the identification unit is specifically configured to:

[0064] Initializing an empty first undirected graph according to the dimension of the distance sparse matrix;

[0065] Traversing the distance sparse matrix to add edges to the empty first undirected graph to obtain a second undirected graph, wherein for each non-zero element in the distance sparse matrix that is less than a preset distance threshold, an edge is added between two corresponding nodes in the first undirected graph;

[0066] A preset connected component algorithm is used to identify all connected components in the graph, and a number of user groups corresponding to each connected component are obtained, wherein each connected component includes a group of interconnected nodes, which are used to represent a group of user groups with similar behavior patterns.

[0067] Optionally, in a sixth implementation of the second aspect of the present invention, the probability estimation model adopts a hidden Markov model, and the abnormal behavior identification device further includes:

[0068] The training module is used to obtain and encode the behavior sequence of normal users in the live broadcast platform to obtain training data; based on the training data, a hidden Markov model is constructed, and the hidden Markov model is trained to obtain a trained probability estimation model.

[0069] A third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned abnormal behavior identification method.

[0070] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned abnormal behavior identification method.

[0071] In the technical solution provided by the present invention, the behavior sequences of each user in the live broadcast platform are first collected, and then the user behavior sequences are measured for similarity, and the user behavior sequences are clustered according to the results of the similarity measurement to identify user groups with similar behavior patterns; then the frequent behavior subsequences of each user group are mined; finally, based on the trained probability prediction model, the generation probability of the frequent behavior subsequences of each user group is calculated. If the frequent behavior subsequences of a certain user group have a high degree of aggregation but an abnormally low generation probability, it is considered that the behavior pattern has potential risks, and the corresponding group is also very likely to be a risk group, thereby realizing the identification of abnormal user behaviors in the live broadcast platform. The present invention improves the recognition accuracy of abnormal user behaviors and also reduces the false recognition rate, thereby effectively resisting fraud in the live broadcast platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 A schematic diagram of an embodiment of a method for identifying abnormal behavior in an embodiment of the present invention;

[0073] Figure 2 A schematic diagram of an abnormal behavior identification device according to an embodiment of the present invention;

[0074] Figure 3 FIG. 1 is a schematic diagram of an embodiment of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0075] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0076] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , an embodiment of the abnormal behavior identification method in the embodiment of the present invention includes:

[0077] 101. Obtain and encode the behavior sequence of each user in the live broadcast platform to obtain a number of user behavior codes corresponding to each of the users;

[0078] In this embodiment, the behavior of each user in the live broadcast platform is called a user behavior event, such as registration, login, binding a mobile phone number, entering a personal center, screenshots, etc. All operations of the same user in the live broadcast platform are combined in order to obtain the behavior sequence of the user. In order to facilitate subsequent data processing, the behavior sequence of each user is further encoded to obtain the user behavior code corresponding to each user.

[0079] In an optional embodiment, the above step 101 further includes:

[0080] 1011. Collect user behavior events of each user in the live broadcast platform and combine them into a behavior sequence;

[0081] 1012. Based on a preset identifier, map and encode each of the user behavior events in the behavior sequence to obtain a plurality of user behavior codes.

[0082] The live broadcast platform will record various behaviors of users on its platform in real time, such as entering the live broadcast room, liking, commenting, purchasing products, sharing, etc. These behaviors are called user behavior events. All behavior events of each user within a period of time (such as a day, a week, or a month) are arranged in chronological order to form a behavior sequence. This behavior sequence reflects the activity trajectory and preferences of users on the live broadcast platform.

[0083] Preset identifiers are a set of predefined rules or mapping tables used to convert user behavior events into unique codes or identifiers. These identifiers can be numbers, letters, or a combination of them, and are used to uniquely identify each behavior event in subsequent analysis and processing. Based on the preset identifiers, each user behavior event in the behavior sequence is mapped to the corresponding code. Each user behavior event is converted into a set of processable sequences of numbers or letters, namely user behavior codes. After mapping and coding, each user's behavior sequence is converted into a series of user behavior codes. These codes can be used for subsequent identification of abnormal user behavior.

[0084] Assuming that the user behavior events collected by user A on the live broadcast platform are: registration, login, change mobile phone number, enter personal center, click my task, screenshot, then map each user behavior event to a unique identifier ID. For example, map the "registration" behavior to 1, the "login" behavior to 2, the "change mobile phone number" behavior to 5, the "enter personal center" behavior to 7, the "click my task" behavior to 8, and the "screenshot" behavior to 12. After the encoding is completed, the user behavior code corresponding to user A is [1, 2, 5, 7, 8, 12].

[0085] By encoding each user behavior event, a data basis can be provided for subsequent processing, which helps save memory and improve operating efficiency. In order to finely distinguish user behaviors, the attributes of each user behavior event can be further spliced, such as 'registration'->'registration_verification code registration', 'gift'->'gift_anchor id', to refine the granularity and expand the dimension of the behavior space.

[0086] 102. Based on the behavior codes of the users, screen and cluster the users to obtain a number of user groups with similar behavior patterns;

[0087] In this embodiment, considering that most user behaviors in the live broadcast platform are normal, in order to reduce the amount of calculation and improve the accuracy of identifying abnormal user behaviors, it is necessary to screen the users, and then cluster the screened users and their behavior sequences to obtain a user group composed of users with similar behavior patterns.

[0088] In an optional embodiment, the above step 102 further includes:

[0089] 1021. Calculate the behavior sequence weight corresponding to each of the user behavior codes respectively;

[0090] In this embodiment, in order to more carefully capture the importance of events in the behavior sequence, a weight allocation method based on the frequency of event occurrence is adopted. The core of this method is to inversely allocate weights according to the frequency of event occurrence, that is, the less an event occurs, the greater its weight in cluster analysis, thereby highlighting uncommon but potentially important behaviors, providing a refined basis for subsequent behavior pattern analysis.

[0091] In one embodiment, the behavior sequence weight is calculated in the following manner:

[0092] (1) merging the behavior sequences of all the users to obtain a long event list containing all user behavior events;

[0093] (2) Counting the number of occurrences of each user behavior event in the long event list and normalizing the number of occurrences to obtain the occurrence frequency of each user behavior event;

[0094] (3) Setting a maximum weight threshold and a minimum weight threshold, and using a preset event weight calculation formula to calculate the weight of each user behavior event based on the frequency of occurrence of each user behavior event;

[0095] (4) Based on the weight of each user behavior event, the behavior sequence weight corresponding to each user behavior code is calculated.

[0096] In this embodiment, the user behavior event data is first summarized, and the behavior sequences of all users are merged to form a long list of all events for frequency calculation; then the frequency of occurrence of each event is calculated, specifically: the number of occurrences of each event is counted and normalized to calculate the frequency of occurrence of each event. Events with high frequency mean common behaviors, and relatively few can indicate abnormal or important behavior patterns. Finally, the highest and lowest weight thresholds are set (for example, from 1 to 10), and the weight of each event is calculated according to the frequency of occurrence of the event. The specific formula is:

[0097]

[0098] Among them, W e represents the weight of user behavior event e, f(e) represents the frequency of user behavior event e, freq high Indicates the highest event frequency among all user behavior events, freq low represents the lowest event frequency among all user behavior events, W high and W lowIndicates the maximum and minimum values ​​of the event weight (both are thresholds), which are custom hyperparameters.

[0099] The above formula for calculating the weight of user behavior events can ensure that the event with the lowest frequency obtains the highest weight value. Through this weight distribution mechanism, the importance of behavior events can be more accurately considered in clustering and subsequent user behavior pattern recognition, thereby enhancing the sensitivity and accuracy of the model for abnormal behavior recognition.

[0100] In this embodiment, after calculating the weight of each user behavior event, it is necessary to further calculate the behavior sequence weight corresponding to each user behavior code.

[0101] There are many ways to calculate the weights of user behavior sequences based on event weights. This embodiment uses two methods: direct summation of event weights in the behavior sequence and averaging. Direct summation focuses on users with longer behavior sequences (for example, a behavior sequence length greater than 20), while averaging focuses on uncommon user behavior events in the behavior sequence.

[0102]

[0103] in, Indicates the weight of the behavior sequence corresponding to the first type of user behavior coding. The length of the behavior sequence of the first type of user behavior coding is greater than the set length threshold. represents the weight of the behavior sequence corresponding to the second type of user behavior coding. There are uncommon user behavior events in the behavior sequence corresponding to the second type of user behavior coding. ei represents the weight of the i-th user behavior event, and n represents the number of user behavior events in the behavior sequence corresponding to the user behavior code.

[0104] 1022. Filter each of the user behavior codes based on a preset weight threshold and a weight of each of the behavior sequences to obtain the filtered user behavior codes;

[0105] In this embodiment, each user behavior code is screened according to the behavior sequence weight corresponding to each user behavior code calculated in the above step 1021, and a portion of users and their behavior sequences with weights lower than the threshold are pre-screened through a threshold, and the user behavior codes corresponding to the remaining users and their behavior sequences enter the next step of calculation.

[0106] 1023. Using a preset similarity measurement algorithm, perform similarity measurement between each of the screened user behavior codes to obtain a similarity measurement result;

[0107] In this embodiment, in order to find users with similar behavior patterns from the behavior sequences of each user, a preset similarity measurement algorithm is used to measure the similarity between the screened user behavior codes, preferably by calculating the edit distance to measure the similarity between different user behavior sequences.

[0108] In an optional embodiment, the above step 1023 further includes:

[0109] (1) selecting two user behavior codes from the screened user behavior codes in sequence;

[0110] (2) Determine whether the two currently selected user behavior codes completely match or are in a containment relationship;

[0111] (3) If there is a complete match or a containment relationship, the edit distance between the two currently selected user behavior codes is determined to be zero and saved in a preset sparse matrix;

[0112] (4) If it is not a complete match or does not belong to a containment relationship, determine whether the length difference between the two currently selected user behavior codes is less than the set edit distance filtering threshold;

[0113] In this embodiment, in order to improve the calculation speed and reduce the amount of calculation, before calculating the edit distance between different user behavior sequences, a preliminary judgment is first made on all user behavior sequences (that is, user behavior codes). Assuming any two character strings (that is, user behavior codes) A ​​and B, where the length of A is m and the length of B is n, the specific contents of the preliminary judgment include:

[0114] The first is to determine whether there is a complete match: if A and B are equal, it means that the two strings are exactly the same and the edit distance is 0;

[0115] The second is to determine whether it is an inclusion relationship: if A is a substring of B or B is a substring of A, it means that one string completely contains the other, so its edit distance can be considered to be 0. Although the actual edit distance may involve deleting unmatched parts, in actual application scenarios, the goal is to find the common continuous behavior sequence substring of the gang and make risk assessments on the substrings. Therefore, the subsequence can be defined with an edit distance of 0.

[0116] The third is to determine whether the length difference is significant: if the length difference between A and B exceeds the set edit distance filter threshold, they will not be included in the edit distance calculation. This can avoid unnecessary edit distance calculations for strings that are too long or too short. That is, this embodiment only calculates the edit distance for user behavior sequences with insignificant length differences.

[0117] (5) If the edit distance is less than the set edit distance filtering threshold, the edit distance between the two currently selected user behavior codes is calculated and saved in the sparse matrix;

[0118] (6) When there is no user behavior code whose edit distance has not been calculated in the screened user behavior codes, the edit distances in the sparse matrix that are lower than a preset edit distance threshold are retained to obtain a distance sparse matrix as a similarity measurement result, wherein each element in the distance sparse matrix represents the distance between two behavior sequences.

[0119] In this embodiment, the following method is specifically used to calculate the edit distance between two different user behavior codes:

[0120] First, assume that there are two arbitrary strings (that is, user behavior codes, which can be regarded as strings) A and B, where the length of A is m and the length of B is n. Initialize the dynamic programming matrix: set the element D[0][0] of the matrix D to 0, that is, the conversion from an empty string to an empty string does not require any operation; initialize the first row D[0][j] of the matrix to j, indicating that the empty string is converted to a string B of length j through an insertion operation; initialize the first column D[i][0] of the matrix to i, indicating that a string A of length i is converted to an empty string through a deletion operation;

[0121] Next, fill in the dynamic programming matrix and calculate each element D[i][j] of the matrix:

[0122] (a) Character matching: If the i-th character of A is the same as the j-th character of B, that is, A[i-1] == B[j-1], then D[i][j] inherits the value of D[i-1][j-1].

[0123] (b) Character mismatch: If the i-th character of A is not the same as the j-th character of B, then D[i][j] takes the minimum value among D[i-1][j]+1, D[i][j-1]+1, and D[i-1][j-1]+1.

[0124] The calculation formula of D[i][j] is as follows:

[0125]

[0126] The lower right corner element D[m][n] of the matrix that is finally filled is the minimum number of edit operations required to convert string A to string B, that is, the edit distance between the two strings. It should be noted that each user behavior code needs to calculate the edit distance with each other user behavior code. That is, assuming the number of user behavior codes is S, S*(S-1) / 2 matrices need to be filled.

[0127] This embodiment preferably uses a sparse matrix to store the edit distances between different users. In addition, this embodiment also sets a reasonable edit distance threshold, retains the edit distances below this threshold, and obtains a distance sparse matrix, thereby filtering out weakly connected behavior sequences and focusing on closely related user groups.

[0128] 1024. Identify the users with similar behavior patterns in the similarity measurement results to obtain a number of user groups.

[0129] In this embodiment, after obtaining the similarity measurement results between each user behavior code pairwise, that is, the edit distance between each user behavior code pairwise, through the above step 1023, these behavior sequences are clustered in combination with the connected component algorithm of graph theory to identify user groups with similar behavior patterns. The connected component algorithm of graph theory is applied to process the distance sparse matrix, and each behavior sequence is regarded as a node in the graph. For a distance less than a preset distance threshold, it indicates that there is a strong correlation between the sequences.

[0130] In an optional embodiment, the above step 1024 further includes:

[0131] (1) Initializing an empty first undirected graph according to the dimension of the distance sparse matrix;

[0132] (2) traversing the distance sparse matrix to add edges to the empty first undirected graph to obtain a second undirected graph, wherein for each non-zero element in the distance sparse matrix that is less than a preset distance threshold, an edge is added between two corresponding nodes in the first undirected graph;

[0133] (3) A preset connected component algorithm is used to identify all connected components in the graph, and a number of user groups corresponding to each connected component are obtained, where each connected component contains a group of interconnected nodes, which are used to represent a group of user groups with similar behavior patterns.

[0134] In this embodiment, an empty undirected graph containing n nodes is first initialized according to the dimension of the distance sparse matrix (assuming it is n×n), and then the distance sparse matrix is ​​traversed, and for each non-zero element in the matrix that is less than a preset distance threshold, an edge is added between two corresponding nodes, and finally a preset connected component algorithm (such as depth-first search (DFS) or breadth-first search (BFS)) is used to identify all connected components in the graph. Each connected component contains a group of interconnected nodes, which represent a group of user groups with similar behavior patterns (including multiple different users, i.e., user groups).

[0135] 103. Obtain user behavior codes of each of the user groups and perform frequent subsequence mining to obtain frequent behavior subsequence codes of the user groups;

[0136] In this embodiment, after obtaining user groups with similar behavior patterns in step 102, it is necessary to further mine frequent behavior subsequences of each user group. Frequent behavior subsequences refer to behavior sequences with a time order in a set of user behavior data that appear more than a preset threshold.

[0137] This embodiment preferably mines the frequent subsequences of each user group through the FP-Growth library. FP-Growth (Frequent Pattern Growth) is an efficient frequent item set mining algorithm suitable for processing large-scale data sets. The FP-Growth algorithm stores frequent item sets by constructing a compact data structure called FP-Tree (Frequent Pattern Tree). The specific implementation method is as follows:

[0138] First, the user behavior codes of each user group are obtained and combined into a behavior sequence database, where each sequence is a user behavior event. Then, FP-Tree is constructed to store frequent itemsets and their occurrence times, which includes: traversing the data sets corresponding to each user group in the sequence database, calculating the support of each user behavior (i.e., the number of times the user behavior appears in the data set). According to the set minimum support threshold, frequent itemsets are screened out. These itemsets will appear as nodes in the subsequent FP-Tree construction; then an empty tree is created as the root node of the FP-Tree, traversing each user behavior in the data set, and inserting the user behavior into the FP-Tree in the order of frequent itemsets. If the same path already exists in the tree, the count of the path is increased; otherwise, a new path is created. At the same time, a header table (HeaderTable) is maintained for fast access to the pointer list of the same items in the FP-Tree. Finally, frequent itemsets (or subsequences) are mined by traversing the FP-Tree, and the mined frequent itemsets (or subsequences) are the frequent behavior subsequence codes of the user group. Specifically, it includes: starting from the bottom item of the header table, mining the conditional pattern base of the header table. The conditional pattern base is a set of paths ending with the element item being searched, and each path is associated with a count value. For each frequent item, a conditional FP tree is constructed based on its conditional pattern base. The conditional FP tree is a subtree of the original FP-Tree, which only contains paths and counts related to the current frequent item. Recursively mine the frequent item sets in the conditional FP tree until no frequent item sets are found. During the recursive process, the frequent item set list is updated each time and a new conditional FP tree is constructed.

[0139] 104. Input the frequent behavior subsequence codes of each user group into a pre-trained probability estimation model, and output the generation probability of the frequent behavior subsequences of each user group;

[0140] 105. Based on the generation probability of the frequent behavior subsequences of each of the user groups, identify whether the corresponding user group has abnormal behavior.

[0141] In this embodiment, the frequent behavior subsequences of each user group mined in the above step 103 include both the frequent behavior subsequences of normal users and the frequent behavior subsequences of abnormal users. Therefore, in order to accurately identify the frequent behavior subsequences of abnormal users and further identify abnormal behaviors, this embodiment trains a probability estimation model based on the behavior sequence of normal users to calculate the generation probability of the frequent behavior subsequences of each user group. If a behavior sequence has a high degree of aggregation but an abnormally low generation probability, it is considered that the behavior pattern has potential risks and the corresponding group is also likely to be a risk group.

[0142] In an optional embodiment, the probability estimation model preferably adopts a hidden Markov model and is obtained by the following training method:

[0143] (1) obtaining and encoding the behavior sequences of normal users on the live broadcast platform to obtain training data;

[0144] (2) constructing a state set, an observation set and model parameters of a hidden Markov model based on the training data, wherein the model parameters include a state transition probability matrix, an observation probability matrix and an initial state probability vector;

[0145] (3) initializing model parameters of the hidden Markov model based on the state set and the observation set;

[0146] (4) Estimate the model parameters of the hidden Markov model using the maximum likelihood estimation method, recalculate the probability of the observation sequence based on the estimated model parameters, and iteratively update the model parameters until the model parameters converge or reach a predetermined number of iterations, thereby obtaining a trained probability estimation model.

[0147] Collect normal user behavior sequences, encode behavior events (for example, map "registration" behavior to 1, "login" behavior to 2, and so on), and preprocess normal user behavior sequences.

[0148] Use the preprocessed data to train the Hidden Markov Model (HMM) and construct the state set, observation set and model parameters of the Hidden Markov Model:

[0149] State set S = {s1, s2, ..., s n}: represents all possible hidden states in the model, and n is the number of hidden states. In actual business scenarios, it refers to different behaviors performed by users on the live broadcast platform, such as registration, login, changing passwords, watching live broadcasts, and giving gifts.

[0150] The observation set O = {o1,o2,...,o m}: represents all possible observations in the model. In actual business scenarios, it refers to the specific data generated by user behavior, such as registration_mobile number, gift_anchor, login_Guangzhou, etc.

[0151] Model parameter 1: state transition probability matrix A = [a ij ] n*n , indicating that it is in state s at time t i The probability of transitioning to state sj at time t+1, a ij =P(q t+1 =s j |q t =s i );

[0152] Model parameter 2: Observation distribution probability matrix B = [b j (k)] n*n , indicating that it is in state s at time t j The probability of generating an observation under the condition of j (k) = P(o t =k|q t =s j );

[0153] Model parameter 3: Initial state probability vector Π = (Π i ), Π i =P(q1=s i ), indicating that time t = 1 is in state s i This embodiment uses a multivariate mixed Gaussian distribution as the probability density function of the observation probability.

[0154] This embodiment preferably uses the maximum likelihood estimation method to estimate the model parameters of the hidden Markov model and update them until the parameters converge or reach a predetermined number of iterations to obtain a trained probability prediction model, and then applies the trained probability prediction model to the user groups obtained by clustering and the frequent subsequences corresponding to each user group, thereby screening out the groups corresponding to the normal behavior sequences, and outputting the user groups corresponding to the behavior sequences with probabilities less than a threshold. The output user groups are the abnormal groups with abnormal behaviors.

[0155] The Hidden Markov Model is a statistical model that assumes that the state of the system is hidden, but each state generates some observations. The transition probabilities between states and the generation probabilities of each state to observations can be learned. By training the Hidden Markov Model on the behavior sequences of normal users, the transition probabilities between various behaviors can be learned. This enables the model to perform probability scoring on new observation sequences, thereby accurately identifying normal and abnormal behavior sequences, and successfully separating the behavior patterns of black market groups from those of normal user groups.

[0156] In this embodiment, the behavior sequences of each user in the live broadcast platform are first collected, and then the user behavior sequences are measured for similarity, and the user behavior sequences are clustered according to the results of the similarity measurement to identify user groups with similar behavior patterns; then the frequent behavior subsequences of each user group are mined; finally, based on the trained probability prediction model, the generation probability of the frequent behavior subsequences of each user group is calculated. If the frequent behavior subsequences of a certain user group have a high degree of aggregation but an abnormally low generation probability, it is considered that the behavior pattern has potential risks, and the corresponding group is also likely to be a risk group, thereby realizing the identification of abnormal user behavior in the live broadcast platform. This embodiment improves the recognition accuracy of abnormal user behavior and also reduces the false recognition rate, thereby effectively resisting fraud in the live broadcast platform.

[0157] The above describes the abnormal behavior recognition method in the embodiment of the present invention. The following describes the abnormal behavior recognition device in the embodiment of the present invention. Figure 2 , an abnormal behavior identification device in an embodiment of the present invention includes:

[0158] The encoding module 201 is used to obtain and encode the behavior sequence of each user in the live broadcast platform to obtain a number of user behavior codes corresponding to each of the users;

[0159] A clustering module 202 is used to screen and cluster the users based on the user behavior codes to obtain a number of user groups with similar behavior patterns;

[0160] A mining module 203 is used to obtain the user behavior code of each user group and perform frequent subsequence mining to obtain the frequent behavior subsequence code of the user group;

[0161] The estimation module 204 is used to input the frequent behavior subsequence encoding of each user group into a pre-trained probability estimation model, and output the generation probability of the frequent behavior subsequence of each user group;

[0162] The identification module 205 is used to identify whether there is abnormal behavior in the corresponding user group based on the generation probability of the frequent behavior subsequence of each user group.

[0163] Optionally, in one embodiment, the encoding module 201 is specifically used for:

[0164] Collect user behavior events of each user in the live broadcast platform and combine them into behavior sequences;

[0165] Based on the preset identifier, each of the user behavior events in the behavior sequence is mapped and encoded to obtain a plurality of user behavior codes.

[0166] Optionally, in one embodiment, the clustering module 202 includes:

[0167] A calculation unit 2021, used to calculate the behavior sequence weight corresponding to each of the user behavior codes;

[0168] A screening unit 2022, configured to screen each of the user behavior codes based on a preset weight threshold and a weight of each of the behavior sequences to obtain the screened user behavior codes;

[0169] The measuring unit 2023 is used to use a preset similarity measurement algorithm to perform similarity measurement between each of the screened user behavior codes to obtain a similarity measurement result;

[0170] The identification unit 2024 is used to identify the users with similar behavior patterns in the similarity measurement results to obtain a plurality of user groups.

[0171] Optionally, in one embodiment, the calculation unit 2021 is specifically used for:

[0172] Merging the behavior sequences of all the users to obtain a long event list containing all user behavior events;

[0173] Counting the number of occurrences of each user behavior event in the long event list and normalizing the number of occurrences to obtain the occurrence frequency of each user behavior event;

[0174] Set the highest weight threshold and the lowest weight threshold, and use the preset event weight calculation formula to calculate the weight of each user behavior event based on the frequency of occurrence of each user behavior event;

[0175] Based on the weight of each user behavior event, the behavior sequence weight corresponding to each user behavior code is calculated.

[0176] Optionally, in one embodiment, the measurement unit 2023 is specifically used for:

[0177] Selecting two user behavior codes from the screened user behavior codes in sequence;

[0178] Determine whether the two currently selected user behavior codes are completely matched or contain each other;

[0179] If there is a complete match or a containment relationship, the edit distance between the two currently selected user behavior codes is determined to be zero and saved in a preset sparse matrix;

[0180] If it is not a complete match or does not belong to a containment relationship, it is determined whether the length difference between the two currently selected user behavior codes is less than the set edit distance filtering threshold;

[0181] If it is less than the set edit distance filtering threshold, the edit distance between the two currently selected user behavior codes is calculated and saved in the sparse matrix;

[0182] When there is no user behavior code whose edit distance has not been calculated in the screened user behavior codes, the edit distances in the sparse matrix that are lower than a preset edit distance threshold are retained to obtain a distance sparse matrix as a similarity measurement result, wherein each element in the distance sparse matrix represents the distance between two behavior sequences.

[0183] Optionally, in one embodiment, the identification unit 2024 is specifically used for:

[0184] Initializing an empty first undirected graph according to the dimension of the distance sparse matrix;

[0185] Traversing the distance sparse matrix to add edges to the empty first undirected graph to obtain a second undirected graph, wherein for each non-zero element in the distance sparse matrix that is less than a preset distance threshold, an edge is added between two corresponding nodes in the first undirected graph;

[0186] A preset connected component algorithm is used to identify all connected components in the graph, and a number of user groups corresponding to each connected component are obtained, wherein each connected component includes a group of interconnected nodes, which are used to represent a group of user groups with similar behavior patterns.

[0187] Optionally, in one embodiment, the probability estimation model adopts a hidden Markov model, and the abnormal behavior identification device further includes:

[0188] The training module 206 is used to obtain and encode the behavior sequence of normal users in the live broadcast platform to obtain training data; based on the training data, a hidden Markov model is constructed, and the hidden Markov model is trained to obtain a trained probability estimation model.

[0189] Since the embodiments of the device part correspond to the embodiments of the above-mentioned method, please refer to the above-mentioned method embodiments for the introduction of the abnormal behavior recognition device provided by the present invention. The present invention will not be repeated here, and it has the same beneficial effects as the above-mentioned abnormal behavior recognition method.

[0190] above Figure 2 The abnormal behavior identification device in the embodiment of the present invention is described in detail from the perspective of modular functional entities, and the computer device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0191] Figure 3 1 is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. The computer device 500 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 510 (for example, one or more processors) and a memory 520, and one or more storage media 530 (for example, one or more mass storage devices) storing application programs 533 or data 532. Among them, the memory 520 and the storage medium 530 can be short-term storage or permanent storage. The program stored in the storage medium 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device 500. Furthermore, the processor 510 may be configured to communicate with the storage medium 530 to execute a series of instruction operations in the storage medium 530 on the computer device 500.

[0192] The computer device 500 may also include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input and output interfaces 560, and / or one or more operating systems 531, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 3 The illustrated computer device structure does not constitute a limitation on the computer device, and may include more or fewer components than illustrated, or combine certain components, or arrange the components differently.

[0193] The present invention also provides a computer device, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the abnormal behavior identification method in the above-mentioned embodiments.

[0194] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the abnormal behavior identification method.

[0195] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0196] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0197] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying abnormal behavior, characterized in that: include: Obtaining and encoding the behavior sequences of each user in the live broadcast platform to obtain a number of user behavior codes corresponding to each of the users; Based on the behavior codes of each user, each user is screened and clustered to obtain several user groups with similar behavior patterns; Obtaining user behavior codes of each of the user groups and performing frequent subsequence mining to obtain frequent behavior subsequence codes of the user groups; Inputting the frequent behavior subsequence encoding of each user group into a pre-trained probability estimation model, and outputting the generation probability of the frequent behavior subsequence of each user group; Based on the generation probability of the frequent behavior subsequences of each of the user groups, it is identified whether the corresponding user group has abnormal behavior.

2. The abnormal behavior identification method according to claim 1, characterized in that: The obtaining and encoding of the behavior sequence of each user in the live broadcast platform to obtain a number of user behavior codes corresponding to each of the users includes: Collect user behavior events of each user in the live broadcast platform and combine them into behavior sequences; Based on the preset identifier, each of the user behavior events in the behavior sequence is mapped and encoded to obtain a plurality of user behavior codes.

3. The abnormal behavior identification method according to claim 1 or 2, characterized in that: Based on the user behavior codes, the users are screened and clustered to obtain several user groups with similar behavior patterns, including: Calculate the behavior sequence weight corresponding to each of the user behavior codes respectively; Based on a preset weight threshold and a weight of each behavior sequence, each user behavior code is screened to obtain the screened user behavior code; Using a preset similarity measurement algorithm, performing similarity measurement between each of the screened user behavior codes to obtain a similarity measurement result; The users having similar behavior patterns in the similarity measurement results are identified to obtain a plurality of user groups.

4. The abnormal behavior identification method according to claim 3, characterized in that: The step of respectively calculating the behavior sequence weights corresponding to the user behavior codes comprises: Merging the behavior sequences of all the users to obtain a long event list containing all user behavior events; Counting the number of occurrences of each user behavior event in the long event list and normalizing the number of occurrences to obtain the occurrence frequency of each user behavior event; Set the highest weight threshold and the lowest weight threshold, and use the preset event weight calculation formula to calculate the weight of each user behavior event based on the frequency of occurrence of each user behavior event; Based on the weight of each user behavior event, the behavior sequence weight corresponding to each user behavior code is calculated.

5. The abnormal behavior identification method according to claim 3, characterized in that: The preset similarity measurement algorithm is used to perform similarity measurement between each of the screened user behavior codes, and the similarity measurement results obtained include: Selecting two user behavior codes from the screened user behavior codes in sequence; Determine whether the two currently selected user behavior codes are completely matched or contain each other; If there is a complete match or a containment relationship, the edit distance between the two currently selected user behavior codes is determined to be zero and saved in a preset sparse matrix; If it is not a complete match or does not belong to a containment relationship, it is determined whether the length difference between the two currently selected user behavior codes is less than the set edit distance filtering threshold; If it is less than the set edit distance filtering threshold, the edit distance between the two currently selected user behavior codes is calculated and saved in the sparse matrix; When there is no user behavior code whose edit distance has not been calculated in the screened user behavior codes, the edit distances in the sparse matrix that are lower than a preset edit distance threshold are retained to obtain a distance sparse matrix as a similarity measurement result, wherein each element in the distance sparse matrix represents the distance between two behavior sequences.

6. The abnormal behavior identification method according to claim 5, characterized in that: The identifying the users with similar behavior patterns in the similarity measurement results to obtain a number of user groups includes: Initializing an empty first undirected graph according to the dimension of the distance sparse matrix; Traversing the distance sparse matrix to add edges to the empty first undirected graph to obtain a second undirected graph, wherein for each non-zero element in the distance sparse matrix that is less than a preset distance threshold, an edge is added between two corresponding nodes in the first undirected graph; A preset connected component algorithm is used to identify all connected components in the graph, and a number of user groups corresponding to each connected component are obtained, wherein each connected component includes a group of interconnected nodes, which are used to represent a group of user groups with similar behavior patterns.

7. The abnormal behavior identification method according to claim 1, characterized in that: The probability estimation model adopts a hidden Markov model and is obtained through the following training method: Obtaining and encoding the behavior sequences of normal users on the live broadcast platform to obtain training data; Based on the training data, a hidden Markov model is constructed, and the hidden Markov model is trained to obtain a trained probability estimation model.

8. An abnormal behavior recognition device, characterized in that: The abnormal behavior identification device comprises: The encoding module is used to obtain and encode the behavior sequence of each user in the live broadcast platform to obtain a number of user behavior codes corresponding to each of the users; A clustering module, for screening and clustering each of the users based on the behavior codes of each of the users, to obtain a number of user groups with similar behavior patterns; A mining module, used to obtain the user behavior code of each user group and perform frequent subsequence mining to obtain the frequent behavior subsequence code of the user group; An estimation module, used to input the frequent behavior subsequence encoding of each user group into a pre-trained probability estimation model, and output the generation probability of the frequent behavior subsequence of each user group; The identification module is used to identify whether there is abnormal behavior in the corresponding user group based on the generation probability of the frequent behavior subsequence of each user group.

9. A computer device, characterized in that: The computer device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the computer device to execute the abnormal behavior identification method according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the abnormal behavior identification method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Website information auditing control method and system

    CN121210772A

  • A method and system for website information review and control

    CN121210772B