Anomaly object recognition method, device, medium, apparatus, and program product

By analyzing user behavior data within a preset time window, constructing a set of similar object behaviors and calculating similarity, abnormal objects are identified. This solves the problems of high threshold sensitivity and delayed diffusion in existing technologies, and enables efficient identification and mining of fraud gangs.

CN116257562BActive Publication Date: 2025-11-21TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111500740.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-11-21
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

Existing technologies for combating fraudulent activities suffer from high threshold sensitivity, making it easy for fraudsters to bypass their strategies. Furthermore, they suffer from delayed dissemination and low coverage, making it difficult to effectively uncover entire criminal groups.

Method used

By analyzing user behavior data within a preset time window, a set of similar object behaviors is constructed. Multiple types of behavioral feature similarity algorithms are used to calculate object similarity. Graph processing is applied to identify abnormal objects. A sliding time window is combined for real-time tracking and gang detection.

Benefits of technology

It improved the coverage of fraudulent user identification, enabled the discovery of abnormal objects in the early stages of the fraud lifecycle, alleviated the problems of high threshold sensitivity and delayed strikes in the diffusion system, and achieved effective detection of fraud gangs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116257562B_ABST
    Figure CN116257562B_ABST
Patent Text Reader

Abstract

The application discloses an abnormal object identification method, device, medium, equipment and program product, which can be applied to scenes such as gang mining, gang characterization and fraud black production of application platforms or application software. The method comprises the following steps: selecting behavior data of each object according to a preset time window to obtain behavior sequence data of each object, and the behavior sequence data comprises at least one object behavior; processing the behavior sequence data of each object to determine an object behavior similar set, wherein the object behavior similar set comprises behavior sequence data similar in behavior between each object; determining object similarity between each object according to the object behavior similar set within the time window; and performing mapping processing on each object according to the object similarity between each object to determine an abnormal object. The method effectively improves the identification coverage of the overall fraudulent user, and to some extent, realizes early discovery of abnormal objects in the pre-stage of the entire fraud life cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a method, apparatus, medium, device, and program product for identifying abnormal objects. Background Technology

[0002] Internet social media platforms possess massive user traffic, attracting numerous fraudsters who use these platforms to scam legitimate users. Currently, security models and strategies are designed to combat these scams. However, the introduction of individuals into these strategies makes threshold controls more sensitive, making it easier for fraudsters to detect these thresholds and increasing the risk of false positives. Furthermore, some anti-fraud systems heavily rely on specific fraudulent seeds for dissemination. Dissemination based on these seeds in the physical environment not only lags behind the actual fraudsters in terms of detection but also has significant limitations in terms of dissemination, failing to truly uncover and dismantle the entire fraudulent group. Summary of the Invention

[0003] This application provides a method, apparatus, medium, device, and program product for identifying abnormal objects, which effectively improves the overall coverage of fraudulent users and, to a certain extent, enables the early detection of abnormal objects in the early stages of the entire fraud lifecycle.

[0004] On the one hand, a method for identifying abnormal objects is provided, the method comprising:

[0005] Behavioral data of each object is selected according to a preset time window to obtain behavioral sequence data of each object, wherein the behavioral sequence data contains at least one object behavior;

[0006] The behavioral sequence data of each object is processed to determine an object behavior similarity set, which includes behavioral sequence data of similar behavior among the objects.

[0007] Within the time window, the object similarity between each object is determined based on the object behavior similarity set;

[0008] The objects are plotted based on their similarity to each other in order to identify anomalous objects.

[0009] On the other hand, an abnormal object identification device is provided, the device comprising:

[0010] The selection unit is used to select behavioral data of each object according to a preset time window to obtain behavioral sequence data of each object, wherein the behavioral sequence data contains at least one object behavior;

[0011] The determining unit is configured to process the behavioral sequence data of the various objects to determine an object behavior similarity set, wherein the object behavior similarity set includes behavioral sequence data showing similar behavior among the various objects; and

[0012] Within the time window, the object similarity between each object is determined based on the object behavior similarity set;

[0013] A composition unit is used to perform composition processing on each object based on the object similarity between the objects, so as to identify abnormal objects.

[0014] On the other hand, a computer-readable storage medium is provided, which stores a computer program adapted for loading by a processor to perform the steps in the abnormal object identification method of any of the above embodiments.

[0015] On the other hand, a computer device is provided, which includes a processor and a memory, wherein a computer program is stored in the memory, and the processor executes the steps in the abnormal object identification method of any of the above embodiments by calling the computer program stored in the memory.

[0016] On the other hand, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps in the method for identifying abnormal objects as described in any of the above embodiments.

[0017] This application selects behavioral data of each object according to a preset time window to obtain behavioral sequence data of each object. The behavioral sequence data contains at least one object behavior. The behavioral sequence data of each object is processed to determine an object behavior similarity set. The object behavior similarity set includes behavioral sequence data with similar behaviors among the objects. Within the time window, the object similarity between each object is determined according to the object behavior similarity set. Based on the object similarity between the objects, the objects are plotted to identify abnormal objects. By using a sliding time window and multiple behavioral feature similarity algorithms to calculate user similarity, the overall coverage of fraudulent user identification is effectively improved, and abnormal objects are detected in advance at the beginning of the entire fraud lifecycle to a certain extent. Compared with the existing technology of filtering by feature thresholds, which is prone to allowing black market operators to test the threshold and bypass the strategy within a certain period of time, this application can effectively track and uncover gangs in real time by processing the object behavior of abnormal objects. In addition, this application utilizes the temporal locality characteristic while integrating the similarity between multiple users, which can effectively alleviate the high threshold sensitivity problem of strategy-based attacks and solve the problems of delayed attacks and low coverage of diffusion systems to a certain extent. Attached Figure Description

[0018] To more clearly illustrate the technical methods in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 Example diagram of the method for identifying abnormal objects provided in the embodiments of this application;

[0020] Figure 2 A flowchart illustrating the method for identifying abnormal objects provided in this application embodiment;

[0021] Figure 3 Example diagram of the method for identifying abnormal objects provided in the embodiments of this application;

[0022] Figure 4 Another example diagram of the method for identifying abnormal objects provided in the embodiments of this application;

[0023] Figure 5 Another flowchart illustrating the method for identifying abnormal objects provided in the embodiments of this application;

[0024] Figure 6 Another example diagram of the method for identifying abnormal objects provided in the embodiments of this application;

[0025] Figure 7 This is a schematic diagram of the structure of the abnormal object recognition device provided in the embodiments of this application;

[0026] Figure 8 This is a schematic structural diagram of the abnormal object identification device provided in the embodiments of this application. Detailed Implementation

[0027] The technical methods in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] First, some of the nouns or terms that appear in the description of the embodiments of this application are explained as follows:

[0029] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0030] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.

[0031] A blockchain system can be a distributed system formed by clients and multiple nodes (any form of computing device connected to the network, such as servers and user terminals) connected through network communication. The nodes form a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In a distributed system, any machine, such as a server or terminal, can join and become a node. A node includes a hardware layer, a middleware layer, an operating system layer, and an application layer.

[0032] The "black market" (also known as cybercrime or fraud black market) refers to illegal activities that use the internet as a medium and network technology as the primary means to pose a potential threat to the security of computer information systems, the order of cyberspace management, and even national security and social and political stability. The cybercrime industry chain refers to the use of internet technology to carry out cyberattacks, steal information, extortion, fraud, theft of money, and the promotion of pornography, gambling, and drugs, as well as the channels and links that provide tools, resources, platforms, and other preparations for these activities, and the channels and links for illegally profiting from them.

[0033] In industries such as finance and the internet, there are numerous application platforms or software. For example, WeChat, as a massive traffic platform, attracts a large number of fraudulent criminals and groups who defraud legitimate users within WeChat.

[0034] Various anti-fraud technologies have been battling wits with cybercriminals for years, developing numerous models and strategies to combat them. However, with the iterative evolution of these efforts, the current fraud cycle differs from the large-scale registration scams of the past. Cybercriminals now utilize legitimate users to help create groups and recruit others, even using them to send fraudulent advertisements to evade existing strategies. As the fight against fraud intensifies, current fraud prevention efforts are encountering certain bottlenecks.

[0035] There are currently two main approaches to combating fraud and cybercrime. One is based on analyzing the behavioral patterns of individual users and then targeting them using combinations of features. To prevent false positives, this approach is highly sensitive to threshold restrictions on these features, allowing cybercriminals to test these thresholds and circumvent the strategy within a certain timeframe. However, it also increases the risk of false positives. The other approach is a diffusion attack strategy, typically based on attacks targeting devices with the same ID and environment. This approach has two significant problems: firstly, it requires a specific "black seed" to trigger the attack; secondly, due to upgraded security strategies, cybercriminals are gradually weakening the correlation between devices and physical environments, resulting in a significant lag. Diffusion attacks based on black seeds not only lag behind cybercriminals in terms of targeting but also have significant limitations in diffusion, failing to effectively and promptly uncover and combat the entire group, resulting in insufficient coverage.

[0036] Please see Figure 1 , Figure 1 This diagram illustrates the lifecycle of a horizontal scam on a social media platform. First, the account is acquired and packaged, including actions such as registration, login, information modification, and sharing on Moments. Second, the relationship building phase begins, including adding friends and joining groups. Then, social communication takes place, such as sending messages through tools like one-on-one chats, group chats, and mass messaging. Finally, the account receives negative feedback, such as being blocked, deleted, removed from groups, or reported.

[0037] After extensive and thorough analysis of the black market, two significant characteristics have been identified. First, a more generalized correlation. Unlike previous black market activities that were significantly linked in physical environments, current fraudulent activities are more dispersed, but they exhibit correlations across multiple dimensions, such as similar nicknames, profile pictures, and obvious similarities in the text used in reported cases. Second, a more generalized temporal locality. Black market activities complete a series of fraudulent actions within specific time windows, such as creating a group, modifying the group name, generating a group QR code, and inviting others to the group. These fraudulent actions may not have a strong chronological order, but they need to be completed within a specific time window, thus exhibiting a generalized temporal locality.

[0038] Based on the two significant characteristics of the black market mentioned above, this application provides a method, device, medium, and equipment for identifying anomalous objects. This alleviates the high threshold sensitivity issue of strategy-based attacks to a certain extent, while also effectively solving the problems of delayed attacks and low coverage in diffusion systems. This application can be applied to scenarios such as gang detection and characterization, internet anti-fraud, and fraud black market activities involving application platforms or software. The application platforms or software can be applied to the internet, social media, and financial sectors, among others. Furthermore, the anomalous object identification method provided in this application can be implemented by constructing a suitable model using artificial intelligence machine learning.

[0039] Specifically, the method in this application embodiment can be executed by a computer device, which can be a terminal or a server or other similar device.

[0040] To better understand the technical methods provided in the embodiments of this application, the following is a brief introduction to the application scenarios to which the technical methods provided in the embodiments of this application are applicable. It should be noted that the application scenarios described below are only for illustrating the embodiments of this application and are not intended to limit the scope. Taking the method for identifying abnormal objects executed by a computer device as an example, the computer device can be a terminal or a server, etc.

[0041] The embodiments of this application can be implemented in conjunction with cloud technology or blockchain network technology. As disclosed in the abnormal object identification method of this application, this data can be stored on the blockchain. For example, the behavioral data of each object, the behavioral sequence data of each object, the object behavior similarity set, object similarity, and abnormal objects can all be stored on the blockchain.

[0042] To facilitate the storage and retrieval of behavioral data, behavioral sequence data, similarity sets, similarity scores, and anomalous objects for each object, the anomalous object identification method may optionally further include: sending the behavioral data, behavioral sequence data, similarity sets, similarity scores, and anomalous objects for each object to a blockchain network, so that the nodes of the blockchain network fill the behavioral data, behavioral sequence data, similarity sets, similarity scores, and anomalous objects for each object into a new block, and appending the new block to the end of the blockchain when consensus is reached on the new block. This embodiment of the application can store the behavioral data, behavioral sequence data, similarity sets, similarity scores, and anomalous objects for each object on the blockchain, achieving record backup. When it is necessary to retrieve an anomalous object, the corresponding anomalous object can be directly and quickly retrieved from the blockchain, thereby improving the efficiency of anomalous object identification.

[0043] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.

[0044] The embodiments of this application provide a method for identifying abnormal objects. The embodiments of this application use a computer device as an example to illustrate the identification method.

[0045] Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for identifying abnormal objects provided in an embodiment of this application. The method includes:

[0046] Step 210: Select the behavior data of each object according to the preset time window to obtain the behavior sequence data of each object. The behavior sequence data of each object includes at least one object behavior.

[0047] Specifically, each object includes multiple objects, such as multiple accounts or users. The behavioral data for each object includes all behavioral data for that object, such as within a certain time period (e.g., one day, one week), or all behavioral data since the user's creation. For example, all behavioral data for a WeChat registered account within one day, including registration, modifications, and social interactions.

[0048] Behavioral data for each object is selected based on a preset time window to obtain behavioral sequence data for each object. Each behavioral sequence data contains at least one object behavior; for example, object 1's behavioral sequence data includes three object behaviors: registration, login, and name modification. The preset time window can be a predefined fixed-size time window or a time window adjustable based on time or actual business needs. For time windows adjustable based on time or actual business needs, for example, if the amount of behavioral data is large in certain time periods, the time window for those periods can be shortened, while if the amount of behavioral data is small in certain time periods, the time window for those periods can be lengthened.

[0049] Optionally, the step of selecting behavioral data of each object according to a preset time window to obtain behavioral sequence data of each object includes:

[0050] Obtain behavioral data for each object;

[0051] The behavioral data of each object is filtered according to the preset time window to obtain behavioral sequence data of multiple time windows, with preset overlap time between the multiple time windows.

[0052] Specifically, the behavior data of each object is obtained. This behavior data can include all the behavior data of each object, such as within a certain period of time, such as a day or a week, or all the behavior data since the user's existence, or the behavior data of multiple objects stored in the database.

[0053] Based on preset time windows, the behavioral data of each object is filtered to extract behavioral sequence data that fits within the time window, resulting in behavioral sequence data for multiple time windows. These time windows have preset overlap times, such as 5 minutes or 10 minutes. This avoids data loss at the boundaries of each time window.

[0054] Please see Figure 3 , Figure 3 The data represents the behavior of three objects. The behavior data within time window 1 is the behavior sequence data for time window 1, and the behavior data selected within time window 2 is the behavior sequence data for time window 2.

[0055] For example, if the preset time window size is 2 hours and the preset overlap time is 10 minutes, starting from 0:00, the first data taken will be the behavioral sequence data of the two-hour interval [00:00, 02:00], and the next data taken will be the behavioral sequence data of the two-hour interval [01:50, 04:00].

[0056] Step 220: Process the behavior sequence data of each object to determine the object behavior similarity set, which includes behavior sequence data of similar behavior among each object.

[0057] Specifically, the object behavior similarity set includes behavioral sequence data of similar behaviors among various objects. This similarity involves cross-comparing multiple behaviors across multiple objects; when several objects exhibit similar behavior, that behavior across those objects can be identified as object behavior similarity. The object behavior similarity set is determined based on all object behaviors in the behavioral sequence data that share the characteristic of object behavior similarity.

[0058] For example, consider objects A and B. Object A's behavior sequence data includes behaviors A1, A2, and A3. Object B's behavior sequence data includes behaviors B1, B2, and B5. After processing, behavior A1 is similar to behavior B1, behavior A2 is similar to behavior B2, while behavior A3 is dissimilar to B1, B2, and B3. Therefore, behavior A1 and behavior B1, and behavior A2 and behavior B2, can be identified as the object behavior similarity set.

[0059] Optionally, the step of processing the behavioral sequence data of each object to determine the object behavior similarity set includes:

[0060] Within a time window, the object behaviors of every two objects in the behavior sequence data are processed to determine the behavior feature set of the object behaviors of the two objects. The behavior feature set contains at least one behavior feature.

[0061] Specifically, within a time window, a behavior sequence data packet contains multiple objects, each object corresponding to one or more object behaviors, and each object behavior possesses one or more behavioral features. Behavioral features represent the attributes and action characteristics of this object behavior, and the behavioral feature set is a collection of behavioral features for each object behavior. For example, behavioral features in social software may include behavior types such as registration, login, greeting, sending messages, joining groups, etc.; the target of the behavior such as an account or a group; behavioral attributes such as IP address and device ID; and behavioral account features such as the object's avatar, friend request text, object name, number of recent greetings, and number of recent group joins.

[0062] Compare the similarity of any two behavioral feature sets between the object behaviors of two objects. If the similarity comparison result is similar, it is determined that the behavioral feature sets between the object behaviors of the two objects are similar.

[0063] Specifically, a behavioral feature set includes one or more behavioral features. The method for comparing the similarity of behavioral feature sets between two objects and determining whether the comparison result is similar may include:

[0064] Two sets of behavioral features are compared one by one. If all behavioral features in the two sets are similar, then the two sets of behavioral features are determined to be similar. For example, if the sets of behavioral features for objects A1 and B1 are {a1, a2, a3} and {b1, b2, b3}, where a1 is similar to b1, a2 is similar to b2, and a3 is similar to b3, then the sets of behavioral features {a1, a2, a3} and {b1, b2, b3} are determined to be similar.

[0065] Optionally, the step of comparing the similarity of the behavioral feature sets between the object behaviors of two objects, and determining that the behavioral feature sets of the object behaviors of the two objects are similar if the similarity comparison result is similar, further includes:

[0066] Retrieve the behavioral characteristics from the set of behavioral characteristics of two objects;

[0067] According to a predetermined comparison strategy, the behavioral features between any two behavioral feature sets are compared. If the comparison results are similar, it is determined that the behavioral features between the object behaviors of the two objects are similar.

[0068] If there is a predetermined proportion of similar behavioral features between any two sets of behavioral features, then the two sets of behavioral features are determined to be similar.

[0069] Specifically, obtaining the behavioral features from the behavioral feature sets of the object behaviors of two objects involves extracting the behavioral features from each of the two behavioral feature sets corresponding to any two object behaviors of the two objects.

[0070] For example, given two objects A and B, whose respective object behaviors are A{A1, A2, A3} and B{B1, B2, B3}, the feature sets of any two object behaviors can be obtained as the feature sets of object behaviors A1 and B2: A1{a1, a2, a3} and B2{b1, b2, b3}. Similarly, for objects A and B, other pairwise comparison feature sets can be obtained as A1-B1, A1-B3, A2-B1, A2-B2, A2-B3, A3-B1, A3-B2, and A3-B3.

[0071] According to a predetermined comparison strategy, the behavioral features between any two behavioral feature sets are compared. If the comparison results are similar, it is determined that the behavioral features between the object behaviors of the two objects are similar.

[0072] Specifically, different comparison strategies are used for different behavioral characteristics, and the results are either similar or dissimilar. Similarity can include a similarity exceeding a preset threshold, or the two characteristics being exactly the same. The specific comparison strategies will be explained in detail below.

[0073] If there is a predetermined proportion of similar behavioral features between any two sets of behavioral features, then the two sets of behavioral features are determined to be similar.

[0074] For example, if the predetermined ratio is set to 80%, then any two behavioral feature sets are considered similar if more than 80% of their behavioral features are similar. It can be understood that any two behavioral feature sets may have different numbers of behavioral features, and therefore the behavioral features themselves may also differ. In this case, a predetermined ratio can be taken from the behavioral feature set with the smallest number of behavioral features, or from the behavioral feature set with the largest number of behavioral features.

[0075] If any two sets of behavioral features are similar, and the time difference between the occurrence of the behaviors of the two objects corresponding to the two sets of behavioral features is within a preset time threshold, then the behaviors of the two objects corresponding to the two sets of behavioral features are determined to be similar.

[0076] Specifically, while determining that two behavioral feature sets are similar, the occurrence time of the object behaviors corresponding to these two behavioral feature sets is detected. For example, if behavioral feature sets {a1, a2, a3} and {b1, b2, b3} are determined to be similar, and the behaviors corresponding to behavioral feature sets {a1, a2, a3} and {b1, b2, b3} are A1 and B1, that is, A1{a1, a2, a3} and B1{b1, b2, b3}, then the occurrence time of object behaviors A1 and B1 is detected.

[0077] Calculate whether the time difference between the occurrences of A1 and B1 is within a preset time threshold. For example, if the preset time threshold is 30 minutes, then if the time difference between the occurrences of A1 and B1 is less than or equal to 30 minutes, it can be determined that the object behaviors A1 and B1 are similar.

[0078] Determining the similarity of behaviors between two objects corresponding to any two sets of behavioral features can be expressed by formula (1):

[0079] A ,T Ai C Ai >= B ,T Bj C Bj >ifC Ai =V Bj and|T Ai -T Bj |≤T sim (1);

[0080] Among them, U A and U B Let i represent user A and user B, and j represent a specific object behavior of user A and user B, respectively. Ai T represents the time when user A's action i occurs. Bj This represents the time when user B's action j occurs, and represents a preset time threshold. C Ai C represents the set of behavioral features of user A object i. Bj The set of behavioral features representing user B object behavior j.

[0081] Determine the object behavior similarity set based on the similar object behaviors identified in the behavior sequence data.

[0082] Specifically, within a time window of behavioral sequence data, there are multiple objects, and each object has multiple behaviors. Therefore, for the behavioral sequence data within the same time window, it is necessary to perform cross-similarity calculations and comparisons on the behaviors of each pair of objects. Then, all pairs of similar behaviors are added to the object behavior similarity set.

[0083] For example, given two objects A and B, whose respective object behaviors are A{A1, A2, A3} and B{B1, B2, B3}, the pairs of object behaviors that are determined to be similar are: A2 and B2, A2 and B3, A3 and B1, and A3 and B2. Then the set of similar behaviors is {A2-B2, A2-B3, A3-B1, A3-B2}.

[0084] Any object's behavior contains one or more behavioral features, and the comparison strategies for comparing the similarity between different behavioral features are different.

[0085] ​​Optionally, behavioral features include object names, and the step of comparing behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes:

[0086] Calculate the first distance between the object names of two objects using edit distance;

[0087] If the first distance is less than a first threshold, the object names of the two objects are determined to be similar.

[0088] Specifically, behavioral characteristics include object names, such as nicknames, accounts, and other textual object names. The first distance between the object names of two objects can be calculated using edit distance. Edit distance is a quantitative measure of the difference between two strings, determined by the minimum number of processing steps required to transform one string into another. The number of processing steps represents the first distance between the two objects.

[0089] For example, if nickname a = 'love' and nickname b = 'lolpe', calculating the edit distance between a and b means determining how many steps are required to change from a to b.

[0090] 1. love->lolve (insert l);

[0091] 2. lolve->lolpe (replace v with p).

[0092] The change from a to b requires two steps, so the edit distance between nicknames a and b is 2.

[0093] If the first distance is less than a first threshold, the object names of two objects are considered to have similar behavior. The first threshold can be set according to the depth of data mining, specific business scenarios, and other practical application scenarios. The larger the first threshold, the lower the similarity requirement. For example, if the first threshold is 3, then when the first distance between object names is less than or equal to 3, the behavioral characteristics of the two object names are considered similar.

[0094] Thus, the similarity between two behavioral features can be determined by comparing the similarity of the account name attribute in the behavioral features.

[0095] Optionally, the behavioral features include the reported text, and the step of comparing the behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes:

[0096] The reported texts of the two objects are processed using a classification model, which is trained based on the reported texts of each object.

[0097] If the classification model outputs the same result for two objects, the reported texts of the two objects are determined to be similar.

[0098] Specifically, behavioral characteristics include the reported text. For example, if an account has been reported before, the corresponding reporting information, i.e., the reported text, will be filled in when reporting. For instance, "This account is a scammer; they join the group to post advertisements and attract people to buy things." Keywords such as "scam," "advertisement," and "buy things" can be extracted. A corresponding classification model can be built based on these characteristics to filter the reported text. When preset sensitive texts are present, different reported texts can be identified based on whether they meet the preset sensitive text criteria. The classification model can be pre-trained based on the reported texts of each object.

[0099] For example, if the reported text 1 is "This account is a fraudster, and the purpose of joining the group is to attract people to place fake orders", then it can be determined as "fake orders" through the corresponding classification model.

[0100] For example, if the reported text 2 is "This account is a scammer, and the person who joined the group is here to post advertisements", then it can be identified as "advertisement" by the corresponding classification model.

[0101] The output results of the reported texts 1 and 2 are "brushing orders" and "advertising" respectively. They can be determined to be dissimilar by a one-to-one correspondence method, or they can be determined to be similar by a certain range, such as setting "brushing orders" and "advertising" to both fall under the category of "fraud".

[0102] In other examples, the reported text can be learned through other machine learning models similar to classification models, such as combining other features such as account origin, account attributes, and operating environment.

[0103] In this way, the similarity between two behavioral characteristics can be determined by comparing the similarity of the reported content in the behavioral characteristics.

[0104] Optionally, the behavioral features include the friend request text. The step of comparing the behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes:

[0105] Perform text recognition on the friend request texts between two objects to determine multiple text features for each object;

[0106] If at least one text feature is the same among multiple text features of two objects, the friend request text of the two objects is determined to be similar.

[0107] Specifically, the Named Entity Recognition (NER) model can be used to identify user friend requests. If two users impersonate multiple entities simultaneously, and these entities overlap, they are considered similar. The NER model is specifically designed to extract names from a sentence.

[0108] For example, if the friend request text of impersonating user 1 is "I am singer a, I am actor b, I am host c", and the friend request text of impersonating user 2 is "I am actor x, I am actor b, I am singer y", then the NER model can determine that impersonating users 1 and 2 have an overlapping part "b". Thus, impersonating users 1 and 2 can be identified as having similar friend request texts based on behavioral characteristics.

[0109] It should be noted that, in addition to the text for adding friends, behavioral features may also include other behavioral attribute features containing text. Multiple text features can also be determined by using the NER model or other models to perform text recognition. If there is at least one identical text feature among the multiple text features of two objects, it can be determined that the text for adding friends of the two objects is similar.

[0110] Thus, similarity can be determined between two behavioral features by comparing the similarity of textual content within the behavioral features.

[0111] Optionally, behavioral features include object avatars, and the step of comparing behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes:

[0112] Compare the image similarity of the headshots of two objects;

[0113] If the image similarity comparison result is less than the second threshold, the two objects are determined to have similar avatars.

[0114] Specifically, in social media software or platforms, profile pictures are linked to accounts, typically represented by static or animated images. Image similarity comparison algorithms can include those that use ID photo algorithms, OCR algorithms, or other image recognition or extraction algorithms for comparison. The second threshold is a similarity threshold setting. If the calculated similarity is 65%, and the second threshold is set to 80%, then the profile pictures of the two objects are not similar. If the calculated similarity is 85%, then the profile pictures of the two objects are similar.

[0115] For example, using an ID photo algorithm to calculate the similarity of two users' profile pictures, if both users have ID photo profile pictures, then the two users can be considered similar. Alternatively, if the similarity of the two users' ID photo profile pictures is less than a second threshold, then the two users can be considered similar.

[0116] It should be noted that, in addition to the object's avatar, behavioral features may also include other image-based behavioral features in other application scenarios. Alternatively, the similarity between the avatars of two objects can be compared, and if the comparison result is less than a second threshold, the avatars of the two objects can be determined to be similar.

[0117] Optionally, the step of comparing behavioral features between any two behavioral feature sets according to a predetermined comparison strategy includes:

[0118] Use an optical character recognition algorithm to obtain the text information contained in the avatars of two objects;

[0119] Calculate the second distance between the text information contained in the object portraits of two objects using edit distance;

[0120] If the second distance is less than the third threshold, the two objects are determined to have similar avatars.

[0121] Specifically, if the avatars of certain objects contain text information, text recognition algorithms, such as Optical Character Recognition (OCR), can be used to identify the text information contained in the avatars of two objects. Then, edit distance is used to calculate a second distance between the text information contained in the avatars of the two objects. If the second distance is less than a third threshold, the avatars of the two objects are determined to be similar. The specific calculation method is the same as the edit distance calculation method in the above embodiment, and will not be repeated here.

[0122] It should be noted that, in addition to the object's avatar, behavioral features may also include other image-based behavioral features containing text information in other application scenarios. Alternatively, text recognition algorithms can be used to obtain the text information contained in the avatars of two objects, and then the edit distance can be used to calculate the second distance between the text information contained in the avatars of the two objects. If the second distance is less than a third threshold, it can be determined that the image-based behavioral features of the two objects are similar.

[0123] Thus, similarity can be determined between two behavioral features by comparing the similarity of the image-related content within the behavioral features.

[0124] Optionally, behavioral features include object type, and the step of comparing behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes:

[0125] If two objects have the same object type, it is determined that the object types of the two objects are similar. The object types include normal account types and abnormal account types.

[0126] Specifically, behavioral characteristics include object type, such as account type. In the social field, in the management of account attributes, there are account types, including normal account types and abnormal account types, such as stolen accounts, accounts bought and sold, and other abnormal account types.

[0127] Two objects are considered similar in type if they are of the same type. For example, if two objects both have accounts that have been hacked, then their object types are similar.

[0128] Thus, it is possible to determine whether two behavioral characteristics are similar by comparing the similarity of the account types of the objects in the behavioral characteristics.

[0129] Optionally, behavioral features include physical attributes, and the step of comparing behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes:

[0130] The physical attributes of the object behaviors of two objects are compared according to a predetermined physical attribute comparison strategy, and the physical attributes of the two objects are determined to be similar based on the comparison results.

[0131] Specifically, physical attributes can include behavioral attributes of the object's behavior or environmental attributes, such as the IP address, device ID, device name, and client version of the electronic device on which the object resides, which are characteristics related to the physical environment.

[0132] For example, physical attribute IP address. IP class similarity comparison can be achieved by detecting the IP address when the object behavior of two objects is initiated. For example, the IP of object behavior 1 is 179.163.118.7 and the IP of object behavior 2 is 179.163.118.9. The first three parts of the two IPs are completely identical, so they can be judged as IP similar, that is, the physical attributes of object behavior 1 and 2 are similar.

[0133] For example, if the physical attribute is a device class, device class similarity comparison can be achieved by detecting device-related information when the object's behavior is initiated. For instance, if the device UUIDs of the two object behaviors are the same, they are considered to be similar in terms of devices, i.e., similar in physical attributes. Alternatively, if two devices are of the same type and both belong to low-end devices, they are also considered to be similar in terms of devices, i.e., similar in physical attributes.

[0134] Thus, by comparing the physical attributes of the environment, it can be determined whether two behavioral characteristics are similar.

[0135] Optionally, the step of comparing behavioral features between any two behavioral feature sets according to a predetermined comparison strategy includes:

[0136] Process the behavioral characteristics of the two objects based on the behavioral characteristic set of the object black database;

[0137] If the behavioral features of two objects are both in the object black database, it is determined that the behavioral features of the two objects are similar.

[0138] The target blacklist includes at least: avatar blacklist, name blacklist, financial blacklist, website blacklist, application blacklist, and device blacklist.

[0139] Specifically, the object blacklist is a collection of abnormal objects accumulated in the history of various application fields, which may include: avatar blacklist, name blacklist, financial blacklist, website blacklist, application blacklist and device blacklist, etc. Among them, the financial blacklist includes bank card number blacklist, stock code / stock name blacklist and financial product-specific blacklist.

[0140] The behavioral features of an object are matched against various object black databases. If any two behavioral features of two objects match data records in the same type of black database, then the behavioral features of the two objects are determined to be similar. The type of behavioral features is not limited and can include any of the aforementioned behavioral features.

[0141] The similarity of behavioral features is determined using an object blacklist, which can be arbitrarily combined with any of the aforementioned behavioral feature similarity methods, and the object blacklist method can be set as the highest priority. For example, if two objects' nicknames are dissimilar according to the edit distance calculation, but both are in the name blacklist, then the behavioral features of the nicknames of the two objects can be determined to be similar.

[0142] For example, a blacklist of profile pictures can be used to determine whether two users are similar. If the profile pictures of the two users match the blacklist of profile pictures that are counterfeit SF Express and counterfeit STO Express respectively, then they are considered similar.

[0143] For example, a nickname blacklist can be used to determine whether two users are similar. If the nicknames of two users match the blacklisted nicknames "Loan Customer Service Xiao Zhou" and "Loan Customer Service Xiao Liu", then they are considered similar.

[0144] For example, by using a device blacklist to determine that two users are using devices commonly used by illicit activities, they are considered similar.

[0145] In this way, behavioral characteristics can be further compared through the object black database, and the results of abnormal objects identified through this application can be newly added to the object black database, so that the identification of abnormal users has the historical reference of the database and improves the efficiency of identification.

[0146] It should be noted that the above methods for calculating the similarity of various behavioral features can be combined arbitrarily. The same type of similarity calculation method can be used to calculate different types of behavioral features, and the same type of behavioral features can also be compared multiple times using different types of similarity calculation methods. Finally, the similarity result is determined according to predetermined rules, such as using priority.

[0147] Optionally, the step of processing the behavioral sequence data of each object to determine the object behavior similarity set further includes:

[0148] The account characteristics of object behavior are statistically analyzed and / or concatenated to preprocess the object behavior in the behavior sequence data of each object;

[0149] Determine the object behavior similarity set based on the preprocessed behavior sequence data.

[0150] Specifically, after obtaining behavioral sequence data within a time window, before comparing behavioral features, the behavioral sequence data of each object can be processed. This includes statistically analyzing and / or concatenating the account characteristics of the object's behavior. Statistical analysis includes calculating certain statistical features, such as the number of times recently added friends or joined groups via QR codes, and calculating the quantity of these behavioral features. Concatenation includes concatenating recently joined group names. For example, some abnormal users may extract a large number of group names by extracting text segments in order to engage in fraudulent activities; concatenating these group names will yield continuous text. Similarly, behavioral features such as the object's profile picture can be concatenated.

[0151] In this way, data can be statistically analyzed and features can be concatenated through preprocessing, making the preprocessed behavioral sequence data more concise and effective, and improving recognition efficiency.

[0152] Step 230: Within the time window, determine the object similarity between each object based on the object behavior similarity set.

[0153] Specifically, an object can contain one or more object behaviors. When multiple object behaviors are similar, i.e., after obtaining the object behavior similarity set between objects, the object similarity between each object can be determined. For example, users A and B have object behaviors A{A1, A2, A3} and B{B1, B2, B3}. The behavior similarity set of users A and B is {A1-B1, A1-B2, A1-B3, A2-B1, A2-B2, A2-B3, A3-B1, A3-B2, A3-B3}, which covers all object behavior pairs of A and B. Therefore, the object similarity can be determined to be 100%.

[0154] Optionally, the step of determining the object similarity between objects based on the object behavior similarity set includes:

[0155] Based on the object behavior similarity set and object behavior weight, the object behavior similarity set of each object is aggregated to determine the object similarity between each object.

[0156] Specifically, the object behavior weights can preset different weights for different behavior characteristics. For example, a higher weight can be set for similar avatars, and a lower weight for similar nicknames, or some heuristic algorithms can be used to define the weights. Among them, the method of using heuristic algorithms to define weights can be: in calculating the user similarity, the similarity has a relatively small contribution to the recognition of the final abnormal user, but at the same time contributes a relatively large number of connected behavior characteristics, multiplied by a penalty coefficient. For example, for the behavior characteristic of similar nicknames, the contribution of similar nicknames to the recognition of the final abnormal user is relatively weaker compared to similar avatars, so a penalty term calculation can be performed on the result of similar nicknames.

[0157] Furthermore, according to the object behavior similarity set and the object behavior weights, aggregating the object behavior similarity sets of each object can be achieved through an aggregation algorithm, expressed as formulas (2) and (3):

[0158]

[0159] Among them, U i and U j represent object i and object j; A i represents the object behavior of object i; A j represents the object behavior of object j; represents the k-th behavior characteristic of the object behavior; w k represents the object behavior weight of the k-th behavior characteristic; δ k is the regularization factor of the k-th behavior characteristic, and the regularization factor δ k can be expressed as formula (3):

[0160]

[0161] Among them, represents the k-th behavior characteristic of the object behavior of object i. represents the k-th behavior characteristic of the object behavior of object j.

[0162] For example, when two user pairs have X times of behavior and Y times of behavior respectively under the behavior characteristic k, and X < Y, the regularization factor is X / Y. In formula (2), the similarity result calculated for the two user pairs under the behavior characteristic k is multiplied by the regularization factor X / Y as the regularization term.

[0163] In an example, the object similarity calculation results of users A, B, and C are as follows:

[0164] object pairs Time window Behavioral characteristics Object similarity AB 1 object name 0 BC 1 Object Avatar 0.7 AB 3 Add friend text 0.1 BC 3 Add friend text 0.9

[0165] As shown in the table above, for the behavioral sequence data of object pair A and B in time windows 1 and 3, the calculated object similarities for different behavioral features (object name and add friend text) are 0 and 0.1 respectively, indicating low similarity. However, for the behavioral sequence data of object pair B and C in time windows 1 and 3, the calculated object similarities for different behavioral features (object avatar and add friend text) are 0.7 and 0.9 respectively, indicating high similarity.

[0166] Step 240: Perform graph construction on each object based on the object similarity between them to identify abnormal objects.

[0167] Optionally, the step of performing graph manipulation on each object based on object similarity to identify anomalous objects includes:

[0168] A graph is constructed based on the object similarity between each object to generate an object similarity graph;

[0169] Graph algorithms are used to analyze object similarity graphs to identify anomalous objects.

[0170] Specifically, the graph is first constructed based on the object similarity between each object. Then, the graph can be analyzed using subgraph mining algorithms such as BotGraph algorithm, single-constraint graph construction, multi-constraint graph construction, graph clustering algorithm, and graph segmentation algorithm to identify abnormal objects, such as uncovering black market gangs.

[0171] For example, please see Figure 4 , Figure 4 This is an example graph constructed based on the object similarity of each object. It utilizes behavioral features such as object name and object avatar, connecting objects with similar names. A subgraph mining algorithm is then used to analyze the graph and identify anomalous objects.

[0172] Please see Figure 5 Optionally, after identifying abnormal objects, abnormal users can also be qualitatively processed, including:

[0173] Step 250: Perform qualitative analysis on the abnormal object to determine its degree of malice;

[0174] Step 260: If the malice level is greater than the fourth threshold, perform the first processing on the abnormal object. The first processing includes banning the object's account or banning the object's functions.

[0175] Step 270: If the malice level is less than or equal to the fourth threshold, the behavior data of the abnormal object is sent to the manual review terminal, and based on the manual review results sent by the manual review terminal, the abnormal object is subjected to a second processing, which includes restoring the object to normal, banning the object's account, or banning the object's functions.

[0176] Specifically, qualitative processing can be divided into automatic qualitative processing and manual qualitative processing. The qualitative processing of the obtained abnormal objects first involves calculating the malice level of an abnormal object, for example, by using the proportion of reported users within the group and the proportion of users registered in malicious regions. A fourth threshold is set to distinguish between automatic and manual qualitative processing. The manual review end can be a corresponding manual management platform, which may be a different device from the computer device executing this application, or it may be the same device. Suspicious objects requiring manual review are displayed through an appropriate user graphical interface, allowing relevant personnel to review suspicious objects by viewing their behavioral data. Furthermore, relevant personnel can input the manual review results into the manual review end through the input window displayed in the user graphical interface, and the manual review end then feeds back the manual review results to the computer device executing this application.

[0177] Groups with a malicious intent level exceeding the fourth threshold will undergo the first step, including direct account bans and function restrictions. Groups with a malicious intent level less than or equal to the fourth threshold will undergo the second step, which involves manual review. The second step includes restoring the target to normal status, banning the target's account, or restricting its functions. After manual assessment, if the review concludes that the identification was incorrect, the target can be restored to normal status; if the review determines it is a group, then the target's account or functions can be banned.

[0178] Please see Figure 6 , Figure 6 A flowchart illustrating the abnormal object identification method of this application is provided. First, behavioral data of each object is acquired and preprocessed. Then, behavioral data of each object is selected according to a preset time window to obtain behavioral sequence data of each object. The behavioral characteristics of each pair of objects are compared for similarity to obtain a set of object behavior similarities. The similarity of objects is determined based on the set of object behavior similarities, and then graph processing is performed based on the object similarity. Qualitative processing, such as gang detection, can be performed to determine whether an object is abnormal. If an abnormal object is determined, corresponding actions are performed, such as banning fraudulent accounts. If qualitative processing cannot determine the object, it is considered a suspicious object and sent for manual review. If the manual review confirms the object as abnormal, corresponding actions are performed, such as banning fraudulent accounts. If the object is determined to be normal, processing can be stopped, or the normal account can be restored.

[0179] Thus, this application selects behavioral data of each object according to a preset time window to obtain behavioral sequence data of each object. The behavioral sequence data contains at least one object behavior. The behavioral sequence data of each object is processed to determine an object behavior similarity set. This object behavior similarity set includes behavioral sequence data of similar behaviors among the objects. Within the time window, the object similarity between objects is determined based on the object behavior similarity set. Based on the object similarity between objects, a graphing process is performed on each object to identify abnormal objects. This has at least the following beneficial effects:

[0180] (1) This application uses a sliding time window and multiple behavioral feature similarity algorithms to calculate the similarity of users, which effectively improves the overall coverage of fraudulent users and to a certain extent realizes the early detection of abnormal objects in the early stage of the entire fraud life cycle.

[0181] (2) Compared to existing technologies that use feature thresholds for screening, which can easily allow malicious actors to test the thresholds and circumvent the strategy within a certain period of time, this application can effectively track and uncover gangs in real time by processing the object behavior of abnormal objects. In addition, this application utilizes the temporal locality characteristic while integrating the similarity between multiple users, which can effectively alleviate the high threshold sensitivity problem of strategy-based attacks and solve the problems of delayed attacks and low coverage in diffusion systems to a certain extent.

[0182] (3) Compared to existing technologies that identify criminal gangs through physical connections such as devices, the accuracy of such identification is low because the hardware connections of criminal gangs are gradually weakening and the methods of combating them are significantly lagging behind. This application determines the similarity of object behavior through soft constraints such as object behavior content, including object avatars, object names, and reported texts. At the same time, it can also combine hard constraints such as the hardware environment of object behavior to determine the similarity of object behavior, thereby conducting gang identification. As criminal gangs develop, they exhibit connections in more content dimensions, so the soft constraints of behavioral similarity can effectively improve the identification efficiency of abnormal objects.

[0183] (4) Within a specific time window, this application calculates the similarity between the behaviors of two objects by using an algorithm based on the similarity of multiple types of behavioral features, providing a more generalizable and extensible similarity clustering framework based on user time sequence.

[0184] (5) Multiple graph algorithms are used for gang detection. The efficiency of gang detection is improved by combining automatic and manual review strategies. Compared with traditional hardware environment association, this framework adopts more content-based association, which effectively improves the detection and crackdown on fraud and malice.

[0185] All the above-mentioned technical methods can be combined in any way to form optional embodiments of this application, and will not be described in detail here.

[0186] To facilitate better implementation of the abnormal object identification method of this application's embodiments, this application also provides an abnormal object identification device. Please refer to... Figure 7 , Figure 7 This is a schematic diagram of the structure of an abnormal object identification device provided in an embodiment of this application. The abnormal object identification device 700 may include:

[0187] The selection unit 710 is used to select the behavior data of each object according to a preset time window to obtain the behavior sequence data of each object, wherein the behavior sequence data contains at least one object behavior.

[0188] The determining unit 720 is used to process the behavior sequence data of each object to determine the object behavior similarity set, which includes behavior sequence data of similar behavior among the objects; and

[0189] Within a time window, the object similarity between objects is determined based on the object behavior similarity set.

[0190] The composition unit 730 is used to perform composition processing on each object based on the object similarity between each object in order to identify abnormal objects.

[0191] Optionally, the selection unit 710 can be used to acquire the behavior data of each object; according to the preset time window, the behavior data of each object is filtered to obtain the behavior sequence data of multiple time windows, and there is a preset overlap time between the multiple time windows.

[0192] Optionally, the determining unit 720 can be used to process the object behaviors of every two objects in the behavior sequence data within a time window to determine the behavior feature set of the object behaviors of the two objects, wherein the behavior feature set contains at least one behavior feature; compare the similarity of any two behavior feature sets between the object behaviors of the two objects, and determine that any two behavior feature sets are similar if the similarity comparison result is similar; determine that the two object behaviors corresponding to any two behavior feature sets are similar if any two behavior feature sets are similar and the occurrence time difference of the two object behaviors corresponding to any two behavior feature sets is within a preset time threshold; and determine the object behavior similarity set based on the similar object behaviors determined in the behavior sequence data.

[0193] Optionally, the determining unit 720 can also be used to obtain behavioral features from the behavioral feature sets of the object behaviors of two objects; compare the behavioral features between any two behavioral feature sets according to a predetermined comparison strategy; if the comparison result is similar, determine that the behavioral features between the object behaviors of the two objects are similar; if there is a predetermined proportion of similar behavioral features between any two behavioral feature sets, determine that any two behavioral feature sets are similar.

[0194] Optionally, the determining unit 720 can also be used to calculate a first distance between the object names of two objects using edit distance; if the first distance is less than a first threshold, the object names of the two objects are determined to be similar.

[0195] Optionally, the determining unit 720 can also be used to process the reported texts of two objects using a classification model, which is trained based on the reported texts of each object; if the reported texts of two objects produce the same output by the classification model, the reported texts of the two objects are determined to be similar.

[0196] Optionally, the determining unit 720 can also be used to perform text recognition on the friend-adding text of two objects to determine multiple text features of each object; if there is at least one identical text feature among the multiple text features of the two objects, the friend-adding text of the two objects is determined to be similar.

[0197] Optionally, the determining unit 720 can also be used to compare the image similarity of the object portraits of two objects; if the comparison result of the image similarity is less than the second threshold, the object portraits of the two objects are determined to be similar.

[0198] Optionally, the determining unit 720 can also be used to obtain the text information contained in the object portraits of two objects using a text recognition algorithm; calculate the second distance between the text information contained in the object portraits of two objects using edit distance; and determine that the object portraits of the two objects are similar if the second distance is less than a third threshold.

[0199] Optionally, the determining unit 720 can also be used to determine that the object types of the two objects are similar when the object types of the two objects are the same, and the object types include normal account types and abnormal account types.

[0200] Optionally, the determining unit 720 can also be used to compare the physical attributes of the object behaviors of two objects according to a predetermined physical attribute comparison strategy, and determine that the physical attributes of the two objects are similar based on the comparison results.

[0201] Optionally, the determining unit 720 can also be used to process the behavioral features of the behavioral feature sets of the object behaviors of two objects based on the object black database; when the behavioral features of the behavioral feature sets of the object behaviors of two objects are all in the object black database, it is determined that the behavioral features between the object behaviors of the two objects are similar; the object black database includes at least: avatar black database, name black database, financial black database, website black database, application black database and device black database.

[0202] Optionally, the determining unit 720 can also be used to statistically analyze and / or concatenate the account features of object behaviors to preprocess the object behaviors in the behavior sequence data of each object; and to determine the object behavior similarity set based on the preprocessed behavior sequence data.

[0203] Optionally, the determining unit 720 can also be used to aggregate the object behavior similarity sets of each object based on the object behavior similarity set and the object behavior weight, so as to determine the object similarity between each object.

[0204] Optionally, the graphing unit 730 can be used to construct a graph based on the object similarity between various objects to generate an object similarity graph; and to analyze the object similarity graph using a graph algorithm to identify abnormal objects.

[0205] Optionally, the abnormal object identification device 700 further includes a qualitative unit 740. The construction unit 740 can be used to perform qualitative analysis on the abnormal object to determine the malice level of the abnormal object; if the malice level is greater than a fourth threshold, the abnormal object is subjected to first processing, which includes banning the object's account or banning the object's functions; if the malice level is less than or equal to the fourth threshold, the abnormal object is subject to manual review to perform second processing, which includes restoring the object to normal, banning the object's account, or banning the object's functions.

[0206] It should be noted that the functions of each module in the abnormal object identification device 700 in this application embodiment can be referred to the specific implementation of any embodiment in the above method embodiments, and will not be repeated here.

[0207] Each unit in the aforementioned abnormal object identification device 700 can be implemented entirely or partially through software, hardware, or a combination thereof. Each unit can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each unit.

[0208] The abnormal object identification device 700 can be integrated, for example, into a terminal or server with storage and a processor, thus possessing computing capabilities; or the abnormal object identification device 700 can be the terminal or server itself. The terminal can be a smartphone, tablet, laptop, smart TV, smart speaker, wearable smart device, personal computer (PC), etc. The terminal can also include a client, which can be a video client, browser client, or instant messaging client, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0209] Figure 8 A schematic structural diagram of the abnormal object recognition device 800 provided in the embodiments of this application is shown below. Figure 8 As shown, the device 800 may include: a communication interface 801, a memory 802, a processor 803, and a communication bus 804. The communication interface 801, memory 802, and processor 803 communicate with each other via the communication bus 804. The communication interface 801 is used for data communication between the device 800 and external devices. The memory 802 can be used to store software programs and modules, and the processor 803 runs the software programs and modules stored in the memory 802, such as the software programs for the corresponding operations in the foregoing method embodiments.

[0210] Optionally, the processor 803 can invoke software programs and modules stored in the memory 802 to perform the following operations:

[0211] Behavioral data of each object is selected according to a preset time window to obtain behavioral sequence data of each object. The behavioral sequence data contains at least one object behavior.

[0212] The behavior sequence data of each object is processed to determine the object behavior similarity set, which includes the behavior sequence data of similar behaviors among the objects;

[0213] Within a time window, the object similarity between objects is determined based on the object behavior similarity set.

[0214] The graph of each object is constructed based on the object similarity between them in order to identify abnormal objects.

[0215] Optionally, the abnormal object identification device 800 can be integrated into a terminal or server with storage and a processor, or the abnormal object identification device 800 can be the terminal or server itself. The terminal can be a smartphone, tablet, laptop, smart TV, smart speaker, wearable smart device, personal computer, or other similar device. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0216] Optionally, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0217] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding process in the abnormal object identification method of the embodiments of this application; for the sake of brevity, it will not be described in detail here.

[0218] This application also provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding process in the abnormal object identification method of the embodiments of this application. For simplicity, further details are omitted here.

[0219] This application also provides a computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding process in the abnormal object identification method of the embodiments of this application. For the sake of brevity, further details are omitted here.

[0220] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0221] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0222] It should be understood that the above-described memory is exemplary and not a limiting description. For example, the memory in the embodiments of this application may also be static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM), etc. That is to say, the memory in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0223] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical method. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0224] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0225] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0226] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the method in this embodiment, depending on actual needs.

[0227] In addition, the functional units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0228] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical methods of this application, essentially or in other words, the parts that contribute to the prior art, or a portion of the technical methods, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0229] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for identifying abnormal objects, characterized in that, The method includes: Behavioral data of each object is selected according to a preset time window to obtain behavioral sequence data of each object, wherein the behavioral sequence data contains at least one object behavior; The behavioral sequence data of each object is processed to determine an object behavior similarity set, which includes behavioral sequence data of similar behavior among the objects. Within the time window, the object similarity between each object is determined based on the object behavior similarity set, including: aggregating the object behavior similarity sets of each object based on the object behavior similarity set, object behavior weights, and regularization factors to determine the object similarity between each object, wherein the regularization factor is determined based on the ratio of the smaller to the larger value of the number of behaviors of two objects under the same behavioral features; The objects are plotted based on their similarity to each other in order to identify anomalous objects.

2. The method according to claim 1, characterized in that, The step of selecting behavioral data of each object according to a preset time window to obtain behavioral sequence data of each object includes: Obtain the behavioral data of each of the objects; According to the preset time window, the behavior data of each object is filtered to obtain the behavior sequence data of multiple time windows, and there is a preset overlap time between the multiple time windows.

3. The method according to claim 1, characterized in that, The process of processing the behavioral sequence data of each object to determine the object behavior similarity set includes: Within one of the time windows, the object behaviors of every two objects in the behavior sequence data are processed to determine the behavior feature set of the object behaviors of the two objects, the behavior feature set containing at least one behavior feature; Compare the similarity of any two behavioral feature sets between the object behaviors of the two objects. If the similarity comparison result is similar, determine that the two behavioral feature sets are similar. If any two sets of behavioral features are similar, and the time difference between the occurrence of the behaviors of the two objects corresponding to the two sets of behavioral features is within a preset time threshold, then the behaviors of the two objects corresponding to the two sets of behavioral features are determined to be similar. The object behavior similarity set is determined based on the similar object behaviors identified in the behavior sequence data.

4. The method according to claim 3, characterized in that, The step of comparing the similarity of any two behavioral feature sets between the object behaviors of the two objects includes: Obtain the behavioral features from the set of behavioral features of the object behaviors of the two objects; According to a predetermined comparison strategy, the behavioral features between any two sets of behavioral features are compared. If the comparison result is similar, it is determined that the behavioral features between the object behaviors of the two objects are similar. If there is a predetermined proportion of similar behavioral features between any two sets of behavioral features, then the two sets of behavioral features are determined to be similar.

5. The method according to claim 4, characterized in that, The behavioral features include object names, and the comparison of behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes: Calculate the first distance between the object names of the two objects using edit distance; If the first distance is less than a first threshold, the object names of the two objects are determined to be similar.

6. The method according to claim 4, characterized in that, The behavioral features include the reported text, and the comparison of behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes: The reported texts of the two objects are processed using a classification model, which is trained based on the reported texts of each object. If the reported texts of the two objects produce the same result after being processed by the classification model, then the reported texts of the two objects are determined to be similar.

7. The method according to claim 4, characterized in that, The behavioral features include friend request text messages. The comparison of behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes: The friend request texts of the two objects are subjected to text recognition to determine multiple text features of each object; If at least one text feature is identical among multiple text features of the two objects, the friend request text of the two objects is determined to be similar.

8. The method according to claim 4, characterized in that, The behavioral features include the object's avatar, and the comparison of behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes: Compare the image similarity of the headshots of two objects; If the comparison result of the image similarity is less than the second threshold, the object avatars of the two objects are determined to be similar.

9. The method according to claim 8, characterized in that, The step of comparing behavioral features between any two behavioral feature sets according to a predetermined comparison strategy further includes: The text information contained in the avatars of the two objects is obtained using a text recognition algorithm; The second distance between the text information contained in the object portraits of the two objects is calculated using edit distance; If the second distance is less than the third threshold, the two objects are determined to have similar avatars.

10. The method according to claim 4, characterized in that, The behavioral features include object type, and the comparison of behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes: If the two objects have the same object type, it is determined that the object types of the object behaviors of the two objects are similar, and the object types include normal account types and abnormal account types.

11. The method according to claim 4, characterized in that, The behavioral features include physical attributes, and the comparison of behavioral features between any two sets of behavioral features according to a predetermined comparison strategy includes: The physical attributes of the object behaviors of the two objects are compared according to a predetermined physical attribute comparison strategy, and the physical attributes of the two objects are determined to be similar based on the comparison results.

12. The method according to claim 4, characterized in that, The step of comparing behavioral features between any two behavioral feature sets according to a predetermined comparison strategy includes: Process the behavioral features of the two objects based on the behavioral feature set of the object behavior of the object black database; If the behavioral features of the two objects are both in the object black database, it is determined that the behavioral features of the two objects are similar. The object blacklist includes at least: avatar blacklist, name blacklist, financial blacklist, website blacklist, application blacklist, and device blacklist.

13. The method according to claim 1, characterized in that, The step of processing the behavioral sequence data of the various objects to determine the object behavior similarity set further includes: The account characteristics of the object behavior are statistically analyzed and / or concatenated to preprocess the object behavior in the behavior sequence data of each object; The object behavior similarity set is determined based on the preprocessed behavior sequence data.

14. The method according to any one of claims 1-13, characterized in that, The step of performing graph construction on each object based on the object similarity between the objects to identify anomalous objects includes: A graph is constructed based on the object similarity between the objects to generate an object similarity graph; The abnormal object is identified by analyzing the similarity graph of the object using a graph algorithm.

15. The method according to any one of claims 1-13, characterized in that, The method also includes: Qualitative analysis is performed on the abnormal object to determine its degree of malice; If the degree of malice is greater than the fourth threshold, the abnormal object is subjected to a first processing, which includes banning the object's account or banning the object's functions. If the degree of malice is less than or equal to the fourth threshold, the behavior data of the abnormal object is sent to the manual review terminal, and based on the manual review result sent by the manual review terminal, the abnormal object is subjected to a second processing, which includes at least one of the following: restoring the object to normal, banning the object's account, and banning the object's functions.

16. A device for identifying abnormal objects, characterized in that, The device includes: The selection unit is used to select behavioral data of each object according to a preset time window to obtain behavioral sequence data of each object, wherein the behavioral sequence data contains at least one object behavior; The determining unit is configured to process the behavioral sequence data of the various objects to determine an object behavior similarity set, wherein the object behavior similarity set includes behavioral sequence data showing similar behavior among the various objects; and Within the time window, the object similarity between each object is determined based on the object behavior similarity set, including: aggregating the object behavior similarity sets of each object based on the object behavior similarity set, object behavior weights, and regularization factors to determine the object similarity between each object, wherein the regularization factor is determined based on the ratio of the smaller to the larger value of the number of behaviors of two objects under the same behavioral features; A composition unit is used to perform composition processing on each object based on the object similarity between the objects, so as to identify abnormal objects.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the steps of the method as described in any one of claims 1-15.

18. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, and the processor executes the steps of the method according to any one of claims 1-15 by calling the computer program stored in the memory.

19. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method described in any one of claims 1-15.

Citation Information

Patent Citations

  • Risk gang identification method and device

    CN110569509A