Abnormal group discovery method and device, computer device and storage medium
Patent Information
- Application Number
- CN202210119294.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-08
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-02-08
AI Technical Summary
这种方式考虑的因素较为单一,不能准确地发现异常团伙
[0040]上述异常团伙发现方法、装置、计算机设备、存储介质和计算机程序产品,利用第一用户集的交易关联关系和传播关系构建第一关系网络,进而对第一关系网络进行社区发现,发掘异常团伙。关系网络构建,不仅考虑了交易关联关系,还进一步考虑了用户间基于商品相关消息的传播关系,这与异常团伙间对商品信息进行传播的特性相符,使得薅羊毛、黄牛类欺诈和作弊团伙通过社交网络沟通和协作的行为被捕捉和识别,从而提高了异常团伙识别的准确度。
Smart Images

Figure CN116630076B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for detecting abnormal groups. Background Technology
[0002] With the rapid development of e-commerce, the exploitation of system loopholes for "coupon hunting" is becoming increasingly common. These "coupon hunters" or "scalpers" are always quick to spot pricing flaws, causing significant losses to merchants. In practice, these "coupon hunters" or "scalpers" often operate in groups; therefore, timely detection and control of these abnormal groups can minimize losses and reduce the impact on legitimate users.
[0003] Anomaly detection typically relies on user information gathered during transactions. However, this method considers only a limited range of factors and cannot accurately identify suspicious groups. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product that can accurately identify abnormal gangs in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides a method for detecting anomalous groups. The method includes:
[0006] Obtain the first user set;
[0007] Based on the transaction data of each first user in the first user set, determine the users in the first user set who have a first associated transaction relationship;
[0008] Based on the social information of each first user in the first user set, obtain the users in the first user set who have a first propagation relationship based on product-related messages;
[0009] Using each first user in the first user set as a node, connect users with a first associated transaction relationship and users with a first propagation relationship to construct a first relationship network;
[0010] Based on the intimacy between the first users who have connections in the first relationship network, community discovery is performed on the first relationship network to identify abnormal communities in the first relationship network.
[0011] Identify abnormal groups based on abnormal communities in the first relationship network.
[0012] Secondly, this application also provides an apparatus for detecting anomalous groups. The apparatus includes:
[0013] The acquisition module is used to acquire the first user set;
[0014] The transaction relationship determination module is used to determine, based on the transaction data of each first user in the first user set, the users in the first user set who have a first associated transaction relationship;
[0015] The propagation relationship determination module is used to obtain users in the first user set who have a first propagation relationship based on product-related messages, according to the social information of each first user in the first user set;
[0016] The construction module is used to connect users with a first associated transaction relationship and users with a first propagation relationship, using each first user in the first user set as a node, to construct a first relationship network;
[0017] The identification module is used to perform community discovery on the first relationship network based on the intimacy between each first user with a connection relationship in the first relationship network, and to identify abnormal communities in the first relationship network.
[0018] The discovery module is used to identify abnormal groups based on abnormal communities in the first relationship network.
[0019] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0020] Obtain the first user set;
[0021] Based on the transaction data of each first user in the first user set, determine the users in the first user set who have a first associated transaction relationship;
[0022] Based on the social information of each first user in the first user set, obtain the users in the first user set who have a first propagation relationship based on product-related messages;
[0023] Using each first user in the first user set as a node, connect users with a first associated transaction relationship and users with a first propagation relationship to construct a first relationship network;
[0024] Based on the intimacy between the first users who have connections in the first relationship network, community discovery is performed on the first relationship network to identify abnormal communities in the first relationship network.
[0025] Identify abnormal groups based on abnormal communities in the first relationship network.
[0026] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0027] Obtain the first user set;
[0028] Based on the transaction data of each first user in the first user set, determine the users in the first user set who have a first associated transaction relationship;
[0029] Based on the social information of each first user in the first user set, obtain the users in the first user set who have a first propagation relationship based on product-related messages;
[0030] Using each first user in the first user set as a node, connect users with a first associated transaction relationship and users with a first propagation relationship to construct a first relationship network;
[0031] Based on the intimacy between the first users who have connections in the first relationship network, community discovery is performed on the first relationship network to identify abnormal communities in the first relationship network.
[0032] Identify abnormal groups based on abnormal communities in the first relationship network.
[0033] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0034] Obtain the first user set;
[0035] Based on the transaction data of each first user in the first user set, determine the users in the first user set who have a first associated transaction relationship;
[0036] Based on the social information of each first user in the first user set, obtain the users in the first user set who have a first propagation relationship based on product-related messages;
[0037] Using each first user in the first user set as a node, connect users with a first associated transaction relationship and users with a first propagation relationship to construct a first relationship network;
[0038] Based on the intimacy between the first users who have connections in the first relationship network, community discovery is performed on the first relationship network to identify abnormal communities in the first relationship network.
[0039] Identify abnormal groups based on abnormal communities in the first relationship network.
[0040] The aforementioned methods, apparatus, computer equipment, storage media, and computer program products for detecting anomalous groups utilize the transaction and propagation relationships of a first user set to construct a first relationship network. This first relationship network is then used for community discovery to identify anomalous groups. The relationship network construction considers not only transaction relationships but also the propagation relationships between users based on product-related information. This aligns with the characteristics of anomalous groups disseminating product information, enabling the capture and identification of fraudulent groups engaging in "wool-pulling," scalping, and other forms of deception through social networks, thereby improving the accuracy of anomalous group identification. Attached Figure Description
[0041] Figure 1 This is a diagram illustrating the application environment of an abnormal gang detection method in one embodiment;
[0042] Figure 2 This is a flowchart illustrating an abnormal gang detection method in one embodiment;
[0043] Figure 3 This is a schematic diagram illustrating the propagation relationship in one embodiment;
[0044] Figure 4 This is a schematic diagram of a relationship network in one embodiment;
[0045] Figure 5 This is a schematic diagram of community discovery in one embodiment;
[0046] Figure 6 This is a flowchart illustrating the process of detecting anomaly groups in another embodiment;
[0047] Figure 7 This is a flowchart illustrating the steps involved in calculating intimacy in one embodiment;
[0048] Figure 8 This is a flowchart illustrating the community discovery steps in one embodiment;
[0049] Figure 9 This is a flowchart illustrating the process of detecting anomalies in one embodiment;
[0050] Figure 10 This is a structural block diagram of an abnormal gang detection device in one embodiment;
[0051] Figure 11 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0053] The abnormal group detection method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. The server obtains a first user set; based on the transaction data of each first user in the first user set, it determines users in the first user set with a first associated transaction relationship; based on the social information of each first user in the first user set, it obtains users in the first user set with a first propagation relationship based on product-related messages; using each first user in the first user set as a node, it connects users with the first associated transaction relationship and users with the first propagation relationship to construct a first relationship network; based on the intimacy between the first users with connections in the first relationship network, it performs community discovery on the first relationship network to identify abnormal communities in the first relationship network; and based on the abnormal communities in the first relationship network, it identifies abnormal groups. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not restricted herein.
[0054] Furthermore, all user data processing in this application requires prior authorization from relevant national authorities and must be conducted legally and compliantly. This application is implemented for the purpose of financial security, preventing financial crimes, and protecting the property of the people.
[0055] In one embodiment, such as Figure 2 As shown, a method for detecting abnormal groups is provided, which can be applied to... Figure 1 Taking the server in the example, the following steps are included:
[0056] Step 202: Obtain the first user set.
[0057] The first user set is a collection of users. It can be a randomly selected group of users from the system's user base to identify abnormal groups within that set. Alternatively, specific user extraction rules can be set to extract users from the first user set and identify abnormal groups within that specific set. For example, target user attributes can be set to extract a first user set that meets the criteria. Experience has shown that abnormal users such as those engaging in "coupon hunting," scalping, and fraud are often younger. Based on this, age criteria can be set to extract a first user set, analyzing users aged 20-40 to identify abnormal groups. Geographic location criteria can also be set to focus on analyzing users in a specific region to identify abnormal groups. Finally, product attribute criteria can be set to analyze users within a specific product category to identify abnormal groups.
[0058] Step 204: Based on the transaction data of each first user in the first user set, determine the users in the first user set who have the first associated transaction relationship.
[0059] Specifically, the order information of each user in the first user set is analyzed to extract the transaction data of each user. This order information can be all orders of a single user or orders for a specific type of product. Based on the order information, relevant transaction data is extracted, including delivery address, order time, payment information, device information, and communication network information.
[0060] The first related transaction relationship refers to the association between the first users as reflected in the transaction information. From an order perspective, each order is independent. After an order is generated, the transaction information related to the order includes the delivery address, order time, payment information, device information, and communication network information. By mining transaction information, related transaction relationships between users can be discovered. For example, orders with the same delivery information are likely to belong to the same user; orders with the same payment information are likely to belong to the same user; orders with the same device information are likely to belong to the same user; and users with orders containing the same network information tend to live in similar environments. These characteristics are consistent with the characteristics of abnormal orders in actual operations. In actual operations, it can be found that in order to maximize "coupon hunting," "coupon hunters" will use the same mobile device to register different accounts to place orders, with the same delivery address.
[0061] Specifically, by analyzing the order information of the first user, orders with the same receiving information, payment information, equipment information, and communication network information are identified as having a first related transaction relationship with the first user, and the specific type of the first related transaction relationship is recorded. Related transaction types specifically include identical receiving information, identical payment information, identical equipment information, and identical network information, etc.
[0062] Step 206: Based on the social information of each first user in the first user set, obtain the users in the first user set who have the first propagation relationship based on product-related messages.
[0063] Specifically, social information refers to interactive information existing on social media, specifically interactive information generated between first-user users. Examples include two first-user users in the same group, two first-user users chatting privately, two first-user users liking the same product, or one first-user user forwarding a product link posted by another first-user user.
[0064] The first propagation relationship refers to the relationship arising from the transmission of product-related messages between users, and it has three attributes: propagation direction, propagation method, and propagation object. Propagation direction refers to the order in which the message is transmitted from the sender to the receiver. Propagation method refers to the interactive method of message transmission, including private chat, group chat, likes, comments, etc. The propagation object includes both the message sender and the message receiver.
[0065] In e-commerce scenarios, groups of users who engage in "coupon hunting," scalpers, and fraudulent activities typically disseminate product information as quickly as possible to other members of their group, such as by sharing product links. Based on this characteristic, the first-level propagation relationships of product-related messages among the first-level users can be obtained by leveraging their social information.
[0066] Specifically, the system obtains product-related interaction information from the social information of the first user set. This includes first users disseminating product-related messages such as product links, product reviews, and product advertisements. Based on this interaction information, the system obtains the direction, method, and target audience of these product-related messages, thus establishing the initial dissemination relationships among the first users in the first user set based on these product-related messages.
[0067] For example, if user A sends a link to product X to users B, C, and D, and user B then forwards the link to user E, the first propagation relationship of the product-related message among these users based on the target attribute is as follows: Figure 3 As shown.
[0068] Among them, the first-level communication relationship has various types based on different communication methods, including private chat, group chat, likes, comments, etc.
[0069] Step 208: Using each first user in the first user set as a node, connect users with first associated transaction relationships and users with first propagation relationships to construct a first relationship network.
[0070] Specifically, each first user in the first user set is treated as a node. If any two first users have a first propagation relationship or a first associated transaction relationship, then the first users with the first propagation relationship or the first associated transaction relationship are connected by an edge. It is understandable that any two first users may simultaneously have both a first associated transaction relationship and a first propagation relationship. For example, if user A shares a product link with user B using the office network, and both users purchase the product, then user A and user B have a first propagation relationship, and also a first associated transaction relationship due to using the same communication network.
[0071] There are various types of first-propagation relationships, including private chats, group chats, likes, and comments. There are also various types of associated transactions, specifically including identical shipping information, identical payment information, identical device information, and identical network information. Therefore, when constructing the first-relationship network, the first-propagation relationship type or first associated transaction type between two specific users is used as the edge between the two users with the relationship. A schematic diagram of the relationship network of one embodiment is shown below. Figure 4 As shown, first users with a first propagation relationship or a first associated transaction relationship are connected via edges. The attributes of the edge connection include the first propagation relationship and / or the associated transaction type. For example, first user A1 and first user A3 have a group chat relationship and an associated transaction relationship based on the same device. It can be inferred that the product link is propagated from user A1 to user A3, and user A1 and user A3 use the same device to place an order for the product.
[0072] Step 210: Based on the intimacy between the first users who have a connection relationship in the first relationship network, perform community discovery on the first relationship network and identify abnormal communities in the first relationship network.
[0073] Intimacy refers to the closeness of the relationship between two first users who are connected in the first relationship network, and it is related to the relationships between the users. The more relationships two users have, the higher the intimacy.
[0074] In order to obtain accurate intimacy between users, a relationship network encompassing all users can be constructed in advance using the transaction information and social relationships of all users. This allows for the acquisition of complete relationships between users and the accurate calculation of intimacy between any two users.
[0075] Specifically, we can use intimacy to assign weights to the edges between any two first users in the first relationship network; the larger the weight, the closer the relationship.
[0076] Community discovery identifies communities by mining the degree of intimacy among users in a relationship network, grouping nodes with relatively strong internal connections into a community. Nodes within the same community are tightly connected, while connections between communities are relatively sparse. For example... Figure 5 As shown, community detection is used to divide a relationship network into three communities.
[0077] The identification process using community detection can be further refined to determine whether a community is anomalous. This can be achieved by pre-setting criteria for identifying anomalous communities. These criteria can include the number of abnormal orders; if the number of abnormal orders reaches a certain percentage, the community is identified as anomalous, or an abnormal group.
[0078] Step 212: Identify the abnormal groups based on the abnormal communities in the first relationship network.
[0079] In this embodiment, a first relationship network is constructed using the transaction data and social relationships of the first user set. Community discovery is performed on the first relationship network, and abnormal communities in the first relationship network are identified as abnormal groups.
[0080] The aforementioned method for detecting anomalous groups utilizes the transaction and propagation relationships of a first user set to construct a first relationship network. This network is then used for community discovery to identify anomalous groups. The relationship network construction considers not only transaction relationships but also the propagation relationships between users based on product-related information. This aligns with the characteristics of anomalous groups disseminating product information, enabling the capture and identification of fraudulent groups engaging in "wool-pulling," scalping, and other forms of deception through social networks, thereby improving the accuracy of anomalous group identification.
[0081] In another embodiment, the abnormal group detection method can be used to detect abnormal groups for a specific product type. Specifically, such as... Figure 6 As shown, it includes the following steps:
[0082] S602, using the transaction data of the target attribute product, obtain the first user set.
[0083] In this context, a target attribute refers to specifying the value of one attribute dimension among multiple attribute dimensions of a product. Using the target attribute, products possessing that attribute value are designated as target attribute products. Specifically, products have different attributes when analyzed from different dimensions. For example, from a category analysis perspective, products may belong to a broad category, such as baby products, household products, or digital products. A specific product category can be used as the target attribute. If the target attribute is baby products, household products, or digital products, using these categories as target attribute products allows obtaining transaction data for baby products, household products, or digital products, enabling the detection of abnormal groups within that specific product category. As another example, from a store analysis perspective, all products from the same store belong to the same store. Therefore, a specific store name can be used as the target attribute. If the target attribute is the store name, using all products under that store name as target attribute products allows obtaining transaction data for all products in that store, enabling the detection of abnormal groups within that specific store. For example, by analyzing whether a product is a best-selling product, and using best-selling products as the target attribute, we can obtain transaction data for best-selling products and discover abnormal groups selling popular products.
[0084] By leveraging target attributes, products sharing common target attributes can be extracted; typically, there are multiple products with the same target attribute. Analyzing transaction data for these target attribute products allows us to identify users who purchased them, forming a first user set. This first user set can then be used to detect unusual groups targeting specific products.
[0085] Step 604: Based on the transaction data of the target attribute products of each first user in the first user set, determine the users in the first user set who have the first associated transaction relationship.
[0086] Specifically, the first user set is the set of users who have purchased products with the target attribute. Users who have purchased products with the target attribute are designated as first users. Transaction data for the target attribute products of each first user is obtained, and this transaction data is used to determine the first associated transaction relationship among the first users. For example, if the target attribute product is a maternity or baby product, then transaction data for maternity or baby products of the first users is obtained to determine the first associated transaction relationship among the first users.
[0087] The specific implementation method of this step is the same as that of step S204, and will not be repeated here.
[0088] Step 606: Based on the social information of each first user in the first user set, obtain the users in the first user set who have a first propagation relationship based on product-related messages.
[0089] If only transactional relationships are considered without considering social relationships, the communication and contact between scalpers, including those working with smaller scalpers, those using aliases, and those operating within the black market, through social networks, cannot be detected in actual e-commerce operations. Therefore, this embodiment also considers social relationships.
[0090] Specifically, the system acquires social information from the first user set related to the target attribute product. This includes first users disseminating messages related to the target attribute product, such as product links, product reviews, and product advertisements. Based on this interaction information, the system acquires the dissemination direction, method, and target audience of these messages, thereby establishing the initial dissemination relationships among the first users in the first user set based on the target attribute product.
[0091] The specific implementation process of this step is the same as that of step S206, and will not be repeated here.
[0092] Step 608: Using each first user in the first user set as a node, connect users with first associated transaction relationships and users with first propagation relationships to construct a first relationship network.
[0093] In this embodiment, a first relationship network of target attribute products is constructed.
[0094] The specific implementation process of this step is the same as that of step S208, and will not be repeated here.
[0095] Step 610: Based on the intimacy between the first users who have a connection relationship in the first relationship network, perform community discovery on the first relationship network and identify abnormal communities in the first relationship network.
[0096] The specific implementation process of this step is the same as that of step S210, and will not be repeated here.
[0097] Step S612: Based on the abnormal communities in the first relationship network, obtain the abnormal groups corresponding to the target attribute products.
[0098] In this embodiment, transaction data of the target attribute product is used to obtain a first user set. A first relationship network is constructed using the transaction and social relationships of this first user set. Therefore, the communities defined based on this first relationship network are the communities corresponding to the target attribute product. Furthermore, the associated products of each community can be mined to pinpoint the preferences of community members. It is understood that an associated product is a specific product within the target attribute product set; it is an element of the target attribute product set. For example, if the target attribute product is a baby product, and mining reveals that most users in community A have purchased a certain type of wet wipes, then wet wipes can be considered an associated product for that community. Using the associated products of each community, the risk profile and preferences of each first user within that community can also be further identified.
[0099] It is understandable that by setting different attribute values for target products, abnormal groups under different attribute products can be obtained.
[0100] In this embodiment, a first user set is obtained by analyzing transaction data of products with target attributes. A first relationship network is constructed using the transaction and social relationships of the first user set. Community discovery is then performed on this first relationship network to identify abnormal groups corresponding to each product with a target attribute. In other words, applying this method to specific products with target attributes allows for the location of abnormal groups associated with that product, enabling precise identification of abnormal groups for different products with different target attributes. This aligns with the cheating patterns of scalpers and bargain hunters targeting specific products.
[0101] Traditional community discovery algorithms assign each customer to a specific group, a simple one-to-one relationship. However, in real-world scenarios, customers may belong to different groups, resulting in a one-to-many relationship. By constructing primary relationship networks for different target product attributes, a single user can correspond to multiple relationship networks, meaning a user might belong to multiple anomalous groups. This aligns with actual business scenarios.
[0102] In another embodiment, the method further includes a step of pre-determining the intimacy level between users. For example... Figure 7 As shown, it includes:
[0103] S702, obtain the second user set, where the first user set is a subset of the second user set.
[0104] The second user set is a larger user set than the first user set, encompassing the first user set. In practical applications, the first user set can be the user set of the target attribute product, in which case the second user set can be the entire platform's user set. Alternatively, the first user set can be the user set of a specific product, in which case the second user set is the user set of the product's category.
[0105] S704, Based on the transaction data of each second user in the second user set, determine the users in the second user set who have a second related transaction relationship.
[0106] Specifically, the implementation of this step is the same as that of step S204, and will not be repeated here.
[0107] S706, based on the social information of each second user in the second user set, obtain the users in the second user set who have a second propagation relationship based on product-related messages.
[0108] Specifically, the implementation of this step is the same as that of step S206, and will not be repeated here.
[0109] S708 uses each second user in the second user set as a node to connect users with second related transaction relationships and users with second propagation relationships, thus constructing a second relationship network.
[0110] Specifically, the implementation of this step is the same as that of step S206, and will not be repeated here.
[0111] It is worth noting that the first user set is a subset of the second user set. Therefore, the second user set includes the first user set, and the constructed second relationship network necessarily covers the first relationship network.
[0112] Compared to the first relationship network, the increased size of the user set inevitably leads to a greater variety of relationships between users. If the second user set is the entire set of users, and a second relationship network is constructed beforehand based on the propagation relationships of all users and the transaction data of all goods, then the second relationship network takes into account all the relationship types between all users.
[0113] S710, determine the intimacy between the second users based on the relationship types and weights between the second users who have a connection relationship in the second relationship network, as well as the total relationship types and weights between the two second users.
[0114] The second relationship network is constructed based on the propagation and transaction relationships of the second user set. The first user set is a subset of the second user set. Therefore, based on the second relationship network, relatively comprehensive user data can be used to calculate the intimacy between users.
[0115] There are various types of relationships between users. If we treat all relationships as a single type of edge and assign uniform weights to the edges, we will ignore all types of relationship information, leading to a deterioration in model performance.
[0116] In this embodiment, each relationship type is assigned a different weight. The weight of each relationship type can be predetermined based on experience with the device; for example, private chat sharing has a higher weight than group chat sharing, and the weight of sharing from the same device is higher than that from the same communication network. Alternatively, the weight of each relationship type can be calculated using a mean encoding method based on a second relationship network.
[0117] The intimacy between two users is calculated based on the weights of each relationship type between them. Specifically, the intimacy between the two users is determined based on the relationship types and weights between the connected users in the second relationship network, as well as the total number of relationship types and their weights between the two users.
[0118] The relationship type between second users refers to the connection type between users with a connection relationship in the second relationship network. All relationship types possessed by a second user and all relationship types diffused outwards from that user. For example, if user A and user B have a private chat propagation relationship and a related transaction relationship on the same device, then private chat propagation and related transaction relationship on the same device are the relationship types between user A and user B. If user A's total outward diffusion of relationship types includes private chat propagation, related transaction relationship on the same device, and group chat sharing relationship, then user A's total relationship types include private chat propagation, related transaction relationship on the same device, and group chat sharing relationship. If user B's total outward diffusion of relationship types includes private chat propagation, related transaction relationship on the same device, and like propagation relationship, then user B's total relationship types include private chat propagation, related transaction relationship on the same device, and like propagation relationship. The intimacy between second users is determined based on the relationship types and weights between each second user with a connection relationship in the second relationship network, and the total relationship types and weights possessed by two second users.
[0119] Specifically, the intimacy between two users is the ratio of the relationship types and weights between the two users to the total number of relationship types and weights possessed by the two users. See below for details:
[0120]
[0121] Among them, P u1→u2 :P u1→u2 ∈P refers to the relationship type between user u1 and user u2, W u1→u2 P represents the weight of the relationship type between user u1 and user u2. u1→u1 :P u1→u1 ∈P refers to all relation types possessed by user U1, W u1 P represents the weights of all relation types possessed by user U1. u2→u2 :P u2→u2 ∈P refers to all relation types possessed by user U2, W u2 The weights of all relation types that user U2 possesses.
[0122] In this embodiment, the intimacy between users is obtained based on different types of relationships between users by utilizing relationship types and their weights. By setting different risk weights for different relationship types, the risk levels of different relationship types are successfully fitted, thereby enabling accurate assessment of intimacy. Since intimacy is determined in advance using the relationship types of all users, when analyzing different scenarios for a specific first user set, the existing intimacy between users can be used directly for analysis without the need for additional model training.
[0123] In another embodiment, the method for determining the weight of each relationship type includes: randomly selecting a preset seed user from all users, calculating the proportion of abnormal seed users among the seed users to obtain a risk benchmark; starting from the abnormal seed users, traversing the relationship network to find the number of abnormal users under each relationship type, and determining the weight of each relationship based on the proportion of abnormal users under each relationship type and the risk benchmark.
[0124] In this embodiment, the seed abnormal users and abnormal users are known abnormal users identified based on existing transaction data. In this embodiment, the weight of each transaction type is determined using the mean coding method to calculate the weight of each type of relationship. For example, private chat sharing has a higher weight, while group chat sharing has a lower weight. A batch of users is randomly selected from the global user base, and the proportion of abnormal seed users is calculated as a risk benchmark. Then, starting from the seed abnormal users, based on each type of relationship, the proportion of abnormal users associated with each relationship is calculated, and the multiple of this proportion to the risk benchmark is used as the weight of each relationship type.
[0125]
[0126] Where y is the label, y=1 represents the anomaly seed label, and r represents each relation type.
[0127] In this embodiment, the mean encoding method is used to determine the weights of different relationship types, treating different relationship types differently and successfully fitting the risk levels corresponding to different relationship types. Furthermore, when calculating intimacy using the weights of relationship types, information from various relationships can be effectively utilized, improving the model's performance.
[0128] In another embodiment, such as Figure 8 As shown, based on the intimacy among the first users with connections in the first relationship network, community detection is performed on the first relationship network to identify abnormal communities in the first relationship network, including:
[0129] S802, based on the intimacy between each first user with a connection relationship in the first relationship network, perform community discovery on the first relationship network to obtain the communities in the first relationship network.
[0130] Specifically, a community detection algorithm is used to detect communities based on the intimacy between any two first users in the first relationship network. The community detection algorithm can employ connected graphs, FastUnfolding, etc. After the first round of community detection, several initially segmented communities can be obtained.
[0131] S804 judges each community based on preset feature dimensions.
[0132] In the initial segmentation of communities, some communities may have a large number of members. However, in actual business scenarios, groups are usually more concentrated, and large groups are rare. Therefore, the initial segmentation of communities may not reflect the actual situation. In this embodiment, each community can be judged based on preset feature dimensions to determine whether it is normal in each preset feature dimension.
[0133] Specifically, the preset feature dimensions include at least one of the following dimensions: the proportion of abnormal orders in the community, the proportion of abnormal seed users in the community, and the proportion of related transaction relationships in the community.
[0134] The percentage of abnormal orders in the community is the ratio of abnormal orders to all order data. Abnormal orders can be identified based on information such as abnormal delivery address, abnormal device, abnormal payment information, or abnormal communication network.
[0135] The percentage of abnormal seed users in a community is the ratio of abnormal seed users to all users in the community. Abnormal seed users are pre-determined, known abnormal users based on existing data. A high percentage of abnormal seed users indicates a greater risk of anomalies within the community.
[0136] The proportion of related-party transactions in a community is the ratio of related-party transactions to all relationship types in the community. Related-party transactions reflect abnormal relationships; if the proportion of abnormal relationships in a community is high, the community has a greater risk of abnormality.
[0137] Each preset feature dimension has established normal and abnormal standards. For example, the upper limit of the feature value corresponds to the abnormal standard, and the lower limit of the feature value corresponds to the abnormal standard. If, in a preset feature dimension, the feature value of a community is greater than the upper limit, then it meets the abnormal standard for that preset feature dimension. If, in a preset feature dimension, the feature value of a community is less than the lower limit, then it meets the normal standard for that preset feature dimension. If, in a preset feature dimension, the feature value of a community is between the upper and lower limits, then it falls between normal and abnormal for that preset feature dimension.
[0138] One or more feature dimensions can be set as evaluation indicators. If one or more feature dimensions are normal, the community is determined to be normal. If one or more feature dimensions are abnormal, the community is determined to be abnormal.
[0139] If any one or more preset feature dimensions meet the abnormality criteria, then step S806 is executed to determine the community as an abnormal community. If any one or more preset feature dimensions meet the normality criteria, then step S810 is executed to determine the community as a normal community.
[0140] If any one or more preset feature dimensions are between the normal standard and the abnormal standard, then step S808 is executed to update the community to the first relationship network, so as to perform community discovery again for the community until any one or more preset feature dimensions of the discovered community meet the abnormal standard or meet the normal standard.
[0141] In this embodiment, instead of refining communities based on their size, selection is based on the dimensions of abnormal features within each community. Communities with indistinct features are selected, and those whose feature values in any one or more feature dimensions fall between the upper and lower limits of the feature value range—that is, those communities whose feature values in any one or more feature dimensions fall between normal and abnormal—are further refined. This yields multiple medium-sized communities, and allows for the determination of whether each community is abnormal, thus improving the accuracy of community segmentation.
[0142] In a specific embodiment, such as Figure 9 As shown, the method for detecting anomalous groups includes two stages:
[0143] The first stage is the calculation of intimacy between users, and the second stage is the analysis of target attribute products to discover corresponding abnormal groups.
[0144] Specifically, such as Figure 9 As shown, the first stage includes three steps:
[0145] Step 1: Take all users as the second user set and construct the second relationship network for the second user set.
[0146] Specifically, a second user set is obtained, and the first user set is a subset of the second user set; based on the transaction data of each second user in the second user set, users in the second user set with second associated transaction relationships are identified; based on the social information of each second user in the second user set, users in the second user set with second propagation relationships based on product-related messages are obtained; using each second user in the second user set as a node, users with second associated transaction relationships and users with second propagation relationships are connected to construct a second relationship network.
[0147] Step 2: Calculate the weights of each relation type using the second relation network.
[0148] Specifically, a preset seed user is randomly selected from the second user set, and the proportion of abnormal seed users in the seed user set is calculated to obtain the risk benchmark. Starting from the abnormal seed users, the second relationship network is traversed to obtain the number of abnormal users under each relationship type. Based on the proportion of abnormal users under each relationship type and the risk benchmark, the weight of each relationship type is determined.
[0149] Step 3: Determine the intimacy between the second users based on the relationship types and weights between the second users who have connections in the second relationship network, as well as the total relationship types and weights between the two second users.
[0150] In the first stage described above, a relationship network is constructed by utilizing the social and propagation relationships of all users. The weight of each relationship type is determined using the relationship network of all users, and then the intimacy between users is calculated using the weight of each relationship.
[0151] In the second phase, the affinity calculated in the first phase is used to analyze the target attribute products and identify the corresponding abnormal groups.
[0152] Specifically, the second phase includes the following steps:
[0153] Step 1: Use the transaction data of the target attribute product to obtain the first user set and construct the first relationship network of the first user set.
[0154] Specifically, by utilizing the transaction data of target attribute products, a first user set is obtained. Based on the transaction data of the target attribute products of each first user in the first user set, users with a first associated transaction relationship in the first user set are identified. Based on the social information of each first user in the first user set, users with a first propagation relationship based on product-related messages in the first user set are obtained. Using each first user in the first user set as a node, users with a first associated transaction relationship and users with a first propagation relationship are connected to construct a first relationship network.
[0155] Understandably, the first user set is a subset of the second user set, and the second user set includes the first user set. Compared to the first relationship network, the second relationship network considers the relationships between all users, thus enabling a global measurement of user relationships and determination of intimacy.
[0156] Step 2: Based on the intimacy between any two first users in the first relationship network, perform community discovery on the first relationship network.
[0157] Specifically, based on the intimacy between the first users who have connections in the first relationship network, community discovery is performed on the first relationship network to obtain the communities in the first relationship network.
[0158] Step 3, group adjustment.
[0159] Specifically, each community is judged based on preset feature dimensions. If any one or more preset feature dimensions meet the abnormality criteria, the community is identified as an abnormal community. If any one or more preset feature dimensions are between the normal and abnormal criteria, the community is detected again until any one or more preset feature dimensions of the detected community meet the abnormality criteria or meet the normal criteria.
[0160] Step 4: Based on the abnormal communities in the first relationship network, obtain the abnormal groups corresponding to the target attribute products.
[0161] In e-commerce scenarios, information dissemination and group formation among users engaging in "coupon hunting" and scalping occur on social networks, along with the coordinated ordering and transactions by numerous multiple accounts. This application's method utilizes a social network comprised of two dimensions: social relationships and transaction data. It defines the weights of heterogeneous user relationships, then constructs a network from social network sharing within a single product category to identify different groups within each product category. By merging groups across different products, it ultimately identifies potential abnormal groups. Early intervention against abnormal customers prevents losses to platform funds and goods, and minimizes the impact on legitimate customers.
[0162] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0163] Based on the same inventive concept, this application also provides an abnormal gang detection device for implementing the abnormal gang detection method described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more abnormal gang detection device embodiments provided below can be found in the limitations of the abnormal gang detection method above, and will not be repeated here.
[0164] In one embodiment, such as Figure 10 As shown, an anomalous gang detection device is provided, comprising:
[0165] Module 1002 is used to obtain the first user set;
[0166] The transaction relationship determination module 1004 is used to determine the users in the first user set who have a first associated transaction relationship based on the transaction data of each first user in the first user set;
[0167] The propagation relationship determination module 1006 is used to obtain users in the first user set who have a first propagation relationship based on product-related messages, according to the social information of each first user in the first user set;
[0168] Module 1008 is used to connect users with first associated transaction relationships and users with first propagation relationships, with each first user in the first user set as a node, to build a first relationship network;
[0169] The identification module 1010 is used to perform community discovery on the first relationship network based on the intimacy between each first user with a connection relationship in the first relationship network, and to identify abnormal communities in the first relationship network.
[0170] The discovery module 1012 is used to identify anomalous groups based on anomalous communities in the first relationship network.
[0171] In another embodiment, the acquisition module is used to acquire a first user set using transaction data of the target attribute product. The transaction data is the transaction data of the target attribute product; the product-related messages are the related messages of the target attribute product. The discovery module is used to obtain the abnormal group corresponding to the target attribute product based on the abnormal community in the first relationship network.
[0172] In another embodiment, the acquisition module 1002 is further configured to acquire a second user set, wherein the first user set is a subset of the second user set;
[0173] The transaction relationship determination module 1004 is also used to determine the users in the second user set who have a second associated transaction relationship based on the transaction data of each second user in the second user set;
[0174] The propagation relationship determination module 1006 is also used to obtain users in the second user set who have a second propagation relationship based on product-related messages, according to the social information of each second user in the second user set;
[0175] The construction module 1008 is also used to connect users with second related transaction relationships and users with second propagation relationships, using each second user in the second user set as a node, to build a second relationship network.
[0176] It also includes an intimacy calculation module, which is used to determine the intimacy between second users based on the relationship types and weights between each second user with a connection in the second relationship network, as well as the total relationship types and weights between two second users.
[0177] In another embodiment, a weight calculation module is also included, which is used to randomly select preset seed users from the second user set, calculate the proportion of abnormal seed users among the seed users to obtain a risk benchmark; starting from the abnormal seed users, traverse the second relationship network to obtain the number of abnormal users under each relationship type, and determine the weight of each relationship type based on the proportion of abnormal users under each relationship type and the risk benchmark.
[0178] In another embodiment, the identification module is used to perform community discovery on the first relationship network based on the intimacy between each first user with a connection relationship in the first relationship network, and obtain the communities in the first relationship network; and to judge each community from preset feature dimensions. If a community meets the abnormal criteria in any one or more preset feature dimensions, then the community is determined to be an abnormal community.
[0179] In another embodiment, the identification module is further configured to perform community discovery again if a community falls between the normal standard and the abnormal standard in any one or more preset feature dimensions, until the discovered community meets the abnormal standard or meets the normal standard in any one or more preset feature dimensions.
[0180] The preset feature dimensions include at least one of the following dimensions: the proportion of abnormal orders in the community, the proportion of abnormal seed users in the community, and the proportion of related transaction relationships in the community.
[0181] The modules in the aforementioned anomalous group detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0182] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 11As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores transaction and social information data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When executed by the processor, the computer program implements a method for detecting anomalous groups.
[0183] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0184] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the abnormal gang detection methods of the above embodiments.
[0185] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the abnormal gang detection methods of the above embodiments.
[0186] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the abnormal gang detection methods of the above embodiments.
[0187] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0188] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0189] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0190] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for detecting abnormal gangs, characterized in that, The method includes: Obtain the second user set; Based on the transaction data of each second user in the second user set, identify the users in the second user set who have a second related transaction relationship; Based on the social information of each second user in the second user set, obtain the users in the second user set who have a second propagation relationship based on product-related messages; Using each second user in the second user set as a node, connect users with second related transaction relationships and users with second propagation relationships to construct a second relationship network; The intimacy between the second users is determined based on the relationship types and weights between the second users who have connections in the second relationship network, as well as the total relationship types and weights between the two second users. Obtain the first user set; the first user set is a subset of the second user set; Based on the transaction data of each first user in the first user set, determine the users in the first user set who have a first associated transaction relationship; Based on the social information of each first user in the first user set, obtain the users in the first user set who have a first propagation relationship based on product-related messages; the first propagation relationship has three attributes: propagation direction, propagation method, and propagation object. The propagation direction refers to the order in which the message sender propagates to the message receiver. The propagation method includes one or more of private chat, group chat, likes, and comments. The propagation object includes the message sender and the message receiver. Using each first user in the first user set as a node, connect users with a first associated transaction relationship and users with a first propagation relationship to construct a first relationship network; Based on the intimacy between the first users who have connections in the second relationship network, community discovery is performed on the first relationship network to identify abnormal communities in the first relationship network; Identify abnormal groups based on abnormal communities in the first relationship network.
2. The method according to claim 1, characterized in that, The acquisition of the first user set includes: Obtain the first user set using transaction data of products with target attributes; The transaction data refers to the transaction data of the target attribute product; the product-related messages refer to the related messages of the target attribute product. The step of determining abnormal groups based on abnormal communities in the first relationship network includes: obtaining abnormal groups corresponding to the target attribute product based on abnormal communities in the first relationship network.
3. The method according to claim 1, characterized in that, The methods for determining the weights of each relation type include: Randomly select preset seed users from the second user set, calculate the proportion of abnormal seed users among the seed users, and obtain the risk benchmark; Starting from the seed abnormal users, traverse the second relationship network to obtain the number of abnormal users under each relationship type. Based on the proportion of abnormal users under each relationship type and the risk benchmark, determine the weight of each relationship type.
4. The method according to claim 1, characterized in that, The step of performing community detection on the first relationship network based on the intimacy between the first users with connections in the second relationship network, and identifying abnormal communities in the first relationship network, includes: Based on the intimacy between the first users who have connections in the second relationship network, community discovery is performed on the first relationship network to obtain the communities in the first relationship network; Each community is judged based on preset feature dimensions. If a community meets the abnormal criteria in any one or more preset feature dimensions, then the community is determined to be an abnormal community.
5. The method according to claim 4, characterized in that, The method further includes: if the community falls between the normal standard and the abnormal standard in any one or more preset feature dimensions, then the community is discovered again until the discovered community meets the abnormal standard or meets the normal standard in any one or more preset feature dimensions.
6. The method according to claim 4 or 5, characterized in that, The preset feature dimensions include at least one of the following dimensions: the proportion of abnormal orders in the community, the proportion of abnormal seed users in the community, and the proportion of related transaction relationships in the community.
7. A device for detecting anomalous gangs, characterized in that, The device includes: The acquisition module is used to acquire the second user set; The transaction relationship determination module is used to determine the users in the second user set who have a second associated transaction relationship based on the transaction data of each second user in the second user set; The propagation relationship determination module is used to obtain users in the second user set who have a second propagation relationship based on product-related messages, according to the social information of each second user in the second user set; The construction module is used to connect users with second related transaction relationships and users with second propagation relationships, using each second user in the second user set as a node, to build a second relationship network; The intimacy calculation module is used to determine the intimacy between the second users based on the relationship types and weights between the second users who have a connection relationship in the second relationship network, as well as the total relationship types and weights between the two second users. The acquisition module is further configured to acquire a first user set; the first user set is a subset of the second user set; The transaction relationship determination module is further configured to determine, based on the transaction data of each first user in the first user set, users in the first user set who have a first associated transaction relationship; The propagation relationship determination module is further configured to obtain, based on the social information of each first user in the first user set, users in the first user set who have a first propagation relationship based on product-related messages; the first propagation relationship has three attributes: propagation direction, propagation method, and propagation object; the propagation direction refers to the order in which the message sender propagates to the message receiver; the propagation method includes one or more of private chat, group chat, likes, and comments; and the propagation object includes the message sender and the message receiver. The construction module is further configured to connect users with a first associated transaction relationship and users with a first propagation relationship, using each first user in the first user set as a node, to construct a first relationship network; The identification module is used to perform community discovery on the first relationship network based on the intimacy between each first user with a connection relationship in the second relationship network, and to identify abnormal communities in the first relationship network. The discovery module is used to identify abnormal groups based on abnormal communities in the first relationship network.
8. The apparatus according to claim 7, characterized in that, The acquisition of the first user set includes: Obtain the first user set using transaction data of products with target attributes; The transaction data refers to the transaction data of the target attribute product; the product-related messages refer to the related messages of the target attribute product. The step of determining abnormal groups based on abnormal communities in the first relationship network includes: obtaining abnormal groups corresponding to the target attribute product based on abnormal communities in the first relationship network.
9. The apparatus according to claim 7, characterized in that, The methods for determining the weights of each relation type include: Randomly select preset seed users from the second user set, calculate the proportion of abnormal seed users among the seed users, and obtain the risk benchmark; Starting from the seed abnormal users, traverse the second relationship network to obtain the number of abnormal users under each relationship type. Based on the proportion of abnormal users under each relationship type and the risk benchmark, determine the weight of each relationship type.
10. The apparatus according to claim 7, characterized in that, The step of performing community detection on the first relationship network based on the intimacy between the first users with connections in the second relationship network, and identifying abnormal communities in the first relationship network, includes: Based on the intimacy between the first users who have connections in the second relationship network, community discovery is performed on the first relationship network to obtain the communities in the first relationship network; Each community is judged based on preset feature dimensions. If a community meets the abnormality criteria in any one or more preset feature dimensions, then the community is determined to be an abnormal community.
11. The apparatus according to claim 10, characterized in that, The device further includes: if the community falls between the normal standard and the abnormal standard in any one or more preset feature dimensions, then the community is detected again until the detected community meets the abnormal standard or meets the normal standard in any one or more preset feature dimensions.
12. The apparatus according to claim 10 or 11, characterized in that, The preset feature dimensions include at least one of the following dimensions: the percentage of abnormal orders in the community, the percentage of abnormal seed users in the community, and the percentage of related transaction relationships in the community.
13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Abnormal user identification method and device
CN107093090A
Device identifier code and social group information-based faked sales detection method and system
CN108038696A
Community division method and system, electronic equipment and storage medium
CN113177854A