Data mining method and device, equipment and storage medium

By obtaining the environment and content characteristics of the group, building a relationship diagram and performing clustering processing, the problem of low efficiency in identifying abnormal groups is solved, and efficient identification of abnormal groups and safe maintenance of the Internet ecosystem is achieved.

CN120259011APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410024944.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively identify and prevent the formation and spread of abnormal groups in the Internet ecosystem, resulting in problems such as spam advertising.

Method used

By obtaining the environment and content characteristics of the group, a relationship diagram is constructed and clustered to identify the exception group.

Benefits of technology

It improves the recognition efficiency of abnormal groups, and can determine multiple groups as abnormal groups at the same time, enhancing the security of the Internet ecosystem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259011A_ABST
    Figure CN120259011A_ABST
Patent Text Reader

Abstract

The invention provides a data mining method and device, equipment and a storage medium, which are used for improving the efficiency of identifying abnormal groups. The method can be applied to the fields of big data, cloud technology, computer technology, security protection and the like. Comprising the steps that a group set is acquired, and groups in the group set comprise a plurality of interaction objects; environment features and content features of each group in the group set are obtained, the environment features are used for indicating group management features of the groups, and the content features are used for indicating features of published contents of interaction objects in the groups; determining an association relationship between groups in the group set based on the environment features and the content features; a relation graph is constructed according to the group set and the incidence relation, nodes of the relation graph are all groups in the group set, and edges of the relation graph are the incidence relation of all the groups in the group set; performing clustering processing on the relation graph to obtain at least one clustering result; and determining an abnormal group set according to the at least one clustering result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a data mining method, apparatus, device, and storage medium. Background Art

[0002] With the maturity and scale of the Internet ecosystem, it has become increasingly common to divert platform users from the public domain scenario to the private domain and establish connections with users in the private domain.

[0003] Private domain gameplay can convert a large amount of traffic into a stable user group and potential benefits. Therefore, some teams attract users to enter groups in free or paid forms by widely posting and spreading eye-catching content (such as pictures, videos, links, etc.) within the Internet ecosystem and leaving contact information. This may lead to the emergence of abnormal groups that post spam advertisements.

[0004] Therefore, in order to maintain the security of the Internet ecosystem, there is an urgent need for a method that can effectively identify abnormal groups. Summary of the Invention

[0005] Embodiments of this application provide a data mining method, apparatus, device, and storage medium for improving the efficiency of identifying abnormal groups.

[0006] In view of this, on the one hand, this application provides a data mining method, including: obtaining a group set, where the groups in the group set include multiple interaction objects; obtaining the environmental features and content features of each group in the group set, where the environmental features are used to indicate the group management features of the group, and the content features are used to indicate the features of the published content of the interaction objects within the group; determining the association relationship between each group in the group set based on the environmental features and the content features; constructing a relationship graph according to the group set and the association relationship, where the nodes of the relationship graph are each group in the group set, and the edges of the relationship graph are the association relationships of each group in the group set; performing clustering processing on the relationship graph to obtain at least one clustering result; and determining the abnormal group set in the group set according to the at least one clustering result.

[0007] On the other hand, this application provides a data mining apparatus, including: an obtaining module, configured to obtain a group set, where the groups in the group set include multiple interaction objects; and obtain the environmental features and content features of each group in the group set, where the environmental features are used to indicate the group management features of the group, and the content features are used to indicate the features of the published content of the interaction objects within the group.

[0008] A processing module, configured to determine the association relationships between the groups in the group set based on the environmental characteristics and the content characteristics; construct a relationship graph according to the group set and the association relationships, where the nodes of the relationship graph are the groups in the group set, and the edges of the relationship graph are the association relationships of the groups in the group set; perform clustering processing on the relationship graph to obtain at least one clustering result; determine an abnormal group set in the group set according to the at least one clustering result.

[0009] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, an acquisition module is configured to acquire a group to be processed and an abnormal group as the group set, where the abnormal group is stored in a database and is an abnormal group determined according to historical records.

[0010] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the processing module is configured to sequentially traverse the at least one clustering result to detect whether there is the abnormal group in the at least one clustering result;

[0011] Use the clustering result in which the abnormal group exists as the abnormal group set.

[0012] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the processing module is configured to sample the at least one clustering result to obtain a sampled clustering result;

[0013] Send the sampled clustering result to an artificial review platform to obtain the abnormal group set.

[0014] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the processing module is configured to acquire the historical records of the groups in the group set, where the historical records include the number of times the group is reported and the number of times the reporting rules are hit;

[0015] Traverse the at least one clustering result according to the historical records to detect whether there is a group with the historical records in the at least one clustering result;

[0016] Use the clustering result in which the group with the historical records exists as the abnormal group set.

[0017] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the acquisition module is configured to acquire the group names, group management personnel information, group owner information, image information, video information, and link information published in the groups in the group set;

[0018] Determine the environmental characteristics of the groups in the group set according to the group names, the group management personnel information, and the group owner information;

[0019] Determine the content features of the environmental features of each group in the group set according to the image information, the video information, and the link information.

[0020] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the obtaining module is configured to match the group names, group manager information, group owner information, image information, video information, and link information published in the group of the first group and the second group to obtain a matching result, where the first group and the second group are included in the group set;

[0021] When the matching result indicates that there is at least one item in common between the first group and the second group, determine that the first group and the second group have an association relationship;

[0022] Traverse each group in the group set in this way to determine the association relationships between the groups in the group set.

[0023] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the processing module is configured to perform clustering processing on the relationship graph according to a community discovery algorithm to obtain the at least one clustering result.

[0024] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the processing module is configured to calculate multiple first modularities by sequentially comparing a third group with other groups in the relationship graph;

[0025] Select two groups corresponding to the largest first modularity to form a first community, where the first community includes the third group;

[0026] Traverse and calculate in this way to obtain a first community graph, where the nodes of the first community graph are communities obtained by pairwise combination of groups;

[0027] Calculate multiple second modularities by sequentially comparing the first community with other communities in the first community graph;

[0028] Select two communities corresponding to the largest second modularity to form a first supernode, where the first supernode includes the first community;

[0029] Traverse and calculate in this way to obtain a second community graph, where the nodes of the second community graph are supernodes obtained by pairwise combination of the communities in the first community graph;

[0030] Repeat the above operations in this way until a convergence condition is reached, and output the community discovery result as the at least one clustering result.

[0031] In a possible design, in another implementation of another aspect of the embodiments of the present application, the processing module is configured to obtain the first edge weight and the sum of the first edge weights. Wherein, the first edge weight is the edge weight between the third group and the fourth group, the sum of the first edge weights is the sum of the edge weights of the groups having an association relationship with the third group, and the fourth group is included in the relationship graph;

[0032] Calculate the first modularity between the third group and the fourth group according to the first edge weight and the sum of the first edge weights;

[0033] Traverse and calculate the multiple first modularities in turn.

[0034] In a possible design, in another implementation of another aspect of the embodiments of the present application, the processing module is configured to obtain the assignments corresponding to the group name, group manager information, group owner information, image information, video information, and link information published in the group between the third group and the fourth group;

[0035] Assign values to the group name, group manager information, group owner information, image information, video information, and link information published in the group respectively to obtain an assignment set;

[0036] Calculate the weights of each assignment in the assignment set to obtain the first edge weight.

[0037] In a possible design, in another implementation of another aspect of the embodiments of the present application, the device further includes a storage module, configured to store the abnormal group level set in the database.

[0038] Another aspect of the present application provides a computer device, including: a memory, a processor, and a bus system;

[0039] Wherein, the memory is used to store programs;

[0040] The processor is configured to execute the programs in the memory, and the processor is configured to execute the methods of the above aspects according to the instructions in the program code;

[0041] The bus system is used to connect the memory and the processor to enable the memory and the processor to communicate with each other.

[0042] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is enabled to execute the methods of the above aspects.

[0043] Another aspect of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.

[0044] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages: Constraints on environmental characteristics and content characteristics of groups are performed, and then association relationships are established for each group according to the constraints and a relationship graph is constructed. Finally, clustering processing is performed on the relationship graph to obtain multiple clustering results, that is, multiple clustering results including multiple groups, and the groups in each clustering result are highly similar. When determining a clustering result as a set of abnormal groups, multiple groups can be determined as abnormal groups at the same time, thereby improving the recognition efficiency of abnormal groups. Description of the Drawings

[0045] Figure 1 It is a schematic diagram of the architecture of an application scenario of the data mining method in the embodiments of the present application;

[0046] Figure 2 It is a schematic diagram of an embodiment of the data mining method in the embodiments of the present application;

[0047] Figure 3 It is a schematic diagram of the constraint conditions of groups in the embodiments of the present application;

[0048] Figure 4 It is a schematic diagram of performing graph calculation using a community discovery algorithm in the embodiments of the present application;

[0049] Figure 5 It is another schematic diagram of performing graph calculation using a community discovery algorithm in the embodiments of the present application;

[0050] Figure 6 It is a schematic flowchart of the data mining method in the embodiments of the present application;

[0051] Figure 7 It is a schematic diagram of the effect of the data mining method in the embodiments of the present application;

[0052] Figure 8 It is a schematic diagram of an embodiment of the data mining device in the embodiments of the present application;

[0053] Figure 9 It is another schematic diagram of an embodiment of the data mining device in the embodiments of the present application;

[0054] Figure 10 It is another schematic diagram of an embodiment of the data mining device in the embodiments of the present application;

[0055] Figure 11 This is a schematic diagram of another embodiment of the data mining device in the embodiments of the present application. Detailed implementation manners

[0056] The embodiments of the present application provide a data mining method, device, equipment, and storage medium for improving the efficiency of identifying abnormal groups.

[0057] In the description and claims of the present application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0058] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.

[0059] With the maturity and scale of the Internet ecosystem, it has become increasingly common to divert platform users from the public domain scenario to the private domain and establish connections with users in the private domain. The private domain gameplay can convert a large amount of traffic into a stable user group and potential benefits. Therefore, some teams attract users to enter groups in a free or paid form by widely publishing and spreading eye-catching content (such as pictures, videos, links, etc.) within the Internet ecosystem and leaving contact information. As a result, there may be abnormal groups that publish spam advertisements. Therefore, in order to maintain the security of the Internet ecosystem, there is an urgent need for a method that can effectively identify abnormal groups.

[0060] To solve the above technical problems, the present application provides the following technical solutions: obtaining a group set, where the groups in the group set include multiple interaction objects; obtaining the environmental characteristics and content characteristics of each group in the group set, where the environmental characteristics are used to indicate the group management characteristics of the group, and the content characteristics are used to indicate the characteristics of the published content of the interaction objects in the group; determining the association relationship between each group in the group set based on the environmental characteristics and the content characteristics; constructing a relationship graph according to the group set and the association relationship, where the nodes of the relationship graph are each group in the group set, and the edges of the relationship graph are the association relationships of each group in the group set; performing clustering processing on the relationship graph to obtain at least one clustering result; and determining an abnormal group set in the group set according to the at least one clustering result. In this way, the environmental characteristics and content characteristics of the groups are constrained, then the association relationships between each group are established and a relationship graph is constructed according to the constraints, and finally multiple clustering results are obtained by performing clustering processing on the relationship graph, that is, multiple clustering results including multiple groups are obtained, and the groups in each clustering result are highly similar. When determining a clustering result as a malicious private domain set, multiple groups can be determined as abnormal groups at the same time, thereby improving the recognition efficiency of abnormal groups.

[0061] For the convenience of understanding, some terms in the present application are described below:

[0062] 1. Graph: A graph composed of several given nodes and edges connecting two points, which can be represented as G(V, E). Where G represents a graph, V is the set of vertices in graph G, and E is the set of edges in graph G. Among them, according to whether E has a direction, the graph can be divided into a directed graph and an undirected graph; and according to whether E has a weight, the graph can be divided into a weighted graph and an unweighted graph.

[0063] 2. Isomorphic network: The nodes in the graph are all of the same type. For example, the network formed by group-group in the present application, and the node types in the network are all groups.

[0064] 3. Heterogeneous network: The nodes in the graph belong to at least two different types. For example, a commercial payment network, where the node types in the network are two, namely merchants and users.

[0065] 4. Group: It can be understood as a set composed of multiple objects. In this embodiment, a group is a kind of group composed of multiple interacting users. For example, the group can be a discussion group in instant communication, etc.

[0066] A data mining method, device, equipment and storage medium provided by an embodiment of the present application are used to improve the efficiency of identifying abnormal groups. The following describes an exemplary application of the electronic device provided by an embodiment of the present application. The electronic device provided by an embodiment of the present application can be implemented as various types of user terminals or as a server.

[0067] The electronic device runs the data mining method provided by the embodiments of the present application to improve the efficiency of identifying abnormal groups.

[0068] The above solution can be applied to many Internet security protection fields. When using the data mining method provided by the embodiments of the present application to help users mine abnormal groups, this method can be implemented as an independent online application and installed in the computer device or the background server used by the user, facilitating the user to use this program to mine abnormal groups.

[0069] In this scenario, obtain multiple groups that are still running, and obtain the environmental characteristics and content characteristics of the groups; determine whether there is an association relationship between each group based on the environmental characteristics and content characteristics; then construct a relationship graph based on the association relationship and the multiple groups; finally, perform clustering processing according to this relationship graph to obtain at least one clustering result; finally, determine the set of abnormal groups in the multiple groups according to this at least one clustering result.

[0070] In an exemplary solution, the data mining method can be applied to a game scenario. For example, obtain multiple discussion groups in the game, perform abnormal identification on the multiple discussion groups (such as whether reporting positions and cheating or colluding to cheat, etc.), and finally notify the identification result to the game server so that the game server performs corresponding processing on the discussion group.

[0071] Of course, in addition to being applied to the above scenarios, the method provided by the embodiments of the present application can also be applied to other scenarios that require abnormal group identification. The embodiments of the present application do not limit the specific application scenarios.

[0072] See Figure 1 , Figure 1It is an optional architecture schematic diagram in an application scenario of the data mining solution provided by the embodiments of the present application. To implement a data mining solution, the terminal device 100 is connected to the server 300 through the network 200, and the server 300 is connected to the database 400. The network 200 can be a wide area network, a local area network, or a combination of the two. The client for implementing the data mining solution is deployed on the terminal device 100. Among them, the client can run on the terminal device 100 in the form of a browser, or can also run on the terminal device 100 in the form of an independent application (APP), etc. The specific presentation form of the client is not limited here. The server 300 involved in the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device 100 can be a smart phone, a tablet computer, a notebook computer, a handheld computer, a personal computer, a smart TV, a smart watch, a vehicle-mounted device, a wearable device, a smart voice interaction device, a smart home appliance, an aircraft, etc., but is not limited thereto. The terminal device 100 and the server 300 can be directly or indirectly connected through the network 200 by wired or wireless communication methods, and the present application does not limit this here. The number of the server 300 and the terminal device 100 is also not limited. The solution provided by the present application can be completed independently by the terminal device 100, or can be completed independently by the server 300, or can also be completed by the cooperation of the terminal device 100 and the server 300. In this regard, the present application does not make specific limitations. Among them, the database 400 can be regarded as an electronic filing cabinet in short - a place for storing electronic files, and users can perform operations such as adding, querying, updating, and deleting data in the files. The so - called "database" is a data set stored together in a certain way, can be shared by multiple users, has the smallest possible redundancy, and is independent of application programs. The Database Management System (DBMS) is a computer software system designed to manage the database, and generally has basic functions such as storage, interception, security guarantee, backup, etc.Database management systems can be classified according to the database models they support, such as relational, Extensible Markup Language (XML); or according to the types of computers they support, such as server clusters, mobile phones; or according to the query languages used, such as Structured Query Language (SQL), XQuery; or according to the performance impulse focus, such as maximum scale, highest running speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories. For example, they can support multiple query languages simultaneously. In this application, the database 400 can be used to store the determined abnormal groups. Of course, the storage location of the determined abnormal groups is not limited to the database. For example, it can also be stored in the terminal device 100, blockchain, or the distributed file system of the server 300, etc.

[0073] It can be understood that the data mining method provided in this application can be executed by the server 300; correspondingly, the data mining device can be set in the server 300. Of course, the data mining method provided in this application can also be executed by the terminal device 100; correspondingly, the data mining device can be set in the terminal device 100. Similarly, the data mining method provided in this application can also be jointly executed by the terminal device 100 and the server 300; correspondingly, the data mining device can be set in the terminal device 100 and the server 300. The embodiments of this application do not impose any restrictions on the device form for executing the data mining method.

[0074] It can be understood that in the specific implementation of this application, it involves data such as the published content of the group, group management personnel information, group owner information, and group name. When the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0075] Combined with the above introduction, the data mining method in this application will be introduced below with the server as the execution entity. Please refer to Figure 2 , an embodiment of the data mining method in the embodiments of this application includes:

[0076] 201. Obtain a group set, where the groups in the group set include multiple interaction objects.

[0077] In this embodiment, the data mining method can be applied to one application software or multiple application software simultaneously. That is, the data mining method can be integrated into an identification function module of an application software; or it can be an independent third-party software that authorizes and manages at least one application software at the same time.

[0078] The following takes the application of this data mining method to an application software, and the execution entity that executes this data mining method is the background server of this application software as an example for illustration:

[0079] The server obtains multiple groups corresponding to the application software as the group set.

[0080] Optionally, the server may obtain the group set at a preset period, and the groups in the group set may be all the groups running within the preset period or the groups filtered according to corresponding conditions. For example, setting 12 hours as a period, assuming that within this time period, there are groups 1, 2, 3, and 4 that have not been audited or new groups 5 and 6 are generated within this time period, then the group set will include groups 1, 2, 3, 4, 5, and 6.

[0081] Optionally, the group set may also introduce groups that have been determined to be abnormal groups, which can facilitate the subsequent determination of the abnormal group set.

[0082] 202. Obtain the environmental characteristics and content characteristics of each group in the group set. The environmental characteristics are used to indicate the group management characteristics of the group, and the content characteristics are used to indicate the characteristics of the published content of the interaction objects within the group.

[0083] After obtaining the group set, the server obtains the environmental characteristics and content characteristics of each group, and uses the environmental characteristics and content characteristics as the constraint conditions of the group.

[0084] As Figure 3 shown, the environmental characteristics can be understood as the group management characteristics of the group, such as the group names, group management personnel information, group member information, and group owner information of each group.

[0085] The content characteristics can be understood as the content characteristics of the content published by each group member (i.e., the interaction objects in this embodiment) within the group. For example, the text information, image information, video information, link information, and domain name information published by the group owner in each group. For example, text, images, or videos containing illegal content; link or domain name information that redirects to illegal websites, etc.

[0086] 203. Determine the association relationship between each group in the group set based on the environmental characteristics and the content characteristics.

[0087] After obtaining the environmental characteristics and the content characteristics, the server matches and compares the environmental characteristics and content characteristics in each group. If there is at least one same, it is determined that there is an association relationship between the two groups.

[0088] In an exemplary solution, there are Group A, Group B, and Group C. The group name of Group A is "Shopping Species" ", the group manager is "Xiaohong", the group owner is "Xiaohong", and the group owner "Xiaohong" has posted purchase links for Product A, Product B, Product C, and Product D; while the group name of Group B is "Shopping Species" ", the group manager is "Xiaoming", the group owner is "Xiaohong", and the group manager "Xiaoming" has posted purchase links for Product C and Product D, as well as Picture A; the group name of Group A is "Patent Group", the group manager is "Xiaogang", the group owner is "Xiaobai", and the group owner "Xiaobai" has posted links to Patent Article A and Patent Article B. According to the above rules, it can be determined that Group A and Group B have an association relationship, while Group C has no association relationship with either Group A or Group B.

[0089] 204. Construct a relationship graph based on the group set and the association relationship. The nodes of the relationship graph are the various groups in the group set, and the edges of the relationship graph are the association relationships of the various groups in the group set.

[0090] In this embodiment, the server constructs the various groups in the group set into a relationship graph according to the association relationships of the various groups in the group set. The nodes of the relationship graph are the various groups in the group set, and the edges of the relationship graph are the association relationships of the various groups in the group set.

[0091] In an exemplary solution, there are Group A, Group B, Group C, Group D, Group E, and Group F. Among them, Group A has an association relationship with Group B and Group C; Group B has an association relationship with Group A and Group C; Group C has an association relationship with Group A, Group B, and Group D; Group D has an association relationship with Group F and Group C; Group E has an association relationship with Group A and Group F; Group F has an association relationship with Group D and Group E. Based on the above description, a relationship graph as shown in Figure 4 can be constructed and generated. Each edge is used to indicate that two groups have an association relationship.

[0092] Among them, the weight value of this edge can be calculated according to the above-mentioned environmental features and content features. In an exemplary solution, values are assigned to each type of environmental feature and each type of content feature, and then weight summation or direct numerical summation or other weight processing is performed based on the assigned values to obtain the evaluation value of this association relationship, and this evaluation value is used as the weight value of the edge of this relationship graph. For example, when the group names are the same, the assigned value is 1, and when they are different, the assigned value is 0; when the group manager information is the same, the assigned value is 1, and when they are different, the assigned value is 0; when the group member information is the same, the assigned value is 1, and when they are different, the assigned value is 0; when the group owner information is the same, the assigned value is 1, and when they are different, the assigned value is 0; when the text information is the same, the assigned value is 1, and when they are different, the assigned value is 0; when the image information is the same, the assigned value is 1, and when they are different, the assigned value is 0; when the video information is the same, the assigned value is 1, and when they are different, the assigned value is 0; when the link information is the same, the assigned value is 1, and when they are different, the assigned value is 0; and when the domain name information is the same, the assigned value is 1, and when they are different, the assigned value is 0; then numerical summation is performed according to the specific situation to obtain the evaluation value of the association relationship between the two groups, and this evaluation value is used as the weight value of the edge of this relationship graph.

[0093] As Figure 4 shown, through the above calculation rules, the weight values of the edges between each group can be obtained. For example, the weight value between group A and group B is 5; the weight value between group C and group B is 2; the weight value between group A and group C is 3; the weight value between group C and group D is 7; the weight value between group A and group E is 1; the weight value between group E and group F is 8; the weight value between group D and group F is 7.

[0094] 205. Perform clustering processing on this relationship graph to obtain at least one clustering result.

[0095] In this embodiment, after constructing this relationship graph, a community discovery algorithm can be used to perform community mining on this relationship graph to obtain at least one clustering result of this relationship graph.

[0096] In this embodiment, this community discovery algorithm can be the louvain algorithm or other community discovery algorithms. There is no specific limitation here as long as clustering processing can be achieved.

[0097] Next, taking the louvain algorithm as an example, a solution for performing clustering processing on this relationship graph to obtain at least one clustering result will be described:

[0098] In the louvain algorithm, modularity is usually used as an indicator to measure whether a community partition is reasonable, that is, the similarity of nodes within the same community is relatively high, while the similarity of nodes in different communities is relatively low. Among them, the definition of this modularity can be as follows:

[0099]

[0100] Among them, A ij represents the weight value of the edge between node i and node j, and K i represents the sum of the weight values of all the edges connected to node i, and K j represents the sum of the weight values of all the edges connected to node j, and c i represents the community to which node i belongs, and c j represents the community to which node j belongs, represents whether they belong to the same community, It should be understood that the definition of this modularity can also have other forms, which are not specifically limited here, as long as the similarity between two nodes can be calculated.

[0101] Therefore, the modularity gain of each node (i.e., group) in this relational graph can be iteratively calculated through the modularity until the iteration termination condition is reached, so as to obtain the at least one clustering result. Based on the definition of this modularity, the modularity increment can be used as an iteration termination condition and a judgment condition for community division. Among them, the definition of this modularity increment can be as follows:

[0102]

[0103] Among them, a is used to indicate the node, C is used to indicate the community, ∑in represents the sum of all edge weight values in community C, and K a,in represents the sum of the weight values of all the edges from node a to community C, and ∑tot represents the sum of the weight values of all the edges connected to community C.

[0104] Based on the above description, in this embodiment, the specific process of this community discovery algorithm can be as follows: The third group is successively calculated with other groups in this relational graph to obtain multiple first modularities; The two groups corresponding to the largest first modularity are selected to form the first community, and this first community includes this third group; Traverse and calculate in turn to obtain the first community graph, and the nodes of this first community graph are the communities obtained by combining groups in pairs; The first community is successively calculated with other communities in this first community graph to obtain multiple second modularities; The two communities corresponding to the largest second modularity are selected to form the first supernode, and this first supernode includes this first community; Traverse and calculate in turn to obtain the second community graph, and the nodes of this second community graph are the supernodes obtained by combining groups in pairs; Repeat the above operations in turn until the convergence condition is reached, and output the community discovery result as the at least one clustering result.

[0105] In an exemplary solution, take Figure 4Taking the shown figure as an example, among them, Group A has an association relationship with Group B and Group C; Group B has an association relationship with Group A and Group C; Group C has an association relationship with Group A, Group B and Group D; Group D has an association relationship with Group F and Group C; Group E has an association relationship with Group A and Group F; Group F has an association relationship with Group D and Group E. Among them, the weight value between Group A and Group B is 5; the weight value between Group C and Group B is 2; the weight value between Group A and Group C is 3; the weight value between Group C and Group D is 7; the weight value between Group A and Group E is 1; the weight value between Group E and Group F is 8; the weight value between Group D and Group F is 7.

[0106] At this time, set as the sum of the weights of the edges of the graph, as a fixed value of 1, then the following modularity set can be calculated:

[0107]

[0108] From the above calculation results, it can be seen that the modularity between Group A and Group B is the largest. Therefore, in the first round, Group A and Group B can be divided into one community.

[0109] Repeat the calculation in this way Figure 4 the modularity between each pair of groups in Figure 5 As shown, that is, Group A and Group B are divided into one community (assumed to be community O), Group C and Group D are divided into one community (assumed to be community R), and Group E and Group F are divided into one community (assumed to be community Y).

[0110] Then in the second-round calculation process, its modularity can be calculated to obtain the following results:

[0111]

[0112] From the above calculation results and the rule of community division when it is greater than 0 and not dividing when it is less than or equal to 0, it can be seen that community division cannot be carried out in the second round, and the three communities obtained by the first-round division can be used as the clustering results for output.

[0113] In this embodiment, in order to ensure the performance of the community discovery algorithm, the clustering results can be manually evaluated periodically. In one possible implementation, for model manual evaluation, this scheme randomly selects N (N = 10) communities and calculates their clustering purity, and the formula is defined as follows:

[0114]

[0115] Among them, N represents the total number of samples, ω represents the clusters obtained by the community discovery algorithm, and c represents the true clusters. The value range of P is [0, 1], and the larger the value, the better the clustering effect.

[0116] 206. Determine the abnormal group set in the group set according to the at least one clustering result.

[0117] After determining the at least one clustering result corresponding to the group set, determine the abnormal group set in the group set from the at least one clustering result according to the corresponding judgment rules.

[0118] In one implementation solution, an abnormal group is introduced into the group set. It can be determined whether the abnormal group exists in the clustering result. If it exists, the clustering result can be directly determined as the malicious private domain set. For example Figure 4 in, group A is marked as an abnormal group, then the communities to which group A and group B belong can be directly determined as the malicious private domain set. In this embodiment, after determining the malicious private domain set, the malicious private domain set can be reported and corresponding processing can be performed. At the same time, the groups in the malicious private domain set can be added to the database as the determined abnormal groups.

[0119] In another implementation solution, send the at least one clustering result to an artificial review platform, and judge the at least one clustering result through the method of artificial review to determine the malicious private domain set in the group set from the at least one clustering result. In order to save labor costs, the at least one clustering result can be sampled according to the modularity value of the community (that is, select the clustering results with the top N modularity rankings), and then the sampled clustering results are sent to the artificial review platform.

[0120] In another implementation manner, obtain the historical records of each group in the group set. At this time, the historical records are used to indicate information such as whether the group has been reported, the number of reports, whether the reporting rules are hit, the number of times the reporting rules are hit, and the number of reporting rules hit, etc.; then detect whether each group in the at least one clustering result has the above historical records, and use the clustering result to which the group with the above historical records belongs as the malicious private domain set. For example, Figure 4 in, group A has a historical record of being reported, then the communities to which group A and group B belong can be directly determined as the malicious private domain set.

[0121] Based on this determination scheme, the malicious private domain set can also be determined according to the number of historical records. For example, Figure 4Among them, Group A has a history of being reported, and the number of records is 4 times; Group F has a history of being reported and a history of hitting the reporting rules, and the number of records is 8 times, while the threshold of the number of records is 5. Then the communities to which Group E and Group F belong can be determined as malicious private domain sets, and the communities to which Group A and Group B belong can be determined as candidate malicious private domain sets, which can be sent to the manual review platform for re-review.

[0122] Based on the above introduction, the data mining method of this application will be described below with the Figure 6 flowchart shown:

[0123] First, introduce the abnormal groups stored in the database into the group set to be identified. It should be understood that the abnormal groups are the identified abnormal groups. Then, construct a relationship graph according to the environmental characteristics and content characteristics of each group in the group set to be identified; then use the community discovery algorithm to obtain at least one clustering result of the relationship graph; finally, determine the malicious private domain set according to at least one clustering result.

[0124] In order to ensure the performance of the community discovery algorithm, the community discovery algorithm can be evaluated by an artificial model. When the model evaluation passes, continue to use the community discovery algorithm. If it does not pass, re-design the community discovery algorithm.

[0125] After using the data mining method provided by this application, the schematic curve diagram of manual quality inspection of abnormal group drainage can be as Figure 7 shown. From the curve trend, it can be seen that the proportion of manual quality inspection of abnormal group drainage has decreased significantly, and the scale of abnormal group drainage has been curbed.

[0126] The data mining device in this application will be described in detail below. Please refer to Figure 8 , Figure 8 which is a schematic diagram of an embodiment of the data mining device in an embodiment of this application. The data mining device 20 includes:

[0127] An acquisition module 201, configured to acquire a group set, where the groups in the group set include multiple interaction objects; acquire the environmental characteristics and content characteristics of each group in the group set, where the environmental characteristics are used to indicate the group management characteristics of the group, and the content characteristics are used to indicate the characteristics of the content published by the interaction objects in the group;

[0128] A processing module 202 is configured to determine the association relationships between the groups in the group set based on the environmental features and the content features; construct a relationship graph according to the group set and the association relationships, where the nodes of the relationship graph are the groups in the group set, and the edges of the relationship graph are the association relationships between the groups in the group set; perform clustering processing on the relationship graph to obtain at least one clustering result; and determine an abnormal group set in the group set according to the at least one clustering result.

[0129] In an embodiment of the present application, a data mining device is provided. By using the above device, constraints on environmental features and content features are imposed on groups, and then association relationships are established for each group according to the constraints and a relationship graph is constructed. Finally, clustering processing is performed on the relationship graph to obtain multiple clustering results, that is, multiple clustering results including multiple groups, and the groups in each clustering result are highly similar. When determining a clustering result as an abnormal group set, multiple groups can be determined as abnormal groups at the same time, thereby improving the recognition efficiency of abnormal groups.

[0130] Optionally, based on the corresponding embodiment above, Figure 8 in another embodiment of the data mining device 20 provided in the embodiment of the present application,

[0131] An acquisition module 201 is configured to acquire a group to be processed and an abnormal group as the group set, where the abnormal group is stored in a database and is an abnormal group determined according to historical records.

[0132] In an embodiment of the present application, a data mining device is provided. By using the above device, the groups determined as abnormal groups according to historical records are added as black seeds to the group set to be recognized. In this way, when judging the abnormal group set according to the clustering result subsequently, the abnormal group set can be recognized more quickly, thereby further improving the recognition efficiency of abnormal groups.

[0133] Optionally, based on the corresponding embodiment above, Figure 8 in another embodiment of the data mining device 20 provided in the embodiment of the present application, the processing module 202 is configured to sequentially traverse the at least one clustering result to detect whether the abnormal group exists in the at least one clustering result;

[0134] Use the clustering result in which the abnormal group exists as the abnormal group set.

[0135] In an embodiment of the present application, a data mining device is provided. By using the above device, the groups determined as abnormal groups according to historical records are added as black seeds to the group set to be recognized. In this way, when judging the abnormal group set according to the clustering result subsequently, the abnormal group set can be recognized more quickly, thereby further improving the recognition efficiency of abnormal groups.

[0136] Optionally, based on the corresponding embodiments described above, in another embodiment of the data mining device 20 provided by the embodiments of the present application, Figure 8 the processing module 202 is configured to sample the at least one clustering result to obtain a sampled clustering result;

[0137] The sampled clustering result is sent to an artificial review platform to obtain the abnormal group set.

[0138] In the embodiments of the present application, a data mining device is provided. By using the above device, multiple clustering results are reviewed through artificial review, thus ensuring the accuracy of the determination result.

[0139] Optionally, based on the corresponding embodiments described above, in another embodiment of the data mining device 20 provided by the embodiments of the present application,

[0140] the processing module 202 is configured to obtain the historical records of each group in the group set, where the historical records include the number of times the group is reported and the number of times the reporting rules are hit; Figure 8 Based on the corresponding embodiments described above, in another embodiment of the data mining device 20 provided by the embodiments of the present application,

[0141] Traverse the at least one clustering result according to the historical records to detect whether there is a group with the historical records in the at least one clustering result;

[0142] The clustering result in which there is a group with the historical records is used as the abnormal group set.

[0143] In the embodiments of the present application, a data mining device is provided. By using the above device, according to the historical records of each group, it is determined whether the clustering result to which it belongs is an abnormal group set, which can reduce labor costs and improve the recognition efficiency of the abnormal group set at the same time.

[0144] Optionally, based on the corresponding embodiments described above, in another embodiment of the data mining device 20 provided by the embodiments of the present application, the obtaining module 201 is configured to obtain the group name, group manager information, group owner information, image information, video information, and link information published in the group for each group in the group set;

[0145] Based on the group name, the group manager information, and the group owner information, determine the environmental characteristics of each group in the group set; Figure 8 Based on the corresponding embodiments described above, in another embodiment of the data mining device 20 provided by the embodiments of the present application, the obtaining module 201 is configured to obtain the group name, group manager information, group owner information, image information, video information, and link information published in the group for each group in the group set;

[0146] Based on the group name, the group manager information, and the group owner information, determine the environmental characteristics of each group in the group set;

[0147] Based on the image information, the video information, and the link information, determine the content characteristics of each group in the group set.

[0148] In an embodiment of the present application, a data mining device is provided. By using the above device, environmental constraints and content constraints are imposed on the group based on a specific environmental feature and content feature, so as to achieve a portrait description of the group, thus providing clustering conditions for subsequent clustering processing, obtaining multiple clustering results including multiple groups, and the groups in each clustering result are highly similar. When determining a clustering result as a set of abnormal groups, multiple groups can be determined as abnormal groups simultaneously, thereby improving the recognition efficiency of abnormal groups.

[0149] Optionally, based on the corresponding embodiment above, in another embodiment of the data mining device 20 provided in the embodiment of the present application, the obtaining module 201 is configured to match the group names, group manager information, group owner information, image information, video information, and link information published in the group between the first group and the second group to obtain a matching result, where the first group and the second group are included in the group set; Figure 8 When the matching result indicates that there is at least one same item between the first group and the second group, it is determined that the first group and the second group have an association relationship;

[0150] Traverse each group in the group set in this way to determine the association relationships between the groups in the group set.

[0151] In an embodiment of the present application, a data mining device is provided. By using the above device, environmental constraints and content constraints are imposed on the group based on a specific environmental feature and content feature, so as to achieve a portrait description of the group, thus providing clustering conditions for subsequent clustering processing, obtaining multiple clustering results including multiple groups, and the groups in each clustering result are highly similar. When determining a clustering result as a set of abnormal groups, multiple groups can be determined as abnormal groups simultaneously, thereby improving the recognition efficiency of abnormal groups.

[0152] Optionally, based on the corresponding embodiment above, in another embodiment of the data mining device 20 provided in the embodiment of the present application,

[0153] The processing module 202 is configured to perform clustering processing on the relationship graph according to the community discovery algorithm to obtain the at least one clustering result. Figure 8 In an embodiment of the present application, a data mining device is provided. By using the above device, the community discovery algorithm is used to perform clustering processing on the constructed relationship graph, so that the solution has a lower time complexity, thereby enabling the solution to be applicable to group scenarios with a large amount of data, and further improving the generalization of the solution.

[0154] Optionally, based on the corresponding embodiment above,

[0155]

[0156] Figure 8 ​​Based on the corresponding embodiments, in another embodiment of the data mining device 20 provided by the embodiments of the present application,

[0157] The processing module 202 is configured to calculate a plurality of first modularities by successively calculating the third group with other groups in the relationship graph;

[0158] Select two groups corresponding to the largest first modularity to form a first community, and the first community includes the third group;

[0159] Perform traversal calculations in this way to obtain a first community graph, and the nodes of the first community graph are communities obtained by pairwise combination of groups;

[0160] Calculate a plurality of second modularities by successively calculating the first community with other communities in the first community graph;

[0161] Select two communities corresponding to the largest second modularity to form a first supernode, and the first supernode includes the first community;

[0162] Perform traversal calculations in this way to obtain a second community graph, and the nodes of the second community graph are supernodes obtained by pairwise combination of communities in the first community graph;

[0163] Repeat the above operations in this way until the convergence condition is reached, and output the community discovery result as the at least one clustering result.

[0164] In the embodiments of the present application, a data mining device is provided. By using the above device, the community discovery algorithm is used for clustering processing on the constructed relationship graph, so that the solution has a lower time complexity, so that the solution can be applied to group scenarios with a large amount of data, thereby improving the generalization of the solution.

[0165] Optionally, based on the corresponding embodiments above, in another embodiment of the data mining device 20 provided by the embodiments of the present application, Figure 8 Based on the corresponding embodiments above, in another embodiment of the data mining device 20 provided by the embodiments of the present application,

[0166] The processing module 202 is configured to obtain a first edge weight and a sum of first edge weights, where the first edge weight is the edge weight between the third group and the fourth group, and the sum of the first edge weights is the sum of the edge weights of the groups having an associated relationship with the third group, and the fourth group is included in the relationship graph;

[0167] Calculate a first modularity between the third group and the fourth group according to the first edge weight and the sum of the first edge weights;

[0168] Perform traversal calculations in this way to obtain the plurality of first modularities.

[0169] In an embodiment of the present application, a data mining device is provided. By using the above device, the constructed relationship graph is clustered using a community discovery algorithm, and the rationality of the community division is judged by modularity. In this way, the solution has a lower time complexity and higher accuracy, so that the solution can be applied to group scenarios with a large amount of data, thereby improving the generalization of the solution.

[0170] Optionally, based on the corresponding embodiment above, Figure 8 in another embodiment of the data mining device 20 provided in the embodiment of the present application,

[0171] The processing module 202 is configured to obtain the assignments corresponding to the group name, group manager information, group owner information, image information, video information, and link information published in the group between the third group and the fourth group;

[0172] Assign values to the group name, the group manager information, the group owner information, the image information, the video information, and the link information published in the group respectively to obtain an assignment set;

[0173] Calculate the weights of each assignment in the assignment set to obtain the first edge weight.

[0174] In an embodiment of the present application, a data mining device is provided. By using the above device, the group is environmentally constrained and content-constrained with a specific environmental feature and content feature, so as to achieve a portrait description of the group. In this way, clustering conditions are provided for subsequent clustering processing, and multiple clustering results including multiple groups are obtained, and the groups in each clustering result are highly similar. When determining a clustering result as an abnormal group set, multiple groups can be determined as abnormal groups at the same time, thereby improving the recognition efficiency of abnormal groups.

[0175] Optionally, based on the corresponding embodiment above, Figure 8 as Figure 9 shown, in another embodiment of the data mining device 20 provided in the embodiment of the present application, the device further includes a storage module 203 for storing the abnormal group set in the database.

[0176] In an embodiment of the present application, a data mining device is provided. By using the above device, the abnormal group database is updated according to the recognition result, and then the abnormal group set is determined according to the black seeds in the abnormal group database. In this way, the number of black seeds can be increased, and the determination conditions for determining the abnormal group set can be increased, thereby improving the recognition efficiency of abnormal groups.

[0177] The data mining device provided by the present application may be a server. Please refer to Figure 10 , Figure 10It is a schematic diagram of a server structure provided by an embodiment of the present application. The server 300 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 can be transient storage or persistent storage. The programs stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.

[0178] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0179] The steps performed by the data mining device in the above embodiments may be based on the Figure 10 shown server structure.

[0180] The data mining device provided by the present application may be a terminal device. Please refer to Figure 11 , for the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. In the embodiments of the present application, a smart phone is taken as an example of the terminal device for illustration:

[0181] Figure 11 What is shown is a block diagram of a part of the structure of a smart phone related to the terminal device provided by the embodiments of the present application. Referring to Figure 11 , the smart phone includes: a radio frequency (RF) circuit 410, a memory 420, an input unit 430, a display unit 440, a sensor 450, an audio circuit 460, a wireless fidelity (WiFi) module 470, a processor 480, and a power supply 490 and other components. Those skilled in the art can understand, Figure 11The smartphone structure shown does not constitute a limitation on the smartphone, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0182] The following will specifically introduce each component of the smartphone in conjunction with Figure 11 :

[0183] The RF circuit 410 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is given to the processor 480 for processing; in addition, it sends the designed uplink data to the base station. Generally, the RF circuit 410 includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 410 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0184] The memory 420 can be used to store software programs and modules. The processor 480 executes various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 420. The memory 420 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, applications required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the smartphone (such as audio data, phone book, etc.). In addition, the memory 420 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0185] The input unit 430 can be used to receive input numeric or character information, and generate key signal inputs related to the user settings and function controls of the smart phone. Specifically, the input unit 430 can include a touch panel 431 and other input devices 432. The touch panel 431, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 431), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 431 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 480, and can receive and execute commands sent by the processor 480. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 431. In addition to the touch panel 431, the input unit 430 can also include other input devices 432. Specifically, the other input devices 432 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0186] The display unit 440 can be used to display information input by the user or information provided to the user and various menus of the smart phone. The display unit 440 can include a display panel 441. Optionally, the display panel 441 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 431 can cover the display panel 441. When the touch panel 431 detects a touch operation thereon or nearby, it is transmitted to the processor 480 to determine the type of touch event. Subsequently, the processor 480 provides corresponding visual output on the display panel 441 according to the type of touch event. Although in Figure 11 the touch panel 431 and the display panel 441 are implemented as two independent components to realize the input and input functions of the smart phone, in some embodiments, the touch panel 431 and the display panel 441 can be integrated to realize the input and output functions of the smart phone.

[0187] The smart phone may further include at least one sensor 450, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 441 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 441 and / or the backlight when the smart phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used in applications for identifying the posture of the smart phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the smart phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0188] The audio circuit 460, the speaker 461, and the microphone 462 can provide an audio interface between the user and the smart phone. The audio circuit 460 can transmit the electrical signal converted from the received audio data to the speaker 461, and the speaker 461 converts it into a sound signal for output; on the other hand, the microphone 462 converts the collected sound signal into an electrical signal, which is received by the audio circuit 460 and then converted into audio data. After the audio data is output to the processor 480 for processing, it is sent through the RF circuit 410 to, for example, another smart phone, or the audio data is output to the memory 420 for further processing.

[0189] WiFi belongs to short - range wireless transmission technology. The smart phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 470, which provides users with wireless broadband Internet access. Although Figure 11 the WiFi module 470 is shown, it can be understood that it does not belong to an essential component of the smart phone and can be omitted entirely within the scope of not changing the essence of the invention according to needs.

[0190] The processor 480 is the control center of the smart phone, connecting various parts of the entire smart phone through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 420, and by calling the data stored in the memory 420, it executes various functions of the smart phone and processes data, thereby monitoring the smart phone as a whole. Optionally, the processor 480 may include one or more processing units; optionally, the processor 480 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 480 either.

[0191] The smart phone further includes a power supply 490 (such as a battery) for powering each component. Optionally, the power supply can be logically connected to the processor 480 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.

[0192] Although not shown, the smart phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0193] The steps performed by the data mining device in the above embodiments may be based on the Figure 11 shown terminal device structure.

[0194] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0195] An embodiment of the present application further provides a computer program product including a program. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0196] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.

[0197] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces, indirect couplings, or communication connections of devices or units, which can be electrical, mechanical, or other forms.

[0198] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0199] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0200] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0201] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A data mining method, characterized in that, Including: Obtain a group set, where the groups in the group set include multiple interaction objects; Obtain the environmental characteristics and content characteristics of each group in the group set, where the environmental characteristics are used to indicate the group management characteristics of the group, and the content characteristics are used to indicate the characteristics of the group publishing content of the interaction objects within the group; Determine the association relationship between each interaction role in the group set based on the environmental characteristics and the content characteristics; Construct a relationship graph according to the group set and the association relationship, where the nodes of the relationship graph community graph are each group in the group set, and the edges of the relationship graph are the association relationships of each group in the group set; Perform clustering processing on the relationship graph community graph to obtain at least one clustering result; Determine an abnormal group set in the group set according to the at least one clustering result; 2. The method according to claim 1, wherein Obtaining the group set includes: Obtain the to-be-processed groups and abnormal groups as the group set, where the abnormal groups are stored in a database and are abnormal groups determined according to historical records.

3. The method according to claim 2, wherein The determining the abnormal group set in the group set according to the at least one clustering result includes: Traverse the at least one clustering result in sequence, and detect whether there are the abnormal groups in the at least one clustering result; Use the clustering result in which the abnormal groups exist as the malicious private domain set.

4. The method according to claim 2, wherein After determining the malicious private domain set in the group set according to the at least one clustering result, the method further includes: Store the malicious private domain set in the database.

5. The method according to claim 1, characterized in that, The determining the abnormal group set in the group set according to the at least one clustering result includes: Sample the at least one clustering result to obtain a sampled clustering result; Send the sampled clustering result to an artificial review platform to obtain the abnormal group set.

6. The method according to claim 1, characterized in that, The determining the abnormal group set in the group set according to the at least one clustering result includes: Obtain the historical records of each group in the group set, where the historical records include the number of times the group is reported and the number of times the reporting rules are hit; Traverse the at least one clustering result according to the historical records, and detect whether there are groups with the historical records in the at least one clustering result; Use the clustering result in which there are groups with the historical records as the abnormal group set.

7. The method according to any one of claims 1 to 6, characterized in that, Obtaining the environmental characteristics and content characteristics of each group in the group set includes: Obtain the group names, group management personnel information, group owner information, image information, video information, and link information published within the groups in the group set; Determine the environmental characteristics of each group in the group set according to the group names, the group management personnel information, and the group owner information; Determine the content characteristics of each group in the group set according to the image information, the video information, and the link information.

8. The method according to claim 7, wherein The determining the association relationship between each group in the group set based on the environmental characteristics and the content characteristics includes: Match the group names, group administrator information, group owner information, image information, video information, and link information published in the first group and the second group to obtain a matching result, where the first group and the second group are included in the group set; When the matching result indicates that there is at least one same item between the first group and the second group, determine that the first group and the second group have an association relationship; Traverse each group in the group set in turn to determine the association relationships between the groups in the group set.

9. The method according to any one of claims 1 to 6, 8, characterized in that, Performing clustering processing on the relationship graph to obtain at least one clustering result includes: Performing clustering processing on the relationship graph according to the community discovery algorithm to obtain the at least one clustering result.

10. The method according to claim 9, characterized in that, Performing clustering processing on the relationship graph according to the community discovery algorithm to obtain the at least one clustering result includes: Calculate multiple first modularities by successively comparing the third group with other groups in the relationship graph; Select two groups corresponding to the largest first modularity to form a first community, and the first community includes the third group; Traverse and calculate in turn to obtain a first community graph, and the nodes of the first community graph are communities obtained by pairwise combination of groups; Calculate multiple second modularities by successively comparing the first community with other communities in the first community graph; Select two communities corresponding to the largest second modularity to form a first supernode, and the first supernode includes the first community; Traverse and calculate in turn to obtain a second community graph, and the nodes of the second community graph are supernodes obtained by pairwise combination of the communities in the first community graph; Repeat the above operations in turn until the convergence condition is reached, and output the community discovery result as the at least one clustering result.

11. The method according to claim 10, characterized in that Calculating multiple first modularities by successively comparing the third group with other groups in the relationship graph includes: Obtain a first edge weight and a sum of first edge weights, where the first edge weight is the edge weight between the third group and the fourth group, the sum of the first edge weights is the sum of the edge weights of the groups having an association relationship with the third group, and the fourth group is included in the relationship graph; Calculate the first modularity between the third group and the fourth group according to the first edge weight and the first edge weight; Traverse and calculate in turn to obtain the multiple first modularities.

12. The method according to claim 11, wherein Obtaining the first edge weight includes: Obtain the assignments corresponding to the group name, group administrator information, group owner information, image information, video information, and link information between the third group and the fourth group; Assign values to the group name, the group administrator information, the group owner information, the image information in the group, the video information, and the link information respectively to obtain an assignment set; Calculate the weights of each assignment in the assignment set to obtain the first edge weight.

13. A data mining device, characterized in that, Includes: An obtaining module, configured to obtain a group set, where the groups in the group include multiple interaction objects; obtain the environmental features and content features of each group in the group set, where the environmental features are used to indicate the group management features of the group, and the content features are used to indicate the features of the published content of the interaction objects in the group; A processing module, configured to determine the association relationships between the various groups in the group set based on the environmental features and the content features; construct a relationship graph according to the group set and the association relationships, where the nodes of the relationship graph are the various groups in the group set, and the edges of the relationship graph are the association relationships of the various groups in the group set; perform clustering processing on the relationship graph to obtain at least one clustering result; determine an abnormal group set in the group set according to the at least one clustering result.

14. A computer device, characterized in that, Comprising: a memory, a processor, and a bus system; wherein, the memory is used for storing programs; the processor is used for executing the programs in the memory, and the processor is used for executing the method according to any one of claims 1 to 12 according to the instructions in the program code; the bus system is used for connecting the memory and the processor, so that the memory and the processor can communicate.

15. A computer-readable storage medium, comprising instructions, which when running on a computer, cause the computer to execute the method according to any one of claims 1 to 12.