A method, apparatus, device, and storage medium for group data analysis
By acquiring operational statistics and object operation data of target groups, the influence level is determined and target objects are filtered. Combined with feature data to identify group categories, the problem of inaccuracy in group analysis caused by text spoofing is solved, and more efficient and reliable group data analysis is achieved.
Patent Information
- Application Number
- CN202210303341.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-03-24
AI Technical Summary
In existing technologies for group data analysis, text data is easily disguised and altered, resulting in insufficient accuracy and reliability of group analysis and difficulty in identifying the true nature of groups.
By acquiring operational statistics and object operation data of the target object group, the influence of the object is determined, multiple target objects are filtered out, and their object characteristic data is obtained. Combining the influence and characteristic data, the group category is determined, reducing interference from irrelevant objects and improving the accuracy of the analysis.
It improves the accuracy and reliability of group category identification, avoids text spoofing interference, accurately identifies the true nature of groups, and enhances the efficiency and reliability of analysis.
Smart Images

Figure CN116881443B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to a group data analysis method, apparatus, device, and storage medium. Background Technology
[0002] With the development of internet technology, the number and complexity of target groups are increasing. Analyzing group-related data is beneficial for network security prevention and anomaly identification, and is an indispensable risk control technique. For example, when conducting group category identification or group risk analysis, relevant data is analyzed to identify the group category or whether there are anomalies and risks.
[0003] Existing technologies typically utilize textual data related to objects within a group for textual analysis to perform category analysis or anomaly identification on group objects, thereby determining the group's category or status. However, textual data is easily disguised and altered. When objects intentionally replace text and evade sensitivity, it can lead to the inability to identify their true nature, thus affecting the accuracy of group analysis. Therefore, a more reliable solution is needed. Summary of the Invention
[0004] To address the problems of existing technologies, this application provides a group data analysis method, apparatus, device, and storage medium. The technical solution is as follows:
[0005] This application provides a group data analysis method, the method comprising:
[0006] Obtain operation statistics data for the target object group and object operation data for each individual object in the target object group;
[0007] The influence degree of each of the multiple objects is determined based on the operation statistics and the object operation data;
[0008] Based on the influence level, multiple target objects are determined from the multiple objects;
[0009] Obtain the object feature data of each of the multiple target objects;
[0010] The group category of the target object group is determined based on the object feature data and influence of each of the multiple target objects.
[0011] Another aspect of this application provides a group data analysis apparatus, the apparatus comprising:
[0012] Data acquisition module: used to acquire operation statistics data of the target object group and object operation data of each object in the target object group;
[0013] Influence Determination Module: Used to determine the influence degree of each of the multiple objects based on the operation statistics data and the object operation data;
[0014] Target object determination module: used to determine multiple target objects from the multiple objects based on the influence degree;
[0015] Feature data acquisition module: used to acquire the object feature data of each of the multiple target objects;
[0016] Group category determination module: used to determine the group category of the target object group based on the object feature data and influence of the multiple target objects.
[0017] On the other hand, a computer device is provided, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the group data analysis method as described above.
[0018] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor to implement the group data analysis method as described above.
[0019] On the other hand, a server is provided, the server including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the group data analysis method as described above.
[0020] On the other hand, a terminal is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the group data analysis method described above.
[0021] On the other hand, a computer program product or computer program is provided, which includes computer instructions that, when executed by a processor, implement the group data analysis method described above.
[0022] The group data analysis method, apparatus, device, storage medium, server, terminal, computer program, and computer program product provided in this application have the following technical effects:
[0023] This application first obtains operational statistics data of a target object group and object operation data of each object within the target object group. Then, it determines the influence degree of each of the multiple objects based on the operational statistics data and the object operation data. Based on the influence degree, it identifies multiple target objects from the multiple objects and obtains object feature data for each of the multiple target objects. Finally, it determines the group category of the target object group based on the object feature data and influence degree of each of the multiple target objects. This reduces interference from irrelevant objects in group data analysis, which is beneficial for improving the accuracy, efficiency, and reliability of group category identification. Furthermore, combining object influence degree and object feature data for group category identification can avoid text spoofing interference, accurately identify the true nature of the group, and improve analysis accuracy. Attached Figure Description
[0024] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application;
[0026] Figure 2 This is a flowchart illustrating a group data analysis method provided in an embodiment of this application;
[0027] Figure 3 This is a relationship diagram of a target object group provided in an embodiment of this application;
[0028] Figure 4 This is a flowchart illustrating another group data analysis method provided in an embodiment of this application;
[0029] Figure 5 This is a flowchart illustrating another group data analysis method provided in an embodiment of this application;
[0030] Figure 6 This is a relationship diagram of another target object group provided in the embodiments of this application;
[0031] Figure 7 This is a flowchart illustrating another group data analysis method provided in an embodiment of this application;
[0032] Figure 8 This is a flowchart illustrating another group data analysis method provided in an embodiment of this application;
[0033] Figure 9This is a relationship diagram of another target object group provided in the embodiments of this application;
[0034] Figure 10 This is a relationship diagram of another target object group provided in the embodiments of this application;
[0035] Figure 11 yes Figure 9 A diagram illustrating the keyword analysis results for the target group.
[0036] Figure 12 This is a schematic diagram of a group data analysis device provided in an embodiment of this application;
[0037] Figure 13 This is a hardware structure block diagram of an electronic device for implementing a group data analysis method, provided in an embodiment of this application. Detailed Implementation
[0038] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout.
[0039] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0040] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0041] Clustering: The process of dividing a collection of physical or abstract objects into multiple classes composed of similar objects is called clustering. A cluster generated by clustering is a set of data objects that are similar to objects within the same cluster and distinct from objects in other clusters. Cluster analysis refers to the analytical process of grouping a collection of physical or abstract objects into multiple classes composed of similar objects.
[0042] K-means clustering algorithm is an iterative clustering analysis algorithm. Its steps are as follows: First, the data is divided into K groups. Then, K objects are randomly selected as initial cluster centers. Next, the distance between each object and each seed cluster center is calculated, and each object is assigned to the nearest cluster center. The cluster centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the cluster centers are recalculated based on the existing objects in the cluster. This process is repeated until a termination condition is met. The termination condition may be that no (or a minimum number) objects are reassigned to different clusters, no (or a minimum number) cluster centers change, or the sum of squared errors reaches a local minimum.
[0043] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0044] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0045] Please see Figure 1 , Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application, such as... Figure 1 As shown, the application environment may include at least terminal 01 and server 02. In practical applications, terminal 01 and server 02 can be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions on this.
[0046] In this application embodiment, server 02 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0047] Specifically, cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology can be applied to various fields, such as medical cloud, cloud IoT, cloud security, cloud education, cloud conferencing, AI cloud services, cloud applications, cloud calling, and cloud social networking. Based on the cloud computing business model, cloud technology distributes computing tasks across a resource pool composed of numerous computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, the resources in the "cloud" are infinitely scalable, readily available, on-demand, expandable, and pay-as-you-go. As a provider of basic cloud computing capabilities, a cloud resource pool (referred to as a cloud platform, generally called IaaS (Infrastructure as a Service)) platform is established, deploying various types of virtual resources within the pool for external customers to choose from. The cloud resource pool mainly includes: computing devices (virtualized machines containing operating systems), storage devices, and network devices.
[0048] Based on logical function, a PaaS (Platform as a Service) layer can be deployed on top of the IaaS layer, and a SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. Alternatively, SaaS can be deployed directly on top of IaaS. PaaS is a platform for running software, such as databases and web containers. SaaS refers to various types of business software, such as web portals and bulk SMS senders. Generally speaking, SaaS and PaaS are upper layers compared to IaaS.
[0049] Specifically, the server 02 mentioned above may include physical devices, such as network communication submodules, processors, and memory, and may also include software running on the physical devices, such as applications.
[0050] Specifically, terminal 01 may include physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart voice interaction devices, smart home appliances, smart wearable devices, and in-vehicle terminal devices, and may also include software running on the physical devices, such as applications.
[0051] In this embodiment, terminal 01 can run a client targeting an object to provide login and operation services, such as transaction services like fund transfers. Server 02 is used to obtain operation statistics data of the object group authorized by the object and object operation data of the object itself, and to filter target objects based on the above data to determine the target objects in the object group. Server 02 is also used to obtain object feature data of the target objects, and then determine the group category based on the object feature data and influence. In addition, server 02 can also be used to provide storage services for the above-mentioned types of data.
[0052] Furthermore, it is understandable that Figure 1 The example shown is merely an application environment for a group data analysis method. This application environment may include more or fewer nodes, and this application does not impose any limitations on it.
[0053] The application environment involved in this application embodiment, or the terminal 01 and server 02 in the application environment, can be a distributed system formed by connecting clients and multiple nodes (any form of computing device accessing the network, such as servers and user terminals) through network communication. The distributed system can be a blockchain system, which can provide the aforementioned group data analysis services and data storage services.
[0054] The following describes a group data analysis method based on the aforementioned application environment. This application's embodiments can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving. Please refer to... Figure 2 , Figure 2 This is a flowchart illustrating a group data analysis method provided in an embodiment of this application. This specification provides method operation steps as shown in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual system or server product execution, the method can be executed sequentially according to the embodiments or drawings, or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as... Figure 2 As shown, the method may include the following steps S201-S209.
[0055] S201: Obtain operation statistics for the target object group and object operation data for each object in the target object group.
[0056] In this embodiment, the object group includes multiple objects. Specifically, the multiple objects in the target object group may include all objects in the target object group. In one embodiment, the object group can be a group that needs to be categorized, and the objects in the object group can be the registered accounts or member identifiers of the members in the group. In practical applications, unique identifiers can be used to mark and distinguish different objects in the target object group. The unique identifiers may include, but are not limited to, the object's account or member identifier. An object can generate operation data on multiple clients, and the object's unique identifier can be used to completely collect this operation data for group data analysis.
[0057] Specifically, the operation statistics for the target object group may include, but are not limited to, group operation intensity and group operation count. Object operation data may include, but are not limited to, object operation intensity and object operation count. Group operation intensity is the total operation intensity corresponding to multiple objects within the target object group, and group operation count is the total number of operations corresponding to multiple objects within the target object group. Specifically, object operation intensity may also include out-degree operation intensity and in-degree operation intensity, and object operation count may also include out-degree operation count and in-degree operation count. Specifically, the operation statistics and object operation data can be operation data within a preset time period, and the operation types may include, but are not limited to, transfers, payments on behalf of others, red envelopes, payment purchases, and logins on the same device.
[0058] In one embodiment, the target object group is a group constructed based on business transaction information, and the operation type is specifically a transfer. Accordingly, the object operation data may include remittance relationship data, specifically transaction data between the remitter and the recipient, and operation data summarized from the transaction data. Specifically, the group operation intensity can be the total transaction amount of all objects in the target object group, and the group operation count can be the total number of transactions of all objects in the target object group; the out-degree operation count can be the amount remitted by the object in the target object group, the in-degree operation intensity can be the amount remitted by the object in the target object group, the out-degree operation count can be the number of times the object remits out of the target object group, and the in-degree operation count can be the number of times the object remits in the target object group.
[0059] S203: Determine the influence of each of the multiple objects based on operation statistics and object operation data.
[0060] In this embodiment, influence characterizes the importance of an object within a target object group, which can be considered as the object's weight within the target object group. It is understood that the higher the influence, the more important the object. Specifically, the first percentage of each object's operation intensity in the target object group's total operation intensity, and the second percentage of each object's operation count in the total number of group operations, are obtained. Then, the influence of each object is calculated based on the first and second percentages, such as by summing and averaging the first and second percentages to obtain the influence.
[0061] In one embodiment, the influence of each of the above objects is calculated using the following formula, where IM j G represents the influence of object j. a For group operation strength, G c P represents the number of group operations. iax P is the intensity of the in-degree operation of object j. icx P represents the number of in-degree operations. sax For the degree of operation intensity, P scx This refers to the number of out-degree operations. The first percentage mentioned above includes the percentage of in-degree intensity. and the proportion of intensity The second proportion includes the proportion of in-degrees. and the percentage of out-degrees
[0062]
[0063] Specifically, taking a transfer as an example, G a G represents the total transaction amount of the group. c P represents the total number of transactions in the group. iax For the amount remitted to object j, P icx P represents the number of inflows. sax For the amount remitted, P scx This refers to the number of outbound remittances.
[0064] S205: Identify multiple target objects from multiple objects based on their influence.
[0065] In this embodiment, after obtaining the influence of each object, the objects in the target object group can be sorted according to their influence. Multiple objects with higher influence are selected as target objects. Then, based on the relevant data, features, and related group data of the target objects, group data analysis is performed to determine the group category. The aforementioned multiple target objects constitute a proper subset of the target object group. This allows for the selection of important objects within the target object group, reducing the amount of group data analysis and resource consumption, improving data analysis efficiency, and preventing unimportant objects in the group from affecting the group data analysis results, thereby improving the accuracy of group category identification. Please refer to... Figure 3 , Figure 3 A relationship diagram of a group of target objects is shown. Each node in the diagram corresponds to an object, and each arrow points to a line that is an edge of the relationship diagram. The nodes in the dashed boxes are the target objects selected using the above scheme.
[0066] In some embodiments, the number of target objects selected is a preset value, which can be set based on actual needs or experience. In other embodiments, the number of target objects is determined based on the group's basic attribute data and the relationships between multiple objects in the target object group. The group's basic attribute data may include, but is not limited to, the number of relationship logs and the number of objects in the target object group. Please refer to [reference needed]. Figure 4 Prior to S203, the method also includes steps S301-S303 for determining the target quantity.
[0067] S301: Obtain the object association density and object association dispersion of the target object group.
[0068] In practical applications, object association density characterizes the density of associations among multiple objects in a target object group, while object association dispersion characterizes the dispersion of these associations. These associations represent the relationships between objects in the target object group, and are determined based on the operational interaction data between them. For example, if object A transfers money to object B, then object A and object B have an association relationship. Using this operational interaction data helps to quickly and accurately determine the associations among multiple objects in the target object group, thus aiding in the subsequent determination of the target quantity.
[0069] In practical applications, please refer to Figure 5 S301 may specifically include the following steps S3011-S3013.
[0070] S3011: Obtain the degree of association between multiple objects and the number of relationship pairs corresponding to the association between multiple objects.
[0071] Specifically, each association can have a corresponding association degree. The association degree can be used to characterize the strength of the relationship between the two objects corresponding to the association. For example, the association degree can be the strength of the relational operation of the association, or the proportion of the relational operation strength of the association in the total group operation strength of the target object group, or the proportion of the relational operation strength of the association in the target operation strength of the target object group. Here, the target operation strength can refer to the total operation strength of the target object group on the operation type of the association. In the relationship graph, the association degree can correspond to the value of the edge. For example, taking the association between object A and object B as an example, the operation type is transfer, and the association is that object A remits money to object B. The association degree corresponding to this association can be the amount of money remitted by object A to object B, or the proportion of the remittance amount of this association in the total remittance amount of the target object group.
[0072] Specifically, the relation logarithm is the number of pairs of objects that form a relationship among all objects in a target object group, such as object A and object B forming a relation pair. Taking a relation graph as an example, the relation logarithm corresponds to the number of edges in the graph. Please refer to [reference needed]. Figure 6 , Figure 6 The relationship diagram shows that the target object group includes object A, object B, object C, object D, and object E. A has relationships with B, C, and D respectively, and the target object group has 5 relationship pairs.
[0073] S3012: The ratio between the number of relation logs and the number of objects is determined as the object association density.
[0074] Specifically, the number of objects is the total number of all objects in the target object group, corresponding to the number of nodes in the relationship graph. Object association density can specifically characterize the comparison between the number of relation pairs and the number of objects in the target object group. It can be understood that the larger the ratio between the number of relation pairs and the number of objects, the greater the object association density, indicating that the relationships between objects in the target object group are tighter, and vice versa.
[0075] In one embodiment, the above-mentioned object association density is calculated using the following formula, where G rp For object association density, G e G is the relation logarithm. n Number of objects.
[0076]
[0077] S3013: Discretely calculate the object association dispersion degree based on the association degree and the logarithm of the association between multiple objects.
[0078] Specifically, object association dispersion can characterize the degree of differentiation between objects with a large number of associated objects and objects with a small number of associated objects in a target object group, that is, whether there are a few objects (nodes) occupying the majority of association relationships (edges) and / or occupying the majority of association degrees.
[0079] In one embodiment, the aforementioned object association dispersion is calculated using the following formula, where G rc Let x represent the dispersion of object associations, and x represent the association degree of an association relationship. This is the average degree of association among all relationships in the target object group.
[0080]
[0081] S303: Determine the target quantity of the target object based on the object association density, object association dispersion, and the number of objects in the target object group.
[0082] In practical applications, the target number of objects to be filtered is calculated based on object association density, object association dispersion, and the total number of objects. Specifically, the proportion of target objects can be obtained based on object association density and object association dispersion, and then the total number of objects can be multiplied by this proportion to obtain the target number. In one embodiment, the aforementioned target number G... imc The following formula is used for calculation.
[0083]
[0084] When the target number is determined based on the aforementioned group basic attribute data and association relationships, S203 may specifically include: determining the target number of target objects from multiple objects based on the degree of influence.
[0085] S207: Obtain the object feature data of each of the multiple target objects.
[0086] In this embodiment, after filtering out multiple target objects, the features of each target object across multiple feature categories are obtained, and then the group characteristics of the target object group are measured based on the features matched by each target object. For example, the overall group risk characteristics of the target object group are measured based on the risk features matched by each target object. Specifically, the multiple feature categories can be set based on the actual needs of group data analysis. For example, in a risk category identification scenario, the multiple feature categories may include, but are not limited to, object account level, object account registration years, and authorized user age, gender, and operation records (such as transaction records). In one embodiment, 25 feature categories may be included, and correspondingly, the object feature data of each target object includes the feature data corresponding to the object in each of the 25 feature categories.
[0087] S209: Determine the group category of the target object group based on the object characteristic data and influence of each of the multiple target objects.
[0088] In this embodiment, based on the actual needs of group data analysis, group categories can represent different meanings. For example, in a risk identification scenario, a group category can represent the risk category or security level of a target group. In a group association attribute identification scenario, a group category can represent the association attribute of a target group, such as belonging to a business group, a work group, or a family and friends group. In one embodiment, a group category represents the remittance event category of a target group, specifically representing the remittance purpose of the group, such as normal transaction remittances and illegal transaction remittances. Illegal transaction remittances can be, for example, money laundering remittances.
[0089] In practical applications, please refer to Figure 7 S209 may include the following steps S2091-S2092.
[0090] S2091: Perform group feature statistics based on the individual object feature data and influence of multiple target objects to obtain the group feature data of the target object group.
[0091] Specifically, based on the influence of each target object, the feature data of all target objects across multiple feature categories can be weighted and summed to obtain the feature data of the target object group in each feature category, thus obtaining the group feature data. This group feature data includes the feature data of the target object group in each feature category. Then, the feature data of the target object group in each feature category is used for category identification to obtain the group category. For example, if there are 25 feature categories, the feature data of the target object group in all 25 feature categories is obtained, and then the group's feature data in all 25 feature categories is used for category identification.
[0092] In one embodiment, the aforementioned group feature data is calculated using the following formula. Where y is the feature category identifier, ty represents the feature category y, IM1 is the influence degree of target object 1, and IM... Gimc For the influence of the target object Ginc, R 1ty R represents the feature data of target object 1 corresponding to feature category ty. Gimcty G represents the feature data of the target object Gimc in feature category ty. ty The feature data of the target feature group on feature category ty, the group feature data includes G t1 G t2 G t3 , ...G ty .
[0093] G ty =IM1*R1ty +IM2*R 2ty +…+IM Gimc-1 *R (Gimc-1)ty +IM Gimc *R Gimcty
[0094] In some embodiments, object feature data includes object feature labels for each feature category of a target object across multiple feature categories; that is, the feature data of a target object in each feature category is the object feature label. Group feature data includes group feature labels for each feature category of a target object group; that is, the feature data of a target object group in each feature category is the group feature label. Accordingly, the acquisition method of group feature data may specifically include: performing a weighted summation of the object feature labels of multiple target objects in each feature category based on the influence of each of the multiple target objects, to obtain the group feature label of the target object group in each feature category. The group feature label can characterize the performance of the target object group in that feature category.
[0095] Taking a scenario with 25 feature categories and 25 risk labels for each target object as an example, the risk labels for target object 1 in the 25 feature categories are R... 1t1 R 1t2 R 1t3 R 1t4 , ...R 1t24 R 1t25 Correspondingly, the group risk label G in the group feature data t1 G t2 G t3 , ...G t25 They can be calculated based on the following formulas respectively.
[0096] G t1 =IM1*R 1t1 +IM2*R 2t1 +…+IM Gimc-1 *R (Gimc-1)t1 +IM Gimc *R Gimct1
[0097] G t2 =IM1*R 1t2 +IM2*R 2t2 +…+IM nGimc-1 *R (Gimc-1)t2 +IM Gimc *R Gimct2
[0098] ...
[0099] G t25 =IM1*R1t25 +IM2*R 2t25 +…+IM Gimc-1 *R (Gimc-1)t25 +IM Gimc *R Gimct25
[0100] S2092: Perform group category identification on the group feature data to obtain the group category of the target object group.
[0101] Specifically, after obtaining the group feature data, the feature data of the target object group in each feature category can be identified. In some embodiments, the target category recognition model can be called to identify the group category of the target object group in each feature category and obtain the group category of the target object group.
[0102] The target category recognition model is a model obtained by training an initial recognition model under constraints based on sample group feature data and sample group categories across multiple feature categories. The sample object groups are historical object groups with group category labels. Specifically, the sample group feature data of the sample object groups across multiple feature categories can be concatenated, such as concatenating the feature labels of each group to obtain concatenated features. These concatenated features are used as input to the initial recognition model, and the corresponding sample group categories are used as the expected output for the constrained training of the aforementioned group category recognition. In some embodiments, the target category recognition model is a multivariate logistic regression model. In one embodiment, there are 25 feature categories, and the group categories include 5 risk categories (F). An example of the sample group feature data of the sample object groups is shown in the table below. Here, F represents the risk category; the larger the number, the higher the risk. Risk category 5 represents a higher risk level than risk category 1.
[0103]
[0104] In other embodiments, cluster analysis can be used to identify the group category based on the group feature labels of the target object group in each feature category. Accordingly, S2092 may specifically include the following steps S20911-S20924.
[0105] S20911: Obtain the group feature labels for each feature category of multiple reference object groups.
[0106] S20912: Based on the group feature labels of multiple reference object groups in each feature category and the group feature labels of the target object group in each feature category, perform cluster analysis on multiple reference object groups and target object groups to obtain multiple group clusters.
[0107] S20913: Identify the target group cluster that includes the target object group from multiple group clusters.
[0108] S20914: Determine the group category of the target object group based on the group category of the reference object group in the target group cluster.
[0109] Specifically, the reference object group is a group of objects whose group categories are known. Using a pre-defined clustering method, cluster analysis is performed on the reference feature data of the reference object group across multiple feature categories and the group feature data of the target object group to classify all groups into multiple group clusters. Then, the group category of the target object group is determined using an unsupervised method. The method for obtaining the reference feature data of the reference object group is similar to that of obtaining the group feature data mentioned above, and will not be repeated here.
[0110] Specifically, the clustering analysis uses the group feature labels of the reference object group and the target object group in each feature category as the clustering analysis objects, resulting in multiple clusters. The object groups within each cluster are of similar categories, and the minimum number of object groups in a cluster can be greater than or equal to 2. After clustering analysis, the cluster containing the target object group is determined as the target cluster. If all reference object groups in the target cluster belong to the same group category, then the group category of that reference object group is used as the group category of the target object group. If the reference object groups in the target cluster belong to different group categories, then the group category with the highest number of reference object groups is used as the group category of the target object group. The clustering analysis method can include, but is not limited to, K-means.
[0111] Based on some or all of the above embodiments, please refer to the embodiments of this application. Figure 8 The method may also include the following steps S401-S409.
[0112] S401: Retrieves the static text data of each of the multiple objects in the target object group.
[0113] In practical applications, static text data can include, but is not limited to, static resource information such as operation interaction record text, object identifier text, and object name obtained through object authorization. Taking transaction operations as an example, static text data can include, but is not limited to, account text information and transaction text information obtained through object authorization, such as account name, account signature, transaction record description, and transaction item name. In business transaction scenarios, static text data can include account static information, which can include, but is not limited to, the account text information of the remitter and the recipient.
[0114] S403: Perform word segmentation on the static text data of each of the multiple objects to obtain word segmentation statistics and the corresponding text word segments for each of the multiple objects.
[0115] In practical applications, word segmentation can be performed on various static text data of each object separately to obtain the corresponding text words for each object. The word segmentation methods used can be dictionary-based word segmentation algorithms, statistical machine learning algorithms, or neural network-based word segmentation algorithms, etc. Statistical processing is then performed on the text words of each object in the target object group to obtain the word segmentation statistics for the target object group. The word segmentation statistics may include, but are not limited to, the total number of words and the number of word categories for the target object group. For example, if the text words of object A and object B both contain "crystal", then "crystal" in object A is counted as one text word number, and "crystal" in object B is also counted as one text word number. "Crystal" is then recorded as one word category, and different text words belong to different word categories.
[0116] S405: Perform vector representation on the text segmentation corresponding to each of the multiple objects to obtain the static text vectors of each object.
[0117] In practical applications, the text segments corresponding to each object are concatenated to obtain a segmentation sequence for each object. Then, the segmentation sequences are vectorized to obtain the static text vectors for each of the multiple objects. Alternatively, each text segment corresponding to each object is first vectorized to obtain a segmentation vector for each text segment. Then, the segmentation vectors of the same object are concatenated to obtain the static text vector for that object.
[0118] S407: Determine the clustering degree of the target object group based on the word segmentation statistics, the static text vectors of multiple objects, and the correlation degree between multiple objects.
[0119] In practical applications, clustering degree characterizes the overall correlation between multiple objects in a target object group. Based on the static text vector of each object in the target object group, the text vector similarity between the two objects corresponding to each correlation can be determined. This similarity can be calculated using existing vector similarity algorithms, such as cosine similarity, and is not limited here. Then, a reference correlation degree is calculated based on the text vector similarity corresponding to each correlation and the correlation degree of that correlation, yielding the reference correlation degree of that correlation. In one embodiment, the reference correlation degree is calculated using the following formula, where AMT... i Gamt represents the relation operation strength of association i (edge i), Gamt represents the total operation strength of the target object group on the operation type of association i, and correspondingly, AMT i / G amt RE represents the degree of association of association i.i E represents the text vector similarity between two objects corresponding to association i. i For reference relevance.
[0120] E i =(AMT) i / G amt )*RE i
[0121] Taking remittance as an example, if relationship i is a remittance from object A to object B, then AMT i Represents the remittance amount for relationship i, Gamt represents the total remittance amount for the target group, and RE i This represents the text vector similarity between object A and object B.
[0122] Furthermore, the clustering degree of the target object group is calculated based on the reference relevance of each relationship corresponding to the target object group and the word segmentation statistics results. Specifically, the reference relevance of each relationship corresponding to the target object group is summed to obtain the relationship clustering degree of the target object group. This summing process can be simple summing or weighted summing, etc.; and the word segmentation clustering degree of the target object group is calculated based on the total number of words and the number of word categories in the word segmentation statistics results. Then, the clustering degree of the target object group is determined based on the relationship clustering degree and the word segmentation clustering degree. In one embodiment, the relationship clustering degree, the word segmentation clustering degree, and the clustering degree of the target object group are calculated using the following formulas, where, μ represents the clustering degree of the target object group n. n It is the degree of word segmentation aggregation. It represents the degree of clustering of the relationship, corresponding to the degree of clustering in the relationship graph. N represents the total number of words corresponding to the target object group, and n represents the number of word categories corresponding to the target object group.
[0123]
[0124] μ n =N / n
[0125]
[0126] S409: If the clustering degree of the target object group meets the preset conditions, the target object group is determined to be an associated object group.
[0127] In practical applications, the clustering degree meeting the preset condition can specifically mean that the clustering degree is greater than or equal to a preset threshold. This preset threshold can be set based on actual needs or empirical values, and is not limited here. For example, the preset threshold can be 0.4. When the clustering degree meets the preset condition, the target object group is determined to be a group constructed based on objects with the same purpose event, i.e., a related object group. When the clustering degree does not meet the preset condition, such as being less than the preset threshold, the target object group is not a related object group, but can be a group constructed based on a random event.
[0128] In some embodiments, steps S401-S409 can be performed before step S201. Accordingly, if the clustering degree of the target object group meets a preset condition and the target object group is determined to be an associated object group, step S401 is executed to perform group data analysis and corresponding group category identification. If the preset condition is not met, group data analysis is not performed. In this way, group filtering can be performed in conjunction with static resource information to solve computational resource issues and avoid resource waste.
[0129] In other embodiments, the results obtained in steps S401-S409 are used as auxiliary results for group data analysis. If the target object group is determined to be a related object group, it indicates that the obtained group category result is highly reliable. If the clustering degree does not meet the preset conditions, it indicates that the group category result is questionable, and further analysis of the target object group can be performed, or a questionable marker and related reminder information can be generated for the target object group for review. Thus, combining static resource information improves the accuracy of group data analysis.
[0130] In some embodiments, the word segmentation statistics results may also include the frequency information of the text segmentation. Based on the frequency information of the text segmentation, keyword analysis is performed on the text segmentation corresponding to each of the multiple objects to obtain the target keywords corresponding to the target object group.
[0131] Specifically, the frequency information of text segmentation represents the probability of the occurrence of the text segmentation in all text segments corresponding to the target object group. The segmentation is sorted according to the frequency information of each text segmentation corresponding to the target object group, and a certain number of segments with the highest sorting are used as keywords. The number of keywords can be preset, or the keywords can be selected from the segments with frequency information higher than the preset probability.
[0132] As can be understood in S401-S409 above, as well as in the keyword analysis step, static text data of all objects in the target object group can be used for analysis. In some cases, such as when used as an auxiliary analysis, only static text data of the target object can be used for analysis to further filter the static text data and improve the reliability of clustering analysis and keyword analysis.
[0133] Based on the above technical solution of this application, it is possible to accurately identify the group category, whether it is a related object group, and the business category of the group. Please refer to... Figure 9-11 The results obtained from the above data analysis methods show that Figure 9 The target object group in the text is the associated object group (the nodes marked with boxes are the target objects). Figure 10 The target object group in the text is not a related object group (the nodes marked with boxes do not share the same destination event, but are related due to a common remitter caused by chance), but a group built based on incidental events (accidental transfers). Furthermore, Figure 11 It shows Figure 9 The keyword analysis results of the target group in the data, through word cloud analysis of the static account information of the target, show that the most frequent words are "crystal", "wholesale" and "jewelry". These three words reflect the strong correlation of the target group and the business category to which the target group is most likely to belong.
[0134] In summary, this application filters out key objects within a group based on operational data related to that group. Then, it uses the object characteristics and influence of these key objects to identify group categories. Furthermore, it combines static resource information related to the group to perform group clustering and keyword analysis, assisting in determining whether the object groups are truly related. Thus, by combining dynamic operational information and static information to identify group categories, such as risk level, suspiciousness, or security level, as well as business category identification, it can accurately and efficiently identify groups and group categories with the same purpose. It can also provide analysis results for groups with ambiguous group or business categories in the analysis results, facilitating review and judgment, thereby improving efficiency.
[0135] This application also provides a group data analysis device 600, such as... Figure 12 As shown, the device may include:
[0136] Data acquisition module 10: used to acquire operation statistics data of the target object group and object operation data of each object in the target object group;
[0137] Impact Determination Module 20: Used to determine the impact of multiple objects based on operation statistics and object operation data;
[0138] Target object determination module 30: used to determine multiple target objects from multiple objects based on influence;
[0139] Feature data acquisition module 40: used to acquire the object feature data of multiple target objects;
[0140] Group Category Determination Module 50: Used to determine the group category of target object groups based on the object characteristic data and influence of multiple target objects.
[0141] In some embodiments, the apparatus may further include:
[0142] Group data acquisition module: used to acquire the object association density and object association dispersion of the target object group before determining multiple target objects from multiple objects based on influence. The object association density characterizes the density of the association between multiple objects in the target object group, and the object association dispersion characterizes the dispersion of the association between multiple objects in the target object group.
[0143] Target Quantity Determination Module: Used to determine the target quantity of a target object based on object association density, object association dispersion, and the number of objects in a target object group.
[0144] Correspondingly, the target object determination module 30 can be specifically used to determine the target number of target objects from multiple objects based on the degree of influence.
[0145] In some embodiments, the group data acquisition module may include:
[0146] The associated data acquisition submodule is used to obtain the degree of association between multiple objects and the number of relationship logs corresponding to the association between multiple objects;
[0147] The association density determination submodule is used to determine the object association density by the ratio between the logarithm of relations and the number of objects.
[0148] The association dispersion determination submodule is used to perform discrete calculations based on the association degree and logarithm of the association relationships between multiple objects to obtain the association dispersion of the objects.
[0149] In some embodiments, the group category determination module 50 may include:
[0150] Feature statistics submodule: used to perform group feature statistics based on the individual object feature data and influence of multiple target objects, and obtain the group feature data of the target object group;
[0151] The first category identification submodule is used to identify the group category of the target object group by analyzing the group feature data.
[0152] In some embodiments, object feature data includes object feature labels for each feature category of a target object in multiple feature categories, and group feature data includes group feature labels for each feature category of a target object group in multiple feature categories; correspondingly, the feature statistics submodule can be specifically used to: perform weighted summation processing on the object feature labels of multiple target objects in each feature category of multiple feature categories according to the influence of each of the multiple target objects, to obtain the group feature labels of the target object group in each feature category.
[0153] In some embodiments, the group category determination module 50 may be specifically used to: call the target category recognition model to perform group category recognition on the group feature labels of the target object group in each feature category, and obtain the group category of the target object group.
[0154] In other embodiments, the group category determination module 50 may include:
[0155] Reference Group Acquisition Submodule: Used to acquire the group feature labels of multiple reference object groups for each feature category. The reference object groups are object groups whose group categories are known.
[0156] Clustering analysis submodule: This module is used to perform clustering analysis on multiple reference object groups and target object groups based on the group feature labels of multiple reference object groups in each feature category and the group feature labels of the target object group in each feature category, to obtain multiple group clusters.
[0157] Target Group Determination Submodule: Used to determine the target group cluster that includes the target object group from multiple group clusters;
[0158] The second category identification submodule is used to determine the group category of the target object group based on the group category of the reference object group in the target group cluster.
[0159] In some embodiments, the apparatus may further include:
[0160] Text data acquisition module: used to acquire the static text data of multiple objects in a target object group;
[0161] Word segmentation module: Used to perform word segmentation on the static text data of multiple objects respectively, and obtain word segmentation statistics results and the text word segmentation corresponding to each of the multiple objects;
[0162] Vector Representation Module: Used to perform vector representation on the text segmentation corresponding to multiple objects respectively, and obtain the static text vectors of each object;
[0163] Clustering Determination Module: This module determines the clustering degree of a target object group based on word segmentation statistics, the static text vectors of multiple objects, and the degree of association between multiple objects.
[0164] Related Object Group Identification Module: Used to identify a target object group as a related object group when the aggregation degree of the target object group meets preset conditions.
[0165] In some embodiments, the word segmentation statistics result includes the frequency information of the text segmentation. Correspondingly, the device may also include a keyword analysis module: used to perform keyword analysis on the text segmentation corresponding to each of the multiple objects based on the frequency information of the text segmentation, so as to obtain the target keywords corresponding to the target object group.
[0166] The apparatus and method embodiments in the device embodiments are based on the same application concept.
[0167] This application provides a group data analysis device. The scheduling device can be a terminal or a server, including a processor and a memory. The memory stores at least one instruction or at least one program. The at least one instruction or at least one program is loaded and executed by the processor to implement the group data analysis method provided in the above method embodiments.
[0168] Memory can be used to store software programs and modules. The processor executes these stored software programs and modules to perform various functional applications and group data analysis. Memory can primarily include a program storage area and a data storage area. The program storage area stores the operating system, application programs required for functionality, etc.; the data storage area stores data created based on device usage, etc. Furthermore, memory can include high-speed random access memory (RAM) and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory can also include a memory controller to provide the processor with access to the memory.
[0169] The methods and embodiments provided in this application can be executed in electronic devices such as mobile terminals, computer terminals, servers, or similar computing devices. Figure 13 This is a hardware structure block diagram of an electronic device for a group data analysis method provided in an embodiment of this application. For example... Figure 13As shown, the electronic device 900 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 910 (CPUs 910 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 930 for storing data, and one or more storage media 920 (e.g., one or more mass storage devices) for storing application programs 923 or data 922. The memory 930 and storage media 920 may be temporary or persistent storage. The program stored in the storage media 920 may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the CPU 910 may be configured to communicate with the storage media 920 and execute a series of instruction operations in the storage media 920 on the electronic device 900. The electronic device 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows Server. TM Mac OS X TM Unix TM Linux™, FreeBSD™, etc.
[0170] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 900. In one example, the input / output interface 940 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 940 may be a radio frequency (RF) module for wireless communication with the Internet.
[0171] Those skilled in the art will understand that Figure 13 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, the electronic device 900 may also include... Figure 13 The more or fewer components shown, or having the same Figure 13 The different configurations shown.
[0172] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a group data analysis method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the group data analysis method provided in the above method embodiment.
[0173] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0174] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.
[0175] As can be seen from the embodiments of the group data analysis method, apparatus, device, server, terminal, storage medium, and program product provided in this application, this application first obtains the operation statistics of the target object group and the object operation data of each of the multiple objects in the target object group, and then determines the influence degree of each of the multiple objects based on the operation statistics and the object operation data; and determines multiple target objects from the multiple objects based on the influence degree, and obtains the object feature data of each of the multiple target objects; then determines the group category of the target object group based on the object feature data and influence degree of each of the multiple target objects. This can reduce the interference caused by irrelevant objects to group data analysis, which is beneficial to improving the accuracy, analysis efficiency, and reliability of group category identification. Furthermore, combining object influence degree and object feature data for group category identification can avoid text spoofing interference, accurately identify the true nature of the group, and improve the analysis accuracy.
[0176] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.
[0177] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device, equipment, and storage medium embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0178] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program instructing the relevant hardware to implement them. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0179] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for analyzing group data, characterized in that, The method includes: Obtain operation statistics data for the target object group and object operation data for each individual object in the target object group; The influence degree of each of the multiple objects is determined based on the operation statistics and the object operation data; Based on the influence level, multiple target objects are determined from the multiple objects; Obtain object feature data for each of the multiple target objects, wherein the object feature data includes object feature labels for each of the target objects in multiple feature categories; Based on the influence of each of the multiple target objects, the object feature labels of the multiple target objects in each of the multiple feature categories are weighted and summed to obtain the group feature labels of the target object group in each feature category, thereby obtaining the group feature data; The group feature data is used to identify the group category, thereby obtaining the group category of the target object group.
2. The method according to claim 1, characterized in that, Before determining multiple target objects from the multiple objects based on the influence degree, the method further includes: Obtain the object association density and object association dispersion of the target object group. The object association density represents the density of the association relationship between multiple objects in the target object group, and the object association dispersion represents the degree of dispersion of the association relationship between multiple objects in the target object group. The target number of the target objects is determined based on the object association density, the object association dispersion, and the number of objects in the target object group. The process of determining multiple target objects from the multiple objects based on the influence degree includes: Based on the influence level, a target number of target objects are determined from the plurality of objects.
3. The method according to claim 2, characterized in that, The process of obtaining the object association density and object association dispersion of the target object group includes: Obtain the degree of association between the multiple objects and the number of relationship logs corresponding to the association between the multiple objects; The ratio between the logarithm of the relationship and the number of objects is determined as the object association density; The object association dispersion degree is obtained by discretely calculating the association degree and the logarithm of the relationship between the multiple objects.
4. The method according to claim 1, characterized in that, The process of identifying the group category of the target object group by performing group category identification on the group feature data includes: The target category recognition model is invoked to identify the group category of the target object group on each feature category, thereby obtaining the group category of the target object group.
5. The method according to claim 1, characterized in that, The process of identifying the group category of the target object group by performing group category identification on the group feature data includes: Obtain group feature labels for multiple reference object groups on each feature category, wherein the reference object groups are object groups whose group categories are known; Based on the group feature labels of the multiple reference object groups on each feature category and the group feature labels of the target object group on each feature category, cluster analysis is performed on the multiple reference object groups and the target object group to obtain multiple group clusters; From the plurality of group clusters, determine the target group cluster that includes the target object group; The group category of the target object group is determined based on the group category of the reference object group in the target group cluster.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: Obtain the static text data of each of the multiple objects in the target object group; Each of the multiple objects' static text data is segmented into words to obtain segmentation statistics and the corresponding text segments for each of the multiple objects. Each of the multiple objects is segmented into words and represented by vectors to obtain the static text vectors of each of the multiple objects. Based on the word segmentation statistics, the static text vectors of each of the multiple objects, and the degree of correlation between the multiple objects, the clustering degree of the target object group is determined. If the aggregation degree of the target object group meets the preset conditions, the target object group is determined to be an associated object group.
7. The method according to claim 6, characterized in that, The word segmentation statistics result includes the frequency information of text word segmentation, and the method further includes: Based on the frequency information of the text segmentation, keyword analysis is performed on the text segmentation corresponding to each of the multiple objects to obtain the target keywords corresponding to the target object group.
8. A group data analysis device, characterized in that, The device includes: Data acquisition module: used to acquire operation statistics data of the target object group and object operation data of each object in the target object group; Influence Determination Module: Used to determine the influence degree of each of the multiple objects based on the operation statistics data and the object operation data; Target object determination module: used to determine multiple target objects from the multiple objects based on the influence degree; Feature data acquisition module: used to acquire object feature data of each of the multiple target objects, the object feature data including object feature labels of the target objects in each of the multiple feature categories; Feature statistics submodule: used to perform weighted summation of object feature labels of the multiple target objects in each of the multiple feature categories according to the influence of each of the multiple target objects, to obtain the group feature label of the target object group in each feature category, and thus obtain group feature data; The first category identification submodule is used to identify the group category of the group feature data to obtain the group category of the target object group.
9. The apparatus according to claim 8, characterized in that, The device further includes: Group data acquisition module: used to acquire the object association density and object association dispersion of the target object group before determining multiple target objects from the multiple objects based on the influence degree. The object association density represents the density of the association relationship between multiple objects in the target object group, and the object association dispersion represents the degree of dispersion of the association relationship between multiple objects in the target object group. Target quantity determination module: used to determine the target quantity of the target object based on the object association density, the object association dispersion, and the number of objects in the target object group; The target object determination module is specifically used to: determine the target number of target objects from the plurality of objects based on the influence degree.
10. The apparatus according to claim 9, characterized in that, The group data acquisition module includes: Related data acquisition submodule: used to acquire the degree of association between the multiple objects and the number of relationship pairs corresponding to the association between the multiple objects; Association density determination submodule: used to determine the association density of the objects by the ratio between the logarithm of the relationship and the number of objects; Association Dispersion Determination Submodule: Used to perform discrete calculations based on the association degree and the logarithm of the relationship between the multiple objects to obtain the association dispersion of the objects.
11. The apparatus according to claim 8, characterized in that, The group category determination module is specifically used for: The target category recognition model is invoked to identify the group category of the target object group on each feature category, thereby obtaining the group category of the target object group.
12. The apparatus according to claim 8, characterized in that, The group category determination module includes: Reference Group Acquisition Submodule: Used to acquire group feature labels of multiple reference object groups on each feature category, wherein the reference object groups are object groups whose group categories are known; Clustering analysis submodule: used to perform clustering analysis on the multiple reference object groups and the target object group based on the group feature labels of the multiple reference object groups on each feature category and the group feature labels of the target object group on each feature category, to obtain multiple group clusters; Target group determination submodule: used to determine the target group cluster that includes the target object group from the plurality of group clusters; The second category identification submodule is used to determine the group category of the target object group based on the group category of the reference object group in the target group cluster.
13. The apparatus according to any one of claims 8-12, characterized in that, The device further includes: Text data acquisition module: used to acquire the static text data of each of the multiple objects in the target object group; Word segmentation module: used to perform word segmentation on the static text data of the multiple objects respectively, and obtain word segmentation statistics results and the text word segmentation corresponding to the multiple objects respectively; Vector representation module: used to perform vector representation on the text segmentation corresponding to each of the multiple objects respectively, to obtain the static text vector of each of the multiple objects; Clustering Determination Module: Used to determine the clustering degree of the target object group based on the word segmentation statistics results, the static text vectors of each of the multiple objects, and the degree of association between the multiple objects; Related object group identification module: used to determine the target object group as a related object group when the aggregation degree of the target object group meets the preset conditions.
14. The apparatus according to claim 13, characterized in that, The word segmentation statistics result includes the frequency information of text word segmentation, and the device also includes a keyword analysis module, used for: Based on the frequency information of the text segmentation, keyword analysis is performed on the text segmentation corresponding to each of the multiple objects to obtain the target keywords corresponding to the target object group.
15. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the group data analysis method as described in any one of claims 1-7.
16. A computer device, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the group data analysis method as described in any one of claims 1-7.
17. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the group data analysis method as described in any one of claims 1-7.
Citation Information
Patent Citations
Method and device for determining abnormal interactive account
CN108295476A