Hierarchical identification method based on big data analysis
By using big data analytics, the system collects data on the financial and non-financial relationships of participants' accounts in pyramid schemes, generates node embedding vectors, and constructs a hierarchical influence radiation model. This solves the problems of low efficiency and insufficient adaptability in pyramid scheme hierarchical identification, and achieves highly accurate hierarchical identification and leader positioning.
Patent Information
- Application Number
- CN202512042388.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-02-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies are inefficient in identifying pyramid scheme hierarchies, lack the ability to identify hidden relationships, and are not dynamic enough to adapt to changing circumstances, making it difficult to identify multi-level networks and rapidly evolving pyramid schemes.
The hierarchical identification method based on big data analysis collects the financial and non-financial connections of participants' accounts as directed edges, uses a hierarchical behavioral causal coding algorithm to generate node embedding vectors, and combines a hierarchical influence radiation model and a real-time dynamic update mechanism to identify communities and high-level leaders.
It has improved the accuracy of identifying the hierarchical structure of pyramid schemes and the reliability of locating leaders, dynamically adapting to changes in network structure and enhancing the targeting of tracking and combating pyramid schemes.
Smart Images

Figure CN121481373A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, specifically to a hierarchical identification method based on big data analysis. Background Technology
[0002] Pyramid schemes are illegal organizations that profit by recruiting members and charging fees. They require participants to recruit others to earn money and have distinct hierarchical structures.
[0003] Patent application number 202410441089.9 discloses a method for identifying suspicious pyramid schemes, comprising: step S1, obtaining the registration information of each organization in the organization set to be identified, and removing previously legitimate organizations from the organization set to be identified based on the registration information of the organizations, wherein the organization set to be identified includes one or more organizations; step S2, collecting evaluation index information for each organization in the organization set to be identified, wherein the evaluation index information includes at least two evaluation indicators from the following: company name, company WeChat official account name, company mini-program name, company product name, company trademark, company shareholder information, and public opinion; step S3, inputting the evaluation indexes of the organizations in the organization set to be identified. Information is fed into a pre-established pyramid scheme identification model, which performs the following steps: calling the scoring algorithm for each evaluation indicator to obtain the score of each organization for each evaluation indicator; merging the scores of all evaluation indicators for each organization to obtain the comprehensive score of each organization; and filtering out a list of suspicious pyramid scheme organizations from the set of organizations to be identified based on the comprehensive scores of the organizations. This application aims to solve the problem that "the transaction flow information of an organization belongs to the privacy information of the enterprise, which requires the cooperation of the bank to obtain. The information is inconvenient to obtain and difficult to implement. Moreover, the identification of illegal pyramid scheme organizations is based on only one dimension of transaction flow, which does not make full use of the common characteristics of new pyramid scheme organizations in multiple dimensions, resulting in low robustness".
[0004] However, existing technologies still have the following shortcomings and deficiencies when it comes to identifying pyramid scheme hierarchies:
[0005] Manual analysis is inefficient: Traditional investigations of pyramid schemes rely on manual analysis of fund flows and personnel relationships. Multi-level networks (tens of thousands of nodes) require weeks of analysis and are prone to missing key hierarchical connections.
[0006] Insufficient identification of hidden relationships: Pyramid schemes evade detection by means of "multi-level transfer accounts" and "indirect downline development," lacking effective methods to quantify the influence of nodes in multiple indirect paths, making it difficult to identify core organizers without direct transfer records;
[0007] Insufficient dynamic adaptability: The structure of pyramid schemes iterates frequently, and existing rules or static models are difficult to adapt to the rapid evolution of network topology, making it impossible to update the hierarchical identification strategy in real time.
[0008] To address this, a hierarchical identification method based on big data analysis is proposed. Summary of the Invention
[0009] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a hierarchical identification method based on big data analysis, which can effectively solve the problems of the existing technology.
[0010] To achieve the above objectives, the present invention is implemented through the following technical solutions;
[0011] This invention discloses a hierarchical identification method based on big data analysis, comprising:
[0012] Using participants' accounts as nodes, directed edges are collected based on both financial and non-financial connections. Edge weights are determined by the impact of associated behaviors on the hierarchical structure. Based on the sequence of associated behaviors of nodes, a hierarchical behavioral causal coding algorithm is used to capture the hierarchical dependency relationship between behavior initiation and response to generate node embedding vectors. A hierarchical influence attenuation coefficient is set that decreases in a stepwise manner with hierarchical distance, so that the embedding vector represents the hierarchical position of the node in the behavioral chain. Based on the node embedding vectors, the collaborative mode of hierarchical behavior is analyzed, and the hierarchical synchronization rate of behavior initiation and response between nodes is calculated. Nodes with a synchronization rate higher than the collaborative benchmark are classified as potential communities. At the same time, based on the hierarchical transmission path of the associated edges in the graph, the functional complementarity of nodes in the hierarchical network is analyzed. Nodes with strong functional complementarity are recorded as the same community. By combining the above two classification results, only nodes that belong to the same community in both hierarchical behavioral collaboration and functional complementarity are retained as the final community members.
[0013] A node-level influence radiation model is constructed to calculate the influence radiation breadth, intensity, and stability of nodes within the community. Based on these three dimensions, the comprehensive value of hierarchical influence is estimated, and nodes are selected as high-level leader nodes. Multi-dimensional behavioral data of nodes is collected to verify their consistency in community detection and leader positioning. The hierarchical matching degree of behavioral features in different dimensions is calculated, and results with matching degrees lower than the consistency benchmark are re-verified. Based on newly collected related behavioral data in real time, the edge weights and node embedding vectors in the related network are continuously and dynamically updated according to a preset period, and the community division and hierarchical influence are recalculated synchronously.
[0014] Furthermore, the node attributes include the frequency of triggering associated behaviors and the number of times hierarchical instructions are transmitted, and the edge attributes include the duration of associated behaviors and the scope of their impact. The establishment of an edge must satisfy the requirement that the contribution rate of associated behaviors to the expansion of the offline network is higher than the group baseline value.
[0015] The fund association refers to: fund flow information a) from one account to multiple accounts split and transferred at a fixed ratio, and fund flow information b) from multiple accounts to the same account merged according to specific rules;
[0016] Based on the condition that the total amount of a single split or merger is greater than the threshold set based on the smallest operating unit of pyramid scheme funds, various types of fund associations are screened.
[0017] The non-funding association refers to: the hierarchical correspondence information x between nodes as training initiators and training participants, and the hierarchical correspondence information y between nodes as resource allocation instruction issuers and instruction receivers;
[0018] In this context, the edge attribute in x indicates whether the training was initiated or participated in, while the edge attribute in y indicates whether the allocation was guided or the allocation was executed.
[0019] Furthermore, the group benchmark value is a critical value determined by behavioral influence regression analysis of historical MLM network data. The critical value corresponds to the minimum value of the parameters of the frequency, duration and coverage of associated behaviors that enable the downstream network to expand effectively. This critical value is the group benchmark value.
[0020] The effective expansion conditions are: the number of new members reaches 15% of the initial number of members and the duration exceeds 20 days;
[0021] The edge weights, when determined, follow the following rules:
[0022] ;
[0023] In the formula: Let be the edge weight from node i to node j; This is the dimensional balance coefficient; The positive / negative hierarchical expansion contribution ratio of node j; Duration of associated behavior; The base time unit; The hierarchical control strength of node j; This represents the maximum level of control within the community. This refers to the hierarchical attenuation coefficient. The hierarchical distance is described by the shortest hierarchical path length; The standard deviation of the behavior sequence; The baseline standard deviation; The entropy weight coefficient; Entropy of behavioral patterns;
[0024] Among them, the entropy weight coefficient The value is set to a range of 0.1 to 0.8, and is consistent with the behavior pattern entropy. Proportional to behavioral pattern entropy The values are calculated based on the probability distribution of different behavior types in the instruction sequence from node i to node j.
[0025] Furthermore, the contribution rate of the associated behavior to the offline network expansion is expressed as follows:
[0026] ;
[0027] In the formula: For behavior Contribution rate to offline network expansion; The total number of valid actions; As weight; For behavior The contribution rate of the i-node hierarchical fission propagation after triggering; For behavior The number of new offline nodes added directly after triggering; These are model parameters; For behavior The depth of the hierarchy; For behavior Total number of triggers; For behavior The set of secondary downstream entities formed after triggering; This represents the activity index of the second-level downstream user k. For behavior The hierarchical conduction length;
[0028] Among them, behavior When conveying actions for instructions, =0.45, behavior When performing resource allocation behavior, =0.35, behavior When training organizational behavior, =0.2.
[0029] Furthermore, the node embedding vector dimension is set to 140 dimensions, which is determined through hierarchical discriminative analysis of behavioral features, wherein the core dimension corresponds to the hierarchical radiation range of the behavior initiation and the degree of coercion of the response.
[0030] The hierarchical discriminative power of the behavioral feature is the ratio of the variance of the distribution of the same feature at different hierarchical nodes to the overall variance of the distribution of the feature across all nodes.
[0031] The core dimension is a specific dimension in the embedded vector used to characterize the number and coverage of the lower-level nodes that can be reached by the hierarchical behavior initiated by the node, as well as the strength of the directive and binding force of the behavior initiated by the node in the process of being responded to by the lower-level nodes. Its value is positively correlated with the breadth of the radiation range and the strength of the forced response.
[0032] Furthermore, the functional complementarity of the nodes in the hierarchical network is quantified by the following formula:
[0033] ;
[0034] In the formula: Score the functional complementarity between nodes i and j; These are the weighting coefficients; Functional type difference matrix; Let i and j be the functional types; For the hierarchy levels of nodes i and j; Let be the shortest path length between nodes i and j; The frequency of interaction between nodes i and j; Let i be the set of child nodes of nodes i and j; It is the set of all subordinate nodes in the network; The total number of functional types; For a common set of subordinate nodes The proportion of functional type u in the middle;
[0035] in, All are positive numbers and their sum is 1. ∈[0,1], representing the degree of difference between function types i and j. middle Represents a set The total number of internal elements is determined by comparing the user-defined threshold for strong functional complementarity with the functional complementarity score between nodes. Nodes with functional complementarity scores greater than the threshold are recorded as nodes with strong functional complementarity.
[0036] Furthermore, in the high-level leader node selection stage, in the core area of the community where the influence radiation breadth is ≥70% and the radiation intensity is ≥80%, nodes with a comprehensive influence value higher than twice the community average and located in the core area are selected as high-level leader nodes.
[0037] Furthermore, the comprehensive value of the hierarchical influence is the sum of the influence's breadth, intensity, and stability.
[0038] The breadth of influence is represented by the proportion of nodes covered by its hierarchical influence, the intensity of influence is represented by the obedience rate of lower-level nodes to its instructions, and the stability of influence is represented by the duration of influence.
[0039] Among them, the influence radiation breadth, radiation intensity and radiation stability are normalized and dimensionless when performing the summation operation.
[0040] Furthermore, the multi-dimensional behavioral data includes fund association, training participation, resource allocation, and a stage for calculating the hierarchical matching degree of behavioral characteristics of different dimensions. At the same time, hierarchical behavioral anomaly identification is introduced to screen nodes whose surface hierarchical behavior does not match their actual influence and exclude such nodes from the candidate qualifications of leaders.
[0041] The hierarchical behavior anomaly identification includes: nodes that continuously participate in training but have no actual authority to initiate instructions;
[0042] The consistency verification logic of the behavioral data in community detection and leader location is expressed as follows: calculate the cosine similarity between the hierarchical feature vector of the node in each behavioral dimension and the results of community detection and leader location; if the weighted average of the cosine similarity of each dimension is higher than the preset consistency threshold, then it is determined that there is consistency.
[0043] The hierarchical matching degree of each dimension of behavioral features is calculated by: mapping each dimension of behavioral features to a hierarchical structure space, calculating the projection similarity matrix of feature vectors on the hierarchical dimension, combining the contribution weight of behavioral features to the hierarchical structure, and obtaining the hierarchical matching degree value by summing the weighted product of hierarchical projection cosine similarity and contribution.
[0044] The similarity matrix elements are calculated by dividing the dot product of the hierarchical projection vectors by the product of the vector magnitudes, and the contribution weights are determined based on the influence of features on the hierarchical instruction transmission path.
[0045] Compared with the known prior art, the technical solution provided by this invention has the following beneficial effects:
[0046] This invention provides a hierarchical identification method based on big data analysis. When applied to the hierarchical identification of pyramid schemes, this method integrates dynamic multidimensional association graphs and behavioral chain feature embedding. Unlike traditional methods that rely solely on fund transfers or simple social relationships, this method utilizes hierarchical behavior collaborative detection and influence radiation models to achieve deep coupling between community segmentation and leader positioning, overcoming the one-sidedness of single-dimensional analysis. Simultaneously, cross-dimensional consistency verification and real-time behavior change adjustment mechanisms can dynamically adapt to the concealment and hierarchical fluidity of pyramid scheme networks, solving the problem that static identification methods struggle to cope with the rapid evolution of network structures. This significantly improves the accuracy of hierarchical identification and the reliability of leader positioning, providing more targeted technical support for tracking and combating pyramid scheme hierarchies. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0048] Figure 1 This is a flowchart illustrating a hierarchical identification method based on big data analysis. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0050] The present invention will be further described below with reference to embodiments.
[0051] Example:
[0052] This embodiment presents a hierarchical identification method based on big data analysis, such as... Figure 1 As shown, it includes:
[0053] Using the accounts of participants in pyramid schemes as nodes, we collect directed edges with and without financial or non-financial connections, and determine the edge weights by examining the impact of related behaviors on the hierarchical structure.
[0054] Node attributes include the frequency of associated behavior triggers and the number of times hierarchical instructions are transmitted. Edge attributes include the duration of associated behavior and the scope of its impact. Furthermore, the establishment of an edge must satisfy the requirement that the contribution rate of associated behavior to the expansion of the offline network is higher than the group baseline value.
[0055] Funds are linked as follows: a) Funds flow information when one account transfers funds to multiple accounts in a fixed proportion; b) Funds flow information when multiple accounts transfer funds to the same account in a combined manner according to specific rules.
[0056] Based on the condition that the total amount of a single split or merger is greater than the threshold set based on the smallest operating unit of pyramid scheme funds, various types of fund associations are screened.
[0057] Non-funding associations are: x, the hierarchical correspondence information between nodes as training initiators and training participants; and y, the hierarchical correspondence information between nodes as resource allocation instruction issuers and instruction receivers.
[0058] Among them, the edge attribute of x is labeled as training initiation or training participation, and the edge attribute of y is labeled as allocation guidance or execution allocation;
[0059] The group baseline value is a critical value determined by behavioral influence regression analysis of historical MLM network data. The critical value corresponds to the minimum value of the parameters of the frequency, duration and coverage of associated behaviors that enable the downstream network to expand effectively. This critical value is the group baseline value.
[0060] It should be noted that during the regression analysis phase, either linear regression or logistic regression can be used to perform the analysis.
[0061] The effective expansion conditions are: the number of new members reaches 15% of the initial number of members and the duration exceeds 20 days;
[0062] When edge weights are determined, they follow the following rules:
[0063] ;
[0064] In the formula: Let be the edge weight from node i to node j; This is the dimensional balance coefficient; The positive / negative hierarchical expansion contribution ratio of node j; Duration of associated behavior; The base time unit; The hierarchical control strength of node j; This represents the maximum level of control within the community. This refers to the hierarchical attenuation coefficient. The hierarchical distance is described by the shortest hierarchical path length; The standard deviation of the behavior sequence; The baseline standard deviation; The entropy weight coefficient; Entropy of behavioral patterns;
[0065] Among them, the entropy weight coefficient The value is set to a range of 0.1 to 0.8, and is consistent with the behavior pattern entropy. Proportional to behavioral pattern entropy The value is calculated based on the probability distribution of different behavior types in the instruction sequence of node i to node j;
[0066] The logic for determining the hierarchical control strength is as follows: the core quantitative indicators are the frequency of instruction transmission by the node, the obedience rate of the subordinate, the number of resource allocation instructions initiated, the proportion of training dominance, and the dominance of the capital flow. The initial score is calculated by weighting the instruction control (0.35), resource allocation (0.25), training dominance (0.15), capital flow control (0.20), and hierarchical expansion contribution (0.05). After normalization and mapping to the [0,1] interval, the score is corrected by combining the hierarchical distance (d) with the attenuation coefficient (λ) (corrected value = initial value × (1-λ×d)).
[0067] The above formula comprehensively considers multiple factors to determine the edge weight from node i to node j. Specifically, the dimensionality balance coefficient balances the influence of different dimensional parameters on the weight; the positive / negative hierarchical expansion contribution ratio of node j reflects its role in hierarchical expansion; the ratio of the duration of the associated behavior to the baseline time unit reflects the impact of the behavior's duration; the ratio of the hierarchical control strength of node j to the maximum control strength within the community reflects the relative level of the node's control strength; the hierarchical decay coefficient and hierarchical distance consider the decay effect in hierarchical transmission; and the relationship between the standard deviation of the behavior sequence and the baseline standard deviation, as well as the entropy weight coefficient and behavior pattern entropy, incorporate factors such as the stability and diversity of behavior patterns. This formula, to a certain extent, overcomes the limitations of determining edge weights by a single factor. By integrating multiple parameters related to hierarchical structure, it can more comprehensively and accurately characterize the strength of associations between nodes, enabling edge weights to more realistically reflect the impact of associated behavior on the hierarchical structure, thus laying a more reliable foundation for subsequent hierarchical identification.
[0068] The contribution rate of associated behavior to the expansion of the offline network is expressed as:
[0069] ;
[0070] In the formula: For behavior Contribution rate to offline network expansion; The total number of valid actions; As weight; For behavior The contribution rate of the i-node hierarchical fission propagation after triggering; For behavior The number of new offline nodes added directly after triggering; These are model parameters; For behavior The depth of the hierarchy; For behavior Total number of triggers; For behavior The set of secondary downstream entities formed after triggering; This represents the activity index of the second-level downstream user k. For behavior The hierarchical conduction length;
[0071] Among them, behavior When conveying actions for instructions, =0.45, behavior When performing resource allocation behavior, =0.35, behavior When training organizational behavior, =0.2;
[0072] The above formula is based on the total number of effective behaviors, combined with the weight of different behaviors, and comprehensively considers factors such as the contribution rate of the hierarchical fission transmission of node i after the behavior is triggered, the number of directly added downstreams, model parameters, hierarchical depth, total number of triggers, set of secondary downstreams, activity index of secondary downstreams, and hierarchical transmission length, to calculate the contribution rate of a specific behavior to the expansion of the downstream network.
[0073] It should be noted that the model parameters The sum should be 1, and it should obey the following rules: , The initial values are set to 0.4, 0.35, and 0.25. Value 'a' represents the immediate impact weight of the number of new downlines directly added after a behavior is triggered, determined based on the proportion of newly added downlines to the initial members (e.g., the contribution to reaching the 15% effective expansion threshold). Value 'b' represents the indirect impact of hierarchical fission transmission, related to the activity index of second-level downlines and the length of hierarchical transmission; the higher the activity and the longer the transmission, the higher the value of 'b'. Value 'c' reflects the cumulative impact of continuous behavior triggering, weighted by the total number of behavior triggers and the depth of the hierarchy (e.g., the continuous impact of higher-level behaviors is more significant). Finally, through regression analysis of behavior and downline expansion in historical pyramid schemes, the ratio of these three values is calibrated to match the actual expansion effect.
[0074] In addition, the activity index of the second-level downstream k ( The value range is [0,1], and it is determined by weighted calculation of four core behavioral indicators:
[0075] The following indicators were used: activity level of fund interaction (weight 0.3, based on the frequency and amount of fund transfers with superiors / peers), depth of training participation (weight 0.2, based on the number and duration of training sessions), efficiency of instruction execution (weight 0.35, based on the timeliness of response to superiors' instructions and the rate of obedience), and contribution to hierarchical expansion (weight 0.15, based on the number of directly added third-level downlines and the duration of their maintenance). These indicators were normalized to the [0,1] interval and then summed according to their weights. Abnormal nodes that only participated in training but had no fund interaction or executed instructions were excluded. Set the value to ≤0.2;
[0076] This formula not only focuses on the increase in the number of downstream users directly caused by the behavior, but also incorporates factors such as hierarchical fission transmission, the depth and breadth of the behavior, and assigns different weights to different types of behavior. It can more meticulously evaluate the actual contribution of each behavior in network expansion and improve the accuracy and pertinence of contribution rate calculation.
[0077] Based on the node-related behavior sequence, a hierarchical behavior causal coding algorithm is used to capture the hierarchical dependency relationship between behavior initiation and response to generate node embedding vectors. A hierarchical influence attenuation coefficient is set that decreases in a stepwise manner with hierarchical distance, so that the embedding vector represents the hierarchical position of the node in the behavior chain.
[0078] The node embedding vector dimension is set to 140 dimensions. The node embedding vector dimension is determined through hierarchical discrimination analysis of behavioral features. The core dimension corresponds to the hierarchical radiation range of the behavior initiation and the degree of coercion of the response.
[0079] The hierarchical discriminative power of a behavioral feature is the ratio of the variance of the distribution of the same feature at different hierarchical nodes to the overall variance of the distribution of the same feature across all nodes.
[0080] The core dimension is a specific dimension in the embedding vector that represents the number and coverage of the lower-level nodes that can be reached by the hierarchical behavior initiated by the node, as well as the strength of the directive and binding force of the behavior initiated by the node in the process of being responded to by the lower-level nodes. The value of this dimension is positively correlated with the breadth of the radiation range and the strength of the forced response.
[0081] Based on node embedding vector analysis of hierarchical behavior collaboration patterns, the hierarchical synchronization rate of behavior initiation and response between nodes is calculated. Nodes with a synchronization rate higher than the collaboration benchmark are classified as potential communities. At the same time, based on the hierarchical transmission path of the associated edges in the graph, the functional complementarity of nodes in the hierarchical network is analyzed. Nodes with strong functional complementarity are recorded as the same community. By combining the above two classification results, only nodes that belong to the same community in both hierarchical behavior collaboration and functional complementarity are retained as the final community members.
[0082] The functional complementarity of nodes in a hierarchical network is quantified by the following formula:
[0083] ;
[0084] In the formula: Score the functional complementarity between nodes i and j; These are the weighting coefficients; Functional type difference matrix; Let i and j be the functional types; For the hierarchy levels of nodes i and j; Let be the shortest path length between nodes i and j; The frequency of interaction between nodes i and j; Let i be the set of child nodes of nodes i and j; It is the set of all subordinate nodes in the network; The total number of functional types; For a common set of subordinate nodes The proportion of functional type u in the middle;
[0085] in, All are positive numbers and their sum is 1. ∈[0,1], representing the degree of difference between function types i and j. middle Represents a set The total number of internal elements is determined by the user-defined threshold for strong functional complementarity and compared with the functional complementarity score between nodes. Nodes with functional complementarity scores greater than the threshold for strong functional complementarity are recorded as nodes with strong functional complementarity.
[0086] Among them, the strong functional complementarity judgment threshold is used as a numerical limit to measure the degree of functional complementarity between nodes. When the functional complementarity index between nodes exceeds the threshold, it is considered that these nodes have a strong functional complementarity relationship in the pyramid scheme organization, which can be used as the basis for judging the same community members.
[0087] The above formula integrates factors such as the functional type difference matrix, node functional type, hierarchical level, shortest path length, interaction frequency, set of subordinate nodes, total number of functional types, and proportion of functional types in the common set of subordinate nodes through weighting coefficients. It quantifies the functional complementarity between nodes, measures the functional complementarity between nodes from multiple dimensions such as functional differences, hierarchical relationship, interaction, and functional distribution of subordinate nodes, and allows users to customize the threshold for strong functional complementarity, which enhances the flexibility and applicability of the formula and can more accurately identify node combinations with strong functional complementarity.
[0088] A node-level influence radiation model is constructed to calculate the influence radiation breadth, intensity, and stability of nodes within the community. Based on the above three-dimensional indicators, the comprehensive value of the hierarchical influence is estimated, and nodes are selected as high-level leader nodes.
[0089] In the high-level leader node selection phase, in the core area of the community where the influence radiation breadth is ≥70% and the radiation intensity is ≥80%, nodes with a comprehensive level influence value higher than twice the community average and located in the core area are selected as high-level leader nodes.
[0090] The comprehensive value of hierarchical influence is the sum of the breadth, intensity, and stability of influence.
[0091] The breadth of influence is represented by the proportion of nodes covered by its hierarchical influence, the intensity of influence is represented by the obedience rate of lower-level nodes to its commands, and the stability of influence is represented by the duration of influence.
[0092] Among them, the influence radiation breadth, radiation intensity and radiation stability are normalized and dimensionless when performing the summation operation.
[0093] Collect multi-dimensional behavioral data of nodes to verify their consistency in community detection and leader location, calculate the hierarchical matching degree of behavioral features in different dimensions, and re-verify results with matching degrees lower than the consistency benchmark.
[0094] Multi-dimensional behavioral data includes financial connections, training participation, and resource allocation. The hierarchical matching degree of behavioral characteristics in different dimensions is calculated. At the same time, hierarchical behavioral anomaly identification is introduced to investigate nodes whose surface hierarchical behavior does not match their actual influence and exclude such nodes from the candidate qualifications of leaders.
[0095] Hierarchical behavioral anomaly identification includes: nodes that continuously participate in training but have no actual authority to initiate instructions;
[0096] The consistency verification logic of behavioral data in community detection and leader localization is expressed as follows: calculate the cosine similarity between the hierarchical feature vector of the node in each behavioral dimension and the results of community detection and leader localization. If the weighted average of the cosine similarity of each dimension is higher than the preset consistency threshold, then it is determined to be consistent.
[0097] When calculating the hierarchical matching degree of behavioral features in each dimension, the following is followed: mapping the behavioral features in each dimension to the hierarchical structure space, calculating the projection similarity matrix of the feature vectors on the hierarchical dimension, combining the contribution weight of the behavioral features to the hierarchical structure, and obtaining the hierarchical matching degree value according to the summation method of the weighted product of the hierarchical projection cosine similarity and the contribution.
[0098] The similarity matrix elements are calculated by dividing the dot product of the hierarchical projection vectors by the product of the vector magnitudes, and the contribution weights are determined based on the influence of the features on the hierarchical instruction transmission path.
[0099] For results with a matching degree lower than the consistency benchmark, a "hierarchical behavior tracing chain" is constructed for re-verification: tracing the source of node behavior in dimensions such as fund association, training participation, and resource allocation, extracting the initiating node, responding node, and transmission path of each dimension's behavior, and analyzing whether different dimensions' behaviors point to the same high-level node (e.g., whether the final destination node of fund transfer is consistent with the initiating node of training instructions); simultaneously establishing a "behavioral contradiction feature library" to record typical patterns with low matching degree (e.g., "weak financial influence but strong training instruction authority" and "high short-term behavior level but low long-term behavior level"). For nodes falling into the feature library, calculating their hierarchical fluctuation coefficient in each dimension's behavior (fluctuation coefficient = difference between different dimension levels / average level), and filtering nodes with a fluctuation coefficient > 0.3 for "behavioral motivation deduction"—by reconstructing the node's behavioral triggering scenario (e.g., the context when the instruction is initiated, the accompanying behavior of fund transfer), determining whether there is any deliberate disguise behavior to hide the level (e.g., using a low fund level to cover up a high instruction level), and finally combining the tracing results, fluctuation coefficient, and motivation deduction to correct the node's hierarchical affiliation and leader candidate qualifications.
[0100] Based on real-time collected new association behavior data, the edge weights and node embedding vectors in the association network are continuously and dynamically updated according to a preset period, and the community division and hierarchical influence are recalculated synchronously.
[0101] In this embodiment, through multi-dimensional correlation analysis and precise weight calculation, the hierarchical structure of pyramid scheme networks can be efficiently identified and high-level leaders can be accurately located. The accuracy of identification is improved by community segmentation and functional complementarity analysis, and the dynamic update mechanism adapts to network changes. The reliability of the results is ensured by multi-dimensional behavioral verification, which can reduce missed and false judgments, provide strong support for the rapid crackdown on pyramid schemes, and significantly improve the efficiency and accuracy of pyramid scheme governance.
[0102] The following is an application example of the method described in the above embodiments:
[0103] I. Network Construction and Edge Selection
[0104] In a certain pyramid scheme case, 300 participating accounts were used as nodes, and the connections between funds and non-funds were collected as directed edges to construct the network.
[0105] Funds linkage: Account A (upstream) splits and transfers funds to Accounts B, C, and D in a 4:3:3 ratio, with a single total amount of 50,000 yuan (higher than the 10,000 yuan threshold set for the smallest operating unit of pyramid scheme funds), forming fund flow information a; Accounts B, C, and D transfer funds to A monthly in a combined manner, with a total amount of 30,000 yuan (higher than the threshold), forming fund flow information b.
[0106] Non-funding association: A, as the training initiator, organizes B and C to participate in the training, forming a hierarchical relationship information x (edge attribute labeled "training initiator" and "training participant"); A issues a resource allocation instruction to B, and B executes the allocation, forming relationship information y (edge attribute labeled "allocation guide" and "execute allocation").
[0107] Edge establishment: The association behavior of A to B lasts for 45 days (higher than the group baseline of 25 days), affects 3 downstream nodes (higher than the baseline of 1), triggers 8 times (higher than the baseline of 3 times), and this behavior causes B's downstream network to add 18% (reaching 15%) and maintains it for 25 days (more than 20 days). The contribution rate is higher than the group baseline value. Therefore, a directed edge from A to B is established.
[0108] II. Quantization of Node and Edge Attributes
[0109] Node attributes: A's associated behavior is triggered 12 times and hierarchical instructions are transmitted 9 times; B's trigger frequency is 5 times and instructions are transmitted 3 times.
[0110] Edge attributes: The association behavior from A to B lasts for 45 days, affecting 3 nodes; the association behavior from B to its downstream nodes lasts for 30 days, affecting 2 nodes.
[0111] III. Edge Weight and Node Embedding Calculation
[0112] Edge weight: Calculated by formula, the edge weight from A to B is 0.82 (taking into account parameters such as B's positive expansion contribution ratio of 1.2, duration of 45 days (baseline time unit is 10 days), B's control strength accounts for 60% of the community's maximum control strength, hierarchical distance of 1, and behavioral sequence standard deviation of 0.3 (baseline standard deviation of 0.5); the edge weight from B to its downstream is 0.56.
[0113] Node embedding vector: Generate a 140-dimensional vector. The core dimension of A shows that its hierarchical radiation range covers 8 nodes, and the execution obedience rate of its subordinates to its instructions reaches 92%; the core dimension of B shows that its radiation range covers 3 nodes, with an obedience rate of 85%.
[0114] IV. Community Segmentation and Functional Complementarity Analysis
[0115] Potential community identification: The behavioral synchronization rate of A, B, and C was 88% (higher than the collaborative benchmark of 70%), and they were classified as potential communities.
[0116] Functional complementarity analysis: The functional complementarity score between A (instruction initiator) and B (execution allocation) is 0.85 (calculated based on functional type difference of 0.2, hierarchical distance of 1, interaction frequency of 15 times / month, and common subordinate ratio of 40%, which is higher than the strong complementarity threshold of 0.7), therefore they are determined to be the same community. After merging, A, B, C, and D form the final community.
[0117] V. Positioning of High-Level Leaders
[0118] The three dimensions of influence are: A's influence reach is 85% (covering 85% of nodes in the community), influence intensity is 90% (subordinate obedience rate), and influence stability is 180 days (duration). The sum of the three after normalization is a comprehensive value of 260.
[0119] Screening results: The comprehensive value is more than twice the community average (120) (240), and A is located in the core area of the community (radiation breadth ≥70%, intensity ≥80%), so A is determined to be a high-level leader node.
[0120] VI. Consistency Verification and Dynamic Update
[0121] Multi-dimensional verification: The weighted average cosine similarity between the hierarchical feature vectors of A in the dimensions of fund flow, training organization, and resource allocation and the location results is 0.89 (higher than the consistency threshold of 0.7), indicating no contradictory features. Node E (continuously involved in training but without the right to initiate instructions) was identified and excluded from the candidate status of leader.
[0122] Dynamic updates: New data is collected monthly, the edge weight from A to B is updated to 0.85, the node embedding vector is adjusted to cover 9 nodes, and the community and leader positioning results remain stable.
[0123] Furthermore, pyramid schemes are currently shifting from offline to online, using the internet as their main platform and leveraging tools such as social media, recruitment platforms, and cryptocurrencies to overcome geographical limitations. Participants do not need to meet in person; they can complete their multi-level marketing through online registration and encrypted payments, making them extremely difficult to detect. They often rent shops in prime locations, using "physical stores + multiple business licenses" as a cover, disguising themselves as entrepreneurial projects, such as health water / health products / Yuan Yu Zhou investment. On the surface, they claim to sell products, but in reality, 90% of their profits come from rebates for developing downline levels.
[0124] This new type of pyramid scheme enhances its concealment through a dual-track model of "online virtual operation + offline physical disguise": online, it relies on social media groups (such as WeChat and QQ groups) to release tiered tasks, and through recruitment platforms (such as a certain recruitment APP) to develop downlines under the guise of "high-salary entrepreneurship" and "partner program". New members are required to complete a specified amount of virtual currency (such as self-created digital currency, USDT, etc.) as an entry fee, and the amount of recharge is directly linked to the level promotion - recharge of 10,000 yuan is "junior agent" and can develop 3 downlines; recharge of 50,000 yuan is "senior agent" and the downline rebate ratio increases to 30%, forming a clear "investment level - level permission" binding relationship.
[0125] To evade fund tracking, they employed a hybrid payment system: some entry fees were transferred peer-to-peer via cryptocurrency wallets, with no clear account entity associated with the funds; others were processed through physical store POS machines by credit card swipes, disguised as "product purchases," but the funds were immediately transferred to a core account upon receipt. Simultaneously, they used multiple business licenses to disperse fund flows, with each license corresponding to different online promotion channels, keeping the transaction volume of a single account below the suspicious threshold. Traditional methods relying solely on single-account fund analysis were insufficient to penetrate their multi-layered fund splitting system.
[0126] This solution penetrates online and offline disguises through multi-dimensional correlation. The technical solution can effectively identify the core levels hidden under the guise of "startup projects." For example, in a certain health water pyramid scheme, by analyzing the virtual currency transfer chain, it was found that the entry fees of all downlines eventually flowed to three virtual currency wallets. Moreover, the social media accounts corresponding to these three wallets have been posting hierarchical tasks for a long time. Their influence covers 85% of the nodes and the command compliance rate reaches 92%. After cross-dimensional verification (the flow of virtual currency matches the information of the actual controller of the physical store), it was finally identified as a high-level leader node, which can be applied to both online and offline pyramid scheme scenarios.
[0127] In summary, the methods described in the above embodiments, by integrating dynamic multidimensional association graphs and behavioral chain feature embedding, differ from the limitations of traditional methods that rely solely on fund transfers or simple social relationships. Utilizing hierarchical behavioral collaborative detection and influence radiation models, they achieve deep coupling between community segmentation and leader positioning, overcoming the one-sidedness of single-dimensional analysis. Simultaneously, cross-dimensional consistency verification and real-time behavioral change adjustment mechanisms can dynamically adapt to the concealment and hierarchical fluidity of pyramid scheme networks, solving the problem that static identification methods struggle to cope with the rapid evolution of network structures. This significantly improves the accuracy of hierarchical identification and the reliability of leader positioning, providing more targeted technical support for tracking and combating pyramid scheme hierarchies.
[0128] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hierarchical identification method based on big data analysis, characterized in that, include: Using the accounts of participants in pyramid schemes as nodes, we collect directed edges with and without financial or non-financial connections, and determine the edge weights by examining the impact of related behaviors on the hierarchical structure. Based on the node-related behavior sequence, a hierarchical behavior causal coding algorithm is used to capture the hierarchical dependency relationship between behavior initiation and response to generate node embedding vectors. A hierarchical influence attenuation coefficient is set that decreases in a stepwise manner with hierarchical distance, so that the embedding vector represents the hierarchical position of the node in the behavior chain. Based on node embedding vector analysis of hierarchical behavior collaboration patterns, the hierarchical synchronization rate of behavior initiation and response between nodes is calculated. Nodes with a synchronization rate higher than the collaboration benchmark are classified as potential communities. At the same time, based on the hierarchical transmission path of the associated edges in the graph, the functional complementarity of nodes in the hierarchical network is analyzed. Nodes with strong functional complementarity are recorded as the same community. By combining the above two classification results, only nodes that belong to the same community in both hierarchical behavior collaboration and functional complementarity are retained as the final community members. A node-level influence radiation model is constructed to calculate the influence radiation breadth, intensity, and stability of nodes within the community. Based on the above three-dimensional indicators, the comprehensive value of the hierarchical influence is estimated, and nodes are selected as high-level leader nodes. Collect multi-dimensional behavioral data of nodes to verify their consistency in community detection and leader location, calculate the hierarchical matching degree of behavioral features in different dimensions, and re-verify results with matching degrees lower than the consistency benchmark. Based on real-time collected new association behavior data, the edge weights and node embedding vectors in the association network are continuously and dynamically updated according to a preset period, and the community division and hierarchical influence are recalculated synchronously.
2. The hierarchical identification method based on big data analysis according to claim 1, characterized in that, The node attributes include the frequency of triggering associated behaviors and the number of times hierarchical instructions are transmitted. The edge attributes include the duration of associated behaviors and the scope of their impact. Furthermore, the establishment of an edge must satisfy the requirement that the contribution rate of associated behaviors to the expansion of the offline network is higher than the group baseline value. The fund association refers to: fund flow information a) from one account to multiple accounts split and transferred at a fixed ratio, and fund flow information b) from multiple accounts to the same account merged according to specific rules; Based on the condition that the total amount of a single split or merger is greater than the threshold set based on the smallest operating unit of pyramid scheme funds, various types of fund associations are screened. The non-funding association refers to: the hierarchical correspondence information x between nodes as training initiators and training participants, and the hierarchical correspondence information y between nodes as resource allocation instruction issuers and instruction receivers; In this context, the edge attribute in x indicates whether the training was initiated or participated in, while the edge attribute in y indicates whether the allocation was guided or the allocation was executed.
3. The hierarchical identification method based on big data analysis according to claim 2, characterized in that, The group benchmark value is a critical value determined by behavioral influence regression analysis of historical MLM network data. The critical value corresponds to the minimum value of the parameters of the frequency, duration and coverage of associated behaviors that enable the downstream network to expand effectively. This critical value is the group benchmark value. The effective expansion conditions are: the number of new members reaches 15% of the initial number of members and the duration exceeds 20 days; The edge weights, when determined, follow the following rules: ; In the formula: Let be the edge weight from node i to node j; This is the dimensional balance coefficient; The positive / negative hierarchical expansion contribution ratio of node j; Duration of associated behavior; The base time unit; The hierarchical control strength of node j; This represents the maximum level of control within the community. This refers to the hierarchical attenuation coefficient. The hierarchical distance is described by the shortest hierarchical path length; The standard deviation of the behavior sequence; The baseline standard deviation; The entropy weight coefficient; Entropy of behavioral patterns; Among them, the entropy weight coefficient The value is set to a range of 0.1 to 0.8, and is consistent with the behavior pattern entropy. Proportional to behavioral pattern entropy The values are calculated based on the probability distribution of different behavior types in the instruction sequence from node i to node j.
4. The hierarchical identification method based on big data analysis according to claim 2, characterized in that, The contribution rate of the associated behavior to the offline network expansion is expressed as follows: ; In the formula: For behavior Contribution rate to offline network expansion; The total number of valid actions; As weight; For behavior The contribution rate of the i-node hierarchical fission propagation after triggering; For behavior The number of new offline nodes added directly after triggering; These are model parameters; For behavior The depth of the hierarchy; For behavior Total number of triggers; For behavior The set of secondary downstream entities formed after triggering; This represents the activity index of the second-level downstream user k. For behavior The hierarchical conduction length; Among them, behavior When conveying actions for instructions, =0.45, behavior When performing resource allocation behavior, =0.35, behavior When training organizational behavior, =0.
2.
5. The hierarchical identification method based on big data analysis according to claim 1, characterized in that, The node embedding vector dimension is set to 140 dimensions. The node embedding vector dimension is determined through hierarchical discrimination analysis of behavioral features, where the core dimension corresponds to the hierarchical radiation range of the behavior and the degree of coercion of the response. The hierarchical discriminative power of the behavioral feature is the ratio of the variance of the distribution of the same feature at different hierarchical nodes to the overall variance of the distribution of the feature across all nodes. The core dimension is a specific dimension in the embedded vector used to characterize the number and coverage of the lower-level nodes that can be reached by the hierarchical behavior initiated by the node, as well as the strength of the directive and binding force of the behavior initiated by the node in the process of being responded to by the lower-level nodes. Its value is positively correlated with the breadth of the radiation range and the strength of the forced response.
6. The hierarchical identification method based on big data analysis according to claim 1, characterized in that, The functional complementarity of the nodes in the hierarchical network is quantified by the following formula: ; In the formula: Score the functional complementarity between nodes i and j; These are the weighting coefficients; Functional type difference matrix; Let i and j be the functional types; For the hierarchy levels of nodes i and j; Let be the shortest path length between nodes i and j; The frequency of interaction between nodes i and j; Let i be the set of child nodes of nodes i and j; It is the set of all subordinate nodes in the network; The total number of functional types; For a common set of subordinate nodes The proportion of functional type u in the middle; in, All are positive numbers and their sum is 1. ∈[0,1], representing the degree of difference between function types i and j. middle Represents a set The total number of internal elements is determined by comparing the user-defined threshold for strong functional complementarity with the functional complementarity score between nodes. Nodes with functional complementarity scores greater than the threshold are recorded as nodes with strong functional complementarity.
7. The hierarchical identification method based on big data analysis according to claim 1, characterized in that, In the high-level leader node selection stage, in the core area of the community where the influence radiation breadth is ≥70% and the radiation intensity is ≥80%, nodes with a comprehensive influence value of more than twice the community average and located in the core area are selected as high-level leader nodes.
8. The hierarchical identification method based on big data analysis according to claim 1, characterized in that, The comprehensive value of the hierarchical influence is the sum of the breadth, intensity, and stability of the influence's reach. The breadth of influence is represented by the proportion of nodes covered by its hierarchical influence, the intensity of influence is represented by the obedience rate of lower-level nodes to its instructions, and the stability of influence is represented by the duration of influence. Among them, the influence radiation breadth, radiation intensity and radiation stability are normalized and dimensionless when performing the summation operation.
9. The hierarchical identification method based on big data analysis according to claim 1, characterized in that, The multi-dimensional behavioral data includes fund association, training participation, and resource allocation. The hierarchical matching degree calculation stage of different dimensions of behavioral characteristics is also introduced. At the same time, hierarchical behavior anomaly identification is introduced to investigate nodes whose surface hierarchical behavior does not match their actual influence and exclude such nodes from the candidate qualifications of leaders. The hierarchical behavior anomaly identification includes: nodes that continuously participate in training but have no actual authority to initiate instructions; The consistency verification logic of the behavioral data in community detection and leader location is expressed as follows: calculate the cosine similarity between the hierarchical feature vector of the node in each behavioral dimension and the results of community detection and leader location; if the weighted average of the cosine similarity of each dimension is higher than the preset consistency threshold, then it is determined that there is consistency. The hierarchical matching degree of each dimension of behavioral features is calculated by: mapping each dimension of behavioral features to a hierarchical structure space, calculating the projection similarity matrix of feature vectors on the hierarchical dimension, combining the contribution weight of behavioral features to the hierarchical structure, and obtaining the hierarchical matching degree value by summing the weighted product of hierarchical projection cosine similarity and contribution. The similarity matrix elements are calculated by dividing the dot product of the hierarchical projection vectors by the product of the vector magnitudes, and the contribution weights are determined based on the influence of features on the hierarchical instruction transmission path.
Citation Information
Patent Citations
Suspicious transmission organization identification method, device and equipment
CN118396641A