Big data comprehensive supervision platform based on intelligent decision

By building a full-link mechanism based on Transformer control of DiffPool structure and isolated forest algorithm, the shortcomings of the existing technology in heterogeneous graph modeling and dynamic risk identification are solved, and an efficient and intelligent supervision platform is realized to adapt to multi-tasking, heterogeneous high-dimensional, and dynamic risk supervision scenarios.

CN120448993AInactive Publication Date: 2025-08-08SHANXI ZHONGWEI INFORMATION ENG CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510634935.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing big data supervision technology has significant flaws in key links such as heterogeneous graph structure modeling, context-driven graph aggregation mechanism, structural disturbance enhancement, comparison training optimization, linkage of abnormal node identification and risk response, and cannot achieve high-precision and strong robust supervision in high-dimensional, multi-task-driven, and dynamic risk evolution scenarios.

Method used

The heterogeneous graph construction module, node feature generation module, structure aggregation module, graph enhancement module, comparison training module, exception detection module and response scheduling module are adopted, and the full link mechanism is built to realize efficient modeling and intelligent response of multi-source regulatory data.

Benefits of technology

The structural expression ability, abnormal identification accuracy and intelligent response scheduling of the regulatory platform have been improved, and the entire process from perceptual modeling to response scheduling has been realized, and the efficient supervision of complex supervision scenarios has been adapted to.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448993A_ABST
    Figure CN120448993A_ABST
Patent Text Reader

Abstract

The invention discloses a big data comprehensive supervision platform based on intelligent decision, and the platform comprises a heterogeneous graph construction module which is used for constructing a supervision graph structure; the node feature generation module is used for generating node embedding representation; the structure aggregation module is used for generating a pooling graph on the basis of a DiffPool structure controlled by Transform; the graph enhancement module is used for generating a graph enhancement sample pair; the comparison training module is used for optimizing embedded representation; the anomaly detection module is used for identifying an abnormal node set; the risk map generation module is used for constructing a risk map structure; and the response scheduling module is used for generating and executing a supervision response instruction. According to the invention, the intelligent level and the structure expression ability of the supervision task are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data supervision and intelligent decision-making technology, and in particular to a big data comprehensive supervision platform based on intelligent decision-making. Background Art

[0002] With the widespread application of big data in urban governance, financial supervision, network security and integrated decision-making, how to achieve efficient perception, association modeling and intelligent identification of complex multi-source data has become a core issue that modern regulatory platforms urgently need to solve. Traditional big data supervision methods mostly use rule template matching, indicator threshold alarms or static graph analysis. Although they have certain adaptability in structured data scenarios, they are still difficult to effectively capture high-dimensional risk factors in the dynamic evolution process when faced with multiple types of regulatory objects, multi-level behavioral events, and diversified management rules. These methods usually rely on predefined rules for trigger analysis, lack the ability to structurally model the context of regulatory tasks, and find it difficult to build a flexible and scalable multi-agent collaborative perception mechanism.

[0003] In recent years, graph neural networks have been gradually introduced into the modeling and analysis of regulatory data. By mapping regulatory objects, behavioral links, and rule logic into graph structures, some studies have made progress in expressiveness and modeling of dependencies between nodes. However, most existing methods use isomorphic graph modeling, which makes it difficult to characterize the heterogeneity of different types of nodes and edges in regulatory semantics. In practical applications, regulatory objects often involve different types of structures such as event subjects, rule nodes, and behavioral sequences. Their interactive behaviors have significant structural differences and semantic dependencies. Traditional unified modeling methods cannot effectively express these complex associations, resulting in loss or offset of embedded representation information. In addition, node attributes and edge attributes are often simplified in existing graph modeling, lacking high-dimensional fusion of interactive behavior frequency, intensity, and time information, resulting in insufficient expression of temporal and dynamic features of graph representation.

[0004] To address the hierarchical abstraction of graph structures, existing methods have attempted to merge nodes using graph pooling mechanisms such as DiffPool and TopKPool to achieve hierarchical representation of the graph. However, these methods face numerous challenges in regulatory tasks. Current mainstream graph pooling mechanisms often generate node allocation matrices based on structural information itself, lacking the ability to model regulatory task types and dynamically adjusting aggregation strategies based on task context. During the pooling process, the attribution relationships between nodes fail to incorporate regulatory semantics and risk sensitivity assessments, leading to redundancy, misclassification, or semantic fragmentation in information aggregation. Furthermore, the construction of the allocation matrix lacks structure-preserving constraints and sparsity control, causing localized damage to the graph structure after aggregation, which in turn affects the perception of the graph's overall structure in subsequent recognition tasks. More importantly, node embedding representations are often static after pooling and cannot effectively reflect embedding changes under perturbations, resulting in insufficient graph structure perception of local perturbations.

[0005] In terms of graph enhancement and optimization training, some studies have attempted to use graph perturbation strategies combined with contrastive learning mechanisms to improve the robustness of the model. However, existing methods mostly remain at low-level processing methods such as node deletion and edge perturbation, and have not formed a complete closed loop of structural perturbation, embedding update and graph contrast optimization. Especially in dynamic risk modeling for regulatory tasks, there are fine-grained semantic differences between graph structure samples, and it is difficult to form an effective training sample space through basic perturbation methods alone. At the same time, most contrastive training strategies do not perform embedding optimization on homologous graph pairs of regulatory graphs, lack modeling of sensitivity to structural changes, and cannot drive the model to actively perceive abnormal behavior or potential deviations.

[0006] In terms of anomaly detection, algorithms based on path splitting mechanisms such as isolation forests have been applied to anomaly recognition in graph embedding space, but their detection performance is still limited by multiple factors. The current anomaly detection process mostly constructs scoring functions based on the density and distribution characteristics of the embedding representation, and fails to integrate them with the attribution relationship, node connection strength and aggregation level information in the pooled graph structure. This makes the scoring results of abnormal nodes susceptible to noise interference and lacks semantic interpretation capabilities. Since the clustering mechanism and connection mapping strategy between structures are not introduced, the detection results are difficult to be directly converted into graph structure results with risk expression significance. The regulatory system still relies on external reasoning or manual evaluation to generate executable risk maps, and the processing chain is incomplete.

[0007] Existing methods also have shortcomings in matching risk result output with task scheduling. Although some systems support generating response instructions through policy templates, most solutions are based on rule table lookup or manually configured paths to complete the response process, and lack linkage with the front-end graph structure processing results. Risk identification results are usually manifested as isolated events or node scores, and it is impossible to build a response path or collaborative trigger chain at the graph structure level, resulting in difficulty in achieving structural-level adaptive control in task scheduling. Especially in multi-task, heterogeneous graph, and high-frequency triggering scenarios, traditional scheduling strategies are difficult to connect with complex risk map outputs, and scheduling accuracy and timeliness are difficult to meet the real-time requirements of supervision.

[0008] To sum up, the existing big data comprehensive supervision technology has significant defects in key links such as heterogeneous graph structure modeling, context-driven graph aggregation mechanism, structural perturbation enhancement, comparative training optimization, abnormal node identification and risk response linkage. It is impossible to form a structural closed loop from perception modeling to response scheduling. When faced with complex supervision scenarios such as heterogeneous high-dimensional, multi-task driven, and dynamic risk evolution, it lacks a unified graph expression optimization and task execution mechanism, which limits its promotion and implementation in high-precision and robust supervision applications. These problems are precisely the core technical pain points that the present invention intends to solve.

[0009] Therefore, how to provide a big data comprehensive supervision platform based on intelligent decision-making is an urgent problem that technical personnel in this field need to solve. Summary of the Invention

[0010] One purpose of the present invention is to propose a big data comprehensive supervision platform based on intelligent decision-making. The present invention integrates heterogeneous graph modeling, graph neural network embedding, DiffPool structure based on Transformer control, graph structure perturbation enhancement, contrastive learning optimization, isolation forest anomaly detection and strategy response generation and other methods. The system constructs a full-link mechanism from multi-source supervision data modeling, structure aggregation, graph enhancement training to risk response scheduling, which has the advantages of strong structural expression ability, high anomaly recognition accuracy and high degree of intelligent response scheduling.

[0011] A big data integrated supervision platform based on intelligent decision-making according to an embodiment of the present invention includes:

[0012] Heterogeneous graph construction module, used to construct a heterogeneous supervision graph structure containing structural attributes and semantic labels, where nodes and edges represent object features and interaction behaviors respectively;

[0013] The node feature generation module is used to fuse structured indicators, behavior records and label information to generate node feature vectors and encode them into node embedding representations through curvature graph network encoding;

[0014] The structure aggregation module is used to load the improved DiffPool structure based on the Transformer control network, generate a node allocation matrix based on the supervision task type and node embedding representation, and perform hierarchical aggregation to obtain a pooling graph structure;

[0015] The graph enhancement module is used to perform node permutation, edge perturbation, and subgraph sampling on heterogeneous supervision graph structures and pooled graph structures to construct graph enhancement sample pairs;

[0016] The contrastive training module is used to perform contrastive training on graph augmentation sample pairs, minimize the embedding differences of homologous graphs, and optimize the pooled graph embedding representation;

[0017] Anomaly detection module, which is used to calculate node anomaly scores based on the isolation forest algorithm and identify abnormal node sets in the embedding space;

[0018] The risk graph generation module is used to perform clustering and edge reconstruction on the abnormal node set, and output a risk graph structure containing node identifiers, edge connections, and embedded representations;

[0019] The response scheduling module is used to match the policy template according to the risk graph structure and the regulatory task type, and generate and execute regulatory response instructions.

[0020] Optionally, modules can be connected using the following methods:

[0021] S1. Construct a heterogeneous supervision graph structure, taking a graph structure composed of different types of nodes and edges as input. Nodes contain structural attributes and semantic labels, and edges represent associations and interactions.

[0022] S2, generate node feature vectors, integrate structured indicators, behavior records and label information, and obtain node embedding representation through curvature graph network coding;

[0023] S3: Load the improved DiffPool structure based on the Transformer control network, dynamically generate the node allocation matrix according to the supervision task type and node embedding representation, perform hierarchical aggregation, and generate a pooling graph structure;

[0024] S4. Perform structural perturbations on the heterogeneous supervision graph structure and the pooled graph structure, generating graph enhancement sample pairs through node replacement, edge perturbation, and subgraph sampling;

[0025] S5. Perform comparative training on graph augmentation sample pairs to minimize the embedding representation differences between homologous sample pairs and optimize the embedding representation of the pooled graph structure.

[0026] S6. Perform anomaly detection on the embedded representation of the pooled graph structure and use the isolation forest algorithm to identify the set of abnormal nodes in the embedded space;

[0027] S7. Perform clustering and connection relationship mapping on the abnormal node set, and output a risk graph structure containing node identifiers, connection edges, and embedding representations;

[0028] S8. Match the strategy template according to the risk graph structure and the regulatory task type, generate regulatory response instructions and schedule their execution.

[0029] Optionally, the improved DiffPool structure includes:

[0030] Introducing a Transformer control network to dynamically generate a node allocation matrix based on the supervision task type and node embedding representation;

[0031] Set structure preservation constraints and sparsity regularization terms to control the continuity and sparsity of the node assignment matrix.

[0032] Optionally, the S1 specifically includes:

[0033] S11. Set a heterogeneous graph node type set, including supervision object nodes, behavior event nodes, and management rule nodes; set an edge type set, including entity interaction edges, rule constraint edges, and historical behavior edges;

[0034] S12. Collecting structural attributes and semantic label information of each node, where structural attributes include node type, node degree value, and node weight, and semantic labels include regulatory domain identifier, risk level label, and timestamp;

[0035] S13. Construct a heterogeneous supervision graph structure, setting the graph structure to G = (V, E, Tn, Te), where V is the node set, E is the edge set, Tn is the node type mapping function, and Te is the edge type mapping function;

[0036] S14. Record the interaction behavior attributes for each edge in the graph structure, including interaction frequency, interaction intensity, and time series index, to form an edge attribute vector set;

[0037] S15. Encode node attributes and edge attributes as structured inputs and uniformly input them into the graph construction module to generate a heterogeneous supervision graph structure containing structural information and semantic information.

[0038] Optionally, the S2 specifically includes:

[0039] S21, extracting the structural attributes, semantic labels, and behavioral records of associated edges of each node in the heterogeneous supervision graph structure to form an original feature set;

[0040] S22. Normalize the numerical indicators in the structural attributes, use one-hot encoding to represent the semantic labels, and construct a standardized feature representation set;

[0041] S23. Arrange the standardized feature representation set into a node feature vector matrix, where each row in the matrix corresponds to a feature vector of a node;

[0042] S24. Input the node feature vector matrix into the curvature graph network model, encode the feature vector through the connection relationship between the nodes and the corresponding curvature information, and generate a node embedding representation matrix.

[0043] Optionally, the S3 specifically includes:

[0044] S31, receiving the node embedding representation matrix and the supervision task type identification vector, constructing the fusion input matrix by splicing the task vector in the feature dimension and copying and extending it to all nodes as the input of the Transformer control network;

[0045] S32, inputting the fused input matrix into the Transformer control network including the attention mapping unit, the residual connection module and the feedforward transformation module, generating a node representation matrix containing contextual semantics, which represents the encoding result of each node under the current supervision task;

[0046] S33. Generate a node assignment matrix based on the node representation matrix, where the probability of each original node belonging to all aggregated nodes is calculated by attention scoring and normalization. The node assignment matrix is used to describe the hierarchical merging relationship in the graph structure.

[0047] S34. Define a regularization loss function to optimize the node allocation matrix. The loss function includes a structure preservation term and a sparsity control term:

[0048]

[0049] in, represents the regularization loss function, λ1 represents the weight coefficient of the structure preservation term, λ2 represents the weight coefficient of the sparsity control term, and p ij represents the belonging probability between the original node and the aggregated node, A ik represents the edge connection weight between two nodes in the original graph, and ε represents the index set of all edges in the original graph structure;

[0050] S35. By combining the main task objective function with the above regularization loss function, an overall optimization objective function is constructed, and the parameters of the Transformer control network are updated using the backpropagation mechanism. The training is iterated until the node allocation matrix converges.

[0051] S36. Perform a matrix multiplication operation on the node allocation matrix after training convergence and the node embedding representation matrix to generate an embedding representation of the aggregated nodes, and construct a pooled graph structure based on the edge connection information and the attribution relationship in the original graph structure to obtain the aggregated node set and edge set.

[0052] Optionally, the S4 specifically includes:

[0053] S41. Construct a perturbation sample generation input set, where the input set includes a node set, an edge set, and a node embedding representation in a heterogeneous supervision graph structure and a pooling graph structure;

[0054] S42, performing a permutation operation on the node set in the heterogeneous supervision graph structure, adjusting the arrangement order of the node embedding representation matrix by randomly rearranging the node indexes, keeping the node type and label content unchanged, and generating a node permutation graph structure;

[0055] S43, performing a perturbation operation on the edge set, setting a perturbation ratio, selecting edge elements that meet the ratio from the edge set, and performing a replacement operation on the start node or the target node to generate an edge perturbation graph structure, with the edge type and edge weight remaining unchanged;

[0056] S44. Perform subgraph sampling operations on the heterogeneous supervision graph structure and the pooled graph structure, set the node sampling number and edge density threshold, control the node category ratio and edge type distribution in the sampling process to be consistent, and generate a subgraph sample set;

[0057] S45. Perform structural embedding difference calculation on each pair of sub-graph samples and define the structural difference function:

[0058]

[0059] Among them, Δ(G a ,G b ) represents the subgraph structure G a With G b The average embedding difference of |V| represents the total number of nodes involved in the comparison. Represents subgraph G a The embedding vector of the i-th node in , Represents subgraph G b The embedding vector of the i-th node in ,||||2 represents the calculation method of the two norm;

[0060] S46. Set upper and lower threshold intervals based on the calculation results of the structural difference function, select graph structure sample pairs that meet the difference requirements, and form a graph enhancement sample pair set.

[0061] Optionally, the S5 specifically includes:

[0062] S51, receiving a set of graph enhancement sample pairs, where each pair of graph structures in the sample pairs includes a perturbed and unperturbed graph structure, a node set, an edge set, and a node embedding representation;

[0063] S52, comparing the node embedding representation of each pair of subgraph samples into the training graph, extracting the graph-level representation vector, and constructing the embedding distance metric between homologous graph samples;

[0064] S53. Set the graph contrast loss function to minimize the difference in embedding vectors between homologous graph samples and define the graph contrast loss function:

[0065]

[0066] in, Represents the graph contrast loss value, g a Represents the graph-level embedding vector of a graph augmentation sample, g b Indicates g a The graph-level embedding vector of another enhanced sample of the same origin, Indicates that it contains g a , sim represents the cosine similarity function, and τ represents the temperature scaling factor;

[0067] S54. Combine the graph contrast loss function and the main task supervision signal to form the target optimization function, and adjust the node embedding representation in the pooled graph structure through joint backpropagation.

[0068] Optionally, the S6 specifically includes:

[0069] S61. Extract the embedded representations of all nodes in the pooled graph structure and construct an anomaly detection input vector set, where each row in the vector set corresponds to the embedded vector of a node;

[0070] S62. Normalize the node embedding representation to unify the feature scale range and ensure that the weights of each dimension feature are consistent during the anomaly detection process;

[0071] S63. Construct an isolation forest model, set the number of subtrees, the number of subsamples, and the maximum tree depth parameters, use a random feature partitioning mechanism to establish a tree structure, perform a path splitting operation on the node embedding representation, and record the average path length;

[0072] S64. Calculate the anomaly score of each node and define the anomaly score function:

[0073]

[0074] Where s(x) represents the anomaly score of the input node embedding vector x, E(h(x)) represents the average path length of node x in the isolation forest, and c(n) represents the normalization constant when the number of samples is n, which is used to adjust the score scale.

[0075] S65. Set an abnormality threshold, mark nodes with abnormality scores higher than the threshold, and generate an abnormal node set. Each element in the set includes a node identifier and an embedded representation.

[0076] Optionally, the S7 specifically includes:

[0077] S71. Receive node identifiers, connection edges, and embedding representation information in the abnormal node set, and construct a data set for clustering, where each item in the data set includes a node number and a corresponding embedding vector;

[0078] S72. Set the number of cluster categories and the distance metric function, perform a clustering operation on the embedded representations in the abnormal node set, and generate a cluster label corresponding to each node, where all labels do not overlap;

[0079] S73, extracting connection edge information in the original graph structure in each cluster, and constructing a connection index table, where the connection index table includes the starting point number, the end point number, and the edge existence identifier of the edge;

[0080] S74. Construct a cluster node association strength matrix based on the connection index table and define an association strength calculation function:

[0081]

[0082] Among them, R ij Indicates the strength of association between node i and node j in the cluster, e ij Indicates the existence of the edge from node i to node j in the connection edge, n represents the number of nodes in the current cluster, and the denominator represents the cumulative weight of the connection edge between node i and all nodes in the cluster;

[0083] S75. Construct a risk graph structure based on the cluster labels, node embedding representations, and connection edge information. The risk graph structure consists of node identifiers, connection edges, and embedding representations in the abnormal node set.

[0084] The beneficial effects of the present invention are:

[0085] This paper overcomes the limitations of existing technologies in the expressiveness of homogeneous modeling by constructing a heterogeneous regulatory graph structure for multiple types of regulatory objects, behavioral events, and management rules, achieving high-dimensional modeling and multi-source fusion of complex regulatory relationships. During the node feature generation phase, the model integrates structured indicators, behavioral records, and semantic labels, and introduces a curvature graph network to encode inter-node associations. This allows the embedded representation to more fully reflect the semantic position and local structural state of the node within the overall graph, significantly enhancing the model's ability to capture the semantics of node behavior.

[0086] During the graph aggregation process, this paper introduces an improved DiffPool structure based on a Transformer control network. This dynamically generates a node allocation matrix based on the supervised task type and node embedding representation, constructing a hierarchical aggregation mechanism with task adaptability. This effectively alleviates the problems of single aggregation strategies and unbalanced classification in traditional pooling methods. Furthermore, structure-preserving constraints and sparsity regularization terms are introduced during the node allocation phase to control the stability and expression sparsity of the graph structure during the aggregation process, improving the pooled graph structure's ability to retain global structural information.

[0087] During the model optimization process, the present invention designed a structural perturbation mechanism based on node permutation, edge perturbation, and subgraph sampling. It also combined a structural difference function to construct graph enhancement sample pairs. This method then optimized and trained the embedding representations of homologous samples using a comparative learning strategy, enhancing the graph model's robustness to minor structural changes and its ability to discriminate abnormal patterns. For anomaly detection, the method employed an isolation forest algorithm combined with a pooled graph embedding representation to calculate node anomaly scores and construct a collection of abnormal nodes in the embedding space, effectively improving the ability to perceive anomalies under unsupervised conditions.

[0088] The present invention generates a risk graph structure containing node identification, connection edges and embedded representation by clustering analysis and connection edge mapping of abnormal node sets, forming a risk expression result with topological interpretation capabilities. In conjunction with the task type-driven policy template matching mechanism, it automatically generates and schedules regulatory response instructions, realizing a closed loop of the entire process from perception modeling, structural abstraction to response execution, and significantly improving the intelligence level, recognition accuracy and response efficiency of the comprehensive regulatory system in the big data environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0090] Figure 1 This is a flow chart of a big data comprehensive supervision platform based on intelligent decision-making proposed by the present invention;

[0091] Figure 2 A flowchart for constructing a heterogeneous supervision map for a big data integrated supervision platform based on intelligent decision-making proposed by the present invention;

[0092] Figure 3 This is a flowchart of anomaly detection and response scheduling for a big data comprehensive supervision platform based on intelligent decision-making proposed by the present invention. DETAILED DESCRIPTION

[0093] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0094] refer to Figure 1-3 , a big data integrated supervision platform based on intelligent decision-making, including:

[0095] Heterogeneous graph construction module, used to construct a heterogeneous supervision graph structure containing structural attributes and semantic labels, where nodes and edges represent object features and interaction behaviors respectively;

[0096] The node feature generation module is used to fuse structured indicators, behavior records and label information to generate node feature vectors and encode them into node embedding representations through curvature graph network encoding;

[0097] The structure aggregation module is used to load the improved DiffPool structure based on the Transformer control network, generate a node allocation matrix based on the supervision task type and node embedding representation, and perform hierarchical aggregation to obtain a pooling graph structure;

[0098] The graph enhancement module is used to perform node permutation, edge perturbation, and subgraph sampling on heterogeneous supervision graph structures and pooled graph structures to construct graph enhancement sample pairs;

[0099] The contrastive training module is used to perform contrastive training on graph augmentation sample pairs, minimize the embedding differences of homologous graphs, and optimize the pooled graph embedding representation;

[0100] Anomaly detection module, which is used to calculate node anomaly scores based on the isolation forest algorithm and identify abnormal node sets in the embedding space;

[0101] The risk graph generation module is used to perform clustering and edge reconstruction on the abnormal node set, and output a risk graph structure containing node identifiers, edge connections, and embedded representations;

[0102] The response scheduling module is used to match the policy template according to the risk graph structure and the regulatory task type, and generate and execute regulatory response instructions.

[0103] The present invention realizes the full process integration from structured modeling to intelligent response execution by constructing a modular platform including heterogeneous graph construction, node feature generation, structure aggregation, graph enhancement, comparative training, anomaly detection, risk graph generation and response scheduling, which significantly improves the structural adaptability and intelligent response capability of the big data supervision system.

[0104] In this embodiment, the modules are connected through the following methods:

[0105] S1. Construct a heterogeneous supervision graph structure, taking a graph structure composed of different types of nodes and edges as input. Nodes contain structural attributes and semantic labels, and edges represent associations and interactions.

[0106] S2, generate node feature vectors, integrate structured indicators, behavior records and label information, and obtain node embedding representation through curvature graph network coding;

[0107] S3: Load the improved DiffPool structure based on the Transformer control network, dynamically generate the node allocation matrix according to the supervision task type and node embedding representation, perform hierarchical aggregation, and generate a pooling graph structure;

[0108] S4. Perform structural perturbations on the heterogeneous supervision graph structure and the pooled graph structure, generating graph enhancement sample pairs through node replacement, edge perturbation, and subgraph sampling;

[0109] S5. Perform comparative training on graph augmentation sample pairs to minimize the embedding representation differences between homologous sample pairs and optimize the embedding representation of the pooled graph structure.

[0110] S6. Perform anomaly detection on the embedded representation of the pooled graph structure and use the isolation forest algorithm to identify the set of abnormal nodes in the embedded space;

[0111] S7. Perform clustering and connection relationship mapping on the abnormal node set, and output a risk graph structure containing node identifiers, connection edges, and embedding representations;

[0112] S8. Match the strategy template according to the risk graph structure and the regulatory task type, generate regulatory response instructions and schedule their execution.

[0113] This invention adopts a process-based approach to carefully define the data flow and computing logic between modules, covering the entire cycle from heterogeneous graph modeling, embedding generation to risk reasoning and response execution, ensuring that the platform has a clear logical closed loop and task linkage capabilities, and improving the system stability and scalability of regulatory tasks.

[0114] In this embodiment, the improved DiffPool structure includes:

[0115] Introducing a Transformer control network to dynamically generate a node allocation matrix based on the supervision task type and node embedding representation;

[0116] Set structure preservation constraints and sparsity regularization terms to control the continuity and sparsity of the node assignment matrix.

[0117] The present invention introduces an improved DiffPool structure based on the Transformer control network, integrates task semantics and structure perception capabilities into the node allocation process, and cooperates with structure preservation and sparsity regularization terms to effectively improve the expression continuity of the aggregate graph structure and the controllability of node allocation.

[0118] In this embodiment, S1 specifically includes:

[0119] S11. Set a heterogeneous graph node type set, including supervision object nodes, behavior event nodes, and management rule nodes; set an edge type set, including entity interaction edges, rule constraint edges, and historical behavior edges;

[0120] S12. Collecting structural attributes and semantic label information of each node, where structural attributes include node type, node degree value, and node weight, and semantic labels include regulatory domain identifier, risk level label, and timestamp;

[0121] S13. Construct a heterogeneous supervision graph structure, setting the graph structure to G = (V, E, Tn, Te), where V is the node set, E is the edge set, Tn is the node type mapping function, and Te is the edge type mapping function;

[0122] S14. Record the interaction behavior attributes for each edge in the graph structure, including interaction frequency, interaction intensity, and time series index, to form an edge attribute vector set;

[0123] S15. Encode node attributes and edge attributes as structured inputs and uniformly input them into the graph construction module to generate a heterogeneous supervision graph structure containing structural information and semantic information.

[0124] The present invention establishes a complete heterogeneous supervision graph structure by setting multiple types of node and edge sets, integrating structural attributes, semantic labels and interactive behavior attributes, enhancing the platform's ability to model complex dependencies between multiple supervision objects, and meeting the needs of high-dimensional heterogeneous data processing.

[0125] In this embodiment, S2 specifically includes:

[0126] S21, extracting the structural attributes, semantic labels, and behavioral records of associated edges of each node in the heterogeneous supervision graph structure to form an original feature set;

[0127] S22. Normalize the numerical indicators in the structural attributes, use one-hot encoding to represent the semantic labels, and construct a standardized feature representation set;

[0128] S23. Arrange the standardized feature representation set into a node feature vector matrix, where each row in the matrix corresponds to a feature vector of a node;

[0129] S24. Input the node feature vector matrix into the curvature graph network model, encode the feature vector through the connection relationship between the nodes and the corresponding curvature information, and generate a node embedding representation matrix.

[0130] The present invention extracts the behavioral records of node structure attributes, semantic labels and their associated edges, and uses a curvature graph network for embedded expression, thereby improving the node representation's ability to jointly perceive the local geometric structure and global graph semantics, and providing more stable feature support for subsequent graph processing tasks.

[0131] In this embodiment, S3 specifically includes:

[0132] S31, receiving the node embedding representation matrix and the supervision task type identification vector, constructing the fusion input matrix by splicing the task vector in the feature dimension and copying and extending it to all nodes as the input of the Transformer control network;

[0133] S32, inputting the fused input matrix into the Transformer control network including the attention mapping unit, the residual connection module and the feedforward transformation module, generating a node representation matrix containing contextual semantics, which represents the encoding result of each node under the current supervision task;

[0134] S33. Generate a node assignment matrix based on the node representation matrix, where the probability of each original node belonging to all aggregated nodes is calculated by attention scoring and normalization. The node assignment matrix is used to describe the hierarchical merging relationship in the graph structure.

[0135] S34. Define a regularization loss function to optimize the node allocation matrix. The loss function includes a structure preservation term and a sparsity control term:

[0136]

[0137] in, represents the regularization loss function, λ1 represents the weight coefficient of the structure preservation term, λ2 represents the weight coefficient of the sparsity control term, and p ij represents the belonging probability between the original node and the aggregated node, A ik represents the edge connection weight between two nodes in the original graph, and ε represents the index set of all edges in the original graph structure;

[0138] S35. By combining the main task objective function with the above regularization loss function, an overall optimization objective function is constructed, and the parameters of the Transformer control network are updated using the backpropagation mechanism. The training is iterated until the node allocation matrix converges.

[0139] S36. Perform a matrix multiplication operation on the node allocation matrix after training convergence and the node embedding representation matrix to generate an embedding representation of the aggregated nodes, and construct a pooled graph structure based on the edge connection information and the attribution relationship in the original graph structure to obtain the aggregated node set and edge set.

[0140] The present invention introduces a task control vector and a node embedding splicing mechanism into the DiffPool structure, and constructs an attribution probability matrix. By constraining the training process of the allocation matrix with structure-preserving terms and sparse regularization terms, the global stability and task adaptability of the node hierarchical merging results are dually guaranteed.

[0141] In this embodiment, the S4 specifically includes:

[0142] S41. Construct a perturbation sample generation input set, where the input set includes a node set, an edge set, and a node embedding representation in a heterogeneous supervision graph structure and a pooling graph structure;

[0143] S42, performing a permutation operation on the node set in the heterogeneous supervision graph structure, adjusting the arrangement order of the node embedding representation matrix by randomly rearranging the node indexes, keeping the node type and label content unchanged, and generating a node permutation graph structure;

[0144] S43, performing a perturbation operation on the edge set, setting a perturbation ratio, selecting edge elements that meet the ratio from the edge set, and performing a replacement operation on the start node or the target node to generate an edge perturbation graph structure, with the edge type and edge weight remaining unchanged;

[0145] S44. Perform subgraph sampling operations on the heterogeneous supervision graph structure and the pooled graph structure, set the node sampling number and edge density threshold, control the node category ratio and edge type distribution in the sampling process to be consistent, and generate a subgraph sample set;

[0146] S45. Perform structural embedding difference calculation on each pair of sub-graph samples and define the structural difference function:

[0147]

[0148] Among them, Δ(G a ,G b ) represents the subgraph structure G a With G b The average embedding difference of |V| represents the total number of nodes involved in the comparison. Represents subgraph G a The embedding vector of the i-th node in , Represents subgraph G b The embedding vector of the i-th node in ,||||2 represents the calculation method of the two norm;

[0149] S46. Set upper and lower threshold intervals based on the calculation results of the structural difference function, select graph structure sample pairs that meet the difference requirements, and form a graph enhancement sample pair set.

[0150] The present invention generates graph enhancement samples by combining node permutation, edge perturbation and subgraph sampling, and uses the structural difference function to screen sample pairs with moderate perturbation amplitude, thereby enhancing the effectiveness of sample pairs in contrastive learning and improving the expression robustness of graph embedding in perturbation scenarios.

[0151] In this embodiment, the S5 specifically includes:

[0152] S51, receiving a set of graph enhancement sample pairs, where each pair of graph structures in the sample pairs includes a perturbed and unperturbed graph structure, a node set, an edge set, and a node embedding representation;

[0153] S52, comparing the node embedding representation of each pair of subgraph samples into the training graph, extracting the graph-level representation vector, and constructing the embedding distance metric between homologous graph samples;

[0154] S53. Set the graph contrast loss function to minimize the difference in embedding vectors between homologous graph samples and define the graph contrast loss function:

[0155]

[0156] in, Represents the graph contrast loss value, g a Represents the graph-level embedding vector of a graph augmentation sample, g b Indicates g a The graph-level embedding vector of another enhanced sample of the same origin, Indicates that it contains g a , sim represents the cosine similarity function, and τ represents the temperature scaling factor;

[0157] S54. Combine the graph contrast loss function and the main task supervision signal to form the target optimization function, and adjust the node embedding representation in the pooled graph structure through joint backpropagation.

[0158] The present invention conducts graph-level comparative training based on graph-augmented sample pairs, introduces the goal of minimizing the similarity of homologous samples, and jointly optimizes graph embedding with the main task supervision signal, thereby improving the model's ability to discriminate structural perturbations and abnormal patterns and strengthening the consistent expression of the embedding space.

[0159] In this embodiment, S6 specifically includes:

[0160] S61. Extract the embedded representations of all nodes in the pooled graph structure and construct an anomaly detection input vector set, where each row in the vector set corresponds to the embedded vector of a node;

[0161] S62. Normalize the node embedding representation to unify the feature scale range and ensure that the weights of each dimension feature are consistent during the anomaly detection process;

[0162] S63. Construct an isolation forest model, set the number of subtrees, the number of subsamples, and the maximum tree depth parameters, use a random feature partitioning mechanism to establish a tree structure, perform a path splitting operation on the node embedding representation, and record the average path length;

[0163] S64. Calculate the anomaly score of each node and define the anomaly score function:

[0164]

[0165] Where s(x) represents the anomaly score of the input node embedding vector x, E(h(x)) represents the average path length of node x in the isolation forest, and c(n) represents the normalization constant when the number of samples is n, which is used to adjust the score scale.

[0166] S65. Set an abnormality threshold, mark nodes with abnormality scores higher than the threshold, and generate an abnormal node set. Each element in the set includes a node identifier and an embedded representation.

[0167] This paper adopts the isolation forest algorithm to perform path splitting and anomaly score calculation on the node embedding representation, combines normalization preprocessing with the score threshold screening mechanism, realizes the accurate identification of potential high-risk nodes under unsupervised conditions, and enhances the platform's anomaly analysis capabilities.

[0168] In this embodiment, the S7 specifically includes:

[0169] S71. Receive node identifiers, connection edges, and embedding representation information in the abnormal node set, and construct a data set for clustering, where each item in the data set includes a node number and a corresponding embedding vector;

[0170] S72. Set the number of cluster categories and the distance metric function, perform a clustering operation on the embedded representations in the abnormal node set, and generate a cluster label corresponding to each node, where all labels do not overlap;

[0171] S73, extracting connection edge information in the original graph structure in each cluster, and constructing a connection index table, where the connection index table includes the starting point number, the end point number, and the edge existence identifier of the edge;

[0172] S74. Construct a cluster node association strength matrix based on the connection index table and define an association strength calculation function:

[0173]

[0174] Among them, Rij Indicates the strength of association between node i and node j in the cluster, e ij Indicates the existence of the edge from node i to node j in the connection edge, n represents the number of nodes in the current cluster, and the denominator represents the cumulative weight of the connection edge between node i and all nodes in the cluster;

[0175] S75. Construct a risk graph structure based on the cluster labels, node embedding representations, and connection edge information. The risk graph structure consists of node identifiers, connection edges, and embedding representations in the abnormal node set.

[0176] The present invention performs clustering and connection edge mapping on abnormal node sets to construct a risk graph structure with structural semantic integrity and node status expression capabilities, and dynamically matches strategy templates based on the risk graph to achieve structure-driven closed-loop control from risk identification to response scheduling.

[0177] Example 1:

[0178] In order to verify the practicality and stability of the big data comprehensive supervision platform based on intelligent decision-making proposed in the present invention in typical grassroots scenarios, five common daily supervision tasks were selected for system deployment and testing, including community electricity consumption anomaly investigation, town-level financial auditing, medical procurement flow monitoring, enterprise risk visit warning and dynamic supervision of construction personnel. These tasks are frequently used in grassroots units, but are often difficult to achieve effective monitoring and automatic response due to heterogeneous data sources, complex object relationships, and lagging recognition mechanisms. The platform of the present invention relies on graph neural networks, structural aggregation mechanisms and comparative training optimization paths, and through unified heterogeneous graph modeling and intelligent response scheduling, it realizes accurate modeling and automatic handling of various supervision scenarios.

[0179] In the task of troubleshooting abnormal electricity consumption in the community, the platform collects meter numbers, electricity consumption curves, resident information and repair records, and constructs nodes including buildings, meters, and residents. Edge attributes record power fluctuations and repair frequency. Through embedding generation and pooled graph structure aggregation, the system combines disturbance enhancement samples with comparative learning mechanism training models to identify multiple node paths whose electricity consumption behaviors are inconsistent with historical records, and triggers inspection response templates to complete on-site inspections.

[0180] In the task of fiscal funds supervision, the system models payment approval nodes, construction projects and enterprise qualifications into a graph structure, and collects the timestamp and approval path of each fund flow. The platform identifies that some payment paths have skipped approvals and repeated bidding by the same enterprise, marks the corresponding nodes in the generated risk graph structure, and automatically schedules freezing instructions and audit reminder tasks.

[0181] In the medical procurement supervision scenario, the system models procurement batches, application units, and supplier nodes, uses procurement frequency and inventory distribution behavior to build edge relationships, and after combining comparative training optimization, identifies several abnormal batches that do not match the qualification update records, and links the supply room to execute the inventory process to improve the controllability and compliance of material flow.

[0182] In the enterprise risk warning task, the platform models legal person nodes, enterprise nodes and change records, analyzes registered capital, change frequency and historical operating abnormal behaviors, and constructs a dynamic relationship diagram. The system identifies potential clusters of shell companies through disturbance graph enhancement and graph comparison mechanisms, and forms a traceable risk path, matching the visit mechanism for on-site verification.

[0183] In the construction personnel supervision task, the platform collects real-name entry and exit records, safety training files and dangerous work application information, builds a multi-dimensional graph structure between personnel, construction sites and time periods, and identifies personnel nodes that have not reported for many days and have been working at heights, and automatically issues rectification and risk warning processes, effectively realizing intelligent safety supervision of construction sites.

[0184] To further illustrate the core performance of the system in different tasks, as shown in Table 1:

[0185] Table 1. Performance data of typical application tasks of intelligent supervision platform

[0186]

[0187] As can be seen from Table 1, the anomaly recognition accuracy of the proposed system in all five types of daily tasks exceeds 90%, indicating that the platform has strong anomaly recognition and structural modeling capabilities. The response generation time is controlled within 3.5 seconds, meeting the application requirements of real-time decision-making. The risk node embedding difference shows the platform's sensitivity to graph perturbations, helping to accurately identify embedded anomaly areas. The strategy matching success rate is stable at above 94%, verifying the execution stability and scheduling reliability of the task type matching mechanism.

[0188] In summary, the platform of the present invention has been verified in various types of daily supervision tasks through unified graph structure modeling and intelligent decision-making mechanisms. It not only effectively solves the problems of information discontinuity and delayed response in traditional processes, but also lays the foundation for subsequent large-scale deployment and multi-task integration. It has good engineering practicality and promotion prospects.

[0189] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A big data comprehensive supervision platform based on intelligent decision-making, characterized by: include: Heterogeneous graph construction module, used to construct a heterogeneous supervision graph structure containing structural attributes and semantic labels, where nodes and edges represent object features and interaction behaviors respectively; The node feature generation module is used to fuse structured indicators, behavior records and label information to generate node feature vectors and encode them into node embedding representations through curvature graph network encoding; The structure aggregation module is used to load the improved DiffPool structure based on the Transformer control network, generate a node allocation matrix based on the supervision task type and node embedding representation, and perform hierarchical aggregation to obtain a pooling graph structure; The graph enhancement module is used to perform node permutation, edge perturbation, and subgraph sampling on heterogeneous supervision graph structures and pooled graph structures to construct graph enhancement sample pairs; The contrastive training module is used to perform contrastive training on graph augmentation sample pairs, minimize the embedding differences of homologous graphs, and optimize the pooled graph embedding representation; Anomaly detection module, which is used to calculate node anomaly scores based on the isolation forest algorithm and identify abnormal node sets in the embedding space; The risk graph generation module is used to perform clustering and edge reconstruction on the abnormal node set, and output a risk graph structure containing node identifiers, edge connections, and embedded representations; The response scheduling module is used to match the policy template according to the risk graph structure and the regulatory task type, and generate and execute regulatory response instructions.

2. A big data comprehensive supervision platform based on intelligent decision-making according to claim 1, characterized in that: The modules are implemented as follows: S1. Construct a heterogeneous supervision graph structure, taking a graph structure composed of different types of nodes and edges as input. Nodes contain structural attributes and semantic labels, and edges represent associations and interactions. S2, generate node feature vectors, integrate structured indicators, behavior records and label information, and obtain node embedding representation through curvature graph network coding; S3: Load the improved DiffPool structure based on the Transformer control network, dynamically generate the node allocation matrix according to the supervision task type and node embedding representation, perform hierarchical aggregation, and generate a pooling graph structure; S4. Perform structural perturbations on the heterogeneous supervision graph structure and the pooled graph structure, generating graph enhancement sample pairs through node replacement, edge perturbation, and subgraph sampling; S5. Perform comparative training on graph augmentation sample pairs to minimize the embedding representation differences between homologous sample pairs and optimize the embedding representation of the pooled graph structure. S6. Perform anomaly detection on the embedded representation of the pooled graph structure and use the isolation forest algorithm to identify the set of abnormal nodes in the embedded space; S7. Perform clustering and connection relationship mapping on the abnormal node set, and output a risk graph structure containing node identifiers, connection edges, and embedding representations; S8. Match the strategy template according to the risk graph structure and the regulatory task type, generate regulatory response instructions and schedule their execution.

3. A big data comprehensive supervision platform based on intelligent decision-making according to claim 2, characterized in that: The improved DiffPool structure includes: Introducing a Transformer control network to dynamically generate a node allocation matrix based on the supervision task type and node embedding representation; Set structure preservation constraints and sparsity regularization terms to control the continuity and sparsity of the node assignment matrix.

4. The big data comprehensive supervision platform based on intelligent decision-making according to claim 2 is characterized in that: Said S1 specifically includes: S11. Set a heterogeneous graph node type set, including supervision object nodes, behavior event nodes, and management rule nodes; set an edge type set, including entity interaction edges, rule constraint edges, and historical behavior edges; S12. Collecting structural attributes and semantic label information of each node, where structural attributes include node type, node degree value, and node weight, and semantic labels include regulatory domain identifier, risk level label, and timestamp; S13. Construct a heterogeneous supervision graph structure, and set the graph structure to G = (V, E, Tn, Te), where V is the node set, E is the edge set, Tn is the node type mapping function, and Te is the edge type mapping function; S14. Record the interaction behavior attributes for each edge in the graph structure, including interaction frequency, interaction intensity, and time series index, to form an edge attribute vector set; S15. Encode node attributes and edge attributes as structured inputs and uniformly input them into the graph construction module to generate a heterogeneous supervision graph structure containing structural information and semantic information.

5. The big data comprehensive supervision platform based on intelligent decision-making according to claim 2 is characterized in that: The S2 specifically includes: S21, extracting the structural attributes, semantic labels, and behavioral records of associated edges of each node in the heterogeneous supervision graph structure to form an original feature set; S22. Normalize the numerical indicators in the structural attributes, use one-hot encoding to represent the semantic labels, and construct a standardized feature representation set; S23. Arrange the standardized feature representation set into a node feature vector matrix, where each row in the matrix corresponds to a feature vector of a node; S24. Input the node feature vector matrix into the curvature graph network model, encode the feature vector through the connection relationship between the nodes and the corresponding curvature information, and generate a node embedding representation matrix.

6. The big data comprehensive supervision platform based on intelligent decision-making according to claim 2 is characterized in that: The S3 specifically includes: S31, receiving the node embedding representation matrix and the supervision task type identification vector, constructing the fusion input matrix by splicing the task vector in the feature dimension and copying and extending it to all nodes as the input of the Transformer control network; S32, inputting the fused input matrix into the Transformer control network including the attention mapping unit, the residual connection module and the feedforward transformation module, generating a node representation matrix containing contextual semantics, which represents the encoding result of each node under the current supervision task; S33. Generate a node assignment matrix based on the node representation matrix, where the probability of each original node belonging to all aggregated nodes is calculated by attention scoring and normalization. The node assignment matrix is used to describe the hierarchical merging relationship in the graph structure. S34. Define a regularization loss function to optimize the node allocation matrix. The loss function includes a structure preservation term and a sparsity control term: in, represents the regularization loss function, λ1 represents the weight coefficient of the structure preservation term, λ2 represents the weight coefficient of the sparsity control term, and p ij represents the belonging probability between the original node and the aggregated node, A ik represents the edge connection weight between two nodes in the original graph, and ε represents the index set of all edges in the original graph structure; S35. By combining the main task objective function with the above regularization loss function, an overall optimization objective function is constructed, and the parameters of the Transformer control network are updated using the backpropagation mechanism. The training is iterated until the node allocation matrix converges. S36. Perform a matrix multiplication operation on the node allocation matrix after training convergence and the node embedding representation matrix to generate an embedding representation of the aggregated nodes, and construct a pooled graph structure based on the edge connection information and the attribution relationship in the original graph structure to obtain the aggregated node set and edge set.

7. The big data comprehensive supervision platform based on intelligent decision-making according to claim 2 is characterized in that: The S4 specifically includes: S41. Construct a perturbation sample generation input set, where the input set includes a node set, an edge set, and a node embedding representation in a heterogeneous supervision graph structure and a pooling graph structure; S42, performing a permutation operation on the node set in the heterogeneous supervision graph structure, adjusting the arrangement order of the node embedding representation matrix by randomly rearranging the node indexes, keeping the node type and label content unchanged, and generating a node permutation graph structure; S43, performing a perturbation operation on the edge set, setting a perturbation ratio, selecting edge elements that meet the ratio from the edge set, and performing a replacement operation on the start node or the target node to generate an edge perturbation graph structure, with the edge type and edge weight remaining unchanged; S44. Perform subgraph sampling operations on the heterogeneous supervision graph structure and the pooled graph structure, set the node sampling number and edge density threshold, control the node category ratio and edge type distribution in the sampling process to be consistent, and generate a subgraph sample set; S45. Perform structural embedding difference calculation on each pair of sub-graph samples and define the structural difference function: Among them, Δ(G a ,G b ) represents the subgraph structure G a With G b The average embedding difference of |V| represents the total number of nodes involved in the comparison. Represents subgraph G a The embedding vector of the i-th node in , Represents subgraph G b The embedding vector of the i-th node in , || ||2 represents the calculation method of the two norm; S46. Set upper and lower threshold intervals based on the calculation results of the structural difference function, select graph structure sample pairs that meet the difference requirements, and form a graph enhancement sample pair set.

8. The big data comprehensive supervision platform based on intelligent decision-making according to claim 2 is characterized in that: The S5 specifically includes: S51, receiving a set of graph enhancement sample pairs, where each pair of graph structures in the sample pairs includes a perturbed and unperturbed graph structure, a node set, an edge set, and a node embedding representation; S52, comparing the node embedding representation of each pair of subgraph samples into the input graph to the training structure, extracting the graph-level representation vector, and constructing the embedding distance metric between homologous graph samples; S53. Set the graph contrast loss function to minimize the difference in embedding vectors between homologous graph samples and define the graph contrast loss function: in, Represents the graph contrast loss value, g a Represents the graph-level embedding vector of a graph augmentation sample, g b Indicates g a The graph-level embedding vector of another enhanced sample of the same origin, Indicates that it contains g a , sim represents the cosine similarity function, and τ represents the temperature scaling factor; S54. Combine the graph contrast loss function and the main task supervision signal to form the target optimization function, and adjust the node embedding representation in the pooled graph structure through joint backpropagation.

9. The big data comprehensive supervision platform based on intelligent decision-making according to claim 2 is characterized in that: The S6 specifically includes: S61. Extract the embedded representations of all nodes in the pooled graph structure and construct an anomaly detection input vector set, where each row in the vector set corresponds to the embedded vector of a node; S62. Normalize the node embedding representation to unify the feature scale range and ensure that the weights of each dimension feature are consistent during the anomaly detection process; S63. Construct an isolation forest model, set the number of subtrees, the number of subsamples, and the maximum tree depth parameters, use a random feature partitioning mechanism to establish a tree structure, perform a path splitting operation on the node embedding representation, and record the average path length; S64. Calculate the anomaly score of each node and define the anomaly score function: Where s(x) represents the anomaly score of the input node embedding vector x, E(h(x)) represents the average path length of node x in the isolation forest, and c(n) represents the normalization constant when the number of samples is n, which is used to adjust the score scale. S65. Set an abnormality threshold, mark nodes with abnormality scores higher than the threshold, and generate an abnormal node set. Each element in the set includes a node identifier and an embedded representation.

10. The big data comprehensive supervision platform based on intelligent decision-making according to claim 2 is characterized in that: The S7 specifically includes: S71. Receive node identifiers, connection edges, and embedding representation information in the abnormal node set, and construct a data set for clustering, where each item in the data set includes a node number and a corresponding embedding vector; S72. Set the number of cluster categories and the distance metric function, perform a clustering operation on the embedded representations in the abnormal node set, and generate a cluster label corresponding to each node, where all labels do not overlap; S73, extracting connection edge information in the original graph structure in each cluster, and constructing a connection index table, where the connection index table includes the starting point number, the end point number, and the edge existence identifier of the edge; S74. Construct a cluster node association strength matrix based on the connection index table and define an association strength calculation function: Among them, R ij Indicates the strength of association between node i and node j in the cluster, e ij Indicates the existence of the edge from node i to node j in the connection edge, n represents the number of nodes in the current cluster, and the denominator represents the cumulative weight of the connection edge between node i and all nodes in the cluster; S75. Construct a risk graph structure based on the cluster labels, node embedding representations, and connection edge information. The risk graph structure consists of node identifiers, connection edges, and embedding representations in the abnormal node set.

Citation Information

Cited By

  • SIM card management platform

    CN122069544A

  • Transaction strategy generation method and system based on graph neural network and reinforcement learning

    CN122367528A