Data flow path rich information-based data cross-border risk monitoring method and device
By constructing a heterogeneous multi-relationship spatiotemporal hypergraph and a collaborative behavior relationship graph, the problem of multi-subject collaborative anomaly detection in cross-border data flow scenarios in existing technologies is solved. This enables global monitoring of complex flow scenarios and identification of hidden behaviors, improving the accuracy and interpretability of anomaly detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2025-12-12
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies struggle to identify multi-entity collaborative anomalies in complex cross-border data flow scenarios. They lack a global perspective and path information, making it difficult to effectively monitor covert detours. Furthermore, existing methods primarily focus on local anomaly detection and lack modeling of dynamic interactions between multiple nodes and entities.
Construct a heterogeneous multi-relationship spatiotemporal hypergraph, dynamically collect cross-border flow path information, extract rich information features of nodes, edges, paths and the whole graph, and mine abnormal patterns through collaborative behavior relationship graph to realize the monitoring of abnormal collaborative behaviors of multiple subjects.
It significantly improves the global modeling capability of cross-border data flow networks, enhances the ability to perceive potential abnormal paths and behaviors, strengthens the detection accuracy and interpretability of multi-entity collaborative anomalies, and enables dynamic monitoring of abnormal events.
Smart Images

Figure CN122020444A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data monitoring technology, specifically relating to a method and device for cross-border data risk monitoring based on rich information about data flow paths. Background Technology
[0002] With the development of the global digital economy, cross-border data flows have become an important support for international trade and other fields. However, they involve core issues such as national security and personal privacy protection. Risk analysis is of great significance for enterprises to implement management regulations and for the state to control the risk situation.
[0003] In existing technologies, some studies rely on methods such as static rule matching, which can detect and control risks in localized anomalies in cross-border data transfer operations. However, they are insufficient to address risk events in complex interactive scenarios caused by frequent cross-border flows, diverse scenarios, and complex business interactions. Some studies construct risk analysis models by combining the flow and interaction relationships between entities, which can achieve risk monitoring and identification. However, they only focus on two types of data entities and their interaction information, resulting in monitoring blind spots and an inability to identify illegal cross-border behaviors involving multiple entities.
[0004] Analysis of relevant technologies reveals that while existing solutions have some effectiveness in areas such as local violation verification and traffic data detection, they generally suffer from limitations: they may struggle to identify complex global workflow scenarios, lack analysis of global abnormal behavior, lack path visibility of other entities during the workflow, or have limited anomaly identification scenarios. Key issues with existing methods include: focusing on local anomaly detection and lacking a global perspective; failing to deeply model dynamic interactions between multiple nodes and entities, resulting in limited anomaly identification capabilities; and path information not being a core feature for proactive risk identification and prediction.
[0005] Currently, the security and compliance of cross-border data flows has become a research hotspot. Existing monitoring methods have problems such as insufficient expression of path information, lack of multi-subject collaborative behavior modeling, and lack of detection of hidden detour behavior, making them unsuitable for complex flow scenarios. Summary of the Invention
[0006] In view of this, the purpose of this invention is to propose a method and apparatus for cross-border data risk monitoring based on data flow path rich information, so as to solve or partially solve the problems mentioned in the background art.
[0007] To achieve the above objectives, in a first aspect, the present invention provides a method for cross-border data risk monitoring based on data flow path-rich information, comprising the following steps: Dynamically and continuously collect cross-border data flow path information, and construct a heterogeneous multi-relationship spatiotemporal hypergraph based on the cross-border data flow path information; Based on the heterogeneous multi-relation spatiotemporal hypergraph, information-rich features of data flow paths are extracted from four levels: nodes, edges, paths, and the global graph. Based on the rich information features, a collaborative behavior relationship graph is constructed. By mining collaborative behavior patterns and performing anomaly identification and judgment, abnormal collaborative behaviors of multiple entities are monitored.
[0008] As a preferred embodiment of a data cross-border risk monitoring method based on rich information about data flow paths, the steps of constructing a heterogeneous multi-relationship spatiotemporal hypergraph based on the data cross-border flow path information include: Data providers, processors, and receivers are abstracted as nodes in the diagram, with each node carrying attributes such as subject type, industry category, geographic region, and security level. Multiple types of edges are established based on data access and transmission behavior. Each edge carries the transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether it is encrypted. The heterogeneous multi-relationship spatiotemporal hypergraph is dynamically updated as cross-border data transactions are executed and time progresses.
[0009] As a preferred approach for cross-border data risk monitoring based on data flow path rich information, the steps for extracting data flow path rich information features at the node level include: Obtain the basic attributes of the nodes; analyze the frequency of data processing, data volume change trend, number of associated entities change trend, transmission method distribution and sensitive data ratio of the nodes within a unit time period, and construct a node behavior profile based on the time window; Graph embedding techniques are used to encode the topological location and neighborhood information of nodes in the network into low-dimensional vectors; The steps for extracting information-rich features of data flow paths from the edge level include: Extract multi-dimensional interactive features including transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether it is encrypted; Calculate the relative position weight of the current edge in the overall path, and the dependency relationship between the current edge and its surrounding edges; Extract the flow direction of edges and the temporal consistency between edges and their upstream and downstream nodes.
[0010] As a preferred approach for cross-border data risk monitoring based on data flow path rich information, the steps for extracting data flow path rich information features from the path level include: By limiting the maximum path hop count threshold, a candidate path set is generated using BFS or DFS search algorithms, excluding loops and known whitelisted paths. Construct path identifiers based on node type, edge type, and access strategy, and define typical path patterns; Aggregate multiple transfer records with the same path pattern in historical data, and statistically analyze the frequency of occurrence, average amount of data transferred, and distribution of transfer time periods. Extract path length, number of transit nodes, coverage of high-risk nodes, and path opacity index; Sequence modeling combined with time-aware embedding is used to encode the temporal order of nodes and edges on the path.
[0011] As a preferred approach for cross-border data risk monitoring based on data flow path rich information, the steps for extracting data flow path rich information features at the layer level include: Statistical analysis includes the number of nodes, the number of edges, the average degree of each node, the degree of the node with the highest degree, the number of superedges, and the relationship type. Analyze graph density, average path length, network diameter, and global clustering coefficient; Extract the distribution ratio of subject type and edge type, relationship diversity index, and node attribute entropy value; Analyze the Gini coefficient, PageRank skewed distribution, cross-border data outbound path centroid, and information flow aggregation degree.
[0012] As a preferred method for cross-border data risk monitoring based on rich information about data flow paths, the steps for constructing a collaborative behavior relationship graph based on the rich information features include: Using data subjects as nodes, collaborative behavior edges are mined based on path structure similarity, behavioral temporal coupling, and transmission destination similarity. The weight of the edge is determined by cross-path co-occurrence, behavioral time window overlap, and semantic or business similarity. A sliding time window mechanism is introduced to summarize and plot the flow events at different time scales; Supplement the data flow links inherited from the original path as structural connection edges.
[0013] As a preferred approach for cross-border data risk monitoring based on information-rich data flow paths, the mining of collaborative behavior patterns includes: A dense subgraph mining algorithm and a spectral clustering method are used to identify candidate cooperative behavior groups from the cooperative behavior relationship graph; Through synergy strength The formula performs a structural evaluation of the collaborative subgraph:
[0014] In the formula, α + β + γ = 1, This represents the average path structure similarity in the subgraphs. For node attribute values, This represents the degree of coordination between path initiation and arrival times, with α, β, and γ being the weights for multi-dimensional feature fusion.
[0015] As a preferred approach for cross-border data risk monitoring based on information-rich data flow paths, anomaly identification and judgment include: Establish a path anomaly scoring function:
[0016] In the formula, Indicates the risk level of path rule matching. The deviation from the path's historical behavior. Indicates path structure anomality. w , y , z For feature weights; Constructing an overall anomaly scoring function for the collaborative subgraph:
[0017] In the formula, μ For path rule risk weights, δ For structural collaborative behavior risk weights, P A set of paths; A sliding time window mechanism is introduced to construct behavioral evolution curves for subjects or paths and identify abnormal fluctuation patterns.
[0018] As a preferred solution for cross-border data risk monitoring based on rich information about data flow paths, each hyperedge contains a node sequence, a temporal attribute based on each flow event occurring at each node, and a behavioral attribute for each flow event. The formula for the dynamic update of the heterogeneous multi-relation spatiotemporal hypergraph over time is:
[0019] In the formula, For the updated supermap, This is the supermap before the update. For the newly added set of nodes, Add a new set of edges.
[0020] As a preferred approach for cross-border data risk monitoring based on information-rich data flow paths, the sliding time window includes 5-minute, 30-minute, and 1-hour scales. The weights of the collaborative behavior edges are given by the formula:
[0021] In the formula, , , These are parameters that are optimized using a data-driven approach. For cross-path co-occurrence, The degree of overlap of behavioral time windows, For semantic or business similarity.
[0022] Secondly, the present invention provides a data cross-border risk monitoring device based on rich information of data flow paths, employing the data cross-border risk monitoring method based on rich information of data flow paths obtained in any possible implementation manner, as described in the first aspect, including: The heterogeneous multi-relationship spatiotemporal hypergraph construction module is used to dynamically and continuously collect cross-border data flow path information and construct a heterogeneous multi-relationship spatiotemporal hypergraph based on the cross-border data flow path information. The rich information feature extraction module is used to extract rich information features of data flow paths from four levels: nodes, edges, paths, and the global graph, based on the heterogeneous multi-relation spatiotemporal hypergraph. The anomaly monitoring module is used to construct a collaborative behavior relationship graph based on the rich information features, and to monitor multi-subject collaborative abnormal behavior by mining collaborative behavior patterns and performing anomaly identification and judgment.
[0023] As a preferred embodiment of a cross-border data risk monitoring device based on rich information about data flow paths, the heterogeneous multi-relationship spatiotemporal hypergraph construction module includes: The node abstraction and attribute configuration submodule is used to abstract data providers, processors and receivers as nodes in the graph. The nodes carry attributes such as subject type, industry category, region of origin and security level. The edge construction and attribute configuration submodule is used to build multiple types of edges based on data access and transmission behavior. The edges carry transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether they are encrypted. The hypergraph dynamic update submodule is used for the dynamic update of the heterogeneous multi-relationship spatiotemporal hypergraph as cross-border data business is executed and time progresses.
[0024] As a preferred embodiment of a cross-border data risk monitoring device based on rich information of data flow paths, the rich information feature extraction module includes a node-level feature extraction submodule and an edge-level feature extraction submodule; The node-level feature extraction submodule includes: The node attribute retrieval unit is used to retrieve the basic attributes of a node. The node dynamic behavior analysis unit is used to analyze the node's data processing frequency, data volume change trend, related subject number change trend, transmission method distribution and sensitive data ratio within a unit time period, and to construct a node behavior profile based on a time window. The node topology coding unit uses graph embedding techniques to encode the topological location and neighborhood information of a node in the network into a low-dimensional vector. The edge-level feature extraction submodule includes: The edge multidimensional feature extraction unit is used to extract multidimensional interactive features such as transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether it is encrypted; The edge relationship calculation unit is used to calculate the relative position weight of the current edge in the overall path, as well as the dependency relationship between the current edge and its surrounding edges; The edge timing and flow analysis unit is used to extract the flow direction of edges and the timing consistency between edges and their upstream and downstream nodes.
[0025] As a preferred embodiment of a cross-border data risk monitoring device based on rich information about data flow paths, the rich information feature extraction module includes a path-level feature extraction submodule, which includes: The candidate path generation unit is used to generate a set of candidate paths by limiting the maximum path hop count threshold and using BFS or DFS search algorithms, excluding loops and known whitelisted paths. The path pattern definition unit is used to construct path identifiers based on node type, edge type, and access strategy, and to define typical path patterns. The path behavior aggregation unit is used to aggregate multiple flow records of the same path pattern in historical data, and to count the frequency of occurrence, average amount of data transmitted, and distribution of transmission time periods. The path structure feature extraction unit is used to extract path length, number of transit nodes, coverage of high-risk nodes, and path opacity index. The path temporal coding unit is used to encode the temporal order of nodes and edges on a path by using sequence modeling combined with time-aware embedding.
[0026] As a preferred embodiment of a cross-border data risk monitoring device based on rich information about data flow paths, the rich information feature extraction module includes a graph-level feature extraction submodule, which includes: The graph structure index statistics unit is used to count the number of nodes, the number of edges, the average degree of nodes, the degree of the node with the highest degree, the number of super edges, and the relationship type. The graph topology attribute analysis unit is used to analyze graph density, average path length, network diameter, and global clustering coefficient. The graph subject and behavior integration analysis unit is used to extract the distribution ratio of subject type and edge type, relationship diversity index, and node attribute entropy value; The graph data transmission centralized analysis unit is used to analyze Gini coefficient, PageRank skewed distribution, cross-border data outbound path centroid, and information flow aggregation degree.
[0027] As a preferred embodiment of a cross-border data risk monitoring device based on rich information about data flow paths, the anomaly monitoring module includes a collaborative behavior relationship graph construction submodule, which includes: The collaborative behavior edge mining unit is used to mine collaborative behavior edges based on path structure similarity, behavior temporal coupling, and transmission destination similarity, with data subjects as nodes. The weight of the edge is determined by cross-path co-occurrence, behavior time window overlap, and semantic or business similarity. The time window processing unit is used to introduce a sliding time window mechanism to summarize and plot the flow events at different time scales. The structural connection edge supplement unit is used to supplement the data flow links inherited from the original path as structural connection edges.
[0028] As a preferred embodiment of a cross-border data risk monitoring device based on rich information about data flow paths, the anomaly monitoring module includes a collaborative behavior pattern mining submodule, which includes: The cooperative behavior group identification unit is used to identify candidate cooperative behavior groups from the cooperative behavior relationship graph using dense subgraph mining algorithm and spectral clustering method; Synergistic subgraph structural evaluation unit, used to assess synergy strength The formula performs a structural evaluation of the collaborative subgraph:
[0029] In the formula, α + β + γ = 1, This represents the average path structure similarity in the subgraphs. For node attribute values, This represents the degree of coordination between path initiation and arrival times, with α, β, and γ being the weights for multi-dimensional feature fusion.
[0030] As a preferred embodiment of a cross-border data risk monitoring device based on rich information about data flow paths, the anomaly monitoring module includes an anomaly identification and judgment submodule, which includes: The path anomaly scoring unit is used to establish the path anomaly scoring function.
[0031] In the formula, Indicates the risk level of path rule matching. The deviation from the path's historical behavior. Indicates path structure anomality. w , y , z For feature weights; The collaborative subgraph anomaly scoring unit is used to construct the overall anomaly scoring function for the collaborative subgraph.
[0032] In the formula, μ For path rule risk weights, δ For structural collaborative behavior risk weights, P A set of paths; The abnormal fluctuation pattern identification unit is used to introduce a sliding time window mechanism to construct a behavior evolution curve for the subject or path and identify abnormal fluctuation patterns.
[0033] As a preferred embodiment of a cross-border data risk monitoring device based on rich information about data flow paths, the anomaly monitoring module includes a hypergraph update and weight calculation submodule, which includes: The hypergraph dynamic update unit is used to dynamically update the heterogeneous multi-relation spatiotemporal hypergraph over time, and the formula is:
[0034] In the formula, For the updated supermap, This is the supermap before the update. For the newly added set of nodes, Add a new set of edges; The cooperative behavior edge weight calculation unit is used to calculate the weights according to the weight formula of the cooperative behavior edges: The weights of the collaborative behavior edges are given by the formula:
[0035] In the formula, , , These are parameters that are optimized using a data-driven approach. For cross-path co-occurrence, The degree of overlap of behavioral time windows, For semantic or business similarity; the sliding time window includes 5-minute, 30-minute, and 1-hour scales.
[0036] As a preferred solution for a cross-border data risk monitoring device based on rich information of data flow paths, each hyperedge in the rich information feature extraction module contains a node sequence, a temporal attribute based on each flow event occurring at each node, and a behavioral attribute of each flow event.
[0037] Thirdly, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the data cross-border risk monitoring method based on rich information of data flow paths, as described in the first aspect or any possible implementation thereof.
[0038] Fourthly, the present invention provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform steps in the data cross-border risk monitoring method based on data flow path rich information in the first aspect or any possible implementation thereof.
[0039] The beneficial effects of the technical solution provided by this invention are as follows: First, by constructing a heterogeneous multi-relation spatiotemporal hypergraph that integrates multi-dimensional attributes and path behaviors, global modeling and real-time updates of cross-border data flow networks are achieved. At the same time, rich information about data flow paths is introduced as the core basis for anomaly identification. It not only integrates the structural attributes of nodes, edges and paths, but also introduces contextual information such as temporal sequence and behavioral semantics, which significantly improves the ability to comprehensively perceive potential abnormal paths and behaviors in cross-border data flows. In particular, it has higher coverage and accuracy in identifying complex multi-hop flows and hidden collaborative paths.
[0040] Second, by introducing a collaborative behavior relationship graph, it effectively identifies abnormal interaction patterns between multiple subjects under collaboration or disguise, solving the problem of insufficient ability of traditional methods to identify potential multi-subject collaborative behaviors. It realizes semantic hierarchical modeling and reasoning of abnormal behaviors, and greatly improves the accuracy and interpretability of multi-subject collaborative anomaly detection.
[0041] Third, by integrating the temporal evolution characteristics of path behavior, dynamic monitoring and rapid identification of abnormal events are achieved. This enables immediate response to explicit risks and, in conjunction with time windows, the determination of the evolution process of latent anomalies, significantly enhancing the system's comprehensive ability to respond to sudden risks and slowly changing hidden dangers. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in this invention or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 A flowchart of a cross-border data risk monitoring method based on rich information about data flow paths, provided in an embodiment of the present invention; Figure 2 The heterogeneous multi-relationship spatiotemporal hypergraph in the data cross-border risk monitoring method based on rich information of data flow paths provided in the embodiments of the present invention; Figure 3 This invention provides a multi-entity collaborative behavior anomaly detection process in a cross-border data risk monitoring method based on rich information about data flow paths, as provided in this embodiment of the invention. Figure 4This is an architecture diagram of a cross-border data risk monitoring device based on rich information about data flow paths, provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0045] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this invention should have the ordinary meaning understood by those skilled in the art to which this invention pertains. The terms "comprising" or "including," or similar words used in the embodiments of this invention, mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects.
[0046] With the rapid development of the global digital economy, cross-border data flows have become a crucial support for international trade, cloud computing, and the Internet of Things. However, cross-border data transmission involves core issues such as national security, personal privacy protection, and industry data sovereignty. Conducting risk analysis during the actual implementation of cross-border data transactions is of great significance for supporting companies exporting data in implementing relevant management regulations and assisting national management departments in controlling the risk situation. To effectively identify cross-border data risks, some existing technologies refer to national standards or relevant laws and regulations, relying on static rule matching or log analysis methods, and employing data classification-based tagging control or designing fixed-strategy abnormal traffic filtering to assess and analyze risky behaviors related to cross-border data flows. This approach can effectively detect and control risks associated with localized anomalies in cross-border data transactions.
[0047] However, with the increasing frequency of cross-border flows, diverse data export scenarios, and complex business interactions, static rule analysis focused on single-point businesses is insufficient to address risk events in complex interactive scenarios. To more comprehensively capture abnormal data export behavior, data flow paths, as crucial supporting information, need to be integrated into anomaly detection schemes to support the discovery of complex risk events. A data flow path refers to the complete link data traverses during cross-border transmission, including the data sender, transmission nodes, transit entities, receivers, and dynamic trajectory information such as the interaction relationships, protocol types, and geographical boundaries of each link. Based on border infrastructure or dedicated monitoring equipment, some studies analyze and model the subject information and interaction relationships of data export senders and receivers, constructing a risk analysis model based on a binary structure for point-to-point flow paths. This model performs similarity-driven correlation analysis on cross-border data flow risk paths, achieving quantitative calculation and risk classification of risk values. While such research can effectively combine inter-subject flow interaction relationships to achieve risk monitoring and identification, this method only focuses on two types of data subjects and their interaction information, lacking actual intra-domain flow path information, i.e., there are monitoring blind spots, and it cannot effectively identify multi-subject related illegal export behaviors. Therefore, how to further improve the flow relationship under the support of the binary network structure in order to achieve refined risk identification and control remains an important problem to be solved.
[0048] To understand the current state of technological development, existing patents were searched, compared, and analyzed, and the following technical information with high relevance to this invention was selected: Existing technical solution 1: CN119603081A, "A Method and System for Verifying Data Cross-border Consistency," discloses a method for verifying data cross-border consistency. This method analyzes cross-border data sent by the data cross-border business system to generate a cross-border data strategy file containing elements such as the cross-border data business scenario, target output location, extraction strategies for each target field, important data categories, sensitive personal information categories, and mapping relationships for non-sensitive personal information categories. When actual cross-border transactions occur, the data processing side consistency verification subsystem uses privacy computing technology to extract the content of each target field corresponding to the cross-border data and performs consistency verification and comparison with the data in the business system without exposing the original cross-border data. This method can effectively perform partial violation verification for the content and behavior of actual cross-border data transactions. However, cross-border transactions are complex, and partial post-event content verification of "already exported data" can only solve some violation problems. When facing complex global flow business scenarios such as multi-entity collaborative cross-border transfers and multi-path split cross-border transfers, this method is difficult to effectively identify violations.
[0049] Existing technical solution 2: CN119048025A, titled "A Method, Platform, Equipment, and Medium for Cross-border Data Circulation Supervision," discloses a method for supervising cross-border data circulation. This method involves real-time collection of cross-border data information, content extraction, attribute feature identification, and further comparison and verification of encrypted cross-border data packets uploaded by users to a cross-border data monitoring platform using corresponding encryption algorithms, keys, and verification algorithms. The verification results are then processed based on the verification results and current pre-set restrictions on cross-border data (e.g., negative lists). This method can detect illegal cross-border content in real-time at the traffic data level, alleviating the difficulty in identifying and verifying encrypted data in real-time. However, this method only performs partial comparison and verification at the content level of the circulating data traffic information, and can only identify some cross-border abnormal events, lacking analysis and consideration of global abnormal behavior. Global-level cross-entity collaborative circulation and behavioral anomalies requiring circulation path information support are difficult to effectively judge using this method.
[0050] Existing technical solution 3: CN118611894A, "A Method and Apparatus for Cross-Border Data Security Management," discloses a method for security detection throughout the entire cross-border data transmission process. This method pre-transmission determines whether to proceed by checking if the requested data exists in a blockchain-based, documented database. During transmission, it analyzes the data in real-time and performs consistency verification with the data in the documented database. It also calculates similarity using regional risk values, network security risk values, and data leakage risk values. If sensitive data content is inconsistent or has low similarity, transmission is terminated. After transmission, it calculates the consistency between the received data from the overseas recipient and the cross-border registration data, thus achieving security detection after data export. This method effectively performs security detection throughout the entire data transmission process, relying on local consistency verification and comparison of the documented database content to determine whether there is a violation. However, this method only judges some abnormal events supported by the data transfer path between the data sender and the overseas recipient, lacking path visibility involving other entities during the transfer process, making it difficult to support global-level risk analysis and violation judgment.
[0051] Existing technical solution 4: CN119646875A, "A Method for Sampling Inspection and Supervision of Outbound Data Based on Blockchain and Proxy Re-encryption," discloses a method for sampling and supervising outbound data. This method utilizes double encryption and the immutability of blockchain to set up a three-stage encryption process for encrypted cross-border data transmission. The first encrypted layer is sampled to form the second encrypted layer. The second encrypted layer is then stored on the blockchain and re-encrypted to form the third encrypted layer. Finally, the data is decrypted on the regulatory side, and data compliance is verified by comparing it with enterprise declaration records, enabling dynamic cross-border data sampling inspection. This method can effectively support secure and efficient post-transfer sampling inspection of outbound data. However, it can only make violation judgments for some data anomalies, possessing a certain ability to detect local anomalies, but lacks the ability to capture risks in complex flow business scenarios. It is insufficient in responding to high-risk, cross-entity transmission scenarios from a global perspective.
[0052] Existing technical solution 5: CN119341846A, "A Method and System for Cross-Border Detection of Sensitive Data Anomalies Based on Traffic Analysis," discloses a method for cross-border detection of sensitive data anomalies. This method analyzes real-time transmitted cross-border data streams, performs a compliance assessment based on matching rules for outbound reporting information to generate a first risk score, constructs a historical behavior baseline model based on historical cross-border transmission behavior, monitors whether the time-series behavior deviates from the historical pattern, calculates a second score, and further generates a final risk value based on a preset comprehensive rule for both assessments, and provides corresponding real-time early warning responses. This method can effectively judge and respond to real-time anomalies in outbound data stream behavior at the source of risk; however, it focuses on anomalies at a local single node and cannot cover global anomalies in outbound behavior, making it difficult to effectively identify anomalies in multi-entity, multi-hop transmission scenarios.
[0053] Existing technical solution 6: CN119830180A, "A Method and System for Anomaly Detection in Cross-border E-commerce Big Data," discloses a method for anomaly detection in transaction data of the cross-border e-commerce industry. This method uses dictionary learning and simulated annealing algorithms to extract feature dictionaries of transaction behavior from historical and current cross-border e-commerce transaction data, calculates reconstruction errors for each, and compares the current error with historical errors to determine anomalies. Furthermore, it calculates the contribution of each transaction feature to the reconstruction error using Shapley values, outputting the feature with the largest contribution as the source of the anomaly. This method can effectively improve the accuracy and robustness of anomaly detection and achieve precise localization of abnormal transaction behavior features. However, this method has limited anomaly identification scenarios, only monitoring anomalies in localized, system-level structured data. If some features are incomplete during actual transmission or if there are overall anomalies in cross-entity transfer transmission, this method is unlikely to be effective.
[0054] Existing technical solution 7: Publication No. CN119743285A, "A Method, Apparatus, Service Mesh, and Computing Device for Abnormal Traffic Management," discloses an abnormal traffic management method. This method monitors the path of request traffic (i.e., the data transmission order between service instances) through the service mesh's mesh proxy and compares it with abnormal paths in an attack source database. If a match is found, the traffic is marked as abnormal, and the abnormal injection point (the starting node of the abnormal path) is determined. A virtual service instance (such as a honeypot container) is created based on the abnormal injection point, and the abnormal traffic is imported into it for isolation, preventing its spread and impact on normal business operations. This method can effectively identify and manage illegal traffic through path identification and matching. However, this method relies on the list in the attack source database for matching and judgment, and can only identify some abnormal illegal flow paths. Moreover, its main purpose is to isolate attack events. In data export scenarios, this solution is difficult to handle complex flow scenarios such as multi-path splitting of cross-border data and multi-entity collaborative illegal export.
[0055] Existing technical solution 8: CN119397073A, titled "A Method, System, Device, and Medium for Tracing the Entire Data Flow of a Visual Data Platform," discloses a method for tracing the entire data flow. This method collects path information for the entire data flow by monitoring the job scheduling process, parsing the input and output configuration information of each job node, performing static analysis of node scripts and code, and utilizing Kafka asynchronous messages and unique identifiers. Based on abnormal events and the flow path, it enables rapid location and visualization of data problems. While this method can effectively and accurately locate nodes involved in abnormal events based on contextual flow path information, its essential purpose is to trace the source of violations based on flow path information, rather than using flow path information as an abnormal event monitoring task. Therefore, it is difficult to apply this method to the identification of abnormal events in cross-border data flow risk monitoring.
[0056] Existing technical solution 9: CN112861967A, "Method and Device for Detecting Abnormal Users in Social Networks Based on Heterogeneous Graph Neural Networks," discloses a method for detecting abnormal users in social networks based on heterogeneous graph neural networks. This method collects user information from social networks and extracts user meta-features, behavioral features, and textual semantic features to construct a heterogeneous information network and designs meta-paths and meta-graphs. It aggregates neighbor node information using the intimacy and similarity between users to determine the user's representation in the social network, and then uses a logistic regression classifier to detect user types and identify abnormal users. While this patented method can effectively detect abnormal users in social networks, it focuses on user representation learning and neighbor aggregation to discover local abnormal events. It lacks the ability to model behavioral chains, temporal evolution, and directional paths in cross-border flow paths from a global perspective, making it difficult to detect complex behaviors such as path detours, split multi-hop transfers, or multi-agent collaborative avoidance strategies.
[0057] Existing technical solution 10: Li Jin et al., based on border monitoring infrastructure, constructed a bipartite network structure including domestic data transmission institutions and overseas data receiving institutions to characterize the cross-border data flow path. Based on the bipartite network model and node attributes, they performed similarity-driven correlation analysis on the risk paths of cross-border data flow, realizing the quantitative calculation of risk values and risk classification. This method can effectively analyze the risks when cross-border data flow actually occurs, supporting the risk classification monitoring of cross-border data flow. However, this method only focuses on two types of data subjects and their interaction information, lacking fine-grained information such as actual intra-domain flow paths and source behavior information, i.e., there are monitoring blind spots, and it cannot effectively identify illegal cross-border behavior involving multiple entities.
[0058] To effectively identify risks during cross-border data transmission, existing research methods for monitoring cross-border data risks have achieved innovative breakthroughs in multiple dimensions, gradually building the capability for risk identification, verification, and monitoring of cross-border data flows. Among existing technologies, risk monitoring methods mainly fall into two categories: one based on static rule matching, tag classification and log analysis, data content consistency verification, and abnormal traffic detection, uses policy configuration and abnormal traffic identification to identify and judge risks, suitable for relatively simple and clearly defined cross-border data transmission scenarios; the other attempts to construct the transmission path and subject interaction relationships of cross-border data, such as bipartite graph structures, data path matching, and full-link tracing methods, analyzing the interaction between the data sender and receiver to identify risk paths. While achieving a certain degree of risk identification and hierarchical monitoring, some solutions also integrate emerging technologies such as privacy computing, blockchain, and graph modeling to improve data protection and anomaly detection capabilities.
[0059] While the aforementioned technologies have achieved some success in identifying single-point anomalies and conducting basic risk analysis, they generally suffer from the following three key problems: First, most existing research methods (technical solutions 1-5) focus on anomaly detection at the "local" or "single-node" level, only able to detect anomalous events in individual links, lacking research on monitoring anomalous behavior from a global data flow perspective. Therefore, they are difficult to effectively and quickly identify anomalous events at the entire link level. Second, existing technical solutions such as 6, 9, and 10, although they have made some attempts in local user modeling or subject relationship characterization and have constructed bipartite network models to achieve a certain degree of anomaly analysis, have not conducted in-depth modeling of the dynamic interaction behavior between multiple nodes and multiple subjects in the cross-border data flow path, lacking characterization and trajectory tracking of intermediate nodes and collaborators. This anomaly identification capability supported by "poor information in the data flow path" is limited, the path granularity is coarse, and it can only support the detection of some anomalous events, making it difficult to identify anomalous behaviors such as multi-subject collaborative cross-border data transfer and data splitting for cross-border data transfer. Third, existing technical solutions such as 7 and 8, although involving path information, are intended for anomaly tracing or attack isolation. Their path information serves as a post-event auxiliary tool and has not become a core feature for proactive risk identification and prediction.
[0060] With the increasing sophistication of global data regulations, the security and compliance of cross-border data flows have become a key issue in the development of the digital economy. How to monitor and identify compliance risks in cross-border data flows in real time has become a research hotspot. Data exporting entities are geographically dispersed, making it difficult to comprehensively cover export risks by relying solely on border controls. Furthermore, diverse business scenarios and complex data interactions among enterprises make it difficult to trace cross-border data paths and identify responsible parties, resulting in the inability to effectively identify or quickly respond to potential global abnormal outbound flows across certain entities. Existing cross-border data monitoring methods mostly focus on single-point outflow events or static strategy matching. Some solutions introduce path interaction information as a basis for risk identification, but these methods have several shortcomings. First, the path information is insufficiently expressed, failing to depict the complex patterns of data transmission through multi-hop paths. Second, there is a lack of modeling for multi-entity collaborative behavior, making it difficult to identify abnormal patterns between related entities. Furthermore, there is a lack of effective detection for concealed bypassing behaviors, such as data evading regulation through multiple intermediate nodes.
[0061] Current technological solutions are ill-suited for handling complex cross-border data flows involving multiple stakeholders, multiple pathways, and multiple stages. Given the increasing trend of compliance risks becoming more path-based, collaborative, and avoidable, there is an urgent need to build a risk monitoring system with rich information about data flow paths as its core support.
[0062] In view of this, this application proposes a method and device for cross-border data risk monitoring based on rich information of multidimensional data flow paths. It integrates multidimensional information such as path topology, subject attributes, and behavioral trajectories, constructs a cross-border data flow network using a heterogeneous multi-relationship spatiotemporal hypergraph, and integrates rich information feature extraction of data flow paths. This effectively identifies path anomalies and collaborative anomaly patterns oriented towards multiple subjects from a global perspective, effectively ensuring data security when leaving the country. The following are the specific contents of the embodiments of this invention.
[0063] See Figure 1 This invention provides a method for monitoring cross-border data risks based on rich information about data flow paths, comprising the following steps: S1. Dynamically and continuously collect cross-border data flow path information, and construct a heterogeneous multi-relationship spatiotemporal hypergraph based on the cross-border data flow path information; S2. Based on the heterogeneous multi-relation spatiotemporal hypergraph, extract data flow path rich information features from four levels: nodes, edges, paths, and the global graph. S3. Construct a collaborative behavior relationship graph based on the rich information features, and monitor multi-subject collaborative abnormal behavior by mining collaborative behavior patterns and performing anomaly identification and judgment.
[0064] This embodiment's design mainly comprises three parts: a heterogeneous multi-relationship spatiotemporal hypergraph, feature extraction based on data flow path rich information, and multi-agent collaborative behavior anomaly identification. In the heterogeneous multi-relationship spatiotemporal hypergraph part, various subjects such as data providers, processors, and receivers are abstracted as nodes in the graph, and multiple types of edges are established based on data access and transmission behaviors; nodes and edges carry multi-dimensional attributes to describe complete semantics. Based on the heterogeneous multi-relationship spatiotemporal hypergraph modeling, in the feature extraction part based on data flow path rich information, this invention designs a multi-level feature extraction framework to characterize data flow features from micro-paths to macro-networks, providing interpretable and quantifiable feature support for subsequent anomaly behavior analysis. In the multi-agent collaborative behavior anomaly detection part, by constructing a collaborative behavior relationship graph, further using graph structure data mining and analysis methods, combined with an anomaly behavior scoring mechanism, accurate determination of anomaly behavior is achieved.
[0065] In this embodiment, step S1, which involves constructing a heterogeneous multi-relationship spatiotemporal hypergraph based on the cross-border data flow path information, includes: S11. Abstract the data provider, processor and receiver as nodes in the diagram. Each node carries attributes such as subject type, industry category, region of origin and security level. S12. Establish multiple types of edges based on data access and transmission behavior. Each edge carries the transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether it is encrypted. S13. The heterogeneous multi-relationship spatiotemporal hypergraph is dynamically updated as cross-border data business is executed and time progresses.
[0066] See Figure 2 Specifically, in the process of constructing a heterogeneous multi-relationship spatiotemporal hypergraph, rich flow path information is first dynamically and continuously collected through methods such as survey analysis, data collection by tracking points, and integration with trusted storage. Furthermore, feature extraction is performed on the collected flow path information. Based on this, this invention uses a heterogeneous multi-relationship spatiotemporal hypergraph to model a cross-border data flow network, ensuring a unified representation of multi-source information and complex relationships in cross-border data flow scenarios. In the graph construction, each node and edge has different multi-dimensional attributes. This method abstracts various entities, including data providers, processors, and receivers, as nodes in a graph. Nodes include attributes such as entity type (individuals, enterprises, public institutions, and service providers), industry category (internet, transportation, finance, and cross-border e-commerce), geographic location, and security level. Edges encompass multiple dimensions, including transmission behavior type (data transmission, storage, and processing), transmission timestamps (data sending time, receiving time, and duration), data volume (including significant and general data sizes), source / destination IP, transmission method (wired / wireless transmission, device transmission, and cloud service transmission), transmission protocol, data sensitivity, and encryption status. Each edge can be linked into a path based on upstream and downstream relationships. The nodes and edge attributes along the complete path collectively characterize the spatiotemporal sequence of data flow. Furthermore, data flow path information collection is continuously performed, and the heterogeneous multi-relationship spatiotemporal hypergraph structure is updated synchronously. The proposed heterogeneous graph model effectively captures the interaction relationships between different entities with spatiotemporal states in data export scenarios, describing their complete semantics. By recording the global data flow path and transmission business behavior details, it provides a rich foundation for subsequent anomaly analysis.
[0067] Specifically, the cross-border data flow network is modeled as a heterogeneous multi-relational spatiotemporal hypergraph with temporal and directional attributes, represented as follows: ,in, Represents a set of nodes (data body). This represents a set of hyperedges (representing a complete data flow path) and carries multi-dimensional attributes. For each hyperedge... , containing a sequence of nodes: an ordered list ( Based on the timing attributes of each flow event occurring at each node and the behavioral attributes of each flow event.
[0068] Furthermore, in the heterogeneous multi-relation spatiotemporal hypergraph, each node This represents a data body. Its attributes are as follows:
[0069] Each edge This represents a complete data flow behavior based on the flow path, and its attributes are as follows:
[0070] Based on this, the time dimension is introduced. Heterogeneous multi-relation spatiotemporal hypergraph It will be dynamically updated as cross-border data transactions are executed and over time:
[0071] In the formula, For the updated supermap, This is the supermap before the update. For the newly added set of nodes, The graph structure and attributes of the newly added edge set are updated in real time as the data acquisition system continues to run, adding new node and edge information.
[0072] In this embodiment, in step S2, abnormal behavior in cross-border data flows is often not only reflected in local single-point events, but also hidden in the evolution of multi-hop paths and data transmission trajectories. Therefore, focusing only on the attributes of a single node or edge is insufficient to fully reveal complex abnormal patterns. To comprehensively support the accurate identification of cross-border data abnormal behavior, this invention, based on the completion of heterogeneous multi-relational spatiotemporal hypergraph modeling, systematically conducts multi-level extraction and optimization selection of information-rich features of data flow paths. Specifically, this includes: node-level features, edge-level features, path-level features, and graph-level global features, to characterize the structure, attributes, and evolutionary features of data flow behavior at a multi-dimensional granularity.
[0073] In one possible embodiment, step S2, the step of extracting information-rich features of the data flow path from the node level, includes: S211. Obtain the basic attributes of the nodes; analyze the data processing frequency, data volume change trend, related subject number change trend, transmission method distribution and sensitive data ratio of the nodes within a unit time period, and construct a node behavior profile based on the time window. S212. Apply graph embedding technology to encode the topological location and neighborhood information of nodes in the network into low-dimensional vectors.
[0074] Specifically, nodes, as active actors in the data processing chain, possess heterogeneous attributes and dynamic behavioral characteristics. Node-level feature extraction not only enhances the understanding of complex relationships between nodes but also provides strong support for subsequent risk identification. In this section, the invention not only focuses on static information such as basic node attributes like subject type, industry category, geographical location, and security level, but also delves into the dynamic behavioral characteristics of each node in the data flow network. For example, by analyzing the frequency of data processing, the trend of data volume changes, the trend of the number of associated subjects, the distribution of transmission methods, and the proportion of sensitive data within a unit of time period, the activity and potential risks of each node can be assessed. Furthermore, by constructing a node behavior profile based on a time window, the flow behavior patterns of individuals in different time periods can be identified, providing a basis for subsequent anomaly analysis. In addition, graph embedding techniques (such as GraphSAGE or Node2Vec) are applied to encode the topological location of nodes in the network and their neighborhood information into low-dimensional vector representations to better capture the functional roles and potential risk patterns of nodes.
[0075] In one possible embodiment, step S2, the step of extracting information-rich features of the data flow path from the edge layer, includes: S221. Extract multi-dimensional interactive features including transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether it is encrypted; S222. Calculate the relative position weight of the current edge in the overall path, and the dependency relationship between the current edge and its surrounding edges; S223. Extract the flow direction of edges and the temporal consistency between edges and upstream and downstream nodes.
[0076] Specifically, edges represent the interactive behaviors between entities and are the core elements of path construction. Their multi-dimensional attributes should be characterized with fine-grained features. In this section, we first extract multi-dimensional interactive features such as transmission behavior type, transmission timestamp, data volume, source / destination IP, transmission method, transmission protocol, data sensitivity, and whether encryption is used. Simultaneously, combined with contextual information, we calculate the relative position weight of the current edge in the overall path (e.g., starting edge, intermediate jump edge, or ending edge) and its dependencies with other edges, forming semantically interpretable behavioral features. Based on this, we should also extract the edge's flow direction (from source node to destination node) and its temporal consistency with upstream and downstream nodes to help identify whether there are abnormal jumps or path avoidance in data flow.
[0077] In one possible embodiment, step S2, the step of extracting data flow path rich information features from the path hierarchy, includes: S231. By limiting the maximum path hop count threshold, use BFS or DFS search algorithms to generate a candidate path set, excluding loops and known whitelisted paths. S232. Construct path identifiers based on node type, edge type, and access strategy, and define typical path patterns; S233. Aggregate multiple transfer records with the same path pattern in historical data, and statistically analyze the frequency of occurrence, average amount of data transferred, and distribution of transfer time periods. S234. Extract path length, number of transit nodes, coverage of high-risk nodes, and path opacity index; S235. Sequence modeling combined with time-aware embedding is used to encode the temporal order of nodes and edges on the path.
[0078] Specifically, a path, as a combination of multiple edges and nodes, is a key unit reflecting the complete semantics of cross-border data flow. This involves two key steps: multi-hop path mining and path structure feature extraction. In multi-hop path mining, firstly, multi-hop transmission paths related to the monitored object are identified; that is, paths originating from a specific starting entity (e.g., a domestic device), passing through several intermediate nodes, and finally reaching the cross-border destination node. For a specific node in the graph, all possible destination paths are searched. By limiting the maximum path hop count threshold L, BFS or DFS search algorithms can be used to efficiently generate a candidate path set, excluding loops and known whitelisted paths. Secondly, path types should be classified. The concept of meta-paths can be introduced, constructing path identifiers based on factors such as node type, edge type, and access strategy, and defining typical path patterns (e.g., terminal → application system → transit service → third party → overseas) to distinguish different flow channels. Further, path behavior clustering analysis should be performed, aggregating multiple flow records with the same path pattern in historical data, and statistically analyzing their frequency of occurrence, average data volume, and transmission time period distribution to provide a statistical basis for subsequent modeling and analysis. In terms of path structure feature extraction, for individual paths, their structural features should be further extracted to characterize their potential risk complexity and transmission concealment. Firstly, features such as path length, number of transit nodes, coverage of high-risk nodes, and path opacity index (identifying the overall path encryption level to measure path concealment) are primarily extracted. Secondly, sequence modeling methods combined with time-aware embedding are used to encode the temporal order of nodes and edges on the path to capture the spatiotemporal continuity and periodic patterns of data flow. Simultaneously, the relationships between each node in the path and its upstream and downstream nodes are considered to identify key steps in specific business processes. The path-level feature extraction method proposed in this invention can not only characterize conventional cross-border data flow behavior patterns but also provide refined feature support and path anomaly clues for subsequent rule-driven and collaborative identification stages.
[0079] In one possible embodiment, step S2, the step of extracting data flow path rich information features from the layer level, includes: S241, Count the number of nodes, number of edges, average degree of nodes, degree of the node with the highest degree, number of superedges, and relation type; S242, Analyze graph density, average path length, network diameter, and global clustering coefficient; S243. Extract the distribution ratio of subject type and edge type, relationship diversity index, and node attribute entropy value; S244. Analyze the Gini coefficient, PageRank skewed distribution, cross-border data outbound path centroid, and information flow aggregation degree.
[0080] Specifically, graph-level feature extraction: After completing the extraction of node, edge, and path-level features, to comprehensively understand the global behavioral patterns and structural evolution trends of cross-border data flow networks, this invention further extracts graph-level features to construct a macroscopic characterization of the entire data flow system. In this section, firstly, structural indicators such as the number of nodes, edges, average node degree, degree of the node with the highest degree, number of superedges, and relation types are statistically analyzed at the global level to characterize the network scale and its internal connection density, assisting in identifying large-scale centralized data transmission and explosive growth. Secondly, topological attribute features such as graph density, average path length, network diameter, and global clustering coefficient are analyzed to measure the "compactness" and "information flow efficiency" of the entire network, detecting potential rapid data diffusion paths and abnormal clustering areas. Furthermore, the distribution ratio of each subject type and edge type, relation diversity index, and node attribute entropy values are extracted from the graph to quantify the degree of integration of different types of subjects and behaviors in the network, identifying the risk of "single-type dominance" or "emergence of atypical structures." Based on this, we analyze indicators such as the Gini coefficient, PageRank skewed distribution, cross-border data outbound path centroid, and information flow aggregation degree to determine whether the data shows a trend of excessive centralized transmission or control by a specific node cluster, in order to help identify "single-point sensitive entities" or "hidden transmission hubs." Graph-level features reflect the overall topological complexity of the network, the distribution characteristics of entities, the concentration of information transmission, and the degree of heterogeneous interaction, which can effectively support the identification of systemic anomalies and the modeling of collaborative behavior.
[0081] In this embodiment, step S3, the step of constructing a collaborative behavior relationship graph based on the rich information features, includes: S31. Using the data subject as a node, collaborative behavior edges are mined based on path structure similarity, behavior temporal coupling, and transmission destination similarity. The weight of the edge is determined by cross-path co-occurrence, behavior time window overlap, and semantic or business similarity. S32. Introduce a sliding time window mechanism to summarize and plot the flow events at different time scales; S33. Supplement the data flow links inherited from the original path as structural connection edges.
[0082] In cross-border data transfer scenarios, some abnormal behaviors are not completed independently by a single entity. Sensitive data flows exhibit characteristics such as "multi-person participation, multi-hop cover-up, and chain-like processing." These abnormal behaviors are constructed and deconstructed through collaboration among multiple entities. While each entity may perform seemingly compliant operations, they form an organized abnormal data flow pattern at the path level to achieve covert transmission, compliance evasion, and dispersal of responsibility. Such behaviors often exhibit fragmentation, strong asynchronicity, and indirect connections between paths. Traditional monitoring methods focused on localized cross-border behaviors or simple path analysis struggle to capture these hidden anomalies under "collaborative behavior." This invention, based on the aforementioned heterogeneous multi-relationship spatiotemporal hypergraph model and a multi-dimensional feature system rich in data flow path information, proposes a multi-entity collaborative behavior anomaly detection method. By defining collaborative identification problems and objectives, constructing a collaborative behavior relationship graph, mining collaborative behavior patterns, and identifying and judging anomalies, this method accurately discovers distributed, highly interconnected, multi-entity global collaborative abnormal behaviors, supporting management departments in controlling the risks of cross-border data flows.
[0083] The "abnormal collaborative behavior" defined in this invention refers to multiple entities that appear legitimate and compliant on the surface collaborating to covertly transfer or forward certain data, making it difficult to determine its abnormality from any single perspective. However, potential abnormal collaborative behavior can be revealed through global path joint analysis. Let... This refers to a set of multiple paths in a cross-border data flow. Each path involves a set of entity nodes and corresponding behavioral information. The identification target for abnormal collaborative behavior is the abnormal consistency or complementary collaborative patterns exhibited by multiple entities in terms of time, space, and path behavior. The key points of identification include: 1) Whether there is significant behavioral coupling or path overlap between different paths. 2) Whether multiple entities jointly form a complete data flow path; individuals may appear compliant, but the overall structure is abnormal. 3) Whether the collaborative behavior has repetitiveness, strategy detour characteristics, or historical collaborative traces.
[0084] See Figure 3 To effectively identify potential multi-entity hidden collaborative behavior chains in cross-border data flow processes and reveal the collaborative risk structure under indirect transmission relationships between entities, this invention introduces a Collaborative Behavioral Relational Graph (CBERP) based on the aforementioned heterogeneous multi-relationship graph modeling and path-rich feature extraction. The construction method is used to model the strength of the collaborative behavior between data subjects in the flow behavior.
[0085] The collaborative behavior graph differs from the data flow graph in the original heterogeneous graph. Its edges do not represent concrete single data transmission behaviors, but rather abstractly represent collaborative behavior patterns between two entities across multiple historical paths, such as high-frequency co-occurrence, overlapping temporal behaviors, or convergent target semantics. Collaborative behaviors can be quantified and mined using features such as path structure similarity, temporal coupling of behaviors, and similarity of transmission destinations. In this graph, if the entities... and If there are significant behavioral coordination characteristics, then in Establish a logical collaborative edge and assign synergy strength This is used to support subsequent collaborative behavior pattern recognition and anomaly detection. The weights of the logical collaborative edges are determined by... , and The components are represented as cross-path co-occurrence ( and The weighting method for edges can be defined as follows: the proportion of paths that appear together in all independent historical flow paths to the total number of paths, the degree of overlap of behavioral time windows, semantic or business similarity (e.g., whether they belong to the same business domain or application group, which can be judged based on characteristics such as data type, data volume, and protocol behavior).
[0086] Among them, parameters It can be optimized through data-driven methods.
[0087] Furthermore, to construct a connected subgraph structure of collaborative behavior relationships, the system can also introduce some structural connection edges, i.e., data flow links inherited from the original path, to ensure the structural integrity of collaborative behavior nodes and upstream and downstream paths. Although these edges do not represent the basis for collaborative behavior scoring, they can serve as a supplement to the graph structure connectivity, assisting in the construction of a complete collaborative propagation graph.
[0088] In the collaborative behavior graph, each edge is embedded with time-series information for subsequent identification of periodic collusion, distributed collaboration, or polling-style data outbound behavior. Based on this, time windows are modeled. A key characteristic of collaborative behavior is "weak temporal synchronicity," meaning that while the actions of the subjects are not completely synchronized, they have a perceptible temporal overlap. Therefore, this invention introduces a sliding time window mechanism: First, a time analysis window is set. Then, all flow events are summarized and plotted within the time window; secondly, multi-scale windows (such as 5 minutes, 30 minutes, 1 hour) are used for progressive analysis to improve the time sensitivity and robustness of the detection.
[0089] In one possible embodiment, in constructing a cooperative behavior relationship graph Building upon this foundation, the system has obtained high-frequency collaborative relationships among entities involved in cross-border data flows and their collaborative behavioral structures in multi-hop paths. However, a collaborative relationship graph alone is insufficient to directly reveal potential risks of anomalous collusion, as some collaborative relationships may originate from legitimate business processes or normal technical scheduling. Therefore, to accurately identify entity combinations and behavioral paths with anomalous risks, this invention further designs a collaborative behavior pattern mining mechanism. This mechanism aims to identify multi-entity combinations that deviate from normal collaborative paradigms, exhibit highly coupled behaviors but abnormally convergent goals, from the collaborative behavior relationship graph. Furthermore, it combines these combinations with behavioral temporal sequences, transmission semantics, and structural features to determine collaborative risks from multiple perspectives.
[0090] Based on the structural topological features of the collaborative behavior relationship graph, potential collaborative subject subgraphs are extracted as candidate pattern units. Considering that abnormal collaborative behaviors often exhibit specific structural patterns, such as chain collaboration, closed-loop collusion, or star-shaped distribution, dense subgraph mining algorithms (such as k-core, k-clique, or connected components under edge weight filtering) are used to screen node combinations with high edge weights, directed closed loops, or convergent path features. In particular, to capture indirect collaborative behaviors with weak connections but high-frequency linkages, a behavior path co-occurrence matrix is introduced. A collaborative similarity metric matrix is constructed by the number of times node pairs co-occur in historical paths and their temporal overlap, and graph partitioning methods such as spectral clustering are used to identify subject groups with high behavioral homogeneity. The above process can not only identify collaborative communities at the structural level but also reveal potential long-term cooperation chains at the path evolution level. The final output is a set of candidate collaborative behavior groups. Members within each group share significant structural or path-level similarities.
[0091] The collaborative behavior aggregation feature calculation stage is mainly used to evaluate the overall behavioral consistency and structural coupling level within the collaborative subgraph. This invention introduces a collaborative behavior aggregation feature calculation method to perform structural evaluation on the collaborative subgraph. This process differs from the calculation of local collaborative strength between node pairs during the collaborative behavior edge establishment process. Instead, it starts from the collaborative subgraph as a whole and measures whether it exhibits consistent and collaborative high-risk behavioral characteristics.
[0092] For each candidate collaborative behavior individual The collaborative characteristics are measured to support the effective identification of risk events, and the collaborative strength is defined. A comprehensive measure of the behavioral coupling and structural overlap between paths is expressed as follows:
[0093] in, This is the average value of the path structure similarity in the subgraph, used to measure whether the same subject combination frequently forms a short chain cooperative structure in multiple paths; The node attribute value measures the node attributes in the collaborative path as well as the sensitivity, volume, protocol type, and other indicators of the data involved, reflecting whether there is a risk of collaborative splitting in scenarios such as "similar attributes + suspicious total amount" of multiple paths; This indicates the degree of coordination between the path initiation and arrival times, assessing whether coordinated behaviors are concentrated in a specific time window, pointing to the abnormal characteristic of "short-term high-intensity linkage".
[0094] Among them, coefficient , representing the weight of multi-dimensional feature fusion. If there is obvious path collusion behavior in the subgraph, then its value is... Significantly higher than ordinary non-cooperative flow path combinations. For each candidate cooperative group... Through calculation This allows for further screening and output of the final set of candidate subgraphs for collaborative behaviors. .
[0095] The collaborative behavior pattern mining mechanism proposed in this invention can effectively capture potential collaborative behaviors of multiple subjects that cannot be covered by traditional rules. In particular, it can identify hidden abnormal patterns with the characteristics of "behavioral convergence + path dispersion + total amount avoidance + goal consistency" in cross-border data export scenarios, and constitute an important supplementary module for anomaly detection from a global path perspective.
[0096] In one possible embodiment, after completing the mining of collaborative behavior patterns, this invention further designs an anomaly identification and judgment mechanism that combines compliance background constraints with path behavior modeling. This aims to achieve multi-dimensional quantitative evaluation and accurate identification of complex anomaly behaviors across multiple granularities and time periods. This module not only supports the direct identification of explicit violations but also uncovers highly concealed collusive anomalies in potential collaborative paths, achieving panoramic detection of anomaly events.
[0097] First, at the path level, for each candidate subgraph... , path set Establish a joint judgment and scoring model, and define each path. The anomaly scoring function is as follows:
[0098] in, This indicates the risk level of matching the path rules, such as whether it violates rules on cross-border sensitive data or whether it bypasses compliant transit points. The deviation from the path's historical behavior is calculated by embedding the current path with its historical behavior into a vector. and mean vector The comparison uses cosine similarity as defined. . It indicates the abnormality of the path structure and measures its rarity compared to the distribution of similar path structures. It can be combined with the path frequency distribution or the rarity score generated by the graph neural network.
[0099] Based on path-level scoring, an overall anomaly score for the collaborative subgraph is further constructed. ,as follows:
[0100] Considering that some "collaborative violations" may not be directly reflected in the path rule scoring (such as exceeding data limits, collaborative deviance, etc.), the design parameters of this invention... and As an independent fusion weighting mechanism, it measures path rule risk and structural collaborative behavior risk separately, and then integrates them into the final anomaly score in a cumulative weighted manner.
[0101] in, , Further calculations are performed using a smoothing function:
[0102]
[0103] in: , and Used to control the growth rate (default value is 2, logic is similar to sigmoid); A value close to 1 indicates a severe path rule violation. A value close to 1 indicates that the collaborative behavior is relatively extreme.
[0104] Furthermore, to further improve the timing recognition capability, this invention introduces a sliding time window mechanism. The process of path extraction, subgraph construction, and anomaly scoring is repeatedly executed over a continuous time period, and a behavior evolution curve is constructed for each subject or path. By detecting abnormal fluctuation patterns across different time periods, complex time trends such as "accumulated risk increase," "periodic surge," or "abnormal drift" can be identified. Trend lines can be modeled using methods such as Simple Moving Average (SMA) and Exponential Weighted Average (EMA).
[0105] If the abnormal evolution curve of a certain entity or path crosses the risk threshold within a continuous time window... These are then marked as high-risk objects for subsequent management. Finally, when outputting the judgment results, this invention not only marks high-scoring paths and their collaborative relationships as abnormal, but also outputs them along with information such as meta-path patterns, subject profiles, historical behaviors, and transmission locations to support post-event evidence collection, compliance audits, and dynamic risk response strategies. Simultaneously, it integrates with a user-adjustable rule base (such as policy changes, sensitive type updates, etc.) to allow for online adjustments to rules and judgment thresholds, enhancing the system's real-time adaptability and strategy dynamism.
[0106] By introducing structured scoring functions, collaborative enhancement mechanisms, and temporal evolution modeling, the anomaly identification and judgment method proposed in this invention can effectively identify multi-subject global collaborative behavior anomalies that are difficult to detect with traditional simple path and local rule verification. It is particularly suitable for cross-border data flow scenarios under path-rich information conditions, and improves the ability to comprehensively perceive and respond to complex risks.
[0107] To further illustrate the application process of this invention in actual cross-border data flow scenarios and its advantages in anomaly identification, the following describes in detail the specific implementation process of the method of this invention using an exemplary business scenario, and deduces the anomaly identification results derived from rich information modeling of data flow paths, multi-dimensional feature extraction, and multi-subject collaborative analysis. This example aims to demonstrate the process logic and risk assessment capabilities of the method of this invention in practical applications.
[0108] (1) Instance scenario setting Suppose that a domestic technology company X has the following data transfer chain when conducting overseas business: Flow chain 1: Enterprise X’s headquarters node A is located in China. It sends user behavior logs (including sensitive data fields) to cloud platform node B located in China through the headquarters’ data scheduling platform.
[0109] Flow chain 2: Node B is a data processing platform responsible for formatting and encrypting logs, and then sending the processing results to two different third-party service nodes C and D (both are outsourced data service providers).
[0110] Flow chain 3: Node C uploads the data it receives to node E (cloud analytics service provider) located overseas.
[0111] Flow chain 4: Node D then transfers the data to another overseas node G through node F (VPN proxy exit).
[0112] In this invention, nodes E and G are both "cooperative analysts" of enterprise X, located in different countries. Superficially, the data transmission behavior between nodes appears independent, and each edge can be interpreted as a legally authorized transmission. However, the total amount of data ultimately sent overseas by nodes C and D exceeds compliance regulations. The method proposed in this invention will take a path-based approach, using the proposed method to identify and analyze multi-hop collaborative behavior in this process, gradually revealing potential abnormal collaborative behaviors. Based on this, the following analysis process will be conducted.
[0113] (2) Constructing a heterogeneous multi-relation spatiotemporal hypergraph In the above scenario, this invention constructs a heterogeneous multi-relationship spatiotemporal hypergraph based on the collected flow events. The node set includes: A (corporate headquarters), B (domestic processing platform), C, D (intermediary service providers), E, G (overseas recipients), and F (VPN proxy exit); the flow behavior set includes: A→B, B→C, B→D, C→E, D→F, and F→G. Each flow event carries attributes such as transmission timestamp, data volume, protocol type, data sensitivity, and whether it is encrypted; node attributes include: subject type (enterprise / outsourcing / overseas service provider), country of origin, industry tag, and security level. Therefore, two hyperedges can be formed, namely: (Including nodes A, B, C, and E) (Including nodes A, B, D, F, and G).
[0114] The graph structure is dynamically maintained and updates over time, forming several complete paths within the observation window.
[0115] (3) Feature extraction based on rich information of data flow path The method of this invention further extracts node-level, edge-level, path-level, and graph-level features to discover potential abnormal structures. For two logical paths, the features are as follows: path :
[0116] path :
[0117] Although the two paths appear to have different data receivers, this method extracts the following key features from them: 1) Path length: Path 2 is significantly longer than path 1, indicating the presence of an additional proxy node F; 2) Path sensitivity clustering: Both paths process sensitive log data from node A, and neither C nor D performs any desensitization or further encryption on the sensitive data during transmission; 3) Path structure similarity: The first part of the two paths is completely identical (A→B), while the middle and later parts show the characteristics of "branching and aggregation", indicating structural synergy; 4) Path goal convergence: Although E and G are geographically different, they have similar service roles, respectively responsible for behavioral data analysis and model feedback, and their business labels and industry affiliations are consistent.
[0118] Based on the feature hints extracted at this stage, B, as a central transit node, distributes sensitive data in multiple directions and ultimately aggregates it to a logically equivalent overseas processing node, which poses a potential risk of bypassing compliance control nodes.
[0119] (5) Construction of collaborative behavior relationship diagram After completing the modeling and multi-dimensional feature extraction of cross-border data flow paths, this invention further introduces a collaborative behavior graph to model the potential collaborative behavior relationships among multiple entities. The goal of the collaborative behavior graph is to capture the implicit collaborative features exhibited by multiple seemingly independent data entities during cross-border flow, such as high-frequency co-occurrence, temporal coupling, or goal consistency, providing a structural foundation for subsequent abnormal pattern recognition.
[0120] The collaborative behavior graph is a weighted graph consisting of the main nodes in the data flow path, collaborative behavior edges, and structural connection edges, where the weight of the collaborative behavior edges is... Reflecting the strength of their collaboration, the weights are determined by... , and The components are represented by cross-path co-occurrence, overlap of behavioral time windows, and semantic or business similarity, respectively.
[0121] Within the sliding time window, the system observed the following cooperative structure: Among them, the node pair (C, D) co-occurs frequently in multiple paths, with highly overlapping transmission times, forming a strong cooperative behavior edge; node B is a common transit point for two paths, connecting to C and D respectively, forming a "cooperative fan-out" pattern; node F is a path The unique hop point between D and G exhibits highly rare and concealed transmission behavior, thus forming a structural cooperative relationship with D and G. Although the node pair (E, G) has no path connection, their service functions are consistent and their target tasks are similar, exhibiting "target coupling" at the behavioral semantic layer. This can be processed in the semantic cooperative graph (but not directly as a behavioral edge in the cooperative behavioral relationship graph). Simultaneously, [the following settings are provided]. , and The parameters are 0.4, 0.3, and 0.3 respectively. Taking the collaborative behavior edge nodes C and D as examples, their weights are calculated. According to the behavior log statistics, the most recent N paths coexist in 60% of the time, and the behavior time periods are highly overlapping, which is calculated to be 0.85. In addition, the transmission destinations are the same and the service types are consistent, with an index of 0.9.
[0122]
[0123] Among them, the cooperation strength between nodes C and D reaches 0.765, which is a significant cooperative behavior edge. Therefore, considering the upstream and downstream path structure, the cooperative behavior relationship graph... The following connections should be established in the middle: 1)
[0124] 2) B→C, B→D: Structural collaborative edges (B is the key collaborative hub) 3) C→D: Behavioral co-occurrence + temporal coupling edge (this is a logical cooperative behavior edge, not a physical transmission edge) 4) D→F→G: Actual transmission behavior of the path (destination convergent nodes, structural collaborative edges) 5) C→E: Actual transmission behavior of the path (destination convergent nodes, structural collaborative edges) Note: A is the initiating node, not a key to collaboration, and is not involved in collaborative behavior modeling here.
[0125] (6) Collaborative behavior aggregation feature mining and calculation The above-mentioned cooperative behavior subgraphs were identified using a dense subgraph mining method (edge filtering + DFS). The conclusion is that this structure exhibits a typical branching + convergence + jump pattern, indicating a high-risk cooperative behavior structure. Further evaluation of the entire cooperative subgraph is needed. To determine whether a high-risk collaborative structure is constituted, this invention further introduces a collaborative behavior aggregation feature scoring function to perform overall modeling of subgraph structure, attributes, and temporal synchronization.
[0126]
[0127] The parameter is set to , , The calculations were performed based on structural characteristics, node attribute values, and time windows. The remaining observation results are as follows: , , .
[0128] This score indicates that the subgraph has highly coupled cooperative behavior characteristics and can serve as a core structural unit for anomaly identification.
[0129] (7) Anomaly identification and judgment This invention further evaluates path-level behavioral anomalies using the following path anomaly scoring function:
[0130] in, , , Further targeting the path and Based on whether the transmission behavior in the path violates compliance rules, the differences in historical behavior distribution, and the path structure pattern, the rule matching risk, behavior deviation, and path structure anomaly are calculated to obtain the following path anomaly scores:
[0131]
[0132] Further calculation of parameters in the anomaly score and To support the final anomaly identification:
[0133]
[0134]
[0135] By fusing path-level anomalies and graph-level collaboration strength, combined with adaptive parameters and A comprehensive risk assessment is performed on the anomalies of the subgraph to obtain an anomaly score:
[0136] If the final score is high, and the system sets the anomaly threshold to 0.5, then this subgraph is identified as a high-risk collaborative behavior structure, triggering an alarm mechanism. The results are also output as shown in Table 1. Table 1. Anomaly Scoring Results
[0137] In addition, to capture the cumulative abnormal behavior of collaborative entities within a continuous time window, this invention introduces an exponentially weighted moving average (EMA) mechanism to track the abnormal trends of each node.
[0138]
[0139] in, This is the smoothing coefficient, with a default value of 0.3. Taking node D as an example... The value is 0.75, and the current subgraph is... ,but:
[0140] This indicates that the node is showing a "gradual deterioration" trend and is marked as a "subject with continuous abnormal behavior". It should be included in the key monitoring list and automatically triggered by the strategy response mechanism to intercept, warn or collect evidence after the fact (the same trend judgment can also be made by using EMA and SMA for a certain path or subgraph).
[0141] This example comprehensively demonstrates how the method of the present invention, based on rich information about data flow paths, effectively identifies hidden abnormal transmission behaviors of multiple entities at the global level through collaborative graph construction, subgraph aggregation scoring, and anomaly identification and analysis. It has high accuracy, interpretability, and practicality.
[0142] It should be noted that the method of this embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this embodiment, and these multiple devices will interact with each other to complete the data cross-border risk monitoring method based on rich information about data flow paths.
[0143] It should be noted that the above description describes some embodiments of the present invention. In some cases, the described actions or steps can be performed in a different order than that shown in the above embodiments and the desired result can still be achieved. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0144] See Figure 4 Based on the same inventive concept, and corresponding to any of the methods in the above embodiments, this invention also provides a cross-border data risk monitoring device based on rich information about data flow paths, comprising: The heterogeneous multi-relationship spatiotemporal hypergraph construction module 100 is used to dynamically and continuously collect cross-border data flow path information and construct a heterogeneous multi-relationship spatiotemporal hypergraph based on the cross-border data flow path information. The rich information feature extraction module 200 is used to extract rich information features of data flow paths from four levels: nodes, edges, paths, and the global graph, based on the heterogeneous multi-relation spatiotemporal hypergraph. The anomaly monitoring module 300 is used to construct a collaborative behavior relationship graph based on the rich information features, and to monitor multi-subject collaborative abnormal behavior by mining collaborative behavior patterns and performing anomaly identification and judgment.
[0145] In this embodiment, the heterogeneous multi-relation spatiotemporal hypergraph construction module 100 includes: The node abstraction and attribute configuration submodule 101 is used to abstract data providers, processors and receivers as nodes in the graph. The nodes carry attributes such as subject type, industry category, region of origin and security level. The edge construction and attribute configuration submodule 102 is used to establish multiple types of edges based on data access and transmission behavior. The edges carry transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether they are encrypted. The hypergraph dynamic update submodule 103 is used for the dynamic update of the heterogeneous multi-relationship spatiotemporal hypergraph as cross-border data business is executed and time progresses.
[0146] In this embodiment, the rich information feature extraction module 200 includes a node-level feature extraction submodule 210 and an edge-level feature extraction submodule 220; The node-level feature extraction submodule 210 includes: Node attribute acquisition unit 211 is used to acquire the basic attributes of a node; The node dynamic behavior analysis unit 212 is used to analyze the node's data processing frequency, data volume change trend, related subject number change trend, transmission mode distribution and sensitive data ratio within a unit time period, and to construct a node behavior profile based on a time window. The node topology coding unit 213 uses graph embedding technology to encode the topological location and neighborhood information of a node in the network into a low-dimensional vector. The edge-level feature extraction submodule 220 includes: The edge multidimensional feature extraction unit 221 is used to extract multidimensional interactive features such as transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether it is encrypted. Edge relationship calculation unit 222 is used to calculate the relative position weight of the current edge in the overall path, and the dependency relationship between the current edge and the surrounding edges; Edge timing and flow analysis unit 223 is used to extract the flow direction of edges and the timing consistency between edges and upstream and downstream nodes.
[0147] In this embodiment, the rich information feature extraction module 200 includes a path-level feature extraction submodule 230, which includes: The candidate path generation unit 231 is used to generate a candidate path set by limiting the maximum path hop count threshold and using a BFS or DFS search algorithm to exclude loops and known whitelisted paths. The path pattern definition unit 232 is used to construct path identifiers based on node type, edge type, and access strategy, and to define typical path patterns. The path behavior aggregation unit 233 is used to aggregate multiple flow records of the same path pattern in historical data and to count the frequency of occurrence, average amount of data transmitted, and distribution of transmission time periods. The path structure feature extraction unit 234 is used to extract path length, number of transit nodes, coverage of high-risk nodes, and path opacity index. The path temporal coding unit 235 is used to encode the temporal order of nodes and edges on a path by using sequence modeling combined with time-aware embedding.
[0148] In this embodiment, the rich information feature extraction module 200 includes a graph-level feature extraction submodule 240, which includes: Graph structure index statistics unit 241 is used to count the number of nodes, the number of edges, the average degree of nodes, the degree value of the node with the highest degree, the number of super edges, and the relationship type. Graph topology attribute analysis unit 242 is used to analyze graph density, average path length, network diameter, and global clustering coefficient; The graph subject and behavior integration analysis unit 243 is used to extract the distribution ratio of subject type and edge type, relationship diversity index, and node attribute entropy value; The graph data transmission centralized analysis unit 244 is used to analyze the Gini coefficient, PageRank skewed distribution, cross-border data outbound path centroid, and information flow aggregation degree.
[0149] In this embodiment, the anomaly monitoring module 300 includes a collaborative behavior relationship graph construction submodule 310, which includes: The collaborative behavior edge mining unit 311 is used to mine collaborative behavior edges based on path structure similarity, behavior temporal coupling, and transmission destination similarity, with the data subject as the node. The weight of the edge is determined by cross-path co-occurrence, behavior time window overlap, and semantic or business similarity. The time window processing unit 312 is used to introduce a sliding time window mechanism to summarize and plot the flow events at different time scales. The structural connection edge supplement unit 313 is used to supplement the data flow link inherited from the original path as a structural connection edge.
[0150] In this embodiment, the anomaly monitoring module 300 includes a collaborative behavior pattern mining submodule 320, which includes: The cooperative behavior group identification unit 321 is used to identify candidate cooperative behavior groups from the cooperative behavior relationship graph using a dense subgraph mining algorithm and a spectral clustering method. Collaborative subgraph structural evaluation unit 322, used to assess collaborative strength The formula performs a structural evaluation of the collaborative subgraph:
[0151] In the formula, α + β + γ = 1, This represents the average path structure similarity in the subgraphs. For node attribute values, This represents the degree of coordination between path initiation and arrival times, with α, β, and γ being the weights for multi-dimensional feature fusion.
[0152] In this embodiment, the anomaly monitoring module 300 includes an anomaly identification and determination submodule 330, which includes: Path anomaly scoring unit 331 is used to establish the path anomaly scoring function:
[0153] In the formula, Indicates the risk level of path rule matching. The deviation from the path's historical behavior. Indicates path structure anomality. w , y , z For feature weights; Collaborative subgraph anomaly scoring unit 332 is used to construct the overall anomaly scoring function for the collaborative subgraph:
[0154] In the formula, μ For path rule risk weights, δ For structural collaborative behavior risk weights, P A set of paths; The abnormal fluctuation pattern identification unit 333 is used to introduce a sliding time window mechanism to construct a behavior evolution curve for the subject or path and identify abnormal fluctuation patterns.
[0155] In this embodiment, the anomaly monitoring module 300 includes a hypergraph update and weight calculation submodule 340, which includes: Hypergraph dynamic update unit 341 is used for the dynamic update of the heterogeneous multi-relation spatiotemporal hypergraph over time, and the formula is:
[0156] In the formula, For the updated supermap, This is the supermap before the update. For the newly added set of nodes, Add a new set of edges; The cooperative behavior edge weight calculation unit 342 is used to calculate the weights according to the weight formula of the cooperative behavior edges: The weights of the collaborative behavior edges are given by the formula:
[0157] In the formula, , , These are parameters that are optimized using a data-driven approach. For cross-path co-occurrence, The degree of overlap of behavioral time windows, For semantic or business similarity; the sliding time window includes 5-minute, 30-minute, and 1-hour scales.
[0158] In one possible embodiment, in the rich information feature extraction module 200, each hyperedge includes a node sequence, a temporal attribute based on each flow event occurring at each node, and a behavioral attribute of each flow event.
[0159] The apparatus described above is used to implement the cross-border data risk monitoring method based on data flow path rich information in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0160] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the cross-border data risk monitoring method based on rich information of data flow paths as described in any of the above embodiments.
[0161] Figure 5 This illustration shows a more specific hardware structure diagram of an electronic device provided in this embodiment. The device may include: a processor 410, a memory 420, an input / output interface 430, a communication interface 440, and a bus 450. The processor 410, memory 420, input / output interface 430, and communication interface 440 are interconnected internally via the bus 450.
[0162] The processor 410 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0163] The memory 420 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 420 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 420 and is called and executed by the processor 410.
[0164] Input / output interface 430 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.
[0165] The communication interface 440 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0166] Bus 450 includes a pathway for transmitting information between various components of the device, such as processor 410, memory 420, input / output interface 430, and communication interface 440.
[0167] It should be noted that although the above-described device only shows the processor 410, memory 420, input / output interface 430, communication interface 440, and bus 450, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0168] The electronic devices described in the above embodiments are used to implement a cross-border data risk monitoring method based on rich information of data flow paths in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0169] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, which are used to cause the computer to execute a data cross-border risk monitoring method based on rich information of data flow paths as described in any of the above embodiments.
[0170] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0171] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute a cross-border data risk monitoring method based on rich information of data flow path as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0172] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the invention is limited to these examples; within the framework of the invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of the different aspects of the embodiments of the invention as described above, which are not provided in detail for the sake of brevity.
[0173] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of the invention, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of the invention, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of the invention will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of the invention, it will be apparent to those skilled in the art that the embodiments of the invention may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0174] Although the invention has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAMDRAM) may be used with the embodiments discussed.
[0175] The embodiments of this invention are intended to cover all such substitutions, modifications, and variations falling within the scope of the claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this invention should be included within the scope of protection of this invention.
Claims
1. A data cross-border risk monitoring method based on rich information about data flow paths, characterized in that, Includes the following steps: Dynamically and continuously collect cross-border data flow path information, and construct a heterogeneous multi-relationship spatiotemporal hypergraph based on the cross-border data flow path information; Based on the heterogeneous multi-relation spatiotemporal hypergraph, information-rich features of data flow paths are extracted from four levels: nodes, edges, paths, and the global graph. Based on the rich information features, a collaborative behavior relationship graph is constructed. By mining collaborative behavior patterns and performing anomaly identification and judgment, abnormal collaborative behaviors of multiple entities are monitored.
2. The data cross-border risk monitoring method based on rich information of data flow paths according to claim 1, characterized in that, The steps for constructing a heterogeneous multi-relationship spatiotemporal hypergraph based on the cross-border data flow path information include: Data providers, processors, and receivers are abstracted as nodes in the diagram, with each node carrying attributes such as subject type, industry category, geographic region, and security level. Multiple types of edges are established based on data access and transmission behavior. Each edge carries the transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether it is encrypted. The heterogeneous multi-relationship spatiotemporal hypergraph is dynamically updated as cross-border data transactions are executed and time progresses.
3. The data cross-border risk monitoring method based on rich information of data flow paths according to claim 1, characterized in that, The steps for extracting information-rich features of data flow paths from the node level include: Obtain the basic attributes of the nodes; analyze the frequency of data processing, data volume change trend, number of associated entities change trend, transmission method distribution and sensitive data ratio of the nodes within a unit time period, and construct a node behavior profile based on the time window; Graph embedding techniques are used to encode the topological location and neighborhood information of nodes in the network into low-dimensional vectors; The steps for extracting information-rich features of data flow paths from the edge level include: Extract multi-dimensional interactive features including transmission behavior type, transmission timestamp, data volume, transmission source / destination IP, transmission method, transmission protocol, data sensitivity, and whether it is encrypted; Calculate the relative position weight of the current edge in the overall path, and the dependency relationship between the current edge and its surrounding edges; Extract the flow direction of edges and the temporal consistency between edges and their upstream and downstream nodes.
4. The data cross-border risk monitoring method based on rich information of data flow paths according to claim 1, characterized in that, The steps for extracting information-rich features of data flow paths from the path hierarchy include: By limiting the maximum path hop count threshold, a candidate path set is generated using BFS or DFS search algorithms, excluding loops and known whitelisted paths. Construct path identifiers based on node type, edge type, and access strategy, and define typical path patterns; Aggregate multiple transfer records with the same path pattern in historical data, and statistically analyze the frequency of occurrence, average amount of data transferred, and distribution of transfer time periods. Extract path length, number of transit nodes, coverage of high-risk nodes, and path opacity index; Sequence modeling combined with time-aware embedding is used to encode the temporal order of nodes and edges on the path.
5. The cross-border data risk monitoring method based on rich information of data flow paths according to claim 1, characterized in that, The steps for extracting rich information features of data flow paths from the layer level include: Statistical analysis includes the number of nodes, the number of edges, the average degree of each node, the degree of the node with the highest degree, the number of superedges, and the relationship type. Analyze graph density, average path length, network diameter, and global clustering coefficient; Extract the distribution ratio of subject type and edge type, relationship diversity index, and node attribute entropy value; Analyze the Gini coefficient, PageRank skewed distribution, cross-border data outbound path centroid, and information flow aggregation degree.
6. The data cross-border risk monitoring method based on rich information of data flow paths according to claim 1, characterized in that, The steps for constructing a collaborative behavior relationship graph based on the rich information features include: Using data subjects as nodes, collaborative behavior edges are mined based on path structure similarity, behavioral temporal coupling, and transmission destination similarity. The weight of the edge is determined by cross-path co-occurrence, behavioral time window overlap, and semantic or business similarity. A sliding time window mechanism is introduced to summarize and plot the flow events at different time scales; Supplement the data flow links inherited from the original path as structural connection edges.
7. The data cross-border risk monitoring method based on rich information of data flow paths according to claim 1, characterized in that, Discovering collaborative behavior patterns includes: A dense subgraph mining algorithm and a spectral clustering method are used to identify candidate cooperative behavior groups from the cooperative behavior relationship graph; Through synergy strength The formula performs a structural evaluation of the collaborative subgraph: ; In the formula, α + β + γ = 1, This represents the average path structure similarity in the subgraphs. For node attribute values, This represents the degree of coordination between the path initiation and arrival times, where α, β, and γ are the weights for multi-dimensional feature fusion. Anomaly identification and judgment include: Establish a path anomaly scoring function: ; In the formula, Indicates the risk level of path rule matching. The deviation from the path's historical behavior. Indicates path structure anomality. w , y , z For feature weights; Constructing an overall anomaly scoring function for the collaborative subgraph: ; In the formula, μ For path rule risk weights, δ For structural collaborative behavior risk weights, P A set of paths; A sliding time window mechanism is introduced to construct behavioral evolution curves for subjects or paths and identify abnormal fluctuation patterns.
8. The data cross-border risk monitoring method based on rich information of data flow paths according to claim 1, characterized in that, Each hyperedge contains a sequence of nodes, a temporal attribute based on each flow event occurring at each node, and a behavioral attribute for each flow event; The formula for the dynamic update of the heterogeneous multi-relation spatiotemporal hypergraph over time is: ; In the formula, For the updated supermap, This is the supermap before the update. For the newly added set of nodes, Add a new set of edges.
9. The data cross-border risk monitoring method based on rich information of data flow paths according to claim 6, characterized in that, The sliding time window includes 5-minute, 30-minute, and 1-hour scales; The weights of the collaborative behavior edges are given by the formula: ; In the formula, , , These are parameters that are optimized using a data-driven approach. For cross-path co-occurrence, The degree of overlap of behavioral time windows, For semantic or business similarity.
10. A data cross-border risk monitoring device based on rich information of data flow paths, employing the data cross-border risk monitoring method based on rich information of data flow paths as described in any one of claims 1 to 9, characterized in that, include: The heterogeneous multi-relationship spatiotemporal hypergraph construction module is used to dynamically and continuously collect cross-border data flow path information and construct a heterogeneous multi-relationship spatiotemporal hypergraph based on the cross-border data flow path information. The rich information feature extraction module is used to extract rich information features of data flow paths from four levels: nodes, edges, paths, and the global graph, based on the heterogeneous multi-relation spatiotemporal hypergraph. The anomaly monitoring module is used to construct a collaborative behavior relationship graph based on the rich information features, and to monitor multi-subject collaborative abnormal behavior by mining collaborative behavior patterns and performing anomaly identification and judgment.