Track portrait generation method and system based on artificial intelligence
The structured data flow map is generated through adaptive spatiotemporal crawler technology, and the spatial-temporal graph neural network and dynamic brain network evolution model are used, combining the cross-modal attention mechanism and nonlinear dynamic equations to generate risk-perceived trajectory portraits, solving the problem of difficult to capture the complex spatiotemporal dependencies of enterprise data flow in traditional methods, and achieving high-precision path prediction and anomaly recognition.
Patent Information
- Application Number
- CN202510598526.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional data analysis methods are difficult to capture the complex spatial and temporal dependencies of enterprise-level data flow and the correlation between actors, and cannot effectively improve enterprise security management, operational optimization and decision-making support.
Structured data flow maps are generated through adaptive spatiotemporal crawler technology, and the spatiotemporal graph neural network and dynamic brain network evolution model are used, combining cross-modal attention mechanisms and nonlinear dynamic equations to generate risk-perceived trajectory portraits.
It realizes high-precision prediction of enterprise data flow paths and identification of potential anomalies, improves the performance capabilities of trajectory portraits, and provides an advanced technical foundation for enterprise security management and decision-making support.
Smart Images

Figure CN120448743A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular to an artificial intelligence-based trajectory profile generation method and system. Background Art
[0002] With the rapid development of information technology and the widespread use of big data, enterprises and institutions across the globe generate massive amounts of heterogeneous data in their daily operations. This data includes user behavior logs, system operation records, and permission change events. Effectively analyzing and understanding the flow patterns, behavioral characteristics, and potential risks of this massive amount of data has become a crucial approach to improving enterprise security, optimizing business processes, and enabling intelligent decision-making. Traditional data analysis methods, which primarily rely on static data statistics and simple rule matching, struggle to capture the complex spatiotemporal dependencies between data and the correlations between actors. Especially in enterprise environments, where data flows are highly dynamic, diverse, and complex, a single analysis technique struggles to fully capture the underlying structure and evolving trends of the data. Summary of the Invention
[0003] The purpose of this invention is to provide a trajectory profile generation method and system based on artificial intelligence to address the deficiencies in the existing technology, improve the performance of trajectory profiles, and provide an advanced technical foundation for enterprise safety management, operation optimization, and decision support.
[0004] One embodiment of the present application provides a method for generating a trajectory profile based on artificial intelligence, the method comprising: Based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources, adaptive spatiotemporal crawler technology is used to extract multi-dimensional data flow characteristics and generate a structured data flow map; Based on the structured data flow graph, a spatiotemporal graph neural network is used to output a high-order semantic feature representation of the data flow by fusing the spatiotemporal dependency of the data flow and the correlation of the operator's behavior; Based on the high-level semantic feature representation, combined with data content sensitivity labels and business scenario context information, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector for the data trajectory; Based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and the potential anomaly diffusion path is predicted through nonlinear dynamic equations to generate a risk-aware trajectory portrait.
[0005] Optionally, the method of extracting multi-dimensional data flow features based on dynamic access logs, data operation behavior records, and permission change events of the enterprise's heterogeneous data sources through adaptive spatiotemporal crawler technology to generate a structured data flow graph includes: Perform multi-source heterogeneous data cleaning based on enterprise database access logs, API call records, and terminal operation logs to remove invalid timestamps and duplicate operation records, and obtain standardized data flow event sequences. Based on the standardized data flow event sequence, the chaos theory coding method is used to perform spatiotemporal coding on the operation subject, data object and permission change, and the spatiotemporal correlation feature matrix is obtained. Based on the spatiotemporal correlation feature matrix, the data flow path is dynamically tracked through adaptive spatiotemporal crawler technology to construct a dynamic correlation map including time dimension, space dimension and operation dimension; Based on the dynamic association graph, the graph structure self-correction algorithm is used to correct the abnormal connection edge weights and generate the final structured data flow graph.
[0006] Optionally, the structured data flow graph is based on the structured data flow graph, and a spatiotemporal graph neural network is used to fuse the spatiotemporal dependency of the data flow and the correlation between the behavior of the operating subject to output a high-order semantic feature representation of the data flow, including: Based on the structured data flow graph, the node features and edge weights of the spatiotemporal graph neural network are initialized. The node features include the identity of the operation subject, the data object type, and the timestamp code. Based on the initialization node features, the spatiotemporal dependencies of data flow are captured through the spatiotemporal convolution layer to obtain the spatiotemporal fusion feature tensor; Based on the spatiotemporal fusion feature tensor, the behavioral association modeling algorithm is used to calculate the behavioral similarity matrix between the operating subjects and generate the subject behavior association features; Based on the subject behavior correlation features, the spatiotemporal features and behavioral features are fused through the dynamic aggregation layer to output a multi-dimensional joint feature vector; According to the multi-dimensional joint feature vector, the feature dimensionality reduction algorithm is used to remove redundant information and generate a high-order semantic feature representation of the data flow.
[0007] Optionally, the step of generating a global embedding vector of the data trajectory by fusing multi-source semantic features through a cross-modal attention mechanism based on the high-order semantic feature representation, combined with data content sensitivity labels and business scenario context information, includes: Based on high-level semantic feature representation, the semantic vector of the data content sensitivity label and the natural language vector of the business scenario context description are extracted; Based on the semantic vector and natural language vector, modality alignment processing is performed through a cross-modal encoder to generate cross-modal alignment features; Based on the cross-modal alignment features, the multi-head attention mechanism is used to calculate the semantic association weights between different modalities and generate an attention weight matrix; Based on the attention weight matrix, the feature fusion layer integrates cross-modal semantic information and outputs a preliminary global embedding vector; According to the preliminary global embedding vector, a generative adversarial network is used to optimize the embedding space distribution and generate the global embedding vector of the final data trajectory.
[0008] Optionally, based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate a risk-aware trajectory portrait, including: Based on the global embedding vector, the nodes and connecting edges of the dynamic brain network are constructed, where the nodes represent the key entities of data flow and the edge weights represent the diffusion probability; Based on the dynamic brain network, nonlinear dynamic equations are used to simulate the path evolution process of data flow and calculate the chaotic diffusion coefficient between nodes; According to the chaotic diffusion coefficient, the path prediction algorithm is used to identify potential abnormal diffusion paths and generate a set of high-risk paths; Based on a set of high-risk paths, the neural radiation field technology is combined to render a three-dimensional trajectory heat map and output a risk-aware trajectory portrait.
[0009] Optionally, the method of simulating the path evolution of data flow based on a dynamic brain network using nonlinear dynamic equations and calculating the chaotic diffusion coefficient between nodes includes: According to the node state vectors and initial weights of the connection edges of the dynamic brain network, the coupling relationship between nodes is established using the Lorentz dynamics equation to generate a nonlinear interaction matrix. Based on the nonlinear interaction matrix, the diffusion sensitivity factor of data flow is introduced, and the node state evolution trajectory is iteratively calculated through the chaotic mapping algorithm to obtain the spatial distribution of chaotic attractors. According to the spatial distribution of chaotic attractors, phase space reconstruction technology is used to extract the trajectory correlation characteristics between nodes and generate trajectory similarity measurement values; Based on the trajectory similarity measure, the uncertainty of the diffusion process is quantified by the maximum Lyapunov exponent algorithm, and the chaotic diffusion coefficient between nodes is output.
[0010] Optionally, identifying potential abnormal diffusion paths through a path prediction algorithm based on the chaotic diffusion coefficient to generate a high-risk path set includes: According to the chaotic diffusion coefficient and the connection edge weight of the dynamic brain network, a path diffusion probability matrix is constructed, in which the probability value is nonlinearly positively correlated with the chaotic diffusion coefficient; Based on the path diffusion probability matrix, the hidden Markov model is used to predict the multi-step transfer path, and the path transfer confidence is calculated in combination with the time decay factor. Based on the path transfer confidence, the candidate abnormal diffusion paths are generated through the Monte Carlo tree search algorithm, and high-risk paths are screened according to the risk threshold; Based on the screened high-risk paths, the path aggregation algorithm is used to merge overlapping path segments and optimize path weights to output the final high-risk path set.
[0011] Another embodiment of the present application provides an artificial intelligence-based trajectory profile generation system, the system comprising: The extraction module is used to extract multi-dimensional data flow characteristics based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources through adaptive spatiotemporal crawler technology to generate a structured data flow map; A fusion module is used to output a high-order semantic feature representation of the data flow by fusing the spatiotemporal dependency of the data flow and the correlation of the operation subject behavior based on the structured data flow graph and using a spatiotemporal graph neural network; A generation module is used to generate a global embedding vector of the data trajectory based on the high-order semantic feature representation, combined with the data content sensitivity label and business scenario context information, and fusing multi-source semantic features through a cross-modal attention mechanism; The simulation module is used to simulate the path evolution process of data flow based on the global embedding vector using a dynamic brain network evolution model, predict potential abnormal diffusion paths through nonlinear dynamic equations, and generate a risk-aware trajectory portrait.
[0012] Yet another embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute any of the above methods when run.
[0013] Yet another embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute any of the above methods.
[0014] Compared with the existing technology, the present invention provides an artificial intelligence-based trajectory portrait generation method, which generates a structured data flow graph based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources; based on the structured data flow graph, the spatiotemporal graph neural network is used to output high-order semantic feature representation of data flow; based on the high-order semantic feature representation, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory; based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate risk-aware trajectory portraits, thereby improving the performance of trajectory portraits and providing an advanced technical foundation for enterprise security management, operation optimization and decision support. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1A hardware structure block diagram of a computer terminal for an artificial intelligence-based trajectory portrait generation method provided in an embodiment of the present invention; Figure 2 A schematic diagram of a process for generating a trajectory profile based on artificial intelligence provided by an embodiment of the present invention; Figure 3 A schematic structural diagram of an artificial intelligence-based trajectory profile generation system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limiting the present invention.
[0017] The embodiment of the present invention first provides a trajectory portrait generation method based on artificial intelligence, which can be applied to electronic devices such as computer terminals, specifically ordinary computers.
[0018] The following describes it in detail by taking running on a computer terminal as an example. Figure 1 The hardware structure block diagram of a computer terminal for a trajectory profile generation method based on artificial intelligence provided by an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0019] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to execute any one of the artificial intelligence-based trajectory profile generation methods.
[0020] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0021] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any one of the trajectory portrait generation methods based on artificial intelligence.
[0022] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0023] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0024] See also Figure 2 , an embodiment of the present invention provides a trajectory portrait generation method based on artificial intelligence, which may include the following steps: S201: Based on the dynamic access logs, data operation behavior records, and permission change events of the enterprise's heterogeneous data sources, the adaptive spatiotemporal crawler technology is used to extract multi-dimensional data flow characteristics and generate a structured data flow graph; Adaptive spatiotemporal crawler technology intelligently collects and processes multi-source heterogeneous enterprise data, including dynamic access logs, operational behavior records, and permission change events. The crawler system automatically adjusts collection strategies based on the spatiotemporal characteristics of the data source, such as using real-time streaming for frequently changing data and batch supplemental collection for historical data. Ultimately, it generates a structured data flow graph that incorporates correlations across time, space, and operational dimensions. This addresses the pain point of unified analysis of multi-source heterogeneous data in enterprise data governance. The structured graph intuitively displays the full picture of data flow, laying the data foundation for subsequent high-level feature extraction and risk prediction. The adaptive nature of the spatiotemporal crawler significantly improves the coverage and timeliness of data collection.
[0025] S202: Based on the structured data flow graph, using a spatiotemporal graph neural network, by fusing the spatiotemporal dependency of the data flow and the correlation of the operator's behavior, output a high-order semantic feature representation of the data flow; A spatiotemporal graph neural network (GNN) is used to perform deep feature learning on data flow graphs. The network structure consists of a spatiotemporal convolutional layer and a dynamic aggregation layer. The spatiotemporal convolutional layer uses three-dimensional convolution kernels to simultaneously capture the temporal sequence patterns and spatial propagation patterns of data flow. The dynamic aggregation layer uses an attention mechanism to fuse the behavioral correlation features between operators, ultimately outputting high-level features with semantic representation capabilities. This breaks the limitation of traditional methods that only analyze single-dimensional features. By jointly modeling spatiotemporal dependencies and behavioral correlations, it more accurately characterizes the complex patterns of data flow. The high-level semantic features provide high-quality input representations for subsequent cross-modal fusion.
[0026] S203, based on the high-order semantic feature representation, combined with the data content sensitivity label and business scenario context information, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory; A cross-modal attention mechanism is designed to integrate three key features: high-level semantic features of data flow, semantic vectors of content sensitivity labels, and natural language descriptions of business scenarios. A multi-head attention layer calculates the correlation weights of features from different modalities, and a generative adversarial network optimizes the spatial distribution of features to generate globally consistent embeddings. This enables a deep fusion of technical features and business knowledge, ensuring that the generated trajectory representations conform to data flow patterns and meet business security requirements. The adversarial training strategy effectively improves the discriminability and generalization capabilities of the embeddings.
[0027] S204: Based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate a risk-aware trajectory portrait.
[0028] The dynamic brain network model is incorporated into data trajectory analysis, simulating the diffusion process of data flow through nonlinear dynamic equations. The model maps data entities to brain network nodes, calculates the diffusion coefficients between nodes using chaos theory, and combines neural radiation field technology to generate a visual risk heat map. This enables the first dynamic simulation prediction of data flow paths, enabling the early detection of potential abnormal diffusion paths that are difficult to detect with traditional methods. The three-dimensional trajectory heat map intuitively displays risk distribution, providing a direct basis for safety decision-making.
[0029] Specifically, based on the dynamic access logs, data operation behavior records, and permission change events of the enterprise's heterogeneous data sources, adaptive spatiotemporal crawler technology is used to extract multi-dimensional data flow characteristics and generate a structured data flow graph, including: Perform multi-source heterogeneous data cleaning based on enterprise database access logs, API call records, and terminal operation logs to remove invalid timestamps and duplicate operation records, and obtain standardized data flow event sequences. The core goal of data cleaning is to unify the format of multi-source heterogeneous data and eliminate noise and redundancy. Taking a company's database access logs, API call records, and terminal operation logs as an example, the raw data may have the following problems: Confused timestamp formats: For example, database logs use Unix timestamps (such as 1630000000), API records use ISO 8601 format (such as "2023-08-25T14:30:00Z"), and terminal logs may contain incomplete dates (such as "Aug25 14:30").
[0030] Repeated operation records: For example, the same terminal may record the same operation multiple times due to network retransmission (such as "File A upload" repeated 3 times).
[0031] Invalid fields: such as missing key fields (the user ID is empty) or illegal characters (such as the "Operation Type" field contains garbled characters).
[0032] The cleaning process is as follows: Timestamp normalization: Use a regular expression to match timestamps in different formats and convert them to nanosecond Unix timestamps (such as 1630000000000). For example, parse "2023-08-25T14:30:00Z" into the Unix timestamp 1692981000000.
[0033] For incomplete timestamps (such as those containing only the date), interpolation is used to complete the timestamp to 00:00:00 on the current day, and the confidence level is marked as low (0.3).
[0034] Deduplication processing: Calculate the MD5 hash value of each record (based on core fields such as operation type, user ID, and data object ID). If the hash values are repeated and the timestamp difference is less than 1 second, the record is considered a duplicate and only the first one is retained.
[0035] For example, the hash values of the three "File A Upload" records in the terminal operation log are the same and the timestamp interval is less than 0.5 seconds. In the end, only the first record is retained.
[0036] Invalid field filtering: Define validation rules: User ID must be in UUID format (such as 550e8400-e29b-41d4-a716-446655440000), and the operation type must belong to a predefined set (such as "read, write, delete, modify").
[0037] Use rule-based filters (such as Apache Nifi's ReplaceText processor) to replace or remove illegal fields. For example, replace the garbled "Operation Type" field with "Unknown" and ignore such records in subsequent analysis.
[0038] Output standardized data flow event sequence: Each event contains the following fields: Event ID: UUID format (e.g., 550e8400-e29b-41d4-a716-446655440000); Timestamp: Unix timestamp (precision 1ns); Operation subject: user ID or service account; Data object: database table name or API endpoint; Operation type: standardized tags (read, write, delete, modify); Confidence level: 0.3 (low) ~ 1.0 (high).
[0039] Based on the standardized data flow event sequence, the chaos theory coding method is used to perform spatiotemporal coding on the operation subject, data object and permission change, and the spatiotemporal correlation feature matrix is obtained. Chaos theory encoding maps discrete events to a continuous spatiotemporal feature space through a nonlinear dynamic system to capture the complex correlations of data flows. The specific steps are as follows: Operation subject code: Each user / service account is assigned a unique initial value x0∈(0,1), and is mapped to the original value by Logistic (chaos equation: x n+1 = μx n (1-x n ), μ=3.99) to generate chaotic sequences.
[0040] For example, if user A's x0=0.3, after 10 iterations, the sequence is [0.3, 0.84, 0.45, 0.88, ...]. The mean of the last five iterations is taken as the encoding vector (e.g., [0.84, 0.45, 0.88, 0.37, 0.92]).
[0041] The permission change event (e.g. user A is granted the administrator role) adjusts the chaotic trajectory by a perturbation factor Δ=0.1 to generate a mutation code (e.g. the original sequence mean is 0.69 → 0.79 after mutation).
[0042] Data object encoding: Perform word embedding (Word2Vec, dimension 128) on the database table name or API endpoint, and then map the embedding vector to three-dimensional space through a chaotic system (such as the Lorenz equation).
[0043] For example, the word embedding of “user table” is mapped to the coordinates (12.3, -5.6, 29.1) after iterating the Lorenz equation.
[0044] Spatiotemporal correlation construction: Define the time and space windows: the time window is 1 hour, and the space window is the data objects within the same business module (such as the "order module" includes the order table, payment API, etc.).
[0045] Within the window, the chaotic coded cosine similarity between the operation subject and the data object is calculated to generate an association strength matrix. For example, the similarity between user A and the "order table" is 0.87, and with the "payment API" is 0.62.
[0046] Output spatiotemporal correlation feature matrix: The matrix dimensions are N×M×K, where: N: the number of operating entities (e.g., 1000 users); M: number of data objects (e.g., 200 tables / APIs); K: Feature dimension (e.g. chaotic code length 5 + spatial coordinate 3 = 8 dimensions).
[0047] The matrix elements represent the association strength between a specific subject and an object within a spatiotemporal window. For example, the association strength between user A and the “order table” within the time window [1630000000000, 1630003600000] is 0.87.
[0048] Based on the spatiotemporal correlation feature matrix, the data flow path is dynamically tracked through adaptive spatiotemporal crawler technology to construct a dynamic correlation map including time dimension, space dimension and operation dimension; The adaptive spatiotemporal crawler dynamically constructs the topological structure of data flow by analyzing the spatiotemporal correlation matrix in real time. Key technologies include: Path tracing algorithm: Based on the improved depth-first search (DFS), edges with high correlation strength (e.g., strength > 0.7) are tracked first.
[0049] Time dimension constraint: Only the association between adjacent time windows (such as t and t+1 hours) is tracked to avoid cross-day noise.
[0050] Spatial dimension constraints: Data objects within the same business module constitute a subgraph, and connections between modules must meet the cross-module operation frequency threshold (e.g., ≥5 times / hour).
[0051] Dynamic weight update: The initial edge weight is the spatiotemporal correlation strength (e.g., 0.87); The weight is dynamically adjusted based on the real-time event stream: if a new operation is detected on the same path (such as user A accessing the "Order Table" again), the weight is increased by Δ=0.1; if it exceeds the threshold (such as weight > 1.0), it is split into multiple paths.
[0052] Graph construction: Node type: operation subject (user), data object (table / API), permission entity (role / policy); Edge Type: Operation edge: user → table, weight = association strength; Permission edge: role → user, weight = permission level (e.g., administrator = 1.0, ordinary user = 0.5); Spatial edge: table → API, weight = call frequency (times / hour).
[0053] Example: User A (node 1) accesses the "Order Table" (node 2) 10 times within time window t, generating an operation edge weight of 0.9; User A is granted the "Administrator" role (node 3), generating a permission edge weight of 1.0; The "Order Table" (node 2) calls the "Payment API" (node 4) 50 times per hour, generating a spatial edge weight of 0.8.
[0054] Based on the dynamic association graph, the graph structure self-correction algorithm is used to correct the abnormal connection edge weights and generate the final structured data flow graph.
[0055] The goal of graph structure self-correction is to eliminate noisy edges (such as false connections caused by occasional operations) and optimize weight accuracy. Key technologies include: Abnormal edge detection: Statistical method: Calculate the Z-score of the edge weight and remove outliers with |Z|>3 (for example, for an edge with weight 0.9, when the mean is 0.4 and the standard deviation is 0.2, Z=2.5 is retained; for an edge with weight 1.5, Z=5.5 is removed).
[0056] Machine learning method: Train the Isolation Forest model to identify low-frequency paths (e.g., <2 times per hour) as abnormal.
[0057] Weight Modifier: Local smoothing: For outlier edge weights, replace them with a sliding average of their neighboring edge weights (window size = 5). For example, an outlier edge weight of 1.5 is replaced with the average of its five neighboring edges, 0.6.
[0058] Global optimization: Graph Convolutional Network (GCN) is used to reconstruct edge weights. The loss function includes mean squared error (prediction weight vs observation weight) and graph Laplace regularization term (enforced smoothness).
[0059] Structured output: Node attributes: chaotic code of the operating subject, spatial coordinates of the data object, and permission level; Edge attributes: weight, operation type, time window; Storage format: Neo4j graph database, supporting Cypher query language.
[0060] Example of the final structured data flow graph: Node 1 (User A): Attributes = {chaos code [0.84, 0.45, 0.88, 0.37, 0.92], Role = Administrator}; Node 2 (order table): attribute = {spatial coordinate (12.3, -5.6, 29.1)}; Edge 1 (User A → Order Table): Attributes = {weight 0.7, operation type = write, time window = 1630000000000~1630003600000}; Edge 2 (User A → Admin Role): Attribute = {Weight 1.0}.
[0061] Specifically, based on the structured data flow graph, a spatiotemporal graph neural network is used to fuse the spatiotemporal dependencies of data flow and the behavioral relevance of the operator to output a high-level semantic feature representation of the data flow, including: Based on the structured data flow graph, the node features and edge weights of the spatiotemporal graph neural network are initialized. The node features include the identity of the operation subject, the data object type, and the timestamp code. The structured data flow graph represents data flow relationships in the form of a graph structure. Nodes represent data entities (such as users, database tables, API interfaces), and edges represent data flow events (such as access, modification, and transmission). The initialization process requires encoding heterogeneous information into numerical features: Operator identity code: One-hot encode the user ID and role (administrator, ordinary user). For example, the administrator is encoded as 1,0, and the ordinary user is encoded as 0,1. Combined with behavioral history (such as login frequency and operation type), a 32-dimensional embedding vector is generated and trained on historical logs using the Word2Vec algorithm. For example, users who frequently perform data exports may be mapped to vectors 0.7, −0.3, ..., 0.2.
[0062] Data object type encoding: Structured data (such as database tables) and unstructured data (such as log files) are categorized and coded using a hierarchical coding strategy. For example, the database table "user_info" is coded as 1,0,0, and the log file "access.log" is coded as 0,1,0; Add data sensitivity labels (such as PII personal identity information marked as 1,0, non-sensitive data as 0,1), and automatically label the data content by matching regular expressions.
[0063] Timestamp encoding: Convert the timestamp to a periodic encoding (sine / cosine function). For example, the time "2023-10-01 14:30" is encoded as: Hour part: sin(14 / 24×2π)=0.96, cos(14 / 24×2π)=0.28; Minute part: sin(30 / 60×2π)=1.0, cos(30 / 60×2π)=0.0; A time decay factor is introduced to give more recent operations a higher weight (e.g., events within 3 days have a weight of 1.0, events within 7 days have a weight of 0.7, and events within 30 days have a weight of 0.3).
[0064] Edge weight initialization: Static weight: Set a basic weight based on the operation type, for example, "read" has a weight of 0.3, "write" has a weight of 0.6, and "delete" has a weight of 0.9; Dynamic weight: Dynamically adjusted based on operation frequency. For example, if a user accesses a table 100 times per hour, the weight increases by 0.1 per access, with an upper limit of 1.0.
[0065] Example: A node feature vector is user embedding (32 dimensions) | data object encoding (3 dimensions) | time encoding (4 dimensions), with a total dimension of 39. In the edge weight matrix, the "write" edge from user A to table B has an initial weight of 0.6, which increases to 0.85 with the number of operations.
[0066] Based on the initialization node features, the spatiotemporal dependencies of data flow are captured through the spatiotemporal convolution layer to obtain the spatiotemporal fusion feature tensor; The spatiotemporal convolution layer is composed of temporal convolution (TCN) and spatial graph convolution (GCN) in parallel, which extract temporal and topological features respectively: Temporal Convolution (TCN): Input: Node features are sliced by time window (e.g., 60 minutes) to form a time series tensor (number of nodes × time step × feature dimension). Convolution kernel: one-dimensional causal convolution, kernel size 3, stride 1, number of channels 64. For example, if a user's access records for three consecutive hours are 0.5, 0.7, and 0.9, the output feature is 0.5×0.2+0.7×0.5+0.9×0.3=0.66; Activation function: GELU (Gaussian Error Linear Unit), which enhances nonlinear expression capabilities.
[0067] Spatial Graph Convolution (GCN): Adjacency matrix: constructed based on edge weights and symmetric normalized (self-loop edges are added and the degree matrix D is normalized to the power of -1 / 2); Graph convolution formula (conceptual description): Aggregate adjacent node features and perform weighted summation. For example, the feature update of node A is: 0.6×A’s own feature + 0.3×adjacent node B’s feature + 0.1×adjacent node C’s feature; Parameter setting: 2 layers of GCN, each layer outputs 64 dimensions, and residual connections are used to prevent gradient disappearance.
[0068] Feature fusion: The temporal features (64 dimensions) output by TCN and the spatial features (64 dimensions) output by GCN are spliced by channel to obtain 128-dimensional spatiotemporal fusion features; Dynamically assign weights to temporal and spatial features through a gating mechanism (such as the sigmoid function). For example, the temporal weight is 0.7 and the spatial weight is 0.3.
[0069] For example, the features of a database table node in a time window are 0.8, 0.2, ..., 0.5 (64 dimensions) after TCN, 0.3, 0.6, ..., 0.1 (64 dimensions) after GCN, and 128-dimensional vector after fusion. After gating and weighting, the final value is 0.7 × 0.8 + 0.3 × 0.3 = 0.65, ..., 0.7 × 0.5 + 0.3 × 0.1 = 0.38.
[0070] Based on the spatiotemporal fusion feature tensor, the behavioral association modeling algorithm is used to calculate the behavioral similarity matrix between the operating subjects and generate the subject behavior association features; The goal of behavioral correlation modeling is to quantify the similarity of operation patterns of different users or systems: Feature Alignment: The spatiotemporal fusion features are grouped by the operation subject (user ID), and the feature dimension of each subject is (number of operations × 128); Pad or truncate variable-length sequences to a fixed length (e.g., the last 100 operations).
[0071] Similarity calculation: Cosine similarity: directly calculates the angle between feature vectors. For example, the similarity between user A and user B is 0.82. Dynamic Time Warping (DTW): Aligns the time offsets of operation sequences and calculates the minimum path distance. For example, the DTW distance between user A's "morning peak access pattern" and user B's "midday batch operations" might be 15.3. Self-attention similarity: Extract sequence-level features through the Transformer encoder and calculate the similarity between CLS tokens.
[0072] Matrix construction: Construct an N×N similarity matrix (N is the number of subjects), where the matrix element S_ij represents the similarity between subjects i and j; Introduce threshold filtering (e.g., similarity < 0.3 is set to 0) to reduce noise correlation.
[0073] Context feature generation: For each subject, extract the mean of the feature vectors of its top-K similar subjects (e.g., K=5) as its associated feature; The associated features are concatenated with the original spatiotemporal features to form enhanced features (128+64=192 dimensions).
[0074] Example: The similarity matrix of user A shows that the similarities with users B and C are 0.8 and 0.6 respectively. The associated feature is (user B's feature × 0.8 + user C's feature × 0.6) / 1.4.
[0075] Based on the subject behavior correlation features, the spatiotemporal features and behavioral features are fused through the dynamic aggregation layer to output a multi-dimensional joint feature vector; The dynamic aggregation layer uses a gating mechanism and cross attention to achieve multi-feature fusion: Gated Weighted Fusion: Input: spatiotemporal features (128 dimensions) and behavioral association features (64 dimensions); Gating signal generation: The weight vector is learned through the fully connected layer, for example, the spatiotemporal feature weight is 0.6 and the behavioral feature weight is 0.4; Weighting formula (conceptual description): fusion feature = 0.6 × spatiotemporal feature + 0.4 × behavioral feature.
[0076] Cross-Attention Mechanism: Query vector: spatiotemporal features; Key-Value vector: behavior association features; Calculate attention scores: for example, the "high-frequency access pattern" in spatiotemporal features and the "batch operation history" in behavioral features; Output: Attention-weighted features normalized by softmax.
[0077] Residual Connection: Add the fused features to the original spatiotemporal features to preserve the underlying information; For example: 192-dimensional feature fusion + 128-dimensional residual connection → 320-dimensional, and then compressed to 256-dimensional through the fully connected layer.
[0078] For example, a user's spatiotemporal features are 0.5, 0.3, ..., its behavior-related features are 0.7, 0.2, ..., and its gating weights are 0.6:0.4. After fusion, the features are 0.5 × 0.6 + 0.7 × 0.4 = 0.58, 0.3 × 0.6 + 0.2 × 0.4 = 0.26, ....
[0079] According to the multi-dimensional joint feature vector, the feature dimensionality reduction algorithm is used to remove redundant information and generate a high-order semantic feature representation of the data flow.
[0080] The dimensionality reduction process needs to retain key semantic information while compressing dimensions to improve computational efficiency: Principal Component Analysis (PCA): Calculate the covariance matrix for the 256-dimensional features and retain the first 50 principal components (explaining 95% of the variance); For example, a certain eigenvector becomes 50-dimensional 2.3,−1.7,...,0.5 after PCA.
[0081] t-SNE visualization guides dimensionality reduction: During the training phase, t-SNE was used to map the 256-dimensional dataset to 2-dimensional datasets to observe the clustering effect. Adjust the retained dimensions based on the visualization results (for example, select 30 dimensions for best separability).
[0082] Autoencoder: Encoder structure: 256 → 128 → 64 → 32, using ReLU activation; Decoder structure: 32 → 64 → 128 → 256, output and input reconstruction; Training objective: minimize reconstruction loss (MSE) and feature sparsity (L1 regularization).
[0083] Business semantic integration: Manually define semantic labels (such as "data leakage risk" and "normal operation and maintenance") and fine-tune the dimensionality reduction model through supervised learning; For example, “high risk” samples are forced to have a distance less than 0.1 in the reduced space.
[0084] For example, after a multi-dimensional joint feature vector is compressed by an autoencoder, the high-order semantic features are 32-dimensional 0.8, −0.3, ..., 0.6, which can be interpreted as "frequent cross-table access and association with suspicious users."
[0085] Specifically, based on the high-order semantic feature representation, combined with the data content sensitivity label and business scenario context information, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory, including: Based on high-level semantic feature representation, the semantic vector of the data content sensitivity label and the natural language vector of the business scenario context description are extracted; High-level semantic feature representation is usually a 512-dimensional feature vector extracted through a spatiotemporal graph neural network, where each dimension corresponds to a specific semantic attribute of the data flow (such as the operator's permission level, data object category, etc.). To extract the semantic vector of the data content sensitivity label, a sensitivity classification model needs to be constructed: Label system definition: Data sensitivity is divided into 5 levels (L1~L5), for example: L1: Public data (such as product instructions); L3: internal data (e.g., meeting minutes); L5: Confidential data (such as customer privacy).
[0086] Semantic vector generation: A variant of the pre-trained language model BERT (such as RoBERTa) is used to encode the label text. For example, the label "L5-Customer Privacy" is encoded into a 768-dimensional vector using RoBERTa, and then the dimension is reduced to 128 dimensions using a fully connected layer.
[0087] Business scenario context processing: The business scenario description text (e.g., "The financial approval process involves departments A and B") is processed through a Bi-LSTM (bidirectional long short-term memory) network to generate a 256-dimensional natural language vector. Specific steps: Word embedding: using GloVe word vectors (300 dimensions); Sequence modeling: Bi-LSTM hidden layer 128-dimensional, output 256-dimensional context vector; Attention Focus: Key entities are strengthened through the self-attention mechanism (e.g., “Financial Approval” has a weight of 0.8, and “Department A” has a weight of 0.6).
[0088] For example, a data object labeled "L4-Contract Amount" may have its semantic vector mapped to high-dimensional features such as "value-sensitive" and "restricted permissions"; the context vector of the business scenario "cross-departmental data sharing" includes semantics such as "collaborative subject" and "data flow direction."
[0089] Based on the semantic vector and natural language vector, modality alignment processing is performed through a cross-modal encoder to generate cross-modal alignment features; The core of the cross-modal encoder is to align the semantic spaces of different modalities through contrastive learning. The encoder architecture adopts a dual-tower structure: Modal Coding Tower: Semantic Vector Tower: 3-layer fully connected network (input 128D → 256D → 256D), ReLU activation; Natural Language Tower: 3-layer fully connected network (input 256D → 512D → 256D), GeLU activation.
[0090] Modal alignment strategy: Hard alignment: Apply cosine similarity loss (margin=0.2) to matching label-scenario pairs (e.g., “L5-Customer Privacy” and “Customer Data Analysis Scenario”); Soft alignment: Dynamically adjust feature weights via a cross-modal attention matrix (128×128).
[0091] Training process: Positive sample construction: Extracting true pairings of labels and scenarios from historical enterprise logs (e.g., "L3-Employee Attendance" corresponds to "HR Monthly Statistics Scenario"); Negative sample sampling: Randomly replace labels or scenarios to generate negative samples (for example, "L3-Employee Attendance" is incorrectly associated with "Supply Chain Scheduling Scenario"); Loss function: InfoNCE loss (temperature coefficient τ = 0.05), batch size 256, optimizer AdamW (learning rate 3e-4).
[0092] Example of aligned features: After a certain "L5-R&D code" label is aligned with the "version iteration test scenario" context, cross-modal features are significantly enhanced in dimensions such as "authority control" and "operation audit."
[0093] Based on the cross-modal alignment features, the multi-head attention mechanism is used to calculate the semantic association weights between different modalities and generate an attention weight matrix; Multi-Head Attention (MHA) is used to capture complex correlations between cross-modal features. Specific parameter configuration: Number of heads and dimensions: 8 heads, each with 32 dimensions (input 256 dimensions → 8 heads × 32 dimensions); Query (Q), key (K), value (V) generation: Semantic vector side: Q = W_q·H_semantic (weight matrix W_q∈R^{256×256}); Natural language side: K = W_k·H_text, V = W_v·H_text (W_k, W_v∈R^{256×256}); Attention calculation: Scaled dot product attention score: Score=Softmax(QK^T / √32); Output: Attention(Q,K,V)=Score·V.
[0094] Weight matrix generation: Each head generates a 32×32 attention weight matrix, which is then concatenated into a 128×128 global weight matrix through a fully connected layer. The weight matrix is sparsely processed (Top-k retention, k=50) to reduce noise interference.
[0095] For example, in the "data leakage risk analysis" scenario, attention head 1 may focus on the association between "sensitive label-L5" and "external transmission operation" (weight 0.9), while head 2 focuses on the relationship between "timestamp dense area" and "batch download behavior" (weight 0.7).
[0096] Based on the attention weight matrix, the feature fusion layer integrates cross-modal semantic information and outputs a preliminary global embedding vector; The feature fusion layer uses a gating mechanism to dynamically weight the contributions of different modalities: Gating weight calculation: Input: Attention weight matrix (128×128) flattened into a 16384-dimensional vector; Gating network: 3-layer fully connected (16384→4096→512→2), outputting semantic gating weight g1 and text gating weight g2 (g1+g2=1); Weighted fusion: Semantic side features: H_semantic' = g1·H_semantic; Text side features: H_text' = g2·H_text; Fusion result: H_fused = H_semantic'⊕ H_text' (⊕ indicates splicing), resulting in a 512-dimensional vector; Dimensionality reduction output: PCA (principal component analysis) is used to compress the 512-dimensional vector to 256 dimensions, retaining 95% of the variance.
[0097] For example, in the "cross-border transmission of highly sensitive data" scenario, the gating weight may be biased towards the semantic side (g1=0.7, g2=0.3), emphasizing the impact of sensitivity labels; while in the "internal collaborative review" scenario, the text side has a higher weight (g1=0.4, g2=0.6).
[0098] According to the preliminary global embedding vector, a generative adversarial network is used to optimize the embedding space distribution and generate the global embedding vector of the final data trajectory.
[0099] Generative Adversarial Networks (GANs) are used to improve the discriminability and robustness of the embedding space: Generator (G) architecture: Input: 256-dimensional preliminary embedding vector + 128-dimensional random noise; Structure: 4-layer fully connected (384→512→512→256), LayerNorm normalization, LeakyReLU (α=0.2) activation; Output: 256-dimensional generated embedding vector.
[0100] Discriminator (D) architecture: Input: 256-dimensional vector (real / generated); Structure: 3-layer fully connected (256→128→64→1), Dropout (p=0.3), Sigmoid output; Adversarial training strategy: Generator objective: minimize the Wasserstein distance between the generated vector and the true distribution; Discriminator goal: maximize the difference between the real sample score and the generated sample score; Optimizer: RMSProp (learning rate G=5e-5, D=1e-4), gradient penalty coefficient λ=10.
[0101] Embedding space optimization: Pattern coverage enhancement: Feature matching loss is used to force the generated vector to align with the true vector in terms of statistical distribution (such as mean and variance); Anomaly suppression: resample the generated samples that the discriminator judges as low confidence (<0.2).
[0102] Example: After 50 rounds of training, the generated embedding vectors show clear clustering in t-SNE visualization (for example, the distance between the "compliant operation" cluster and the "abnormal access" cluster is significant), and the accuracy of adversarial example detection increases to 98%.
[0103] Specifically, based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential anomaly diffusion paths are predicted through nonlinear dynamic equations to generate a risk-aware trajectory portrait, including: Based on the global embedding vector, the nodes and connecting edges of the dynamic brain network are constructed, where the nodes represent the key entities of data flow and the edge weights represent the diffusion probability; The construction of the dynamic brain network is based on the global embedding vector (each entity corresponds to a 512-dimensional vector) and is achieved through the following steps: Key Entity Extraction: Filter out frequently accessed entities (such as database tables, API interfaces, and user accounts) from the global embedding vector. For example, the key entities of a financial enterprise include "Customer Information Table" (embedding vector V1), "Transaction Record API" (V2), and "Risk Control Administrator Account" (V3).
[0104] Node initialization: Each entity is mapped to a network node, and the node attributes include: Type label (data table, API, user, etc.); Frequency of visits (e.g., the average daily visit count for the “Customer Information Form” is 12,000); Sensitivity score (0-1, based on data content labels, such as customer information sensitivity 0.95).
[0105] Edge weight calculation: Based on the statistical diffusion probability of historical data flow paths. For example, if the "risk control administrator account" has a 60% probability of triggering access to the "customer information table" in the past 30 days, the edge weight between the two is set to 0.6. The weight calculation formula is: Use Python's NetworkX library to build a graph structure with approximately 500-1000 nodes and 3000-5000 edges.
[0106] Example: In an e-commerce scenario, the edge weight between the nodes "Order Database" and "Payment Gateway API" is 0.75, indicating that 75% of order data flows will trigger payment operations.
[0107] Based on the dynamic brain network, nonlinear dynamic equations are used to simulate the path evolution process of data flow and calculate the chaotic diffusion coefficient between nodes; The nonlinear dynamics model uses the modified Lorenz equations to simulate the chaotic characteristics of data flow: Parameterization of the dynamic equations: Classical Lorentz parameters: σ = 10 (Prantl number), ρ = 28 (Rayleigh number), β = 8 / 3; Introduce a data diffusion sensitivity factor α (0.1-0.5) and adjust the equation to reflect differences in business scenarios: Chaotic Map Iteration: Use the fourth-order Runge-Kutta method (RK4) to solve the differential equation with a time step of Δt=0.01 and 1000 iterations. For example, the initial state is x=0.1, y=0, z=0, and the node state trajectory (x(t), y(t), z(t)) is generated after iteration.
[0108] Chaotic diffusion coefficient calculation: Maximum Lyapunov Exponent (LLE): quantifies the sensitivity of the system to initial conditions. It is calculated as the logarithmic mean of the trajectory separation rate, with a threshold of 0.5. If LLE > 0.5, it is considered chaotic diffusion; Phase space reconstruction: The univariate sequence is expanded into a three-dimensional phase space by time delay embedding (delay τ = 10, embedding dimension m = 3), and the trajectory correlation is calculated (e.g., a Pearson correlation coefficient > 0.7 indicates a strong correlation).
[0109] For example, in a supply chain system, the LLE between the "inventory database" and the "logistics scheduling API" is 0.63, indicating that there is a significant chaotic diffusion risk between the two.
[0110] According to the chaotic diffusion coefficient, the path prediction algorithm is used to identify potential abnormal diffusion paths and generate a set of high-risk paths; The path prediction algorithm combines the Hidden Markov Model (HMM) and Monte Carlo Tree Search (MCTS): Hidden Markov Model (HMM) construction: State collection: network nodes (e.g. S1 = customer table, S2 = payment API); Observation sequence: historical data flow path (e.g. S1→S2→S3); Transition probability matrix A: Dynamically adjusted based on the chaotic diffusion coefficient. For example, if the diffusion coefficient from S1 to S2 is 0.6, then A[S1][S2]=0.6; Emission probability matrix B: Normalized node sensitivity scores (e.g., S1 sensitivity 0.9 → B[S1] = 0.9 / Σsensitivity).
[0111] Multi-step transition prediction: Viterbi algorithm: calculates the most likely path (e.g., the next 3 steps: S1 → S2 → S4 → S5); Time decay factor: λ=0.95, the confidence of long-term transfer decays exponentially (confidence at step k = λ^k).
[0112] Monte Carlo Tree Search (MCTS): Number of simulations: 1000 times, each simulation depth is 10 steps; UCB1 formula: The profit weight of selecting a child node = Q(s,a) + c√(lnN(s) / n(s,a)), and the exploration coefficient c=1.5; Risk threshold: Pathway comprehensive risk score > 0.7 (score = diffusion coefficient × sensitivity × path length attenuation factor).
[0113] Example: The predicted comprehensive risk score of the path "User Account A → Order Database → Log Server" is 0.82 (diffusion coefficient 0.6 × sensitivity 0.9 × length decay 0.8 = 0.432, normalized to 0.82), and is marked as a high-risk path.
[0114] Based on a set of high-risk paths, the neural radiation field technology is combined to render a three-dimensional trajectory heat map and output a risk-aware trajectory portrait.
[0115] Neural Radiance Field (NeRF) technology maps path data into a 3D heat map: Data voxelization: Spatial grid division: Map network nodes to a 512×512×512 voxel space, where each voxel corresponds to a physical area (such as a computer room or cloud server area); Pathway strength encoding: the voxel value of a high-risk path = diffusion coefficient × 10 (e.g., coefficient 0.6 → voxel value 6).
[0116] NeRF model training: Input: Multi-view 2D risk heatmap (projected from X / Y / Z axes); Network structure: 5-layer MLP (512 neurons per layer), activation function ReLU; Loss function: mean squared error (MSE) + regularization term (weight 0.01).
[0117] Heatmap rendering: Color mapping: low risk (blue, RGB=0,0,255) → high risk (red, RGB=255,0,0); Transparency control: higher voxel values result in more opacity (α = voxel value / 10).
[0118] Interactive visualization: WebGL framework: Three.js library enables real-time rendering on the browser side; Click to query: Users can click on a hot zone to display path details (e.g., "Zone X: Risk score 0.82, affecting database A and API B").
[0119] Example: The trajectory profile of a banking system shows that the northwest area of the data center (voxel coordinates x=300-350, y=200-250, z=100) appears dark red, corresponding to the high-risk path cluster of "customer credit data → risk control model → external API".
[0120] Specifically, based on the dynamic brain network, nonlinear dynamic equations are used to simulate the path evolution process of data flow and calculate the chaotic diffusion coefficient between nodes, including: According to the node state vectors and initial weights of the connection edges of the dynamic brain network, the coupling relationship between nodes is established using the Lorentz dynamics equation to generate a nonlinear interaction matrix. The node state vector of a dynamic brain network contains multidimensional characteristics of key entities in data flow. For example, a data node might include the following parameters: access frequency (times / second), sensitivity level (1-5), and the timestamp of the most recent operation (in milliseconds). The initial weight of the connecting edge is based on the statistical probability of historical data flow paths. For example, the initial weight of 0.7 for node A to node B indicates that 70% of data flows have flowed through this path in the past.
[0121] The Lorentz dynamics equation is used here to model the nonlinear coupling relationship between nodes. The traditional Lorentz equation is used to describe the three-dimensional chaotic system of atmospheric convection. Its parameters σ (Prantl number), ρ (Rayleigh number), and β (geometric parameter) are redefined as data flow characteristics: σ=10: represents the diffusion rate of data flow. A larger value means faster data transmission between different nodes. ρ=28: reflects the sensitivity of the node state. A higher value means that the node state is more susceptible to small disturbances. β=8 / 3: The dissipative characteristic of the control system is used to balance the stability and chaos of data flow.
[0122] The state evolution of each node is calculated as follows: State initialization: The initial state vector of node i is (x_i, y_i, z_i). For example, the initial values of a node are (1.2, 0.8, 5.3), which correspond to the data inflow rate, processing delay, and queue length, respectively. Coupling relationship establishment: If nodes i and j are connected, the coupling strength is determined by the edge weight w_ij. For example, when w_ij = 0.7, the x component of node i is affected by the x component of node j, and the weight coefficient is 0.7 × σ; Iterative calculation: A fourth-order Runge-Kutta algorithm (with a step size of Δt = 0.01 seconds) is used to iteratively update the states of all nodes. For example, after 100 iterations, the state of node i changes to (2.5, -1.1, 6.8), indicating a surge in the data inflow rate and negative processing delays (possibly indicating queue congestion).
[0123] The resulting nonlinear interaction matrix has dimensions N × N (N is the number of nodes), with the element M_ij representing the coupling strength from node i to node j. For example, in a matrix where M_23 = 0.9, this indicates that node 2 has a very strong interaction with node 3, potentially forming a high-risk diffusion channel.
[0124] Based on the nonlinear interaction matrix, the diffusion sensitivity factor of data flow is introduced, and the node state evolution trajectory is iteratively calculated through the chaotic mapping algorithm to obtain the spatial distribution of chaotic attractors. The Diffusion Sensitivity Factor (DSF) is used to quantify the node's tendency to spread abnormal data. Its value range is 0 to 1 and is determined by the data sensitivity label and the operation behavior history. For example: DSF=0.3: ordinary file transfer node, low transmission risk; DSF=0.8: Nodes containing user privacy data have a high risk of dissemination.
[0125] The chaos mapping algorithm adopts the improved Logistic mapping, and its equation is: Among them: r=4.0: ensures that the system is in a chaotic state; α=0.2: controls the influence weight of the diffusion sensitivity factor; x_n∈[0,1]: normalized node state parameter.
[0126] Iterative calculation process: Parameter injection: Map the coupling strength of the nonlinear interaction matrix to the initial x value. For example, M_ij=0.7 corresponds to x_0=0.7; Chaotic evolution: Each node is independently iterated 1000 times to generate a state sequence. For example, a node sequence of [0.7, 0.84, 0.5376, ...] shows typical chaotic fluctuations. Attractor extraction: Identify attractor structures through phase space projection (e.g., constructing a two-dimensional graph of x_n and x_{n+1}). For example, attractors for high-risk nodes exhibit divergent trajectories, while attractors for low-risk nodes tend to be circular and closed.
[0127] The spatial distribution of chaotic attractors is stored as a three-dimensional point cloud, with each point representing the state of a node at a specific moment. For example, in one distribution, high-risk nodes cluster near the coordinates (0.8, 0.6, 1.2), forming an "abnormal cluster" that provides a basis for subsequent path prediction.
[0128] According to the spatial distribution of chaotic attractors, phase space reconstruction technology is used to extract the trajectory correlation characteristics between nodes and generate trajectory similarity measurement values; Phase space reconstruction technology uses the time delay embedding method (Takens theorem) to expand the single variable time series into a high-dimensional phase space trajectory. Specific parameter settings: Embedding dimension m = 3: automatically selected based on the false nearest neighbor method (FNN); Time delay τ=10: determined based on the mutual information minimization criterion.
[0129] Reconstruction steps: Trajectory generation: For each node's chaotic sequence x_n, generate a phase space vector X_k = (x_k, x_{k+τ}, x_{k+2τ}). For example, after reconstructing a node sequence, we obtain the trajectory point set {(0.7, 0.65, 0.72), (0.65, 0.72, 0.68), ...}. Correlation calculation: The Dynamic Time Warping (DTW) algorithm is used to quantify the trajectory similarity between nodes i and j. For example, the DTW distance between nodes A and B is 1.2, while that between nodes A and C is 3.5, indicating that the behavior patterns of A and B are more similar. Metric normalization: Map the DTW distance to a similarity metric S_ij=1 / (1+DTW) between 0 and 1. For example, DTW=1.2 corresponds to S_ij=0.45.
[0130] The trajectory similarity measurement matrix has an N×N dimension and is used to identify potential abnormal propagation paths. For example, S_23=0.8 indicates that the trajectories of nodes 2 and 3 are highly similar and may form a diffusion chain.
[0131] Based on the trajectory similarity measure, the uncertainty of the diffusion process is quantified by the maximum Lyapunov exponent algorithm, and the chaotic diffusion coefficient between nodes is output.
[0132] The Maximum Lyapunov Exponent (MLE) is used to characterize the system's sensitivity to initial conditions (the degree of chaos). The calculation process is as follows: Neighboring trajectory selection: Find neighboring trajectories with an initial distance ε=1e-4 for each node trajectory in phase space. For example, the initial distance between the reference trajectory and the neighboring trajectory of node A is 0.0001; Divergence tracking: Iteratively track the distance Δ(t) between two trajectories and record its exponential growth rate. For example, after 100 iterations, Δ(t) = 0.012, and the growth rate is calculated as λ = ln(Δ(t) / ε) / t≈0.45; Exponential convergence: Repeat the above process 100 times and take the average to obtain the MLE value λ_i for node i. For example, for a high-risk node, λ_i = 0.82 (> 0 indicates chaos), and for a low-risk node, λ_i = -0.03 (stable).
[0133] The chaotic diffusion coefficient CDC_ij is calculated by the following formula: S_ij: trajectory similarity measure; λ_i,λ_j: MLE values of nodes i and j.
[0134] Output example: The CDC between nodes 2 and 3 is 0.8×(0.82+0.76) / 2=0.632, indicating a high risk of diffusion. The CDC between nodes 5 and 7 is 0.3×(0.12+0.05) / 2=0.0255, indicating a negligible risk.
[0135] This coefficient matrix will be directly input into the path prediction module to guide the identification of high-risk paths and the generation of blocking strategies.
[0136] Specifically, based on the chaotic diffusion coefficient, a path prediction algorithm is used to identify potential abnormal diffusion paths and generate a set of high-risk paths, including: According to the chaotic diffusion coefficient and the connection edge weight of the dynamic brain network, a path diffusion probability matrix is constructed, in which the probability value is nonlinearly positively correlated with the chaotic diffusion coefficient; The edge weights of the dynamic brain network represent the initial strength of the association between data entities (e.g., the frequency of user A accessing database B), while the chaotic diffusion coefficient quantifies the uncertainty of the spread of abnormal patterns along the edges. To construct the path diffusion probability matrix, the two need to be combined: Nonlinear mapping function design: A piecewise exponential function is used to correlate the chaotic diffusion coefficient and probability value. For example, when the chaotic diffusion coefficient is ≤0.3, the probability = coefficient × 0.8; when the coefficient is >0.3, the probability = 0.24 + (coefficient - 0.3) × 2.5, ensuring that high diffusion coefficient paths significantly increase the probability.
[0137] Weight normalization: Softmax normalizes the weights of all outgoing edges from the same source node so that the total probability sums to 1. For example, if node X has three outgoing edges with original weights [0.6, 0.3, 0.1], the probability distribution after normalization is [0.55, 0.27, 0.18].
[0138] Dynamic matrix update: The probability matrix is recalculated every 5 minutes, and weights are adjusted based on real-time data flow events (such as sudden high-frequency access). For example, if node Y is accessed 10 times in the last minute, its outbound edge weight is temporarily increased by 30%.
[0139] For example, in a dynamic brain network of a financial system, the chaotic diffusion coefficient from node P (payment system) to node Q (user database) is 0.45, and the initial edge weight is 0.7. After nonlinear mapping, the diffusion probability = 0.24 + (0.45-0.3) × 2.5 = 0.615. After softmax normalization (assuming the total weight of similar edges is 2.1), the final probability = 0.615 / 2.1 ≈ 0.293.
[0140] Based on the path diffusion probability matrix, the hidden Markov model is used to predict the multi-step transfer path, and the path transfer confidence is calculated in combination with the time decay factor. The Hidden Markov Model (HMM) considers the data flow path as a sequence of states, where the states correspond to nodes in the dynamic brain network and the observed values are the characteristics of the operation behavior (such as access type and data volume). State transition probability setting: Directly use the path diffusion probability matrix as the HMM transition probability matrix. For example, the probability of node A→B is 0.3, and the probability of B→C is 0.5.
[0141] Observation probability modeling: Based on historical data, the typical operation characteristics of each node are statistically analyzed. For example, the observation probability distribution of node D (customer information database) is: query operation 60%, modification operation 30%, and deletion operation 10%.
[0142] Time decay factor introduction: Define the decay coefficient λ = 0.95, indicating that the confidence decays by 5% after each transfer step. For example, the confidence of the path A→B→C = 0.3(A→B) × 0.5(B→C) × λ² = 0.3 × 0.5 × 0.9025 ≈ 0.135.
[0143] Multi-step path prediction: Use the forward-backward algorithm to calculate the joint probability of an N-step (e.g., N = 5) transition path. For example, the confidence level of the predicted path A→B→C→D = 0.3×0.5×0.4×λ³ = 0.3×0.5×0.4×0.857≈0.051.
[0144] Confidence dynamic correction: Real-time behavior matching: If unusual operations are observed at node C (such as batch export at 3 a.m.), the corresponding path confidence is increased by 20%; Risk event superposition: When external attack features (such as SQL injection attempts) are detected, the confidence of the relevant path is increased by an additional 0.15.
[0145] Based on the path transfer confidence, the candidate abnormal diffusion paths are generated through the Monte Carlo tree search algorithm, and high-risk paths are screened according to the risk threshold; Monte Carlo Tree Search (MCTS) explores high-risk areas by simulating a large number of possible paths. The specific steps are: Tree structure initialization: The root node is the starting point of the currently detected anomaly (such as a hacked user account), and each layer of child nodes corresponds to the data entity that may be accessed by the next hop.
[0146] Node expansion strategy: Exploration and Exploitation Balance: The UCB (Upper Confidence Bound) formula is used to select child nodes, with an exploration coefficient C of 1.4. UCB value = node confidence + C × √(ln parent node visit count / child node visit count); Depth limit: The maximum search depth is set to 10 hops to avoid unlimited expansion.
[0147] Simulation and backtracking: Random simulation: Random paths are expanded for incompletely expanded nodes until a termination condition is reached (e.g., accessing a sensitive database); Confidence backtracking: Simulation results (such as path risk scores) are updated backwards to all nodes along the path. For example, if a path A→B→C is marked as high risk in the simulation, the number of visits to node B will be increased by 1, and the total risk score will be increased by 0.8.
[0148] High-risk path screening: Set a risk threshold of θ = 0.7 (confidence ≥ 0.7 indicates high risk) and extract all paths from leaf nodes that exceed the threshold. For example, the path X→Y→Z has a confidence of 0.82 and is included in the candidate set.
[0149] Dynamic pruning optimization: Local pruning: If the confidence of all child nodes of a node is less than 0.4, then stop expanding; Path merging: Merge paths with the same prefix (such as A→B→C and A→B→D) to reduce redundant calculations.
[0150] Based on the screened high-risk paths, the path aggregation algorithm is used to merge overlapping path segments and optimize path weights to output the final high-risk path set.
[0151] The path aggregation algorithm merges similar paths into more representative risk patterns by identifying common nodes and edges: Overlapping Path Detection: Prefix tree construction: insert all paths into the prefix tree (Trie tree) in node sequence and count the frequency of common prefixes. For example, if paths A→B→C and A→B→D share the prefix A→B, the frequency count is +2; Community detection: The Louvain algorithm is used to identify densely connected subgraphs in paths. For example, the path set {A→B→C, B→C→D, C→D→E} is clustered into community 1 (around node C).
[0152] Path merging rules: Weight superposition: The weight of the merged path = the sum of the weights of each path × the overlap coefficient (number of overlapping nodes / total number of nodes). For example, merging A→B→C (weight 0.8) and A→B→D (weight 0.6), with an overlap coefficient of 2 / 3, results in a combined weight of (0.8+0.6)×(2 / 3)=0.93. Key node extraction: retain high-frequency nodes (such as nodes that appear more than 5 times) as path representatives.
[0153] Weight optimization: PageRank adjustment: Redistribute path weights based on the importance of the node in the entire network. For example, if node B has a PageRank value of 0.15, the weight of its path will be increased by 10%; Time-sensitive decay: The weight of old paths (such as those older than 24 hours) decays by 2% every hour.
[0154] Final output format: The high-risk path collection is stored in the form of a JSON array. Each path contains a node sequence, a combined weight, and a last active timestamp.
[0155] It can be seen that a structured data flow graph is generated based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources; based on the structured data flow graph, the spatiotemporal graph neural network is used to output high-order semantic feature representations of data flow; based on the high-order semantic feature representation, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory; based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate risk-aware trajectory portraits, thereby improving the performance of trajectory portraits and providing an advanced technical foundation for enterprise security management, operation optimization and decision support.
[0156] Another embodiment of the present invention provides a trajectory profile generation system based on artificial intelligence, see Figure 3, the system may include: Extraction module 301 is used to extract multi-dimensional data flow features based on dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources through adaptive spatiotemporal crawler technology to generate a structured data flow graph; A fusion module 302 is configured to output a high-level semantic feature representation of the data flow by fusing the spatiotemporal dependency of the data flow and the correlation between the behavior of the operator based on the structured data flow graph and using a spatiotemporal graph neural network; A generation module 303 is configured to generate a global embedding vector of the data trajectory by fusing multi-source semantic features through a cross-modal attention mechanism based on the high-order semantic feature representation, combined with the data content sensitivity label and business scenario context information; The simulation module 304 is used to simulate the path evolution process of data flow based on the global embedding vector using a dynamic brain network evolution model, predict potential abnormal diffusion paths through nonlinear dynamic equations, and generate a risk-aware trajectory portrait.
[0157] It can be seen that a structured data flow graph is generated based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources; based on the structured data flow graph, the spatiotemporal graph neural network is used to output high-order semantic feature representations of data flow; based on the high-order semantic feature representation, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory; based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate risk-aware trajectory portraits, thereby improving the performance of trajectory portraits and providing an advanced technical foundation for enterprise security management, operation optimization and decision support.
[0158] An embodiment of the present invention further provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps of any one of the above method embodiments when running.
[0159] Specifically, in this embodiment, the above-mentioned storage medium may be configured to store a computer program for executing the following steps: S201: Based on the dynamic access logs, data operation behavior records, and permission change events of the enterprise's heterogeneous data sources, the adaptive spatiotemporal crawler technology is used to extract multi-dimensional data flow characteristics and generate a structured data flow graph; S202: Based on the structured data flow graph, using a spatiotemporal graph neural network, by fusing the spatiotemporal dependency of the data flow and the correlation of the operator's behavior, output a high-order semantic feature representation of the data flow; S203, based on the high-order semantic feature representation, combined with the data content sensitivity label and business scenario context information, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory; S204: Based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate a risk-aware trajectory portrait.
[0160] It can be seen that a structured data flow graph is generated based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources; based on the structured data flow graph, the spatiotemporal graph neural network is used to output high-order semantic feature representations of data flow; based on the high-order semantic feature representation, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory; based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate risk-aware trajectory portraits, thereby improving the performance of trajectory portraits and providing an advanced technical foundation for enterprise security management, operation optimization and decision support.
[0161] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.
[0162] Specifically, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0163] Specifically, in this embodiment, the processor may be configured to execute the following steps through a computer program: S201: Based on the dynamic access logs, data operation behavior records, and permission change events of the enterprise's heterogeneous data sources, the adaptive spatiotemporal crawler technology is used to extract multi-dimensional data flow characteristics and generate a structured data flow graph; S202: Based on the structured data flow graph, using a spatiotemporal graph neural network, by fusing the spatiotemporal dependency of the data flow and the correlation of the operator's behavior, output a high-order semantic feature representation of the data flow; S203, based on the high-order semantic feature representation, combined with the data content sensitivity label and business scenario context information, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory; S204: Based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate a risk-aware trajectory portrait.
[0164] It can be seen that a structured data flow graph is generated based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources; based on the structured data flow graph, the spatiotemporal graph neural network is used to output high-order semantic feature representations of data flow; based on the high-order semantic feature representation, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector of the data trajectory; based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate risk-aware trajectory portraits, thereby improving the performance of trajectory portraits and providing an advanced technical foundation for enterprise security management, operation optimization and decision support.
[0165] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings. The above is only a preferred embodiment of the present invention, but the scope of implementation of the present invention is not limited to what is shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which do not exceed the spirit covered by the description and drawings, should be within the scope of protection of the present invention.
Claims
1. A trajectory portrait generation method based on artificial intelligence, characterized in that: The method comprises: Based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources, adaptive spatiotemporal crawler technology is used to extract multi-dimensional data flow characteristics and generate a structured data flow map; Based on the structured data flow graph, a spatiotemporal graph neural network is used to output a high-order semantic feature representation of the data flow by fusing the spatiotemporal dependency of the data flow and the correlation of the operator's behavior; Based on the high-level semantic feature representation, combined with data content sensitivity labels and business scenario context information, multi-source semantic features are fused through a cross-modal attention mechanism to generate a global embedding vector for the data trajectory; Based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and the potential anomaly diffusion path is predicted through nonlinear dynamic equations to generate a risk-aware trajectory portrait.
2. The method according to claim 1, characterized in that The method extracts multi-dimensional data flow features based on the dynamic access logs, data operation behavior records, and permission change events of the enterprise's heterogeneous data sources through adaptive spatiotemporal crawler technology to generate a structured data flow graph, including: Perform multi-source heterogeneous data cleaning based on enterprise database access logs, API call records, and terminal operation logs to remove invalid timestamps and duplicate operation records, and obtain standardized data flow event sequences. Based on the standardized data flow event sequence, the chaos theory coding method is used to perform spatiotemporal coding on the operation subject, data object and permission change, and the spatiotemporal correlation feature matrix is obtained. Based on the spatiotemporal correlation feature matrix, the data flow path is dynamically tracked through adaptive spatiotemporal crawler technology to construct a dynamic correlation map including time dimension, space dimension and operation dimension; Based on the dynamic association graph, the graph structure self-correction algorithm is used to correct the abnormal connection edge weights and generate the final structured data flow graph.
3. The method according to claim 2, characterized in that Based on the structured data flow graph, the spatiotemporal graph neural network is used to fuse the spatiotemporal dependency of data flow and the correlation of operator behavior to output a high-order semantic feature representation of data flow, including: Based on the structured data flow graph, the node features and edge weights of the spatiotemporal graph neural network are initialized. The node features include the identity of the operation subject, the data object type, and the timestamp code. Based on the initialization node features, the spatiotemporal dependencies of data flow are captured through the spatiotemporal convolution layer to obtain the spatiotemporal fusion feature tensor; Based on the spatiotemporal fusion feature tensor, the behavioral association modeling algorithm is used to calculate the behavioral similarity matrix between the operating subjects and generate the subject behavior association features; Based on the subject behavior correlation features, the spatiotemporal features and behavioral features are fused through the dynamic aggregation layer to output a multi-dimensional joint feature vector; According to the multi-dimensional joint feature vector, the feature dimensionality reduction algorithm is used to remove redundant information and generate a high-order semantic feature representation of the data flow.
4. The method according to claim 3, characterized in that The method generates a global embedding vector of the data trajectory by fusing multi-source semantic features through a cross-modal attention mechanism based on the high-order semantic feature representation, combined with data content sensitivity labels and business scenario context information, including: Based on high-level semantic feature representation, the semantic vector of the data content sensitivity label and the natural language vector of the business scenario context description are extracted; Based on the semantic vector and natural language vector, modality alignment processing is performed through a cross-modal encoder to generate cross-modal alignment features; Based on the cross-modal alignment features, the multi-head attention mechanism is used to calculate the semantic association weights between different modalities and generate an attention weight matrix; Based on the attention weight matrix, the feature fusion layer integrates cross-modal semantic information and outputs a preliminary global embedding vector; According to the preliminary global embedding vector, a generative adversarial network is used to optimize the embedding space distribution and generate the global embedding vector of the final data trajectory.
5. The method according to claim 4, characterized in that Based on the global embedding vector, a dynamic brain network evolution model is used to simulate the path evolution process of data flow, and potential abnormal diffusion paths are predicted through nonlinear dynamic equations to generate a risk-aware trajectory portrait, including: Based on the global embedding vector, the nodes and connecting edges of the dynamic brain network are constructed, where the nodes represent the key entities of data flow and the edge weights represent the diffusion probability; Based on the dynamic brain network, nonlinear dynamic equations are used to simulate the path evolution process of data flow and calculate the chaotic diffusion coefficient between nodes; According to the chaotic diffusion coefficient, the path prediction algorithm is used to identify potential abnormal diffusion paths and generate a set of high-risk paths; Based on a set of high-risk paths, the neural radiation field technology is combined to render a three-dimensional trajectory heat map and output a risk-aware trajectory portrait.
6. The method according to claim 5, characterized in that The method is based on a dynamic brain network and uses nonlinear dynamic equations to simulate the path evolution process of data flow and calculate the chaotic diffusion coefficient between nodes, including: According to the node state vectors and initial weights of the connection edges of the dynamic brain network, the coupling relationship between nodes is established using the Lorentz dynamics equation to generate a nonlinear interaction matrix. Based on the nonlinear interaction matrix, the diffusion sensitivity factor of data flow is introduced, and the node state evolution trajectory is iteratively calculated through the chaotic mapping algorithm to obtain the spatial distribution of chaotic attractors. According to the spatial distribution of chaotic attractors, phase space reconstruction technology is used to extract the trajectory correlation characteristics between nodes and generate trajectory similarity measurement values; Based on the trajectory similarity measure, the uncertainty of the diffusion process is quantified by the maximum Lyapunov exponent algorithm, and the chaotic diffusion coefficient between nodes is output.
7. The method according to claim 6, characterized in that The method of identifying potential abnormal diffusion paths through a path prediction algorithm based on the chaotic diffusion coefficient and generating a high-risk path set includes: According to the chaotic diffusion coefficient and the connection edge weight of the dynamic brain network, a path diffusion probability matrix is constructed, in which the probability value is nonlinearly positively correlated with the chaotic diffusion coefficient; Based on the path diffusion probability matrix, the hidden Markov model is used to predict the multi-step transfer path, and the path transfer confidence is calculated in combination with the time decay factor. Based on the path transfer confidence, the candidate abnormal diffusion paths are generated through the Monte Carlo tree search algorithm, and high-risk paths are screened according to the risk threshold; Based on the screened high-risk paths, the path aggregation algorithm is used to merge overlapping path segments and optimize path weights to output the final high-risk path set.
8. An artificial intelligence-based trajectory portrait generation system, characterized in that: The system comprises: The extraction module is used to extract multi-dimensional data flow characteristics based on the dynamic access logs, data operation behavior records and permission change events of the enterprise's heterogeneous data sources through adaptive spatiotemporal crawler technology to generate a structured data flow map; A fusion module is used to output a high-order semantic feature representation of the data flow by fusing the spatiotemporal dependency of the data flow and the correlation of the operation subject behavior based on the structured data flow graph and using a spatiotemporal graph neural network; A generation module is used to generate a global embedding vector of the data trajectory based on the high-order semantic feature representation, combined with the data content sensitivity label and business scenario context information, and fusing multi-source semantic features through a cross-modal attention mechanism; The simulation module is used to simulate the path evolution process of data flow based on the global embedding vector using a dynamic brain network evolution model, predict potential abnormal diffusion paths through nonlinear dynamic equations, and generate a risk-aware trajectory portrait.
9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when run.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.
Citation Information
Cited By
Optical fiber sensing voiceprint feature analysis model construction method based on composite neural network
CN120636468A
Internet of Things equipment binding method and system based on AI behavior data
CN120785758A
Police service studying and judging method based on behavior feature recognition
CN120832603A
BIM-based cable bridge routing autonomous optimization visualization system
CN120976440A
Insurance user portrait generation method and device based on PageRank and mutual information
CN121190112A