Big data information network security adaptive protection method based on trusted data space
By constructing a data flow potential energy model and baseline flow map in a trusted data space, dynamically generating decoy data, monitoring and reconstructing the data topology in real time, and performing unsupervised clustering and intent counterbalancing, the problem of delayed response and static strategies to new attacks in existing technologies is solved, achieving efficient and proactive network security protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 张晨
- Filing Date
- 2026-02-25
- Publication Date
- 2026-06-02
AI Technical Summary
Existing network security protection systems are unable to respond promptly to new, low-frequency, slow attacks in trusted data spaces, lack awareness of data flow status, and static security strategies cannot adapt to dynamic data sharing scenarios. Attacker costs only increase linearly, resulting in lagging and inefficient security protection.
By modeling data flow potential energy and constructing baseline flow maps, we can dynamically generate logically related data, deploy highly interactive dynamic decoys, monitor data flow status in real time, perform unsupervised clustering analysis of attacker behavior, reconstruct data space topology for dynamic data slicing and isolation, generate customized data and scenarios for intent countermeasures, and optimize immune memory and potential baseline.
It achieves proactive defense against new low-frequency, slow attacks, improves the efficiency and depth of data security flow, ensures business continuity and security in dynamic sharing scenarios, strips away superficial behavior to focus on the real target, and seamlessly integrates dynamic isolation and sharing, resulting in an exponential increase in the operational costs for attackers.
Smart Images

Figure CN122137602A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network security technology, specifically referring to an adaptive protection method for big data information network security based on trusted data space. Background Technology
[0002] With the accelerated marketization of data elements, the trusted data space, as a core infrastructure for achieving data usability without visibility, controllability, and measurability, is becoming a key carrier for cross-institutional and cross-domain data collaboration. Its core characteristic lies in supporting fine-grained, high-frequency, and controlled dynamic data sharing and collaborative computing among multiple participants, while ensuring data sovereignty and privacy. However, existing network security protection systems have revealed systemic flaws when facing the new security requirements of the trusted data space, making it difficult to support the secure and efficient flow of data elements.
[0003] However, existing adaptive protection methods for big data information network security still have certain shortcomings. Existing technologies are slow to respond to new, low-frequency, and slow attacks, and cannot respond to unknown threats in a timely manner, resulting in a significant time window for security protection. At the same time, existing technologies only focus on the security of network pipelines and server nodes, lacking awareness of data content, context, and state transitions when data flows between multiple parties, creating data blind spots and failing to detect covert data theft. In addition, security strategies mostly adopt static authorization mechanisms, which cannot adapt to the dynamic and controlled data sharing scenarios required by trusted data spaces, leading to the sacrifice of security or impact on business efficiency in high-frequency data collaboration. Finally, existing technologies have failed to build a dynamic environment in which the cost to attackers increases exponentially with time, space, and attack depth; the cost to attackers only increases linearly, failing to effectively undermine the economic motivation for attacks, and turning security protection into a passive response rather than proactive countermeasure. These shortcomings together mean that existing technologies cannot meet the dual requirements of security and sharing in trusted data spaces, making the flow of data elements face both security risks and limitations imposed by inefficient protection mechanisms. Therefore, this paper proposes an adaptive protection method for big data information network security based on trusted data spaces. Summary of the Invention
[0004] The purpose of this invention is to provide an adaptive protection method for big data information network security based on trusted data space, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an adaptive protection method for big data information network security based on trusted data space, comprising the following steps:
[0006] S1. Based on the spatial attributes and historical flow of data, perform data flow potential modeling and baseline flow map construction through calculation and recording;
[0007] S2. Based on the data flow potential energy model, highly interactive dynamic decoy deployment is carried out by dynamically generating logically related data;
[0008] S3. Based on the baseline flow map, the potential energy field abnormal disturbance is sensed by real-time monitoring of the data flow status.
[0009] S4. Based on the perceived potential energy field disturbance and the triggered dynamic decoy, perform unsupervised clustering analysis of the attacker's behavior sequence to determine the attack intent.
[0010] S5. Based on the attack intent identified through clustering, perform dynamic data space slicing and isolation based on intent by reconstructing the data space topology;
[0011] S6. In the isolated data subspace, based on the attack intent, customized data and scenarios are generated and fed back to perform intent countermeasures and adaptive strategy generation.
[0012] S7. Based on complete attack lifecycle data and countermeasures results, immune memory and potential baseline evolution are carried out through knowledge accumulation and model optimization.
[0013] Preferably, in step S1, a data flow monitoring module is deployed in the trusted data space to record the flow path, traffic, frequency, and access sequence of all data entities. The collected flow data is cleaned to remove outliers and noise, and the data format is unified. Based on business characteristics, the historical flow data is divided into time windows. Based on the business scenario, weights are assigned to the four potential energy dimensions, and the current potential energy value is calculated for each data object.
[0014] Based on the data flow graph standard, four basic elements of the baseline flow graph are defined: external entities, processing procedures, data storage, and data flow. Normal data flow paths are extracted from historical flow data to form typical data flow patterns. Based on the extracted flow paths, a data space topology is constructed. According to business complexity, the baseline flow graph is divided into three layers: top-level, middle-level, and bottom-level, forming a complete baseline flow graph. By comparing historical data, abnormal flow patterns are identified and eliminated, and the stability and representativeness of the baseline flow graph are evaluated. The data potential model is associated and integrated with the baseline flow graph to establish a mapping relationship between data potential and flow paths. The constructed data flow potential model and the baseline flow graph are uniformly stored in the metadata management repository of the data space.
[0015] Preferably, in step S2, based on the data flow potential energy model, data entities with potential energy values exceeding a preset threshold are selected from the trusted data space, and the business logic context of the selected high potential energy data entities is parsed.
[0016] Based on the business context analysis results, decoy data is dynamically generated: using a preset business template, false but logically consistent values are filled in. The decoy data shares the same business relationships with the real data, enhancing the business value of the decoy. Hidden markers are implanted in the access interface and flow path of the decoy data. The generated decoy data is deployed in a mixed manner with the real data in the data space. The complete deployment information of the decoy is recorded.
[0017] Preferably, in step S3, a lightweight flow monitoring agent is deployed at key nodes across the entire data link in the trusted data space to capture all data flow events. For the captured flow events, four types of core features are dynamically extracted, including: flow direction features, traffic features, permission features, and sequence features. Based on the data entity ID in the current flow event, the corresponding baseline pattern is loaded in real time from the baseline flow graph library.
[0018] The potential energy value of the target data entity is mapped to the access permission level onto the potential energy gradient coordinate system. When the flow direction violates the potential energy gradient, it is marked as a gradient anomaly. A deep comparison is performed on the flow path and access sequence to calculate the similarity between the current flow path and the baseline path.
[0019] ,
[0020] In the formula, Indicates path similarity. This represents the size of the common nodes of the current path P and the baseline path B. This indicates the length of the current path P. Indicates the length of baseline path B. This represents the maximum value between the current path length and the baseline path length. Represents the order sensitivity coefficient. This represents the absolute value of the difference between the position index of node c in the current path and the baseline path. exp represents the exponential function. If the similarity is lower than the preset threshold, it is judged as an abnormal path detour.
[0021] Preferably, in step S3, the current access sequence is input into the baseline sequence model to calculate the sequence deviation, which is implemented as follows:
[0022] ,
[0023] In the formula, This represents the sequence deviation, where k represents the length of the currently accessed sequence. This represents the i-th data entity in the sequence. Indicates baseline Markov chain charging arrive The transition probability, This represents the average of all transition probabilities at the baseline. The entropy of the baseline sequence is represented by... denoted by entropy sensitivity coefficient, log represents the logarithmic function. If the deviation exceeds the threshold, it is judged as an atypical anomaly in the sequence.
[0024] Preferably, in step S4, events containing implanted decoy tags are selected from potential energy field disturbance events as attack trigger points. Upon decoy triggering, the deep monitoring module is immediately activated to record all subsequent attack sequences, converting the original sequence into high-dimensional feature vectors, and performing real-time clustering on these feature vectors.
[0025] ,
[0026] In the formula, v represents the feature vector of the attack behavior sequence. Indicates the length of the attack sequence. This represents the type of the i-th operation in the sequence. Indicates the operation type The dynamic business embedding vector is used to output clusters.
[0027] Preferably, in step S4, obtaining clusters and calculating the center vector of each cluster is implemented as follows:
[0028] ,
[0029] In the formula, This represents the center vector of cluster j. Represents the set of all feature vectors in cluster j. Indicates the size of cluster j; analyzes the core behavioral pattern of the cluster center vector, and automatically maps intent tags based on a preset business knowledge base;
[0030] The confidence score is calculated for each intent label, implemented as follows:
[0031] ,
[0032] In the formula, Let L represent the intent confidence of cluster j, and L represent the set of business intents. A dynamic intent template representing business intent l. The Euclidean norm of a vector.
[0033] Preferably, in step S5, the output attack intent tags are mapped to a specific set of data entities; a dynamic slicing strategy is generated based on the intent target set, the isolation range is determined according to the intent granularity, and the topological path of the target data entities in the data space is calculated, which is implemented as follows:
[0034] ,
[0035] In the formula, Represents the topological path of the target data entity. This represents the set of typical flow path nodes for target data entities in the baseline flow map. The potential energy value of node c is represented by c, and the topological coordinate vector of node c is represented by c. Slice routing rules are generated based on the topological path, and the slice execution order is dynamically adjusted based on the urgency of the attack intent.
[0036] Preferably, in step S5, a topology-level reconstruction is performed at the data space control layer, the data flow routing table is modified, the access path of the target data entity is redirected from the main space to the isolated subspace, an isolated subspace is created at the data space network layer, and the target data entity and its associated path are included therein. The traffic scheduling mechanism of the data space is switched during the intervals between attacker operations.
[0037] Preferably, in step S6, within the isolated data subspace, the attack intent tags clustered in step S4 are dynamically matched with a predefined countermeasure strategy library to transform the intent semantics into business-aware strategy logic; simultaneously, the attacker's behavior sequence in the subspace is monitored in real time, and the trap strength is dynamically adjusted based on the reaction efficiency; the system synchronously quantifies the strategy effect, including trap recognition rate, time cost increase, and resource consumption multiple, and generates structured data packets.
[0038] Preferably, in step S7, by deeply analyzing the entire lifecycle data of the attack, the attack features are structurally precipitated into a knowledge base, the four-dimensional weight allocation of the data flow potential energy model is dynamically optimized, and the business context association depth of the dynamic decoy generation logic is updated. At the same time, based on the intent clustering results of the attack sequence, the order sensitivity coefficient in the path similarity formula and the entropy sensitivity coefficient in the sequence deviation formula are adaptively adjusted to make the anomaly detection more in line with the new attack mode. Finally, the optimized potential energy model, decoy logic and parameter configuration are precipitated into an immune memory library, so that the system can automatically apply historical attack experience in subsequent defenses.
[0039] Compared with the prior art, the beneficial effects of the present invention are:
[0040] 1. This invention solves the core problems of existing technologies, such as delayed response to new low-frequency slow attacks, lack of awareness of data flow status, contradiction between static security strategies and dynamic data sharing, and insufficient improvement of attacker costs, by constructing a trusted data space adaptive immune protection method based on data flow potential energy and intent balance. It realizes the transformation from passive boundary protection to active in vivo immunity, enabling the trusted data space to have integrated capabilities of prediction, defense, response, and evolution, significantly improving the efficiency and depth of data security flow, and ensuring the dual goals of business continuity and security in dynamic sharing scenarios.
[0041] 2. This invention uses unsupervised clustering analysis to analyze attack behavior sequences, abstract the attacker's tactical intent, strip away surface behaviors to focus on the real target, enabling the system to understand the attacker's purpose rather than tools, avoiding misjudgments and inefficiencies caused by the inability to understand intent, and achieving a qualitative change from behavior capture to intent insight;
[0042] 3. This invention resolves the contradiction between static isolation and dynamic sharing by reconstructing the data space topology according to the attack intent, performing dynamic data space slicing and isolation, and realizing flexible isolation of "data does not move, space moves". Attackers are trapped in a dedicated sandbox, real data is not exposed, business flow is uninterrupted, and an execution environment is provided for intent countermeasures, so that security protection and business sharing are seamlessly integrated and the impact of isolation on collaboration efficiency is eliminated.
[0043] 4. This invention generates customized data and scenarios based on attack intent in an isolated subspace to counteract the intent. Through dynamic traps, the cost of each operation by the attacker increases exponentially, resulting in losses for successful operations and undermining the economics of the attack. At the same time, the strategy effect feedback optimizes the system, forming a self-evolving closed loop, which greatly improves the initiative and sustainability of the protection, and traps the attacker in a self-set trap. Attached Figure Description
[0044] Figure 1 The following is the operational flow of the adaptive protection method for big data information network security based on trusted data space according to the present invention. Figure 1 ;
[0045] Figure 2 The following is the operational flow of the adaptive protection method for big data information network security based on trusted data space according to the present invention. Figure 2 ;
[0046] Figure 3 The following is the operational flow of the adaptive protection method for big data information network security based on trusted data space according to the present invention. Figure 3 ;
[0047] Figure 4 The following is the operational flow of the adaptive protection method for big data information network security based on trusted data space according to the present invention. Figure 4 ;
[0048] Figure 5 The following is the operational flow of the adaptive protection method for big data information network security based on trusted data space according to the present invention. Figure 5 . Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Example
[0051] Please see Figures 1-5 As shown, the present invention provides a technical solution comprising the following steps:
[0052] S1. Based on the spatial attributes and historical flow of data, perform data flow potential modeling and baseline flow map construction through calculation and recording;
[0053] S2. Based on the data flow potential energy model, highly interactive dynamic decoy deployment is carried out by dynamically generating logically related data;
[0054] S3. Based on the baseline flow map, the potential energy field abnormal disturbance is sensed by real-time monitoring of the data flow status.
[0055] S4. Based on the perceived potential energy field disturbance and the triggered dynamic decoy, perform unsupervised clustering analysis of the attacker's behavior sequence to determine the attack intent.
[0056] S5. Based on the attack intent identified through clustering, perform dynamic data space slicing and isolation based on intent by reconstructing the data space topology;
[0057] S6. In the isolated data subspace, based on the attack intent, customized data and scenarios are generated and fed back to perform intent countermeasures and adaptive strategy generation.
[0058] S7. Based on complete attack lifecycle data and countermeasures results, immune memory and potential baseline evolution are carried out through knowledge accumulation and model optimization.
[0059] In this embodiment, in step S1, a data flow monitoring module is deployed in the trusted data space to record historical data such as the flow path, traffic, frequency, and access sequence of all data entities. The collected flow data is cleaned to remove outliers and noise, and the data format is unified. Based on business characteristics, the historical flow data is divided into time windows. Based on the business scenario, weights are assigned to four potential energy dimensions, such as sensitivity weight 40%, authority weight 30%, demand popularity 20%, and historical entropy value 10%. For each data object, its current potential energy value is calculated as follows: Potential energy value = sensitivity × sensitivity weight + authority × authority weight + demand popularity × demand popularity weight + historical entropy value × historical entropy value weight.
[0060] Specifically, a dynamic update mechanism for preset potential energy values is used to periodically recalculate and update the potential energy values of data objects based on real-time data flow. Based on the data flow graph standard, four basic elements of the baseline flow graph are defined: external entities, such as data sources and destinations; processing processes, such as data processing nodes; data storage, such as data residence locations; and data flow, such as data flow paths. Normalized data flow paths are extracted from historical flow data to form typical patterns of data flow.
[0061] Specifically, based on the extracted flow path, a topological structure of the data space is constructed, including key information such as data flow direction, traffic volume, and access frequency. According to business complexity, the baseline flow map is divided into three levels: top layer (system level), middle layer (module level), and bottom layer (data object level), forming a complete baseline flow map.
[0062] By comparing historical data, abnormal flow patterns are identified and eliminated, and the stability and representativeness of the baseline flow map are evaluated. A regular update mechanism for the baseline flow map is established, and the data potential energy model is associated and integrated with the baseline flow map to establish a mapping relationship between data potential energy and flow path. The potential energy change gradient on the data flow path is analyzed, potential abnormal flow patterns are identified, and the constructed data flow potential energy model and baseline flow map are uniformly stored in the metadata management library of the data space.
[0063] In this embodiment, in step S2, based on the data flow potential energy model, data entities with potential energy values exceeding a preset threshold are selected from the trusted data space. For the selected high-potential-energy data entities, their business logic context is parsed, including:
[0064] Data types, such as patient medical records and stock transaction details;
[0065] Related entity relationships, such as medical records being associated with hospital departments, or transactions being associated with customer accounts;
[0066] Typical access patterns, such as high-frequency query times and frequently used API interfaces.
[0067] Specifically, based on the business context analysis results, decoy data is dynamically generated: using preset business templates, such as medical data templates containing standard fields, and filling in false but logically consistent values, the decoy data shares the same business relationships with the real data, such as associating fictitious medical records with real departments and fictitious transactions with real customer accounts, to enhance the business value of the decoy, such as adding high-value tags and simulating real-time updates to increase the probability of attackers clicking.
[0068] Specifically, stealth markers are implanted in the access interface and flow path of the decoy data:
[0069] Access interface tagging: Embed a unique identifier, such as a specific HTTP header field X-Detour-Token, in the API endpoint or data access interface to mark it as a decoy;
[0070] Flow path marking: Add invisible metadata, such as Detour-Flag:TRUE, to the data flow path log to record the flow trajectory of the decoy;
[0071] Tag invisibility: The tag is designed so that attackers cannot detect it through conventional means, such as coexisting with legitimate request headers and having no additional traffic characteristics.
[0072] In this embodiment, the generated decoy data and real data are deployed together in the data space:
[0073] Distribution strategy: The bait is randomly distributed according to business logic. For example, bait is placed at 5% of the traffic points around high-potential data to avoid concentrated exposure.
[0074] Deployment timing: Dynamically insert decoys during off-peak data access periods, such as nighttime business downtime, to minimize the impact on normal business operations;
[0075] Hybrid verification: After deployment, small-scale tests are conducted to verify the logical consistency between the decoy and the real data (such as query response time and field matching degree);
[0076] Specifically, record the complete deployment information of the decoy:
[0077] Metadata storage: The decoy location, tagging information, and generation context are stored in the metadata management repository;
[0078] Status monitoring: Real-time tracking of the decoy's access activity, such as the number of times it is accessed and the response time, and dynamic adjustment of deployment density;
[0079] Expiration mechanism: Set the validity period of the bait, such as 24 hours, and automatically recycle it after the expiration to avoid long-term invalid bait occupying resources.
[0080] In this embodiment, in step S3, a lightweight flow monitoring agent is deployed at key nodes across the entire data chain in the trusted data space. These key nodes, such as data interfaces, gateways, and storage layers, capture all data flow events, including data entity IDs, source / target participant identities, access timestamps, request types, permission levels, data volume, and path node sequences. For the captured flow events, four core features are dynamically extracted, including:
[0081] Flow characteristics: data flow direction (source → destination) and path node sequence;
[0082] Traffic characteristics: data volume per unit time, request frequency, and data entity access frequency;
[0083] Access control characteristics: A comparison between the access level of the access party and the potential value of the target data entity;
[0084] Sequence characteristics: time intervals between consecutive access behaviors and data entity access order patterns.
[0085] Specifically, based on the data entity ID in the current flow event, the corresponding baseline pattern is loaded in real time from the baseline flow graph library:
[0086] Baseline hierarchical invocation: Invokes the underlying baseline at the data entity level, such as the flow pattern of a single medical record, rather than the global baseline;
[0087] Dynamic matching: Select the latest baseline version based on a time window to ensure that it reflects the current business norm.
[0088] The potential energy value of the target data entity is mapped to the access permission level of the party to the potential energy gradient coordinate system. When the flow direction violates the potential energy gradient, such as a party with low permission requesting high potential energy data, or the difference in potential energy value exceeds the threshold, it is marked as a gradient anomaly. The threshold is adaptive based on the business scenario.
[0089] A deep alignment of the flow path and access sequence is performed to calculate the similarity between the current flow path and the baseline path, which is implemented as follows:
[0090] ,
[0091] In the formula, Indicates path similarity. This represents the size of the common nodes between the current path P and the baseline path B, and the number of common nodes. This represents the length and number of nodes of the current path P. This indicates the length and number of nodes of baseline path B. This represents the maximum value between the current path length and the baseline path length. Represents the order sensitivity coefficient. This represents the absolute value of the difference between the position indices of node c in the current path and the baseline path. for , This represents the index of node c in the current path P. This represents the position index of node c in the baseline path B, and exp represents the exponential function. If the similarity is lower than the preset threshold, it is determined to be an abnormal path detour.
[0092] In this embodiment, step S3 involves inputting the current access sequence into the baseline sequence model and calculating the sequence deviation, which is implemented as follows:
[0093] ,
[0094] In the formula, This represents the sequence deviation, where k represents the length of the currently accessed sequence and the number of data entities. This represents the i-th data entity in the sequence. Indicates baseline Markov chain charging arrive The transition probability, This represents the average of all transition probabilities at the baseline. The entropy of the baseline sequence is represented by... for This indicates the randomness of the baseline flow. denoted by entropy sensitivity coefficient, log represents the logarithmic function. If the deviation exceeds the threshold, it is judged as an atypical anomaly in the sequence.
[0095] Specifically, the detected abnormal features are aggregated, and gradient, path, and sequence abnormal events are aggregated by time window. When the number of aggregated events exceeds the dynamic threshold, such as three consecutive abnormalities or single multiple abnormalities, a potential energy field disturbance alarm is triggered. If the disturbance event involves an implanted decoy tag, such as Detour-Flag: TRUE, it is preferentially marked as a decoy trigger event.
[0096] In this embodiment, in step S4, events containing implanted decoy tags are selected from potential energy field disturbance events as attack trigger points. After the decoy is triggered, the deep monitoring module is immediately activated to record all subsequent operation sequences of the attacker, including operation type, key attributes, and sequence integrity.
[0097] The original operation sequence is converted into a high-dimensional feature vector. A temporal embedding model, such as the Transformer encoder, generates a fixed-length feature vector. Real-time clustering is then performed on these feature vectors. This is implemented as follows:
[0098] ,
[0099] In the formula, v represents the feature vector of the attack behavior sequence. Indicates the length of the attack sequence. This represents the type of the i-th operation in the sequence. Indicates the operation type The dynamic business embedding vector is used to output clusters.
[0100] In this embodiment, step S4 involves obtaining clusters and calculating the center vector of each cluster, which is implemented as follows:
[0101] ,
[0102] In the formula, This represents the center vector of cluster j. Let v represent the set of all feature vectors in cluster j. Indicates the size of cluster j; analyzes the core behavioral pattern of the cluster center vector, and automatically maps intent tags based on a preset business knowledge base;
[0103] Specifically, the clustering results are mapped to attack intentions that are understandable to the business: a confidence score is calculated for each intention label, implemented as follows:
[0104] ,
[0105] In the formula, This represents the intent confidence level of cluster j, ranging from [0,1]. A larger value indicates a higher intent match. L represents the set of business intents. A dynamic intent template representing business intent l. The Euclidean norm of a vector.
[0106] In this embodiment, in step S5, the output attack intent label is mapped to a specific data entity set:
[0107] Real data: Extract high-potential data entities that are the target of the intent, such as Class A medical records and financial transaction data;
[0108] Decoy data: Related to dynamic decoys of the same type deployed in S2, such as Class A false data generated in S2;
[0109] Related entities: Automatically expand related data entities, such as patient records and transaction accounts corresponding to Class A data.
[0110] Specifically, based on the intent target set, a dynamic slicing strategy is generated, the isolation range is determined according to the intent granularity, and the topological path of the target data entity in the data space is calculated, which is implemented as follows:
[0111] ,
[0112] In the formula, Represents the topological path of the target data entity. This represents the set of typical flow path nodes for target data entities in the baseline flow map. The potential energy value of node c is represented by c, and the topological coordinate vector of node c is represented by c. Slice routing rules are generated based on the topological path, and the slice execution order is dynamically adjusted based on the urgency of the attack intent.
[0113] In this embodiment, in step S5, a topology-level reconstruction is performed at the data space control layer, the data flow routing table is modified, the access path of the target data entity is redirected from the main space to the isolated subspace, and an isolated subspace is created at the data space network layer. The subspace is logically isolated but not physically isolated, and the target data entity and its associated path are included in it. The traffic scheduling mechanism of the data space, such as traffic control based on the service mesh, is switched during the attacker's operation interval.
[0114] Specifically, deploy an attacker-exclusive environment in an isolated subspace:
[0115] Decoy Filling: Inject dynamic decoys of the same type generated by S2 to ensure they match the attacker's intent;
[0116] Behavior monitoring: Activate the S4 deep monitoring module to capture all operations within the subspace in real time;
[0117] Policy injection: Pre-set S6 intent check policy, so that the attacker's successful operation in the subspace is actually a system trap;
[0118] Specifically, within the isolated subspace, ensure that the entire process of an attacker's operations is traceable:
[0119] Operation log: Records all actions of the attacker in the subspace, including timestamps, operation sequences, and data content;
[0120] No-interference mechanism: Data response within the subspace simulates the real environment, such as download speed and API response time being consistent with the main space, to prevent attackers from noticing any anomalies;
[0121] Real-time feedback: Push behavior logs to the S6 module for dynamic generation of checks and balances strategies;
[0122] Perform post-isolation verification:
[0123] Main space verification: Ensures that the data flow of other participants in the main data space is not affected, such as normal users being able to access Class A data;
[0124] Isolation domain monitoring: Real-time monitoring of resource utilization in isolated subspaces, such as CPU and bandwidth, to avoid impacting overall system performance;
[0125] Automatic revert mechanism: If the attacker does not perform any further operations in the subspace, such as after a 5-minute timeout, the subspace will be automatically reclaimed and the main space topology will be restored.
[0126] In this embodiment, in step S6, within the isolated data subspace, attack intent tags clustered in step S4 are dynamically matched with a predefined countermeasure strategy library to transform the intent semantics into business-aware strategy logic. For example, for theft intent, medical records containing fictitious patient IDs are generated and embedded with invisible tracking watermarks and conditional triggering logic. Simultaneously, the attacker's behavior sequence in the subspace is monitored in real time, and the trap strength is dynamically adjusted based on reaction efficiency. For example, when the attacker quickly identifies the watermark, the density is automatically increased by 10%, ensuring that each round of operation consumes additional resources and triggers an exponential increase in cost. For example, the cost for the first time is 1, and the cost for the tenth time is 1024, making the attacker's successful operation a negative gain. For example, data destruction is triggered after downloading fake data.
[0127] The system synchronously quantifies the effectiveness of the strategy, including trap identification rate, time cost increase, and resource consumption multiplier. It generates structured data packets and pushes them to the S7 module, which are then stored as immune memory to optimize the S1 potential energy model weights and the S2 decoy generation logic. This achieves a closed loop of intent control, dynamic feedback, and system evolution, causing attackers to continuously fall into self-set traps in an isolated subspace. The system's capabilities increase exponentially with the number of attacks, completely undermining the economics of attacks and ensuring zero business interruption.
[0128] Preferably, in step S7, by deeply analyzing the attack's entire lifecycle data, the attack features are structurally precipitated into a knowledge base, the four-dimensional weight allocation of the data flow potential energy model is dynamically optimized, and the business context association depth of the dynamic decoy generation logic is updated.
[0129] Based on the intent clustering results of the attack sequence, the order sensitivity coefficient in the path similarity formula and the entropy sensitivity coefficient in the sequence deviation formula are adaptively adjusted to make anomaly detection more in line with new attack patterns. Finally, the optimized potential energy model, decoy logic and parameter configuration are deposited into an immune memory library, enabling the system to automatically apply historical attack experience in subsequent defenses, achieve exponential evolution of the security baseline, completely eliminate passive lag, and achieve the closed-loop evolution goal of strong immunity after a single attack.
[0130] Working principle: By deploying a data flow monitoring module within the trusted data space, the system systematically records and cleans the historical flow data of data entities, including information such as path, traffic, frequency, and access sequence; based on the business scenario, it dynamically allocates weights for sensitivity, authority, demand intensity, and historical entropy, calculates the dynamic potential value of each data object, and establishes a baseline flow graph associated with it; this graph defines four basic elements—external entities, processing procedures, data storage, and data flow—using data flow graph standards, extracts normal flow patterns from historical data, and constructs a hierarchical topology structure;
[0131] By constructing a potential energy model, the system automatically filters high-potential-energy data entities and deeply analyzes their business context, including data types, related entity relationships, and typical access patterns. It dynamically generates logically closely related decoy data, ensuring that the content conforms to business logic and shares the same relationships as real data, while enhancing business value to increase attractiveness. Hidden markers are implanted in the access interfaces and flow paths of the decoy data to ensure the markers are imperceptible. The decoys are randomly distributed according to business logic and deployed during off-peak periods. After verifying logical consistency, deployment information is fully recorded. Lightweight monitoring agents are deployed at key nodes throughout the data space to capture data flow events in real time and extract flow direction, traffic, permissions, and sequences. Four core features: The system dynamically loads the corresponding data entity flow patterns in the baseline flow graph, maps potential energy values and permission levels to the potential energy gradient coordinate system, and detects gradient violations, path deviations, and sequence anomalies; by aggregating multi-dimensional abnormal events, it triggers potential energy field disturbance alarms and prioritizes marking decoy trigger events; the system selects potential energy field disturbance events containing decoy markers as attack behavior trigger points, immediately activates the deep monitoring module to record the attacker's subsequent operation sequence; the sequence is converted into a high-dimensional feature vector, and an unsupervised clustering algorithm is applied to analyze the behavior pattern and calculate the cluster center vector; based on the business knowledge base, the center vector is parsed, automatically mapped to attack intent labels that are understandable to the business, and the confidence level is calculated;
[0132] The system maps attack intent tags to specific data entity sets, including real data, decoy data, and related entities. Based on the granularity of the intent, it generates dynamic slicing strategies, reconstructs the data space topology, and redirects target data paths from the main space to isolated subspaces. In the isolated subspaces, a dedicated environment is deployed, injecting similar decoys, activating behavior monitoring, and pre-setting countermeasures. After execution, it verifies that the main space's business flow is unaffected. In the isolated subspaces, the system dynamically matches a countermeasures library based on attack intent tags, transforming intent semantics into business-aware strategy logic and generating customized data containing traps. It monitors attack behavior sequences in real time, dynamically adjusting trap strength to consume additional resources, causing attack costs to increase exponentially. Successful attacker operations result in negative returns; the system quantifies the strategy's effectiveness and provides feedback for optimization. The system deeply analyzes the entire attack lifecycle data, including perturbation patterns, intent clustering, and countermeasure effects, and structurally stores this data in a knowledge base. It dynamically optimizes the four-dimensional weight allocation of the potential energy model and the business association depth of the decoy generation logic, while adaptively adjusting anomaly detection parameters. The optimization results are stored in an immune memory library, enabling the system to automatically apply historical experience in subsequent defenses, and the security baseline increases exponentially with the number of attacks.
[0133] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their likenesses.
[0134] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.
Claims
1. An adaptive protection method for big data information network security based on trusted data space, characterized in that, Includes the following steps: S1. Based on the spatial attributes and historical flow of data, perform data flow potential modeling and baseline flow map construction through calculation and recording; S2. Based on the data flow potential energy model, highly interactive dynamic decoy deployment is carried out by dynamically generating logically related data; S3. Based on the baseline flow map, the potential energy field abnormal disturbance is sensed by real-time monitoring of the data flow status. S4. Based on the perceived potential energy field disturbance and the triggered dynamic decoy, perform unsupervised clustering analysis of the attacker's behavior sequence to determine the attack intent. S5. Based on the attack intent identified through clustering, perform dynamic data space slicing and isolation based on intent by reconstructing the data space topology; S6. In the isolated data subspace, based on the attack intent, customized data and scenarios are generated and fed back to perform intent countermeasures and adaptive strategy generation. S7. Based on complete attack lifecycle data and countermeasures results, immune memory and potential baseline evolution are carried out through knowledge accumulation and model optimization.
2. The adaptive protection method for big data information network security based on trusted data space according to claim 1, characterized in that: In S1, a data flow monitoring module is deployed in the trusted data space to record the flow path, traffic, frequency, and access sequence of all data entities. The collected flow data is cleaned to remove outliers and noise, and the data format is unified. Based on business characteristics, the historical flow data is divided into time windows. Based on the business scenario, weights are assigned to the four potential energy dimensions. For each data object, its current potential energy value is calculated. Based on the data flow graph standard, four basic elements of the baseline flow graph are defined: external entities, processing procedures, data storage, and data flow. Normal data flow paths are extracted from historical flow data to form typical data flow patterns. Based on the extracted flow paths, a data space topology is constructed. According to business complexity, the baseline flow graph is divided into three layers: top-level, middle-level, and bottom-level, forming a complete baseline flow graph. By comparing historical data, abnormal flow patterns are identified and eliminated, and the stability and representativeness of the baseline flow graph are evaluated. The data potential model is associated and integrated with the baseline flow graph to establish a mapping relationship between data potential and flow paths. The constructed data flow potential model and the baseline flow graph are uniformly stored in the metadata management repository of the data space.
3. The adaptive protection method for big data information network security based on trusted data space according to claim 1, characterized in that: In S2, based on the data flow potential energy model, data entities with potential energy values exceeding a preset threshold are selected from the trusted data space, and their business logic context is analyzed for the selected high potential energy data entities. Based on the business context analysis results, decoy data is dynamically generated: using a preset business template, false but logically consistent values are filled in. The decoy data shares the same business relationship with the real data, enhancing the business value of the decoy. Hidden markers are implanted in the access interface and flow path of the decoy data. The generated decoy data is mixed with real data and deployed in the data space; and the complete deployment information of the decoy is recorded.
4. The adaptive protection method for big data information network security based on trusted data space according to claim 1, characterized in that: In S3, a lightweight flow monitoring agent is deployed at key nodes across the entire data flow space to capture all data flow events. For the captured flow events, four types of core features are dynamically extracted, including: flow direction features, traffic features, permission features, and sequence features. Based on the data entity ID in the current flow event, the corresponding baseline pattern is loaded in real time from the baseline flow graph library. The potential energy value of the target data entity is mapped to the access permission level onto the potential energy gradient coordinate system. When the flow direction violates the potential energy gradient, it is marked as a gradient anomaly. A deep comparison is performed on the flow path and access sequence to calculate the similarity between the current flow path and the baseline path. , In the formula, Indicates path similarity. This represents the size of the common nodes of the current path P and the baseline path B. This indicates the length of the current path P. Indicates the length of baseline path B. This represents the maximum value between the current path length and the baseline path length. Represents the order sensitivity coefficient. This represents the absolute value of the difference between the position index of node c in the current path and the baseline path. exp represents the exponential function. If the similarity is lower than the preset threshold, it is judged as an abnormal path detour.
5. The adaptive protection method for big data information network security based on trusted data space according to claim 4, characterized in that: In step S3, the current access sequence is input into the baseline sequence model to calculate the sequence deviation, which is implemented as follows: , In the formula, This represents the sequence deviation, where k represents the length of the currently accessed sequence. This represents the i-th data entity in the sequence. Indicates baseline Markov chain charging arrive The transition probability, This represents the average of all transition probabilities at the baseline. The entropy of the baseline sequence is represented by... denoted by entropy sensitivity coefficient, log represents the logarithmic function. If the deviation exceeds the threshold, it is judged as an atypical anomaly in the sequence.
6. The adaptive protection method for big data information network security based on trusted data space according to claim 1, characterized in that: In step S4, events containing implanted decoy markers are selected from potential energy field disturbance events as attack trigger points. Upon decoy triggering, the deep monitoring module is immediately activated to record all subsequent attack sequences. The original attack sequences are converted into high-dimensional feature vectors, and real-time clustering is performed on these feature vectors. , In the formula, v represents the feature vector of the attack behavior sequence. Indicates the length of the attack sequence. This represents the type of the i-th operation in the sequence. Indicates the operation type The dynamic business embedding vector is used to output clusters.
7. The adaptive protection method for big data information network security based on trusted data space according to claim 1, characterized in that: In step S4, clusters are obtained, and the center vector of each cluster is calculated, which is implemented as follows: , In the formula, This represents the center vector of cluster j. Represents the set of all feature vectors in cluster j. Indicates the size of cluster j; analyzes the core behavioral pattern of the cluster center vector, and automatically maps intent tags based on a preset business knowledge base; The confidence score is calculated for each intent label, implemented as follows: , In the formula, Let L represent the intent confidence of cluster j, and L represent the set of business intents. A dynamic intent template representing business intent l. The Euclidean norm of a vector.
8. The adaptive protection method for big data information network security based on trusted data space according to claim 1, characterized in that: In step S5, the output attack intent tags are mapped to specific data entity sets; based on the intent target set, a dynamic slicing strategy is generated, the isolation range is determined according to the intent granularity, and the topological path of the target data entities in the data space is calculated, which is implemented as follows: , In the formula, Represents the topological path of the target data entity. This represents the set of typical flow path nodes for target data entities in the baseline flow map. The potential energy value of node c is represented by c, and the topological coordinate vector of node c is represented by c. Slice routing rules are generated based on the topological path, and the slice execution order is dynamically adjusted based on the urgency of the attack intent.
9. The adaptive protection method for big data information network security based on trusted data space according to claim 8, characterized in that: In S5, a topology-level reconstruction is performed at the data space control layer, the data flow routing table is modified, and the access path of the target data entity is redirected from the main space to the isolated subspace. An isolated subspace is created at the data space network layer, and the target data entity and its associated path are included in it. The traffic scheduling mechanism of the data space is switched during the intervals between attacker operations.
10. The adaptive protection method for big data information network security based on trusted data space according to claim 1, characterized in that: In step S6, within the isolated data subspace, attack intent tags clustered in step S4 are dynamically matched with a predefined countermeasure strategy library to transform intent semantics into business-aware strategy logic. Simultaneously, the attacker's behavior sequence in the subspace is monitored in real time, and the trap strength is dynamically adjusted based on reaction efficiency. The system synchronously quantifies the strategy effect, including trap recognition rate, time cost increase, and resource consumption multiple, and generates structured data packets.