Method and device for constructing data security agent based on large model
By building a data security agent system based on large models, using the collaborative work of edge agents and cloud, we realize the hidden risk identification of multi-source heterogeneous data in dynamic environments and effective defense of new attack modes, solving the problem of insufficient identification and defense in the existing technology.
Patent Information
- Application Number
- CN202510964330.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-11
AI Technical Summary
The prior art is difficult to effectively identify hidden risks in multi-source heterogeneous data in dynamic environments, and its defense capabilities against new attack modes are insufficient.
Build a data security agent system based on large models, detect potential attacks through edge agents and initiate consensus requests. The cloud uses the blockchain-stored risk response strategy inference rule database and multiple edge agents to vote to determine the nature of the attack, generate adversarial samples, quantify risk levels through Bayesian networks, and formulate dynamic defense strategies.
It ensures cognitive consistency of attacks in dynamic and complex environments, reduces excessive defense caused by false positives, realizes adaptive learning and dynamically formulates risk response strategies, and improves the defense capabilities of new attack modes.
Smart Images

Figure CN120474842A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information security technology, and in particular to a method and device for constructing a data security intelligent body based on a large model. Background Art
[0002] With the acceleration of digital transformation, data security threats are becoming more diverse and dynamic. Traditional data security protection mechanisms based on rules or manual intervention face challenges such as response delays and lagging knowledge updates when dealing with complex attack scenarios.
[0003] While large language models have shown promise in areas such as natural language processing and intelligent decision-making, their application in data security remains limited. Conventional algorithms struggle to effectively identify hidden risks in multi-source, heterogeneous data, and their lack of adaptive learning capabilities in dynamic environments prevents them from effectively defending against new attack vectors. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a method for constructing a data security intelligent body based on a large model, so as to solve the technical problem that it is impossible to effectively defend against new attack modes in a dynamic environment.
[0005] In a first aspect, an embodiment of the present application provides a method for constructing a data security agent based on a large model, the method being applied to a data security agent system, the system comprising: multiple edge agents and a cloud, the method comprising: When the current edge agent detects an unknown potential attack, it initiates a consensus request to the cloud, which includes security threat elements and evidence data; The cloud calls at least three other edge agents to determine that the unknown potential attack is a real new unknown attack through voting based on the risk response strategy inference rule library stored in the blockchain and the consensus request; The cloud performs incremental learning based on the security threat factors and the evidence data to generate new adversarial samples. Taking the security threat factors, the evidence data and the new adversarial samples as input, the cloud-based Bayesian network quantifies the risk level of the new unknown attack. Combining the risk level and the security threat factors, a risk response strategy is formulated, and at least one other edge intelligent agent is dispatched to perform dynamic defense operations.
[0006] Optionally, calling at least three other edge agents to vote to determine that the unknown potential attack is a real new unknown attack includes: Calling at least three other edge agents to vote using the PBFT (Practical Byzantine Fault Tolerance) algorithm, and determining that the unknown potential attack is a real new unknown attack when two-thirds of the at least three other edge agents determine that the unknown potential attack is a real attack; When the risk response strategy customized by the cloud includes a high-risk defense operation, the cloud notifies multiple edge agents scheduled in the risk response strategy to execute the high-risk defense operation through alliance chain signature authorization, and the high-risk defense operation includes: at least one of: global service interruption and global network disconnection.
[0007] Optionally, the method further includes: When the current edge agent mines high-frequency attack features through the edge domain large model, it determines that there is a potential attack. The current edge agent performs lightweight feature encoding on the high-frequency attack features to generate security threat elements of the potential attack, wherein the security threat elements include: attack entry, attack stage features, The current edge agent uses the attack entrance as is the root node of the edge domain Bayesian network, the attack stage feature is used as the intermediate node of the edge domain Bayesian network, and the potential attack is used as the hidden variable node; Obtaining evidence data of the current edge agent as an observation value of the edge domain Bayesian network, quantifying the risk level of the potential attack, and uploading the risk level as a local network parameter of the edge domain Bayesian network to the cloud; The cloud updates the global network parameters of the cloud Bayesian network on the cloud based on the local network parameters of the edge domain Bayesian network through the blockchain.
[0008] Optionally, the cloud updates the global network parameters of the cloud Bayesian network of the cloud based on the local network parameters of the edge domain Bayesian network through the blockchain, including: The cloud learns the local network parameters of the edge domain Bayesian network through the teacher-student network to obtain updated global network parameters of the cloud Bayesian network, and writes the updated records of the local network parameters and the global network parameters into the blockchain, and / or, The current edge agent encrypts the local network parameters and uploads them to the cloud. The cloud aggregates the local network parameters of all edge domain Bayesian networks through federated learning to obtain the global network parameters of the cloud Bayesian network. The global network parameters are sent to all edge agents, and the update records of the global network parameters are written into the blockchain.
[0009] Optionally, the security rule constraints include initial role allocation rules determined according to high-frequency attack characteristics, The initial role allocation rules include: The cloud dispatches at least three edge agents according to the type of attack entry uploaded by the current edge agent to perform asset analysis, threat hunting and threat assessment respectively. The cloud formulates an optimized risk response strategy based on the asset analysis results, threat hunting results and threat assessment results fed back by the at least three edge agents, and distributes the optimized risk response strategy to the scheduled edge agents for execution, wherein threat hunting includes detecting potential attacks through a data security knowledge graph, the asset analysis results include the scale of data assets involved in the potential attack, and the threat assessment includes detecting and assessing the risk level of the potential attack.
[0010] Optionally, when the risk response strategy customized in the cloud is a non-high-risk defense operation, the scheduling at least one other edge agent to perform a dynamic defense operation includes: performing traffic cleaning on the current edge agent; The step of performing traffic cleaning on the current edge agent includes: Transfer the business of the current edge agent to the idle edge agent, Obtaining attack phase characteristics of the current edge agent; Perform at least one of the following defense operations on the current edge agent according to the attack phase characteristics: Algorithm reinforcement, IP blocking, traffic analysis, and permission restriction.
[0011] Optionally, scheduling at least one other edge agent to perform a dynamic defense operation includes: The cloud generates a lightweight on-chain attack fingerprint based on the risk level and security threat factors of the new unknown attack; The cloud determines at least one other edge agent involved in the restored cross-domain attack chain in the blockchain through the threat diffusion speed in the security threat elements, pushes the on-chain attack fingerprint and risk response strategy to each edge agent involved in the restored cross-domain attack chain, and instructs each edge agent involved in the restored cross-domain attack chain to perform defense operations in accordance with the risk response strategy.
[0012] Optionally, when the current edge agent mines high-frequency attack features through the edge domain large model embedded with security rule constraints, the method further includes: The current edge agent determines whether the high-frequency attack feature belongs to an unknown potential attack based on the local edge domain large model and data security knowledge graph. If the high-frequency attack feature belongs to an unknown potential feature, local incremental learning is triggered locally to generate a new adversarial sample. The new adversarial sample is input into the adversarial sample library, and the attack stage features are extracted from the high-frequency attack feature. A graph computing engine is used to construct a dynamic attack graph based on the attack stage features to quantify the level of the new security threat and the threat diffusion speed. Inputting the dynamic attack graph into the edge domain big model, and generating a local risk defense strategy based on the level of the newly added security threat and the threat diffusion speed; Determining, by means of an edge domain Bayesian network, whether a utility analysis result corresponding to the local risk defense strategy meets business requirements, wherein the business requirements include at least one of a service interruption risk, a compliance risk, and a service performance degradation risk; When the utility analysis result indicates that the local risk defense strategy meets the business requirements, the local network parameters of the edge domain Bayesian network are updated, and the dynamic attack graph, the optimized local risk defense strategy, and the updated local network parameters of the edge domain Bayesian network are uploaded to the cloud via the blockchain, so that the cloud updates the cloud Bayesian network; When the utility analysis result indicates that the local risk defense strategy does not meet business requirements, a consensus request is initiated to the cloud.
[0013] Optionally, judging whether the utility analysis result corresponding to the local risk defense strategy meets the business requirements by using an edge domain Bayesian network includes: The edge agent is capable of defining candidate defensive actions using an edge domain Bayesian network; performing a utility analysis on the defensive action; Through real-time pruning search, the defensive action combination that maximizes the expected utility analysis result is selected as the optimal strategy; Based on the update of the edge domain leaf-Bassian network, the optimal strategy is calibrated for posterior probability, and the calibrated optimal strategy is used as the local risk defense strategy.
[0014] In a second aspect, an embodiment of the present application provides an electronic device, including: processor; A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions implement the above method when executed by the processor.
[0015] The beneficial effects brought about by the technical solution provided by the embodiments of the present application include at least: when the current edge agent detects an unknown potential attack, at least three other edge agents on the blockchain vote to determine that the unknown potential attack is a real new unknown attack, so as to ensure the cognitive consistency of multiple edge agents on dynamic and complex attack scenarios, and avoid excessive defense caused by false alarms of a small number of edge agents; when at least three other edge agents vote to determine that the unknown potential attack is a real new unknown attack, the cloud performs incremental learning based on the security threat factors and evidence data sent by the current edge agent, and quantifies the risk level of the new unknown attack through the cloud Bayesian network, and formulates a risk response strategy based on the risk level and security threat factors of the new unknown attack, so that the cloud can adaptively learn the data of the current new unknown attack, and can also reduce the illusion of cognition caused by the fact that the cloud only has a simulated environment but no real business, and dynamically formulates a risk response strategy that matches the new unknown attack based on the data of the new unknown attack to perform dynamic defense operations, improve the adaptability of security defense operations in dynamic and complex environments, and effectively defend against new attack modes. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 This is a method for constructing a data security intelligent entity based on a large model provided in an embodiment of the present application; Figure 2 This is part of the steps of a method for constructing a data security intelligent entity based on a large model provided in an embodiment of the present application; Figure 3 These are some steps of a method for constructing a data security intelligent entity based on a large model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] The present application will be described in detail below in conjunction with the specific embodiments shown in the accompanying drawings, but these embodiments do not limit the present application. Structural, methodological, or functional changes made by ordinary technicians in this field based on these embodiments are included in the scope of protection of the present application.
[0019] The present invention provides a method for constructing a data security agent based on a large model. The method is applied to a data security agent system, which includes multiple edge agents and a cloud. An edge agent is an agent that can perceive the environment and take actions to achieve specific goals. It has autonomy, adaptability, and interaction capabilities. Figure 1 As shown, the method for constructing a data security agent based on a large model includes the following steps: Step S101: When the current edge agent detects an unknown potential attack, it initiates a consensus request to the cloud, which includes security threat elements and evidence data.
[0020] Optionally, an edge domain large model is deployed in the edge agent, and the edge domain large model is used to identify attacks. If the edge domain large model of the current edge agent detects an attack, the edge domain large model determines whether there are abnormal events in the system log of the current edge agent, abnormal network traffic, etc. based on local static evidence and / or dynamic evidence. If the static evidence proves that there is an abnormality, the type of attack can be matched according to the corresponding security rules. If the dynamic evidence proves that there is an abnormality, the edge domain large model can be used to deduce which attack the evidence belongs to. If neither static evidence nor dynamic evidence can determine the type of attack, the attack detected by the current edge agent can be considered an unknown potential attack.
[0021] Exemplarily, static evidence may refer to evidence obtained by direct matching based on security rules, such as evidence obtained by matching data such as files, processes, logs, and network protocols based on an existing vulnerability library.
[0022] For example, abnormal DNS lookup records correspond to DNS hijacking and DNS cache poisoning in security rules.
[0023] For another example, by analyzing the attack signature or file hash value corresponding to the malware list, the malware name corresponding to the attack signature or file hash value in the security rule is found.
[0024] Dynamic evidence can be derived from a knowledge graph and / or adversarial network using a large edge domain model to identify security threats. In this case, security rules may not be able to directly match evidence of the security threat. However, the knowledge graph and / or adversarial network can be used to obtain attack entry points and attack stage characteristics, and the large edge domain model can be used to deduce evidence of the security threat.
[0025] Step S102: The cloud calls at least three other edge agents based on the risk response strategy inference rule base and consensus request stored in the blockchain to determine whether the unknown potential attack is a real new unknown attack through voting.
[0026] Among them, if the current edge agent detects an unknown potential attack and initiates a consensus request to the cloud, the cloud calls at least three other edge agents to vote to determine that the unknown potential attack is a real new unknown attack. That is, the cloud adopts distributed verification to ensure the consistency of multiple edge agents' cognition of complex unknown potential attacks, which can avoid excessive defense caused by false alarms of edge agents.
[0027] Step S103: The cloud performs incremental learning based on security threat factors and evidence data to generate new adversarial samples. Taking security threat factors, evidence data and new adversarial samples as input, the cloud-based Bayesian network quantifies the risk level of new unknown attacks. Combining the risk level and security threat factors, a risk response strategy is formulated, and at least one other edge intelligent agent is dispatched to perform dynamic defense operations.
[0028] Optionally, after determining that an unknown potential attack is a new type of unknown attack, the cloud can call at least three other edge agents to vote to determine whether the unknown potential attack is a real new type of unknown attack. If it is a real new type of unknown attack, the cloud starts incremental learning to generate new adversarial samples. Furthermore, the cloud Bayesian network takes security threat factors, evidence data, and new adversarial samples as input to quantify the risk level of the new unknown attack. Among them, the risk level of the new unknown attack quantified by the Bayesian network can include high risk, medium risk, and low risk. Specifically, the CVSS (Common Vulnerability Scoring System) score can be used to finally determine whether the risk level of the new unknown attack is high risk, medium risk, or low risk.
[0029] In the above-mentioned application embodiment, when the current edge agent detects an unknown potential attack and sends a consensus request to the cloud, at least three other edge agents on the blockchain vote to determine that the unknown potential attack is a real new unknown attack, which triggers the cloud to perform incremental learning. By having at least three other edge agents vote to determine whether it is a real new unknown attack, it can ensure that multiple edge agents have consistent cognition of dynamic and complex attack scenarios, and avoid excessive defense caused by false alarms of a small number of edge agents; the cloud performs incremental learning based on security threat factors and evidence data to generate new adversarial samples, and quantifies the risk level of the new unknown attack through the cloud Bayesian network, and formulates a risk response strategy based on the risk level and security threat factors of the new unknown attack, so that the cloud can adaptively learn the current new unknown attack data, and dynamically formulate a risk response strategy that matches the new unknown attack based on the new unknown attack data, so as to perform dynamic defense operations, improve the adaptability of security defense operations in dynamic and complex environments, and effectively defend against new attack patterns.
[0030] In one embodiment, in step S102, the cloud calls at least three other edge agents to vote to determine whether the unknown potential attack is a real new unknown attack, specifically including: Call at least three other edge agents to vote through the Practical Byzantine Fault Tolerance (PBFT) algorithm. If two-thirds of the edge agents in at least three other edge agents determine that the unknown potential attack is a real attack, the unknown potential attack is determined to be a real new unknown attack. When the risk response strategy customized in the cloud includes high-risk defense operations, the cloud notifies multiple edge agents scheduled in the risk response strategy to perform high-risk defense operations through alliance chain signature authorization. The high-risk defense operations include: at least one of global service interruption and global network disconnection.
[0031] Specifically, each edge agent, acting as a node on the consortium chain, detects an unknown potential attack and initiates a consensus request to the cloud. The participating nodes (e.g., ≥4 nodes) vote using the Practical Byzantine Fault Tolerance (PBFT) algorithm. If ≥2 / 3 of the nodes confirm, the unknown potential attack is considered a genuine new, unknown attack. Furthermore, high-risk defense operations (such as global service interruption or global network disconnection) require signature authorization from a majority of the consortium chain nodes to prevent single node triggering errors.
[0032] It should be noted that the cloud-based large model can simulate the business impact after the execution of the risk response strategy (such as the probability of service interruption and compliance risk) and generate a risk score. If the score exceeds the threshold (such as >70%), it can automatically switch to a milder risk response strategy (such as current limiting instead of blocking).
[0033] In one embodiment, the edge agent is provided with an edge domain large model, and the edge agent can pre-train the edge domain large model. For example, when training the edge domain large model, historical attack data (such as penetration test reports and payloads captured by honeypots) is collected and categorized by attack pattern, such as SQL (Structured Query Language) injection patterns, XSS (Cross Site Scripting) attack patterns, unauthorized access attack patterns, and data sharing and leakage attack patterns. Feature vectors of the historical attack data are annotated according to the attack pattern, such as the grammatical structure of the injection statement, the statistical characteristics of abnormal traffic, and the statistical paradigm of data leakage.
[0034] The edge agent uses a generative adversarial network and / or mutation algorithm to process historical attack data to generate adversarial samples. This historical attack data includes different types of attacks, and the methods for generating adversarial samples vary for each type of attack. Specifically, for text-based attacks, the edge agent injects malicious fragments into query statements to generate adversarial samples. For traffic-based attacks, the edge agent inserts covert attack payloads (such as DNS Domain Name System tunnel data) into legitimate traffic to generate adversarial samples. For data leakage attacks, after identity and permission authentication, the edge agent inserts a small portion of request data that exceeds the user's permission range to generate adversarial samples. The adversarial samples generated from the historical attack data are mixed with legitimate data to construct a library of adversarial samples that closely resemble real-world attacks. This data serves as training data for the edge-domain large model in the edge agent. It should be noted that during subsequent use of the trained edge-domain large model, the edge agent can process newly collected attack data to generate new adversarial samples. These new adversarial samples can be added to the adversarial sample library for updating and training the edge-domain large model, thereby improving its robustness against covert attacks.
[0035] In one embodiment, the edge domain large model includes a heterogeneous data perception module and a threat identification module. The heterogeneous data perception module is used to parse heterogeneous data from multiple sources, such as IoT terminals, to extract data features and identify sensitive data. The heterogeneous data perception module integrates a multi-application layer protocol parsing engine and can parse structured data (such as database records), semi-structured data (such as logs in JSON or XML format), and unstructured data (such as network traffic packets, images, and video data). The threat identification module is used to perform threat identification analysis on identified sensitive data.
[0036] Optionally, when extracting data features, the heterogeneous data perception module in the edge domain large model uses different extraction methods for data in different formats.
[0037] For example, ANTLR (Another Tool for Language Recognition) is used to perform syntax and semantic analysis on structured data to extract key fields in the structured data, such as user operations (addition / deletion / modification / query) on patient medical records in the hospital system.
[0038] For semi-structured data, regular expression matching and natural language processing are used to extract key fields in the semi-structured data, such as timestamp, user ID, operation content, database password information, software and hardware password module access control code information, etc.
[0039] For unstructured data, the source IP, destination IP, port, protocol, and payload features are extracted through application layer protocol analysis.
[0040] Optionally, when identifying sensitive data, the heterogeneous data perception module in the edge domain large model builds a hybrid recognition model based on a rules engine and a deep learning model. This model identifies sensitive fields in the data, such as passwords, access control information, keys, and important business data (such as core area monitoring information, personal privacy data such as ID card numbers and bank card numbers, government reports, and corporate financial statements). It then classifies and levels sensitive data by combining compliance labels in the data security knowledge base (for user information and corporate data). The rules engine can be one that uses regular expression matching to identify sensitive fields, and the deep learning model can be a Bidirectional Long Short-Term Memory with Conditional Random Fields (BLSTM+CRF) model.
[0041] Optionally, when performing threat identification analysis on identified sensitive data, the threat identification module in the edge domain large model can retrieve a pre-generated data security knowledge graph and use entity linking technology to align the sensitive data with the data security knowledge graph. If an abnormal alignment result is obtained, a graph neural network is used to calculate and reason about the similarity of multi-source heterogeneous data with historical attack data to determine whether the current behavior is an attack. Entity relationships can be pre-extracted from data security standard documents (e.g., "User A has permission B, requests asset C, uses protection measure D," so "faces risk E") to generate the data security knowledge graph.
[0042] See also Figure 2 As shown, in one embodiment, before step S101, the method for constructing a data security agent based on a large model further includes the following steps: Step S201: When the current edge agent mines high-frequency attack features through the edge domain large model embedded with security rule constraints, it determines that there is a potential attack. The current edge agent performs lightweight feature encoding on the high-frequency attack features to generate security threat elements of the potential attack, where the security threat elements include: attack entry and attack stage features.
[0043] The integration of security rule constraints with the edge domain big model to obtain the edge domain big model embedded with security rule constraints can be achieved in the following ways: The compliance policies in security rules (such as prohibiting cross-provincial transmission of video surveillance data) are converted into logical expressions, and the converted logical expressions are encoded into executable constraints (i.e., security rule constraints) using a logic programming framework.
[0044] Collect and calculate entity data in sensitive data, such as access users, access frequency, access time distribution, access geography, etc., and use time series analysis models (such as LSTM, Long Short-Term Memory) to detect abnormal operation sequences in entity data, such as high-frequency data export in a short period of time, unauthorized user operations, data access and interface calls during abnormal time periods, and access from abnormal geographic locations.
[0045] For example, security rule constraints can be applied to entity data in the following ways: By introducing security rule constraint weights into the attention layer of the edge domain large model, the security constraint rules pay different attention to different data in the abnormal operation sequence in different scenarios. For example, in the cross-provincial data circulation scenario, the focus on the geographic location field in the abnormal operation sequence is strengthened.
[0046] If the current edge agent identifies high-frequency attack signatures (such as the behavior patterns of specific malicious IP addresses) through the edge domain large model embedded with security rule constraints, it will determine the presence of a potential threat. Furthermore, the current edge agent can use hashing to perform lightweight feature encoding on these high-frequency attack signatures to generate security threat factors for the potential attack. These security threat factors include attack entry points and attack phase characteristics. Attack entry points include kernel vulnerabilities, plaintext communications, weak passwords, and low-level security algorithms. Attack phase characteristics include the attack stage, threat spread speed, and impact range.
[0047] Step S202: The current edge agent uses the attack entry as the root node of the edge domain Bayesian network, the attack stage feature as the intermediate node of the edge domain Bayesian network, and the potential attack as the hidden variable node.
[0048] The current edge agent can define edge-domain Bayesian network nodes based on entity data in the data security knowledge graph (such as user roles, data assets, protection measures, risks, attack methods, vulnerabilities, asset value, etc.). Entity data can include, but is not limited to, network security, data security, information security, and privacy security data.
[0049] For example: the root node includes the attack entry point (such as kernel vulnerabilities, plaintext communication, weak passwords and low-level security algorithms, etc.); the intermediate node includes the attack stage characteristics, which include the attack stage (such as data detection, unauthorized use, etc.), the scope of impact (such as data leakage, service interruption, etc.) and the threat spread speed; the leaf node includes the risk level (high risk / medium risk / low risk).
[0050] Current edge agents can define conditional probabilistic relationships between nodes based on historical attack data and threat intelligence. For example, P(data breach | unencrypted sensitive data) = 0.90, indicating the probability of data breach when sensitive data is encrypted; and P(service interruption | peak attack traffic > 10Gbps) = 0.95, indicating the probability of service interruption when the peak attack traffic is > 10Gbps. Current edge agents can extract statistics such as vulnerability exploitation frequency and attack pattern distribution from the data security knowledge graph to initialize root node probabilities. They can also dynamically adjust probabilities based on real-time data reported by edge nodes (such as device exposure and log alert frequency).
[0051] Hidden variable nodes are introduced to detect potential attacks (such as zero-day vulnerabilities and exploits of historical algorithms by the latest cryptographic techniques). Current edge agents can estimate the probability of these hidden variable nodes using the Markov Chain Monte Carlo method (MCMC). For example, if a sudden increase in the frequency of device log alarms, A, is detected but no matching vulnerability is found in the data security knowledge graph, a hidden variable, H, is introduced to indicate a potential attack. Based on empirical evidence, the likelihood function P(A|H) is dynamically adjusted to 80% (assuming an 80% increase in A when H exists). Using the MCMC algorithm, samples are taken from the posterior distribution P(H|A), yielding a posterior probability of P(H|A) = 0.65. This indicates that given a sudden increase in A, the probability of H existing is 65%. Exposure evidence E (e.g., an open high-risk port) is introduced to further update the posterior probability. The new posterior probability, P(H|E,A), is 0.82. This indicates that given the simultaneous occurrence of A and E, the probability of H existing rises to 82%, triggering a zero-day vulnerability alert.
[0052] Step S203: Obtain evidence data of the current edge agent as the observation value of the edge domain Bayesian network, quantify the risk level of potential attacks, and upload the risk level as a local network parameter of the edge domain Bayesian network to the cloud.
[0053] When the current edge agent determines a potential attack, it obtains evidence data from the current edge agent. This evidence data includes static and dynamic evidence. Static evidence includes asset attributes (such as data classification and grading information, equipment, system architecture, etc.) and compliance policies (such as legal and regulatory provisions). Dynamic evidence includes real-time alerts (such as abnormal logins, sudden traffic increases, and abnormal data flows) and threat intelligence (such as CVE vulnerability scores).
[0054] The evidence data of the edge agent is mapped to the observation values of the edge domain Bayesian network. The edge domain Bayesian network can estimate at least the probability of the leaf node, namely the risk probability, based on the MCMC algorithm. It should be noted that the MCMC algorithm is an approximate algorithm and estimates the probability through random sampling. When the current edge agent determines that there is a potential attack, in order to more accurately quantify the risk level of the potential attack based on the evidence data of the current edge agent, the edge domain Bayesian network can use an exact inference algorithm such as the junction tree algorithm to determine the risk probability of the leaf node. When the risk probability of the leaf node is obtained, combined with the business impact, the risk probability of the leaf node is converted into a risk value Risk = P (risk occurrence) × asset value × impact coefficient; the risk level of the potential attack is determined based on the size of the risk value. The current edge agent uploads the risk level as a local network parameter of the edge domain Bayesian network to the cloud.
[0055] Step S204: The cloud updates the global network parameters of the cloud Bayesian network on the cloud based on the local network parameters of the edge domain Bayesian network through the blockchain.
[0056] In one embodiment, in step S204, the cloud updates the global network parameters of the cloud Bayesian network on the cloud side based on the local network parameters of the edge domain Bayesian network through the blockchain, specifically including: The first method is that the cloud learns the local network parameters of the edge domain Bayesian network through the teacher-student network, obtains the updated global network parameters of the cloud Bayesian network, and writes the update records of the local network parameters and the global network parameters into the blockchain.
[0057] Specifically, the current edge agent, acting as the student network in a teacher-student network, uses local data to train an edge-domain Bayesian network and outputs the local network parameters of the edge-domain Bayesian network. The cloud, acting as the teacher network in the teacher-student network, aggregates and distills the local network parameters of the edge-domain Bayesian network to obtain updated global network parameters of the cloud-based Bayesian network. The distillation process preserves key domain features (such as attack pattern weights and rule constraint logic) and filters out noisy data in the edge-domain Bayesian network. The cloud-based agent can write the updated global network parameters of the cloud-based Bayesian network to the blockchain, allowing the edge agent to retrieve the updated global network parameters of the cloud-based Bayesian network from the blockchain and update the local network parameters of the edge-domain Bayesian network.
[0058] In the second method, the current edge agent encrypts the local network parameters and uploads them to the cloud. The cloud aggregates the local network parameters of all edge domain Bayesian networks through federated learning to obtain the global network parameters of the cloud Bayesian network. The global network parameters are then distributed to all edge agents, and the update records of the global network parameters are written into the blockchain.
[0059] For example, the edge agent currently trains an edge Bayesian network based on local data, obtains the local network parameters of the edge Bayesian network, and uploads them to the cloud in encrypted form. The cloud aggregates the local network parameters of all edge Bayesian networks through federated learning to obtain the global network parameters of the cloud Bayesian network. This is then distributed to all edge agents to synchronize the edge Bayesian network. The cloud also writes updated records of the global network parameters (such as the CP table version and utility function weights) to the blockchain to ensure transparency and traceability.
[0060] It should be noted that the above two methods can be used separately or in combination to implement the updating of the global network parameters of the cloud-based Bayesian network based on the local network parameters of the edge domain Bayesian network, which is not limited here.
[0061] In one embodiment, when the current edge agent mines high-frequency attack features through the edge domain large model embedded with security rule constraints, the method for constructing a data security agent based on the large model further includes the following steps: Figure 3 As shown: Step S301: The current edge agent determines whether the high-frequency attack features belong to unknown potential attacks based on the local edge domain large model and data security knowledge graph. When the high-frequency attack features belong to unknown potential features, local incremental learning is triggered locally to generate new adversarial samples. The new adversarial samples are input into the adversarial sample library, and the attack stage features are extracted from the high-frequency attack features. The graph computing engine is used to construct a dynamic attack graph based on the attack stage features to quantify the level of new security threats and the speed of threat diffusion.
[0062] For example, when the current edge agent identifies high-frequency attack signatures (such as the behavior patterns of specific malicious IP addresses) as potential unknown attacks, it automatically triggers local model incremental learning, generates new adversarial examples in real time, and injects them into the adversarial example library. Based on the ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) framework, it extracts attack phase characteristics from multi-source data (logs, traffic, threat intelligence) and / or high-frequency attack signatures. A graph computing engine is used to construct a dynamic attack graph, quantifying the threat level (such as CVSS score) and the speed of threat spread.
[0063] Among them, the graph computing engine builds a dynamic attack graph, which can extract attack stage features from multi-source data and / or high-frequency attack features, represent and store attack stage features in graph form, use graph algorithms in real time to identify attack paths, associate attack events, and update the attack graph in real time to obtain a dynamic attack graph.
[0064] Step S302: Input the dynamic attack graph into the edge domain big model, and generate a local risk defense strategy based on the level of the newly added security threat and the threat diffusion speed.
[0065] For example, "isolated server A" is converted into a graph node, and its corresponding defense actions and impact range are obtained. Multimodal deduction is performed on the changes in its impact range and the losses suffered by data assets. The defense map is dynamically updated and the effects are evaluated, such as simulating the evolution of attacks (such as "after isolation, the attacker switches to server B"), thereby generating a local risk response strategy.
[0066] Step S303: Using the edge domain Bayesian network, determine whether the utility analysis result corresponding to the local risk defense strategy meets the business requirements, where the business requirements include at least one of service interruption risk, compliance risk, and service performance degradation risk.
[0067] Step S304: When the utility analysis results indicate that the local risk defense strategy meets the business requirements, the local network parameters of the edge domain Bayesian network are updated, and the dynamic attack graph, the optimized local risk defense strategy, and the updated local network parameters of the edge domain Bayesian network are uploaded to the cloud through the blockchain, so that the cloud updates the cloud Bayesian network.
[0068] Optionally, the edge agent can use the edge domain Bayesian network to model and analyze the utility of defense actions, select the optimal strategy, and calibrate the posterior probability of the selected optimal strategy, and use the calibrated optimal strategy as the local risk defense strategy.
[0069] The defense action modeling and effectiveness analysis specifically include: Define candidate defense actions, such as blocking IP addresses, isolating devices, starting backups, and strengthening algorithms; Associate each candidate defense action with an execution cost and effectiveness probability. The execution cost can be resource consumption or service interruption time, while the effectiveness probability can be the probability that executing the corresponding defense action can reduce the risk. For example, executing the IP blocking defense action can reduce the attack risk by 80%. Based on multi-objective optimization, a utility function is defined, and the utility analysis results of the defensive actions are determined based on the utility function. Multi-objective optimization can be risk reduction goal optimization and cost control goal optimization. The utility function can be U(a)=α*ΔRisk-β*Cost(a), where α represents the weight coefficient of the risk reduction goal and β represents the weight coefficient of the cost control goal. The two weight coefficients can be dynamically adjusted according to the priority of compliance requirements and business requirements.
[0070] For example, in an emergency business scenario in the financial industry, α uses a higher weight coefficient and β uses a lower weight coefficient; In the school homework grading system for the teaching industry, α uses a lower weight coefficient and β uses a higher weight coefficient. Since edge agents in the same industry have different locations in the data security agent system, the weight coefficients of α and β can also be adjusted accordingly.
[0071] Selecting the optimal strategy may include expanding the edge-domain Bayesian network into an influence diagram consisting of decision nodes and utility nodes to construct a causal relationship between defense actions and utility analysis results. For example, a real-time pruning search can be used to select the defense action combination that maximizes the expected utility analysis result as the optimal strategy. Pruning search is a method that eliminates potential optimal choices during the search process. Specifically, an initial set of defense action combinations can be generated, and the expected utility (e.g., cost, reliability) of each defense action combination in the initial set can be determined. Upper and / or lower bounds on the expected utility can be set. If a defense action combination does not meet the upper and / or lower bounds, the branch containing the defense action combination can be pruned. Furthermore, constraints can be set and the remaining defense action combinations can be determined to meet them. If not, the corresponding branch can be pruned, and the defense action combination corresponding to the remaining branch can be selected as the optimal strategy. For example, for APT attacks, the defense action combination of "isolating the device, enabling traffic analysis, and restricting permissions" can be selected. For false positive scenarios, the defense action combination of "logging only and manual review" can be selected.
[0072] The posterior probability calibration of the selected optimal strategy may include: Record the risk changes after the actual defense action is executed, such as whether the attack is terminated or whether losses occur; reversely update the local network parameters in the marginal Bayesian network. For example, if the attack stops after the actual execution of the "block IP" defense action, increase P(attack mitigation | block IP), that is, the probability of attack mitigation after executing the blocking IP defense action; if the actual execution of the "isolate device" defense action causes business interruption, increase the cost coefficient β of the defense action.
[0073] It should be noted that posterior probability calibration of the selected optimal strategy can reversely update the local network parameters in the edge domain Bayesian network, further updating the dynamic attack map and optimizing the local risk defense strategy. The updated dynamic attack map, optimized local risk defense strategy, and updated local network parameters of the edge domain Bayesian network are uploaded to the cloud via the blockchain, allowing the cloud to update the cloud Bayesian network.
[0074] Step S305: When the utility analysis result indicates that the local risk defense strategy does not meet the business requirements, a consensus request is initiated to the cloud.
[0075] Optionally, when the utility analysis results indicate that the local risk defense strategy does not meet business requirements, one way is to reversely update the local network parameters in the edge domain Bayesian network, update the dynamic attack graph, or optimize the local risk defense strategy. Another optional way is to initiate a consensus request to the cloud, and the cloud will formulate a more reasonable risk response strategy in response to the consensus request.
[0076] In one embodiment, when the cloud-customized risk response strategy is not a high-risk defense operation, step S101 schedules at least one other edge agent to perform a dynamic defense operation, including: performing traffic cleaning on the current edge agent. The traffic cleaning on the current edge agent specifically includes: Transfer the business of the current edge agent to the idle edge agent, Obtain the attack phase characteristics of the current edge agent; Perform at least one of the following defensive operations on the current edge agent based on the attack phase characteristics: Algorithm reinforcement, IP blocking, traffic analysis, and permission restriction.
[0077] In one embodiment, scheduling at least one other edge agent to perform dynamic defense operations in step S101 includes: The cloud generates lightweight on-chain attack fingerprints based on the risk level and security threat factors of new unknown attacks; The cloud determines at least one other edge agent involved in the restoration cross-domain attack chain in the blockchain through the threat diffusion speed in the security threat elements, pushes the on-chain attack fingerprint and risk response strategy to each edge agent involved in the restoration cross-domain attack chain, and instructs each edge agent involved in the restoration cross-domain attack chain to perform defense operations in accordance with the risk response strategy.
[0078] The data security agent system deploys a private blockchain based on Hyperledger Fabric, with cloud-based and multiple edge agents serving as nodes. Security threat factors (such as kernel vulnerabilities, plaintext communications, weak passwords, low-level security algorithms, attack phases, threat spread speed, and impact range) are recorded on the blockchain, with Merkle trees used for integrity verification. The blockchain also defines on-chain rules (such as requiring verification by ≥3 nodes), access permissions (e.g., only nodes can query specific threat fingerprints), and supports on-demand subscription intelligence push.
[0079] For example, the cloud generates a lightweight on-chain attack fingerprint based on the risk level and security threat factors of the new unknown attack. The on-chain attack fingerprint is a unique identifier of the new unknown attack, facilitating rapid matching and dissemination. The cloud uses the threat diffusion speed analysis within the security threat factors to determine the potential scope of the new unknown attack within the blockchain network. Based on the spatiotemporal correlation of the on-chain attack fingerprint, the cloud restores the cross-domain attack chain (such as the ransomware propagation path), identifies other edge agents involved in the restored cross-domain attack chain within the blockchain, and instructs each edge agent involved in the restored cross-domain attack chain to execute defensive actions according to the risk response strategy.
[0080] In one embodiment, the security rule constraints in step S201 include initial role assignment rules determined based on high-frequency attack characteristics, wherein the initial role assignment rules specifically include: The cloud dispatches at least three edge agents based on the type of attack entry uploaded by the current edge agent to perform asset analysis, threat hunting, and threat assessment respectively. The cloud formulates an optimized risk response strategy based on the asset analysis results, threat hunting results, and threat assessment results fed back by at least three edge agents, and distributes the optimized risk response strategy to the dispatched edge agents for execution. Among them, threat hunting includes detecting potential attacks through data security knowledge graphs, asset analysis results include the scale of data assets involved in potential attacks, and threat assessment includes detecting and assessing the risk level of potential attacks.
[0081] Optionally, a dynamic role orchestration algorithm can be used, combining heuristic rules with reinforcement learning to adaptively assign five edge agent roles (such as data asset analysis, threat hunting, threat assessment, policy optimization, and execution edge agents) based on real-time situational awareness, forming an optimal collaborative link. Heuristic rules are a decision-making technique based on empirical judgment.
[0082] Adaptive allocation of five edge agent roles specifically involves: initial role assignment based on attack event characteristics and predefined strategies in a heuristic rule base; dynamic adjustment of the collaborative topology based on at least the state space and the corresponding action space; and balanced allocation and failover of edge agents. The heuristic rule base can be a pre-generated role allocation scheme based on previous real-time status, such as threat level and resource load, to facilitate subsequent initial role assignment based on attack event characteristics and the role allocation scheme in the heuristic rule base.
[0083] Optionally, the attack event characteristics include at least one of the attack type, impact scope, and confidence level. The rule library pre-defined policies include different initial role combinations corresponding to different attack scenarios. For example, if the attack scenario is data leakage risk, the initial role combination is: data asset analysis → threat hunting → threat assessment → policy optimization → execution edge agent.
[0084] Optionally, the state space is: S = {{attack stage}, {node load}, {business priority}, {historical collaboration efficiency}}, for example, S = {{lateral movement}, {edge node load 70%}, {core database}, {previous task delay 200ms}}. Optionally, the action space includes at least one of the following: adding, deleting, or replacing edge agents; adjusting the initial role combination; adding parallel edge agents to respond to attacks; or replacing high-load edge agents with backup execution edge agents. The high-load edge agent may be the edge agent that detected the attack.
[0085] Furthermore, edge agents are evenly distributed and failover is performed. For example, the resource utilization of each edge agent is monitored. If the delay of multiple consecutive tasks exceeds 100ms, a redistribution is triggered, for example, migrating tasks from a high-loaded edge agent to an idle edge agent. After redistribution, the status of the edge agents is uploaded to the cloud, allowing the cloud to adjust risk response strategies.
[0086] In the embodiment of the present application, edge agents may include but are not limited to embedded devices, industrial computers, servers, etc. The cloud and edge agents may use the same hardware or different hardware, which is not limited here.
[0087] An embodiment of the present application provides an electronic device, including: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method for constructing the data security intelligent body based on the large model is implemented.
[0088] The embodiment of the present application provides a device for constructing a data security intelligent entity based on a large model, comprising processor, A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method for constructing a data security intelligent body based on a large model is implemented.
[0089] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), but may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0090] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0091] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are fully or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via infrared, microwave, or other means. A computer-readable storage medium can be any available medium accessible by a computer, or a data storage device such as a server or data center that contains a collection of one or more available media. Available media can include magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), solid-state drives, etc.
[0092] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0093] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0094] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0095] The communication interface is used for communication between the above-mentioned electronic device and other devices. The memory may include a random access memory (RAM) and may also include a non-volatile memory (NVM), such as at least one disk storage. Optionally, the memory may also be at least one storage device located away from the aforementioned processor. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0096] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0097] Each embodiment in this specification is described in a related manner. Similar portions between the embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so their description is relatively simple. For related portions, refer to the description of the method embodiments.
[0098] The above are only preferred embodiments of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.
Claims
1. A method for constructing a data security agent based on a large model, characterized in that: The method is applied to a data security agent system, the system comprising: a plurality of edge agents and a cloud, the method comprising: When the current edge agent detects an unknown potential attack through the edge domain large model embedded with security rule constraints, it initiates a consensus request to the cloud, which includes security threat elements and evidence data; The cloud calls at least three other edge agents to determine that the unknown potential attack is a real new unknown attack through voting based on the risk response strategy inference rule library stored in the blockchain and the consensus request; The cloud performs incremental learning based on the security threat factors and the evidence data to generate new adversarial samples. Taking the security threat factors, the evidence data and the new adversarial samples as input, the cloud-based Bayesian network quantifies the risk level of the new unknown attack. Combining the risk level and the security threat factors, a risk response strategy is formulated, and at least one other edge intelligent agent is dispatched to perform dynamic defense operations.
2. The method according to claim 1, wherein Call at least three other edge agents to vote to determine that the unknown potential attack is a real new unknown attack, including: Call at least three other edge agents to vote through the Practical Byzantine Fault Tolerance (PBFT) algorithm, and when two-thirds of the at least three other edge agents determine that the unknown potential attack is a real attack, determine that the unknown potential attack is a real new unknown attack; When the risk response strategy customized by the cloud includes a high-risk defense operation, the cloud notifies multiple edge agents scheduled in the risk response strategy to execute the high-risk defense operation through alliance chain signature authorization, and the high-risk defense operation includes: at least one of: global service interruption and global network disconnection.
3. The method according to claim 1, wherein The method further comprises: When the current edge agent mines high-frequency attack features through the edge domain large model, it determines that there is a potential attack. The current edge agent performs lightweight feature encoding on the high-frequency attack features to generate security threat elements of the potential attack, wherein the security threat elements include: attack entry, attack stage features, The current edge agent uses the attack entry as the root node of the edge domain Bayesian network, the attack stage feature as the intermediate node of the edge domain Bayesian network, and the potential attack as the hidden variable node; Obtaining evidence data of the current edge agent as an observation value of the edge domain Bayesian network, quantifying the risk level of the potential attack, and uploading the risk level as a local network parameter of the edge domain Bayesian network to the cloud; The cloud updates the global network parameters of the cloud Bayesian network on the cloud based on the local network parameters of the edge domain Bayesian network through the blockchain.
4. The method according to claim 3, wherein The cloud updates the global network parameters of the cloud Bayesian network of the cloud based on the local network parameters of the edge domain Bayesian network through the blockchain, including: The cloud learns the local network parameters of the edge domain Bayesian network through the teacher-student network to obtain updated global network parameters of the cloud Bayesian network, and writes the updated records of the local network parameters and the global network parameters into the blockchain, and / or, The current edge agent encrypts the local network parameters and uploads them to the cloud. The cloud aggregates the local network parameters of all edge domain Bayesian networks through federated learning to obtain the global network parameters of the cloud Bayesian network. The global network parameters are sent to all edge agents, and the update records of the global network parameters are written into the blockchain.
5. The method according to claim 3, wherein The security rule constraints include initial role allocation rules determined according to high-frequency attack characteristics, The initial role allocation rules include: The cloud dispatches at least three edge agents according to the type of attack entry uploaded by the current edge agent to perform asset analysis, threat hunting and threat assessment respectively. The cloud formulates an optimized risk response strategy based on the asset analysis results, threat hunting results and threat assessment results fed back by the at least three edge agents, and distributes the optimized risk response strategy to the scheduled edge agents for execution, wherein threat hunting includes detecting potential attacks through a data security knowledge graph, the asset analysis results include the scale of data assets involved in the potential attack, and the threat assessment includes detecting and assessing the risk level of the potential attack.
6. The method according to claim 2, wherein When the risk response strategy customized in the cloud is a non-high-risk defense operation, the scheduling at least one other edge agent to perform a dynamic defense operation includes: performing traffic cleaning on the current edge agent; The step of performing traffic cleaning on the current edge agent includes: Transfer the business of the current edge agent to the idle edge agent, Obtaining attack phase characteristics of the current edge agent; Perform at least one of the following defense operations on the current edge agent according to the attack phase characteristics: Algorithm reinforcement, IP blocking, traffic analysis, and permission restriction.
7. The method according to claim 1, wherein The scheduling of at least one other edge agent to perform a dynamic defense operation includes: The cloud generates a lightweight on-chain attack fingerprint based on the risk level and security threat factors of the new unknown attack; The cloud determines at least one other edge agent involved in the restored cross-domain attack chain in the blockchain through the threat diffusion speed in the security threat elements, pushes the on-chain attack fingerprint and risk response strategy to each edge agent involved in the restored cross-domain attack chain, and instructs each edge agent involved in the restored cross-domain attack chain to perform defense operations in accordance with the risk response strategy.
8. The method according to claim 3, wherein When the current edge agent mines high-frequency attack features through the edge domain large model embedded with security rule constraints, the method further includes: The current edge agent determines whether the high-frequency attack feature belongs to an unknown potential attack based on the local edge domain large model and data security knowledge graph. If the high-frequency attack feature belongs to an unknown potential feature, local incremental learning is triggered locally to generate a new adversarial sample. The new adversarial sample is input into the adversarial sample library, and the attack stage features are extracted from the high-frequency attack feature. A graph computing engine is used to construct a dynamic attack graph based on the attack stage features to quantify the level of the new security threat and the threat diffusion speed. Inputting the dynamic attack graph into the edge domain big model, and generating a local risk defense strategy based on the level of the newly added security threat and the threat diffusion speed; Determining, by means of an edge domain Bayesian network, whether a utility analysis result corresponding to the local risk defense strategy meets business requirements, wherein the business requirements include at least one of a service interruption risk, a compliance risk, and a service performance degradation risk; When the utility analysis result indicates that the local risk defense strategy meets the business requirements, the local network parameters of the edge domain Bayesian network are updated, and the dynamic attack graph, the optimized local risk defense strategy, and the updated local network parameters of the edge domain Bayesian network are uploaded to the cloud via the blockchain, so that the cloud updates the cloud Bayesian network; When the utility analysis result indicates that the local risk defense strategy does not meet business requirements, a consensus request is initiated to the cloud.
9. The method according to claim 8, wherein Using the edge domain Bayesian network, determine whether the utility analysis result corresponding to the local risk defense strategy meets the business requirements, including: The edge agent is capable of defining candidate defensive actions using an edge domain Bayesian network; performing a utility analysis on the defensive action; Through real-time pruning search, the defensive action combination that maximizes the expected utility analysis result is selected as the optimal strategy; Based on the update of the edge domain leaf-Bassian network, the optimal strategy is calibrated for posterior probability, and the calibrated optimal strategy is used as the local risk defense strategy.
10. A device for constructing a data security intelligent agent based on a large model, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Edge intelligent DDoS positive defense method and system
CN118101267A
A deep protocol network security adjustment method based on dynamic defense
CN119743322A
Network security intelligent management and control system based on big data
CN120110786A
Internet of Things data security collaborative defense system based on edge computing
CN120263548A