A security situation awareness and dynamic defense system based on multi-dimensional data association
Through a security situation awareness and dynamic defense system based on multi-dimensional data correlation, the problems of misjudgment and network instability in the processing of multi-dimensional heterogeneous situation data in cross-border financial clearing networks have been solved, and stable defense of core routing nodes and business continuity assurance have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHAANXI BENNDACHANG NETWORK TECHNOLOGY CO LTD
- Filing Date
- 2026-04-23
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies struggle to effectively handle multidimensional and heterogeneous situational data in cross-border financial clearing networks, leading to misjudgment risks exceeding preset tolerance levels. They also fail to accurately identify network conditions and lack a tiered defense mechanism that balances topological stability, business continuity, and policy delivery efficiency, making them prone to network instability.
A security situation awareness and dynamic defense system with multidimensional data association is adopted. The data acquisition module acquires multidimensional heterogeneous situation data, the core processing module extracts network topology features and quantifies defense trust entropy, the result generation module outputs risk contribution prediction value, and the feedback adaptive module generates dynamic micro-segmentation or static route cleaning strategy based on the prediction value to optimize the defense strategy and ensure the continuous service availability of core routing nodes.
Accurately identify false flag alarm storm interference data, reduce the risk of misjudgment, accurately reflect the degree of network risk, reduce the number of nodes undergoing topology changes, avoid network instability caused by frequent defense actions, and improve the service assurance capabilities of core routing nodes.
Smart Images

Figure CN122093192B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security and communication network protection technology, specifically a security situation awareness and dynamic defense system based on multi-dimensional data association. Background Technology
[0002] In the security protection of cross-border financial clearing networks, the ability to perceive network risks based on multi-source situational data and generate defense strategies in a coordinated manner has gradually become an important technical means to ensure the continuous availability of core clearing links.
[0003] Currently, most security protection solutions support the collection and analysis of traffic logs, terminal behavior, access control records, or alarm information. If it is necessary to take into account both risk prediction and policy adaptive adjustment under complex network topology, it is usually necessary to correlate multiple types of heterogeneous data and dynamically adjust defense actions in combination with changes in network status in order to achieve continuous protection of core business links.
[0004] However, when processing multi-dimensional heterogeneous situational data involving cross-regional and cross-latency scenarios with noise, the above methods are easily affected by fragmented, highly realistic abnormal traffic and alarm storms, leading to inaccurate identification of the actual network status and a risk of misjudgment exceeding the preset tolerance range. Furthermore, existing processing methods often focus more on anomaly detection itself, failing to simultaneously characterize structural instability factors such as access control resource pressure, routing table oscillations, and service-oriented feedback. This results in an inability to effectively reflect the degree of risk of the network approaching the continuous service interruption boundary, leading to insufficient basis for policy switching. Moreover, when risks escalate or management channels become congested, related solutions typically lack a hierarchical defense and convergence mechanism that balances topology stability, service continuity, and policy issuance efficiency. This can easily cause network instability due to overly frequent defense actions, resulting in weak capabilities to ensure continuous service at core routing nodes. Summary of the Invention
[0005] The purpose of this invention is to provide a security situation awareness and dynamic defense system based on multi-dimensional data association, solving the following technical problems:
[0006] It avoids the amplification of network topology oscillations caused by the repeated changes in isolation boundaries of the defense system induced by highly realistic false flag alarm storms. It is also easier to achieve a dynamic adaptive balance between threat interception and structural stability. When the risk is controllable, the system can give full play to the fine-grained advantage of dynamic micro-isolation. When it approaches the failure boundary of continuous service interruption, it can actively trigger policy degradation to reduce the intensity of action and prioritize the continuous service availability of core routing nodes.
[0007] The objective of this invention can be achieved through the following technical solutions:
[0008] A security situation awareness and dynamic defense system based on multi-dimensional data correlation includes:
[0009] The data acquisition module is used to acquire multidimensional heterogeneous situational data reported by each service node in the communication network, wherein the communication network includes the service nodes and the core routing node, and the multidimensional heterogeneous situational data includes traffic logs and terminal behavior characteristics with network latency characteristics and noise attributes.
[0010] The core processing module is used to extract network topology features based on the multidimensional heterogeneous situational data and the preset communication protocol feature association model, and to quantitatively calculate the defense trust entropy of the current network topology. The defense trust entropy is used to characterize the tolerance of the service node to the defense strategy and the vulnerability index of the network topology features.
[0011] The result generation module is used to output a predicted risk contribution value for the service blocking risk of the core routing node based on the defense trust entropy and the preset failure boundary prediction model.
[0012] The feedback adaptive module is used to determine the logical relationship between the predicted risk contribution value and the preset danger threshold.
[0013] If the predicted risk contribution value is less than the danger threshold, a first defense strategy is generated, wherein the first defense strategy includes a dynamic micro-segmentation strategy based on multi-agent deep reinforcement learning.
[0014] If the predicted risk contribution value is greater than or equal to the danger threshold, a policy degradation mechanism is triggered, and a second defense strategy is generated. The second defense strategy includes a static route cleaning strategy, which is used to converge the current network topology to a preset baseline routing path to reduce the number of network topology change nodes.
[0015] Preferably, after the data acquisition module acquires the multidimensional heterogeneous situational data reported by each service node in the communication network, the core processing module is further used for:
[0016] Based on a preset abnormal signature feature library, the multidimensional heterogeneous situational data is matched to separate fragmented traffic data with abnormal signatures.
[0017] Extract the communication frequency characteristics and protocol simulation characteristics of the fragmented traffic data;
[0018] If the communication frequency characteristic is greater than a preset frequency threshold and the protocol simulation characteristic is greater than a preset simulation threshold, then the fragmented traffic data is determined to be false flag alarm storm interference data; otherwise, it is determined to be valid alarm data.
[0019] The false flag alarm storm interference data is removed from the multidimensional heterogeneous situation data to generate purified multidimensional heterogeneous situation data, and the purified multidimensional heterogeneous situation data is input into the communication protocol feature association model.
[0020] Preferably, the core processing module quantifies and calculates the defense trust entropy of the current network topology, specifically for:
[0021] Collect the current dynamic access control list full load rate and routing table update frequency of the core routing node;
[0022] Collect the current Transmission Control Protocol timeout retransmission storm index of the service node;
[0023] The dynamic access control list full load rate, the routing table update frequency, and the transmission control protocol timeout retransmission storm index are each assigned a preset weight, and then weighted and summed to calculate the defense trust entropy.
[0024] Preferably, the result generation module, based on the defense trust entropy and the preset failure boundary prediction model, outputs a predicted risk contribution value for the service blocking risk of core routing nodes in the communication network, specifically used for:
[0025] Input the defense trust entropy into the failure boundary prediction model;
[0026] Calculate the probability value of continuous communication interruption occurring within a preset time window for the network topology characteristics;
[0027] The probability value is converted into the risk contribution prediction value through a preset mapping function and then output.
[0028] Preferably, the feedback adaptive module generates the first defense strategy, specifically for:
[0029] Based on the multidimensional heterogeneous situational data, malicious traffic features and legitimate business traffic features are extracted; the first defense strategy is generated by a preset multi-agent deep reinforcement learning model, wherein the preset multi-agent deep reinforcement learning model includes multiple agents deployed in the terminal domain, regional aggregation domain and core gateway domain respectively, and the action set of the multiple agents includes session-level isolation, service segment rate limiting and bypass traffic diversion;
[0030] The malicious traffic features and the legitimate business traffic features are input into a multi-agent deep reinforcement learning model;
[0031] With the joint optimization objectives of maximizing the malicious traffic interception rate and maximizing the legitimate business network throughput, the dynamic micro-segmentation strategy is output.
[0032] Preferably, the feedback adaptive module generates the second defense strategy, specifically for:
[0033] In response to the triggering of the policy downgrade mechanism, the issuance of the first defense policy is suspended;
[0034] Extract the baseline routing table data of the core routing node;
[0035] Based on the baseline routing table data, a static route cleaning strategy is generated, wherein the execution priority of the static route cleaning strategy is higher than that of the dynamic micro-segmentation strategy, and the static route cleaning strategy is used to reduce the defense trust entropy.
[0036] Preferably, the system further includes a policy distribution module, which is used for:
[0037] Obtain the in-band management channel bandwidth utilization rate of the communication network;
[0038] Determine whether the bandwidth utilization rate is greater than a preset congestion threshold:
[0039] If the congestion threshold is exceeded, the first defense strategy or the second defense strategy is compressed, and the delivery operation is executed based on the compressed instructions.
[0040] If the value is less than or equal to the congestion threshold, the distribution operation is executed directly.
[0041] Preferably, the communication network is a cross-border financial clearing network, the business node is a financial clearing terminal, and the core routing node is a clearing gateway node.
[0042] The beneficial effects of this invention are:
[0043] 1. This invention extracts the communication frequency characteristics and protocol simulation characteristics of fragmented traffic data to accurately identify and remove false flag alarm storm interference data from multidimensional heterogeneous situational data; this mechanism effectively solves the problem that the system is susceptible to interference from fragmented, highly realistic abnormal traffic, improves the purity of input data and the accuracy of identifying the real network status, and significantly reduces the risk of misjudgment.
[0044] 2. This invention calculates the defense trust entropy by weighting the dynamic access control list full load rate, routing table update frequency, and transmission control protocol timeout retransmission storm index, and outputs the risk contribution prediction value by combining the failure boundary prediction model. This mechanism successfully characterizes the structural instability factors of the network, accurately reflects the risk level of the network approaching the continuous service blocking boundary, and provides a quantitative basis for the dynamic switching of defense strategies.
[0045] 3. This invention adaptively switches between dynamic micro-segmentation strategy and static route cleaning strategy based on risk prediction value, effectively reducing the number of topology change nodes under high risk, and compressing and issuing defense strategy instructions when the in-band management channel is congested. This mechanism avoids network oscillations caused by frequent defense actions, takes into account topology stability and policy issuance efficiency, and significantly improves the ability to guarantee continuous services of core routing nodes. Attached Figure Description
[0046] The invention will now be further described with reference to the accompanying drawings.
[0047] Figure 1 This is a schematic diagram of a security situation awareness and dynamic defense system based on multidimensional data association, provided as an embodiment of this application. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] Please see Figure 1 A security situation awareness and dynamic defense system based on multidimensional data association includes: a data acquisition module, used to acquire multidimensional heterogeneous situation data reported by each service node in the communication network, wherein the communication network includes the service node and the core routing node, wherein the multidimensional heterogeneous situation data includes traffic logs and terminal behavior features with network latency characteristics and noise attributes;
[0050] The core processing module is used to extract network topology features based on the multidimensional heterogeneous situational data and a preset communication protocol feature association model, and to quantitatively calculate the defense trust entropy of the current network topology. The defense trust entropy is used to characterize the tolerance of the service node to the defense strategy and the vulnerability index of the network topology features. The result generation module is used to output the risk contribution prediction value of the service blocking risk of the core routing node based on the defense trust entropy and a preset failure boundary prediction model.
[0051] The feedback adaptive module is used to determine the logical relationship between the predicted risk contribution value and the preset danger threshold: if the predicted risk contribution value is less than the danger threshold, a first defense strategy is generated, wherein the first defense strategy includes a dynamic micro-segmentation strategy based on multi-agent deep reinforcement learning; if the predicted risk contribution value is greater than or equal to the danger threshold, a strategy degradation mechanism is triggered, and a second defense strategy is generated, wherein the second defense strategy includes a static route cleaning strategy, and the static route cleaning strategy is used to converge the current network topology to a preset baseline routing path to reduce the number of nodes changing the network topology.
[0052] This embodiment provides a security situation awareness and dynamic defense mechanism based on multi-dimensional data association; specifically, the system is deployed in a cross-regional clearing network, which consists of multiple business nodes and at least one core routing node; the business nodes can be regional clearing front-end machines, message access terminals, risk control audit terminals or reconciliation processing terminals, and the core routing node can be a backbone gateway carrying the clearing main link;
[0053] During daily operation, the system continuously receives traffic logs, session status, terminal behavior trajectories, access control hit records, and alarm summaries reported by each node, and aggregates these data from different sources, arriving at different times, and with different noise levels into multidimensional heterogeneous situational data. It should be noted that the preset danger threshold, preset frequency threshold, preset simulation threshold, and preset congestion threshold involved in this system are not fixed absolute values, but are dynamically calculated and set by the system during the initialization phase based on the historical baseline data of the communication network.
[0054] Specifically, the system extracts data from networks in near-terminal conditions where there are no attacks and services are operating stably. The distribution characteristics of relevant indicators within a natural day, such as normal distribution or long-tail distribution, are used as the upper boundary of the confidence interval covering normal business fluctuations as the initial threshold. In the actual operation of the system, the operation and maintenance personnel can also manually fine-tune the above thresholds in combination with the upper limit of the hardware performance stress test of specific financial clearing nodes, so as to ensure that each threshold can truly reflect the current network topology's tolerance limit.
[0055] Specifically, the data acquisition module does not only receive a single result of whether it is abnormal, but also retains information such as the arrival delay of the data, the reporting timestamp offset, the degree of log fragmentation, and the terminal behavior context;
[0056] The technical principle is that in a wide area clearing network, there is often link jitter between edge sites and central scheduling nodes. If all data is regarded as generated at the same time, it is easy to mistake delayed normal records for sudden anomalies, thereby inducing excessive defense.
[0057] To this end, the acquisition module can be encapsulated layer by layer according to the original record layer, node aggregation layer and regional aggregation layer. The original record layer retains basic traffic attributes such as five-tuple, session direction and connection maintenance time. The node aggregation layer supplements terminal behavior characteristics such as terminal process calls, login state switching and certificate handshake behavior. The regional aggregation layer adds link delay label and noise confidence label.
[0058] After receiving this data, the core processing module calls the preset communication protocol feature association model to extract network topology features; the communication protocol feature association model realizes the association and matching of sequential relationships by constructing a finite state machine or a communication sequence knowledge graph;
[0059] The correlation here is not simply comparing whether fields are the same, but rather identifying whether the network structure is in a stable state based on the connection between the preceding and following parts of the communication process;
[0060] Let's take a specific network communication scenario as an example: Assume that the message fragments from three service nodes are denoted as the first message fragment, the second message fragment, and the third message fragment, respectively. The first message fragment represents the terminal initiating a session request to the regional front-end machine, the second message fragment represents the regional front-end machine forwarding an authentication message to the core gateway, and the third message fragment represents the core gateway returning a routing response. If the normal association is the first message fragment → the second message fragment → the third message fragment, then the system internally forms a set of directional connection descriptions.
[0061] If the actual data collected is that multiple abnormal mimicry fragments are derived from the first message fragment, such as derivative fragment A, derivative fragment B, and derivative fragment C, and these fragments correspond to multiple conflicting subsequent responses, it indicates that the policy switching or attack disturbance in the network has disrupted the original communication loop.
[0062] Based on this, the core processing module extracts topological features such as node connection stability, forwarding path change frequency, critical link detour count, and control plane oscillation trend. Among them, the node connection stability is quantified by extracting the relative ratio of the number of sessions that have not experienced abnormal interruptions within a preset time period to the total number of established sessions; the forwarding path change frequency is obtained by statistically analyzing the number of routing table hops of core routing nodes per unit time.
[0063] The number of critical link detours is obtained by extracting the number of redundant hop nodes in the current actual data packet forwarding path compared to the preset baseline routing path; the control plane oscillation trend is characterized by calculating the time series variance of the number of rule addition and deletion instructions in the dynamic access control list within multiple consecutive time windows;
[0064] Based on the above, the system further quantifies and calculates the defense trust entropy; the defense trust entropy is used to characterize the network state in the following two dimensions: one is the tolerance of business nodes to the current defense strategy, and the other is the vulnerability of the network topology under continuous regulation; if business nodes frequently encounter session disconnection, repeated rerouting of authentication links, and repeated handshakes for whitelisted services, it indicates that the nodes are beginning to have difficulty maintaining their expectations for the stability of the current defense scheduling.
[0065] If the core path switches frequently, the routing table is updated frequently, and access control rules are added or deleted frequently, it indicates that the topology itself is inherently fragile in engineering terms. The system integrates these characteristics into a single defense trust entropy value, which serves as the core input for subsequent predictions.
[0066] Based on the entropy value and the failure boundary prediction model, the result generation module outputs the risk contribution prediction value of the core routing node service interruption risk. Here, the failure boundary corresponds to the risk trend of continuous interruption of core clearing services for more than a preset duration. The prediction model does not require direct judgment of whether failure will occur immediately, but rather measures the degree of contribution of the current network state to future continuous interruption events. Specifically, when the network can still forward services but there are already retransmission backlogs, control plane jitter, and unexpected route detours, the system can give a high risk contribution value in advance, instead of waiting for the link to be interrupted before responding.
[0067] The feedback adaptive module selects the defense path based on the relationship between the predicted value and the danger threshold; when the predicted value is lower than the danger threshold, the system has structural resilience and can adopt a dynamic micro-segmentation strategy to implement more granular access constraints on suspicious business flows, terminal groups or regional links.
[0068] At this point, the multi-agent deep reinforcement learning model takes the perspective of different nodes or regions and makes coordinated decisions on how to intercept malicious traffic as much as possible while minimizing the impact on legitimate clearing services. If the predicted value reaches or exceeds the danger threshold, it indicates that the network is close to structural instability. At this point, continuing to pursue fine-grained dynamic optimization is more likely to exacerbate the oscillation. Therefore, the policy degradation mechanism is triggered, the dynamic changes are suspended, and a static route cleaning policy is generated instead.
[0069] Static route cleaning strategy achieves critical path convergence with fewer topology change nodes. Its core goal is not to pursue the finest-grained blocking, but to prioritize ensuring that the main cleaning channel does not become unstable due to the defense action itself.
[0070] To avoid ambiguity in the reference of terms in different paragraphs, in this article, the terms "clearing main link," "clearing main channel," "main path," and "core clearing channel" all refer to the same type of backbone forwarding path that carries the continuous forwarding of key services of core routing nodes, unless otherwise specified.
[0071] The core gateway, backbone gateway, and headquarters clearing gateway are all used as scenario-based examples of core routing nodes unless otherwise specified; the dynamic micro-segmentation strategy and the first defense strategy are inclusive, with the former being the core strategy content of the latter;
[0072] The static route cleaning strategy and the second defense strategy are also inclusive, with the former being the core strategy of the latter; the aforementioned uniformity of terminology is only used to eliminate differences in description and does not change the scope of each technical feature.
[0073] Furthermore, if some of the collected node data suffers from packet loss, timestamp drift, or reporting interruption, the system can first maintain the topological baseline within the most recent stable window and mark the missing nodes as insufficiently observed, thus avoiding their direct inclusion in high-risk inferences.
[0074] If the core processing module finds that there is a lack of sufficient protocol connection between the current data, such as only isolated fragments that cannot restore the basic communication chain, the system will not immediately trigger large-scale dynamic isolation, but will first narrow it down to the vicinity of the known anomaly source for local defense; if the risk contribution is close to the preset critical judgment range of the threshold, the system can enter a short observation state, prioritize postponement of policy changes for secondary nodes, and first protect the core routing node and its adjacent links to reduce network shock caused by false triggers.
[0075] For example, during peak operating hours of a cross-border financial clearing network, clearing messages are continuously exchanged between terminals in the Asian region, front-end servers in the European region, and the headquarters clearing gateway; the system receives logs reported from multiple regions, some of which arrive late due to cross-border link delays, while some fragmented alarms indicate that some terminals are attempting abnormal handshakes.
[0076] After the core processing module correlated these records, it found that although the malicious probe traffic had increased, the main clearing link remained basically stable. Therefore, the defense trust entropy was within a controllable range, and the risk contribution output by the result generation module was lower than the danger threshold. As a result, the system issued a dynamic micro-isolation strategy to implement fine isolation only for suspicious source terminals and adjacent business segments.
[0077] As the attack continued to evolve, the core gateway began to experience continuous rule table jitter, repeated reconstruction of legitimate liquidation connections, and retransmission accumulation in multiple regional terminals. At this point, the system reassessed and determined that the risk contribution level had risen above the danger threshold, thus triggering policy downgrading. It stopped further expanding the scope of dynamic micro-segmentation and instead generated a static route cleaning policy based on the stable baseline route, retaining only the few changes necessary for the main liquidation path, thereby suppressing topology oscillations.
[0078] The purpose of this step is to manage threat interception and structural stability within the same scheduling framework, enabling the system to exert the flexibility of dynamic defense when risks are controllable, and to proactively reduce the intensity of actions when approaching the failure boundary, thereby prioritizing the continuous service availability of core routing nodes.
[0079] In this embodiment of the invention, after the data acquisition module acquires the multidimensional heterogeneous situational data reported by each service node in the communication network, the core processing module is further used to: match the multidimensional heterogeneous situational data based on a preset abnormal signature feature library, so as to separate out fragmented traffic data with abnormal signatures.
[0080] Extract the communication frequency characteristics and protocol simulation characteristics of the fragmented traffic data; if the communication frequency characteristics are greater than a preset frequency threshold and the protocol simulation characteristics are greater than a preset simulation threshold, then the fragmented traffic data is determined to be false flag alarm storm interference data; otherwise, it is determined to be valid alarm data.
[0081] The false flag alarm storm interference data is removed from the multidimensional heterogeneous situation data to generate purified multidimensional heterogeneous situation data, and the purified multidimensional heterogeneous situation data is input into the communication protocol feature association model.
[0082] This embodiment provides a purification and processing mechanism for false flag alarm storms; specifically, in the scenario of continuous operation of the aforementioned cross-border financial clearing network, when the system relies solely on the original multidimensional heterogeneous situational data for correlation analysis, it is easily affected by highly realistic fragmented traffic.
[0083] This type of traffic does not directly impact the throughput of the main business, but rather manifests as tiny communication fragments that have a certain similarity to both abnormal communication characteristics and the baseline of legitimate business communication, inducing the system to repeatedly change the micro-isolation boundary, thereby amplifying topology fluctuations; therefore, before entering the communication protocol feature association model, a data purification step for false flag alarm storms should be added.
[0084] Specifically, the core processing module first calls the preset abnormal signature feature library to perform an initial screening of the collected data; the abnormal signatures here may include features such as atypical handshake sequence, abnormal certificate fields, abnormal port access rhythm, unbalanced packet length distribution, and fake heartbeat mode.
[0085] All of the above features are generated based on rule matching thresholds. For example, the abnormal port access rhythm is manifested as the frequency of port probes initiated by the same source IP to the core routing node within a preset short window exceeding the upper limit of the variance of the baseline frequency. The imbalance in message length distribution is manifested as the deviation of the payload length of application layer messages from the preset expected value of the message length of the standard clearing protocol in a continuous communication sequence being greater than a preset percentage.
[0086] The purpose of the initial screening is not to immediately classify it as an attack, but to separate fragmented traffic with abnormal traces and prevent them from being mixed with normal main chain messages in the same association pool.
[0087] The system extracts two key features from these fragmented traffic: communication frequency features and protocol simulation features. The communication frequency features reflect the density of these fragments recurring within a short time window. High-frequency dissemination often means that someone is trying to continuously pollute the perception layer with a large number of tiny perturbations.
[0088] Protocol fidelity features reflect the degree to which these fragments mimic the appearance of legitimate communication, such as whether they maintain similar header structures, similar handshake fields, and similar session rhythms; real abnormal traffic usually does not meet standard protocol specifications, while the communication characteristics of fake flag interference data are close to the normal business baseline and only contain abnormal features below a preset proportion to mislead the defense system.
[0089] To facilitate understanding, a simplified implementation scenario is given: Suppose that three groups of fragments are separated from the regional node at a certain time: the first type of fragment, the second type of fragment, and the third type of fragment; the first type of fragment appears more than ten times in a short period of time, and each time it imitates the initial handshake format of the clearing message, only retaining a weak abnormal signature in the subsequent fields;
[0090] Although the second type of fragments have abnormal signatures, they are sparse and there are obvious missing items in the protocol fields; the third type of fragments are one-time burst fragments with messy protocol formats; based on this, the system can regard the first type of fragments as high-frequency and highly realistic interference candidates, while the second and third types of fragments are closer to real alarms or low-quality noise.
[0091] If a set of fragments simultaneously meets both the conditions of high frequency and high simulation, the system determines it to be false flag alarm storm interference data and removes it from the original situation data; if it does not meet both conditions, it is retained as valid alarm data for subsequent analysis.
[0092] Such judgment logic has engineering rationale: if an attacker only sends low-fidelity abnormal fragments, they are usually easily intercepted by traditional rule bases and it is difficult to cause deep disturbances to situational awareness.
[0093] The real danger lies in those high-frequency, continuous fragmented traffic that closely resembles legitimate business traffic. They may not necessarily disrupt a single session, but they can cause the model to misperceive the abnormal density of the entire network, leading to continuous high-frequency switching or abnormal oscillation of defense strategies.
[0094] Furthermore, the aforementioned purified multidimensional heterogeneous situational data still belongs to the subsequent form of the multidimensional heterogeneous situational data after the purification steps of this embodiment. Its data source, data type and subsequent associated objects have not changed. Only the identified false flag alarm storm interference data is removed from the original input set to improve the purity of the input.
[0095] Therefore, when it comes to extracting topological features, malicious traffic features, or legitimate business traffic features based on multidimensional heterogeneous situational data, if such steps occur after the purification steps in this embodiment, they can be performed by default based on the purified multidimensional heterogeneous situational data, and do not constitute a switch of the technical object name.
[0096] Furthermore, if some fragment traffic has a high communication frequency but the protocol simulation cannot be effectively determined, such as missing key fields or incomplete sampling, the system will place it in the review area instead of directly removing it to avoid mistakenly deleting real attack alarms.
[0097] Conversely, if the simulation level is high but the frequency of occurrence is extremely low, it is more likely to be an occasional protocol compatibility phenomenon or a single detection, and should not be hastily identified as storm interference; for the data in the area to be verified, it can be confirmed again in conjunction with subsequent windows; if multiple consecutive windows still maintain high frequency and high simulation, then the removal should be performed again;
[0098] If the abnormal signature feature library cannot match certain new fragments temporarily, these fragments will be retained as unknown abnormal fragments and enter the next stage to avoid missed detections due to the lag in the feature library.
[0099] For example, when the aforementioned clearing network enters the intraday peak settlement phase, the system observes that multiple edge financial terminals continuously report minor anomaly logs; these logs are very similar in appearance to normal clearing handshakes, but they reappear at a high density in a very short period of time and are distributed in multiple areas;
[0100] If it is directly incorporated into the original situational data, the communication protocol feature association model will mistakenly believe that multiple regions are experiencing real anomalies at the same time, thereby driving large-scale micro-isolation; through the purification steps of this embodiment, the system identifies a batch of segments that have both high-frequency propagation and high protocol simulation characteristics, and therefore marks them as false flag alarm storm interference data and removes them from the input set, retaining only the remaining alarm records that are truly capable of handling.
[0101] The purpose of this step is to first restore the purity of the situational awareness input, and then carry out topological correlation and risk prediction, so as to achieve the preemptive suppression of the chain reaction of misjudgments caused by false flag interference.
[0102] In this embodiment of the invention, the core processing module quantifies and calculates the defense trust entropy of the current network topology. Specifically, it is used to: collect the current dynamic access control list full load rate and routing table update frequency of the core routing node; collect the current transmission control protocol timeout retransmission storm index of the service node; assign preset weights to the dynamic access control list full load rate, the routing table update frequency, and the transmission control protocol timeout retransmission storm index, and then perform weighted summation to calculate the defense trust entropy.
[0103] This embodiment provides a quantitative mechanism for defense trust entropy; specifically, after the aforementioned purification process is introduced, although the quality of input data is improved, it is still insufficient to determine the defense strength solely based on whether abnormal traffic is detected, as this is not enough to reflect what kind of instability boundary the network is approaching.
[0104] Especially in clearing private networks, what truly causes continuous service interruptions across the entire network is often not the attack traffic itself, but rather the simultaneous fatigue and jitter of the access control table, routing control plane, and service transport layer; therefore, this embodiment uses three of the most direct state variables in engineering to jointly characterize the defense trust entropy.
[0105] Specifically, the first state variable is the full load rate of the dynamic access control list of the core routing node; the access control list is used to carry whitelists, blacklists, temporary isolation items and fine-grained session restriction rules;
[0106] When the system is near full load for an extended period, it means that the system is already maintaining a security boundary by stacking rules in a high-intensity manner. Once a new dynamic isolation command is introduced, the device may need to frequently replace existing rules, causing legitimate sessions that should be stable to be forcibly interrupted or rebuilt. Therefore, the full load rate reflects whether the core control resources have sufficient reserves.
[0107] The second state variable is the routing table update frequency. Routing table updates are normal within a certain range, but if the updates are too frequent, it indicates that the network path is continuously converging, switching, or being forced to detour. For backbone nodes that are responsible for forwarding clearing messages, excessively fast routing updates will directly increase latency jitter and may even cause session interruptions. The update frequency reflects whether the stability of the network topology is declining.
[0108] The third state variable is the Transmission Control Protocol Timeout Retransmission Storm Index of the service node; the Transmission Control Protocol Timeout Retransmission Storm Index is used to characterize the technical phenomenon that a large number of duplicate messages are sent by the service node in a short period of time. Specifically, the service node repeatedly sends the same or similar messages in a short period of time due to unstable upstream path, policy blocking, queue congestion or handshake interruption.
[0109] This phenomenon is extremely disruptive to financial clearing operations because the terminal side will repeatedly send confirmation, resend, and session reconstruction requests to ensure successful transaction submission, further straining already tight link and control plane resources; this index can be seen as a reverse feedback from the business side to defense strategies and topology oscillations.
[0110] The system aggregates the above three types of state variables to form a defense trust entropy. The weighting here does not emphasize the specific mathematical form, but rather the primary and secondary relationship in an engineering sense: if the control plane of the core routing node is close to its limit, its risk weight should be higher than that of the isolated anomaly of ordinary edge nodes; if the terminal retransmission storm has obviously spread, it means that the business side's tolerance for the current defense behavior is rapidly decreasing, and its impact cannot be regarded as secondary noise; specifically, this entropy value represents whether the network can still withstand continued aggressive scheduling.
[0111] To avoid the quantification process remaining at the level of abstract description, the system can first convert the three state variables into pressure values in the range of 0 to 1 during engineering implementation; among them, the full load rate of the dynamic access control list can be directly taken as the ratio of the currently used rule items to the upper limit of available rule items; the update frequency of the routing table can be calculated according to the degree of deviation of the number of updates within the preset observation window from the number of updates of the stable baseline.
[0112] The Transmission Control Protocol (TCP) timeout retransmission storm index can be calculated by combining the number of timeout retransmission packets, the number of repeated session reconstructions, and the proportion of affected terminals within the same service node within the same window; after completing the unified scaling, a weighted summation is then performed to obtain the defense trust entropy. In this calculation, the normalized stress value corresponding to the full load rate of the dynamic access control list is denoted as... The normalized pressure value corresponding to the routing table update frequency is denoted as The normalized stress value corresponding to the Transmission Control Protocol timeout retransmission storm index is denoted as The weights corresponding to the three are denoted as follows: , , and satisfy Then we have:
[0113]
[0114] in, Indicates multiplication operation. This represents addition. In a directly reproducible implementation, if the current access control list utilization rate of the core routing node reaches 0.82, the routing table update frequency is converted to 0.67 after baseline comparison, the retransmission storm index of the service node is converted to 0.74, and considering the network's priority requirements for control plane stability, the preset... 0.40, 0.35 A value of 0.25 yields a higher defense trust entropy.
[0115] Correspondingly, if the three state variables in another observation window are only 0.35, 0.22 and 0.18 respectively, then even if there is local abnormal traffic, the obtained defense trust entropy is still in a low range, indicating that the network still has sufficient capacity to continue to perform fine-grained dynamic adjustment.
[0116] Furthermore, the aforementioned weights can be pre-set based on the deployment scenario rather than arbitrarily shifted online; for example, in cross-border financial clearing networks, the systemic risk caused by the exhaustion of core gateway control plane resources is usually higher than the short-term jitter in edge areas, so the full load rate of the dynamic access control list can be assigned the highest weight.
[0117] If the system is running during the regional centralized settlement period, the impact of repeated terminal submissions on business continuity is significantly amplified. In this case, the upper edge of the weight of the retransmission storm index can be temporarily increased. However, this adjustment must be subject to a preset range to avoid the distortion of entropy values in adjacent windows due to weight jumps in the same network.
[0118] Furthermore, if some service nodes are temporarily unable to report the retransmission storm index due to link interruption, the system can first use the aggregated status of other nodes in the same area for compensation estimation, but will not extrapolate the compensation value to the entire network indefinitely; if the core routing node is in the maintenance window, resulting in a naturally high routing update frequency, the system can combine the maintenance flag to isolate and interpret the status, avoiding misjudging the planned switchover as an attack-induced shock.
[0119] If there is a contradiction among the three state variables, such as the access control list and routing table being stable, but the retransmission storm index suddenly rising significantly, the system can prioritize checking the terminal access side or application layer for duplicate submission issues, rather than immediately expanding the network isolation scope.
[0120] Furthermore, if a certain state variable exceeds a reasonable upper limit due to sampling anomalies, the system can truncate it to a preset maximum value before participating in the calculation, thus avoiding the amplification of the overall entropy value by a single invalid sampled data; if two or more core state variables are missing in the same window, the system will maintain the entropy value result of the previous stable window and mark the current result as having limited credibility, for subsequent prediction modules to use with caution.
[0121] For example, after the aforementioned storm of false flag alarms was partially filtered out, the system continued to observe that the access control rules of the headquarters clearing gateway were approaching the device's capacity limit. At the same time, due to multiple regional detours, the density of routing update events increased significantly. In addition, clearing terminals in Europe and Southeast Asia successively reported repeated submissions after session timeouts.
[0122] Based on this, the system determined that the current problem was no longer just whether there was suspicious traffic, but that the defense action itself had put perceptible pressure on the network structure. Therefore, the defense trust entropy was raised to a high level and entered a key warning state.
[0123] The purpose of this step is to unify control surface pressure, topological oscillations, and service-side rebounds into a single measure of network structural resilience, thereby enabling early detection of signs of instability.
[0124] In this embodiment of the invention, the result generation module outputs a predicted risk contribution value for the service blocking risk of core routing nodes in the communication network based on the defense trust entropy and the preset failure boundary prediction model. Specifically, it is used to: input the defense trust entropy into the failure boundary prediction model.
[0125] Calculate the probability value of continuous communication interruption occurring within a preset time window based on the network topology characteristics; convert the probability value into the predicted risk contribution value through a preset mapping function and output it.
[0126] This embodiment provides a risk contribution prediction mechanism for failure boundaries; specifically, after obtaining the defense trust entropy, the system does not directly equate the entropy value with whether to immediately switch the defense strategy, but further combines it with the failure boundary prediction model to determine its contribution to the continuous service interruption of the core routing node.
[0127] The reason for this is that the same high defense trust entropy may correspond to completely different business consequences: sometimes it is just a brief shock in a local area, but sometimes it indicates that the core clearing channel is about to enter a long-term and irreversible blockage state. The two need to be treated differently.
[0128] In detail, the failure boundary prediction model uses defense trust entropy and network topology characteristics as joint inputs, and focuses on observing whether core routing nodes may encounter continuous communication interruption within a preset time window; here, continuous communication interruption emphasizes that the core business link cannot be stably forwarded continuously, rather than a few instantaneous packet losses or short-term retries.
[0129] Financial clearing operations are highly time-sensitive. If the main path is continuously blocked for a certain period of time, transaction receipts, fund position confirmations, and reconciliation results will accumulate, with consequences far greater than short-term interruptions in general office networks. Therefore, the model focuses on whether the business can be maintained continuously, rather than whether individual messages arrive successfully.
[0130] To avoid confusion with the responsibility boundaries of the preceding modules, in the specific implementation, the defense trust entropy input failure boundary prediction model can be understood as follows: the defense trust entropy is used as the main driving input into the model, while the network topology features are used as accompanying state variables pre-extracted by the core processing module and sent into the model to participate in the probability calculation.
[0131] The model starts the risk assessment process with the defense trust entropy, and calculates the probability of the entropy value inducing continuous communication blockage under the current topology state, combined with the extracted topology features; it also makes the data flow relationship of entropy value-topology-blockage consequences clearer.
[0132] In engineering implementation, the failure boundary prediction model can adopt a hierarchical judgment logic based on time windows: first read the defense trust entropy of the current window, and then read the key topology features within the same window. The key topology features include at least node connection stability, forwarding path change frequency, key link detour number and control plane oscillation trend.
[0133] Based on whether these characteristics continue to deteriorate, rather than whether they are momentary anomalies, the probability of continuous communication disruptions occurring within a preset time window is estimated. Specifically, if only a single detour or short-term jitter occurs, but the main path recovers quickly, the probability value is not significantly increased. If multiple windows continuously show frequent path switching, failure of legitimate session closure, and accumulation of control surface oscillations, the probability value is gradually increased.
[0134] To facilitate understanding, a specific micro-level implementation scenario can be constructed: Suppose that the system records three types of topological characteristics within an observation window. The first type is characterized by frequent switching between adjacent links of the core gateway, the second type is characterized by multiple fallback routes at secondary nodes on the core path, and the third type is characterized by a legitimate session being established and then quickly disconnected.
[0135] When these features are input into the model along with the aforementioned higher defense trust entropy, the model gives a higher probability of continuous blocking within the future window; after mapping, a higher risk contribution prediction value is obtained; in another case, although there are local anomalies and several fragment alarms in the network, the core path does not experience continuous link reconstruction, and legitimate sessions can still close the loop, the model will give a lower blocking probability, corresponding to a lower risk contribution.
[0136] The mapping function mentioned here serves as the business interpretation layer. Because device state variables, topology features, and blocking probabilities belong to different levels of information, it is not intuitive to use them directly with the policy module. After mapping, the system can use a single risk contribution index to drive decision-making, so that subsequent modules only need to compare the index with the danger threshold to complete the policy diversion.
[0137] For example, the mapping function can adopt a segmented mapping approach: when the probability of continuous communication interruption is in a low range, the risk contribution is increased only slowly; when the probability enters the medium-high range, the slope of the risk contribution increase increases, thereby enabling the system to identify the state approaching the failure boundary earlier, without waiting for the service to be completely interrupted before triggering the switchover; in a specific engineering implementation, the preset mapping function is expressed as the following segmented logic: let the probability value of continuous communication interruption be... The predicted risk contribution value is ;when When the interval is lower, the mapping function is:
[0138]
[0139] in, For a smaller smoothing coefficient, for example ;when When entering the middle to high range, the mapping function changes to:
[0140]
[0141] in ,For example , for The boundary value at that time is used to accelerate the increase of the risk prediction value; when hour, It was directly set to the maximum saturation value at the highest risk level;
[0142] Furthermore, to avoid the prediction module from forming unexplainable discontinuous mutations, the system can simultaneously retain an explanation label each time it outputs a risk contribution prediction value. The explanation label at least indicates which type of factor mainly raises the prediction value, such as control surface pressure, path detour, or legal session loop failure.
[0143] This explanation label does not change the prediction process, but it helps the subsequent feedback adaptive module to determine whether to prioritize local isolation or stable path convergence, thereby ensuring that the strategy switching has a traceable engineering basis.
[0144] Furthermore, if the prediction model detects significant defects in the input features, such as core link topology data being missing beyond a preset range, the system will not output an overly aggressive high-risk judgment, but will instead mark the prediction result as having limited credibility and prioritize the collection of key link information.
[0145] If the network is in a known planned maintenance window, resulting in a short-term concentration of topology switching and session interruptions, the model can combine maintenance indicators to reduce the contribution of these features to the blocking probability, avoiding misjudging planned behavior as approaching the failure boundary. If the risk contribution is near the threshold and fluctuates up and down in multiple consecutive windows, the system can adopt a continuous confirmation mechanism, switching to the degraded mode only when the dangerous conditions are met continuously, in order to reduce frequent reversals caused by instantaneous jitter.
[0146] Furthermore, if the blocking probability output by a certain window is high, but the interpretation label shows that it mainly comes from a single edge area and does not affect the adjacent links of the core routing node, the system can keep observing instead of directly increasing the risk contribution of the entire network, so as to avoid excessive amplification of local anomalies.
[0147] For example, in the aforementioned clearing network, the control plane of the headquarters gateway was already under high pressure, and multiple regional business terminals were experiencing continuous retransmissions; the system further observed that the two backbone links adjacent to the core clearing path began to take turns undertaking forwarding tasks, and after the legitimate clearing session was established, it repeatedly went back and reconnected.
[0148] Based on this, the failure boundary prediction model judges that the probability of continuous service interruption of the core gateway increases significantly in the next critical window, and thus outputs a high-risk contribution prediction value. This result does not mean that the network has failed, but rather that continuing to maintain high-frequency dynamic defense is likely to push the system to the continuous interruption boundary.
[0149] The purpose of this step is to transform the abstract network instability state into an actionable risk indicator oriented towards the consequences of failure, thereby enabling proactive triggering of policy switching rather than reactive remediation.
[0150] In this embodiment of the invention, the feedback adaptive module generates the first defense strategy, specifically used for: extracting malicious traffic features and legitimate business traffic features based on the multidimensional heterogeneous situational data; the first defense strategy is generated by a preset multi-agent deep reinforcement learning model, wherein the preset multi-agent deep reinforcement learning model includes multiple agents deployed in the terminal domain, the regional aggregation domain and the core gateway domain respectively, and the action set of the multiple agents includes session-level isolation, service segment rate limiting and bypass traffic redirection;
[0151] The malicious traffic features and the legitimate service traffic features are input into a multi-agent deep reinforcement learning model; with the joint optimization objective of maximizing the malicious traffic interception rate and maximizing the legitimate service network throughput, the dynamic micro-segmentation strategy is output.
[0152] This embodiment provides a dynamic micro-segmentation generation mechanism. Specifically, when the aforementioned risk contribution is still below the danger threshold, it indicates that the network topology has sufficient resilience to withstand fine-grained control. At this time, the system does not need to immediately adopt a conservative static defense, but can use a multi-agent deep reinforcement learning model to generate a dynamic micro-segmentation strategy to more accurately distinguish between malicious traffic and legitimate clearing services.
[0153] Specifically, the feedback adaptive module extracts malicious traffic features and legitimate business traffic features from the purified multi-dimensional heterogeneous situational data; malicious traffic features may include abnormal detection rhythms, unauthorized destination combinations, abnormal session maintenance, abnormal certificate usage, and behavioral chain jumps, etc.
[0154] Legitimate business traffic characteristics focus more on the periodicity of clearing business, the directionality of reconciliation messages, the two-way confirmation loop, and the stable calling patterns between whitelisted systems. The reason for extracting both types of characteristics at the same time is that in the clearing private network, it is not enough to only determine whether the traffic has malicious attack characteristics; it is also necessary to identify whether it has key business characteristics. Otherwise, the system may mistakenly damage the main clearing traffic in order to increase the interception rate.
[0155] Multi-agent deep reinforcement learning models can deploy multiple agents by region, by link segment, or by node role; for example, edge terminal domain agents are suitable for identifying abnormal terminal behavior, regional aggregation domain agents are suitable for identifying cross-node propagation trends, and core gateway domain agents are more concerned with the main path load and control plane pressure.
[0156] Each intelligent agent makes collaborative decisions under the same goal: to improve the malicious traffic interception rate while maintaining the legitimate business network throughput as much as possible; the output is not a single blocking or allowing instruction, but forms a dynamic micro-segmentation strategy, such as restricting a certain terminal group to access a certain business segment, shortening the life cycle of suspicious sessions, raising the authentication threshold, limiting the rate of specific protocol branches, and guiding suspicious sources to a bypass cleaning path, etc.
[0157] To ensure the model's behavior is clearly engineering-executable, the state inputs, action sets, and feedback signals of each agent can be structurally defined in the specific implementation. Among them, the state inputs include at least: the intensity of malicious traffic characteristics in this area, the throughput level of legitimate business traffic, the latency changes of critical links, the number of currently applied isolation rules, and the collaborative labels reported by neighboring agents.
[0158] The action set includes at least: maintaining the status quo, performing session-level isolation on a specified terminal group, rate limiting on a specified service segment, raising the verification threshold of a specific authentication link, redirecting suspicious traffic to bypass cleaning, and withdrawing local isolation actions that have been proven to have excessive side effects; in this way, the model does not output strategies out of thin air, but makes choices within a limited and interpretable action space;
[0159] Furthermore, the joint optimization objective can be implemented as an immediate reward calculation for the effect of each round of actions; a directly deployable reward format can be written as... The immediate return is determined by the malicious traffic interception rate, the legitimate business throughput maintenance rate, and the technical overhead; if... The target weight corresponding to the malicious traffic blocking rate is represented by... This represents the target weight corresponding to the legitimate business throughput retention rate, in order to Let the suppression weights corresponding to the technical overhead be:
[0160]
[0161] in, Indicates multiplication operation. This represents the addition operation. This represents the subtraction operation; The malicious traffic blocking rate is... The throughput maintenance rate for the aforementioned legitimate business. This refers to policy perturbation overhead; policy perturbation overhead is used to characterize the number of newly added isolation rules, the number of session reconstructions triggered, and the additional pressure on the adjacent topology of the core routing node; and the policy perturbation overhead The dimensionless converted value after normalization;
[0162] Under the same training round or the same deployment configuration, , , To maintain the preset constant, a specific implementation scenario can be constructed for ease of understanding: Assume that there are a first legitimate business flow, a second legitimate business flow, a first suspicious flow, and a second suspicious flow in the clearing network; the edge agent determines that the terminal from which the first suspicious flow originates has abnormal process calls, the regional agent discovers that the first suspicious flow and the second suspicious flow exhibit synchronous detection rhythms at different sites, and the core agent confirms that the first legitimate business flow and the second legitimate business flow are carrying important clearing transaction receipts;
[0163] After synthesis, the dynamic micro-segmentation strategy output by the model does not shut down the entire area exit, but only applies fine-grained isolation to the terminal groups to which the first and second suspicious flows belong and adjacent suspicious session segments, while reserving the main channel for priority forwarding of the first and second legitimate service flows; thus maintaining the preset legitimate service network throughput and effectively suppressing the propagation of attack traffic.
[0164] Compared to fixed rules, the advantage of dynamic micro-segmentation is that it can continuously adjust the boundaries based on current observations. For example, when the attack is mainly concentrated on the terminal access side, the isolation range can be close to the terminal group. When the anomaly begins to spread across regions, the isolation range can be raised to the regional convergence layer. However, as long as the risk contribution has not exceeded the danger threshold, it will not be temporarily changed due to the instantaneous fluctuations of a single observation window.
[0165] The system still prioritizes a defense philosophy of local isolation and global availability;
[0166] To avoid actions that cancel each other out or even amplify each other among multiple agents, the system can also set cooperative constraints: if the edge agent suggests expanding the isolation, and the core gateway domain agent detects that the throughput of the main path is approaching the protection lower limit, the system will prioritize low-disturbance actions, such as rate limiting or bypass diversion, rather than directly cutting off the relevant service segments.
[0167] If multiple agents in a region simultaneously propose repetitive actions to the same terminal group, the system merges them into a unified strategy to avoid the same object being issued overlapping rules multiple times. The above constraints do not change the basic structure of the dynamic micro-isolation strategy output by the multi-agent model, but only add engineering consistency convergence to the output results.
[0168] Furthermore, if the multi-agent model lacks sufficient observation data in certain areas, the system can revert to the preset safety baseline rules and not force the output of high-confidence dynamic isolation actions; if the isolation scheme given by the model conflicts with the critical business whitelist, the whitelist protection rules take precedence, and the system only allows low-interference measures that do not interrupt the critical clearing loop; if a strategy causes a significant deterioration in the latency of legitimate business after trial execution, the system can immediately revert to the local isolation boundary and re-evaluate adjacent strategies to avoid amplifying structural instability due to local misjudgments when the risk is still controllable.
[0169] Furthermore, if the model repeatedly gives oscillating actions of isolation-withdrawal-re-isolation to the same object within multiple consecutive decision windows, the system can trigger a short-term freeze mechanism. During the freeze period, only fine-tuning of the action amplitude is allowed, and the isolation range is not allowed to be expanded, so as to avoid turning the strategy itself into a new source of disturbance.
[0170] For example, in the aforementioned network, even after purification, some regional terminals were still found to initiate atypical session probes to the core clearing gateway, but the main business links were generally stable and the risk contribution was below the danger threshold; the system then called the multi-agent model: the edge side identified that the behavior of a certain batch of terminals was inconsistent with the historical clearing terminals, the regional side found that their communication frequency was abnormally concentrated, and the core side confirmed that these sessions were not critical settlement channels.
[0171] Based on this, the model outputs a dynamic micro-segmentation strategy, which restricts relevant terminals from accessing certain sensitive routing segments while maintaining the high-priority forwarding of normal clearing messages; in this way, it can achieve fine-grained blocking of suspicious traffic without interrupting the continuity of legitimate clearing.
[0172] The purpose of this step is to maximize both interception effectiveness and service throughput by using fine-grained and interconnected defense methods when the network still has room for adjustment, thereby achieving dynamic defense with strong local handling and minimal global disruption.
[0173] In this embodiment of the invention, the feedback adaptive module generates the second defense strategy, specifically for: suspending the issuance of the first defense strategy in response to the triggering of the strategy degradation mechanism; extracting the baseline routing table data of the core routing node; and generating the static route cleaning strategy based on the baseline routing table data, wherein the execution priority of the static route cleaning strategy is higher than that of the dynamic micro-segmentation strategy, and the static route cleaning strategy is used to reduce the defense trust entropy.
[0174] This embodiment provides a static route cleaning mechanism under policy degradation. Specifically, although the aforementioned dynamic micro-segmentation can provide good fine-grained defense when the risk is controllable, its premise is that the network topology can still withstand continuous policy updates. Once the defense trust entropy increases and the risk contribution reaches the dangerous threshold, continuing to frequently issue dynamic isolation actions may become a triggering factor that further aggravates the control plane load of core routing nodes and topology oscillations.
[0175] Therefore, this embodiment introduces a strategy degradation mechanism under such extreme conditions, shifting the focus of defense from more refined interception to stabilizing the main path first;
[0176] Specifically, when the degradation conditions are met, the feedback adaptive module pauses the issuance of dynamic micro-isolation strategies. This pause does not necessarily mean revoking isolation rules that have been stably executed and have no side effects, but rather stopping the addition of new or highly extensible dynamic changes to avoid continued oscillations in the control plane.
[0177] The system extracts baseline routing table data from core routing nodes; this baseline routing table is usually derived from historical stable operating windows and reflects the set of most reliable forwarding paths for the main clearing service under non-attack, non-congestion, and non-maintenance conditions.
[0178] Based on this baseline routing table, the system generates a static route cleaning strategy. The so-called cleaning is to converge the current chaotic, temporary, overly refined or conflicting paths and strategies back to a small number of reliable backbone paths, prioritize the retention of fixed forwarding paths that carry the core business of clearing, and remove or suppress temporary routing items that cause detours, backtracking and repeated switching.
[0179] Because such strategies rely on stable baselines rather than real-time exploration, they are more interpretable and cause significantly fewer network topology changes than dynamic micro-segmentation strategies. In terms of execution priority, static route cleaning should be higher than dynamic micro-segmentation to ensure that when the system is close to the failure boundary, there will be no conflict between cleaning being performed while dynamic strategies continue to cause disturbances.
[0180] To facilitate understanding, a specific implementation scenario can be provided: Suppose that the original dynamic micro-segmentation had already issued fine-grained restrictions on nodes N1, N2, N3, and N4, causing the main liquidation path to frequently change routes between multiple nodes; after the system enters a degraded state, it will no longer extend these fine-grained restrictions, but will instead revert to the stable baseline route, retaining only the main path from the regional front-end machine through N2 to the core gateway, and clearing the high-frequency change items used for temporary detours on N3 and N4;
[0181] While this may reduce the fine-grained blocking intensity of some edge probe traffic, it can quickly reduce topology oscillations and suppress retransmission storms and session backoffs.
[0182] The mechanism by which static route cleaning reduces defense trust entropy is as follows: on the one hand, it reduces continuous changes to the control plane, alleviating the fatigue of access control lists and routing tables; on the other hand, it allows business nodes to see stable and predictable forwarding paths again, reducing the uncertainty caused by frequent reconnection of terminals; for financial clearing systems, the stable restoration of communication continuity is itself part of risk control.
[0183] Furthermore, if the baseline routing table data is at risk of becoming outdated, for example, if the network topology has recently undergone structural changes due to equipment replacement, the system can first perform reachability verification on the baseline routes, and then perform cleaning if the verification passes; if the verification fails, it will then select the most recent verified and stable alternative baseline.
[0184] If the intensity of a local attack temporarily increases after dynamic micro-segmentation is suspended, the system can retain only the necessary minimum isolation set near the attack source without resuming global dynamic expansion; if core services still experience continuous rollbacks after static route cleaning is executed, the system can further shrink to a smaller range of critical service survival strategies, prioritizing the survival of the main liquidation link.
[0185] For example, in the aforementioned cross-border clearing network, as the storm of false flag alarms continues to escalate, the access control rules of the headquarters gateway are constantly being added to and deleted from, and the dynamic micro-segmentation actions of multiple regions overlap, causing legitimate clearing sessions to be frequently rebuilt.
[0186] The system determined that it was approaching the continuous blocking boundary, so it triggered a policy downgrade, stopped issuing dynamic micro-segmentation commands to more nodes, and extracted the baseline routing table formed by the core gateway in the previous stable clearing day. Based on this baseline, the system converged to a small number of fixed main paths, cleaned up many temporary detours and duplicate isolation items, thereby restoring the core clearing channel to stability.
[0187] The purpose of this step is to proactively reduce some fine-grained interception intensity during high-risk phases in exchange for smaller topology disturbances and higher business continuity, thereby achieving proactive avoidance of failure boundaries.
[0188] In this embodiment of the invention, the system further includes a policy distribution module, which is used to: obtain the in-band management channel bandwidth utilization rate of the communication network;
[0189] Determine whether the bandwidth occupancy rate is greater than a preset congestion threshold: if it is greater than the congestion threshold, perform instruction compression processing on the first defense strategy or the second defense strategy, and execute the delivery operation based on the compressed instruction; if it is less than or equal to the congestion threshold, directly execute the delivery operation.
[0190] This embodiment provides a channel adaptive mechanism for policy distribution; specifically, in the aforementioned implementation stages, regardless of whether the system chooses dynamic micro-segmentation or static route cleaning, the policy ultimately needs to be delivered to the relevant network nodes through the in-band management channel.
[0191] However, in cross-regional clearing private networks, in-band management channels often share some basic transmission resources with business links. Once business peaks or attacks cause channel congestion, if complete, lengthy, and item-by-item policy instructions are still issued, it will not only delay the execution of defenses, but may also further squeeze the transmission space of core services. Therefore, this embodiment introduces a switching of the issuance method based on bandwidth utilization.
[0192] Specifically, the policy delivery module continuously monitors the bandwidth utilization of the in-band management channel and compares it with the congestion threshold. When the utilization rate is within the normal range, it indicates that the channel still has sufficient transmission capacity, and the system can directly deliver the complete policy, including the complete list of nodes, action type, scope of application, priority, and rollback flag. This makes it easy for each node to parse and execute the policy as is and execute it quickly.
[0193] When the occupancy rate exceeds the congestion threshold, it indicates that the management channel has entered a high-pressure state. At this time, the system compresses the instructions to be issued for policy execution. The compression here is not limited to byte-level compression, but emphasizes the convergence of control semantics. For example, when multiple adjacent nodes execute the same isolation action, the node-by-node instructions can be merged into a region-level action.
[0194] For recurring rule prefixes, target network segments, and priority descriptions, index references can be used instead of sending the entire list repeatedly. For static route cleaning strategies, only change summaries relative to the baseline route can be sent instead of transmitting the complete routing table. This can significantly reduce the occupation of management channels and shorten the delivery time of critical policies.
[0195] To facilitate understanding, a simplified implementation scenario can be provided: Suppose that the original dynamic micro-segmentation strategy needs to issue four similar instructions to four nodes, namely, restricting the first terminal group from accessing the first service segment, restricting the first terminal group from accessing the second service segment, restricting the first terminal group from accessing the third service segment, and restricting the first terminal group from accessing the fourth service segment.
[0196] When the management channel is congested, the system can compress it into a single merge instruction that restricts access to the first service segment to the fourth service segment for the first terminal group, and attach a unified priority identifier.
[0197] For example, static route cleaning originally required distributing a complete route snapshot, but now it can be changed to only sending a summary command to restore to the third baseline version and delete the seventh, eighth, and ninth temporary entries;
[0198] Each node can complete the restoration based on the pre-shared baseline version number; to avoid ambiguity with other symbols in the previous example, the aforementioned first terminal group only represents the identifier of a terminal group to be restricted, the first service segment to the fourth service segment only represents the identifier of consecutive target service segments, the third baseline version only represents the shared baseline route version number, and the seventh to ninth temporary items only represent the temporary route item numbers to be deleted.
[0199] The above notation is only used to illustrate the merging expression method when compressing and distributing data, and does not refer to network topology characteristics, probability characteristics, or other pre-calculated variables;
[0200] Furthermore, if the occupancy rate of the management channel has not exceeded the congestion threshold, but is close to the upper limit of the threshold, the system can prioritize the compression of low-urgency strategies while retaining the complete expression of high-urgency strategies.
[0201] If the compressed command fails to be parsed at a node, the node can automatically request the minimum necessary retransmission segment instead of requiring a full retransmission strategy to avoid further exacerbating congestion. If the in-band management channel experiences severe congestion or even partial interruption, the system can prioritize sending the minimum keep-alive command to the core routing node and critical adjacent nodes, with secondary regional nodes retransmitting it later, to ensure that the critical clearing main path is restored to stability first.
[0202] For example, when the aforementioned network has entered a high-risk phase, the system decides to issue a static route cleaning policy. However, at this time, the cross-border in-band management channel is already close to congestion due to a large number of status reports and control messages. The policy issuance module detects that the occupancy rate exceeds the preset threshold, so instead of sending the complete multi-node cleaning details, it issues a compression command to the core gateway and several key nodes in the form of a baseline version number plus a change summary.
[0203] Because the command length is significantly shortened, critical nodes are able to converge routes in a timely manner, avoiding the management messages themselves slowing down the recovery of core business operations;
[0204] The purpose of this step is to ensure that the defense strategy can still be delivered to critical nodes in a timely and reliable manner under restricted transmission conditions, thereby achieving the adaptation of the distribution to be more convergent as the channel becomes more congested.
[0205] In this embodiment of the invention, the communication network is a cross-border financial clearing network, the business node is a financial clearing terminal, and the core routing node is a clearing gateway node.
[0206] This embodiment provides a deployment method for cross-border financial clearing networks. Specifically, when the aforementioned system is applied to cross-border financial clearing scenarios, the business nodes are financial clearing terminals distributed in different jurisdictions, different settlement centers, and different bank access domains, and the core routing node is the clearing gateway node that carries the main channel for cross-border clearing.
[0207] This scenario is highly adaptable to the solution of the present invention because it simultaneously possesses characteristics such as high timeliness, high reliability, high resistance, and high severity of failure.
[0208] Specifically, messages in cross-border financial clearing networks not only handle fund transfers, receipt confirmations, risk checks, and reconciliation instruction transmissions, but also often cross dedicated lines, encrypted tunnels, and compliance boundaries in different regions.
[0209] Therefore, the network inherently presents the following challenges: First, the large latency differences between regional nodes make it impossible for logs and alarms to be fully synchronized; second, the clearing gateway node is a critical unit, and if its control plane resources are dragged down by policy fluctuations, it will directly affect the flow of funds in multiple regions; third, the business side has extremely high requirements for continuity, and even a short interruption can trigger transaction queuing, rollback, and manual compensation.
[0210] Fourth, attackers tend to contaminate the monitoring and scheduling infrastructure by disguising legitimate business operations rather than directly launching easily identifiable large-scale traffic surges; the multi-dimensional heterogeneous data collection, fake flag interference purification, defense trust entropy assessment, failure boundary prediction, dynamic micro-isolation, policy downgrading and compression distribution in the aforementioned embodiments are precisely designed to address these industry-specific constraints.
[0211] Under this deployment method, financial clearing terminals can report situational data in groups according to region or institution attributes, while clearing gateway nodes additionally report access control table pressure, routing convergence status, main path health, and key business forwarding statistics. The system can be deployed in an independent security dispatch center or adopt a primary and backup dual-center structure to ensure that if a control node in a certain region is abnormal, the other control node can still continue to complete risk assessment and policy distribution.
[0212] Furthermore, if some regions cannot upload the complete original message due to compliance restrictions, they can upload only the de-identified communication feature summary and terminal behavior tag. The system can still complete the topology situation analysis based on the protocol association. If the leased lines in some regions have fixed delays exceeding the preset delay threshold, the data in that region can use a more lenient time alignment window to avoid being misjudged as abnormal due to natural delays.
[0213] If an individual clearing terminal suspends service due to local maintenance, the system can temporarily remove it from the main evaluation set of defense trust entropy to prevent maintenance behavior from disturbing the overall judgment.
[0214] For example, in a multinational financial clearing network covering Asia, Europe and the Middle East, multiple financial clearing terminals are connected to the headquarters clearing gateway through regional front-end systems;
[0215] During peak business hours, attackers deploy highly realistic fragmented traffic in several edge areas, attempting to induce the system to frequently switch access control and routing policies; the system cleans up the reported data, identifies and removes false flag alarm storm fragments; it calculates the defense trust entropy by combining gateway control plane load, routing update frequency and terminal retransmission status, and predicts its risk contribution to the continuous blocking of core services.
[0216] When the risk is low, the system implements dynamic micro-segmentation, restricting only a few suspicious terminal groups; when the risk increases, the system switches to static route cleaning and issues compressed management instructions to the clearing gateway and key nodes to ensure that the cross-border clearing main channel is restored to stability as soon as possible.
[0217] The purpose of this step is to implement the aforementioned technical aspects in financial clearing scenarios that require high risk, high impact, and high continuity, thereby achieving security protection and continuous operation assurance for the core links of cross-border clearing business.
[0218] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A security situation awareness and dynamic defense system based on multi-dimensional data association, characterized in that, include: The data acquisition module is used to acquire multidimensional heterogeneous situational data reported by each service node in the communication network, wherein the communication network includes the service nodes and the core routing node, and the multidimensional heterogeneous situational data includes traffic logs and terminal behavior characteristics with network latency characteristics and noise attributes. The core processing module is used to extract network topology features based on the multidimensional heterogeneous situational data and the preset communication protocol feature association model, and to quantitatively calculate the defense trust entropy of the current network topology. The defense trust entropy is used to characterize the tolerance of the service node to the defense strategy and the vulnerability index of the network topology features. The result generation module is used to output a predicted risk contribution value for the service blocking risk of the core routing node based on the defense trust entropy and the preset failure boundary prediction model. The feedback adaptive module is used to determine the logical relationship between the predicted risk contribution value and the preset danger threshold. If the predicted risk contribution value is less than the danger threshold, a first defense strategy is generated, wherein the first defense strategy includes a dynamic micro-segmentation strategy based on multi-agent deep reinforcement learning. If the predicted risk contribution value is greater than or equal to the danger threshold, a policy degradation mechanism is triggered, and a second defense strategy is generated. The second defense strategy includes a static route cleaning strategy, which is used to converge the current network topology to a preset baseline routing path to reduce the number of network topology change nodes. After the data acquisition module acquires the multidimensional heterogeneous situational data reported by each service node in the communication network, the core processing module is further used for: Based on a preset abnormal signature feature library, the multidimensional heterogeneous situational data is matched to separate fragmented traffic data with abnormal signatures. Extract the communication frequency characteristics and protocol simulation characteristics of the fragmented traffic data; If the communication frequency characteristic is greater than a preset frequency threshold and the protocol simulation characteristic is greater than a preset simulation threshold, then the fragmented traffic data is determined to be false flag alarm storm interference data; otherwise, it is determined to be valid alarm data. The false flag alarm storm interference data is removed from the multidimensional heterogeneous situation data to generate purified multidimensional heterogeneous situation data, and the purified multidimensional heterogeneous situation data is input into the communication protocol feature association model. The core processing module quantifies and calculates the defense trust entropy of the current network topology, specifically for: Collect the current dynamic access control list full load rate and routing table update frequency of the core routing node; Collect the current Transmission Control Protocol timeout retransmission storm index of the service node; The dynamic access control list full load rate, the routing table update frequency, and the transmission control protocol timeout retransmission storm index are each assigned a preset weight, and then weighted summation is performed to calculate the defense trust entropy. The result generation module, based on the defense trust entropy and a preset failure boundary prediction model, outputs a predicted risk contribution value for the service blocking risk of core routing nodes in the communication network, specifically used for: Input the defense trust entropy into the failure boundary prediction model; Calculate the probability value of continuous communication interruption occurring within a preset time window for the network topology characteristics; The probability value is converted into the predicted risk contribution value through a preset mapping function and then output. The feedback adaptive module generates the first defense strategy, specifically for: extracting malicious traffic features and legitimate business traffic features based on the multidimensional heterogeneous situational data; the first defense strategy is generated by a preset multi-agent deep reinforcement learning model, wherein the preset multi-agent deep reinforcement learning model includes multiple agents deployed in the terminal domain, regional aggregation domain and core gateway domain respectively, and the action set of the multiple agents includes session-level isolation, service segment rate limiting and bypass traffic redirection; The malicious traffic features and the legitimate business traffic features are input into the multi-agent deep reinforcement learning model; With the joint optimization objectives of maximizing the malicious traffic interception rate and maximizing the legitimate business network throughput, the dynamic micro-segmentation strategy is output.
2. The security situation awareness and dynamic defense system based on multi-dimensional data association according to claim 1, characterized in that, The feedback adaptive module generates the second defense strategy, specifically for: In response to the triggering of the policy downgrade mechanism, the issuance of the first defense policy is suspended; Extract the baseline routing table data of the core routing node; Based on the baseline routing table data, a static route cleaning strategy is generated, wherein the execution priority of the static route cleaning strategy is higher than that of the dynamic micro-segmentation strategy, and the static route cleaning strategy is used to reduce the defense trust entropy.
3. The security situation awareness and dynamic defense system based on multi-dimensional data association according to any one of claims 1 to 2, characterized in that, The system also includes a policy distribution module, which is used for: Obtain the in-band management channel bandwidth utilization rate of the communication network; Determine whether the bandwidth utilization rate is greater than a preset congestion threshold: If the congestion threshold is exceeded, the first defense strategy or the second defense strategy is compressed, and the delivery operation is executed based on the compressed instructions. If the value is less than or equal to the congestion threshold, the distribution operation is executed directly.
4. The security situation awareness and dynamic defense system based on multi-dimensional data association according to claim 1, characterized in that, The communication network is a cross-border financial clearing network, the business node is a financial clearing terminal, and the core routing node is a clearing gateway node.