Multi-model game method for multi-scene intelligent safety management

By using multi-domain data processing and game template routing, step-level risk signals are generated and online game solutions are performed. This solves the problems of game model mismatch and risk inconsistency in intelligent security management, and enables the system to be stable and respond quickly in dynamic environments.

CN121365734APending Publication Date: 2026-01-20ZHONGDA HOSPITAL SOUTHEAST UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511505920.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing intelligent security management methods suffer from problems such as mismatch between game models and scenarios, inconsistent risk injection timing, and lack of early warning of equilibrium point stability when facing dynamically changing adversarial environments, leading to unstable decision-making and delayed response.

Method used

By acquiring original data from multiple domains to generate information structure features, performing game template routing and decision-sensitive subspace compression, generating step-level risk signals, and conducting online game solving and equilibrium vulnerability governance, the system outputs security decisions and governance strategies, including cross-temporal risk injection, dynamic risk measurement, and equilibrium vulnerability index monitoring.

Benefits of technology

It enhances the decision-making adaptability, stability, and robustness of intelligent systems in dynamic adversarial environments, ensuring the effectiveness of decisions and rapid response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365734A_ABST
    Figure CN121365734A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-model game method for multi-scene intelligent safety management. The method comprises the following steps: acquiring multi-domain original data and generating information structure features; the heterogeneous indexes are uniformly mapped into information loss components based on the features, and stable routing of the game template is carried out in combination with sequential statistics and hysteretic decision; performing game payment-oriented decision sensitive subspace compression on the unified semantic state output after routing; a dynamic risk-to-go (dynamic risk surplus) time difference method is adopted to generate step risk signals meeting total amount conservation, and the step risk signals are injected into a compression space; and through online estimation of the spectral radius of the Jacobian matrix of the optimal response operator, an equilibrium vulnerability index is generated, and stability management and closed-loop adjustment of an equilibrium point are realized. According to the invention, the decision adaptability, stability and robustness of the intelligent system in a dynamic confrontation environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence security, and relates to a multi-model game method for intelligent security management in multiple scenarios. BACKGROUND

[0002] In intelligent systems such as network security, content risk control, and automatic driving, security management is facing an increasingly complex and dynamic confrontation environment. The strategy, intensity, and camouflage of attackers or risk events are constantly evolving, which puts higher requirements on the adaptive ability and decision robustness of the system. Therefore, researching a security management method that can face multiple scenarios and achieve intelligent decision-making and stable control is of great practical significance for improving the reliability and security of key intelligent systems.

[0003] Currently, for intelligent security management in such a confrontation environment, existing research has used game theory for modeling. Common methods are based on specific scenario assumptions to establish fixed game models, for example, applying a Stackelberg security game model in a resource attack-defense scenario. To adapt to environmental changes, some solutions use heuristic methods based on rules or simple thresholds to switch between a few pre-set models. In the representation of decision state space, existing technologies usually use general dimension reduction methods such as principal component analysis (PCA) or autoencoders to process high-dimensional original state data. In risk management, to incorporate long-period risk considerations (such as avoiding extreme losses) into decision-making, some methods use risk measures such as conditional value at risk (CVaR) as a penalty term, which is added to the single-step reward function through simple weighting or averaging.

[0004] However, existing methods still face several deep technical challenges in handling scenario dynamics, risk temporal consistency, and decision stability. Due to the instantaneous changes in information structure (such as data missing rate, observation delay), relying on fixed game models or simple heuristic switching rules can easily lead to model mismatch with the actual scenario, or produce frequent and unstable model jitter at the decision boundary, reducing the effectiveness of decision-making. Simply averaging long-period risks as step-level penalties violates the temporal consistency principle of dynamic programming in theory, resulting in fragmented global risks that cannot guarantee total conservation and are prone to dramatic shocks in the strategy learning process due to distorted risk signals. Even if the system solves the strategy equilibrium point, existing technologies lack forward-looking indicators that can calculate the local stability of this equilibrium point online and quickly, leading to a lag in the system's response to the risk of equilibrium collapse that may be triggered by minor disturbances, making it difficult to actively and timely manage. SUMMARY

[0005] The application provides a multi-model game method for intelligent safety management in multiple scenes, which solves the problems of game model mismatching with scenes, inconsistent risk injection time and lack of equilibrium point stability warning in existing methods.

[0006] The technical scheme, according to one aspect of the application, a multi-model game method for intelligent safety management in multiple scenes, comprises:

[0007] Obtaining multi-domain original data, and generating information structure features based on the multi-domain original data;

[0008] According to the information structure features, performing game template routing, outputting selected game templates, unified semantic states and unified action sets;

[0009] Receiving the unified semantic states, performing decision-sensitive subspace compression, and outputting compressed states;

[0010] Receiving the compressed states, performing cross-time domain risk injection, and generating step-level risk signals;

[0011] Receiving and based on the compressed states, the unified action sets, the step-level risk signals and the game templates, performing online game solving and equilibrium vulnerability management, and outputting safety decisions and management strategies and deploying.

[0012] According to one aspect of the application, the step of performing cross-time domain risk injection to generate step-level risk signals comprises:

[0013] Constructing a conditional distribution about a future loss sequence, and setting a time-consistent dynamic risk measure;

[0014] Based on the conditional distribution and the dynamic risk measure, evaluating the dynamic risk-to-go at the current time;

[0015] Calculating the conditional expectation of the dynamic risk-to-go at the next time under the current information condition;

[0016] Differencing the dynamic risk-to-go at the current time and the conditional expectation of the dynamic risk-to-go at the next time to generate the step-level risk signal.

[0017] According to one aspect of the application, the step of evaluating the dynamic risk-to-go at the current time comprises:

[0018] Setting a terminal boundary condition, and performing backward recursion from the terminal boundary condition to calculate the dynamic risk-to-go at each time in the time sequence;

[0019] The method further comprises: after generating the step-level risk signal, checking whether the sum of the step-level risk signal in the time sequence is conserved with the overall risk at the initial time calculated by the backward recursion.

[0020] According to an aspect of the present application, the dynamic risk measure comprises an adaptive confidence level, and the generation of the adaptive confidence level comprises:

[0021] Analyzing key risk indicators derived from multi-domain raw data to identify event mutations and divide time series into event segments and stationary segments;

[0022] In operation, the adaptive confidence level is determined according to the fluctuation intensity of the event segment.

[0023] According to an aspect of the present application, after generating the step-level risk signal, the following steps are further included:

[0024] Based on the compressed state, the contribution coefficients of each dimension inside are calculated;

[0025] According to the contribution coefficients, the step-level risk signal is scaled and allocated to construct a step-level risk penalty for the game solution target;

[0026] And the violation rate of the risk constraint is monitored to update the dual multipliers used to weigh the step-level risk penalty online.

[0027] According to an aspect of the present application, the step of balancing vulnerability governance and outputting governance strategies comprises:

[0028] According to the latest strategy and payoff matrix output by the online game solution, a local linear approximation of the best response operator in the game is constructed to obtain its Jacobian matrix;

[0029] The spectral radius of the Jacobian matrix is estimated;

[0030] Based on the estimation result of the spectral radius, the equilibrium vulnerability index is generated, and when the index exceeds a preset threshold, the governance strategy is generated.

[0031] According to an aspect of the present application, the step of generating the equilibrium vulnerability index further comprises:

[0032] Based on the strategy and payoff trajectory, the short-window regret slope is calculated;

[0033] The equilibrium vulnerability index is a weighted combination of the estimation result of the spectral radius and the short-window regret slope;

[0034] When the index exceeds the preset threshold, the generated governance strategy includes at least one of the following: freezing the route of the game template, widening the hysteresis interval for selecting the game template, applying low-pass filtering to the step-level risk signal, or reducing the update step of the dual multiplier.

[0035] According to an aspect of the present application, the step of routing the game template according to the information structure characteristics comprises:

[0036] Map the heterogeneous indexes in the information structure features into information loss components, the heterogeneous indexes including at least two of the following: intention posterior entropy, data missing rate, link delay, and distribution drift;

[0037] Aggregate the information loss components to generate an information structure index;

[0038] Select a game template according to the information structure index, and perform game template routing.

[0039] According to an aspect of the present application, the step of mapping the heterogeneous indexes into information loss components is based on log-likelihood ratio or relative entropy;

[0040] The step of selecting a game template according to the information structure index includes:

[0041] The information structure index is time accumulated in a sequential statistical manner to obtain a sequential statistic;

[0042] The sequential statistic is judged by using a threshold with a hysteresis interval to select a game template.

[0043] According to an aspect of the present application, the step of performing decision-sensitive subspace compression and outputting a compressed state includes:

[0044] Evaluate the sensitivity of each state dimension in the unified semantic state to the safety utility or payment of the game to form a sensitivity spectrum;

[0045] Project the unified semantic state according to the sensitivity spectrum to output a compressed state.

[0046] According to an aspect of the present application, the step of projecting the unified semantic state according to the sensitivity spectrum includes:

[0047] Self-sampling and voting integration are performed on the sensitivity spectrum to generate a robust dimension ranking;

[0048] When the robust dimension ranking changes beyond a threshold, update the key dimension set used for projection according to a set ranking hysteresis threshold.

[0049] The ranking hysteresis threshold is set to update the key dimension set used for projection only when the robust dimension ranking changes significantly.

[0050] According to an aspect of the present application, the step of projecting the unified semantic state according to the sensitivity spectrum further includes:

[0051] Construct a projection operator based on the key dimension set;

[0052] Calculate the condition number of the subspace defined by the projection operator;

[0053] And ensure that the condition number does not exceed the preset upper limit, to generate a numerically stable compression state.

[0054] According to one aspect of the present application, before mapping the data missing rate into the information loss component, the method further comprises:

[0055] Identifying whether the missing mechanism of the data is non-random missing;

[0056] And when the missing mechanism is non-random missing, using a selection model correction or importance weighting method to correct the bias of the data missing rate.

[0057] According to one aspect of the present application, the step of performing decision-sensitive subspace compression and outputting the compressed state comprises:

[0058] Evaluating the sensitivity of each state dimension in the unified semantic state to the security utility or payment of the game to form a sensitivity spectrum;

[0059] And projecting the unified semantic state according to the sensitivity spectrum to generate the compressed state.

[0060] According to one aspect of the present application, the step of projecting the unified semantic state according to the sensitivity spectrum comprises:

[0061] Self-sampling and voting integration of the sensitivity spectrum to generate a robust dimension ranking;

[0062] And setting a ranking hysteresis threshold to update the key dimension set used for projection only when the robust dimension ranking changes significantly.

[0063] According to one aspect of the present application, the step of updating the key dimension set used for projection further comprises:

[0064] Applying a diversity constraint to limit the number of dimensions selected into the key dimension set from the same semantic cluster based on the semantic similarity between dimensions, thereby ensuring balanced distribution of the key dimension set among different semantic categories.

[0065] According to one aspect of the present application, after generating and applying the governance strategy, the method further comprises:

[0066] Continuously monitoring the estimated result of the spectral radius and the short-window regret slope;

[0067] When the estimated result of the spectral radius returns to the safe interval and the short-window regret slope turns to be insignificant, gradually undoing the actions contained in the governance strategy.

[0068] According to one aspect of the present application, the step of aggregating the information loss component to generate the information structure index comprises:

[0069] A joint aggregation model including linear terms and quadratic interaction terms is adopted;

[0070] The linear terms are used to evaluate the independent influence of each information loss component, and the quadratic interaction terms are used to capture the coupling effect between different information loss components.

[0071] The beneficial effects, through the above technical solutions, the intelligent system can improve the decision adaptability, stability and robustness in dynamic confrontation environment. BRIEF DESCRIPTION OF DRAWINGS

[0072] Figure 1 is the overall process schematic diagram of the multi-model game method for multi-scene intelligent security management;

[0073] Figure 2 is the process schematic diagram of generating step risk signal by cross-time domain risk injection;

[0074] Figure 3 is the process schematic diagram of outputting governance strategy by balanced vulnerability governance;

[0075] Figure 4 is the process schematic diagram of game template routing according to information structure characteristics;

[0076] Figure 5 is the process schematic diagram of outputting compressed state by decision sensitive subspace compression. DETAILED DESCRIPTION

[0077] In order to make the person skilled in the art better understand the present application scheme, the technical scheme in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0078] It should be noted that the terms "first", "second" and the like in the specification and above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein.

[0079] It should be particularly pointed out that, in order to clearly show the step flow of the present application, serial numbers are marked for each step in the description. These serial numbers are only for the convenience of explanation, and do not limit the execution order of the steps. In actual operation, according to the technical needs of the specific implementation scene, the steps can be executed in a different order from that shown in the description, and in some cases parallel processing between steps can also be realized.

[0080] Embodiment one, describes the overall flow of the multi-model game method for multi-scene intelligent security management. In a specific implementation scene, for example, a network security defense system, the system needs to cope with attack behaviors under different types and different information conditions.

[0081] Specifically, as shown in Figure 1 , the steps include:

[0082] Step S101, acquiring multi-domain original data, and generating information structure features based on the multi-domain original data.

[0083] In this embodiment, the multi-domain original data can include user-generated content data in the content security domain, traffic data in the network telemetry domain, user behavior records in the access log domain, and posterior probability distribution of attack intent recognition model in the model output domain, etc. After acquiring the data, data cleaning is performed through time alignment, outlier elimination and field standardization, etc. to form a unified data stream. On this basis, the system will calculate two types of features: one is the information structure feature, such as the intent posterior entropy for measuring intent uncertainty, the data missing rate reflecting data integrity, the link delay quantile representing information timeliness and the distribution drift degree measuring the speed of environmental change; the second is the feature for risk assessment. Specifically, through the rolling statistics method, the key risk indicators (KRI) are continuously calculated, and the historical loss samples in the long time window are aggregated.

[0084] Step S102, according to the information structure features, performing game template routing, outputting the selected game template, unified semantic state and unified action set.

[0085] According to the generated information structure features, the information completeness and the antagonistic nature of the current scene are judged, and the most suitable game model is dynamically selected.

[0086] For example, when the information is close to complete observability, the system routes to the Stackelberg security game model; when the defense side does not fully grasp the information of the attacker, it routes to the Bayesian Stackelberg game model; when the behavior of the attacker can be regarded as a kind of strategic signal, the signal game model is selected. After selecting the template, in order to provide a standardized interface for subsequent steps, this step will also map the internal private state and action of the selected template to a globally unified, template-independent unified semantic state and unified action set.

[0087] Step S103, receive the unified semantic state, perform decision-sensitive subspace compression, and output the compressed state.

[0088] Considering that the dimension of the unified semantic state can be very high (i.e. super high dimension), directly calculating in this space will result in low efficiency and instability. Therefore, this step identifies a number of key state dimensions that have the greatest impact on the final game decision (i.e. security utility or payment), and projects the original unified semantic state onto a low-dimensional subspace composed of these key dimensions to obtain the compressed state. This dimension reduction method is not based on traditional variance interpretation (such as PCA), but directly faces the payment structure of the game, while reducing the computational complexity, it retains the important information required for decision-making.

[0089] Step S104, receive the compressed state, perform cross-time domain risk injection, and generate step-by-step risk signals.

[0090] In security management, the system not only optimizes the current benefits, but also must avoid possible major losses in the future (i.e. long cycle risk, such as conditional value at risk CVaR). However, long cycle risk cannot be directly used in step-by-step decision optimization algorithms. This step uses a time-consistent dynamic risk measurement method to decompose the overall long cycle risk into a series of step-by-step risk signals that can be directly used for each step decision, in a way that satisfies the total conservation. The signal will then be injected into the subsequent game solving process.

[0091] Step S105, receive and based on the compressed state, the unified action set, the step-by-step risk signal and the game template, perform online game solving and equilibrium vulnerability management, output security decision and management strategy and deploy.

[0092] Under the game template determined in the foregoing steps, in the compressed low-dimensional state space, combined with the step-by-step risk signal, online strategy optimization is performed to approximately solve the current Nash equilibrium point and form an executable security decision. In order to prevent the system from collapsing near the equilibrium point due to small perturbations, this step will calculate the equilibrium vulnerability index in real time, and once the index exceeds the threshold, it will immediately generate a management strategy (for example, temporarily freeze the switching of the template) and deploy it.

[0093] Exemplarily, the method can further comprise the steps of receiving a governance policy, and adjusting the game template routing, the decision-sensitive subspace compression and the cross-time domain risk injection on-line according to the governance policy.

[0094] The governance policy generated in S105 is written back to the preceding S102, S103 and S104 modules in real time to intervene and adjust their behaviors. For example, a governance policy can require S102 to widen the hysteresis interval for selecting a game template to increase stability, or require S104 to apply stronger low-pass filtering to the step-level risk signal to suppress disturbances.

[0095] The embodiment builds a complete closed loop from data perception, model selection, space compression, risk injection to governance solution. After S105 outputs a safe decision and is deployed, its execution trajectory and governance log are written back to the unified data bus as part of the original data, becoming the starting point of data for the next decision cycle, realizing continuous and adaptive closed-loop optimization.

[0096] Embodiment two, this embodiment details the specific process of game template routing according to information structure characteristics, solves the problems of how to quantify the quality indicators of information from different sources and how to avoid frequent switching near the decision boundary, and realizes accurate and stable routing of game templates.

[0097] Specifically, as shown in Figure 4 , this step comprises:

[0098] In step S201, the heterogeneous indicators in the information structure characteristics are uniformly mapped into information loss components, and the heterogeneous indicators include at least two of the intention posterior entropy, the data missing rate, the link delay and the distribution drift.

[0099] The information loss component is a dimensionless value, which is used to measure the quality degradation of the current information indicators compared with the ideal benchmark. In this embodiment, this mapping process is preferably performed after the dimensionless and robust filtering preprocessing of the information structure characteristics, to obtain a standardized feature vector that is resistant to noise and avoids the bias caused by direct comparison of different dimensions.

[0100] Specifically, the mapping of the heterogeneous indicators into the information loss components is realized based on the log-likelihood ratio or relative entropy. For example:

[0101] For the intention posterior entropy, the baseline distribution q _base (representing the entropy when there is no attack or random guessing) is taken as a reference, and the current posterior distribution p _now is calculated relative to q _baseThe relative entropy (Kullback-Leibler divergence) yields the information deficit component z. _entropy z _entropy The larger the value, the higher the uncertainty of the current intention and the more severe the information loss.

[0102] The data missing rate can be modeled as a Bernoulli channel with an ideal sampling rate p. _base Using 100% as a baseline, calculate the log-likelihood ratio corresponding to the currently observed missing rate to obtain the information loss component z. _missing .

[0103] For link delay, with baseline delay distribution g _base For reference, the log-likelihood ratio of the currently observed lag quantile (e.g., the 95th quantile) under this baseline distribution is calculated to obtain the information loss component z. _Latency This value reflects the loss of effectiveness due to outdated information.

[0104] To address distribution drift, the distribution p of the historical reference data is calculated. _base With the current data distribution p _now The relative entropy or its robust approximation (such as JS divergence) between the two yields the information loss component z. _drift .

[0105] As a preferred implementation, to address situations where data is sparse or its distribution is unknown in the real world, this step can also employ a more robust estimation method. For example, when samples are scarce, kernel density estimation can be used for density-related calculations (such as relative entropy), and adaptive bandwidth selection and extremum robustness processing can be performed to obtain a conservative upper bound on the information loss, avoiding underestimation of information loss due to insufficient data.

[0106] Furthermore, before mapping the data missing rate to an information loss component, this method also includes: identifying whether the missing mechanism of the data is non-random missing; and when the missing mechanism is non-random missing, using a selection model correction or importance weighting method to correct the bias of the data missing rate.

[0107] This step addresses the technical problem that systematic data gaps (non-random gaps) contain more hidden information about the system's state than random gaps, and failing to distinguish between them can severely mislead information quality assessments. For example, the fact that a certain type of attacker always disables telemetry reporting before launching a specific attack is itself a signal. In practice, the system uses statistical tests and other methods to determine the gap mechanism. If it is identified as a non-random gap, then in calculating z... _missing First, the missing rate is corrected by selection model (such as Heckman selection model) or importance weighting to compensate for the systematic bias it introduces.

[0108] Step S202, aggregate the information deficit components to generate an information structure index. In this embodiment, the aggregation step employs a joint aggregation model containing linear terms and quadratic interaction terms.

[0109] Specifically, the multiple information deficit components obtained in S201 are constituted into a vector:

[0110] z _vector =[z _entropy ,z _missing ,z _Latency ,z _drift ]。

[0111] The final information structure index I _value is calculated by:

[0112] I _value =W T *z _vector +z _vector T *M*z _vector ; where W is a linear weight vector, W T *z _vector is a linear term for assessing the independent impact of each information deficit component; M is an interaction matrix, z _vector T *M*z _vector is a quadratic interaction term for capturing the coupling effect between different information deficit components. For example, high latency and high distribution drift occurring simultaneously can mean that the attacker is conducting environmental probing while implementing jamming, whose combined risk is much greater than the sum of the two independent risks, and this coupling effect can be captured by the corresponding off-diagonal element in the interaction matrix M. The linear weight vector W and the interaction matrix M can be derived from offline calibration and fine-tuned online in small steps.

[0113] Step S203, select a game template according to the information structure index, and perform game template routing.

[0114] In this embodiment, this selection step is not simply comparing the instantaneous I _value with a threshold value, but adopts a more stable sequential statistics and hysteresis decision mechanism. Specifically, the information structure index is time-accumulated in a sequential statistics manner to obtain a sequential statistic. For example, the cumulative sum of generalized likelihood ratios can be used to accumulate I _value at consecutive time steps, and an exponential decay factor is added to forget outdated information, to obtain the sequential statistic I _seqThis time accumulation method enhances the sensitivity to persistent information structure anomalies while filtering out transient glitch fluctuations. Further, the threshold with hysteresis band is used to make decision on the sequential statistics, and select the game template.

[0115] Specifically, the system will preset at least two thresholds, a high threshold T _H (the first threshold) and a low threshold T _L (the second threshold smaller than the first threshold) (T _H > T _L ), and a hysteresis band is formed between them. The decision rule is as follows:

[0116] When I _seq crosses T _H from bottom to top, the system switches to the game template with less complete information (for example, from Stackelberg to Bayesian Stackelberg).

[0117] When I _seq crosses T _L from top to bottom, the system switches back to the game template with more complete information.

[0118] When I _seq fluctuates between T _L and T _H , the system keeps the current template unchanged. The decision mechanism with hysteresis band effectively avoids the problem of frequent switching of game templates (i.e. jitter) caused by slight fluctuations of information structure index near a single threshold, and improves the stability of the whole system.

[0119] As an enhanced function linked to subsequent governance steps, this embodiment can also evaluate the upper bound of the worst-case strategy loss caused by possible misrouting and record it in the governance log after generating I _value and I _seq each time. This upper bound will be used in the equilibrium vulnerability governance module of Embodiment Six to determine whether temporary freezing of routes or dynamic widening of hysteresis band is needed.

[0120] Embodiment Three, this embodiment describes the following process in detail: output the selected game template, unified semantic state and unified action set. Solve the state space and action set heterogeneity problem caused by dynamically selecting different game templates, and provide a stable and standard internal interface for the whole method.

[0121] Optionally, in this embodiment, after step S102 (or S203) selects a specific game template, the system will perform template unification mapping and semantic alignment, which specifically includes:

[0122] State semantics alignment is performed to generate a unified semantic state. This step is to map the private state representation within any selected template (template private state) to a predefined, global unified, template-agnostic standardized state space (unified semantic state). The unified semantic state is a vector space with fixed dimensions and constant physical meaning for each dimension. For example, the 1st dimension of this space always represents the attacker capability assessment score, the 2nd dimension always represents the current network load rate, and so on.

[0123] In implementation, the alignment process is performed according to a preset invariant list. The invariant list defines the system core attributes and constraints that must remain unchanged during the mapping process, such as the total amount of resource quota, absolute forbidden items of security policy, reachability boundary of actions, etc. The list provides basic constraints for subsequent mapping rules.

[0124] On this basis, the alignment process is implemented in one or more of the following ways:

[0125] Field alignment is performed, specifically: matching the fields in the template private state (e.g., attacker_type_posterior in Bayesian game) with the semantic closest dimension in the unified semantic state (e.g., attacker intention posterior distribution).

[0126] Unit conversion is performed, including: converting the units used in the template (e.g., delay in milliseconds) to standard units (e.g., seconds) to eliminate dimensional differences.

[0127] Value domain standardization is performed, including: normalizing or Z-score standardizing (standard deviation standardization) different ranges of numerical values to a standard interval, such as [0, 1] or [-1, 1].

[0128] Decomposition and aggregation are performed, including: for composite fields that cannot be directly one-to-one mapped, decomposition or aggregation strategies are used. For example, the host security score field in the template may need to be decomposed into multiple dimensions such as CPU utilization, memory occupancy, and vulnerability number in the unified semantic state; conversely, multiple template private telemetry indicators may need to be weighted and aggregated into a system health dimension in the unified semantic state.

[0129] Further, action semantic alignment and reachability check are performed to generate a unified action set. Similar to state alignment, this step maps the action set private to a template (e.g., actions {Block_IP, Throttle_Port} for a firewall template) to a predefined, globally unified action set. Reachability check is performed after the mapping, i.e., verifying that the mapped action sequence still reaches the same feasible state region as the original template in the unified semantic space, ensuring that the mapping does not break the intrinsic dynamics logic of the original model.

[0130] As a preferred implementation, to ensure the reliability and safety of the mapping process, the embodiment further includes the following quality assurance steps:

[0131] Constraint consistency check is performed after the generation of the unified semantic state and action set, the system rechecks all the safety constraints defined in the invariant list. For example, checking whether the numerical bounds, logical mutual exclusion relationships (such as not being able to simultaneously enable two conflicting defense strategies) are inadvertently relaxed after mapping. If any constraint is detected to be relaxed, the system will automatically back off the mapping strategy, for example, preferentially using a more conservative aggregation method or tightening the value range to ensure safety.

[0132] Further, compilation loss evaluation and upper limit control are performed. It can be understood that any mapping may introduce information loss, i.e., compilation loss. This step calculates the maximum compilation loss introduced by the current mapping scheme online through replay testing or difference evaluation on a small batch of samples, and compares it with the acceptable upper limit preset by the business. If the loss exceeds the upper limit, it indicates that the current mapping scheme (e.g., a certain aggregation weight setting) may cause decision information distortion, and the system will trigger a rollback to switch to a more faithful but possibly more complex mapping scheme.

[0133] Through the above steps, the embodiment generates a standardized state and action interface, and through multiple verification and quality control mechanisms, the reliability, safety and fidelity of the interface are ensured.

[0134] Embodiment Four, detailed description of the specific implementation process of decision-sensitive subspace compression. This embodiment receives the unified semantic state generated by Embodiment Three, solves the dimension disaster problem caused by high-dimensional state space, i.e., high computational cost, sparse samples, and unstable learning. The basis for dimension reduction is not the statistical variance of data, but the actual impact of each state dimension on the final game payoff or safety utility.

[0135] Specifically, as shown in Figure 5 This step includes:

[0136] Step S401, evaluate the sensitivity of each state dimension in the unified semantic state to the security utility or payoff of the game to form a sensitivity profile. The sensitivity profile is a data structure that assigns a comprehensive score to each dimension of the unified semantic state, quantifying its importance. Preferably, the score includes not only the instantaneous sensitivity, but also its rate of change and confidence.

[0137] In implementation, the evaluation process is as follows: the system extracts a sequence of state-action-reward triples from the recent time window as the evaluation sample. For each state dimension, a small perturbation budget is set, and the sensitivity is measured dimension by dimension using finite difference method or influence function method. For example, for dimension d, a small positive perturbation +ε and a small negative perturbation -ε are applied to it while keeping other dimensions unchanged, and the expected change in the game payoff is observed and calculated, which is the initial value of the sensitivity of dimension d. Exponential smoothing is performed on the time series of the sensitivity to obtain a more stable sensitivity and its rate of change. Bootstrap method can be used to give a confidence interval for the sensitivity of each dimension, and the sensitivity, rate of change, and confidence are combined into the final sensitivity profile.

[0138] The above method directly quantifies the causal relationship between each state dimension and the final decision goal, so that the dimensions retained later are valuable dimensions for decision-making.

[0139] Step S402, project the unified semantic state according to the sensitivity profile to output the compressed state. This step is based on the sensitivity profile generated in S401, filters out the key dimensions, and constructs a projection operator to map the high-dimensional unified semantic state to a low-dimensional subspace composed of key dimensions.

[0140] In a preferred embodiment, to ensure robustness and effectiveness, the projection process includes the following sub-steps:

[0141] Bootstrap sampling and voting integration are performed on the sensitivity profile to generate a robust dimension ranking. That is, the system will resample the sensitivity profile multiple times (e.g. 100 times), and each time a dimension importance ranking is generated. The rankings of each round are integrated by voting method (e.g. Borda count method) to obtain a final dimension ranking that is not sensitive to occasional samples and very robust.

[0142] On this basis, a ranking hysteresis threshold is set to update the key dimension set used for projection when the robust dimension ranking changes. The frequent jitter of the key dimension set can be suppressed to enhance the stability of the system. Specifically, only when the dimension not currently in the key set has a robust ranking that exceeds the lowest ranked dimension in the set, and the number of ranks exceeds the preset rank hysteresis threshold (e.g. 5 ranks), will the dimension be added to the key dimension set.

[0143] Further, in order to ensure the interpretability and generalization ability of the projected state, the step of updating the key dimension set for projection further comprises: imposing a diversity constraint to limit the number of dimensions selected from the same semantic cluster into the key dimension set based on the semantic similarity between dimensions, and ensuring the balanced distribution of the key dimension set among different semantic categories. For example, all state dimensions can be pre-divided into semantic clusters such as network traffic features, host behavior features, and attack load features. When screening the key dimensions, the system will impose constraints, such as stipulating that no more than 5 dimensions can be selected from the network traffic feature cluster. This avoids the loss of perception of other aspects due to the concentration of key dimensions in a single semantic category caused by high correlation between features.

[0144] On this basis, a projection operator is constructed based on the key dimension set, the condition number of the subspace defined by the projection operator is calculated, and the condition number is ensured not to exceed the preset upper limit, so as to generate a numerically stable compressed state. The condition number is an index for measuring the numerical stability of a mathematical problem, and a large condition number means that a small input error will be amplified sharply, resulting in unreliable calculation results, which may lead to gradient explosion or disappearance in machine learning. In this step, the system generates an orthogonal basis of the projection operator based on the covariance matrix of the key dimension set through a regularized decomposition method (such as principal component analysis with a regularization term). The condition number of the subspace formed by the basis is calculated immediately after generation. If the condition number exceeds the preset upper limit (for example, 1000), it indicates that the subspace is ill-conditioned and unstable. At this time, the system will trigger a rollback mechanism, such as increasing the regularization strength or discarding the key dimension with the lowest contribution, and re-construct and calculate until the condition number meets the requirements.

[0145] Optionally, after generating a new projection operator, the change in learning variance before and after projection can also be evaluated by replaying a small batch of historical samples. If the variance is found to be abnormally amplified, the system can be rolled back to the last version of the stable projection operator to ensure smooth iteration of the entire learning system.

[0146] This embodiment not only realizes efficient dimension reduction, but also ensures the high quality of the compressed state from multiple angles such as the robustness, diversity, and stability of dimension screening, and the numerical stability of the final projection.

[0147] Embodiment Five, this embodiment describes the process of cross-time domain risk injection in step S104 of embodiment one in detail. It solves the technical problem of how to reasonably incorporate long-period, complex risk metrics (e.g., conditional value at risk CVaR) into a step-by-step online learning and decision-making framework. Existing methods usually use simple averaging processing, which can destroy the consistency of risk in the time dimension, leading to oscillation of the strategy learning process or distortion of the perception of rare risk events. This embodiment solves the above problems through a dynamic, risk-difference method that satisfies the total conservation. Specifically, as shown in Figure 2 This step includes:

[0148] Step S501, construct the conditional distribution of the future loss sequence, and set a time-consistent dynamic risk metric. The system estimates the probability distribution of future losses under different conditions based on the long-time window historical loss samples and key risk indicators aggregated in embodiment one (S101) through non-parametric methods such as kernel smoothing.

[0149] As a preferred implementation, this process includes an adaptive confidence level generation mechanism, which includes: analyzing key risk indicators derived from multi-domain raw data, identifying event mutations, and dividing time series into event segments and stationary segments; in operation, according to the volatility intensity of the event segment, determine the adaptive confidence level.

[0150] Specifically, the system monitors key risk indicators (e.g., the frequency of high-risk attack types) online through a change point detection algorithm, and when a mutation is detected, marks the subsequent period of time as an event segment, and the rest as a stationary segment. Accordingly, the system will generate an adaptive confidence level α _adapt that changes over time. In the stationary segment, α _adapt maintains a baseline level (e.g., 0.95); while in the event segment, α _adapt will be dynamically adjusted higher (e.g., to 0.99), indicating that the system should adopt a more conservative and risk-averse attitude during this period. To enhance the accuracy of extreme event assessment, the system can also introduce extreme value theory or quantile correction in the tail of the loss distribution in the event segment for robust processing.

[0151] Further, a time-consistent dynamic risk metric is selected, such as the nested conditional value at risk (Nested Conditional VaR / CVaR). Time consistency is a key mathematical property of this risk metric, which ensures that the evaluation of future risk at any time point does not contradict the evaluation of more future risk at a future time point, and is the theoretical basis for subsequent dynamic programming and recursive calculation.

[0152] Step S502: Based on the conditional distribution and dynamic risk measurement, assess the dynamic risk-to-go (or dynamic risk residual) at the current moment. Dynamic risk-to-go (denoted as risk) _to_go_t The definition is: given all available information up to time t, using the dynamic risk metric selected in S501, the loss sequence L from the current time t to the future endpoint T is calculated. _t_to_T The total amount of risk assessed. Unlike traditional expected loss (cost-to-go), dynamic risk-to-go assesses risk rather than expectation, and is better able to capture the impact of low-probability, high-loss events.

[0153] In a preferred embodiment, assessing the dynamic risk-to-go at the current moment includes: setting endpoint boundary conditions and performing backward recursion starting from the endpoint boundary conditions to calculate the dynamic risk-to-go at each moment in the time series. Specifically, the boundary condition for endpoint T is set as risk. _to_go_t =0, because at the final moment, there is no future loss, so the risk is zero. Starting from t=T-1, the time axis is recursively extended to the current moment. At each moment t, the system uses the future loss conditional distribution constructed by S501 under the information conditions of that moment t to calculate the risk. _to_go_t To ensure numerical stability during the computation, the system can trigger a rollback mechanism when numerical divergence or abnormal tail weights are detected during the recursion process. This mechanism can temporarily reduce the tail amplification factor, shorten the recursion window, or further increase α. _adapt The conservatism continues until the calculation returns to stability.

[0154] Step S503: Calculate the conditional expectation of the dynamic risk-to-go at the next time step under the current information conditions. This step calculates E. _t [risk _to_go_{t+1} ], that is, the expected value of the total risk at time t and for the next time t+1.

[0155] Step S504: Differentiate the current dynamic risk-to-go value with the conditional expectation of the next dynamic risk-to-go value to generate a step-level risk signal. Step-level risk signal raw _penalty At time t, it is defined as: raw _penalty_t =risk _to_go_t -E _t [risk _to_go_{t+1} This mathematically guarantees time consistency. The immediate risk that needs to be borne at the current moment is equal to the total future risk at the current moment, minus the future risk that is expected to remain at the next moment after deducting the uncertainty of the current step.

[0156] According to an aspect of the present application, the method of cross-time domain risk injection further comprises: after generating the step-wise risk signal, checking whether the summation of the step-wise risk signal over the time series is conserved with the overall risk at the initial time calculated by the backward recursion.

[0157] Specifically, the system will distribute the computed raw _penalty_t risk signal over the time series, and verify whether the summation is equal to the overall risk at the initial time within an error threshold. _to_go_0 If the conservation is not satisfied, it is possible that the parameters of the conditional distribution estimation (such as the kernel bandwidth) are not appropriate, and the system will back off and adjust the parameters until the conservation error meets the requirement.

[0158] Step S505, the generated step-wise risk signal is processed and injected into the game solving process. After generating the raw _penalty risk signal, preferably, a first-order lag filter is applied to smooth the high-frequency noise and avoid excessive disturbance of the policy gradient, and the variance before and after filtering is recorded to ensure that key abnormal signals are not lost.

[0159] Further, the method further comprises: based on the compressed state, calculating the contribution coefficients of each dimension inside; according to the contribution coefficients, performing scale allocation on the step-wise risk signal to construct the step-wise risk penalty injected into the game solving objective; and monitoring the violation rate of the risk constraint to update the dual multiplier used to weigh the step-wise risk penalty online.

[0160] Specifically, the system reads the compressed state and its projection operator generated in embodiment four, calculates the contribution coefficients of each dimension in the compressed space (for example, based on variance contribution or sensitivity accumulation). The scalar raw _penalty signal is allocated according to these contribution coefficients to form a step-wise risk penalty vector with the same dimension as the compressed state, while a total conservation constraint is applied to ensure that the total risk after allocation remains unchanged.

[0161] Further, the system will continuously monitor the actual violation rate of the risk constraint, and adaptively adjust the dual multiplier used to weigh the risk penalty term according to the violation rate and the training stability.

[0162] For example, when the violation rate rises, the dual multiplier is increased to strengthen the penalty. Optionally, the update step of the policy and the update step of the dual multiplier can be set to different time scales, the former is smaller to ensure the learning stability, and the latter is larger to quickly respond to constraint violations, and the update scheduling of the two time scales can effectively avoid the high-frequency oscillation of the policy caused by the adjustment of the risk constraint.

[0163] Through the above steps, the long-period risk is injected into the step-by-step learning in a time-consistent manner in this embodiment, and through a series of designs such as adaptive confidence level, total amount conservation verification, multi-time scale updating, and accurate allocation to decision-sensitive dimensions, the accuracy, stability and effectiveness of the risk injection process are ensured.

[0164] Embodiment six, this embodiment describes in detail the process of online game solving and equilibrium vulnerability governance in step S105 of embodiment one. In the adversarial game, it is not enough to only solve the Nash equilibrium point, some equilibrium points may be very fragile or unstable, and a small strategic exploration of the opponent or a slight disturbance of the environment may lead to a sharp change in strategy, and even system collapse. This embodiment provides an online computable and forward-looking equilibrium point stability early warning mechanism, and automatically executes closed-loop governance when identifying vulnerability risks.

[0165] In a specific scenario, the system continuously updates the strategy and approaches the equilibrium point through online solving algorithms (such as online gradient descent or virtual game) under the guidance of the risk signal output by embodiment five. In this process, the governance module of this embodiment will run synchronously. As shown in Figure 3 The equilibrium vulnerability governance is performed, and the governance strategy is output, specifically including:

[0166] Step S601, according to the latest strategy and payment matrix output by the online game solving, a local linear approximation of the best response operator in the game is constructed to obtain its Jacobian matrix.

[0167] The best response operator is a function or mapping, whose input is the opponent's strategy, and whose output is the optimal strategy that can obtain the maximum benefit. The Jacobian matrix of the operator is a linear approximation of the operator near the current strategy point, which describes the sensitivity of the best response strategy to the small changes in the opponent's strategy. In specific implementation, numerical difference can be performed by applying a small disturbance to the opponent's strategy, or automatic differentiation tools can be used to quickly approximate the Jacobian matrix.

[0168] Step S602, estimate the spectral radius of the Jacobian matrix to quantify the local stability of the current strategy equilibrium point.

[0169] Spectrum radius refers to the absolute value of the largest eigenvalue of a matrix. In the dynamic analysis of game theory, the spectrum radius of the Jacobian matrix of the best response operator reflects the local amplification factor in the process of policy iteration. If the spectrum radius is less than 1, it means that any small deviation of the strategy will be gradually attenuated in the iteration, and the system will converge to the equilibrium point, and the equilibrium is stable; otherwise, if the spectrum radius is greater than 1, a small deviation will be amplified, leading to policy divergence, and the equilibrium is unstable. Since it is too costly to calculate the Jacobian matrix and solve all its eigenvalues in a joint field, this embodiment uses an efficient numerical approximation method, such as the damped power iteration method, which can quickly estimate the spectrum radius without explicitly constructing the entire Jacobian matrix. An alternative solution is to switch to the Chebyshev polynomial acceleration method when the local compression mapping condition is not met (i.e., the spectrum radius may be greater than 1), resulting in poor convergence of the power iteration method. A reliable upper bound estimate of the spectrum radius can be obtained, which is sufficient to drive the subsequent governance gate.

[0170] Step S603, based on the estimated result of the spectrum radius, an equilibrium vulnerability index is generated, and a governance strategy is generated when the index exceeds a preset threshold.

[0171] In a preferred embodiment, the step of generating the equilibrium vulnerability index further comprises: calculating a short window regret slope based on the strategy and the payoff trajectory; and the equilibrium vulnerability index is a weighted combination of the estimated result of the spectrum radius and the short window regret slope.

[0172] Wherein, the short window regret slope refers to the change rate of the cumulative regret of the party within the latest time window, which is calculated by linear regression. A positive regret slope indicates that the party's strategy is deviating from the optimal response, which is another important indicator of instability. The final equilibrium vulnerability index (EVI _index ) is obtained by weighted combination of the spectrum radius approximation and the regret slope.

[0173] Specifically, the system will set two thresholds: a lower observation threshold and a higher governance threshold. When the EVI _index exceeds the observation threshold, the system enters the observation state and records potential vulnerable points; when it exceeds the governance threshold, the governance action is triggered immediately.

[0174] When the index exceeds the preset threshold, the generated governance strategy includes at least one of the following: freezing the route of the game template, widening the hysteresis interval for selecting the game template, applying low-pass filtering to the step-level risk signal, or reducing the update step size of the dual multiplier.

[0175] The governance actions are collaboratively designed to cool down and stabilize the system from different levels. Specifically, temporarily prohibiting the game template switching in Embodiment Two avoids introducing additional non-stationarity at the critical stability point due to model switching. Dynamically adjusting the high threshold T _H and the low threshold T _L of the hysteresis decision in Embodiment Two to improve the stability margin of the decision. Applying stronger filtering to the stepwise risk signal generated in Embodiment Five to suppress the influence of high-frequency risk disturbance on the strategy gradient. Slowing down the update speed of the dual multipliers in Embodiment Five to reduce the tightening pressure of the risk constraint and give the policy learning more relaxed adjustment space.

[0176] After generating and applying the governance strategy, the method further includes: continuously monitoring the estimated result of the spectral radius and the short-window regret slope; and gradually revoking the actions contained in the governance strategy when the estimated result of the spectral radius returns to the safe interval and the short-window regret slope becomes insignificant.

[0177] Specifically, the system continuously monitors the decrease of EVI _index . Only when the approximation of the spectral radius clearly falls back to the safe interval (e.g., less than 0.9) and the regret slope is no longer statistically significantly different from zero, the system gradually and in a preset order revokes the governance actions started in S603 until the normal operation state is restored.

[0178] Through the above steps, the embodiment constructs a complete equilibrium point stability closed-loop governance system from online monitoring, quantitative early warning to collaborative governance and smooth exit, and moves the system stability guarantee from the post-response to the proactive and calculable intervention before the critical point occurs, thereby improving the robustness of the entire intelligent safety management method.

[0179] Embodiment Seven describes a specific case of cross-time-domain risk consistency injection method, and details the backward recursive calculation of dynamic risk-to-go and the differential generation process of stepwise risk signal.

[0180] Assume a simplified safety decision-making scenario with a timeline of 3 discrete time steps, i.e., t=1, 2, 3, where t=3 is the terminal point T. The system needs to evaluate the future risk at the beginning of each time step.

[0181] The Conditional Value-at-Risk (CVaR) is chosen as the time-consistent dynamic risk measure in this case. CVaR _α (L) represents the expected value of the loss L exceeding its VaR part at a confidence level α, i.e., the average loss under the worst (1-α) probability.

[0182] According to external intelligence, the system identifies t=1 as a stationary segment and t=2 as an event segment (e.g., a new type of attack is expected to emerge). Thus, the adaptive confidence level a _adapt is set as follows:

[0183] At t=1, a _1 = 0.90 (a more conventional risk aversion level is taken).

[0184] At t=2, a _2 = 0.95 (in the event segment, the risk aversion level is raised, and it is more conservative).

[0185] For simplicity of calculation, it is assumed that at each time point, there is a discrete, known conditional probability distribution for the loss L in the next step:

[0186] Under the information condition at t=1, the predictive distribution of the loss L _2 at t=2 is: P(L _2 = 10) = 0.8; P(L _2 = 100) = 0.2.

[0187] Under the information condition at t=2, the predictive distribution of the loss L _3 at t=3 is (in the event segment, the tail is heavier): P(L _3 = 20) = 0.9; P(L _3 = 500) = 0.1.

[0188] For simplicity, it is assumed that each step loss is conditionally independent. The calculation process is as follows:

[0189] Step A: Calculate the dynamic risk-to-go at each time point;

[0190] Time t=3 (end point T): According to the end boundary condition, there is no future loss, so: risk _to_go_3 = 0;

[0191] Time t=2: At t=2, the future loss sequence only contains L _3 . At this time, the system is in the event segment, and the confidence level a _2 = 0.95 is adopted. risk _to_go_2 = CVaR _0.95 (L _3 ), calculate CVaR _0.95 (L _3 ):

[0192] The loss distribution is {20 (probability 0.9), 500 (probability 0.1)}.

[0193] 1-a _2 = 0.05. The worst 5% of loss events are the part of loss 500.

[0194] Since P(L _3 =500)=0.1>0.05, VaR _0.95 (L _3 )=20.

[0195] CVaR _0.95 (L _3 ) is the expected value of the worst 5% of the distribution of L _3 . In this discrete case, the tail is completely made of the event of loss 500. Therefore, risk _to_go_2 =500.

[0196] At time t=1: At t=1, the future loss sequence is L _2 ,L _3 . At this time, the system is in a stationary segment, with a confidence level a _1 =0.90. The total future loss L _total_from_1 =L _2 +L _3 . We need to calculate its distribution:

[0197] P(L _total_from_1 =10+20=30)=0.8*0.9=0.72;

[0198] P(L _total_from_1 =10+500=510)=0.8*0.1=0.08;

[0199] P(L _total_from_1 =100+20=120)=0.2*0.9=0.18;

[0200] P(L _total_from_1 =100+500=600)=0.2*0.1=0.02 risk _to_go_1 =CVaR _0.90 (L _total_from_1 );

[0201] Calculate CVaR _0.90 (L _total_from_1 ): 1-a _1 =0.10. We need to find the worst 10% of the loss events.

[0202] Sort by loss from largest to smallest: {600 (probability 0.02), 510 (probability 0.08), 120 (probability 0.18), 30 (probability 0.72)}.

[0203] The worst events are loss 600 (probability 0.02) and loss 510 (probability 0.08), and the sum of their probabilities is 0.02 + 0.08 = 0.10. Thus, the worst 10% tail is exactly composed of these two events.

[0204] CVaR _0.90 That is, the conditional expectation of these two events:

[0205] E[L|Lintail] = (600*0.02 + 510*0.08) / (0.02 + 0.08) = (12 + 40.8) / 0.1 = 52.8 / 0.1 = 528.

[0206] Thus, risk _to_go_1 = 528.

[0207] Step B: Compute the stepwise risk signal;

[0208] The stepwise risk signal at time t = 2: raw _penalty_2 = risk _to_go_2 - E _2 [risk _to_go_3 ];

[0209] risk _to_go_2 = 500. risk _to_go_3 = 0 is a deterministic value, so its expectation E _2 [risk _to_go_3 ] under any information condition is 0. raw _penalty_2 = 500 - 0 = 500.

[0210] The stepwise risk signal at time t = 1: raw _penalty_1 = risk _to_go_1 - E _1 [risk _to_go_2 ]; risk _to_go_1 = 528. risk _to_go_2 = 500 is a deterministic value that can already be computed at time t = 1 (because it only depends on the information and the distribution of L _3 at time t = 2, which are known at time t = 1 under the assumption that they are known), so E _1 [risk _to_go_2 ] = 500. raw _penalty_1 = 528 - 500 = 28.

[0211] Step C: Check the total amount conservation;

[0212] The overall risk at the initial time is risk _to_go_1 = 528. The sum of the stepwise risk signals is:

[0213] raw _penalty_1 +raw _penalty_2 =28 + 500 = 528. Since the two are equal, 528 = 528, the total amount is conserved, and this is verified.

[0214] This case study describes the following: 1) How to calculate the dynamic risk-to-go at each time point through backward recursion; 2) How to calculate the risk-to-go at each time point based on the adaptive confidence level (α). _1 =0.90, α _2 =0.95) Adjust the conservatism of risk assessment; 3) How to generate step-level risk signals (28 and 500) through time difference; 4) The sum of the generated step-level risk signals equals the initial total risk, satisfying time consistency and total amount conservation.

[0215] This application employs the I-Router mechanism, which uniformly maps heterogeneous indicators such as posterior entropy of intent and data missing rate into information loss components, and aggregates them using a model containing linear and quadratic interaction terms, providing a scientific and quantitative basis for model selection. By using sequential statistical accumulation of information and employing a threshold with hysteresis intervals for decision-making, it suppresses the problem of frequent model switching caused by fluctuations in information indicators around the threshold, thus solving the problems of model mismatch and decision jitter.

[0216] This application defines the step-level risk signal as the difference between the conditional expectation of the current dynamic risk -to-go and the next dynamic risk -to-go through the Tτ mechanism. Based on the definitions of backward recursion and time difference, the temporal consistency of risk decomposition is theoretically guaranteed. Simultaneously, through a total amount conservation check step, the sum of the decomposed step-level risks is made equal to the initial global risk, avoiding distortion of the risk signal. This solves the problem of temporal inconsistency in risk injection.

[0217] This application employs the EVI mechanism, which constructs a local linear approximation (Jacobi matrix) of the optimal response operator and uses efficient numerical methods such as damped power iteration to estimate its spectral radius online, providing a forward-looking, online-calculated stability quantification index. Based on this index, the system can provide early warning when the equilibrium point becomes vulnerable but has not yet collapsed, and automatically trigger a series of closed-loop governance actions such as freezing routes and enhancing filtering, realizing a transformation from passive response to active defense.

[0218] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A multi-scene-oriented intelligent security management multi-model game method, characterized in that, The method comprises: acquiring multi-domain raw data and generating information structure features based on the multi-domain raw data; performing game template routing based on the information structure features, outputting selected game templates, unified semantic states, and unified action sets; receiving the unified semantic states, performing decision-sensitive subspace compression, and outputting compressed states; receiving the compressed states, performing cross-time-domain risk injection, and generating step-level risk signals; receiving and based on the compressed states, the unified action sets, the step-level risk signals, and the game templates, performing online game solving and equilibrium vulnerability management, and outputting safety decisions and management strategies.

2. The method of claim 1, wherein, The method of performing cross-time-domain risk injection and generating step-level risk signals comprises: constructing a conditional distribution of a future loss sequence and setting a time-consistent dynamic risk measure; based on the conditional distribution and the dynamic risk measure, evaluating the dynamic risk-to-go at the current time; calculating the conditional expectation of the dynamic risk-to-go at the next time under the current information condition; differencing the dynamic risk-to-go at the current time and the conditional expectation of the dynamic risk-to-go at the next time to generate the step-level risk signal.

3. The method of claim 2, wherein, The method of evaluating the dynamic risk-to-go at the current time comprises: setting a terminal boundary condition and performing backward recursion from the terminal boundary condition to calculate the dynamic risk-to-go at each time in the time sequence; The method further comprises: after generating the step-level risk signal, verifying whether the sum of the step-level risk signals in the time sequence is conserved with the overall risk at the initial time calculated by the backward recursion.

4. The method of claim 2, wherein, The dynamic risk measure includes an adaptive confidence level, and the generation of the adaptive confidence level comprises: analyzing key risk indicators derived from multi-domain raw data, identifying event mutations, and dividing the time sequence into event segments and stationary segments; in operation, determining the adaptive confidence level according to the fluctuation intensity of the event segment.

5. The method of claim 2, wherein, After generating the step-level risk signal, the method further comprises: based on the compressed states, calculating the contribution coefficients of each dimension inside the compressed states; based on the contribution coefficients, performing scale allocation on the step-level risk signal to construct a step-level risk penalty for injecting game solving objectives; monitoring the violation rate of the risk constraint and updating the dual multipliers used to weigh the step-level risk penalty based on the violation rate.

6. The method of claim 1, wherein, The method of performing equilibrium vulnerability management and outputting management strategies comprises: based on the latest strategies and payoff matrices output by the online game solving, constructing a local linear approximation of the best response operator in the game to obtain its Jacobian matrix; estimating the spectral radius of the Jacobian matrix; based on the estimation result of the spectral radius, generating an equilibrium vulnerability index, and generating a management strategy when the index exceeds a preset threshold.

7. The method of claim 6, wherein, The method of generating the equilibrium vulnerability index further comprises: based on the strategy and payoff trajectory, calculating a short-window regret slope; the equilibrium vulnerability index is a weighted combination of the estimation result of the spectral radius and the short-window regret slope; when the index exceeds the preset threshold, the generated management strategy comprises at least one of the following: freezing the routing of the game template, widening the hysteresis interval for selecting the game template, applying low-pass filtering to the step-level risk signal, or reducing the update step of the dual multiplier.

8. The method of claim 1, wherein, The method of performing game template routing based on the information structure features comprises: The heterogeneous indexes in the information structure feature are uniformly mapped into information loss components, the heterogeneous indexes including at least two of the following: intention posterior entropy, data missing rate, link delay, and distribution drift; The information structure index is generated by fusing the information loss components; The game template is selected according to the information structure index, and the game template routing is performed.

9. The method of claim 8, wherein, The uniform mapping of the heterogeneous indexes into the information loss components is realized based on the log-likelihood ratio or relative entropy; And, the selection of the game template according to the information structure index includes: The sequential statistics method is used to accumulate the information structure index in time to obtain a sequential statistic; The sequential statistic is judged by using a threshold with a hysteresis interval to select the game template.

10. The method of claim 1, wherein, The decision-sensitive subspace compression is performed to output a compressed state, including: The sensitivity of each state dimension in the unified semantic state to the safety utility or payment of the game is evaluated to form a sensitivity spectrum; The unified semantic state is projected according to the sensitivity spectrum to output the compressed state.

Citation Information

Cited By

  • Track prediction model robustness enhancement method based on dynamic subspace projection decomposition

    CN121637493A

  • Multi-agent-based chronic disease risk management and hierarchical diagnosis decision-making method

    CN122050821A

  • A Trusted Routing Method and System for Large Language Models

    CN122413423A