A large model-based network operation state anomaly management system and method
Patent Information
- Application Number
- CN202611014963.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-09
AI Technical Summary
综上所述,现有网络异常管理技术在检测与定位的协同性、未知故障的应对能力、修复操作的安全保障以及大模型与物理规律的融合深度等方面均存在明显不足,亟需一种能够打通感知-定位-决策-验证全链路、兼具物理可解释性与数学安全保障的异常管理方法
1、本发明通过将网络运行数据映射至高维特征空间,以历史无故障特征库为基准,利用主成分分析的特征值熵变实时计算当前状态相对于历史正常状态的偏离程度,并采用动态更新的曲率阈值进行异常判定,能够在异常尚未形成明显指标越限之前即捕捉到系统状态的微小漂移,实现早期预警,同时有效降低误报率;
Smart Images

Figure CN122533920B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network anomaly management technology, specifically a network operation status anomaly management system and method based on a large model. Background Technology
[0002] With the rapid development of technologies such as 5G, cloud computing, and microservices, the structural complexity and technical component diversity of network systems continue to increase, making traditional operation and maintenance methods insufficient to meet the requirements of real-time monitoring and efficient operation and maintenance. Currently, the network operation and maintenance field mainly suffers from the following technical shortcomings: Existing network anomaly management mainly relies on two technical approaches: one is a fixed rule and threshold alarm system based on expert experience. This type of method has low root cause localization accuracy when facing cross-professional cascading failures, and the rules are severely lacking in adaptability, making it difficult to cope with dynamically changing network environments; the other is anomaly detection models based on machine learning or deep learning. This type of method highly depends on a large amount of high-quality labeled data for supervised training, but the operation and maintenance field lacks sufficient labeled samples, resulting in insufficient model generalization ability and long cross-scenario adaptation cycle. While recent research has attempted to introduce large language models into network operations and maintenance, enabling alarm interpretation and operational suggestion generation through natural language interaction, such applications essentially remain at the level of text-based question-and-answer. The large model serves merely as an add-on tool for "knowledge retrieval and text generation," neither participating in the numerical perception of network status, nor engaging in the quantitative calculation of root cause localization, nor undertaking the mathematical verification of repair strategies. A deeper problem lies in the fundamental paradigm conflict between the probabilistic generation mechanism of the large model and the deterministic physical laws governing network operation. In summary, existing network anomaly management technologies have significant shortcomings in terms of the coordination of detection and localization, the ability to respond to unknown faults, the security of repair operations, and the depth of integration between large models and physical laws. There is an urgent need for an anomaly management method that can connect the entire link of perception-localization-decision-verification and has both physical interpretability and mathematical security. Summary of the Invention
[0003] The purpose of this invention is to provide a network operation status anomaly management system and method based on a large model, so as to solve the problems raised in the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a method for managing network operation status anomalies based on a large model, specifically including the following steps: Step S1: Obtain multi-source time-series data of network operation status, map the multi-source time-series data onto a representation manifold in a high-dimensional feature space, calculate the local geometric curvature change of the representation manifold between the current time and historical time; if the local geometric curvature change exceeds a preset dynamic curvature threshold, determine that the network has an anomaly and generate an anomaly trigger signal. Step S2: Model based on queue occupancy length and signal propagation delay respectively, and use the modeling results as a differentiable operator layer to construct a virtual forward computation graph; the differentiable operator layer is used to establish a differentiable mapping relationship between network input state variables and network output performance index variables; Step S3: In response to the anomaly trigger signal, the deviation between the actual observed value and the expected standard value of the network output performance index variable is defined as the loss function; the loss function is backpropagated through the virtual forward computation graph to obtain the loss gradient value corresponding to each network node; based on the spatial distribution characteristics of the loss gradient value, the set of abnormal root cause nodes is selected from all the network nodes. Step S4: Based on the set of abnormal root cause nodes and the current network state variables, call the preset large language model to output the recommended iterative target strategy in the preset control action space.
[0005] Further, in step S1, multi-source time-series data of the network's operating status is acquired, and the multi-source time-series data is mapped onto a representation manifold in a high-dimensional feature space. The local geometric curvature change of the representation manifold between the current time and historical time is calculated. When the local geometric curvature change exceeds a preset dynamic curvature threshold, an anomaly is determined to have occurred in the network, and an anomaly trigger signal is generated. Specifically: Step S1-1: Collect multi-source time-series data from each switch and router in the target physical network. For the i-th node in the network, the raw data vector collected at time t is represented as: x i (t)=[q i (t),d i (t),u i (t),l i (t),δ i [(t)]; where q i (t) represents the instantaneous length occupied by the output queue; d i (t) represents the average forwarding delay of data packets; u i (t) represents the port bandwidth utilization; i (t) represents the packet loss rate; δ i (t) represents the link error rate; the data from all N network nodes at consecutive T time points are stacked along the time dimension to construct the original observation tensor: X(t)∈R T×N×5 Where T is the length of the sliding time window; Step S1-2: Map the original data vector from the low-dimensional physical space to a high-dimensional feature space of dimension H, which is H=64 in this embodiment, to obtain the system state representation matrix: Z(t)=f emb (X(t))∈R N×H Z(t) represents the current state point; the set of all possible system states Z in the high-dimensional space is defined as a closed space M; It should be noted that the set of all possible system states in the high-dimensional space is a closed space, that is, the low-dimensional geometry formed by the trajectory of the network's normal operating state in this space is a closed space. Step S1-3: On the M, take the current state point Z(t) and its K nearest neighbors from adjacent historical times, where K is the logarithm of the number of network nodes N, i.e. ; Calculate the covariance matrix C=(Z) of point Z(t). k -Z') T (Z k -Z') / K; where Z k Let C represent the high-dimensional feature space of the k-th nearest neighbor, k∈K; Z' represents the mean of the high-dimensional feature space in history; perform eigenvalue decomposition on C to obtain H eigenvalues λ1≥λ2≥...≥λH; substitute the eigenvalues into the standardized curvature estimation formula to obtain a scalar characterizing the degree of local curvature, denoted as curvature value Ric(t), and define the change in local geometric curvature between the current time and the previous time as: ΔRic(t)=|Ric(t)-Ric(t-1)|; Step S1-4: Obtain the historical curvature value sequence {Ric(t-1),Ric(t-2),...,Ric(tW)}; where Ric(t-1),Ric(t-2),...,Ric(tW) represent the W consecutive curvature values before time t; calculate the moving average μRic and standard deviation σRic of the sequence in real time; set the dynamic curvature threshold as: Threshold(t)=μRic+β×σRic; where β is an adjustable sensitivity coefficient; when ΔRic(t)>Threshold(t), it is determined that the current network operation state has deviated from the normal manifold and an anomaly has occurred, and the system immediately generates a binary anomaly trigger signal Flag=1.
[0006] Furthermore, in step S2, models are performed based on queue occupancy length and signal propagation delay, and the modeling results are used as a differentiable operator layer to construct a virtual forward computation graph; the differentiable operator layer is used to establish a differentiable mapping relationship between network input state variables and network output performance index variables; specifically: The queue occupancy length variation of each network node in the target physical network is modeled as a first-order ordinary differential equation discretized computational unit with time as the independent variable; the signal propagation delay of each physical link in the target physical network is modeled as a linear mapping computational unit with link length as the independent variable; the first-order ordinary differential equation discretized computational unit and the linear mapping computational unit are embedded as fixed, non-trainable operator layers into a virtual forward computation graph composed of fully connected layers and residual connections; the input layer of the virtual forward computation graph is defined as the adjustable configuration parameter vector of each network node, and the output layer is the prediction performance index vector of the entire network; The composite mapping function from the input layer to the output layer in the virtual forward computation graph can be analyzed using the chain rule for the partial derivatives of any variable in the input layer.
[0007] Furthermore, the specific steps of modeling the queue occupancy length variation of each network node in the target physical network as a first-order ordinary differential equation with time as the independent variable into a discretized computational unit include: For any network node i, the predicted queue occupancy length q at discrete sampling time p i (p+1) is: q i (p+1)=q i (p)+[λ i (p)-μ i (p)]×ΔT; where λ i (p) represents the input flow rate value of network node i at discrete sampling time p, μ i (p) is the output link service rate value of the network node i at the discrete sampling time p, and ΔT is the preset discretization time step; The recursive equation relates to the output link service rate value μ. i The first-order partial derivative of (p) is -ΔT, and the partial derivative is always a negative constant when the queue length is greater than zero.
[0008] Furthermore, the step of modeling the signal propagation delay of each physical link in the target physical network as a linear mapping calculation unit with link length as the independent variable specifically includes: Construct an adjacency matrix A of the target physical network, wherein the adjacency matrix A is an N×N square matrix, and the elements A ij =1 indicates that there is a direct physical link between node i and node j, A ij =0 indicates that no direct physical link exists; construct the link length matrix L of the target physical network, where the link length matrix L is an N×N square matrix, and the elements L ij The value of is the actual length in kilometers of the physical link between node i and node j, when A ij When =0, the corresponding L ijThe value is zero; the signal propagation delay matrix T prop The characterization formula is denoted as: T prop =A⊙L / v; where the symbol ⊙ indicates the element-wise multiplication of the matrix; v is the constant propagation speed of electromagnetic waves in the physical medium; The signal propagation delay matrix T prop Regarding any non-zero element L in the link length matrix L ij The partial derivative is a constant 1 / v.
[0009] Furthermore, the composite mapping function of the virtual forward computation graph is expressed as: y pred =D(sin,Φ); where sin is an input layer variable, including the output link service rate values {μ1,μ2,...,μ} of all network nodes. N}; Φ is the set of trainable weight parameters for the fully connected layer; y pred For output layer variables, including the predicted end-to-end latency T for the entire network. prop ; The composite mapping function D is continuous and differentiable everywhere in the input domain, and its gradient vector with respect to any input variable θ∈{sin,Φ} The propagation channel used as the loss function for backpropagation in step S3.
[0010] Further, in step S3, in response to the anomaly trigger signal, the deviation between the actual observed value and the expected standard value of the network output performance index variable is defined as a loss function; the loss function is backpropagated through the virtual forward computation graph to obtain the loss gradient value corresponding to each network node; based on the spatial distribution characteristics of the loss gradient values, a set of anomaly root cause nodes is selected from all the network nodes; specifically: Step S3-1: Based on the generated anomaly trigger signal, the system collects the actual observed values y of the performance index variables in the physical network at the time of the anomaly. obs And retrieve the expected standard value y under the same business load from the performance database. targe Define the deviation as the mean squared error loss function MSE: MSE = ||y obs -y targe ||2 2 ; Step S3-2: Using the scalar of the mean squared error loss function as the simulated endpoint of the virtual forward computation graph, trigger the backpropagation operation of the constructed composite mapping function D, and propagate the gradient back from the output to the input layer by layer according to the chain rule in calculus: ; where θ represents the internal state hidden variable h of each physical network node in the ergodic composite mapping function D. i , i=1,2,...,N; For each network node i, the L2 norm of its corresponding latent variable gradient vector is taken as the causal contribution gradient value g of that node. i : ; Calculate the gradient values {g1, g2, ..., g} of all N network nodes. N Based on the spatial distribution characteristics of the nodes, an adaptive threshold screening method is used to normalize the gradient values and select nodes whose normalized gradient values are greater than a preset proportional threshold to form the abnormal root cause node set Groot.
[0011] It should be noted that the internal state hidden variable h i This includes at least one of the following: current output link service rate, current input traffic arrival rate, current average processing latency of data packets within the node, current chip / optical module operating temperature, current port bit error rate, current queue congestion window pressure coefficient, and current output queue instantaneous occupancy length.
[0012] Furthermore, in step S4, based on the set of abnormal root cause nodes and the current network state variables, a preset large language model is invoked to output a recommended iterative target strategy within a preset control action space, specifically as follows: Step S4-1: Vectorize and concatenate the located abnormal root cause node set Groot and the current network's global state variable scurrent to construct the standard prompt word embedding vector of the large model; call the pre-trained large language model, which has a high-dimensional continuous control action space B. Each dimension in the action space B corresponds to an adjustable parameter in the network, and the value range is normalized to [0,1]. Step S4-2: The large language model generates initial candidate regulation strategies b0∈B; then it enters the iterative optimization loop: the current strategy b x Input the composite mapping function D, perform forward inference, and obtain the result at b x Output layer variable y under the strategy pred (b x )=D(b x And recalculate the new loss function value MSE. new =||y pred (b x )-y target ||2 2 ; MSE new As a hard constraint, it is fed back to the generative decoder of the large language model, and the policy gradient is calculated. The large language model is guided to adjust its strategy in the next round of output to reduce loss; the iterative objective function is defined as: Π=argmin b∈B ||D(b)-y target ||22 ; Π represents the iterative target strategy; D(b) represents the output layer variable based on the composite mapping function D under strategy b; Step S4-3: Repeat the process. When the rate of change of the loss function between two adjacent iterations is less than a preset threshold, or when the preset maximum number of iterations is reached, terminate the iteration and output the iteration target strategy Π.
[0013] A network operation status anomaly management system based on a large model, comprising: The data acquisition and sensing module is configured to collect multi-source operational data from network nodes, map it to the feature space and build a historical feature library, calculate the deviation of the current state from the historical normal state in real time, and generate an abnormal trigger signal when the deviation exceeds the dynamic threshold. The differentiable digital twin module is configured to construct a virtual forward computation graph consisting of a fixed physical operator layer and a trainable residual compensation layer connected in series. It is used to simulate the differentiable mapping relationship between the network state from the input configuration to the output performance index, supporting both forward inference prediction and backpropagation differentiation. The root cause localization module is configured to respond to anomaly triggering signals, use the deviation between the actual observed value and the expected standard value as a loss function, obtain the contribution of each node to the deviation through backpropagation of the differentiable digital twin, and thereby filter out the set of abnormal root cause nodes. The strategy generation module is configured with a finite number of preset discrete atomic actions, and the strategy with the smallest deviation between the deduction result and the expected value is used as the iterative target strategy.
[0014] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention maps network operation data to a high-dimensional feature space, uses a historical fault-free feature library as a benchmark, and calculates the deviation of the current state from the historical normal state in real time using the entropy change of the eigenvalues of principal component analysis. It also uses a dynamically updated curvature threshold for anomaly determination, which can capture the slight drift of the system state before the anomaly forms a significant index exceeding the limit, thus achieving early warning and effectively reducing the false alarm rate. 2. This invention defines network performance deviation as a loss function and uses backpropagation of differentiable twins to backpropagate the loss gradient to each network node, using the magnitude of the gradient vector as the causal contribution of each node to the performance deviation. Compared with traditional root cause localization methods based on alarm association rule mining or knowledge graph search, this invention transforms root cause localization into a tensor back-differentiation operation, the complexity of which does not increase exponentially with the growth of network size, and can achieve millisecond-level root cause localization. Moreover, the gradient backpropagation is based on the mathematical chain rule of physical twins, is not limited by unknown fault modes, and has zero-sample localization capability. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating a network operation status anomaly management method based on a large model according to the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example: Figure 1 As shown, the present invention provides a technical solution, a method for managing network operation status anomalies based on a large model, which specifically includes the following steps: Step S1: Obtain multi-source time-series data of network operation status, map the multi-source time-series data onto a representation manifold in a high-dimensional feature space, calculate the local geometric curvature change of the representation manifold between the current time and historical time; if the local geometric curvature change exceeds a preset dynamic curvature threshold, determine that the network has an anomaly and generate an anomaly trigger signal. Step S2: Model based on queue occupancy length and signal propagation delay respectively, and use the modeling results as a differentiable operator layer to construct a virtual forward computation graph; the differentiable operator layer is used to establish a differentiable mapping relationship between network input state variables and network output performance index variables; Step S3: In response to the anomaly trigger signal, the deviation between the actual observed value and the expected standard value of the network output performance index variable is defined as the loss function; the loss function is backpropagated through the virtual forward computation graph to obtain the loss gradient value corresponding to each network node; based on the spatial distribution characteristics of the loss gradient value, the set of abnormal root cause nodes is selected from all the network nodes. Step S4: Based on the set of abnormal root cause nodes and the current network state variables, call the preset large language model to output the recommended iterative target strategy in the preset control action space.
[0018] Further, in step S1, multi-source time-series data of the network's operating status is acquired, and the multi-source time-series data is mapped onto a representation manifold in a high-dimensional feature space. The local geometric curvature change of the representation manifold between the current time and historical time is calculated. When the local geometric curvature change exceeds a preset dynamic curvature threshold, an anomaly is determined to have occurred in the network, and an anomaly trigger signal is generated. Specifically: Step S1-1: Collect multi-source time-series data from each switch and router in the target physical network using the Simple Network Management Protocol (SNMP) at a preset sampling period. For the i-th node in the network, the raw data vector collected at time t is represented as: x i (t)=[q i (t),d i (t),u i (t),l i (t),δ i [(t)]; where q i (t) represents the instantaneous length of the output queue, expressed in bytes; d i (t) represents the average packet forwarding delay, expressed in microseconds; u i (t) represents port bandwidth utilization, expressed as a percentage; i (t) represents the packet loss rate, expressed in ppm; δ i (t) represents the link error rate, uniformly expressed as a percentage; data from all N network nodes at consecutive T time points are stacked along the time dimension to construct the original observation tensor: X(t)∈R T×N×5 Where T is the length of the sliding time window, such as 120 sampling points in this embodiment.
[0019] Step S1-2: Map the original data vector from the low-dimensional physical space to a high-dimensional feature space of dimension H, which is H=64 in this embodiment, to obtain the system state representation matrix: Z(t)=f emb (X(t))∈R N×H Z(t) represents the current state point; the set of all possible system states Z in the high-dimensional space is defined as a closed space M; It should be noted that the set of all possible system states in the high-dimensional space is a closed space, that is, the low-dimensional geometry formed by the trajectory of the network's normal operating state in this space is a closed space. Step S1-3: On the M, take the current state point Z(t) and its K nearest neighbors from adjacent historical times, where K is the logarithm of the number of network nodes N, i.e. ; Calculate the covariance matrix C=(Z) of point Z(t). k -Z') T (Z k -Z') / K; where Z kLet C represent the high-dimensional feature space of the k-th nearest neighbor, k∈K; Z' represents the mean of the high-dimensional feature space in history; perform eigenvalue decomposition on C to obtain H eigenvalues λ1≥λ2≥...≥λH; substitute the eigenvalues into the standardized curvature estimation formula to obtain a scalar characterizing the degree of local curvature, denoted as curvature value Ric(t), and define the change in local geometric curvature between the current time and the previous time as: ΔRic(t)=|Ric(t)-Ric(t-1)|; In one embodiment of the present invention, the embedding function f is... emb Based on the deterministic dimensionality reduction mapping using graph attention mechanism, the first column element z of Z(t) i The vector formed by (t) is represented as: z i (t)=ReLU(Σ j∈N(i) α ij ×G×x j (t)); Where G represents the trainable weight matrix; the attention coefficient α ij Determined by the ratio of the port rates at both ends of the node; x j (t) represents the 5-dimensional original observation tensor of node j at time t; Calculate the Nth column element of Z(t) one by one; transpose it as a row vector and concatenate them to form the matrix Z(t).
[0020] In one embodiment of the present invention, the eigenvalue entropy change obtained through local principal component analysis is used as the curvature value Ric(t). Specifically, the K=30 nearest neighbors of the current feature point Z(t) in the closed space M are selected to form a matrix Q∈R. 30×64 ; Centering Q by subtracting the historical mean from Q yields Q', and the covariance matrix C = (Q')T(Q') / 30 is calculated; 64 eigenvalues λ1, λ2, ..., λ64 are obtained; then the curvature value Ric(t) is characterized as: Ric(t) = -Σ h=1 64 λh'log(λh'+e); where h represents the identifier of the eigenvalue, h∈[1,64]; λh' represents the weighted value of the eigenvalue; e represents the smallest positive number to prevent division by zero; λh'=λh / (λ1+λ2+...+λ64+e).
[0021] In one embodiment of the present invention, the closed space M in the code implementation refers only to a finite and bounded historical normal state feature library, specifically defined as: M={Z normal (t) | Feature matrix of all sampling points during the historical fault-free period}; Preferably, when the system is running online, it is only necessary to calculate the variance of the Euclidean distance distribution from the current feature point Z(t) to the nearest neighbor in the feature library, as the local geometric curvature change.
[0022] Step S1-4: Obtain the historical curvature value sequence {Ric(t-1),Ric(t-2),...,Ric(tW)}; where Ric(t-1),Ric(t-2),...,Ric(tW) represent the W consecutive curvature values before time t; in this embodiment, W is taken as 60 sampling points, and the moving average μRic and standard deviation σRic of the sequence are calculated in real time; the dynamic curvature threshold is set as: Threshold(t)=μRic+β×σRic; where β is an adjustable sensitivity coefficient, and in this embodiment, the value of β is in the range of 3.0~5.0, which can be adjusted according to the actual situation and is not limited here; when ΔRic(t)>Threshold(t), it is determined that the current network operation state deviates from the normal manifold and an anomaly occurs, and the system immediately generates a binary anomaly trigger signal Flag=1.
[0023] Furthermore, in step S2, models are performed based on queue occupancy length and signal propagation delay, and the modeling results are used as a differentiable operator layer to construct a virtual forward computation graph; the differentiable operator layer is used to establish a differentiable mapping relationship between network input state variables and network output performance index variables; specifically: The queue occupancy length variation of each network node in the target physical network is modeled as a first-order ordinary differential equation discretized computational unit with time as the independent variable; the signal propagation delay of each physical link in the target physical network is modeled as a linear mapping computational unit with link length as the independent variable; the first-order ordinary differential equation discretized computational unit and the linear mapping computational unit are embedded as fixed, non-trainable operator layers into a virtual forward computation graph composed of fully connected layers and residual connections; the input layer of the virtual forward computation graph is defined as the adjustable configuration parameter vector of each network node, and the output layer is the prediction performance index vector of the entire network; The composite mapping function from the input layer to the output layer in the virtual forward computation graph can be analyzed using the chain rule for the partial derivatives of any variable in the input layer.
[0024] Furthermore, the specific steps of modeling the queue occupancy length variation of each network node in the target physical network as a first-order ordinary differential equation with time as the independent variable into a discretized computational unit include: For any network node i, the predicted queue occupancy length q at discrete sampling time p i (p+1) is: q i (p+1)=q i (p)+[λ i (p)-μ i (p)]×ΔT; where λ i (p) represents the input flow rate value of network node i at discrete sampling time p, μi (p) is the output link service rate value of the network node i at the discrete sampling time p, and ΔT is the preset discretization time step; The recursive equation relates to the output link service rate value μ. i The first-order partial derivative of (p) is -ΔT, and the partial derivative is always a negative constant when the queue length is greater than zero.
[0025] Furthermore, the step of modeling the signal propagation delay of each physical link in the target physical network as a linear mapping calculation unit with link length as the independent variable specifically includes: Construct an adjacency matrix A of the target physical network, wherein the adjacency matrix A is an N×N square matrix, and the elements A ij =1 indicates that there is a direct physical link between node i and node j, A ij =0 indicates that no direct physical link exists; construct the link length matrix L of the target physical network, where the link length matrix L is an N×N square matrix, and the elements L ij The value of is the actual length in kilometers of the physical link between node i and node j, when A ij When =0, the corresponding L ij The value is zero; the signal propagation delay matrix T prop The characterization formula is denoted as: T prop =A⊙L / v; where the symbol ⊙ indicates the element-wise multiplication of the matrix; v is the constant propagation speed of electromagnetic waves in the physical medium; The signal propagation delay matrix T prop Regarding any non-zero element L in the link length matrix L ij The partial derivative is a constant 1 / v.
[0026] Furthermore, the composite mapping function of the virtual forward computation graph is expressed as: y pred =D(sin,Φ); where sin is an input layer variable, including the output link service rate values {μ1,μ2,...,μ} of all network nodes. N}; Φ is the set of trainable weight parameters for the fully connected layer; y pred For output layer variables, including the predicted end-to-end latency T for the entire network. prop ; The composite mapping function D is continuous and differentiable everywhere in the input domain, and its gradient vector with respect to any input variable θ∈{sin,Φ} The propagation channel used as the loss function for backpropagation in step S3.
[0027] In one embodiment of the present invention, The set of service rates output by all nodes in the network at the current time p is denoted as {μ1(p), μ2(p), ..., μ...}N (p)}; Based on the constructed signal propagation delay matrix T prop and the preset routing split weight matrix R, R ji This represents the proportion of traffic sent from node j to node i to the total output traffic of node j, and is a fixed constant. Based on the input service rate, calculate the actual upstream traffic rate λ received by each network node. i (p); Among them, T prop (j,i) represents the signal propagation delay matrix T. prop The signal propagation delay from node j to node i; The rounding up sign indicates the delay of converting continuous propagation delay into discrete sampling steps; Based on the queue length q from the previous time step i (p) λ calculated in this unit i (p) Current input service rate μ i (p); Preset discretization step size ΔT; Based on the pre-built model q i (p+1)=q i (p)+[λ i (p)-μ i (p)]×ΔT calculates the predicted queue occupancy length q i (p+1); Based on the obtained new queue length q i (p+1), Service Rate μ i (p+1); The predicted time delay base value d is calculated. i,base (p+1)=q i (p+1) / μ i (p+1); The unit is seconds for calculation; The queue length deviation and rate of change of the current global state are concatenated into a feature vector e(p); Φ is the set of trainable weight parameters for the fully connected layer. In this embodiment, Φ includes the weight matrix W1, W2 and biases b1, b2 of a single hidden layer fully connected network; the number of hidden layer neurons is fixed at 4N, which is 4 times the number of physical nodes. Characterized as: Δd i (p+1)=W2×ReLU(W1×e(p)+b1)+b2; The composite mapping function is denoted as: y pred =D(sin,Φ)=[d i,base ,Δd i ] i=1N .
[0028] It is important to note that this embodiment employs a standard fully connected layer with the ReLU activation function. While ReLU is traditionally non-differentiable at x=0, in this engineering implementation, it is replaced with the Softplus function ln(1+e^(-x) / x). x This function is infinitely differentiable over the entire real number field, and its approximation error to ReLU is less than 10. -6 .
[0029] Further, in step S3, in response to the anomaly trigger signal, the deviation between the actual observed value and the expected standard value of the network output performance index variable is defined as a loss function; the loss function is backpropagated through the virtual forward computation graph to obtain the loss gradient value corresponding to each network node; based on the spatial distribution characteristics of the loss gradient values, a set of anomaly root cause nodes is selected from all the network nodes; specifically: Step S3-1: Based on the generated anomaly trigger signal, the system collects the actual observed values y of the performance index variables in the physical network at the time of the anomaly. obs And retrieve the expected standard value y under the same business load from the performance database. targe The expected value is derived from the weighted average of historical data from the same period without failures; the deviation is defined as the mean squared error loss function MSE: MSE = ||y obs -y targe ||2 2 ; Step S3-2: Using the scalar of the mean squared error loss function as the simulated endpoint of the virtual forward computation graph, trigger the backpropagation operation of the constructed composite mapping function D, and propagate the gradient back from the output to the input layer by layer according to the chain rule in calculus: ; where θ represents the internal state hidden variable h of each physical network node in the ergodic composite mapping function D. i Let i = 1, 2, ..., N; for each network node i, take the L2 norm of its corresponding latent variable gradient vector as the causal contribution gradient value g of that node. i : ; Calculate the gradient values {g1, g2, ..., g} of all N network nodes. N Based on the spatial distribution characteristics of the nodes, an adaptive threshold screening method is used to normalize the gradient values and select nodes whose normalized gradient values are greater than a preset proportional threshold to form the abnormal root cause node set Groot.
[0030] It should be noted that the internal state hidden variable h iThis includes at least one of the following: current output link service rate, current input traffic arrival rate, current average processing latency of data packets within the node, current chip / optical module operating temperature, current port bit error rate, current queue congestion window pressure coefficient, and current output queue instantaneous occupancy length.
[0031] Because the filtering process varies slightly depending on the network's throughput, input and output data in different scenarios, this application does not impose the same constraint. Filtering can be done according to the actual situation. In this embodiment, the top 20% of network nodes or those with a normalized value > 0.7 can be considered as the set of abnormal root cause nodes. The physical meaning of the abnormal root cause node set Groot is that the larger the gradient magnitude, the greater the causal responsibility of the latent variables of the network node for the final performance deviation.
[0032] Furthermore, in step S4, based on the set of abnormal root cause nodes and the current network state variables, a preset large language model is invoked to output a recommended iterative target strategy within a preset control action space, specifically as follows: Step S4-1: Vectorize and concatenate the located abnormal root cause node set Groot and the current network's global state variable scurrent to construct the standard prompt word embedding vector of the large model; call the pre-trained large language model, which has a high-dimensional continuous control action space B. Each dimension in the action space B corresponds to an adjustable parameter in the network, and the value range is normalized to [0,1]. The global state variable scurrent includes the current configuration parameters of all nodes; The adjustable parameters include interface rate limiting percentage, QoS queue weight value, and ECMP load sharing threshold.
[0033] Step S4-2: The large language model generates initial candidate regulation strategies b0∈B; then it enters the iterative optimization loop: the current strategy b x Input the composite mapping function D, perform forward inference, and obtain the result at b x Output layer variable y under the strategy pred (b x )=D(b x And recalculate the new loss function value MSE. new =||y pred (b x )-y target ||2 2 ; MSE new As a hard constraint, it is fed back to the generative decoder of the large language model, and the policy gradient is calculated. The large language model is guided to adjust its strategy in the next round of output to reduce loss; the iterative objective function is defined as: Π=argmin b∈B ||D(b)-y target ||2 2 ; Π represents the iterative target strategy; D(b) represents the output layer variable based on the composite mapping function D under strategy b; Step S4-3: Repeat the process. When the rate of change of the loss function between two adjacent iterations is less than a preset threshold, or when the preset maximum number of iterations is reached, terminate the iteration and output the iteration target strategy Π.
[0034] A network operation status anomaly management system based on a large model, comprising: The data acquisition and sensing module is configured to collect multi-source operational data from network nodes, map it to the feature space and build a historical feature library, calculate the deviation of the current state from the historical normal state in real time, and generate an abnormal trigger signal when the deviation exceeds the dynamic threshold. The differentiable digital twin module is configured to construct a virtual forward computation graph consisting of a fixed physical operator layer and a trainable residual compensation layer connected in series. It is used to simulate the differentiable mapping relationship between the network state from the input configuration to the output performance index, supporting both forward inference prediction and backpropagation differentiation. The root cause localization module is configured to respond to anomaly triggering signals, use the deviation between the actual observed value and the expected standard value as a loss function, obtain the contribution of each node to the deviation through backpropagation of the differentiable digital twin, and thereby filter out the set of abnormal root cause nodes. The strategy generation module is configured with a finite number of preset discrete atomic actions, and the strategy with the smallest deviation between the deduction result and the expected value is used as the iterative target strategy.
[0035] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A method for managing network operational status anomalies based on a large model, characterized in that: Specifically, the steps include the following: Step S1: Obtain multi-source time-series data of network operation status, map the multi-source time-series data onto a representation manifold in a high-dimensional feature space, calculate the local geometric curvature change of the representation manifold between the current time and historical time; if the local geometric curvature change exceeds a preset dynamic curvature threshold, determine that the network has an anomaly and generate an anomaly trigger signal. Step S2: Model based on queue occupancy length and signal propagation delay respectively, and use the modeling results as a differentiable operator layer to construct a virtual forward computation graph; the differentiable operator layer is used to establish a differentiable mapping relationship between network input state variables and network output performance index variables; Step S3: In response to the anomaly trigger signal, the deviation between the actual observed value and the expected standard value of the network output performance index variable is defined as a loss function; the loss function is backpropagated through the virtual forward computation graph to obtain the loss gradient value corresponding to each network node; based on the spatial distribution characteristics of the loss gradient values, a set of anomaly root cause nodes is selected from all the network nodes, specifically: Step S3-1: Based on the generated anomaly trigger signal, the system collects the actual observed values y of the performance index variables in the physical network at the time of the anomaly. obs And retrieve the expected standard value y under the same business load from the performance database. targe Define the deviation as the mean squared error loss function MSE: MSE = ||y obs -y targe ||2 2 ; Step S3-2: Using the scalar of the mean squared error loss function as the simulation endpoint of the virtual forward computation graph, trigger the backpropagation operation of the constructed composite mapping function D. According to the chain rule in calculus, the gradient is propagated back layer by layer from the output to the input: where θ represents the internal state hidden variable h of each physical network node in the composite mapping function D. i Let i = 1, 2, ..., N; for each network node i, take the L2 norm of its corresponding latent variable gradient vector as the causal contribution gradient value g of that node. i :; Calculate the gradient values {g1, g2, ..., g} of all N network nodes. N The spatial distribution characteristics of the nodes are analyzed using an adaptive threshold filtering method. The gradient values are normalized, and nodes with normalized gradient values greater than a preset proportional threshold are selected to form the abnormal root cause node set Groot. Step S4: Based on the set of abnormal root cause nodes and the current network state variables, call the preset large language model to output the recommended iterative target strategy in the preset control action space, specifically: Step S4-1: Vectorize and concatenate the identified abnormal root cause node set Groot and the current network's global state variables to construct the standard prompt word embedding vector of the large language model; call the pre-trained large language model, which has a high-dimensional continuous control action space B preset inside. Each dimension in the action space B corresponds to an adjustable parameter in the network, and the value range is normalized to [0,1]. Step S4-2: The large language model generates initial candidate regulation strategies b0∈B; then it enters the iterative optimization loop: the current strategy b x Input the composite mapping function D, perform forward inference, and obtain the result at b x Output layer variable y under the strategy pred (b x )=D(b x And recalculate the new loss function value MSE. new =||y pred (b x )-y target ||2 2 ; MSE new The hard constraint is fed back to the generator decoder of the large language model; the iterative objective function is defined as: Π=argmin b∈B ||D(b)-y target ||2 2 ; Π represents the iterative target strategy; D(b) represents the output layer variable based on the composite mapping function D under strategy b; Step S4-3: Repeat the process. When the rate of change of the loss function between two adjacent iterations is less than a preset threshold, or when the preset maximum number of iterations is reached, terminate the iteration and output the iteration target strategy Π.
2. The network operation status anomaly management method based on a large model according to claim 1, characterized in that: In step S1, multi-source time-series data of the network's operating status is acquired, and the multi-source time-series data is mapped onto a representation manifold in a high-dimensional feature space. The local geometric curvature change of the representation manifold between the current time and historical time is calculated. When the local geometric curvature change exceeds a preset dynamic curvature threshold, an anomaly is determined to have occurred in the network, and an anomaly trigger signal is generated. Specifically: Step S1-1: Collect multi-source time-series data from each switch and router in the target physical network at a preset sampling period. For the i-th node in the network, the original data vector x collected at time t is... i (t) is represented as: x i (t)=[q i (t),d i (t),u i (t),l i (t),δ i [(t)]; where q i (t) represents the instantaneous length of the output queue, d i (t) represents the average forwarding delay of data packets, u i (t) represents the port bandwidth utilization, l i (t) represents the packet loss rate, δ i (t) represents the link error rate; the data from all N network nodes at T consecutive time points are stacked along the time dimension to construct the original observation tensor X(t): X(t)∈R T×N×5 Where T is the length of the sliding time window. Step S1-2: Map the original data vector from the low-dimensional physical space to a high-dimensional feature space of dimension H to obtain the system state representation matrix: Z(t) = f emb (X(t))∈R N×H Z(t) represents the current state point; the set of all possible system states Z in the high-dimensional space is defined as a closed space M; Step S1-3: On the M, take the current state point Z(t) and its K nearest neighbors from adjacent historical times, and calculate the covariance matrix C=(Z... k -Z') T (Z k -Z') / K; where Z k Let C represent the high-dimensional feature space of the k-th nearest neighbor, k∈K; Z' represents the mean of the high-dimensional feature space in history; perform eigenvalue decomposition on C to obtain H eigenvalues λ1≥λ2≥...≥λH; substitute the eigenvalues into the standardized curvature estimation formula to obtain a scalar characterizing the degree of local curvature, denoted as curvature value Ric(t), and define the change in local geometric curvature between the current time and the previous time as: ΔRic(t)=|Ric(t)-Ric(t-1)|; Step S1-4: Obtain the historical curvature value sequence {Ric(t-1),Ric(t-2),...,Ric(tW)}; where Ric(t-1),Ric(t-2),...,Ric(tW) represent the W consecutive curvature values before time t, and calculate the moving average μRic and standard deviation σRic of the sequence in real time; set the dynamic curvature threshold Threshold(t) as: Threshold(t)=μRic+β×σRic; where β is an adjustable sensitivity coefficient; when ΔRic(t)>Threshold(t), it is determined that the current network operation state deviates from the normal manifold and an anomaly occurs, and the system immediately generates a binary anomaly trigger signal Flag=1.
3. The network operation status anomaly management method based on a large model according to claim 2, characterized in that: In step S2, models are performed based on queue occupancy length and signal propagation delay, and the modeling results are used as a differentiable operator layer to construct a virtual forward computation graph; the differentiable operator layer is used to establish a differentiable mapping relationship between network input state variables and network output performance index variables; specifically: The queue occupancy length variation of each network node in the target physical network is modeled as a first-order ordinary differential equation discretization computation unit with time as the independent variable; the signal propagation delay of each physical link in the target physical network is modeled as a linear mapping computation unit with link length as the independent variable; the first-order ordinary differential equation discretization computation unit and the linear mapping computation unit are embedded as fixed, untrainable operator layers into a virtual forward computation graph composed of fully connected layers and residual connections; The input layer of the virtual forward computation graph is defined as the adjustable configuration parameter vector of each network node, and the output layer is the prediction performance index vector of the entire network. The composite mapping function from the input layer to the output layer in the virtual forward computation graph can be analyzed using the chain rule for the partial derivatives of any variable in the input layer.
4. The network operation status anomaly management method based on a large model according to claim 3, characterized in that: The discretization computation unit that models the queue length variation of each network node in the target physical network as a first-order ordinary differential equation with time as the independent variable specifically includes: For any network node i, the predicted queue occupancy length q at discrete sampling time p i (p+1) is: q i (p+1)=q i (p)+[λ i (p)-μ i (p)]×ΔT; where, q i (p) represents the instantaneous queue length at discrete sampling time p; λ i (p) represents the input flow rate value of network node i at discrete sampling time p, μ i (p) represents the output link service rate value of network node i at discrete sampling time p, and ΔT is the preset discretization time step.
5. The network operation status anomaly management method based on a large model according to claim 4, characterized in that: Furthermore, the step of modeling the signal propagation delay of each physical link in the target physical network as a linear mapping calculation unit with link length as the independent variable specifically includes: Construct an adjacency matrix A of the target physical network, wherein the adjacency matrix A is an N×N square matrix, and the elements A ij =1 indicates that there is a direct physical link between node i and node j, A ij =0 indicates that no direct physical link exists; construct the link length matrix L of the target physical network, where the link length matrix L is an N×N square matrix, and the elements L ij The value of is the actual length in kilometers of the physical link between node i and node j, when A ij When =0, the corresponding L ij The value is zero; the signal propagation delay matrix T prop The characterization formula is denoted as: T prop =A⊙L / v; where the symbol ⊙ indicates the multiplication of corresponding elements of the matrix; v is the constant propagation speed of electromagnetic waves in the physical medium.
6. The network operation status anomaly management method based on a large model according to claim 5, characterized in that: The composite mapping function of the virtual forward computation graph is expressed as: y pred =D(sin,Φ); where sin is an input layer variable, including the output link service rate values {μ1,μ2,...,μ} of all network nodes. N }; Φ is the set of trainable weight parameters for the fully connected layer; y pred For output layer variables, including the predicted end-to-end latency T for the entire network. prop .
7. A network operation status anomaly management system based on a large model, employing the network operation status anomaly management method based on a large model as described in any one of claims 1-6, characterized in that: include: The data acquisition and sensing module is configured to collect multi-source operational data from network nodes, map it to the feature space and build a historical feature library, calculate the deviation of the current state from the historical normal state in real time, and generate an abnormal trigger signal when the deviation exceeds the dynamic threshold. The differentiable digital twin module is configured to construct a virtual forward computation graph consisting of a fixed physical operator layer and a trainable residual compensation layer connected in series. It is used to simulate the differentiable mapping relationship between the network state from the input configuration to the output performance index, supporting both forward inference prediction and backpropagation differentiation. The root cause localization module is configured to respond to anomaly triggering signals, use the deviation between the actual observed value and the expected standard value as a loss function, obtain the contribution of each node to the deviation through backpropagation of the differentiable digital twin, and thereby filter out the set of abnormal root cause nodes. The strategy generation module is configured with a finite number of preset discrete atomic actions, and the strategy with the smallest deviation between the deduction result and the expected value is used as the iterative target strategy.
Citation Information
Patent Citations
Industrial network security situation prediction method and system based on generative large model
CN121388958A
Graph structure feature-based routing optimization method and system
WO2024037136A1