Network configuration information change method and system

By constructing a dynamic simulation model of network mirroring, the problem of synchronizing logical correctness and performance prediction in SDN network configuration changes is solved. It enables quantitative assessment of performance impacts such as latency and packet loss before changes, ensuring seamless service switching and improving the level of operation and maintenance automation.

CN122293508APending Publication Date: 2026-06-26JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINAN INSPUR DATA TECH CO LTD
Filing Date
2026-05-26
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies cannot simultaneously verify logical correctness and performance when SDN network configuration changes, resulting in a large error between the evaluation results and the actual performance, and lacking the ability to perform real-time synchronous performance prediction and closed-loop optimization.

Method used

A dynamic simulation model of the network mirror, which is updated synchronously with the physical network, is constructed. By using graph neural networks and network calculus techniques, minimal change operation information is generated, and risk prediction is performed in terms of logical correctness and performance impact. The model parameters are calibrated in real time to achieve closed-loop feedback optimization.

Benefits of technology

Quantitatively assess the performance impact before network configuration changes to ensure seamless service switching, significantly reduce the error between assessment results and actual performance, and improve the security of network changes and the level of operational automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122293508A_ABST
    Figure CN122293508A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for changing network configuration information, relating to the field of cloud technology. The method includes constructing a dynamic network mirroring model that is synchronously updated with the physical network and has predictive capabilities based on network entity configuration and network operation data; generating minimum change operation information for changing from the current network to the target network based on business requirement description information; performing multi-dimensional risk prediction on the minimum change operation information based on this model; and deploying the minimum change operation information to the physical network according to dependencies when the risk prediction results meet the conditions and while maintaining consistency of business traffic during the change period. The measured performance data after deployment is compared with the predicted network performance data, and the model is calibrated using the deviation data. This invention can solve the problem of large errors between the network configuration change evaluation results and actual performance in related technologies, and can significantly improve the security of the network configuration change execution process and the automation level of network operation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud technology, and in particular to a method and system for changing network configuration information. Background Technology

[0002] With the large-scale deployment of SDN (Software Defined Networking) in data centers, network resource allocation strategies change frequently and carry high risks. When network configurations change, related technologies cannot simultaneously verify the logical correctness of the changes and the resulting performance improvements. Furthermore, they cannot utilize real-world operational data after the changes to revise the assessment criteria used before the next change, leading to significant discrepancies between network change assessment results and actual performance.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] This invention provides a method and system for changing network configuration information. Without interrupting services, it quantitatively assesses the impact of network configuration changes on network latency and packet loss before the network configuration goes live. The assessment criteria are adjusted based on the actual performance data collected after the configuration change goes live, so that the assessment results before the next network configuration change are closer to the actual situation.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a method for changing network configuration information, comprising: Based on network entity configuration and network operation data, a network mirror dynamic simulation model is constructed that is updated synchronously with the physical network and is used to output network performance change information based on network configuration change information. Based on the business requirement description, generate the minimum change operation information from the current network state to the target network state; the minimum change operation information should at least include the configuration change data and the dependencies between the configuration change items; Based on the network mirror dynamic simulation model, the minimum change operation information is risk-predicted from at least the dimensions of network logical correctness and network performance impact. When the risk prediction results meet the preset release conditions, the minimum change operation information is deployed to the physical network according to the dependency relationship while maintaining the consistency of business traffic during network changes. The measured performance data of the physical network that performs the minimum change operation is compared with the network performance change information output by the network mirror dynamic simulation model, and the deviation information obtained from the comparison is used to calibrate the network mirror dynamic simulation model.

[0006] Another aspect of the present invention provides a network configuration information change system, comprising: Memory, used to store computer programs; A processor is configured to implement the steps of the network configuration information change method described above when executing the computer program.

[0007] The advantages of the technical solution provided by this invention are as follows: By constructing a network mirror dynamic simulation model that is synchronously updated with the physical network and has performance prediction capabilities, a high-fidelity digital benchmark is provided for subsequent change evaluation. On this basis, minimal change operation information containing only necessary changes is generated according to business needs, and the dependencies between operations are clearly defined, effectively reducing the scope of network disturbance caused by changes. Subsequently, the model is used to predict changes from multiple dimensions such as logical correctness and performance impact. Changes are only deployed according to dependencies when the prediction results meet the conditions, ensuring uninterrupted business traffic. This quantitative screening of potential logical errors and performance degradation is completed before the change goes live, avoiding potential failures caused by direct deployment. Finally, the measured performance data collected after deployment is compared with the model output, and the model parameters are calibrated using deviation feedback, so that the model's prediction accuracy for subsequent changes is gradually improved. Therefore, this invention enables quantitative assessment of the performance impacts such as latency and packet loss before network configuration changes, ensures seamless service switching during change execution, and continuously optimizes the assessment criteria through closed-loop feedback. This significantly reduces the error between the change assessment results and actual performance, making the assessment results before the next network configuration change closer to reality, thus improving the security of network changes and the level of operational automation. Furthermore, this invention also provides a corresponding implementation system for the network configuration information change method, which has corresponding advantages. Attached Figure Description

[0008] To more clearly illustrate the technical solutions of the present invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 A flowchart illustrating a method for changing network configuration information provided by the present invention; Figure 2 This is a schematic diagram of the network mirror dynamic simulation model construction process provided by the present invention; Figure 3 This is a schematic diagram of the heterogeneous network state diagram construction process provided by the present invention; Figure 4 A schematic diagram of the node state update method for heterogeneous network state graph provided by the present invention; Figure 5 This is a schematic diagram of the multidimensional physical constraint embedding provided by the present invention; Figure 6 This is a structural framework diagram of an exemplary embodiment of the network configuration information changing device provided by the present invention. Figure 7 A schematic diagram of an exemplary embodiment of the network configuration information change system provided by the present invention; Figure 8 A schematic diagram of the hardware composition framework applicable to the network configuration information change method provided by the present invention; Figure 9 This is a schematic diagram of the framework of the network configuration information modification method provided by the present invention in an exemplary application scenario. Detailed Implementation

[0010] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. In this specification and the aforementioned drawings, the terms "first," "second," "third," "fourth," etc., are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0011] With the rapid development of technologies such as cloud computing, big data, and artificial intelligence, the scale and complexity of data center networks are increasing daily, and the demands on network performance and stability from the services they support are becoming increasingly stringent. SDN, through its control-forward separation architecture, provides a theoretical basis for centralized control and automated management of network resources. In private cloud or virtualization platforms in production environments, any change to network resource allocation strategies (such as adjusting equal-cost multipath route weights, modifying queue scheduling parameters, or updating access control lists) can have unpredictable impacts on service traffic.

[0012] When modifying network configuration information, one approach is to verify the network using an offline simulation platform. This method involves building a simulation environment identical to the physical network topology (e.g., using tools like Mininet or GNS3), running modification scripts within it, and checking the network reachability and basic functionality after the changes. While this method can detect some deterministic logical errors, such as routing loops or configuration syntax errors, it suffers from insufficient fidelity. First, offline simulation environments struggle to fully replicate the complex, dynamic, and self-similar real-world traffic characteristics of production networks, failing to simulate the massive, heterogeneous, and highly dynamic service traffic of the real world. Therefore, predictions of performance issues (latency, jitter, packet loss) are virtually ineffective. Second, the simulation model is disconnected from the physical network's state; any instantaneous changes in the physical network (such as link congestion or device CPU fluctuations) cannot be reflected in the simulation. This leads to a significant discrepancy between the final simulation results and actual online performance, and many potential performance issues (such as micro-burst congestion and inter-application latency jitter) cannot be detected before deployment. This offline simulation-based approach fails to meet the practical needs of operations and maintenance personnel. Another approach uses network configuration analysis tools such as Batfish and Veriflow to abstract network device configurations and topology into a mathematical model. This model uses formal methods to exhaustively enumerate all possible packet forwarding paths to rigorously prove whether the network satisfies certain invariants, such as the inability of hosts A and B to communicate. While this method can detect 100% of configuration errors that violate predefined logical rules, providing a very high guarantee of logical correctness, formal verification is essentially static analysis and cannot provide answers about network performance. For example, formal verification can output whether a packet reaches the receiver, but it cannot output the time required from sending the packet to receiving it, nor can it predict the degree of network queue backlog during traffic surges. Yet another approach uses a monitoring platform that can visualize network topology, traffic, and alarms in real time. This platform can collect telemetry data from the physical network and present it in a digital dashboard, providing richer and more real-time network visibility compared to traditional network management. However, these methods cannot predict events that haven't occurred, especially lacking the ability to predict "what-if" scenarios—that is, what the network performance would be like if a certain change were made. In other words, they lack the ability to precisely quantify and extrapolate, making it difficult to quantitatively assess the scope and potential risks of changes. A seemingly harmless ACL (Access Control List) rule modification might trigger unexpected business disruptions or security vulnerabilities. The change process also lacks fine-grained canary releases and reliable automatic circuit breaker mechanisms; once a problem occurs, the impact is widespread and the recovery time is long.This approach merely achieves a digital mirror, lacking a computationally achievable performance model capable of learning and reasoning, and it lacks an automated framework that integrates predictions with risk assessment, security execution, and closed-loop optimization. Therefore, it is not a true digital twin. After network policy deployment, its operational effectiveness relies on manual monitoring and passive responses. It lacks the closed-loop feedback capability to automatically compare real-time online operational data with expected targets and continuously optimize itself accordingly. Network performance optimization is often delayed and reactive, rather than proactive and predictive.

[0013] As can be seen from the above, related technologies cannot guarantee logical correctness, accurately predict performance, execute securely, and continuously evolve. Therefore, this invention constructs a highly reliable network mirror dynamic simulation model that is synchronized with the physical network in real time. Based on this model, an automated closed-loop control method integrating policy generation, online verification, risk quantification, secure execution, and feedback optimization is established. This significantly improves the intelligence level, security, and operational efficiency of network resource allocation. Before network configuration changes go live, the correctness of the configuration logic is checked simultaneously, and quantitative predictions of network latency, packet loss, and other performance characteristics after the changes are made. This solves the problems of insufficient pre-deployment verification of resource allocation strategies in related SDN technologies, high deployment risks, and a lack of continuous self-optimization capabilities based on actual test results. The various non-limiting embodiments of this invention are described in detail below with reference to the accompanying drawings and specific implementation methods. Please refer to [link to relevant documentation] first. Figure 1 According to the network configuration information modification method provided by the present invention, it can be implemented as a computer program product, installed and run on the processor of a network control server or cloud management platform, for the purpose of implementing automated evaluation, secure execution and closed-loop optimization processing of network configuration changes. In some embodiments of the method, the method includes the following steps: S101: Based on network entity configuration and network operation data, construct a network mirror dynamic simulation model that is updated synchronously with the physical network and is used to output network performance change information based on network configuration change information.

[0014] Network entity configuration refers to the fixed configuration data of all devices in the physical network (such as switches and routers), including topology, port parameters, routing rules, and access control lists. Network operation data includes various data generated during the operation of the physical network, such as traffic volume, forwarding latency, link bandwidth utilization, and device operating status. The network mirror dynamic simulation model is updated synchronously with the physical network in real time. After inputting changes to the network configuration, it can output a network model showing changes in performance such as network latency, packet loss rate, and throughput. By fusing this data and utilizing graph neural networks and network calculus techniques, a digital model that can map the physical network state in real time and predict the performance after changes is established.

[0015] S102: Based on the business requirement description information, generate the minimum change operation information from the current network state to the target network state.

[0016] The business requirement description information can be network optimization goals input by the administrator, such as improving service throughput, reducing end-to-end latency, and ensuring network isolation. For example, it could guarantee end-to-end latency of less than 20 milliseconds for video conferencing services. It can also be network optimization requirements automatically triggered by the system. The minimum change operation information refers to the minimum configuration modifications required to adjust the network from the current state to the target state. This includes at least the configuration change data and the dependencies between configuration change items. Configuration change data could include modifying a switch's routing table entry or adjusting queue priorities. Dependencies between configuration change items could include adding bandwidth limits before switching traffic. After parsing this requirement in this step, and considering the current network state, the required configuration items are identified through comparison. The execution order constraints between these configuration items are analyzed to form a minimum change set, avoiding the disruption caused by a full configuration rollout.

[0017] S103: Based on the network mirror dynamic simulation model, perform risk prediction on the minimum change operation information from at least the dimensions of network logical correctness and network performance impact. When the risk prediction results meet the preset release conditions, and while maintaining the consistency of business traffic during network changes, deploy the minimum change operation information to the physical network according to the dependency relationship.

[0018] The network logic correctness dimension verifies whether configuration changes will cause logical problems such as routing loops, forwarding black holes, service connectivity interruptions, and breaches of security isolation rules. The network performance impact dimension assesses whether performance indicators such as network latency, jitter, packet loss rate, and queue backlog meet business requirements after configuration changes. Preset deployment conditions can be: no logical errors in verification, performance risk level below a set threshold, and model prediction reliability meeting requirements. Service traffic consistency ensures that data packets follow either the old or new rules throughout the network change process, without packet loss or out-of-order delivery due to mid-way rule switching. In other words, maintaining service traffic consistency means ensuring a smooth transition between old and new paths during network changes, without packet loss or out-of-order delivery. When the prediction results meet the preset deployment conditions (e.g., no logical errors and performance indicators within the business's acceptable range), the minimum change operation information is deployed to the physical network according to dependencies, while maintaining service traffic consistency during network changes. The risk prediction result can be determined by the verification results of the network logic correctness dimension and the verification results of the network performance impact dimension with equal weight. Different weight ratios can also be set, and the weighted comprehensive result can be used as the risk prediction result. If the risk prediction result is lower than the release threshold, it is determined that the preset release conditions are met. It can also generate a change license file that includes a credibility score, batch deployment suggestions, revocation trigger conditions, and revocation operation points.

[0019] S104: Compare the measured performance data of the physical network that performs the minimum change operation with the network performance change information output by the network mirror dynamic simulation model, and use the deviation information obtained from the comparison to calibrate the network mirror dynamic simulation model.

[0020] The physical network measured performance data includes actual latency, packet loss, and throughput data collected by network monitoring tools after the configuration change corresponding to the minimum change operation information is issued and the matching configuration modification operation is executed. Deviation information is the difference between the performance data predicted by the network mirror dynamic extrapolation model and the physical network measured performance data. For example, if the actual collected latency is 2 milliseconds higher than the predicted value, this deviation is used as a training sample to incrementally learn the model, update the model parameters, and make the next prediction more accurate. Through this closed-loop feedback, the model can continuously adapt to dynamic changes in the network, improving the accuracy of subsequent change assessments.

[0021] In the technical solution provided in this embodiment, a high-fidelity digital benchmark is provided for subsequent change evaluation by constructing a network mirror dynamic simulation model that is synchronously updated with the physical network and has performance prediction capabilities. Based on this, minimal change operation information containing only necessary changes is generated according to business needs, and the dependencies between operations are clearly defined, effectively reducing the scope of network disturbance caused by changes. Subsequently, the model is used to predict changes from multiple dimensions such as logical correctness and performance impact. Changes are deployed according to dependencies only when the prediction results meet the conditions, ensuring uninterrupted business traffic. This quantitative screening of potential logical errors and performance degradation is completed before the change goes live, avoiding potential failures caused by direct deployment. Finally, the measured performance data collected after deployment is compared with the model output, and the model parameters are calibrated using deviation feedback, so that the model's prediction accuracy for subsequent changes is gradually improved. Therefore, this invention enables quantitative assessment of the performance impact of network configuration changes, such as latency and packet loss, before changes are made. It ensures seamless service switching during change execution and continuously optimizes the assessment criteria through closed-loop feedback. This significantly reduces the error between the change assessment results and the actual performance, making the assessment results before the next network configuration change closer to the real situation, and improving the security of network changes and the level of operation and maintenance automation.

[0022] In the above embodiments, no limitation is made on how the model is constructed. This embodiment further defines the construction process of the network mirror dynamic inference model, such as... Figure 2 As shown, it includes the following steps: The network mirror dynamic inference model comprises a data layer and a performance prediction layer. The data layer is the network mirror graph, generated based on network entity configuration and network operation data. It uses network entity objects as graph nodes and the existence of links between graph nodes as graph edges, following a graph structure construction method. The performance prediction layer includes a target graph neural network, a physical constraint layer, a network computation layer, and an output layer. The target graph neural network is a graph neural network structure with spatial convolutional layers followed by cascaded temporal processing layers; its input is a heterogeneous network state graph. The heterogeneous network state graph is constructed based on the network mirror graph, using multiple types of network objects as graph nodes and determining node edges based on the connections between network objects. In the target graph neural network, the states of various graph nodes are updated in stages according to the relationships between different types of graph nodes in the heterogeneous network state graph and the upstream and downstream dependencies of message passing. During the state update process of each graph node, the physical constraint layer embeds matching physical constraints according to the network object type corresponding to each type of graph node. The physical constraints include: updating graph node types involving queuing behavior using differentiable queuing theory approximation formulas, setting physical upper limits for graph node types involving resource capacity, and setting lower limits for graph node types involving end-to-end cumulative performance, which are composed of the sum of propagation delay and processing delay. The network computation layer runs in parallel with the target graph neural network, providing physical boundary constraints for the target graph neural network and outputting analytical baseline values. The output layer determines the network performance change prediction results based on the analytical baseline values ​​and the predicted values ​​output by the target graph neural network.

[0023] In this embodiment, to ensure the realism of the network mirroring dynamic simulation model, data needs to be collected from the production network in real time and comprehensively. For static data, standardized network management interfaces, such as model-based network configuration protocols (e.g., NETCONF) or network management interfaces (e.g., gNMI), can be used to actively obtain the network device topology, port configuration, routing protocol configuration, access control lists, etc., which constitute the skeleton of the network mirroring dynamic simulation model. By monitoring updates to routing protocols (e.g., BGP-LS, Border Gateway Protocol-Link State), the routing table information and reachability changes of the entire network can be monitored in real time. By collecting network traffic matrices, such as using sFlow and NetFlow, the end-to-end traffic distribution and magnitude in the network can be understood. By deploying INT (In-band Network Telemetry) technology, KPIs (Key Performance Indicators) such as packet forwarding paths, end-to-end latency, hop-by-hop queue depth, jitter, and congestion status can be monitored. Precise measurement of performance indicators (key performance metrics). In-band network telemetry embeds telemetry commands and metadata directly into ordinary data packets, which are then transmitted along with the packets in the network. When a data packet passes through a network device that supports this technology, the device adds its own operational status (such as outgoing port ID, queue occupancy, forwarding latency, etc.) to the telemetry header of the data packet, enabling the receiving end to obtain a complete forwarding history and performance profile of the data packet. All collected data, whether configuration changes or performance metrics, are assigned precise timestamps and aligned on a unified timeline to form a time-series database that can be used for historical playback and correlation analysis.

[0024] Once this data is collected, a network mirror graph can be constructed as the data layer, based on network entity configuration and network operation data. Network entity objects (such as switches, routers, and server network cards) are used as graph nodes, and the existence of physical links (such as fiber optic connections) or logical links (such as VXLAN tunnels) between nodes are used as graph edges. This graph synchronizes the physical network state in real time. To address the insufficient accuracy of traditional analytical mathematical models (such as M / M / 1 queuing theory) when handling complex self-similar traffic, and the poor generalization ability of general deep learning models when lacking large-scale labeled data, this embodiment designs a performance prediction layer based on a hybrid modeling architecture that is specifically adapted to the characteristics of data center networks and incorporates prior knowledge of queuing theory. A heterogeneous network state graph is further constructed based on the network mirror graph. Nodes in this graph are not homogeneous, and node edges are determined based on physical connections and traffic mapping relationships. In the target graph neural network, the state of various nodes is updated in stages according to the relationships between different types of nodes in the heterogeneous network state graph and the upstream and downstream dependencies of message passing. During node state updates, the physical constraint layer embeds corresponding physical constraints based on the node type: for queue nodes, it uses a differentiable queuing theory approximation formula to update the delay; for device nodes, it sets a physical bandwidth upper limit; and for flow nodes, it sets a lower bound on the sum of propagation delay and processing delay. Due to the strong time-varying nature of network traffic, a temporal processing unit (such as a gated recurrent unit or a Transformer Encoder) is cascaded after the spatial convolutional layer of the target graph neural network. This temporal processing layer is responsible for processing the time series of each node's features over time, capturing the temporal dependence of traffic bursts. For example, learning the rapid growth rate of the queue level at the current moment predicts a high probability of packet loss in the next moment. The network computation layer runs in parallel with the target graph neural network, providing physical boundary constraints (such as the upper bound of the worst-case delay) and outputting analytical baseline values ​​(such as the theoretical delay calculated by the classic queuing model). The output layer adds the analytical baseline values ​​to the prediction residuals output by the target graph neural network to obtain the final prediction result of network performance changes.

[0025] As can be seen from the above, this embodiment solves the problem of unreliable predictions caused by the lack of physical constraints in pure data-driven models, and improves the prediction accuracy and robustness of the model by integrating prior physical knowledge.

[0026] Based on the above embodiments, this embodiment further defines the construction process of the network mirror map, which may include the following: Treating network devices or their ports as network entity objects, create corresponding graph nodes for each network entity object and configure node attributes for each graph node. Node attributes include the port's physical capacity, buffer size, queue scheduling algorithm type, and congestion management mechanism. If the first and second network entity objects have physical or logical links, set graph edges for the first and second network entity objects and configure edge attributes for each graph edge. Edge attributes include the link's bandwidth, propagation delay baseline, historical average packet loss rate, and the identifier of the shared risk group to which it belongs. Configure graph attributes for the network mirror graph based on the currently effective traffic matrix, the resource allocation weight of the service slice, service level agreement requirements, and global traffic shaping or rate limiting policies.

[0027] In this embodiment, network devices (such as switches and routers) or their ports are treated as network entity objects. A graph node is created for each network entity object, and node attributes are configured for each node. Node attributes may include, for example, the physical capacity of the port (e.g., 1Gbps), buffer size (e.g., 10MB), queue scheduling algorithm type (e.g., strict priority, weighted fair queue), and congestion management method (e.g., ECN (Electronic Network Congestion Control Protocol based on explicit feedback)). If there is a physical link (e.g., optical fiber) or a logical link (e.g., VXLAN (Virtual Extensible Local Area Network) tunnel) between two network entity objects, a graph edge is established between the two nodes, and edge attributes are configured for each edge. Edge attributes may include, for example, the link bandwidth (e.g., 100Gbps), propagation delay baseline (e.g., light speed propagation delay), historical average packet loss rate (e.g., 0.01%), and the identifier of the shared risk group to which it belongs (e.g., multiple optical fibers in the same optical cable). In addition, graph attributes can be configured for the entire network mirror graph based on the currently effective traffic matrix (end-to-end traffic distribution), the resource allocation weight of service slices (such as the bandwidth ratio of different tenants), service level agreement requirements (such as maximum latency and minimum bandwidth), and global traffic shaping or rate limiting policies.

[0028] As can be seen from the above, this embodiment solves the problem of how to organize network static and dynamic information into a graph format. The network mirror graph is not only a static snapshot, but it will be dynamically updated in real time as telemetry data flows in, becoming a precise mirror of the physical network world in the digital space, providing a structured data foundation for subsequent prediction.

[0029] Based on the above embodiments, this embodiment further defines the construction process of the heterogeneous network state graph, such as... Figure 3 As shown, it may include the following: Network devices, network service quality, and service traffic characteristics are used as network objects, and the vertices of the heterogeneous network state graph are determined as network device nodes, queue nodes, and flow nodes. Based on the physical connections between network devices, the mapping relationship between traffic and queues, and the internal affiliation relationship between devices and queues, it is determined whether there are node edges between the vertices of the heterogeneous network state graph. Network device characteristics representing the computing power and forwarding capacity of network devices are configured as the initial characteristic data of network device nodes, queue characteristics representing the buffer depth of physical ports and queue scheduling behavior are configured as the initial characteristic data of queue nodes, and service flow session characteristics representing end-to-end service flow are configured as the initial characteristic data of flow nodes.

[0030] In this embodiment, the heterogeneous network state graph does not simply abstract the network into homogeneous nodes, but instead constructs a heterogeneous graph G=(V, E) containing three types of entity nodes. The node set V of the heterogeneous network state graph G can be divided into: network device nodes (representing the computation and table lookup capabilities of switches / routers), queue nodes (representing the cache depth and scheduling behavior of physical ports), and flow nodes (representing the end-to-end service flow session characteristics). The edge set E can include: physical connection edges (device-device), logical mapping edges (flow-queue, indicating which queue a flow is scheduled to), and internal topology edges (device-queue, indicating cache affiliation). That is, network devices (such as switches and routers), network service quality (such as queue priority and cache depth), and service traffic characteristics are taken as network objects, and the graph vertices are determined as network device nodes, queue nodes, and flow nodes. Based on the physical connections between network devices (such as network cable connections), the mapping relationship between traffic and queues (i.e., which queue a certain service flow is scheduled to), and the internal affiliation relationship between devices and queues (such as the queue on a port belonging to that device), it is determined whether there are node edges between each vertex. High-dimensional feature encoding is performed on the aforementioned nodes. The initial feature data configured for network device nodes can be a vector consisting of current CPU utilization, memory utilization, basic forwarding latency, total forwarding table capacity, current flow table entry occupancy, and device role identifier. In other words, the feature data includes computing capability status: current CPU utilization, memory utilization, and basic forwarding latency; table lookup capability status: total forwarding table / flow table capacity, and current flow table entry occupancy (or tri-state content-addressable memory or static random access memory usage); and basic device attributes: device role (such as Spine / Leaf identifier) ​​and total switching capacity / backplane bandwidth. The queue node can be associated with its corresponding network device node. The initial feature data configured for the queue node can be a vector consisting of the identifier of its network device node, port physical capacity, buffer depth, current water level, drop threshold, and queue scheduling algorithm type. An initial feature vector is configured for each flow node, and the flow node is associated with the sequence of queue nodes it traverses on its forwarding path. The initial feature data for the flow node can be a vector consisting of the source address hash value, destination address hash value, average packet length, packet rate distribution parameters, and the sequence of device node identifiers it traverses. Correspondingly, the graph node update process of the heterogeneous network state graph is as follows: based on the logical mapping relationship between queue nodes and flow nodes, the internal topological relationship between network device nodes and queue nodes, and the path association relationship between flow nodes and nodes on their forwarding path, the various graph vertices of the heterogeneous network state graph are updated in stages. That is, based on the logical mapping relationship between queue nodes and flow nodes (i.e., which queue the flow is mapped to), the internal topological relationship between network device nodes and queue nodes (i.e., which device the queue belongs to), and the path association relationship between flow nodes and nodes on their forwarding path, the various graph vertices of the heterogeneous network state graph are updated in stages.

[0031] As can be seen from the above, this embodiment solves the problem of how to uniformly model heterogeneous network elements as a graph structure, providing accurate input for graph neural networks.

[0032] Based on the above embodiments, unlike traditional full-neighbor aggregation, this embodiment also provides a method for updating the graph node state of a heterogeneous network state graph, such as... Figure 4 As shown, it may include the following: The graph nodes are network device nodes, queue nodes, and flow nodes. During the queue node update phase, for each queue node, the first flow node pointing to the current queue node through a logical mapping relationship is determined. Based on the traffic data of each first flow node, the processing capacity of the network device node to which the current queue node belongs, and the link capacity characteristics of the current queue node, the current characteristic data of the current queue node is updated. During the network device node update phase, for each network device node, the first queue node connected to the current network device node is determined. Based on the status information of each first queue node and global routing control messages, the current characteristic data of the current network device node is updated. During the flow node update phase, for each flow node, the second queue nodes and second network device nodes traversed on the current flow node's forwarding path are determined. Based on the status feedback information of each second queue node and each second network device node, the queuing delay, processing delay, and packet loss probability along the path are accumulated to update the current characteristic data of the current flow node.

[0033] In this heterogeneous network state graph, the state update process for various graph nodes is a multi-stage information interaction process. Queue node updates involve congestion-aware aggregation. During state updates, queue nodes first aggregate the messages from flow nodes that point to them via logical mapping edges. This aggregation logic simulates the physical process of multiplexing in the network, calculating the sum of all traffic bursts flowing into the physical queue and performing a non-linear transformation based on the processing capacity characteristics of its parent device node and its own link capacity characteristics to update the queue's current congestion level and queuing delay state. Device node updates involve load feedback aggregation. Device nodes aggregate the state messages of all connected child queue nodes, as well as the current global routing control messages, to assess changes in the overall backplane bandwidth pressure, CPU load, and lookup latency of the device, thereby updating the overall computing and forwarding capability state of the device node. Flow node updates involve end-to-end performance mapping. After the network device node updates, the flow node aggregates the state feedback from all device nodes and queue nodes along its forwarding path. This step simulates the hop-by-hop transmission of data packets in the network, accumulating and mapping the queuing delay, processing delay, and packet loss probability along the way to the characteristics of the flow node, and finally using it to output end-to-end SLA (Service Level Agreement) predictions (such as end-to-end latency and jitter).

[0034] In this embodiment, during the queue node update phase, for each queue node, all flow nodes pointing to that queue node through logical mapping relationships are determined (i.e., which service flows enter the queue). Then, based on the traffic data of these flow nodes (e.g., total rate, burst level), the processing capacity of the network device node to which the current queue node belongs (e.g., the scheduling rate of the forwarding processor), and the link capacity characteristics of the current queue node (e.g., egress bandwidth), the current characteristic data of the queue node is updated, such as updating the queue length and queuing latency. During the network device node update phase, for each network device node, all queue nodes connected to that device node are determined (i.e., all port queues on that device). Based on the status information of these queue nodes (e.g., queue congestion level) and global routing control messages (e.g., routing table changes), the current characteristic data of the network device node is updated, such as updating the device CPU load and forwarding table occupancy. During the flow node update phase, for each flow node, all queue nodes and network device nodes traversed on the flow node's forwarding path are determined. Based on the status information fed back by these nodes, the queuing delay, processing delay and packet loss probability along the way are accumulated to update the current characteristic data of the flow node, such as updating end-to-end latency and packet loss rate.

[0035] As can be seen from the above, this embodiment solves the problem of how to interact and update the state between multiple types of nodes, and realizes the dynamic evolution simulation of network state.

[0036] Based on the above embodiments, the present invention also provides an implementation process for embedding corresponding physical constraints for different node types, such as... Figure 5 As shown, it may include the following: If the network object type corresponds to a queue node, the link utilization is calculated based on the aggregated traffic of the previous layer. The correction coefficient and basic deviation are determined based on the current traffic burst characteristics and long-tail distribution characteristics of the target graph neural network. The queuing delay characteristics are updated using the correction coefficient to correct the link utilization and the basic deviation. If the network object type corresponds to a network device node, the instantaneous throughput characteristic of its aggregation is forcibly limited to not exceeding the physical backplane bandwidth of the corresponding network device node, and a lower bound constraint is set for the processing delay characteristic of the corresponding network device node. The lower bound constraint is that the processing delay of the network device node is greater than or equal to the minimum physical delay. If the network object type corresponds to a flow node, a physical lower bound is forcibly set for the end-to-end delay prediction value of the flow node. The physical lower bound is that the predicted end-to-end total delay is greater than the sum of the light speed propagation delay of all physical links traversed by the corresponding flow node and the minimum processing delay of all devices.

[0037] In this embodiment, if the network object type corresponds to a queue node, it has a queuing delay constraint. First, the link utilization rate (i.e., the ratio of actual traffic to link bandwidth) is calculated based on the aggregated traffic of the previous layer. Then, based on the current traffic burst characteristics (such as whether the traffic is bursty) and long-tail distribution characteristics (such as whether the traffic distribution deviates from the exponential distribution) learned by the target graph neural network, the correction coefficient and the base deviation are dynamically determined. The link utilization rate is corrected using the correction coefficient, and the base deviation is added to obtain the updated value of the queuing delay. For example, the queuing delay Delay can be calculated using the relationship Delay≈α·ρ / (1-ρ)+β, where ρ is the link utilization rate, α is the correction coefficient, and β is the base deviation. α and β are not fixed constants but are dynamically learned by the neural network based on the current traffic burst characteristics and long-tail distribution characteristics. This design enables the model to provide a reasonable lower bound on latency based on physical formulas even in zero-sample or few-sample situations (such as unprecedented traffic bursts), avoiding the physical paradoxes that pure artificial intelligence models may encounter on out-of-distribution (OOD) data, such as predicting negative latency.

[0038] If the network object type corresponds to a network device node, it has capacity and processing physical limit constraints. When updating the device node state, physical limit constraints on computing power and bandwidth are forcibly applied, such as limiting its aggregated instantaneous throughput characteristic to not exceed the physical backplane bandwidth of the device (e.g., 100Gbps), and setting a hard lower bound constraint for the processing latency characteristic, that is, the processing latency must be greater than or equal to the minimum physical latency of the underlying table lookup (e.g., 50 nanoseconds), to prevent the model output from having an unreasonable processing speed that violates the hardware limits.

[0039] If the network object type corresponds to a flow node, it has end-to-end accumulation and conservation constraints. Forced flow feature updates follow the flow conservation law and the delay accumulation principle. When fitting end-to-end performance, the predicted total delay of the flow node is subject to a physical lower bound constraint; that is, the predicted end-to-end total delay must be greater than the sum of the light propagation delay of all physical links traversed by the flow and the minimum processing delay of all devices. For example, if the flow passes through two fiber optic links, each with a light propagation delay of 1 millisecond, and the minimum processing delay of each device is 10 microseconds, then the lower bound for the total delay is 2.02 milliseconds.

[0040] As can be seen from the above, this embodiment, by setting multi-dimensional physical constraints, enables the model to provide a reasonable lower bound for latency based on physical formulas and boundaries even in cases with zero or few samples, such as unprecedented traffic bursts. This completely avoids the physical paradoxes that pure artificial intelligence models may encounter with out-of-distribution data, such as predicting negative latency.

[0041] Furthermore, this embodiment also defines how the network computation layer provides physical boundary constraints, which may include the following: The network computation layer determines the upper bound of latency and the upper bound of backlog that match the current network environment based on the current traffic characteristics, and uses the upper bound of latency and the upper bound of backlog as physical boundary constraints, which are then input into the target graph neural network. The target graph neural network adds the physical boundary constraints to the loss function or uses them as the upper limit of the output truncation.

[0042] To compensate for the uncertainty in extreme value prediction of deep learning models and improve the computational efficiency of the system, this embodiment utilizes a dual-engine architecture consisting of a network calculus layer and a target graph neural network model. The network calculus layer is, for example, a pure mathematical analytical engine based on DNC (Deterministic Network Calculus) and MVA (Mean Value Analysis). Based on current traffic characteristics (such as arrival curves and service curves), the network calculus layer uses deterministic network calculus theory to calculate upper bounds on latency (i.e., worst-case latency) and backlog (i.e., worst-case queue length) that match the current network environment. These two upper bounds are then used as physical boundary constraints and input into the target graph neural network. During training or inference, the target graph neural network adds these boundaries to the loss function (e.g., imposing a penalty when the predicted value exceeds the boundary) or directly uses them as the truncation upper limit of the output layer (i.e., the output value must not exceed this boundary). For example, if network computation calculates that the upper limit of the latency of a certain link is 25 milliseconds, then the latency value predicted by GNN (Graph Neural Network) will be limited to between 0 and 25 milliseconds.

[0043] As can be seen from the above, this embodiment ensures through the network computation layer that no matter how the target graph neural network model reasones, its prediction results will never violate basic physical laws. For example, the prediction delay cannot be lower than the propagation delay at the speed of light, nor can it be higher than the theoretical maximum value under the token bucket constraint, thus solving the problem that deep learning models may fail in extreme value prediction.

[0044] Based on the above embodiments, the present invention further defines a multi-task output implementation method for the output layer, which may include the following: The output layer outputs the mean latency prediction, variance latency prediction, jitter prediction, and packet loss rate prediction for the target service flow. Among them, the variance latency prediction represents the prediction uncertainty. When the network is in an unknown abnormal state, the value of the variance latency prediction increases. Based on the mean latency prediction, the predicted packet loss rate, and the preset service level protocol threshold, the performance risk score is determined, and based on the jitter prediction and variance latency prediction, the stability score is calculated through an inverse proportional mapping function.

[0045] The output layer is designed as a multi-task head, capable of simultaneously predicting end-to-end latency distribution, jitter, and packet loss rate for a specific service flow (i.e., the target service flow). For example, a hybrid density network or Monte Carlo Dropout (random deactivation) technique can be used to output the probability distribution of the prediction results (latency prediction mean μ, latency prediction variance σ²), resulting in four output metrics: latency prediction mean μ, latency prediction variance σ², jitter prediction value, and packet loss rate prediction value. The latency prediction variance represents the uncertainty of the prediction. When the network is in an unknown abnormal state (such as a traffic burst mode that has never occurred before), the model will automatically output a very large variance value, triggering a defense mechanism. Based on the latency prediction mean, packet loss rate prediction value, and preset service level agreement thresholds (e.g., SLA requirements of latency less than 10 milliseconds and packet loss rate less than 0.1%), a performance risk score is calculated. The performance risk score corresponds to the distance between the absolute value of the model output (such as average latency or packet loss rate) and the preset SLA threshold. This distance can be obtained using a normalization function; for example, the closer the predicted value is to or exceeds the SLA threshold, the higher the performance risk score. Simultaneously, the stability score corresponds to the volatility of the model output and the uncertainty / variance of the prediction results. Based on the jitter prediction value and the latency prediction variance, the stability score can be calculated using an inverse proportional mapping function. The greater the jitter, the wider the variance, and the broader the distribution range, the more unstable the network performance, and the lower the stability score. These two scores are subsequently used to comprehensively determine whether to allow changes to go live. This embodiment solves the problem of how to quantitatively assess the reliability and risk level of prediction results, providing a basis for automated decision-making.

[0046] Furthermore, when in the target stage, the output layer outputs the parsed baseline value as the prediction result of network performance change; when not in the target stage, the output layer adds the parsed baseline value to the residual of the target graph neural network output as the prediction result of network performance change.

[0047] The target stage refers to the phase where the training data sample size of the target graph neural network is lower than a preset sample size threshold (e.g., less than 1000 samples). In this stage, the GNN has not yet been fully trained, and its output residual may be close to 0 or have extremely low confidence. That is, during the system cold start or data scarcity phase, the analytical solutions of classic queuing models such as M / G / 1 / PS are directly used as the baseline value for performance prediction. The output layer directly uses the analytical baseline value provided by the network computation layer as the network performance change prediction result, without relying on the GNN's prediction. The target graph neural network is configured to learn the residual between the network's measured performance value and the analytical baseline value. Since the variance of the residual between the learned measured value and the analytical baseline value is usually much smaller than the variance of the original index, the training difficulty of the target graph neural network is significantly reduced, accelerating convergence. Furthermore, the analytical part represents general physical laws, while the neural network part represents the hardware noise and nonlinear characteristics of specific devices, thus endowing the target graph neural network with extremely strong interpretability. If the target graph neural network is not in the target stage, that is, in the mature stage with abundant data, it has already learned complex nonlinear features. At this point, the target graph neural network can output very accurate residuals. The output layer adds the analytical baseline value to the residuals output by the target graph neural network to obtain an accurate final prediction.

[0048] As can be seen from the above, this embodiment solves the problem of inaccurate predictions in the cold start phase of the model. In the cold start phase, the theoretical baseline is used to ensure the availability of predictions, and in the mature phase, residuals are combined to improve accuracy, thus balancing the model deployment speed and prediction accuracy.

[0049] Furthermore, the present invention also provides an implementation process for generating minimal change operation information from the current network state to the target network state based on business requirement description information, which may include the following: The process involves parsing business requirement descriptions to obtain the source address group, destination address group, and service level agreement (SLA) requirements of the business flow; reading the current network resource occupancy status, topology connection data, and effective network configuration data from the network mirror dynamic deduction model to determine the current network state; determining the target network state based on SLA requirements, tenant security isolation requirements, network device processing capacity limitations, and network bandwidth resource limits; comparing the current network state with the target network state item by item to identify changed target network configuration items and organizing the target configuration operations performed on each target network configuration item into a configuration operation set as configuration change data; using each target configuration operation as a relation node, analyzing the execution order constraints and resource mutual exclusion relationships between the target configuration operations in the configuration operation set to determine whether there are connection edges between relation nodes, generating a dependency graph; and generating minimum change operation information based on the configuration operation set and dependency graph.

[0050] In this embodiment, firstly, the business requirement description information is parsed, such as providing an end-to-end latency guarantee of less than 5 milliseconds for tenant A's online transaction business. The source address group (e.g., 192.168.1.0 / 24), destination address group (e.g., 10.0.0.0 / 24), and service level agreement (SLA) requirements (latency <5ms) of the business flow are extracted from this information. Then, the current network resource utilization status (e.g., link utilization), topology connection data (inter-node connection relationships), and effective network configuration data (e.g., routing tables, ACLs) are read from the data layer of the network mirror dynamic simulation model as the current network state. Next, based on SLA requirements, tenant security isolation requirements (e.g., tenant A cannot access tenant B), network device processing capacity limitations (e.g., a switch supports a maximum of 1000 flow tables), and network bandwidth resource limits (e.g., total link bandwidth), the target network state is determined through an optimization algorithm. The current network state is compared item by item with the target network state to identify changed target network configuration items, such as changing the route of a flow from port 1 to port 2. The operations (such as adding, deleting, and modifying) corresponding to each target configuration item are organized into a configuration operation set, which serves as configuration change data. Then, using each configuration operation as a relation node, the execution order constraints (e.g., bandwidth limits must be added before traffic switching) and resource mutual exclusion relationships (such as a flow cannot be modified by two operations simultaneously) between these operations are analyzed. This determines whether there are connecting edges between relation nodes, generating a directed acyclic graph (DAG) as a dependency graph. In this DAG graph structure, nodes represent various change items in the configuration operation set (e.g., issuing specific flow table entries on a switch or modifying the queue configuration of a port). Edges represent the dependencies and execution order between these change items. The direction of the edges explicitly indicates which operation must be completed before another operation; for example, an alternative path must be established before traffic can be switched, thus ensuring that the change process does not cause transient routing loops or network congestion. Finally, based on the configuration operation set and the dependency graph, minimal change operation information is generated.

[0051] As can be seen from the above, this embodiment can solve the problem of how to automatically derive specific and minimal change instructions from high-level intentions, generate minimal and orderly change operations, reduce network disturbances, reduce the risk of change failures, and accurately match business needs.

[0052] Furthermore, the configuration operation set can be divided into multiple sub-change sets according to the execution order indicated by the connection edges of the dependency graph. There are no dependencies between the target configuration operations in each sub-change set, or the dependencies have been resolved sequentially within the sub-set. There is a sequential execution order between different sub-change sets. The target configuration operations of each sub-change set are executed in the sequential execution order. After all the target configuration operations of each sub-change set are executed, a cooling-off observation phase is entered. During the cooling-off observation phase, the target performance indicators of the network are monitored to see if they exceed the preset performance indicator thresholds. If they do not exceed the thresholds, the target configuration operations of the next sub-change set are executed. If they exceed the thresholds, the execution of the remaining sub-change sets is terminated, and a rollback operation is triggered.

[0053] In this embodiment, after the dependency graph is generated, the configuration operation set is divided into multiple sub-change sets according to the execution order indicated by the connecting edges in the graph. The division principle is: there are no dependencies between the configuration operations in each sub-change set, or the dependencies have been resolved through the order within the subset; there is a strict execution order between different sub-change sets. For example, if operation A must precede operation B, then A and B belong to different subsets. Then, the configuration operations in each sub-change set are executed sequentially according to the execution order. After all operations in each sub-change set are executed, a cooling-off observation phase is entered. During the cooling-off observation phase, key network performance indicators (such as latency and packet loss rate) are monitored to see if they exceed preset thresholds. If the indicators are normal, the next sub-change set is executed; if the indicators are abnormal (such as a sudden spike in latency), the remaining sub-change sets are immediately terminated, and a rollback operation is triggered to restore the network to its state before the change. This embodiment solves the problem that large-scale changes may cause cascading failures, and ensures the security of changes through a batch-based, observable, and rollback-enabled mechanism.

[0054] For example, a multi-objective policy generator can transform high-level user intentions, such as those of administrators, or system-triggered optimization requirements into specific and executable network device configuration changes. The multi-objective policy generator is an intelligent planning engine. Administrators can input business requirements, such as maximizing data throughput for a big data analytics cluster while ensuring end-to-end latency for online transactions is below 5 milliseconds. The multi-objective policy generator comprehensively considers the current network resource status (from the network mirror graph), business performance objectives (throughput, latency, packet loss), and necessary constraints (such as security isolation between tenants, cost budgets, and device capacity limitations). Utilizing hybrid solution techniques, such as combining classic integer programming algorithms with modern heuristic search or deep reinforcement learning algorithms, it calculates one or more candidate solutions. The output of the multi-objective policy generator is not a complete, entirely new network configuration, but rather a minimal change operation information. This minimal change operation information is a type of policy differential information, describing the minimum set of changes required to move from the current state to the target state. For example, to optimize the path of a service flow, the differential packet might contain only three instructions: the first is to delete the old routing table entry pointing to port 1 on switch A; the second is to add a new routing table entry pointing to port 2 on switch A; and the third is to raise the priority of the queue associated with this service flow from normal to high on switch B. Policy differential information ensures that the scope of changes is minimized, reducing network disturbances and the risk of introducing new errors. Furthermore, each minimal change operation directly corresponds to a specific optimization goal, facilitating subsequent inspection and understanding. The subsequent verification and execution layers only need to handle a small number of changes, rather than a full configuration, resulting in faster efficiency. While generating minimal change operation information, the generator also analyzes the inherent dependencies between change items and generates a dependency graph. For example, if the bandwidth limit of a link needs to be increased first, and then a large traffic flow is switched to it, there is a dependency between these two operations. This dependency graph is an important basis for the subsequent security execution layer to schedule changes, ensuring that change operations are executed in the correct order and avoiding transient forwarding black holes, loops, or congestion. Meanwhile, based on dependency graphs, complex changes can be broken down into a series of independent, smaller changes that can be distributed in batches (i.e., windowed), further enhancing the stability and controllability of the change process.

[0055] Furthermore, this invention also provides an implementation method for multi-dimensional risk prediction, which may include the following: The configuration change data in the minimum change operation information is applied to a formal logical model representing the current network configuration. The formal logical model is used to simulate packet forwarding paths, and logical errors are detected during the simulation. The logical error type and its associated network device identifier and forwarding table entry identifier are output as logical error detection results. Logical errors include packet forwarding loops, packets being dropped at unexpected nodes, service connectivity interruptions, or violations of security isolation rules. The configuration change data in the minimum change operation information is applied to a network mirror dynamic simulation model as a digital sandbox representing the future state of the network. Preset traffic patterns are loaded into the digital sandbox, and the network mirror dynamic simulation model is used to determine the network performance changes of the configuration change data under different traffic patterns as a performance impact estimate. Traffic patterns include regular traffic, burst traffic, and link interruption scenarios.

[0056] In terms of logical correctness, the configuration change data from the minimum change operation information is applied to a formal logical model representing the current network configuration. This formal logical model abstracts configuration rules (such as routing tables and ACLs) into Boolean logical expressions. This model is used to simulate packet forwarding paths, detecting logical errors during the simulation, including: packet forwarding loops (e.g., A to B to A), packets being dropped at unexpected nodes (e.g., no matching forwarding entry), inter-service connectivity interruptions (e.g., firewall rules blocking legitimate traffic), or violations of security isolation rules (e.g., tenant A accessing tenant B's resources). The detected logical error types and their associated network device and forwarding entry identifiers are output. In terms of performance impact, the configuration change data is applied to a dynamic network mirroring model, forming a digital sandbox representing the future state of the network. Preset traffic patterns are loaded into the digital sandbox, including regular traffic (normal load), burst traffic (e.g., a sudden 3x increase), and link interruption scenarios (e.g., a fiber optic cable being disconnected). By using a network mirroring dynamic simulation model, the performance changes of configuration change data under different traffic patterns are calculated, such as latency, packet loss rate, and jitter, as a performance impact estimate. This embodiment addresses the problem of how to comprehensively assess the risks of changes, while covering both logical correctness and performance performance.

[0057] Furthermore, before sending the configuration change data in the minimum change operation information to the physical network, the system intercepts configuration pending execution instructions that are in the sending or pending sending state; it parses the matching fields of the configuration pending execution instructions to determine the target data packets affected by the instructions; it incrementally updates the forwarding paths corresponding to the target data packets on the global data plane state graph maintained in memory, and determines the new forwarding paths after the configuration pending execution instructions take effect; it logically compares the characteristics of the new forwarding paths with preset network invariants, which are Boolean logic rules satisfied by the network; if the characteristics of the new forwarding paths do not conform to the network invariants, it prevents the issuance of configuration pending execution instructions and outputs a non-compliance message.

[0058] In this embodiment, before sending configuration change data to the physical network, configuration execution instructions that are in the sending or pending sending state are intercepted. The matching fields of the instruction (such as IP prefix and port number) are parsed to determine the target data packet set affected by the instruction (i.e., which data packets will be matched by this rule). A global data plane state graph is maintained in memory, reflecting all currently effective forwarding paths in real time. Only the forwarding paths corresponding to the target data packets are incrementally updated, i.e., old paths are erased and new paths are derived, without recalculating the entire network. The new forwarding path after the instruction takes effect is determined. Then, the characteristics of the new forwarding path (such as the device sequence and outgoing port on the path) are logically compared with preset network invariants. Network invariants are Boolean logic rules that the network must always satisfy, such as that a packet from tenant A can never reach the gateway of tenant B. If the characteristics of the new forwarding path do not conform to the network invariant (e.g., the new path will reach the gateway of tenant B), the instruction is blocked from being sent, and a non-compliance message is output. This embodiment solves the computational efficiency problem of real-time verification, achieving millisecond-level online verification through incremental calculation.

[0059] Considering the inevitable discrepancies between the network mirror dynamic simulation model and the real world, quantifying and managing these discrepancies is a prerequisite for achieving reliable closed-loop control. This embodiment also provides a method for determining the credibility of the network mirror dynamic simulation model, which may include the following: The statistical target is determined by quantifying the degree of conformity between the predicted network performance changes of the network mirror dynamic inference model and the actual effects after the corresponding network configuration changes are implemented within the historical verification period. Based on the error index between the predicted network performance changes and the measured network performance changes, the quantified value of the coverage between the training data of the network mirror dynamic inference model and the network scenario to be predicted, and the detection score and conformity quantified value of whether the network state of the network scenario to be predicted deviates from the known distribution, the credibility score of the network mirror dynamic inference model is determined. Based on the credibility score, the weight values ​​of the logic error detection results and the performance impact prediction values ​​are adjusted. The weight value of the performance impact prediction value is directly proportional to the credibility score, while the weight value of the logic error detection results is inversely proportional to the credibility score.

[0060] In this embodiment, the degree of conformity between the predicted network performance changes of the network mirror dynamic simulation model and the actual effect after the corresponding network configuration changes are implemented within the historical verification period (e.g., the past 30 days) is quantified. For example, the mean absolute percentage error between predicted latency and actual latency is calculated. Then, the credibility score of the network mirror dynamic simulation model is calculated according to the following four dimensions: First, the performance prediction value of the network mirror dynamic simulation model is continuously compared with the real performance value obtained through in-band network telemetry and other means. The error index between the model prediction value and the measured value (e.g., mean absolute percentage error, root mean square error) is calculated, and it is converted into a score of 0-1 through an inverse proportional mapping function (e.g., 1 / (1+mean absolute percentage error) or 1 / (1+root mean square error)). Second, the degree of coverage between the model's training data and the network scenario to be predicted is quantified. For example, the frequency of the current traffic features appearing in the training set, or the percentage of overlap of the current network configuration features in the historical training set is calculated. The coverage quantification value, such as the scenario coverage rate, is used to evaluate whether the data trained by the current model covers the scenario to be evaluated. For example, if you want to evaluate a strategy for high-speed data backup traffic, but such traffic patterns are lacking in historical training data, then the model's prediction reliability in this scenario will be low. Third, there's the detection score for whether the network state of the network scenario to be predicted deviates from the known distribution, i.e., the OOD detection score. This can be achieved using statistical or machine learning methods (such as feature kernel density estimation) to determine whether the current network state or the strategy to be evaluated significantly deviates from the model's known distribution range. For example, the probability value (0%-100%) that the current input traffic features belong to the known training data distribution can be calculated as the OOD detection score. An OOD event (such as an unprecedented traffic burst pattern) means that the model's prediction may no longer be reliable. Fourth, there's the consistency quantification value, which refers to the degree of consistency between the actual effect and the expectations after the most recent K changes made to the model prediction based on network mirror dynamic extrapolation. The scores of these four dimensions can be directly summed or weighted summed to obtain a reliability score between 0 and 1.

[0061] During the multi-dimensional verification process, the weights of the logic error detection results and the performance impact predictions can be adjusted according to the confidence score. The weight of the performance impact prediction is directly proportional to the confidence score: the higher the confidence score, the more it relies on performance prediction. The weight of the logic error detection results is inversely proportional to the confidence score: the lower the confidence score, the more it relies on deterministic logic verification.

[0062] As can be seen from the above, this embodiment solves the problem of how to quantify model reliability and adaptively adjust decision-making strategies. By dynamically adjusting the risk assessment weights, it emphasizes performance when the model is reliable and logic when it is unreliable, thereby improving the rationality of risk decisions.

[0063] In practical applications, an online verification and risk classification gateway can be set up. All configuration change data in the minimum change operation information must be verified by the online verification and risk classification gateway before being sent to the real device. The online verification and risk classification gateway can integrate three parallel verification channels, namely the intent verification channel, the invariant online guardian channel, and the performance playback and evaluation channel, and perform adaptive arbitration based on the credibility of the network mirror dynamic inference model.

[0064] The intent verification channel is used for offline structural risk analysis of configuration change data. It receives configuration change data and applies it to a formal logical model representing the current network state. By simulating packet forwarding paths and policy matching processes within this model, a series of deterministic logical errors can be detected in advance. The intent verification channel verifies the logical correctness of the change and outputs a structural risk score. Logical errors may include, for example, routing loops: whether the change introduces paths that cause packets to loop infinitely in the network; forwarding black holes: whether the change causes packets to be dropped at unexpected nodes; reachability breaches: whether the change will unexpectedly disrupt network connectivity between critical services; and isolation policy violations: whether the change will break preset security isolation rules, such as allowing a virtual machine in a test environment to access a production database. The formal logical model can be constructed through the data layer of a network mirror dynamic derivation model. For example, static configuration data in the network (such as topology connections, routing table entries, ACL access control lists, etc.) is extracted, and these configuration rules written by users or controllers are abstracted and transformed into mathematical expressions based on Boolean logic or packet header space to obtain the formal logical model. Formal logic models can be rigorously proven using mathematical solvers to determine whether configuration rules contain structural errors.

[0065] The invariant online daemon channel provides real-time data plane consistency protection, while bypassing the monitoring on the path where the SDN controller issues commands to network devices. Here, "online" and "monitoring" do not mean the policy has already been issued to the physical switch and started processing real service traffic, but rather that control-level interception occurs before this. Monitoring occurs on the communication path between the SDN controller and the underlying physical devices, acting as a pre-commit hook or proxy gateway. When the controller is about to issue a command to the device, this daemon channel intercepts or bypasses and mirrors it before the command actually reaches the device hardware and takes effect. At this point, the command is for the command about to be issued and has not yet actually taken effect on the physical network data plane. For each configuration command about to be issued, the invariant online daemon channel calculates the overall impact of the command on the forwarding behavior of the entire network packets in real time and incrementally, and checks whether it violates a predefined set of network invariants. These invariants are fundamental rules that the network must always satisfy, such as any data packet sent to the public network must pass through a firewall, and two virtual machines belonging to different tenants must never communicate with each other, ensuring that network security and functional attributes are maintained even during dynamic changes. The invariant online guardian channel outputs an invariant risk score. To achieve millisecond-level real-time and incremental calculations, instead of recalculating the entire network configuration, an incremental graph calculation method based on packet header space or packet equivalence classes is used: First, the affected packets are extracted to determine the incremental range: When the SDN controller is about to issue a new instruction, such as modifying the outgoing port of a certain IP network segment, the system first parses the matching field of the instruction to calculate which part of the packet equivalence classes (i.e., the set of packets with the same characteristics and undergoing the same forwarding behavior) in the entire network are affected by this new rule. Then, local path deduction, i.e., the incremental graph update process, is performed: A currently effective global data plane state graph is maintained in memory. The system does not recalculate the entire network, but only the affected packets, erasing their old paths on the state graph, and deducing their new forwarding paths according to the new instructions. For example, if the original path went through port A, the deduction now suggests that it will go through port B and the subsequent hop nodes. Other massive service flow paths unaffected by this instruction do not participate in the final logical intersection comparison, i.e., the invariant verification step: After obtaining the locally new path after the incremental update, the system immediately performs a mathematical logic comparison between the characteristics of this new path and predefined network invariants, using rules defined by Boolean logic expressions, such as "tenant A's packet path must never reach tenant B's gateway." If the path attributes conflict with the invariant logic, a violation is determined.

[0066] The performance replay and evaluation channel assesses the impact of changes on network performance, especially under stress and abnormal conditions. This channel injects the configuration change data to be verified into the network mirror graph, and then conducts a series of virtual drills in this digital sandbox, such as regular performance prediction: whether the new strategy can meet the SLA requirements of various services under estimated future traffic patterns; stress and disturbance testing: simulating various extreme or high-risk scenarios, such as: sudden interruption of critical links, core switch failure, micro-bursts of traffic increasing several times instantaneously, and specific patterns of distributed denial-of-service attacks. Using a model built with S101 to initiate digital sandbox simulations, leveraging the spatiotemporal capture capability of the target graph neural network for micro-burst traffic and the physical boundary constraints of the parsing module, this channel can not only simulate regular traffic but also quantitatively predict the performance (such as latency, jitter, and packet loss rate) and stability of the strategy under these different scenarios at the millisecond level. Finally, the performance replay and evaluation channel outputs a performance risk score, a stability score, and provides the uncertainty range of the prediction results. During probe flow verification, the measured performance metrics of the probe flow are quickly compared with the prediction baseline and uncertainty range generated in the risk prediction results. Only when the measured result falls within the confidence interval of the model prediction is the probe considered successful and real traffic changes continue. Without an uncertainty range, the execution layer cannot determine whether the measured latency of 6.2ms is a normal, reasonable fluctuation or a serious performance deviation relative to the predicted mean of 6ms. Furthermore, the stability score can be used to determine the rollback trigger conditions and the threshold used for rollback operations. If the prediction model indicates poor stability (i.e., severe jitter) after the strategy is deployed, the system will automatically tighten the circuit breaker threshold monitored online in the risk prediction results, making the defense mechanism more sensitive. In the multi-task output stage, if the value representing prediction uncertainty is extremely high, it means that the current network is in an abnormal state unknown to the model, and the system will automatically trigger the defensive control of the verification gateway.

[0067] Finally, the online verification and risk grading gateway does not simply average or sum the risk scores of the three channels. Instead, it performs an adaptive weighted synthesis based on the credibility score of the network mirror dynamic inference model. For example, a high credibility score indicates that the network mirror dynamic inference model's performance prediction is highly reliable. In this case, the weight of the performance risk score in the overall risk assessment is significantly increased. Conversely, a low credibility score, such as when the target graph neural network detects input traffic characteristics triggering an OOD alarm, or when the prediction result reaches the physical boundary set by the parsing module, indicates that the performance prediction may be inaccurate. In this case, the weight of the performance risk score is automatically reduced, relying more on deterministic logical verification results from static verification and invariant guardian channels. The release threshold for changes may also be automatically raised, such as requiring more manual review. Dynamically adjustable quantization thresholds can be pre-set, for example, setting a high confidence threshold T_high = 0.8 and a low confidence threshold T_low = 0.5. When the credibility score is ≥ T_high (high), the prediction of the network mirror dynamic inference model is considered highly reliable, and the weight of the performance risk score is significantly increased during unified risk synthesis. When the confidence score is ≤ T_low (low), it is determined that there is a significant risk of prediction bias. For example, if the prediction error is greater than 30% for 24 consecutive hours, the weight of the performance risk score is automatically reduced, and the result of the deterministic intent verification channel is relied upon entirely. The release threshold is automatically raised, such as forcing manual review. After weighted synthesis, the online verification and risk classification gateway can generate a standardized, machine-readable risk prediction result. The risk prediction result serves as the assessment report and execution basis for this change. Its content may include: the unique identifier of the change, the applicable business and network scope, the detailed risk assessment results of the three verification channels, the final comprehensive risk level, the dynamically calculated confidence score, the system's recommended canary release plan (e.g., recommending 5 batches, switching 10% of the traffic in each batch, with a 5-minute cooldown observation between batches), clear rollback trigger conditions and rollback operation points, and the overall confidence level of this assessment. The unique identifier for this change is neither a single policy code nor a single network device ID, but rather a globally unique transaction ID (such as a UUID or automated serial number) generated for the entire batch of configuration change data distribution events. This is because a single high-level intent change typically spans multiple network devices and includes multiple instruction sets. To manage this entire batch of operations as an indivisible transaction, the system must assign it a globally unified ID. The business scope is directly parsed and extracted from the business intent input through the northbound interface. For example, when an administrator submits an intent, it includes the source / destination IP address group, tenant ID, application protocol, or port number. By parsing this metadata, the system can accurately pinpoint which part of the business flow the change targets.The extracted service flow features are projected onto the network mirror map. Through path calculation algorithms, the system accurately tracks which specific switches (nodes) the service flow will pass through, and which physical ports and queues (edges) it will occupy. This specific topology impact range is the network scope. Rollback trigger conditions are calculated jointly by the predicted performance metrics output from the performance playback and evaluation channel and the SLA threshold of the service input: the system extracts the average performance and reasonable fluctuation range predicted by the target graph neural network, and automatically generates monitoring and alarm rules based on the SLA threshold. For example, if the predicted latency is 6ms and the service SLA upper limit is 10ms, the system will automatically generate conditions such as triggering rollback when the actual latency measured by INT telemetry exceeds 8ms for three consecutive probe cycles. Rollback operation points are derived from the minimum change operation information. Because configuration change data records precise atomic-level changes (such as which flow table entry on which switch was modified), the system automatically captures a snapshot of the current state of these specific devices before issuing the rollback operation. The rollback operation points are these snapshot nodes and the automatically generated reverse recovery command (undoing the configuration change data operation). The overall confidence level is a mathematical quantification of the comprehensive reliability of this gateway risk assessment. It is a percentage value obtained by mathematically mapping the comprehensive confidence score (reflecting the model's grasp of the current scenario) and the prediction uncertainty (i.e., the variance of the probability distribution) associated with the end-to-end performance of the target graph neural network.

[0068] Furthermore, to ensure the secure, stable, and controllable deployment of network configuration policies in the physical network, this embodiment does not directly update or distribute them haphazardly. Instead, before switching service traffic, it distributes probe rules to the network devices that the service traffic passes through, and selects a portion of the service traffic to copy as probe traffic or injects synthetic probe traffic. It guides the probe traffic through the new configuration path described by the configuration change data and collects latency and packet loss data of the probe traffic on the new configuration path. It compares the latency and packet loss data with the performance predictions output by the network mirror dynamic simulation model. If the deviation between the latency and packet loss data and the performance predictions is within a preset deviation range, it starts switching the service traffic to the new configuration path.

[0069] For example, when switching service traffic to a newly configured path, a new forwarding table entry is pre-installed on the network devices involved in the path change, and this new forwarding table entry is set to an inactive state (i.e., not involved in actual forwarding). On the entry device of the data flow (such as the first switch), a path identifier is added to the arriving data packets (e.g., a tag is added to the packet header). This causes data packets with added path identifiers to be forwarded according to the new path, while data packets without added path identifiers continue to be forwarded according to the original path. In this way, newly arriving traffic takes the new path, while old traffic already in the network continues to take the old path, avoiding interruption. After waiting for a period of time to ensure that all old data packets without added path identifiers have left the network, typically waiting for the maximum transmission time, i.e., waiting for all data packets without added path identifiers to leave the network, the old forwarding table entry is removed, and the new forwarding table entry is modified to an active state. At this point, all traffic automatically takes the new path.

[0070] The purpose of the probe rules is to copy a portion of the traffic selected from the business traffic (e.g., 1% based on the source IP hash) as probe traffic, or to proactively inject synthetic probe traffic with characteristics similar to real business traffic. This probe traffic is guided through the new configuration path described in the configuration change data (i.e., the path to be deployed), while in-band network telemetry is used to collect latency and packet loss data on the new path. The collected latency and packet loss data are compared with the performance estimates output by the network mirroring dynamic simulation model (e.g., predicted latency of 6 milliseconds, with a deviation range of ±1 millisecond). If the measured data falls within the predicted deviation range (e.g., measured latency of 6.2 milliseconds), the verification is considered successful, and the switch of all business traffic to the new configuration path begins; if the measured data exceeds the deviation range (e.g., measured latency of 15 milliseconds), the deployment is immediately terminated, and an alert is sent to the operations and maintenance personnel.

[0071] In this embodiment, to avoid momentary packet loss, out-of-order delivery, or loops caused by inconsistencies between old and new rules when updating network device forwarding table entries, this embodiment strictly follows the dependencies between configuration change items attached to the configuration change data to arrange the order of instructions issued on each device. For path switching of a single data stream, a two-phase commit method of "install before switching, mark before recycling" is adopted: On the device involved in the path change, the new forwarding table entries are pre-installed but not yet enabled. On the entry device of the data stream, newly arriving data packets are marked so that they begin to follow the new forwarding path. At the same time, old data packets that are still in transit in the network and have not been marked continue to be forwarded along the old path. After waiting for a sufficiently long time to ensure that all old data packets have left the network, instructions are sent to the device to safely remove the old forwarding table entries that are no longer in use. Through this fine-grained control, precise switching can be achieved, minimizing the impact on service traffic. For changes with high risk or wide impact, the system will not be fully deployed all at once, but will adopt a gradual, canary release process. The process is as follows: Before officially switching any real user traffic, a set of probe rules is first issued to network devices. The probe rules will copy a small portion (e.g., 1%) of the service traffic, or the system will actively inject some synthetic probe streams with characteristics similar to real service traffic. These copied or injected traffic (i.e., probe streams) will be guided to a new policy path for processing. The probe streams carry in-band network telemetry (INT) instructions in the packet header. After traversing the new path, the system can parse detailed, hop-by-hop real performance data from the packet header. The measured INT performance indicators of the probe streams are quickly compared with the prediction baseline and uncertainty range generated by the prediction value of the target graph neural network in the risk prediction results. Only when the measured results fall within the confidence interval of the model prediction, for example, the deviation between the measured latency and the prediction value of the target graph neural network is less than the residual threshold, will the probe be considered successful, and the next step, i.e., switching the first batch of real service traffic, will begin. If the actual test results deviate significantly from the prediction, the release process will be terminated immediately, and an alert will be sent to the operations and maintenance personnel.

[0072] As can be seen from the above, this embodiment solves the problem of how to verify the performance of a new path without affecting real business operations, and solves the problems of packet loss, out-of-order delivery, or loops that may occur during the switching between old and new paths. It ensures that network configuration strategies are deployed securely, stably, and controllably in the physical network, and achieves lossless business switching.

[0073] Furthermore, this invention not only aims to safely implement changes, but also to build a self-optimizing system capable of learning from experience and continuously evolving. Based on this, this embodiment also provides an implementation method for calibrating a network mirror dynamic inference model using measured data, which may include the following: The residual between the measured performance data of the physical network and the information on network performance changes is used as a new training sample; the new training sample is used to perform incremental training or online calibration of the network mirror dynamic inference model; the output confidence of the network mirror dynamic inference model is updated according to the calibrated model.

[0074] In this embodiment, after the network configuration change policy is implemented, in-band network telemetry and other methods can be used to continuously monitor the performance of the new network configuration policy in a real environment, obtaining physical network measured performance data. The physical network measured performance data is then continuously compared with the network performance change information output by the network mirror dynamic inference model before the change. The resulting deviation between prediction and measurement (i.e., residual) is fed back to the data layer of the network mirror dynamic inference model as new training samples. This is used for incremental training or online calibration of the weights and correction coefficients of the target graph neural network and the network computation layer. Simultaneously, based on the calibrated model update confidence score or model confidence, the prediction variance or confidence score is recalculated. For example, the residual (difference) between the physical network measured performance data (such as actual latency and packet loss rate collected after deployment) and the network performance change information output by the network mirror dynamic inference model (such as predicted latency) is used as a new training sample, and the correction parameters and basic deviation are dynamically adjusted using the measured data. Incremental training refers to fine-tuning model parameters with new samples without completely retraining, while online calibration refers to updating statistics in the model in real time, such as mean and variance.

[0075] As can be seen from the above, in this embodiment, the network mirror dynamic inference model can continuously learn and adapt to unexpected subtle changes in the network. Its prediction accuracy will gradually improve over time, and the confidence score will also change dynamically. This solves the problem of the accumulation of deviation between the model and the real environment over time, and realizes the continuous evolution of the network mirror dynamic inference model.

[0076] Furthermore, in this embodiment, the complete record of each change event can be structured and stored in the historical record repository. The complete record includes business requirement description information, minimum change operation information, and the actual effect after deployment. When new business requirement description information is received, similar successful change records are retrieved from the historical record repository and used as the initial reference data for generating minimum change operation information.

[0077] In this embodiment, the complete record of each change event is structured and stored in the historical record repository. Each complete record may include: business requirement description information (such as providing low latency guarantee for video services), minimum change operation information (such as a specific set of configuration instructions), and the actual effect after deployment (such as measured latency, whether it was successful, etc.).

[0078] The historical record repository can be used for knowledge retrieval and reuse, knowledge distillation and transfer learning, and champion-challenger evaluation. Knowledge retrieval and reuse means that when similar network scenarios or business needs are encountered in the future, successful strategies from the historical record repository can be retrieved as a starting point for new strategies, significantly accelerating the speed and quality of strategy generation. Knowledge distillation and transfer learning refers to the large amount of high-quality problem-solution-effect triplet data pairs accumulated in the historical record repository, which can be used to distill smaller, more efficient strategy generation models, or to transfer optimization experience from a mature network environment to a newly built network environment, achieving rapid cold start of the model. Champion-challenger evaluation means that the system can periodically introduce new strategy generation algorithms (challengers) and have them compete back-to-back with the currently used best algorithm (champion) against historical scenarios in the repository. Only when a challenger proves to systematically outperform the champion will it be adopted as a new online service model. When new business requirement descriptions are received, similar successful change records are first retrieved from the historical record repository. Similarity can be matched based on features such as business type, source and destination address ranges, and SLA requirements in the requirement description. The most similar successful change record retrieved is used as the initial reference for generating new minimal change operation information. For example, its configuration operation set can be directly reused, and then fine-tuned according to the current network status. This can greatly speed up policy generation and draw on historical success experience. This embodiment solves the problem of low efficiency in repetitive changes and realizes the accumulation and reuse of knowledge.

[0079] It should be noted that there is no strict order of execution between the steps in this invention. As long as they conform to the logical order, these steps can be executed simultaneously or in a certain preset order. Figure 1 This is just an illustrative example and does not mean that this is the only possible execution order.

[0080] This invention also provides a corresponding apparatus for the network configuration information modification method, further enhancing the method's practicality. The apparatus can be described from both a functional module perspective and a hardware perspective. The following describes the network configuration information modification apparatus provided by this invention, which is used to implement the network configuration information modification method provided by this invention. In this embodiment, the network configuration information modification apparatus may include or be divided into one or more program modules. These program modules are stored in a storage medium and executed by one or more processors to complete the network configuration information modification method disclosed in the embodiment. The program module referred to in this embodiment is a series of computer program instruction segments capable of performing a specific function, which is more suitable than the program itself for describing the execution process of the network configuration information modification apparatus in the storage medium. The following description will specifically introduce the functions of each program module in this embodiment. The network configuration information modification apparatus described below can be referred to in correspondence with the network configuration information modification method described above.

[0081] From the perspective of functional modules, see Figure 6 , Figure 6 This is a structural diagram of the network configuration information changing device provided in this embodiment under one specific implementation. The device may include: The model building module 601 is used to build a dynamic simulation model of network mirroring that is updated synchronously with the physical network based on network entity configuration and network operation data, and is used to output network performance change information based on network configuration change information.

[0082] The minimum operation determination module 602 is used to generate minimum change operation information from the current network state to the target network state based on the business requirement description information; the minimum change operation information includes at least configuration change data and the dependencies between configuration change items.

[0083] The multi-dimensional verification module 603 is used to predict the risks of minimal change operation information based on the network mirror dynamic inference model, at least from the dimensions of network logic correctness and network performance impact.

[0084] The strategy deployment module 604 is used to deploy minimum change operation information to the physical network according to dependencies, while maintaining the consistency of business traffic during network changes, when the risk prediction results meet the preset release conditions.

[0085] The continuous optimization module 605 is used to compare the measured performance data of the physical network that performs the minimum change operation with the network performance change information output by the network mirror dynamic simulation model, and to use the deviation information obtained from the comparison to calibrate the network mirror dynamic simulation model.

[0086] The network configuration information changing device mentioned above is described from the perspective of functional modules. Furthermore, the present invention also provides an electronic device, which is described from the perspective of hardware. The electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described network configuration information changing method embodiments.

[0087] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the network configuration information change method when running.

[0088] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0089] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described network configuration information change method embodiments.

[0090] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described network configuration information change method embodiments.

[0091] Finally, the present invention also provides a network configuration information change system, see [link to relevant documentation]. Figure 7 and Figure 8 The system may include a memory 701 for storing a computer program, i.e., software code implementing the aforementioned method embodiments; and a processor 702 for executing the computer program to implement the steps of the network configuration information changing method described in any of the above embodiments. The system may be a standalone server or integrated into a cloud management platform.

[0092] For example, an exemplary workflow for a network configuration information change system could be: The policy lifecycle begins with a specific triggering event. This triggering event can be a high-level business requirement submitted by the cloud management platform administrator, such as allocating high bandwidth for a newly launched AI training service and ensuring online payment latency is below 10 milliseconds; or it can be an automatically triggered optimization requirement triggered by system monitoring of an impending service level agreement breach, persistent network congestion, or device failure. Upon receiving the trigger signal and optimization objective, the policy generator interacts with the network mirror graph, using a multi-objective optimization algorithm to calculate one or more candidate configuration change data, and attaches an update dependency graph to each configuration change data to form minimal change operation information. Subsequently, this minimal change operation information is sent to the online verification and risk grading gateway. The online verification and risk grading gateway simultaneously initiates three parallel verification channels: the intent verification channel checks for routing loops, forwarding black holes, or security policy violations; the invariant guardian channel verifies in real time whether the upcoming instructions violate fundamental rules that the network must always satisfy; and the performance replay and evaluation channel loads the configuration change data into a digital sandbox formed by the network mirror dynamic simulation model, simulating normal traffic, burst traffic, and link interruption scenarios to predict performance indicators such as end-to-end latency, jitter, and packet loss rate. The gateway combines dynamically calculated trust scores, adaptively weights and synthesizes the final risk assessment result, and issues a risk prediction result that includes recommendations for canary releases.

[0093] If the overall risk level of the risk prediction results exceeds the preset automation threshold (e.g., too low a confidence score or too high a prediction performance risk), the release process will be terminated and transferred to manual review; otherwise, it will enter the automatic execution phase. The security execution layer first verifies the probe flow: it issues probe rules to network devices, copies a small portion of business traffic as a probe flow or injects it into a synthetic probe flow, guides it through the new configuration path, and collects measured latency and packet loss data for comparison with model prediction values. After confirmation, it strictly follows the gray-scale scheme and dependency graph in the risk prediction results to deploy the configuration change data to the physical network in batches and gradually, setting a cooling-off observation period between each batch to monitor network stability. During the entire execution process, if any monitoring indicator (such as latency collected by in-band network telemetry or queue depth) reaches the rollback threshold defined in the risk prediction results, the system immediately triggers the circuit breaker mechanism, stops the release, and automatically rolls back, restoring the network to its stable state before the change.

[0094] After the change is successfully implemented and running stably, the closed-loop self-optimization layer continuously compares the measured performance data with the predicted values ​​output by the network mirror dynamic inference model. The resulting prediction-measured deviation is used as new training samples for incremental training or online model calibration. At the same time, the complete event record from business intent to successful implementation (including minimum change operation information, change license documents, measured results, etc.) is structured and stored in the historical record repository.

[0095] The network mirror dynamic extrapolation model also has a lifecycle from construction, calibration, versioning to retirement. The system creates a new version for each major network architecture change or model retraining. During operation, if the system finds that the confidence score of a certain version of the model is consistently below an acceptable threshold (e.g., prediction error greater than 30% for 24 consecutive hours), it considers that the version has drifted too much and cannot reliably guide decision-making. The system automatically marks it as retired and triggers a completely new model construction and training process based on the latest full network snapshot and telemetry data to generate a new version more adapted to the current network conditions. Through this dual closed-loop design, this invention achieves full automation from policy generation, verification, execution to effect verification, and enables the model to continuously evolve.

[0096] To ensure reliability under extreme conditions, the system also features multi-layered anomaly handling paths. Automated checkpoints are set up at multiple key nodes. For example, if the overall risk is too high at the online verification and risk grading gateway, it automatically switches to manual review; if the probe flow verification fails, the release is terminated; and if performance metrics exceed preset thresholds during canary releases, a rollback is triggered. Triggering any checkpoint will cause the automated process to pause or rollback, ensuring that problematic changes are not pushed to the production environment. When the invariant guardian channel or intent verification channel detects logical conflicts (such as potential routing loops or security isolation violations), the system not only blocks the change but also generates a precise diagnostic report based on a formal logical model and dependency graph, identifying the network device, forwarding entry, and dependency chain causing the conflict, assisting operations personnel in quickly troubleshooting. For issues with decreased credibility of the network mirror dynamic simulation model, the system performs attribution analysis, classifying it into different types such as model error (requiring retraining), missing telemetry data (requiring inspection of the acquisition link), or abnormal device behavior (requiring diagnosis of specific devices), and triggers corresponding automated or semi-automated governance processes.

[0097] In extreme cases, operations and maintenance personnel have the highest privileges. The system provides a global emergency stop switch, which, once triggered, immediately halts all ongoing policy generation and change activities and restores the entire network's policies to a predefined, absolutely secure baseline state. Simultaneously, the system provides complete, tamper-proof audit logs and decision timelines, detailing all automated operations and decision-making processes from the generation of business intent to the emergency stop, for post-event review and analysis.

[0098] In summary, this invention's adaptive risk assessment mechanism integrates three key technologies—static logic verification, real-time invariant protection, and dynamic performance prediction—into a unified verification gateway, constructing a complete process from business intent to change execution and closed-loop learning. Through mechanisms such as minimal change operation information, consistent updates, and probe flow verification, it ensures the refinement, security, and controllability of the change process. Furthermore, through a historical record repository and continuous verification calibration, it endows the system with long-term self-optimization capabilities, providing an automated solution for software-defined network management in private cloud environments that combines intelligence, high reliability, and strong robustness.

[0099] Furthermore, to facilitate the deployment of this invention in a private cloud environment, this embodiment interfaces with the upper-layer cloud management platform and the lower-layer network infrastructure via a hardware communication interface. Correspondingly, such as... Figure 9 As shown, the network configuration information change system also includes a first hardware communication interface, such as an Ethernet interface or a PCIe interface, for communicating with the cloud management platform, and a second hardware communication interface, such as another Ethernet interface or a dedicated management interface, for communicating with physical network devices. The processor receives service requirement description information through the first hardware communication interface and sends configuration change data to the physical network devices through the second hardware communication interface. The two interfaces can be independent physical network cards or different virtual interfaces on the same network card.

[0100] For example, in cloud computing, traffic from outside the cloud to inside the cloud is typically referred to as southbound, while traffic from inside the cloud to outside the cloud is northbound. Correspondingly, the first hardware communication interface is the northbound interface, a highly abstract service-oriented interface based on user intent, targeting cloud operating systems (such as OpenStack, VMware vSphere, etc., which deploy and manage containers within lightweight virtual machines) or CMP (Cloud Management Platform). Cloud platform administrators do not need to concern themselves with the complex details of the underlying SDN; they only need to submit business requirements through this interface. For example, an application programming interface (API) call to create a high-security video conferencing network might only include the tenant ID, a list of participant IP addresses, and SLA requirements (such as latency <20ms, bandwidth >100Mbps). Once the system receives the intent, it automatically completes the entire process of policy generation, verification, execution, and optimization, and feeds back key statuses (such as risk prediction results and canary release progress) to the cloud management platform via callbacks or query interfaces. The second hardware communication interface is a southbound interface, facing the underlying network control and data plane. It is responsible for translating the unified, abstract policy commands from the upper layers into configurations that can be understood by devices from different manufacturers and using different technologies. This layer is compatible with SDN controllers (such as open network operating systems), virtual switches (such as Open vSwitch (OVS)), and programmable hardware (such as switch processors that support P4 language (Programming Protocol-independent Packet Processors)) through a pluggable adapter mode. It interacts through the P4 runtime interface to support advanced functions such as in-band network telemetry.

[0101] As can be seen from the above, this embodiment solves the problem of how the system can interface with the external environment and clarifies the hardware implementation of the north-south interface.

[0102] Finally, to enable those skilled in the art to better understand the technical solution of the present invention, the present invention also provides an exemplary application embodiment, which may include the following: The private cloud platform adopts a software-defined networking (SDN) architecture. The operations team was tasked with allocating high-bandwidth resources to the newly launched AI training service while ensuring that the end-to-end latency of the existing online payment service is strictly controlled within 10 milliseconds. The cloud management platform administrator submitted high-level intents through the northbound application programming interface (API), namely, to provide maximum bandwidth for traffic from the source address group (AI_Cluster) to the destination address group (Storage_Cluster), while ensuring that the end-to-end latency of the service (Online_Pay) is below 10 milliseconds according to the Service Level Agreement (SLA).

[0103] Upon receiving the intent, the system first queries the network mirror dynamic simulation model for information such as the current network topology, link load, queue depth, and effective policies. The built-in policy generator recognizes this as a coexistence optimization problem and uses a solution algorithm to calculate a candidate minimum change operation. This operation includes three configuration change data points: applying an equivalent multipath routing policy to AI training traffic on leaf switches Switch-L1 and Switch-L2 to utilize multiple uplinks; mapping the differential service code point value of the service flow to a high-priority queue on all switch port queues involving online payment service paths; and setting an initial shaping rate for AI training traffic on the leaf switch egress queues to prevent AI training traffic from affecting other services. Simultaneously, the system generates a dependency graph, indicating that the priority mapping operation must be completed before or simultaneously with the routing policy and shaping operation to prioritize critical services.

[0104] Minimal change operation information is sent to the online verification and risk classification gateway, triggering three parallel channels. The intent verification channel performs logical analysis of routing policies and priority mappings to confirm that it will not introduce routing loops, disrupt the basic reachability of online payment services, or violate security isolation policies between tenants, outputting a low structural risk score. The invariant online guardian channel pre-configures network invariants, ensuring that any traffic marked as online payment cannot be downgraded or dropped. Channel analysis confirms that the change conforms to this invariant, outputting an extremely low invariant risk score. The performance replay and evaluation channel applies configuration change data to a dynamic network mirroring model, forming a digital sandbox of future states. Historical traffic data is loaded to simulate the highly bursty traffic patterns of AI training services and the short, high-frequency request patterns of online payment services. A physically-guided graph neural network is launched for simulation, and the model captures the possibility that micro-bursts in AI training traffic may cause momentary congestion on the spine switch. Model prediction: Without intervention, there is a 20% probability that the end-to-end latency of online payment transactions will exceed 10 milliseconds, reaching 12 to 15 milliseconds. After applying the priority strategy, the 99.9th percentile latency will stabilize at around 6 milliseconds, while the total throughput of AI training transactions reaches 80 gigabits per second. The parallel network computation module calculates the upper bound of the payment transaction latency in the worst case as 9.5 milliseconds, providing a deterministic boundary for graph neural network prediction. The channel comprehensive output performance risk is divided into low to medium risk levels, with a high stability score, and a prediction uncertainty range is given, for example, the predicted payment transaction latency is 6 milliseconds plus or minus 0.5 milliseconds.

[0105] The online verification and risk grading gateway queries the credibility score of the network mirror dynamic inference model. Due to the recent collection of a large amount of traffic data for similar scenarios via in-band network telemetry, and the small residual between the model's predictions and actual measurements, the current credibility score is 0.92. Based on this high credibility score, the online verification and risk grading gateway assigns a higher weight to the performance risk score when synthesizing the total risk, ultimately calculating a low overall risk level. The system issues a machine-readable change permission document as a risk prediction result that meets the preset release conditions. This document includes the change identifier, details of the minimum change operation, risk scores for each channel, the overall risk level, the credibility score, and a system recommendation: adopt a probe flow verification plus two-stage release mode; the rollback trigger condition is that the measured end-to-end latency of the online payment business exceeds 8 milliseconds three times consecutively; the rollback operation point is the revocation of all related policies.

[0106] After receiving the change permission document, the security execution layer continues the automated process due to the low risk level. First, probe flow verification is performed: probe rules are sent to the ingress switch via the programming interface, replicating 1% of online payment production traffic and marking it as a probe flow. These probe flows are directed to the path applying the new policy, carrying in-band network telemetry instructions, and the header upon return includes hop-by-hop actual latency and queue information. The system parses the telemetry report and finds that the average latency of the probe flows is 6.2 milliseconds, falling entirely within the predicted 6 milliseconds plus or minus 0.5 milliseconds confidence interval, thus verifying success. Subsequently, according to the dependency graph, priority mapping operations are first issued, followed by routing policies and shaping operations. Initially, only 10% of AI training traffic is switched to the new policy path, and the system enters a 5-minute cooldown observation period, continuously monitoring the latency of online payment transactions and the throughput of AI transactions; all indicators meet expectations. Then, the remaining 90% of AI training traffic is switched to the new policy path, and the network status remains stable.

[0107] After the strategy goes live, the closed-loop self-optimization layer continues to operate. It continuously collects the actual latency of online payment transactions through in-band network telemetry, stabilizing at 6.2 milliseconds, and compares this with the 6 milliseconds predicted by the graph neural network model. This stable 0.2-millisecond deviation serves as a new training sample, feeding back into the network mirror dynamic inference model. The system backend triggers an online fine-tuning of the model, adjusting the weights and correction coefficients of some neurons in the graph neural network to improve prediction accuracy for similar scenarios in the future. Simultaneously, the entire event from intent to successful launch, including minimal change operation information, change license documents, and actual test results, is fully stored as a success case in the historical record repository. In the future, if other businesses require high bandwidth, the system can directly retrieve and reuse this successful model from the repository.

[0108] Through the above implementation methods, this invention transforms a complex, high-risk network change into a highly automated, measurable, controllable, and learnable intelligent process. Compared with related technologies, this invention transforms change risks into quantifiable risk prediction results through cross-validation across three dimensions: static, invariant, and performance prediction. It utilizes a probe flow mechanism to conduct experiments without affecting real business operations, ensuring verification before implementation. Combined with canary releases and automatic rollback mechanisms, it significantly improves the security and stability of network changes. Through a dual-engine architecture integrating physically guided graph neural networks and network computation, it achieves accurate quantitative prediction of performance indicators such as latency, jitter, and packet loss rate, upgrading network management from passive response to proactive prediction. End-to-end automation from intent understanding, policy generation, verification, execution to verification and calibration significantly improves operational efficiency and reduces labor costs and human error. Continuous residual calibration and historical record reuse build a self-optimizing capability for sustainable evolution. By introducing a credibility score to quantify model reliability and adaptively adjusting decision-making strategies, it achieves quantitative management of uncertainty, improving efficiency while ensuring security and robustness.

[0109] The foregoing has provided a detailed description of a network configuration information modification method and system provided by the present invention. The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Whether the units and algorithm steps of the various examples described in the disclosed embodiments are executed by electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, and such implementations should not be considered beyond the scope of the present invention. Several improvements and modifications can be made to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A method for changing network configuration information, characterized in that, include: Based on network entity configuration and network operation data, a network mirror dynamic simulation model is constructed that is updated synchronously with the physical network and is used to output network performance change information based on network configuration change information. Based on the business requirement description, generate minimum change operation information for changing from the current network state to the target network state; the minimum change operation information includes at least configuration change data and the dependencies between configuration change items; Based on the network mirroring dynamic simulation model, the minimum change operation information is risk-predicted from at least the dimensions of network logical correctness and network performance impact. When the risk prediction results meet the preset release conditions, the minimum change operation information is deployed to the physical network according to the dependency relationship while maintaining the consistency of business traffic during network changes. The measured physical network performance data of the physical network that performs the minimum change operation information is compared with the network performance change information output by the network mirror dynamic simulation model, and the deviation information obtained from the comparison is used to calibrate the network mirror dynamic simulation model.

2. The network configuration information modification method according to claim 1, characterized in that, The network mirroring dynamic simulation model includes a data layer and a performance prediction layer, and its construction process includes: Based on network entity configuration and network operation data, a network mirror graph is constructed as the data layer, with network entity objects as graph nodes and the existence of links between graph nodes as graph edges. A performance prediction layer is constructed based on the target graph neural network, physical constraint layer, network computation layer and output layer. The target graph neural network is a graph neural network structure with spatial convolutional layers followed by cascaded temporal processing layers. Based on the network mirror graph, a heterogeneous network state graph is constructed as the input to the performance prediction layer, using multiple types of network objects as graph nodes and determining node edges based on the connections between network objects. In the target graph neural network, the state of various graph nodes is updated in stages according to the association between various graph nodes in the heterogeneous network state graph and the upstream and downstream dependency order of message passing; During the state update process of each graph node, the physical constraint layer embeds matching physical constraint conditions according to the network object type corresponding to each type of graph node. The network computation layer runs in parallel with the target graph neural network, providing physical boundary constraints for the target graph neural network and outputting analytical baseline values; The output layer determines the network performance change prediction result based on the analytical baseline value and the predicted value output by the target graph neural network.

3. The network configuration information modification method according to claim 2, characterized in that, The process of constructing the heterogeneous network state diagram includes: Network devices, network service quality, and service traffic characteristics are taken as network objects, and the graph vertices of the heterogeneous network state graph are determined as network device nodes, queue nodes, and flow nodes. Based on the physical connections between network devices, the mapping relationship between traffic and queues, and the internal affiliation relationship between devices and queues, determine whether there are node edges between the vertices of the heterogeneous network state diagram; The network device features representing the computing and forwarding capabilities of the network device are configured as the initial feature data of the network device node; the queue features representing the cache depth and queue scheduling behavior of the physical port are configured as the initial feature data of the queue node; and the service flow session features representing the end-to-end flow are configured as the initial feature data of the flow node. Accordingly, the graph node update process of the heterogeneous network state graph is as follows: based on the logical mapping relationship between the queue node and the flow node, the internal topology relationship between the network device node and the queue node, and the path association relationship between the flow node and the nodes on its forwarding path, the various graph vertices of the heterogeneous network state graph are updated in stages.

4. The network configuration information modification method according to claim 2, characterized in that, The graph nodes are network device nodes, queue nodes, and flow nodes. The process of updating the state of various graph nodes in stages includes: During the queue node update phase, for each queue node, the first-line node that points to the current queue node through a logical mapping relationship is determined. Based on the traffic data of each first-line node, the processing capacity of the network device node to which the current queue node belongs, and the link capacity characteristics of the current queue node, the current characteristic data of the current queue node is updated. During the network device node update phase, for each network device node, the first queue node connected to the current network device node is determined, and the current characteristic data of the current network device node is updated according to the status information of each first queue node and the global routing control message. During the flow node update phase, for each flow node, the second queue node and the second network device node traversed on the current flow node's forwarding path are determined. Based on the status feedback information of each second queue node and each second network device node, the queuing delay, processing delay, and packet loss probability along the way are accumulated to update the current characteristic data of the current flow node.

5. The method for changing network configuration information according to claim 2, characterized in that, The graph nodes are network device nodes, queue nodes, and flow nodes. Based on the network object type corresponding to each type of graph node, matching physical constraints are embedded, including: If the network object type corresponds to a queue node, the link utilization is calculated based on the aggregated traffic of the previous layer. The correction coefficient and the basic deviation are determined based on the current traffic burst characteristics and long-tail distribution characteristics of the target graph neural network. The queuing delay characteristics are updated using the correction coefficient to correct the link utilization and the basic deviation. If the network object type corresponds to a network device node, the instantaneous throughput characteristic of its aggregation is forcibly limited to not exceeding the physical backplane bandwidth of the corresponding network device node, and a lower bound constraint condition for the processing latency characteristic of the corresponding network device node is set; the lower bound constraint condition for latency is that the processing latency of the network device node is greater than or equal to the minimum physical latency. If the network object type corresponds to a flow node, a physical lower bound is forcibly set for the end-to-end delay prediction value of the flow node. The physical lower bound is the sum of the predicted end-to-end total delay and the light speed propagation delay of all physical links traversed by the corresponding flow node and the minimum processing delay of all devices.

6. The network configuration information modification method according to claim 2, characterized in that, The output layer outputs the mean latency prediction, variance latency prediction, jitter prediction, and packet loss rate prediction for the target service flow; wherein, the variance latency prediction represents the prediction uncertainty, and the value of the variance latency prediction increases when the network is in an unknown abnormal state. Based on the predicted mean latency, the predicted packet loss rate, and the preset service level agreement threshold, a performance risk score is determined, and a stability score is calculated using an inverse proportional mapping function based on the predicted jitter and the predicted latency variance.

7. The method for changing network configuration information according to claim 2, characterized in that, The network computation layer determines the upper bound of latency and the upper bound of backlog that match the current network environment based on the current traffic characteristics, and inputs the upper bound of latency and the upper bound of backlog as physical boundary constraints into the target graph neural network. The target graph neural network adds the physical boundary constraints to the loss function or as an upper limit for output truncation.

8. The method for changing network configuration information according to claim 2, characterized in that, The output layer, when in the target stage, outputs the analytical baseline value as the network performance change prediction result; when not in the target stage, it adds the analytical baseline value to the residual output of the target graph neural network and outputs it as the network performance change prediction result. The target stage is the stage where the training data sample size of the target graph neural network is lower than a preset sample size threshold, and the target graph neural network is configured to learn the residual between the measured performance value of the network and the analytical baseline value.

9. The method for changing network configuration information according to claim 2, characterized in that, The process of constructing the network mirror map includes: Treat network devices or their ports as network entity objects, create corresponding graph nodes for each network entity object, and configure node attributes for each graph node; the node attributes include the port's physical capacity, buffer size, queue scheduling algorithm type, and congestion management mechanism; If the first network entity object and the second network entity object have physical links or logical links, then set graph edges for the first network entity object and the second network entity object, and configure edge attributes for each graph edge. Configure graph attributes for the network mirror graph based on the currently effective traffic matrix, resource allocation weights of service slices, service level agreement requirements, and global traffic shaping or rate limiting policies.

10. The network configuration information modification method according to any one of claims 1 to 9, characterized in that, Based on the business requirement description, generate the minimum change operation information from the current network state to the target network state, including: Parse the business requirement description information to obtain the source address group, destination address group, and service level agreement requirements of the business flow; The current network status is obtained from the network mirror dynamic simulation model, including the current network resource usage status, topology connection data, and effective network configuration data. Determine the target network status based on service level agreement requirements, security isolation requirements between tenants, network device processing capacity limitations, and network bandwidth resource limits; The current network state is compared with the target network state item by item to identify the target network configuration items that have changed, and the target configuration operations performed on each target network configuration item are organized into a configuration operation set as configuration change data. Using each target configuration operation as a relation node, the execution order constraints and resource mutual exclusion relationships between the target configuration operations in the configuration operation set are analyzed to determine whether there are connecting edges between the relation nodes, and a dependency graph is generated. Based on the configuration operation set and the dependency graph, generate minimal change operation information.

11. The network configuration information modification method according to claim 10, characterized in that, After generating the dependency graph, the following steps are also included: Based on the execution order indicated by the connecting edges of the dependency graph, the configuration operation set is divided into multiple sub-change sets; there are no dependencies between the target configuration operations in each sub-change set or the dependencies have been resolved sequentially within the sub-set, and there is a sequential execution order between different sub-change sets; The target configuration operations for each sub-change set are executed sequentially according to the order described. After all target configuration operations in a sub-change set are completed, a cooling-off observation phase is entered. During the cooling-off observation phase, the target performance indicators of the network are monitored to see if they exceed the preset performance indicator thresholds. If they do not exceed the thresholds, the target configuration operations in the next sub-change set are executed. If they exceed the thresholds, the execution of the remaining sub-change sets is terminated, and a rollback operation is triggered.

12. The network configuration information modification method according to any one of claims 1 to 9, characterized in that, Risk prediction should be performed on the minimum change operation information from at least the dimensions of network logical correctness and network performance impact, including: The configuration change data in the minimum change operation information is applied to the formal logical model representing the current network configuration; The formal logic model is used to simulate packet forwarding paths, and logical errors are detected during the simulation. The logical error type and its associated network device identifier and forwarding table entry identifier are output as logical error detection results. The logical errors include packet forwarding loops, packets being dropped at unexpected nodes, inter-service connectivity interruptions, or violations of security isolation rules. The configuration change data in the minimum change operation information is applied to the network mirror dynamic simulation model to serve as a digital sandbox representing the future state of the network. A preset traffic pattern is loaded in the digital sandbox, and the network performance change information of the configuration change data under different traffic patterns is determined by the network mirror dynamic simulation model, so as to serve as the performance impact estimate; the traffic pattern includes regular traffic, burst traffic and link interruption scenarios.

13. The network configuration information modification method according to claim 12, characterized in that, Before deploying the minimum change operation information to the physical network according to the aforementioned dependencies, the method further includes: Before sending the configuration change data in the minimum change operation information to the physical network, intercept the configuration waiting execution instruction that is in the sending state or waiting to be sent state; Parse the matching fields of the configuration waiting to execute instruction to determine the target data packet affected by the configuration waiting to execute instruction; On the global data plane state graph maintained in memory, the forwarding path corresponding to the target data packet is incrementally updated, and a new forwarding path is determined after the configuration waits for the execution instruction to take effect; The characteristics of the new forwarding path are logically compared with a preset network invariant, which is a Boolean logic rule satisfied by the network. If the characteristics of the new forwarding path do not conform to the network invariant, the configuration is prevented from waiting for the execution instruction to be issued, and a non-compliance prompt message is output.

14. The network configuration information modification method according to claim 12, characterized in that, After performing risk prediction on the minimum change operation information from at least the dimensions of network logical correctness and network performance impact, it also includes: The quantitative value of the degree of conformity between the predicted network performance changes of the network mirror dynamic simulation model and the actual effect after the corresponding network configuration changes are implemented during the historical verification period of the statistical target. The credibility score of the network mirror dynamic inference model is determined based on the error index between the predicted network performance change results and the measured network performance change values, the quantification value of the coverage between the training data of the network mirror dynamic inference model and the network scene to be predicted, the detection score of whether the network state of the network scene to be predicted deviates from the known distribution, and the quantification value of the conformity. Based on the confidence score, the weight values ​​of the logic error detection result and the performance impact estimate are adjusted respectively, wherein the weight value of the performance impact estimate is directly proportional to the confidence score, and the weight value of the logic error detection result is inversely proportional to the confidence score.

15. The network configuration information modification method according to any one of claims 1 to 9, characterized in that, While maintaining consistency of service traffic during network changes, the minimum change operation information is deployed to the physical network according to the aforementioned dependencies, including: Before switching the service traffic, a probe rule is sent to the network devices through which the service traffic passes, and a portion of the traffic is selected from the service traffic and copied as probe traffic or injected as synthetic probe traffic; The probe traffic is guided through the new configuration path described by the configuration change data, and latency and packet loss data of the probe traffic on the new configuration path are collected. The latency, packet loss data, and performance predictions output by the network mirror dynamic simulation model are compared. If the deviations between the latency, packet loss data, and performance predictions are within a preset deviation range, the service traffic is switched to the new configuration path.

16. The method for changing network configuration information according to claim 15, characterized in that, Switching the service traffic to the newly configured path includes: Pre-install the new forwarding table entry on network devices involved in path changes, and set the new forwarding table entry to an inactive state; At the entry device of the data stream, a path identifier is added to the arriving data packets so that the data packets with the added path identifier are forwarded according to the new path, while the data packets without the added path identifier continue to be forwarded according to the original path. After all packets without path identifiers leave the network, remove the old forwarding table entry and modify the new forwarding table entry to be active.

17. The network configuration information modification method according to any one of claims 1 to 9, characterized in that, The network mirror dynamic inference model is calibrated using the deviation information obtained from the comparison, including: The residual between the measured performance data of the physical network and the network performance change information is used as a new training sample. The newly added training samples are used to perform incremental training or online calibration of the network mirror dynamic inference model; The output confidence of the network mirror dynamic inference model is updated based on the calibrated model.

18. The network configuration information modification method according to claim 17, characterized in that, Also includes: The complete record of each change event is structured and stored in the historical record repository. The complete record includes business requirement description information, minimum change operation information, and the actual effect after deployment. When a new business requirement description is received, similar successful change records are retrieved from the historical record repository and used as initial reference data for generating minimal change operation information.

19. A network configuration information modification system, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the network configuration information change method as described in any one of claims 1 to 18 when executing the computer program.

20. The network configuration information change system according to claim 19, characterized in that, It also includes a first hardware communication interface for communicating with the cloud management platform, and a second hardware communication interface for communicating with physical network devices; The processor receives the service requirement description information through the first hardware communication interface and sends configuration change data to the physical network device through the second hardware communication interface.