A large model reinforcement learning network configuration generation method based on verification feedback

By using a large-model reinforcement learning network configuration generation method based on validation feedback, and by optimizing the communication network configuration using graph neural networks and causal graphs, the high policy generation latency and insufficient adaptability of existing technologies are solved, and real-time response and dynamic optimization of network configuration are achieved.

CN120525019BActive Publication Date: 2026-03-17HUBEI THREE GORGES POLYTECHNIC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510463855.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2026-03-17
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing automated configuration technologies for communication networks rely on predefined policy templates, which are difficult to cope with sudden demands, have high policy generation latency and poor accuracy, and cannot adapt to changes in network topology and migration of business models in real time.

Method used

A large-scale reinforcement learning network configuration generation method based on validation feedback is adopted. Semantic parsing is performed through graph neural networks, and policy optimization is performed by combining causal graphs and digital twin systems. The network state is updated in real time, and the configuration policy is adjusted through reinforcement learning.

Benefits of technology

It enables real-time response and dynamic optimization of network configuration, improves the accuracy of policy generation and the adaptability of the system, reduces the risk of QoS degradation, and enhances the flexibility and stability of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525019B_ABST
    Figure CN120525019B_ABST
Patent Text Reader

Abstract

The application discloses a kind of big model reinforcement learning network configuration generation methods based on verification feedback, input network phenomenon and state content are converted into semantic and action sequence by network configuration semantic analysis, based on semantic action sequence, through hybrid action space strategy generation and neural symbol collaborative reinforcement learning model, the configuration framework and parameter that meet the requirements are generated, the network configuration information generated is verified and fed back in digital twin system, formal verification and performance simulation are carried out in virtual environment, real network scene is simulated, the correctness and performance of configuration are comprehensively evaluated, and then feedback signal containing multidimensional information is generated, strategy model is modified and optimized according to digital twin verification, finally, through reward mechanism, intelligent agent is guided to adjust high-entropy configuration item, and optimization configuration generation big model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and network communication technology, and in particular relates to a method for generating configurations for large-scale reinforcement learning networks based on verification feedback. Background Technology

[0002] With the rapid development of digital technology, communication networks are playing an increasingly important role in social life and industrial operations. From enterprise office networks to large data center networks supporting online services, and then to mobile networks used daily by people, their scale and complexity are growing exponentially. As network traffic increases, network configuration management also faces enormous challenges. Traditional network configuration management mainly relies on manual operation, which is not only inefficient and prone to human error, but also unable to cope with large-scale, dynamically changing network environments. Even small networks involve multiple devices and complex network policies, making the configuration process cumbersome and time-consuming. For large data centers and carrier networks, the difficulty of configuration management increases exponentially. Furthermore, with the widespread application of artificial intelligence, 5G, the Internet of Things, and cloud computing, the demand for network dynamism and flexibility has increased significantly. Networks need to be rapidly adjusted and reconfigured based on real-time business needs and changes in user behavior.

[0003] Existing automated configuration technologies for communication networks rely on predefined policy templates. These fixed templates struggle to handle sudden surges in demand, and policy activation delays exceeding service tolerance thresholds (measured at 12-15 seconds) occur during traffic spikes or topology changes, leading to a significant increase in QoS degradation risks. While open-source verification tools (such as Batfish) can detect configuration conflicts, their offline analysis mode prevents the embedding of a configuration generation loop. Current research employs deep learning models (e.g., GNN-based models) for communication network configuration generation; however, model training relies on historical static datasets, resulting in temporal mismatch with dynamic network environments and causing policy generation accuracy to decrease with network state changes. Therefore, a large-scale reinforcement learning network configuration generation method based on verification feedback is needed to address these issues. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a large-model reinforcement learning network configuration generation method based on verification feedback. It aims to solve the problems of existing technologies for automated network communication configuration that rely on predefined templates, are relatively rigid, cannot cope with sudden needs outside the templates, have high policy generation latency, are prone to system risks, and have poor accuracy. The invention has the characteristics of obtaining network status feedback through real-time environmental interaction, breaking through the response speed limitation of static rule base, and enabling the system to continuously adapt to network topology evolution and business model migration.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0006] A method for generating configurations for large-scale reinforcement learning networks based on validation feedback includes the following steps:

[0007] S1, semantic parsing based on graph neural network configuration:

[0008] The input unstructured business requirements are transformed into semantic action sequences. By constructing a formal ontology and dynamic knowledge graph, the network state is updated in real time by combining a temporal graph convolutional network.

[0009] S2, a hybrid action space construction strategy generation model for communication network configuration:

[0010] A hierarchical neural symbolic reinforcement learning method is adopted to collaboratively optimize symbolic rules and continuous parameters, generating a configuration framework that includes device parameters and link connections.

[0011] S3, Policy optimization based on communication network configuration using cause-effect graph computation:

[0012] The generated configuration is input into the digital twin system for simulation verification. The causal entropy of the nodes is calculated based on the causal graph, high-entropy nodes are selected, and the optimization strategy is carried out through reinforcement learning.

[0013] S4, Optimize the strategy generation model:

[0014] Based on causal entropy values ​​and feedback signals, the configuration of high-entropy nodes is adjusted, and the configuration strategy is adjusted through reinforcement learning algorithms, thereby optimizing the overall operation of the network and ultimately generating a large model with optimized configuration.

[0015] Preferably, in step S1, semantic parsing specifically includes:

[0016] A formal ontology library is constructed using the OWL-based method, and relationships between devices, protocols, and topology entities are defined.

[0017] Real-time updates of link load and device performance status using time-series graph convolutional networks:

[0018] ;

[0019] In the formula, Indicates at time step t The hidden state matrix of nodes in the time graph. For activation function, For the normalized adjacency matrix, To include node-level features, For learnable graph / node feature transformation matrices, b This is the paranoia vector.

[0020] Preferably, in step S2, strategy generation specifically includes:

[0021] Calculate the probability of sign rule selection using probabilistic logic programming formulas:

[0022] ;

[0023] In the formula, Indicates the state Next, select symbol rules The probability of adopting a specific routing strategy or symbolic rule under a specific network state; This represents the current operating status of the network. This represents a specific network configuration rule. This represents the total number of rules in the symbol rule set; It is a function that measures the state. Below, symbol rules The score or fit; the higher the score, the greater the likelihood that the rule will be selected in the current state.

[0024] Preferably, the strategy generation also includes a dynamic weight fusion module, which automatically adjusts the contribution ratio of symbol rules and continuous parameter optimization according to the degree of network load fluctuation.

[0025] Preferably, in step S3, strategy optimization specifically includes:

[0026] Each configuration item in the network, including bandwidth allocation and routing settings, is treated as a node. The interrelationships between configuration items are considered as edges. Build a cause-effect graph for configuration items Perform causal analysis and extract configuration item dependencies from the verification logs;

[0027] For each configuration item node Determine the set of all possible states or events related to this. And estimate each state or event. The probability of occurrence; the causal entropy of a node is calculated using the following formula:

[0028] ;

[0029] In the formula, Represents a node causal entropy, It is a random variable, representing the relationship with the node. A related state or event, It is with nodes The set of all possible states or events related to this. Indicates at node In the relevant context, state or event The probability of occurrence.

[0030] Preferably, in step S4, model optimization specifically includes:

[0031] By adjusting the configuration strategy through reinforcement learning algorithms, the overall operating state of the network is optimized, and finally, a large model is generated through optimized configuration.

[0032] Preferably, the verification process of the digital twin system includes:

[0033] Formal verification of configuration conflict detection, and performance simulation of real network scenarios;

[0034] Generate multi-dimensional feedback signals including latency, throughput, and packet loss rate for policy optimization.

[0035] Preferably, the construction of the cause-effect graph further includes:

[0036] Extract configuration item dependencies from verification logs and dynamically update edge EE weights based on historical data and real-time status;

[0037] The incremental model update algorithm adapts to network topology evolution and business model migration.

[0038] Preferably, the reward mechanism for reinforcement learning includes:

[0039] The reward function is set according to the causal entropy value of the node, and the configuration adjustment actions of high-entropy nodes are optimized first.

[0040] The reward value is positively correlated with network performance metrics, including QoS and load balancing, and negatively correlated with the number of configuration conflicts.

[0041] Preferably, Natural Language Processing (NLP) is introduced in the semantic parsing stage to parse the ambiguous semantics in unstructured requirements; the robustness of the graph neural network is enhanced through adversarial training to avoid parsing bias caused by noisy data.

[0042] The beneficial effects of this invention are as follows:

[0043] 1. This method firstly transforms the input network phenomena and state content into semantic and action sequences that can be understood by the agent in reinforcement learning through network configuration semantic parsing. Then, based on these semantic action sequences, a hybrid action space strategy is generated using a neural symbolic cooperative reinforcement learning model to fully consider network constraints and generate a configuration framework and parameters that meet the requirements, providing a specific implementation plan for network configuration. Subsequently, the generated network configuration information is verified and feedback is provided in a digital twin system. In a virtual environment, formal verification and performance simulation are used to simulate real network scenarios, comprehensively evaluate the correctness and performance of the configuration, and generate feedback signals containing multi-dimensional information. Finally, in the configuration strategy optimization stage, the strategy model is corrected and optimized based on the feedback obtained from the digital twin verification. The configuration strategy is adjusted through reinforcement learning algorithms to optimize the overall network operation state, and finally, the optimized configuration generates a large model.

[0044] 2. This method possesses the capability of "environmental perception-intelligent decision-making-closed-loop verification". It applies a reinforcement learning-driven dynamic policy generation mechanism and obtains network status feedback through real-time environmental interaction, breaking through the response speed limitation of static rule base. At the same time, it develops an incremental model update algorithm, enabling the system to continuously adapt to network topology evolution and business model migration. Attached Figure Description

[0045] Figure 1 This is a flowchart of the present invention;

[0046] Figure 2 This is a schematic diagram of the entire process in an embodiment of the present invention. Detailed Implementation

[0047] Example 1:

[0048] like Figure 1 As shown, a method for generating configurations for large-scale reinforcement learning networks based on validation feedback includes the following steps:

[0049] S1, semantic parsing based on graph neural network configuration:

[0050] The input unstructured business requirements are transformed into semantic action sequences. By constructing a formal ontology and dynamic knowledge graph, the network state is updated in real time by combining a temporal graph convolutional network.

[0051] S2, a hybrid action space construction strategy generation model for communication network configuration:

[0052] A hierarchical neural symbolic reinforcement learning method is adopted to collaboratively optimize symbolic rules and continuous parameters, generating a configuration framework that includes device parameters and link connections.

[0053] S3, Policy optimization based on communication network configuration using cause-effect graph computation:

[0054] The generated configuration is input into the digital twin system for simulation verification. The causal entropy of the nodes is calculated based on the causal graph, high-entropy nodes are selected, and the optimization strategy is carried out through reinforcement learning.

[0055] S4, Optimize the strategy generation model:

[0056] Based on causal entropy values ​​and feedback signals, the configuration of high-entropy nodes is adjusted, and the configuration strategy is adjusted through reinforcement learning algorithms, thereby optimizing the overall operation of the network and ultimately generating a large model with optimized configuration.

[0057] Example 2:

[0058] like Figure 2 As shown, this embodiment proposes a digital twin verification feedback reinforcement learning optimization framework that integrates semantic parsing, hybrid action space policy generation, digital twin verification, and feedback-driven digital twin verification. Through innovative technologies such as neural symbolic collaboration, causal reasoning, and quantum acceleration, it achieves autonomous generation and dynamic optimization of network configuration.

[0059] Preferably, semantic parsing specifically includes:

[0060] A formal ontology library is constructed using the OWL-based method, and relationships between devices, protocols, and topology entities are defined.

[0061] Real-time updates of link load and device performance status using time-series graph convolutional networks:

[0062] ;

[0063] In the formula, Indicates at time step t The hidden state matrix of nodes in the time graph. For activation function, For the normalized adjacency matrix, To include node-level features, For learnable graph / node feature transformation matrices, b This is the paranoia vector.

[0064] Preferably, strategy generation specifically includes:

[0065] Calculate the probability of sign rule selection using probabilistic logic programming formulas:

[0066] ;

[0067] In the formula, Indicates the state Next, select symbol rules The probability of adopting a specific routing strategy or symbolic rule under a specific network state; This represents the current operating status of the network. This represents a specific network configuration rule. This represents the total number of rules in the symbol rule set; It is a function that measures the state. Below, symbol rules The score or fit; the higher the score, the greater the likelihood that the rule will be selected in the current state.

[0068] Preferably, the strategy generation also includes a dynamic weight fusion module, which automatically adjusts the contribution ratio of symbol rules and continuous parameter optimization according to the degree of network load fluctuation.

[0069] Preferably, strategy optimization specifically includes:

[0070] Each configuration item in the network, including bandwidth allocation and routing settings, is treated as a node. The interrelationships between configuration items are considered as edges. Build a cause-effect graph for configuration items Perform causal analysis and extract configuration item dependencies from the verification logs;

[0071] For each configuration item node Determine the set of all possible states or events related to this. And estimate each state or event. The probability of occurrence; the causal entropy of a node is calculated using the following formula:

[0072] ;

[0073] In the formula, Represents a node causal entropy, It is a random variable, representing the relationship with the node. A related state or event, It is with nodes The set of all possible states or events related to this. Indicates at node In the relevant context, state or event The probability of occurrence.

[0074] Preferably, model optimization specifically includes:

[0075] By adjusting the configuration strategy through reinforcement learning algorithms, the overall operating state of the network is optimized, and finally, the optimized configuration is used to generate a large model.

[0076] Preferably, the verification process of the digital twin system includes:

[0077] Formal verification of configuration conflict detection, and performance simulation of real network scenarios;

[0078] Generate multi-dimensional feedback signals including latency, throughput, and packet loss rate for policy optimization.

[0079] Preferably, the construction of the cause-effect graph further includes:

[0080] Extract configuration item dependencies from verification logs and dynamically update edge EE weights based on historical data and real-time status;

[0081] The incremental model update algorithm adapts to network topology evolution and business model migration.

[0082] Preferably, the reward mechanism for reinforcement learning includes:

[0083] The reward function is set according to the causal entropy value of the node, and the configuration adjustment actions of high-entropy nodes are optimized first.

[0084] The reward value is positively correlated with network performance metrics, including QoS and load balancing, and negatively correlated with the number of configuration conflicts.

[0085] Preferably, Natural Language Processing (NLP) is introduced in the semantic parsing stage to parse the ambiguous semantics in unstructured requirements; the robustness of the graph neural network is enhanced through adversarial training to avoid parsing bias caused by noisy data.

[0086] Example 3: Data Center Network Burst Traffic Scheduling Scenario

[0087] Example Background: In a data center with a Layer 3 Fat-Tree topology network architecture housing 200 servers, the access layer directly connects to all 200 servers. Under normal circumstances, network traffic between servers remains at a relatively stable baseline level, with each server requiring approximately 200Mbps of bandwidth. However, in sudden situations, such as emergency video conferences, network traffic at the access layer surges dramatically. The normal network traffic flow between servers is disrupted, and the bandwidth requirement per server jumps abruptly from the baseline 200Mbps to 800Mbps. The access layer switches face a massive traffic surge and must handle this explosive data growth. The core layer has a multi-path routing mechanism. When sudden traffic from the access layer converges at the core layer, the core layer switches need to make rapid and accurate decisions. The core layer switches must select the optimal transmission path for the massive data traffic based on factors such as real-time network status and link load, ensuring fast and stable data transmission, avoiding network congestion and data loss, and guaranteeing the smooth operation of sudden video conferences.

[0088] Implementation process:

[0089] (1) Semantic parsing

[0090] First, semantic parsing of the network status is performed. The input is an unstructured requirement: "Prioritize video conferencing bandwidth and automatically avoid high-load links." This requirement is rather abstract, requiring analysis of the input text using a graph neural network-based network configuration semantic parsing method. This analysis aims to understand the user's intent at the semantic level and generate a specific semantic action sequence: [Bandwidth allocation priority: Video conferencing > 80%; Link load threshold: 75%; Routing strategy: ECMP optimization]. "Bandwidth allocation priority: Video conferencing > 80%" means that during network bandwidth allocation, the bandwidth allocated to video conferencing should exceed 80% of the total bandwidth to prioritize video conferencing. "Link load threshold: 75%" sets a critical value for link load; when the link load reaches or exceeds 75%, it is considered a high-load link, and the system will automatically avoid it. "Routeing strategy: ECMP optimization" specifies that an equal-cost multipath (ECMP) optimization strategy will be used to achieve reasonable traffic allocation in routing selection.

[0091] In this process, a temporal graph convolutional network was used to predict link states. The network parameters were set as follows: node feature dimension 64, adjacency matrix update frequency 1Hz. Through learning and training on historical link state data, the final prediction results provided a reliable basis for subsequent policy generation and network optimization.

[0092] (2) Strategy generation

[0093] After semantic parsing is completed, the policy generation stage begins. This stage mainly includes three parts: symbol rule selection, continuous parameter optimization, and generation of specific configurations.

[0094] Symbol rule selection: The system's rule base contains 32 predefined routing policies, which are pre-defined based on different network scenarios and requirements. Under the current load surge, after extensive experiments and data analysis, it was found that the ECMP optimization rule has a higher probability of selection than the fixed selection probability of the traditional static rule base. The ECMP optimization rule can more effectively cope with network changes and achieve better traffic distribution and load balancing, and therefore is preferred as the current routing policy.

[0095] Continuous parameter optimization: To further optimize network performance, the bandwidth allocation weighting coefficient needs to be dynamically adjusted. During actual operation, the bandwidth allocation weighting coefficient is increased based on changes in network traffic, and the response time of the entire adjustment process is controlled within <50ms. This rapid dynamic adjustment can adapt to fluctuations in network traffic in a timely manner, ensuring reasonable bandwidth allocation under different network load conditions and guaranteeing the normal operation of critical services such as video conferencing.

[0096] Specific configuration generation: Based on the above symbol rule selection and continuous parameter optimization, the final specific network configuration is generated: {Guaranteed bandwidth for video conferencing: 650Mbps, Redundant paths: 3, QoS level: Gold}. Here, "Guaranteed bandwidth for video conferencing: 650Mbps" specifies the minimum guaranteed bandwidth provided for video conferencing to ensure smooth operation; "Redundant paths: 3" sets up 3 redundant paths, allowing data to be transmitted through these redundant paths when the main path fails or becomes congested, improving network reliability; "QoS level: Gold" assigns a high quality of service level to video conferencing services, guaranteeing their priority transmission within the network.

[0097] (3) Strategy optimization

[0098] After the policy is generated, it needs to be optimized to ensure its best performance in a real network environment. Policy optimization is mainly achieved through two aspects: digital twin verification metrics and causal entropy analysis.

[0099] Digital Twin Verification Metrics: A digital twin system was used to compare and analyze key network performance indicators before and after policy optimization. Regarding latency, the network latency before optimization was greater than after optimization. After optimization, data transmission speeds in the network are faster, providing a smoother user experience. Regarding throughput, optimization significantly improved the network's data transmission capacity, better meeting users' demands for high-bandwidth services. Regarding packet loss rate, optimization reduced data loss and improved data transmission reliability.

[0100] Causal entropy analysis: Causal entropy analysis was performed on the core switch. The causal entropy value reflects the complexity and uncertainty of the switch in the network. When this value is too high, it can lead to problems such as network jitter. By optimizing its ECMP weight allocation, network jitter was effectively reduced. This shows that through causal entropy analysis and targeted optimization measures, network stability and performance can be further improved.

[0101] (4) Model optimization

[0102] First, high-entropy nodes are screened by determining a causal entropy threshold using Monte Carlo simulation. Monte Carlo simulation is a statistical method based on random sampling that simulates various network states through numerous random experiments to determine a suitable threshold. Then, causal entropy is calculated for all nodes in the network, and nodes with causal entropy greater than the threshold are selected. To guide the agent to prioritize optimization of high-entropy nodes, a dynamic reward function is designed, focusing on improving network service quality; a penalty coefficient is used to penalize configuration conflict behavior. For high-entropy nodes, a 3x reward weight is applied, making the agent more inclined to operate on these nodes during optimization. When the agent's actions reduce the causal entropy of high-entropy nodes, improve network QoS, and reduce configuration conflicts, it will receive a higher reward, thus incentivizing the agent to continuously explore optimal optimization strategies and achieve model optimization.

[0103] After comprehensive optimization of the network system, core performance indicators were significantly improved. Network latency was reduced from over 50ms to less than 30ms, throughput increased to 8.2Tbps, an improvement of approximately 37%, and packet loss rate was effectively controlled from 1.2% to within 0.15%. Critical business operations were strengthened, video conferencing bandwidth was stably maintained above 650Mbps, and three redundant paths were established to achieve 99.99% transmission reliability. Service quality met Gold-level standards, with latency jitter less than 5ms.

Claims

1. A large model reinforcement learning network configuration generation method based on verification feedback, characterized in that, Comprise the following steps: S1, semantic parsing based on network configuration of graph neural network: Convert the input unstructured business requirements into a semantic action sequence, build a formal ontology library and a dynamic knowledge graph, and update the network state in real time by combining a time series graph convolution network; S2, a hybrid action space construction strategy generation model for communication network configuration, the strategy generation model includes symbol rule selection, continuous parameter optimization, and generation of specific network configuration, specifically including: Using a hierarchical neural symbolic reinforcement learning method, the updated network state in step S1 is input to cooperatively optimize the symbol rule and the continuous parameter, and generate a specific network configuration containing device parameters and link connections; Calculate the symbol rule selection probability through the probabilistic logic programming formula: ; wherein, represents the probability of selecting a symbol rule under state ; represents the current operational state of the network, represents a specific network configuration rule, represents the total number of rules in the symbol rule set; is a function that measures the score or fitness of a symbol rule under state ; the higher the score, the more likely the rule will be selected under the current state. Through the dynamic weight fusion module, the contribution proportion of the symbol rule and the continuous parameter optimization is automatically adjusted according to the network load fluctuation degree; S3, calculate the causal entropy based on the causal graph: Input the generated specific network configuration into the digital twin system for simulation verification, calculate the node causal entropy based on the causal graph: Each configuration item in the network, including bandwidth allocation and routing settings, is included as a node The mutual influence relationship between configuration items is included as an edge A configuration item causal diagram is constructed Causal analysis is performed, and configuration item dependency relationships are extracted from the verification log; For each configuration item node determine the set of all possible states or events relevant and estimate the probability of each state or event occurring; calculate the causal entropy of the node by the formula: ; wherein represents a node is a random variable representing some state or event related to node is the set of all possible states or events related to node represents the probability of state or event occurring in the context of node ​​​​ S4, strategy generation model optimization: Based on the causal entropy value and the feedback signal, adjust the high-entropy node configuration, adjust the configuration strategy through the reinforcement learning algorithm, and optimize the overall network operation state to finally optimize the configuration to generate a large model; The reward mechanism of reinforcement learning includes: Set the reward function according to the node causal entropy value, and preferentially optimize the configuration adjustment action of the high-entropy node; The reward value is positively correlated with network performance indicators including QoS and load balancing degree, and is negatively correlated with the number of configuration conflicts; Specifically including: First, high-entropy node screening is performed, and the causal entropy threshold is determined by the Monte Carlo simulation method; Then, the causal entropy of all nodes in the network is calculated, and the nodes with a causal entropy greater than the threshold are screened out; In order to guide the agent to preferentially optimize the high-entropy node, a dynamic reward function is designed, which focuses on the improvement of network service quality; The penalty coefficient is used to punish the configuration conflict behavior; For high-entropy nodes, a reward weight of 3 times is applied, so that the agent is more inclined to operate these nodes during optimization; When the agent's action can reduce the causal entropy of the high-entropy node, improve the network QoS, and reduce the configuration conflict, it will obtain a higher reward, thereby encouraging the agent to continuously explore the optimal optimization strategy, and realizing model optimization.

2. The method of claim 1, wherein, In step S1, semantic parsing specifically includes: Build a formal ontology library based on the OWL method, and define the device, protocol and topology entity relationship; Use the time series graph convolution network to update the link load and device performance state in real time: ; In the formula, Indicates at time step t The hidden state matrix of nodes in the time graph. For activation function, For the normalized adjacency matrix, To include node-level features, For learnable graph / node feature transformation matrices, b This is the bias vector.

3. The method of claim 1, wherein, In step S4, model optimization specifically includes: Adjust the configuration strategy through the reinforcement learning algorithm, thereby optimizing the overall network operation state, and finally optimizing the configuration to generate a large model.

4. The method of claim 1, wherein, The verification process of the digital twin system includes: Formal verification of configuration conflict detection and performance simulation of real network scenarios; Generate multi-dimensional feedback signals including delay, throughput and packet loss rate for strategy optimization.

5. The method of claim 1, wherein, The construction of the causal graph also includes: Extract the configuration item dependency relationship from the verification log, and dynamically update the weight of edge E based on historical data and real-time state; By incremental model updating algorithm, the network topology evolution and business mode migration are adapted.

6. The method of claim 1, wherein, In the semantic analysis stage, natural language processing (NLP) is introduced to analyze the ambiguous semantics in unstructured requirements. Through adversarial training, the robustness of the graph neural network is enhanced to avoid analysis deviation caused by noisy data.

Citation Information

Patent Citations

  • Virtual network mapping method based on maximum entropy and multiple agents

    CN118827421A

  • Mathematical twinborn construction method based on large language model and reinforcement learning

    CN118862642A