Network element configuration generation method and device, network equipment, medium and program product
Through structured processing of network element configuration data and diffusion model training, combined with reinforcement learning and reward function optimization, a customized configuration model is generated, which solves the problems of network element configuration errors and complex network management, and achieves efficient and accurate network configuration.
Patent Information
- Application Number
- CN202510686606.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, network element configuration errors lead to network interruption, manual configuration is difficult to meet the needs of complex network architectures, and heterogeneous cloud network management is complex.
By structuring the network element configuration data, custom configuration models are generated using diffusion models and reinforcement learning, and optimization is combined with reward functions that quantify network performance to generate configuration strategies that meet network needs.
It improves the efficiency and accuracy of network configuration, reduces the cost and error risk of manual configuration, and adapts to changes in complex network architectures.
Smart Images

Figure CN120342864A_ABST
Abstract
Description
Background Art
[0002] In the field of network management, generating network element configurations is a key task to ensure the efficient operation of the network. The parameter settings of network devices such as routers and switches are crucial. However, since device suppliers have specific models, the configuration work requires professionals to spend a lot of energy studying user manuals, collecting adaptation commands, verifying configuration templates, and accurately mapping template parameters to the controller database. During this process, even a single incorrect ACL (Access Control List) configuration can directly lead to network outages. In addition, with the rapid development of heterogeneous cloud networks, the scope of network management not only includes traditional network devices but also a large number of computing and storage devices. Therefore, to adapt to complex network architectures and prevent network problems caused by manual configuration errors, a network self-configuration solution is urgently needed.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a method for generating network element configurations, a configuration device, a network device, a storage medium, and a computer program product, which at least to some extent overcome the problems in the related art that configuration errors have a great impact on the network and manual configuration is difficult to meet complex network architectures.
[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0006] According to one aspect of the present disclosure, a method for generating network element configurations is provided, including: structuring configuration data for configuring the operation of multiple network element devices to obtain configuration records in a corresponding format; selecting a diffusion model that matches the corresponding format, and using the configuration records as training data to train the diffusion model to obtain an initial configuration model; performing reinforcement learning on the initial configuration model based on a reward function for quantifying network performance to obtain a network element configuration model; and performing demand adjustment on the network element configuration model based on the obtained configuration demand information to obtain a customized configuration model, and deploying the customized configuration model to the required network.
[0007] In an embodiment of the present disclosure, performing reinforcement learning on the initial configuration model based on a reward function for quantifying network performance to obtain a network element configuration model includes: configuring the reward function; and performing reinforcement learning on the initial configuration model based on the denoising diffusion policy optimization and the reward function to obtain the network element configuration model.
[0008] In one embodiment of the present disclosure, configuring the reward function includes: configuring the throughput, latency, and packet loss rate of the network element device and the carried network link as quantization metrics of the network performance; configuring weights corresponding to the throughput, the latency, and the packet loss rate respectively based on the application scenario of the network element device; and determining the reward function based on the quantization metrics and the corresponding weights.
[0009] In one embodiment of the present disclosure, performing reinforcement learning on the initial configuration model based on the denoising diffusion strategy optimization and the reward function to obtain the network element configuration model includes: generating a configuration action based on the initial network state of the network element device and the initial configuration model, and executing the configuration action on the network element device to modify the configuration parameters of the network element device based on the configuration action; calculating the reward function based on the execution result of the configuration action to obtain a reward value; modifying the network state based on the modification result of the configuration parameters of the network element device to obtain the modified network state and the corresponding state transition information; and performing a cyclic reinforcement learning operation on the initial configuration model based on the modified network state, the reward value, and the state transition information until the network element configuration model is obtained.
[0010] In one embodiment of the present disclosure, generating a configuration action based on the initial network state of the network element device and the initial configuration model includes: configuring the initial network state based on the network environment where the network element device is located; configuring a diffusion model policy generator based on the initial configuration model; generating a configuration policy in the initial network state by the diffusion model policy generator, where the configuration policy is used to determine the parameter adjustment direction for the network element device; and generating the configuration action based on the configuration policy and the current network state.
[0011] In one embodiment of the present disclosure, generating the configuration action based on the configuration policy and the current network state includes: integrating the configuration policy and the state information of the current network state to obtain an integrated feature, extracting the state feature related to the action from the integrated feature, and constructing a constraint condition for action generation; mapping the configuration policy and the constraint condition to the action latent space to obtain a guiding signal for the action denoising process; and performing multi-step denoising operations based on the guiding signal to obtain the configuration action.
[0012] In one embodiment of the present disclosure, performing a cyclic reinforcement learning operation on the initial configuration model based on the modified network state, the reward value, and the state transition information until the network element configuration model is obtained includes:
[0013] In a reinforcement learning operation cycle, if the loop reinforcement learning operation does not meet the constraint conditions, determine the modified network state as the current network state; extract sampling data from the state transition information to calculate the policy gradient of the initial configuration model based on the sampling data; based on obtaining the policy gradient, adjust the model parameters of the initial configuration model to update the configuration diffusion model policy generator based on the adjusted initial configuration model to update the configuration action, and repeat the reinforcement learning operation cycle until the obtained reward value reaches the reward threshold or the number of reinforcement learning operation cycles reaches the number threshold, and determine the adjusted initial configuration model as the network element configuration model.
[0014] In an embodiment of the present disclosure, performing requirement adjustment on the network element configuration model based on the obtained configuration requirement information to obtain a customized configuration model, including: identifying the requirement type of the requirement information; generating initial configuration information matching the requirement type based on a text prompt guiding model; adjusting the weight of the reward function to match the requirement type; receiving feedback information of the network element device on the initial configuration information; adjusting the model parameters of the network element configuration model based on the feedback information and the adjusted reward function to obtain the customized configuration model.
[0015] In an embodiment of the present disclosure, structuring configuration data for configuring the operation of multiple network element devices to obtain a configuration record in a corresponding format, including: collecting multi-source configuration data from different network element devices; performing normalization processing on continuous data in the configuration data and performing enumeration processing on discrete data in the configuration data; adopting a list-type parameter type, using row identifiers to configure samples, and constructing the configuration record in a table format based on the processed configuration data.
[0016] In an embodiment of the present disclosure, selecting a diffusion model matching the corresponding format, and using the configuration record as training data to perform model training on the diffusion model to obtain an initial configuration model, including: determining a diffusion model applicable to the table format; performing model training on the diffusion model based on the configuration record to enable the diffusion model to learn the dependency relationship and distribution characteristics between the configuration data, so that the trained initial configuration model outputs configuration samples.
[0017] According to another aspect of the present disclosure, there is provided an apparatus for generating network element configurations, including: a processing module configured to perform structured processing on configuration data for configuring the operation of multiple network element devices to obtain configuration records in a corresponding format; a training module configured to select a diffusion model that matches the corresponding format, and use the configuration records as training data to perform model training on the diffusion model to obtain an initial configuration model; a reinforcement learning module configured to perform reinforcement learning on the initial configuration model based on a reward function for quantifying network performance to obtain a network element configuration model; and an adjustment module configured to perform requirement adjustment on the network element configuration model based on the obtained configuration requirement information to obtain a customized configuration model, and deploy the customized configuration model to a required network.
[0018] According to still another aspect of the present disclosure, there is provided a network device, including: a processor; and a memory for storing executable instructions of the processor; the processor is configured to execute the method for generating network element configurations in the first aspect above by executing the executable instructions.
[0019] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for generating network element configurations above is implemented.
[0020] According to yet another aspect of the present disclosure, there is provided a computer program product having a computer program stored thereon, and when the computer program is executed by a processor, the method for generating network element configurations above is implemented.
[0021] The solution for generating network element configurations provided by the embodiments of the present disclosure performs structured processing on the configuration data of multiple network element devices to provide a data basis for model training, selects a diffusion model with a matching data format for training, which is conducive to effectively learning the internal rules of the configuration data and ensuring the rationality of the generated configuration. Combining the reinforcement learning process with the reward function, with the optimization of network performance as the goal, to improve the practicality of the model generation strategy. Obtaining a customized configuration model based on requirement adjustment can realize personalized network configuration requirements, thereby realizing the full-process optimization of network element configuration from data processing, model training to personalized generation, which is conducive to improving the efficiency, accuracy of network configuration and the adaptability to complex network architectures, and reducing the manual configuration cost and error risk.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 A flowchart showing a method for generating a network element configuration in an embodiment of the present disclosure;
[0025] Figure 2 A flowchart showing another method for generating a network element configuration in an embodiment of the present disclosure;
[0026] Figure 3 A flowchart showing yet another method for generating a network element configuration in an embodiment of the present disclosure;
[0027] Figure 4 A flowchart showing yet another method for generating a network element configuration in an embodiment of the present disclosure;
[0028] Figure 5 A flowchart showing yet another method for generating a network element configuration in an embodiment of the present disclosure;
[0029] Figure 6 A schematic diagram showing a device for generating a network element configuration in an embodiment of the present disclosure;
[0030] Figure 7 A block diagram showing the structure of a computer device in an embodiment of the present disclosure. Detailed implementation manners
[0031] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.
[0032] In addition, the accompanying drawings are only schematic diagrams of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0033] Network element configuration generation is a key task in network management, which involves setting parameters for devices such as routers and switches to ensure the efficient operation of the network. Due to vendor-specific device models, a large amount of expert effort is required to study user manuals, collect appropriate commands, verify configuration templates, and map template parameters to the controller database. During this process, even a single ACL configuration error can cause network outages. Additionally, considering the growing heterogeneous cloud network, which also requires managing a large number of computing and storage devices, a unified natural language configuration interface is crucial for simplifying the configuration process and enabling self-configuring networks.
[0034] In this disclosure, a network element configuration scheme can be obtained by combining a diffusion model and reinforcement learning. Among them, the diffusion model is a generative model that learns the data distribution by gradually adding and removing noise. Reinforcement learning is a machine learning method that learns decisions by interacting with the environment and is suitable for optimizing dynamic systems. Combining the two can provide an innovative solution for automated network element configuration generation. Specifically, first, data preparation is carried out, that is, an effective network element configuration data set is collected to ensure that the data covers different device types and network scenarios. Secondly, diffusion model training is performed, that is, the diffusion model is used to learn the configuration distribution and generate new configurations. Further, reinforcement learning optimization is carried out. A reward function is defined based on network performance metrics (such as throughput and latency), and the configuration is evaluated using a network simulator or a real environment. Finally, configuration generation and application are performed. The optimized diffusion model is used to generate new configurations and deploy them to the network.
[0035] For ease of understanding, several terms involved in this application are first explained below.
[0036] Diffusion model: In machine learning, a diffusion model or diffusion probability model is a class of latent variable models and is a Markov chain trained using variational estimation. The goal of the diffusion model is to learn the latent structure of a data set by modeling the way data points diffuse in the latent space.
[0037] Reinforcement learning (abbreviated as RL in English: Reinforcement learning) is a field in machine learning that emphasizes how to act based on the environment to obtain the maximum expected benefit.
[0038] Next, each step of the method for generating network element configuration in this exemplary embodiment will be described in more detail in conjunction with the accompanying drawings and embodiments.
[0039] Figure 1 The flowchart of a method for generating network element configuration in an embodiment of the present disclosure is shown.
[0040] As Figure 1 shown, according to an embodiment of the present disclosure, the method for generating network element configuration includes:
[0041] Step S102: Structurally process the configuration data for configuring the operation of multiple network elements to obtain configuration records in the corresponding format.
[0042] In some embodiments, the configuration data refers to the original configuration information collected from various network elements such as routers, switches, and firewalls, including device parameters, interface settings, routing protocols, etc. Additionally, the configuration data can cover different network scenarios such as enterprise networks and data center networks, as well as various device roles such as edge devices and core devices.
[0043] In some embodiments, the structural processing is to remove invalid, incorrect, or duplicate data through data cleaning, normalize continuous parameters, encode discrete parameters, and convert the unstructured or semi-structured original data into a two-dimensional table form, making the data have structured feature dimensions for subsequent model processing.
[0044] Step S104: Select a diffusion model that matches the corresponding format, and use the configuration record as training data to train the diffusion model to obtain an initial configuration model.
[0045] In some embodiments, the diffusion model that matches the corresponding format can be TabDDPM applicable to tabular data. This model can handle the data structure of a mixture of continuous and discrete parameters and learn the distribution law of the data through the forward diffusion and reverse denoising mechanisms.
[0046] In some embodiments, the initial configuration model is a model obtained by training the diffusion model with the structured configuration record. Since the parameter dependence relationship and distribution characteristics of the network element configuration data have been initially mastered, it can generate network element configuration strategies that are syntactically and semantically reasonable from random noise.
[0047] Step S106: Perform reinforcement learning on the initial configuration model based on the reward function that quantifies network performance to obtain a network element configuration model.
[0048] In some embodiments, the reward function is a function constructed based on network performance metrics such as throughput, latency, and packet loss rate. By setting weights for different metrics, it quantitatively evaluates the impact of the configuration strategy on network performance.
[0049] In some embodiments, the reinforcement learning can be oriented by the reward function, allowing the initial configuration model to continuously try different configuration strategies in a simulated or actual network environment and adjust the model parameters according to the reward feedback.
[0050] In some embodiments, the network element configuration model is a model optimized through reinforcement learning. Compared with the initial configuration model, the network element configuration model better meets the network performance requirements and obtains more refined network element configuration parameters.
[0051] Step S108, adjust the requirements of the network element configuration model based on the obtained configuration requirement information to obtain a customized configuration model, so as to deploy the customized configuration model to the required network.
[0052] In some embodiments, requirement adjustment means that according to specific requirements, such as "low latency priority", "applicable to 5G base stations", etc., through means such as conditional injection, parameter fine-tuning, and model structure adjustment, the generation process of the network element configuration model is inclined towards meeting specific requirements.
[0053] In some embodiments, the customized configuration model is a model adjusted for specific requirements. It can output network element configurations that meet specific network scenarios, device types, or performance requirements, realizing personalized customization of configurations.
[0054] In this embodiment, by structuring the configuration data of multiple network element devices, a data basis is provided for model training. A diffusion model with a matching data format is selected for training, which is beneficial to effectively learning the internal laws of the configuration data, ensuring the rationality of the generated configurations. Combining the reinforcement learning process with the reward function, with the goal of optimizing network performance, to improve the practicality of the model generation strategy. Based on the requirement adjustment to obtain the customized configuration model, personalized network configuration requirements can be realized, thus realizing the full-process optimization of network element configuration from data processing, model training to personalized generation, which is conducive to improving the efficiency, accuracy of network configuration and the adaptability to complex network architectures, and reducing the manual configuration cost and error risk.
[0055] In an embodiment of the present disclosure, performing reinforcement learning on the initial configuration model based on a reward function for quantifying network performance to obtain a network element configuration model, including:
[0056] Configuring a reward function; performing reinforcement learning on the initial configuration model based on the denoising diffusion strategy optimization and the reward function to obtain a network element configuration model.
[0057] In an embodiment of the present disclosure, configuring the reward function includes: configuring the throughput, latency, and packet loss rate of the network element device and the network link it carries as quantization metrics of network performance; respectively configuring the weights corresponding to the throughput, latency, and packet loss rate based on the application scenarios of the network element device; determining the reward function based on the quantization metrics and the corresponding weights.
[0058] In some embodiments, configuring the reward function is to convert network performance metrics (such as throughput, latency, packet loss rate, etc.) into quantifiable numerical feedback. By assigning weight coefficients to different metrics (such as increasing the weight of the latency metric in a low-latency scenario), a reward function in the form of Equation (1) is constructed to evaluate the pros and cons of the configuration strategy.
[0059] R = ω1 * throughput + ω2 * latency + ω3 * packet loss rate (1)
[0060] In some embodiments, Denoising Diffusion Policy Optimization (DDPO) refers to combining the generation ability of a diffusion model with the optimization mechanism of reinforcement learning. During the reverse denoising process of the diffusion model, instead of solely relying on the training data distribution, the gradient information of the reward function is introduced, and the model parameters are adjusted by maximizing the cumulative reward to guide the model to generate a better configuration strategy.
[0061] In this embodiment, by configuring the reward function, the network performance requirements are transformed into clear optimization goals, closely associating the model training direction with the actual network operation effect. Combining with the reinforcement learning mechanism based on DDPO, a reward orientation is incorporated into the denoising process of the diffusion model. This not only retains the powerful modeling ability of the diffusion model for complex configuration distributions but also endows the model with the flexibility to dynamically optimize the strategy according to real-time feedback, enabling the finally obtained network element configuration model to generate configuration strategies that better meet the actual network requirements and significantly improve network performance.
[0062] As Figure 2 shown, in an embodiment of the present disclosure, reinforcement learning is performed on an initial configuration model based on denoising diffusion policy optimization and a reward function to obtain a network element configuration model, including:
[0063] Step S202: Generate a configuration action based on the initial network state and the initial configuration model of the network element device, and execute the configuration action on the network element device to modify the configuration parameters of the network element device based on the configuration action.
[0064] In some embodiments, the initial network state includes information such as the current configuration parameters, network topology structure, and performance metrics of the network element device, which is a current snapshot of the network operation. After inputting the initial network state into the initial configuration model, the model generates a configuration action through the reverse denoising process, combining the current network state based on the learned configuration rules. For example, adjusting the routing protocol parameters of a router, modifying the port rate of a switch, etc. The configuration action refers to a specific operation to optimize the network state. By executing these actions on the network element device, the configuration parameters of the device are directly modified, and thus the operation state of the network can be changed.
[0065] Step S204: Calculate the reward function based on the execution result of the configuration action to obtain a reward value.
[0066] In some embodiments, when the configuration action is executed on the network element device, an actual network operation result will be generated, such as whether the network throughput is improved, whether the latency is reduced, etc. Substitute the result into the reward function, and by calculating the sum of the weighted values of each performance metric, the reward value is obtained. Among them, a positive reward value indicates that the configuration action has a positive impact on network performance, and a negative reward value indicates a negative impact, thereby quantitatively evaluating the pros and cons of the configuration action.
[0067] Step S206: Modify the network state based on the modification result of the configuration parameters of the network element device to obtain the modified network state and the corresponding state transition information.
[0068] In some embodiments, if the configuration action modifies the configuration parameters of the device, the operating state of the network will also change accordingly. For example, modifying the routing table parameters of a router will affect the forwarding path of data packets, thereby changing the traffic distribution and latency of the network.
[0069] Integrate the modified network parameters and related performance metrics to obtain the modified network state. Additionally, record the change process and related information from the network state before the execution of the configuration action to the modified network state to form the state transition information.
[0070] In some embodiments, the state transition information is the core basis for updating the model parameters of the reinforcement learning algorithm (such as diffusion policy optimization). From the current state S t to the next state S t+1 , the state transition information includes:
[0071] The current state S t , that is, the network state before the execution of the action, the configuration action A t , the reward value r t and the next state S t+1 , that is, the network state after the execution of the action, such as the modified configuration parameters, the new traffic distribution, etc.
[0072] Step S208: Perform cyclic reinforcement learning operations on the initial configuration model based on the modified network state, the reward value, and the state transition information until the network element configuration model is obtained.
[0073] In some embodiments, using the modified network state as the new input, combined with the reward value and the state transition information, calculate the update direction and amplitude of the model parameters through the reinforcement learning algorithm (such as the policy gradient method). The reward value reflects the quality of the configuration action. If the reward value is high, increase the generation probability of the corresponding configuration policy. If the reward value is low, decrease the generation probability of the corresponding policy. The state transition information records the complete process of the network state change, helping the model understand the causal relationship between different configuration actions and network state changes. By continuously repeating this process, that is, continuously inputting the new network state into the model, calculating the reward value and the state transition information, and updating the model parameters, cyclic reinforcement learning operations are performed.
[0074] In this embodiment, by generating and executing configuration actions based on the initial network state and the initial configuration model, and using the reward function to quantitatively evaluate the execution results of the configuration actions, the optimization goal of network performance becomes more clear. Record the state transition information and perform cyclic reinforcement learning based on this, enabling the model to continuously learn and optimize from the feedback of the actual network operation, gradually adapt to the complex and changing network environment, enabling the network element configuration model to continuously evolve, generate configuration strategies that more closely match the actual network requirements, effectively improve the network throughput, reduce latency, and reduce the packet loss rate, and achieve the efficient utilization and intelligent management of network resources.
[0075] As Figure 3 shown, in one embodiment of the present disclosure, generating a configuration action based on the initial network state and the initial configuration model of a network element device includes:
[0076] Step S302, configuring the initial network state based on the network environment where the network element device is located.
[0077] In some embodiments, the initial network state can be understood as a benchmark representation of the network environment, and needs to be constructed by comprehensively considering the network element device type (such as routers, switches), network scenarios (such as enterprise networks, data centers), topological structures, historical configuration data, and typical performance indicators.
[0078] Step S304, configuring the diffusion model policy generator based on the initial configuration model.
[0079] In some embodiments, since the initial configuration model has learned the distribution and dependency relationships of the configuration parameters, by loading the model parameters of this model, the policy generator inherits the modeling ability of the initial configuration model for data distribution and is adapted to the task of "generating configuration policies".
[0080] Step S306, generating a configuration policy in the initial network state by the diffusion model policy generator, and the configuration policy is used to determine the parameter adjustment direction for the network element device.
[0081] In some embodiments, after inputting the initial network state into the diffusion model policy generator, the model can gradually recover a reasonable configuration policy through the reverse denoising process.
[0082] Step S308, generating a configuration action based on the current network state by the configuration policy.
[0083] In some embodiments, the current network state is dynamic data collected in real time (such as current load, real-time topology). As an abstract adjustment rule, the configuration policy needs to be combined with the current state to generate specific actions. For example, if the policy is "when the load increases by 10%, the bandwidth increases by 5 MHz", and the current load has increased compared to the initial state, then a specific action of "increasing the bandwidth by 15 MHz" is generated. By multiplying or mapping the generalization rule of the policy with the specific values of the real-time state, the transformation from "direction" to "operation" is achieved.
[0084] In this embodiment, the policy generator configuration based on the initial configuration model reuses the modeling ability of the diffusion model for complex data distributions to quickly construct a policy generation framework. The abstract policy generation enables the model to output a generalized parameter adjustment direction, and the combination of the policy and the current state to generate actions endows the system with dynamic response capabilities, which not only ensures the rationality and generalization of the configuration policy but also improves the pertinence and timeliness of the configuration actions, ultimately facilitating the improvement of the automation level and optimization effect of network element configuration and reducing the manual intervention cost.
[0085] As Figure 4 shown, in an embodiment of the present disclosure, generating a configuration action based on the current network state by a configuration policy includes:
[0086] Step S402, integrating the configuration policy and the state information of the current network state to obtain an integrated feature.
[0087] In some embodiments, the configuration policy includes the direction and amplitude of parameter adjustment, and the current network state includes real-time parameter values. By operations such as splicing and weighting, the feature vectors of the two are combined to form an integrated feature including the adjustment intention and the real-time state.
[0088] Step S404, extracting the state features related to the action from the integrated feature to construct the constraint conditions for action generation.
[0089] In some embodiments, key information directly affecting the feasibility of the action is screened out from the integrated feature, such as device type (determining the supported configuration commands), interface type (limiting the bandwidth adjustment range), current load (triggering adjustment strategies with different priorities), etc., and is transformed into constraint conditions.
[0090] Step S406, mapping the configuration policy and the constraint conditions to the action latent space to obtain the guiding signal for the action denoising process.
[0091] In some embodiments, the abstract adjustment direction and constraint conditions of the configuration policy are converted into a low-dimensional guidance signal (such as a conditional embedding in vector form) through an encoder network (such as an MLP). During the reverse denoising process of the diffusion model for action generation, this signal serves as an additional input to guide the model to generate actions that conform to the policy direction and satisfy the constraints. For example, in each denoising step, the action parameters are forced to converge towards the direction of the guidance signal through a loss function.
[0092] Step S408, perform multi-step denoising operations based on the guidance signal to obtain the configuration action.
[0093] In some embodiments, in the diffusion model for action generation, starting from random noise, through T-step reverse denoising iterations, the guidance signal is gradually incorporated into the optimization process of the noise vector. In each step, the denoising direction is adjusted according to the guidance signal. For example, under the guidance of "low latency first", the noise components of the queue scheduling parameters are preferentially adjusted until a specific action that conforms to the policy and constraints is generated. Multi-step denoising ensures the rationality and fineness of the action through progressive optimization.
[0094] In this embodiment, by integrating the configuration policy with the state information of the current network state, state features related to actions are extracted to construct constraint conditions, and then the configuration policy and constraint conditions are mapped to the guidance signal for action denoising. Finally, multi-step denoising is performed based on the guidance signal to obtain the configuration action, realizing the transformation from the configuration policy to specific executable actions. With the help of the guidance signal and the multi-step denoising mechanism, the parameter details are gradually optimized during the action generation process, enabling the configuration action to not only meet the performance optimization requirements but also adapt to the dynamic changes of the network environment, effectively improving the accuracy of the network element configuration action, reducing the complexity and error risk of manual configuration, and enhancing the network's adaptive ability to real-time loads and different scenarios.
[0095] In an embodiment of the present disclosure, perform loop reinforcement learning operations on the initial configuration model based on the modified network state, reward value, and state transition information until the network element configuration model is obtained, including:
[0096] In a reinforcement learning operation cycle, if the loop reinforcement learning operation does not meet the constraint conditions, the modified network state is determined as the current network state.
[0097] In some embodiments, in each reinforcement learning cycle, if the preset constraint conditions are not met (such as the improvement of the reward value not reaching the expectation), the modified network state (the new state after performing the action) is redefined as the "current network state" and used as the input for the next round of iteration, forming a closed loop for state update.
[0098] Extract sampling data from the state transition information to calculate the policy gradient of the initial configuration model based on the sampling data.
[0099] In some embodiments, valid samples (such as configuration action sequences with high reward values) are extracted from the state transition information (recording the complete trajectory of state-action-reward), the gradient direction of the model parameters is calculated through the policy gradient algorithm, and the contribution degree of different configuration policies to the reward value is quantitatively evaluated.
[0100] Based on the obtained policy gradient, the model parameters of the initial configuration model are adjusted to update the configuration diffusion model policy generator based on the adjusted initial configuration model to update the configuration action.
[0101] In some embodiments, according to the calculated policy gradient, the parameters of the initial configuration model are adjusted to increase the probability that the model generates high-reward-value policies, and the diffusion model policy generator is updated synchronously to ensure that the output configuration policy is consistent with the optimized model.
[0102] Repeat the reinforcement learning operation cycle until the obtained reward value reaches the reward threshold or the number of reinforcement learning operation cycles reaches the number threshold, and determine the adjusted initial configuration model as the network element configuration model.
[0103] In some embodiments, repeat the above cycle until the reward value reaches the preset threshold (such as a 30% reduction in latency) or reaches the maximum number of iterations, indicating that the model has fully learned the policy for optimizing network performance at this time and is solidified as the network element configuration model.
[0104] In this embodiment, through cyclic reinforcement learning operations, in each cycle, according to the modified network state, reward value, and state transition information, it is judged whether the constraint conditions are met to update the current network state, sampling data is extracted to calculate the policy gradient, and then the initial configuration model parameters are adjusted and the configuration action is updated until the reward value reaches the threshold or the number of cycles reaches the upper limit, and then the network element configuration model is determined, enabling the model to continuously learn from the actual feedback of network operation, adjust the parameters through the policy gradient, and gradually improve the quality of the generated configuration policy, which is beneficial to enhancing the adaptability of the network element configuration model to complex dynamic network environments.
[0105] In an embodiment of the present disclosure, demand adjustment is performed on the network element configuration model based on the obtained configuration demand information to obtain a customized configuration model, including:
[0106] Identifying the demand type of the demand information; guiding the model to generate initial configuration information matching the demand type based on text prompts; adjusting the weight of the reward function to match the demand type; receiving feedback information of the network element device on the initial configuration information; adjusting the model parameters of the network element configuration model based on the feedback information and the adjusted reward function to obtain a customized configuration model.
[0107] In some embodiments, by parsing keywords in the requirement text (such as "low latency", "high reliability", and "cost optimization", etc.) and classifying them in combination with predefined requirement type tags (such as performance optimization, security enhancement, resource conservation), the unstructured requirements are transformed into actionable type identifiers.
[0108] In some embodiments, based on the identified requirement type, a specific text prompt (such as "Generate router configurations suitable for low latency scenarios") is constructed as a conditional input to the diffusion model, and using the context learning ability of the model, it is guided to generate initial configuration information that meets the requirement type (such as preferentially selecting the OSPF protocol, adjusting queue scheduling parameters, etc.).
[0109] In some embodiments, the weights of each performance metric in the reward function are dynamically adjusted according to the requirement type. For example, for "low latency" requirements, the weight of the latency metric can be increased and the throughput weight can be decreased to make the reward function more focused on the target requirements.
[0110] In some embodiments, after deploying the initial configuration information to the network element device, actual operation data (such as latency values, packet loss rates) are collected, compared with the expected target to generate feedback information, and the feedback information is input into the reinforcement learning framework. Combining with the adjusted reward function, the policy gradient is calculated to update the parameters of the network element configuration model, so that the model is more biased towards configuration policies that meet specific requirements in subsequent generations, forming a customized configuration model.
[0111] In this embodiment, according to specific requirement input conditions, such as "low latency first" or "suitable for edge routers", the model adjusts the generation process according to the conditions and outputs customized configurations. The generated configurations not only meet the syntax requirements of network devices, but also improve the intelligence level of network configuration, enabling the system to automatically understand business requirements and generate optimized configuration policies, thereby shortening the deployment cycle.
[0112] As Figure 5 shown, a method for generating a network element configuration according to another embodiment of the present disclosure includes:
[0113] Step S502, initialize the configuration environment state.
[0114] Step S504, configure the diffusion model policy generator based on the configuration model, and the diffusion model policy generator generates configuration policies.
[0115] Step S506, input the current state and generate configuration actions through multi-step denoising.
[0116] Step S508, execute the generated configuration actions, and the network environment returns a reward value and a new state.
[0117] Step S510, detect whether the end condition is met. If the detection result is "yes", go to step S518; if the detection result is "no", take the new state as the current state and go to step S512.
[0118] Step S512, store the state transition information between the old and new states as experience in the buffer.
[0119] Step S514, extract sampling samples from the buffer batch.
[0120] Step S516, calculate the policy gradient based on the sampling samples, update the model parameters of the configuration model based on the policy gradient, and return to step S504.
[0121] Step S518, generate a policy configuration model.
[0122] In this embodiment, the method for generating network element configuration by combining a diffusion model and reinforcement learning reduces the configuration time, reduces the risk of human errors, improves the consistency and accuracy of device configuration, and ensures the stability and reliability of the existing network for network devices.
[0123] In an embodiment of the present disclosure, the configuration data for configuring the operation of multiple network element devices is structured to obtain configuration records in a corresponding format, including:
[0124] Collect multi-source configuration data from different network element devices; perform normalization processing on continuous data in the configuration data and enumeration processing on discrete data in the configuration data; adopt a list-type parameter type, use row identifiers for configuration samples, and construct a table-format configuration record based on the processed configuration data.
[0125] In some embodiments, collect a configuration dataset of valid network elements (devices such as routers and switches), ensure that the data covers different device types and network scenarios, extract configuration records from existing network devices (such as routers and switches), or obtain historical data using a network management database. The data covers multiple network types (such as enterprise networks and data center networks) and device scenarios (such as edge devices and core routers) to enhance the generalization ability of the model. Finally, output a structured configuration dataset as the training basis for the diffusion model.
[0126] In this embodiment, by collecting multi-source configuration data from different network elements, normalizing continuous data and enumerating discrete data, and constructing a configuration record in tabular format with a list-type parameter type and row identifier configuration samples, the standardization and integration of heterogeneous configuration data are achieved. The multi-source data collection ensures coverage of different device types and network scenarios, enhancing data diversity. Among them, the normalization of continuous data and the enumeration of discrete data eliminate unit and format differences, enabling the data to have a unified quantization standard. The tabular construction endows the data with a clear row and column structure, facilitating subsequent processing such as diffusion models, thus providing a high-quality structured data basis for the training of network element configuration models.
[0127] In one embodiment of the present disclosure, a diffusion model matching the corresponding format is selected to perform model training on the diffusion model using the configuration record as training data, obtaining an initial configuration model, including:
[0128] Determine a diffusion model applicable to the tabular format; perform model training on the diffusion model based on the configuration record, so that the diffusion model learns the dependency relationships and distribution characteristics among the configuration data, so that the trained initial configuration model outputs configuration samples.
[0129] In some embodiments, for the mixed characteristics of tabular format data (coexistence of continuous / discrete parameters and complex dependency relationships among parameters), a diffusion model specifically designed for tabular data, such as TabDDPM, is selected. This type of model gradually adds noise to the original data through the forward diffusion process, converting it into pure noise, and then recovers the data from the noise through the reverse denoising process, and can effectively capture the joint distribution and dependency structure among different parameters in the tabular data.
[0130] In some embodiments, the tabular format configuration record is divided into a training set, a validation set, and a test set by row. Each row contains features such as device type, interface parameters, and protocol configuration. The training process for each training sample includes a diffusion process and a denoising process. The diffusion process includes converting it into noise data through a T-step noise addition process, and the model learns the distribution characteristics of the configuration data. The denoising process includes training a denoising network and updating network parameters.
[0131] In some embodiments, through a shared network layer and an attention mechanism, the model learns the dependency relationships among different parameters. After training, configuration samples are generated from random noise through the reverse denoising process, and the generated configuration samples are subjected to legality verification and formatting, and converted into executable configuration commands.
[0132] In this embodiment, a diffusion model is used to learn the configuration distribution and generate new configurations. The diffusion model is utilized to capture the dependencies and distribution characteristics among configuration parameters, generating configuration samples with reasonable structures and valid contents. The diffusion model TabDDPM (Tabular Denoising Diffusion Probabilistic Model) applicable to tabular data is adopted. TabDDPM can effectively handle mixed data of continuous and discrete parameters, being suitable for the complexity of network element configurations. The input of the model is configuration data in tabular form, and the output is the generated configuration samples. After training is completed, the diffusion model can generate new network element configurations from random noise, and these configurations are similar to historical valid configurations in terms of syntax and semantics, having high feasibility.
[0133] It should be noted that the above-mentioned drawings are merely schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0134] Next, refer to Figure 6 to describe the network element configuration generation device 600 according to the embodiments of the present disclosure. Figure 6 The shown network element configuration generation device 600 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0135] The network element configuration generation device 600 is presented in the form of a hardware module. The components of the network element configuration generation device 600 may include, but are not limited to: a processing module 602, which is used to structurally process the configuration data for configuring the operation of multiple network element devices to obtain configuration records in corresponding formats; a training module 604, which is used to select a diffusion model that matches the corresponding format and use the configuration records as training data to train the diffusion model to obtain an initial configuration model; a reinforcement learning module 606, which is used to perform reinforcement learning on the initial configuration model based on a reward function for quantifying network performance to obtain a network element configuration model; and an adjustment module 608, which is used to perform demand adjustment on the network element configuration model based on the obtained configuration requirement information to obtain a customized configuration model, and then deploy the customized configuration model to the required network.
[0136] Those skilled in the art to which the present disclosure pertains can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to herein as "circuit", "module", or "system".
[0137] The following will refer to Figure 7 to describe the electronic device 700 according to this embodiment of the present disclosure. It may be a network device or a terminal. Figure 7 The illustrated electronic device 700 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0138] As Figure 7 shown, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: the at least one processing unit 710 described above, the at least one storage unit 720 described above, and a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710).
[0139] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification. For example, the processing unit 710 may execute as Figure 1 described in the solution.
[0140] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 7201 and / or a cache 7202, and may further include a read-only storage unit (ROM) 7203.
[0141] The storage unit 720 may further include a program / utility 7204 having a set (at least one) of program modules 7205. Such program modules 7205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0142] The bus 730 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0143] The electronic device 700 can also communicate with one or more external devices 770 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device (such as a router, a modem, etc.) that enables the electronic device 700 to communicate with one or more other computing devices. Such communication can be carried out through the input / output (I / O) interface 750. Moreover, the electronic device 700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 760. As shown in the figure, the network adapter 760 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0144] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0145] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to enable the electronic device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0146] The program product for implementing the above method according to the embodiments of the present disclosure can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on an electronic device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that includes or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device.
[0147] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0148] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0149] The program code included on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.
[0150] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0151] It should be noted that although several modules or units of the devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.
[0152] In addition, although the various steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in that specific order, or that all of the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0153] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of the present disclosure.
[0154] After considering the specification and practicing the disclosure herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A method for generating network element configuration, characterized in that, Including: Structurally process the configuration data used to configure the operation of multiple network elements to obtain configuration records in a corresponding format; Select a diffusion model that matches the corresponding format, and use the configuration records as training data to train the diffusion model to obtain an initial configuration model; Perform reinforcement learning on the initial configuration model based on a reward function that quantifies network performance to obtain a network element configuration model; Perform demand adjustment on the network element configuration model based on the obtained configuration requirement information to obtain a customized configuration model, and deploy the customized configuration model to the required network.
2. The method for generating the network element configuration according to claim 1, wherein Performing reinforcement learning on the initial configuration model based on a reward function that quantifies network performance to obtain a network element configuration model, including: Configure the reward function; Perform reinforcement learning on the initial configuration model based on denoising diffusion policy optimization and the reward function to obtain the network element configuration model.
3. The method for generating the network element configuration according to claim 2, wherein Configuring the reward function includes: Configure the throughput, latency, and packet loss rate of the network element device and the network link it carries as quantization metrics for the network performance; Configure the weights corresponding to the throughput, the latency, and the packet loss rate respectively based on the application scenario of the network element device; Determine the reward function based on the quantization metrics and the corresponding weights.
4. The method for generating the network element configuration according to claim 2, wherein Performing reinforcement learning on the initial configuration model based on denoising diffusion policy optimization and the reward function to obtain the network element configuration model, including: Generate a configuration action based on the initial network state of the network element device and the initial configuration model, and execute the configuration action on the network element device to modify the configuration parameters of the network element device based on the configuration action; Calculate the reward function based on the execution result of the configuration action to obtain a reward value; Modify the network state based on the modification result of the configuration parameters of the network element device to obtain a modified network state and corresponding state transition information; Perform a cyclic reinforcement learning operation on the initial configuration model based on the modified network state, the reward value, and the state transition information until the network element configuration model is obtained.
5. The method for generating the network element configuration according to claim 4, wherein Generating a configuration action based on the initial network state of the network element device and the initial configuration model includes: Configure the initial network state based on the network environment where the network element device is located; Configure a diffusion model policy generator based on the initial configuration model; Generate a configuration policy in the initial network state by the diffusion model policy generator, and the configuration policy is used to determine the parameter adjustment direction of the network element device; Generate the configuration action based on the current network state by the configuration policy.
6. The method for generating the network element configuration according to claim 5, wherein Generating the configuration action based on the current network state by the configuration policy includes: Integrate the configuration policy and the state information of the current network state to obtain an integrated feature; Extract the state features related to the action in the integrated feature to construct the constraint conditions for action generation; Map the configuration policy and the constraint conditions to the action latent space to obtain a guiding signal for the action denoising process; Perform multiple-step denoising operations based on the guiding signal to obtain the configuration action.
7. The method for generating the network element configuration according to claim 5, wherein Perform cyclic reinforcement learning operations on the initial configuration model based on the modified network state, the reward value, and the state transition information until the network element configuration model is obtained, including: In a reinforcement learning operation cycle, if the cyclic reinforcement learning operation does not meet the constraint conditions, determine the modified network state as the current network state; Extract sampling data from the state transition information to calculate the policy gradient of the initial configuration model based on the sampling data; Based on obtaining the policy gradient, adjust the model parameters of the initial configuration model to update the configuration diffusion model policy generator based on the adjusted initial configuration model and update the configuration action; Repeat the reinforcement learning operation cycle until the obtained reward value reaches the reward threshold or the number of reinforcement learning operation cycles reaches the number threshold, and determine the adjusted initial configuration model as the network element configuration model.
8. The method for generating the network element configuration according to claim 1, wherein Perform requirement adjustment on the network element configuration model based on the obtained configuration requirement information to obtain a customized configuration model, including: Identify the requirement type of the requirement information; Generate initial configuration information matching the requirement type based on the text prompt guidance model; Adjust the weight of the reward function to match the requirement type; Receive the feedback information of the network element device on the initial configuration information; Adjust the model parameters of the network element configuration model based on the feedback information and the adjusted reward function to obtain the customized configuration model.
9. The method for generating the network element configuration according to claim 1, wherein Structurally process the configuration data for configuring the operation of multiple network element devices to obtain a configuration record in a corresponding format, including: Collect multi-source configuration data from different network element devices; Perform normalization processing on the continuous data in the configuration data and perform enumeration processing on the discrete data in the configuration data; Adopt a list-type parameter type, use row identifiers to configure samples, and construct the configuration record in a table format based on the processed configuration data.
10. The method for generating the network element configuration according to claim 9, wherein Select a diffusion model matching the corresponding format to perform model training on the diffusion model using the configuration record as training data to obtain an initial configuration model, including: Determine the diffusion model applicable to the table format; Perform model training on the diffusion model based on the configuration record so that the diffusion model learns the dependency relationship and distribution characteristics between the configuration data, so that the trained initial configuration model outputs configuration samples.
11. A generating device for network element configuration, characterized in that, Include: A processing module for structurally processing the configuration data for configuring the operation of multiple network element devices to obtain a configuration record in a corresponding format; A training module for selecting a diffusion model matching the corresponding format to perform model training on the diffusion model using the configuration record as training data to obtain an initial configuration model; A reinforcement learning module for performing reinforcement learning on the initial configuration model based on a reward function for quantifying network performance to obtain a network element configuration model; An adjustment module for performing requirement adjustment on the network element configuration model based on the obtained configuration requirement information to obtain a customized configuration model to deploy the customized configuration model to the required network.
12. A network device, characterized in that, Include: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the generation method of the network element configuration according to any one of claims 1 to 10 by executing the executable instructions.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the generation method of the network element configuration according to any one of claims 1 to 10.
14. A computer program product having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the generation method of the network element configuration according to any one of claims 1 to 10.