Network simulation method, device and system based on digital twinning
By using a network simulation method based on digital twins, and leveraging edge computing nodes and a network cloud decision-making platform, the real-time operational data of physical layer device groups is analyzed and simulated. This solves the problem of large discrepancies between simulation results and actual conditions in existing technologies, and achieves full-architecture simulation and efficient fault solution verification.
Patent Information
- Application Number
- CN202510996938.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-21
AI Technical Summary
Existing network simulation technologies cannot achieve full-architecture simulation of actual network environments, resulting in significant deviations between simulation results and reality.
A network simulation method based on digital twins is adopted. By utilizing edge computing nodes and a network cloud decision-making platform, the system performs preliminary analysis by receiving real-time operating data from the physical layer device group, generates fault solutions, and uses digital twin topology modules to conduct simulation tests to verify the feasibility of the solutions. Finally, control commands are executed in the physical layer device group.
It achieves full architecture simulation of the actual network environment, reduces the deviation between simulation results and actual operation, improves the accuracy and real-time performance of simulation results, and reduces cloud load pressure.
Smart Images

Figure CN121000613A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network simulation, and particularly relates to a network simulation method, device and system based on digital twinning. BACKGROUND
[0002] With the rapid development of network cloud technology, network architecture is increasingly complex. Although the existing network simulation technology can simulate the network behavior of devices in the actual network environment (such as issuing device control instructions) to a certain extent, it cannot realize full-architecture simulation with the actual network environment, resulting in a large deviation between the simulation result and the actual situation. SUMMARY
[0003] The present application provides a network simulation method, device and system based on digital twinning, aiming to solve the problem of large deviation between the simulation result and the actual situation of the existing network simulation technology.
[0004] In a first aspect, the present application provides a network simulation method based on digital twinning, based on an edge computing node, comprising: receiving real-time running data of a physical layer device group in an actual network environment; performing preliminary analysis on the real-time running data to obtain a preliminary analysis result; in the case where it is determined according to the preliminary analysis result that the physical layer device group has a fault risk, uploading the real-time running data and the preliminary analysis result to a network cloud decision platform; receiving a fault solution issued by the network cloud decision platform; performing simulation test on the fault solution by using a digital twinning topology module to obtain a test result; uploading the test result to the network cloud decision platform for the network cloud decision platform to determine the feasibility of the fault solution; wherein the digital twinning topology module is a digital twin of the actual network environment of at least part of the devices of the physical layer device group.
[0005] As an embodiment, the real-time running data of the physical layer device group is preliminarily analyzed to obtain a preliminary analysis result, specifically comprising: extracting key features of the real-time running data; generating state prediction data of the physical layer device group at the next time step based on the key features as the preliminary analysis result.
[0006] As an embodiment, the real-time running data is received by using a differential synchronization protocol.
[0007] As an embodiment, the differential synchronization protocol adopts a dual-channel redundancy mechanism, the main channel adopts the WebSocket communication protocol to realize the transmission of real-time streams, and the standby channel adopts the message queue telemetry transfer protocol to realize offline transmission and differential update.
[0008] In a second aspect, the present application also provides a network simulation method based on digital twinning, based on a network cloud decision platform, comprising: receiving real-time running data of a physical layer device group uploaded by any edge computing node and a preliminary analysis result of the real-time running data, the preliminary analysis result showing that the physical layer device group has a fault risk; performing root cause analysis according to the real-time running data and the preliminary analysis result; generating a fault solution scheme according to the root cause analysis result; downloading the fault solution scheme to all edge computing nodes for each edge computing node to simulate and test the fault solution scheme by using a digital twinning topology module; receiving test results of all edge computing nodes; determining the feasibility of the fault solution scheme according to the test results; in the case that the fault solution scheme is feasible, downloading a control instruction to the physical layer device group according to the fault solution scheme for the physical layer device group to execute the control instruction and verify the effect of the fault solution scheme; wherein the digital twinning topology module is a digital twin of an actual network environment of at least part of the devices of the physical layer device group.
[0009] As an embodiment, the network simulation method based on digital twinning further comprises: optimizing a resource scheduling module of the network cloud decision platform based on a deep reinforcement learning algorithm.
[0010] In a third aspect, the present application also provides a network simulation device based on digital twinning, based on an edge computing node, comprising a first receiving module, a preliminary analysis module, a first uploading module, a second receiving module, a test module and a second uploading module; The first receiving module is used to receive real-time running data of a physical layer device group in an actual network environment; The preliminary analysis module is used to preliminarily analyze the real-time running data to obtain a preliminary analysis result; The first uploading module is used to upload the real-time running data and the preliminary analysis result to a network cloud decision platform in the case that the physical layer device group is determined to have a fault risk according to the preliminary analysis result; The second receiving module is used to receive a fault solution scheme downloaded by the network cloud decision platform; The test module is used to simulate and test the fault solution scheme by using a digital twinning topology module to obtain a test result; The second uploading module is configured to upload the test result to the network cloud decision platform, so that the network cloud decision platform determines the feasibility of the fault solution; The digital twin topology module is a digital twin of an actual network environment of at least part of devices of the physical layer device group.
[0011] In a fourth aspect, the application further provides a network simulation device based on digital twin, based on a network cloud decision platform, comprising a third receiving module, a root cause analysis module, a solution generation module, a first issuing module, a fourth receiving module, a feasibility determination module and a second issuing module; The third receiving module is configured to receive real-time running data of the physical layer device group uploaded by any edge computing node and preliminary analysis results of the real-time running data, and the preliminary analysis results show that the physical layer device group has a fault risk; The root cause analysis module is configured to perform root cause analysis according to the real-time running data and the preliminary analysis results; The solution generation module is configured to generate a fault solution according to the root cause analysis results; The first issuing module is configured to issue the fault solution to all edge computing nodes, so that each edge computing node simulates and tests the fault solution by using the digital twin topology module; The fourth receiving module is configured to receive test results of all edge computing nodes; The feasibility determination module is configured to determine the feasibility of the fault solution according to the test results; The second issuing module is configured to issue a control instruction to the physical layer device group according to the fault solution in the case that the fault solution has feasibility, so that the physical layer device group executes the control instruction to verify the effect of the fault solution; The digital twin topology module is a digital twin of an actual network environment of at least part of devices of the physical layer device group.
[0012] In a fifth aspect, the application further provides a network simulation system based on digital twin, comprising a physical layer device group, at least one edge computing node and a network cloud decision platform; The physical layer device group is configured to collect real-time running data and transmit the real-time running data to the corresponding edge computing node, and receive and execute a control instruction issued by the network cloud decision platform; The edge computing node is configured to preliminarily analyze the real-time running data to obtain preliminary analysis results, upload the real-time running data and the preliminary analysis results to the network cloud decision platform in the case that the physical layer device group has a fault risk according to the preliminary analysis results, simulate and test a fault solution issued by the network cloud decision platform by using a digital twin topology module, and upload test results to the network cloud decision platform; The network cloud decision platform is used for root cause analysis according to real-time operation data and preliminary analysis results uploaded by the edge computing nodes, and generates a fault solution; the fault solution is issued to all edge computing nodes; in the case that the fault solution is feasible, control instructions are issued to the physical layer device group according to the fault solution; The digital twin topology module is a digital twin of an actual network environment of at least part of the devices in the physical layer device group.
[0013] In a sixth aspect, the present application also provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements any of the network simulation methods based on digital twin when executing the computer program.
[0014] In a seventh aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement any of the network simulation methods based on digital twin.
[0015] In an eighth aspect, the present application also provides a computer program product, which comprises a computer program, and the computer program is executable on a processor to implement any of the network simulation methods based on digital twin. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 is one of the structural schematic diagrams of the network simulation system based on digital twin provided by the present application; Figure 2 is one of the flow schematic diagrams of the network simulation method based on digital twin provided by the present application; Figure 3 is the second flow schematic diagram of the network simulation method based on digital twin provided by the present application; Figure 4 is the third flow schematic diagram of the network simulation method based on digital twin provided by the present application; Figure 5 is one of the structural schematic diagrams of the network simulation device based on digital twin provided by the present application; Figure 6 is the second structural schematic diagram of the network simulation device based on digital twin provided by the present application; Figure 7 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0018] For the purposes of the present application, the technical solutions and advantages, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0019] It should be noted that in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the sentence "comprises a" does not exclude the presence of another identical element in the process, method, article or device comprising the element.
[0020] The terms "first", "second", and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a class, and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / ", generally means that the front and rear associated objects are in a "or" relationship.
[0021] The embodiments of the present application will be described below in conjunction with Figures 1 to 7 The network simulation method, device and system based on digital twinning provided by the present application are described.
[0022] As shown in Figure 1 The network simulation system based on digital twinning provided by the present application includes a physical layer device group 110, at least one edge computing node 120 and a network cloud decision platform 130.
[0023] The physical layer device group 110 includes a plurality of devices, and the devices are network connected with each other to realize mutual communication, whereby the physical layer device group forms an actual network environment.
[0024] The physical layer device group 110 is configured to collect real-time running data and transmit the real-time running data to the corresponding edge computing node; receive and execute the control instructions issued by the network cloud decision platform.
[0025] One or more edge computing nodes can be deployed in one physical layer device group. In the case of deploying multiple edge computing nodes, each edge computing node is responsible for processing the data of a part of devices in the physical layer device group. In the case of deploying one edge computing node, the edge computing node is responsible for processing the data of all devices in the physical layer device group.
[0026] Each edge computing node 120 is configured to: perform preliminary analysis on the real-time running data uploaded by the corresponding devices to obtain preliminary analysis results; in the case where it is determined according to the preliminary analysis results that there is a risk of failure in the physical layer device group, upload the real-time running data and the preliminary analysis results to the network cloud decision platform; simulate and test the failure solution issued by the network cloud decision platform by using a digital twin topology module; and upload the test results to the network cloud decision platform.
[0027] The digital twin topology module is a digital twin of the actual network environment of at least a part of the devices in the physical layer device group. Specifically, in the case of deploying multiple edge computing nodes, the digital twin topology module of each edge computing node is a digital twin of the actual network environment of a part of the devices in the physical layer device group. In the case of deploying one edge computing node, the digital twin topology module of the edge computing node is a digital twin of the actual network environment of all devices in the physical layer device group.
[0028] In one possible implementation, the digital twin topology module of the edge computing node is a lightweight digital twin, which matches the hardware resources of the edge device. The use of a lightweight digital twin topology module makes the calculation faster, shortens the response time, meets the real-time requirement, simplifies the deployment and update process, and improves the operation and maintenance efficiency.
[0029] In one possible implementation, the simulator (such as a load generator based on Prometheus+Grafana) of the digital twin topology module is trained using the historical data of the Kubernetes cluster to simulate scenarios such as node failure and traffic burst.
[0030] The network cloud decision platform 130 is configured to: perform root cause analysis according to the real-time running data and the preliminary analysis results uploaded by the edge computing nodes, and generate a failure solution; issue the failure solution to all edge computing nodes; and in the case where the failure solution is feasible, issue a control instruction to the physical layer device group according to the failure solution.
[0031] It should be noted that the network cloud decision platform 130 can simultaneously receive the real-time running data and the preliminary analysis results of multiple edge computing nodes, can jointly perform root cause analysis on these data, and can give a global failure solution to solve multiple failure risks.
[0032] The embodiment of the application simulates the running results of the physical layer device group in the entire actual network environment by using the digital twin topology module of all edge computing nodes to simulate and test the fault solution generated by the network cloud decision platform, determines the feasibility of the fault solution, and realizes full-architecture simulation based on the actual network environment, thereby reducing the deviation between the simulation results and the actual running conditions. Meanwhile, by performing simulation testing through the edge computing nodes, the load pressure of the cloud can be reduced, and the real-time performance of the decision can be improved. The application is suitable for high-dynamic demand scenarios such as cloud gaming and industrial Internet of Things.
[0033] In a possible implementation, the edge computing node and the physical layer device group realize state synchronization through a differential synchronization protocol. When transmitting real-time running data, the physical layer device group only transmits the data that has changed compared with the last transmitted data (i.e., incremental data) instead of complete data, thereby reducing the data transmission amount by more than 50%, saving bandwidth resources, improving the reliability of data transmission, and greatly improving the synchronization efficiency, reducing the delay from seconds to milliseconds.
[0034] In a possible implementation, the differential synchronization protocol adopts a dual-channel redundancy mechanism, the main channel adopts a WebSocket communication protocol to realize real-time stream transmission, and the standby channel adopts a Message Queuing Telemetry Transport (MQTT) protocol to realize offline transmission and differential update.
[0035] Generally, the main channel is used for real-time transmission of data, but in a few special cases (such as network interruption and data interruption), the standby channel has a message queue retransmission mechanism to store the data packets lost during the network interruption, and retransmits or continues the transmission through the differential update mechanism after the network is restored. Through the cooperative work of the WebSocket and MQTT dual channels, the data can be completely and real-time synchronized under special network environments (such as network interruption and delay).
[0036] The embodiment of the application ensures the integrity of data transmission between the edge computing node and the physical layer device group through the cooperative work of the WebSocket and MQTT dual channels.
[0037] In a possible implementation, the digital twin topology module of the edge computing node updates the states of the virtual nodes of each device in the digital twin topology module based on the real-time running data of the physical layer device group, so that, during the simulation testing of the fault solution, the digital twin topology module does not need to first simulate the real-time state of the device, but can directly enter the testing, thereby improving the efficiency of the simulation testing.
[0038] In a possible implementation, the network cloud decision platform deploys a digital twin topology module of the entire physical layer device group, and the edge computing node or the physical layer device group uploads real-time running data to the network cloud decision platform, the network cloud decision platform updates the state of the virtual node of each device in the digital twin topology module in real time, and visualizes the state of the entire digital twin topology module for the user to view. In this way, the state of the physical layer device group is intuitively displayed through the digital twin topology module of the network cloud decision platform, and the user can quickly master the system situation.
[0039] Based on the above, the application also provides a network simulation method based on digital twinning. It should be noted that the network simulation method based on digital twinning provided by the embodiments of the application is implemented based on the network simulation system based on digital twinning.
[0040] Figure 2 is one of the flowcharts of the network simulation method based on digital twinning provided by the application.
[0041] As shown in Figure 2 , the network simulation method based on digital twinning provided by the application comprises: S210: The physical layer device group in the actual network environment collects real-time running data and transmits the real-time running data to the corresponding edge computing node. The real-time running data includes temperature, vibration, network traffic, etc.
[0042] S220: The edge computing node preliminarily analyzes the real-time running data to obtain a preliminary analysis result (for example, a temperature change trend, a network traffic continuity, etc.), and determines that there is a fault risk in the physical layer device group according to the preliminary analysis result. If so, S230 is executed (see, for example, 2); otherwise, return to S210.
[0043] S230: The edge computing node uploads the real-time running data and the preliminary analysis result to the network cloud decision platform.
[0044] S240: After receiving the real-time running data and the preliminary analysis result, the network cloud decision platform performs root cause analysis according to the real-time running data and the preliminary analysis result, generates a fault solution (for example, reducing the working frequency of the target device, shutting down the device, etc.), and issues the fault solution to all edge computing nodes.
[0045] S250: All edge computing nodes simulate and test the fault solution issued by the network cloud decision platform by using a digital twin topology module.
[0046] It should be noted that the adjustment of individual devices in the physical layer device group (for example, adjusting the hardware working frequency, turning off the device, etc.) may affect the operation of the entire physical layer device group, therefore, in this step, all edge computing nodes of the physical layer device group are used to simulate and test the fault solution to realize the simulation of the entire actual network environment, so that the simulation result is closer to the result of actual operation.
[0047] S260: All edge computing nodes upload the test results to the network cloud decision platform.
[0048] S270: The network cloud decision platform determines the feasibility of the fault solution according to the test results of all edge computing nodes. If the fault solution is feasible, S280 is executed (see, for example, 2); otherwise, return to S240.
[0049] S280: According to the fault solution, determine the control instruction (such as controlling the target device to reduce the working frequency or controlling the target device to turn off, etc.) and send the control instruction to the physical layer device group for the physical layer device group to execute the control instruction and verify the effect of the fault solution. If the verification result is that the fault solution does not achieve the expected effect, return to step S240.
[0050] Specifically, the network cloud decision platform verifies the effect of the fault solution by judging whether the edge computing node continues to upload the same fault risk. For example, in step S230, if the edge computing node uploads the risk of device A having temperature exceeding the standard, in step S280, the network cloud decision platform judges whether the preliminary analysis result of device A having the risk of temperature exceeding the standard is received within a preset time since the control instruction is sent to the physical layer device group. If the preliminary analysis result of the temperature exceeding the standard risk is received within the preset time, the fault solution does not achieve the expected effect; otherwise, the fault solution achieves the expected effect.
[0051] Based on the above, Figure 3 is a flowchart of the network simulation method based on digital twinning provided by the present application, which corresponds to the edge computing node.
[0052] As Figure 3 shown, the network simulation method based on digital twinning comprises: S310: Receive real-time running data of the physical layer device group in the actual network environment.
[0053] S320: Preliminary analysis is performed on the real-time running data to obtain a preliminary analysis result.
[0054] S330: In the case where it is determined according to the preliminary analysis result that the physical layer device group has a fault risk, upload the real-time running data and the preliminary analysis result to the network cloud decision platform.
[0055] S340: receiving a fault solution issued by the network cloud decision platform.
[0056] S350: simulating and testing the fault solution by using the digital twin topology module to obtain a test result.
[0057] S360: uploading the test result to the network cloud decision platform for the network cloud decision platform to determine the feasibility of the fault solution.
[0058] Embodiments of the present application utilize the digital twin topology module of the edge computing node to perform simulation testing. Since the topology structure in the digital twin topology module is the same as that of the physical layer device group, the simulation testing can highly simulate the actual operation of the device, greatly reducing the deviation between network simulation and actual situation.
[0059] In one possible implementation, in step S320, the real-time operation data of the physical layer device group is preliminarily analyzed to obtain a preliminary analysis result, which specifically includes: S3201: extracting key features of the real-time operation data.
[0060] In one possible implementation, when the key features of the real-time operation data are extracted, the space-time features (time information and spatial relationship) in the real-time operation data are compressed into key features that can reflect the change trend of each index, such as temperature rising for 3 times in succession.
[0061] S3202: generating state prediction data of the physical layer device group at the next time step based on the key features as the preliminary analysis result.
[0062] In one possible implementation, the extraction of the key features and the state prediction are implemented by using an LSTM-GAN hybrid neural network.
[0063] First, the space-time features are encoded by using a Long Short-Term Memory (LSTM) model to compress them into key features, so that more accurate prediction can be made with less data in the following prediction process. The input data of the LSTM network model is the real-time operation data, and the output data is the key features. The trained LSTM network model learns the long-term dependence relationship of the historical time series data of the physical layer device group (for example, temperature rising for 3 times in succession may cause overheating of the device), and the extraction of the key features is based on the model. The LSTM network model solves the problems of inaccurate prediction and modeling lag of the traditional digital twin topology module.
[0064] Subsequently, the GAN generator in Generative Adversarial Networks (GANs) is used to generate state prediction data (e.g., predicting the temperature for the next 5 seconds) based on key features, conforming to the distribution of real data. GANs can address the problems of insufficient training data and lack of extreme conditions in traditional digital twin topology modules. Simultaneously, the GAN discriminator in GANs, by distinguishing between real and generated data, drives the optimization of the GAN generator, thus solving the overfitting problem of traditional digital twin topology modules.
[0065] In one possible implementation, the prediction mechanism is as follows: ; in The device state vector determined for the current key features. G This is a GAN generator network. The prediction mechanism is based on historical data (such as...). Using current data (i.e., device time-series data), we can predict the future state of a device (e.g., predict the temperature in the next few minutes or hours).
[0066] This application embodiment learns the long-term dependencies of historical time-series data of physical layer device groups through an LSTM network model, and extracts key features based on the model, using fewer features to represent device status. It also uses generative adversarial networks to solve the problems of insufficient training data, missing extreme conditions, and overfitting in traditional digital twin topology modules, thereby improving prediction accuracy. At the same time, the LSTM-GAN hybrid neural network can achieve high real-time performance in preliminary analysis.
[0067] In one possible implementation, in step S310, a differential synchronization protocol is used to receive real-time running data.
[0068] This application embodiment reduces the amount of data transmission between edge computing nodes and physical layer device groups by more than 50% through differential synchronization protocol, saving bandwidth resources and greatly improving synchronization efficiency, reducing latency from the second level to the millisecond level.
[0069] Figure 4 This is the third flowchart of the network simulation method based on digital twins provided in this application, and the flowchart corresponds to the network cloud decision-making platform.
[0070] like Figure 4 As shown, network simulation methods based on digital twins include: S410: Receives real-time operating data of the physical layer device group uploaded by any edge computing node, as well as preliminary analysis results of the real-time operating data. The preliminary analysis results indicate that the physical layer device group has a risk of failure.
[0071] S420: performing root cause analysis according to the real-time operation data and the preliminary analysis result.
[0072] In a possible implementation, the network cloud decision platform performs root cause analysis by using the LSTM-GAN hybrid neural network in step S320 described above. For details, refer to the description of step S320, which will not be repeated here. The LSTM-GAN hybrid neural network can achieve high real-time performance of root cause analysis and improve the prediction accuracy.
[0073] S430: generating a fault solution according to the root cause analysis result.
[0074] S440: issuing the fault solution to all edge computing nodes, so that each edge computing node simulates and tests the fault solution by using the digital twin topology module.
[0075] S450: receiving the test results of all edge computing nodes.
[0076] S460: determining the feasibility of the fault solution according to the test results.
[0077] S470: in the case where the fault solution is feasible, issuing a control instruction to the physical layer device group according to the fault solution, so that the physical layer device group executes the control instruction to verify the effect of the fault solution. If the verification result is that the fault solution does not achieve the expected effect, return to step S420.
[0078] In the embodiment, the network cloud decision platform simultaneously simulates and tests the same fault solution by using all edge computing nodes, so as to simulate the running of the full architecture of the physical layer device group in the fault solution, making the simulation result more similar to the actual running result. In the case where the test of the edge computing node shows that the solution is feasible, the solution is verified by the physical layer device group, further improving the accuracy of network simulation.
[0079] In a possible implementation, the network cloud decision platform adopts a container deployment strategy (such as a Kubernetes cluster) to package an application and all dependent codes, system tools, libraries, and the like into a lightweight and portable "container image", which can be run in any environment (for example, an edge computing node) by one key without reconfiguration. The resource scheduling module is used to implement the resource scheduling of the container. According to the real-time changes of the use demand of the cloud service application, the process of dynamically allocating the resources related to the container and optimizing and adjusting the deployment of the container in the network cloud decision platform is performed. For example, when it is found that the access volume of a cloud service application carried by a container is increased and the resources are in shortage at a certain period, the resource scheduling module will take some measures to migrate the container to a server with low resource utilization.
[0080] In a possible implementation, the resource scheduling module includes a monitoring module, an intelligent decision module, and an execution module.
[0081] The monitoring module is responsible for collecting real-time state information of the cluster and the container, including CPU utilization, memory utilization, disk input / output, network bandwidth, request delay, Service Level Objective (SLO) violation times, and the like.
[0082] The intelligent decision module learns the optimal resource scheduling strategy according to the real-time state information collected by the monitoring module, and generates scheduling instructions. For example, when it is predicted that a device is about to fail, the resource scheduling module can migrate the tasks on the device to other devices by adjusting the resources of the container to which the application of the device belongs, to ensure the continuity of the business.
[0083] The execution module is responsible for executing the scheduling instructions generated by the intelligent decision module, including container migration, resource adjustment, and the like.
[0084] On the basis described above, in a possible implementation, the network simulation method based on digital twinning further includes: optimizing the resource scheduling module of the network cloud decision platform based on a deep reinforcement learning algorithm.
[0085] The embodiments of the present application optimize the resource scheduling algorithm of the container through the resource scheduling module, improve the reliability of resource scheduling, can effectively improve the resource utilization, reduce the service delay, and improve the service quality.
[0086] In a possible implementation, the optimization problem of the resource scheduling algorithm is modeled as a Markov Decision Process (MDP), including the following elements: a: state: describes the real-time state of the cluster and the container, including: node resources: CPU utilization , memory utilization , disk input / output , network bandwidth of the node , and the like; container demand: request CPU , memory (j ), resource request and runtime of the container, and the like; Quality of Service (QoS): current request delay , SLO violation times , delay, concurrency, and the like; System state: pending container queue length q、 Historical load trend , node failure network delay, etc.
[0087] The mathematical formula of the state is: ; b: action: describes the operation of resource scheduling, including: Container migration: migrate containers from one node to another node; Resource adjustment: adjust the CPU and memory resources of the container.
[0088] c: reward: used to evaluate the pros and cons of the scheduling strategy, which can be designed according to resource utilization, service delay, SLO violation times, etc. The reward function is: ; Where, , , is the weight value. The SLA satisfaction rate is the achievement rate of the service level agreement (Service Level Agreement).
[0089] In one possible implementation, the resource scheduling module of the network cloud decision platform is optimized based on a deep reinforcement learning algorithm, which specifically includes: Loop the following steps until the algorithm converges: P1: initialization: initialize the policy network and value network of the deep reinforcement learning algorithm.
[0090] In one possible implementation, the deep reinforcement learning algorithm used is the Proximal Policy Optimization (PPO) algorithm.
[0091] The main role of the policy network is to generate a probability distribution of an action (a) according to the current state (s), or directly output a deterministic action. The policy network essentially learns a policy function that describes the probability of taking action a given state s. The input of the policy network is the state of the current environment. The device state can be the original observation value (such as image pixels), or a pre-processed feature vector. The output of the policy network is a discrete action space or a continuous action space. The training goal of the policy network is to maximize the expected return, that is, by adjusting the network parameters, the network cloud decision platform can take actions that obtain higher cumulative rewards.
[0092] The main role of the value network is to evaluate the expected return that the network cloud decision platform can obtain in a given state. The value network is essentially a scalar value representing the value function V(s) of state s, which describes the expected value of future cumulative rewards in state s. The value network is also known as the critic, which evaluates the goodness of the current policy. The input of the value network is the current state of the environment, which is the same as the policy network. The output of the value network is the value function V(s). This value is an estimate of the future cumulative reward, which is used to evaluate the goodness of the current state. The training goal of the value network is to make its output value function V(s) as close as possible to the true cumulative reward.
[0093] In the PPO algorithm, the policy network is responsible for selecting actions, and the value network is responsible for evaluating states. The policy network generates actions based on the current state, and then the environment provides rewards and the next state. The value network evaluates the value of the current state, which is used to guide the update of the policy network.
[0094] P2: State observation: Obtain real-time state information of clusters and containers from the monitoring module.
[0095] P3: Action selection: Select an action according to the current policy network.
[0096] P4: Action execution: Execute the selected action and update the state of the cluster and container.
[0097] P5: Reward calculation: Calculate the reward value according to the updated state.
[0098] P6: Data storage: Store state, action, reward, and other data into the experience replay buffer.
[0099] P7: Policy update: Sample data from the experience replay buffer and update the policy network and value network using the PPO algorithm.
[0100] The core formula of the PPO algorithm is as follows: ; represents the total loss function of policy update, which is used to guide the adjustment of the policy in the real environment; represents the clipping loss function used in the PPO algorithm, which improves the stability of training by limiting the changes between the old and new policies; is a hyperparameter that balances the importance of clipping loss and KL divergence term, which controls the weight of the KL divergence term in the total loss; represents the true policy (policy under the current policy network parameters) under the parameters of the policy network The KL (Kullback-Leibler) divergence between the two probability distributions is used to ensure that the policy in the real environment does not deviate too far from the policy learned in the simulated environment.
[0101] in, ; in, Indicates time step t The expected value is usually estimated by averaging a batch of samples; Indicates the state Use strategy Select Action The resulting probability ratio is calculated using the new strategy. Compared to the old strategy The ratio of the probabilities of choosing that action, i.e.: ; This represents the estimated value of the advantage function at time step [number]. t The advantage function estimate measures the performance of a state. Take action below Benefits relative to the average level. Indicates to Cut to ensure The value in and The pruning operation is the core of the PPO algorithm, used to limit the magnitude of policy updates and prevent the policy from changing too much in a single iteration. It is a hyperparameter, typically ranging from 0.1 to 0.3. min represents the minimum function, used to calculate the minimum value between the clipped objective function value and the original objective function value.
[0102] The embodiments of this application realize the adaptability and automatic dynamic optimization of the resource scheduling algorithm through deep reinforcement learning algorithm, so as to achieve a cloud resource utilization rate of 92% and a 40% reduction in SLA default rate.
[0103] Based on the above, this application also provides a network simulation device based on digital twins. The network simulation device based on digital twins and the aforementioned network simulation method based on digital twins can be referred to and corresponded to each other.
[0104] Figure 5 This is one of the structural schematic diagrams of the network simulation device 500 based on digital twins provided in this application, corresponding to an edge computing node. For example... Figure 5 As shown, the network simulation device includes: a first receiving module 510, a preliminary analysis module 520, a first uploading module 530, a second receiving module 540, a testing module 550, and a second uploading module 560.
[0105] The first receiving module 510 is configured to receive real-time running data of the group of physical layer devices in the actual network environment.
[0106] The preliminary analysis module 520 is configured to perform preliminary analysis on the real-time running data to obtain a preliminary analysis result.
[0107] The first uploading module 530 is configured to upload the real-time running data and the preliminary analysis result to the network cloud decision platform if it is determined according to the preliminary analysis result that the group of physical layer devices has a risk of failure.
[0108] The second receiving module 540 is configured to receive a failure solution issued by the network cloud decision platform.
[0109] The testing module 550 is configured to perform simulation testing on the failure solution by using the digital twin topology module to obtain a testing result.
[0110] The second uploading module 560 is configured to upload the testing result to the network cloud decision platform, so that the network cloud decision platform determines the feasibility of the failure solution.
[0111] The digital twin topology module is a digital twin of the actual network environment of at least part of the devices in the group of physical layer devices.
[0112] The embodiment of the application performs simulation testing by using the digital twin topology module of the edge computing node. Since the topology structure in the digital twin topology module is the same as the topology structure of the group of physical layer devices, the simulation testing can highly simulate the running condition of the actual devices, and greatly reduces the deviation between network simulation and actual condition.
[0113] In a possible implementation manner, the preliminary analysis module is specifically configured to: extract key features of the real-time running data.
[0114] generate state prediction data of the group of physical layer devices at a next time step based on the key features, as the preliminary analysis result.
[0115] In a possible implementation manner, the LSTM-GAN hybrid neural network is adopted to implement the key feature extraction and the state prediction.
[0116] The embodiment of the application learns the long-term dependence relationship of the historical time series data of the group of physical layer devices by using the LSTM network model, and extracts the key features based on the model, so that the device state is represented by fewer features. The generation adversarial network is used to solve the problems of insufficient training data, missing extreme working conditions and overfitting of the traditional digital twin topology module, so as to improve the prediction accuracy. Meanwhile, the LSTM-GAN hybrid neural network can realize high real-time performance of the preliminary analysis.
[0117] In a possible implementation, the first receiving module receives the real-time running data in a differential synchronization protocol.
[0118] The embodiments of the present application reduce the data transmission amount between the edge computing nodes and the physical layer device group by more than 50% through the differential synchronization protocol, save bandwidth resources, and greatly improve the synchronization efficiency, reducing the delay from seconds to milliseconds.
[0119] Figure 6 is a structure diagram of a network simulation device 600 provided by the present application based on digital twinning, corresponding to an edge computing node. As shown in Figure 6 The network simulation device includes a third receiving module 610, a root cause analysis module 620, a scheme generation module 630, a first issuing module 640, a fourth receiving module 650, a feasibility determination module 660, and a second issuing module 670.
[0120] The third receiving module 610 is configured to receive the real-time running data of the physical layer device group uploaded by any edge computing node and the preliminary analysis result of the real-time running data. The preliminary analysis result shows that the physical layer device group has a fault risk.
[0121] The root cause analysis module 620 is configured to perform root cause analysis according to the real-time running data and the preliminary analysis result.
[0122] The scheme generation module 630 is configured to generate a fault solution according to the root cause analysis result.
[0123] The first issuing module 640 is configured to issue the fault solution to all edge computing nodes, so that each edge computing node simulates and tests the fault solution by using a digital twinning topology module.
[0124] The fourth receiving module 650 is configured to receive the test results of all edge computing nodes.
[0125] The feasibility determination module 660 is configured to determine the feasibility of the fault solution according to the test results.
[0126] The second issuing module 670 is configured to issue a control instruction to the physical layer device group according to the fault solution in the case that the fault solution has feasibility, so that the physical layer device group executes the control instruction to verify the effect of the fault solution.
[0127] In the embodiments of the present application, the network cloud decision platform simultaneously simulates and tests the same fault solution by using all edge computing nodes, so as to simulate the running of the whole architecture of the physical layer device group in the fault solution, so that the simulation result is more similar to the actual running result, and in the case that the test of the edge computing node shows that the solution is feasible, the solution is verified by the physical layer device group, thereby further improving the accuracy of network simulation.
[0128] In a possible implementation, the network simulation method further includes an optimization module configured to optimize the resource scheduling module of the network cloud decision platform based on a deep reinforcement learning algorithm.
[0129] The embodiments of the present application optimize the resource scheduling algorithm of the resource scheduling module, improve the reliability of resource scheduling, effectively improve the resource utilization rate, reduce service delay, and improve service quality.
[0130] In a possible implementation, the optimization module is specifically configured to: initialize a policy network and a value network of the deep reinforcement learning algorithm; obtain real-time state information of the cluster and the container from the monitoring module.
[0131] select an action according to the current policy network.
[0132] execute the selected action and update the state of the cluster and the container.
[0133] calculate a reward value according to the updated state.
[0134] store the state, the action, the reward and the like into an experience replay buffer.
[0135] sample data from the experience replay buffer and update the policy network and the value network using a PPO algorithm.
[0136] The embodiments of the present application realize the adaptability and automatic dynamic optimization of the resource scheduling algorithm through the deep reinforcement learning algorithm, so that the cloud resource utilization rate reaches 92%, and the SLA violation rate decreases by 40%.
[0137] Figure 7 is a structural schematic diagram of an electronic device provided by the present application, such as Figure 7As shown, the electronic device can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 complete communications with each other through the communications bus 740. The processor 710 can invoke the logic instructions in the memory 730 to execute any of the above-described network simulation methods based on digital twinning based on an edge computing node or a network cloud decision platform.
[0138] In addition, the logic instructions in the memory 730 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0139] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer readable storage medium, and the computer program includes program instructions, when the program instructions are executed by a computer, the computer can execute any of the above-described network simulation methods based on digital twinning based on an edge computing node or a network cloud decision platform.
[0140] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement any of the above-described network simulation methods based on digital twinning based on an edge computing node or a network cloud decision platform.
[0141] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement it without creative labor.
[0142] Those skilled in the art can clearly understand the implementation of the embodiments by means of software and necessary general hardware platforms through the description of the above embodiments, and of course, the embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in the sense of contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a read-only memory (ROM) / random access memory (RAM), a magnetic disk, an optical disc, or the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the methods.
[0143] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features thereof; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A network simulation method based on digital twinning, characterized in that, The edge computing node comprises: receiving real-time running data of a physical layer device group in an actual network environment; performing preliminary analysis on the real-time running data to obtain preliminary analysis results; in the case where it is determined according to the preliminary analysis results that the physical layer device group has a fault risk, uploading the real-time running data and the preliminary analysis results to a network cloud decision platform; receiving a fault solution issued by the network cloud decision platform; performing simulation test on the fault solution by using a digital twin topology module to obtain test results; uploading the test results to the network cloud decision platform for the network cloud decision platform to determine the feasibility of the fault solution; wherein the digital twin topology module is a digital twin of an actual network environment of at least part of devices of the physical layer device group.
2. The digital-twin-based network simulation method of claim 1, wherein, The preliminary analysis on the real-time running data of the physical layer device group to obtain preliminary analysis results specifically comprises: extracting key features of the real-time running data; generating state prediction data of the physical layer device group at a next time step based on the key features as the preliminary analysis results.
3. The digital-twin-based network simulation method of claim 1, wherein, The real-time running data is received by using a differential synchronization protocol.
4. The digital-twin-based network simulation method of claim 3, wherein, The differential synchronization protocol adopts a dual-channel redundancy mechanism, a main channel uses a WebSocket communication protocol to realize real-time stream transmission, and a standby channel uses a message queue telemetry transmission protocol to realize network recovery and differential update.
5. A network simulation method based on digital twinning, characterized in that, The network cloud decision platform comprises: receiving real-time running data of a physical layer device group uploaded by any edge computing node and preliminary analysis results of the real-time running data, the preliminary analysis results showing that the physical layer device group has a fault risk; performing root cause analysis according to the real-time running data and the preliminary analysis results; generating a fault solution according to the root cause analysis results; issuing the fault solution to all edge computing nodes for each edge computing node to perform simulation test on the fault solution by using a digital twin topology module; receiving test results of all edge computing nodes; determining the feasibility of the fault solution according to the test results; in the case where the fault solution has feasibility, issuing a control instruction to the physical layer device group according to the fault solution for the physical layer device group to execute the control instruction to verify the effect of the fault solution; wherein the digital twin topology module is a digital twin of an actual network environment of at least part of devices of the physical layer device group.
6. The digital-twin-based network simulation method of claim 5, wherein, It further comprises: optimizing a resource scheduling module of the network cloud decision platform based on a deep reinforcement learning algorithm.
7. A network simulation apparatus based on digital twinning, characterized by, The edge computing node comprises a first receiving module, a preliminary analysis module, a first uploading module, a second receiving module, a test module and a second uploading module; the first receiving module is used for receiving real-time running data of a physical layer device group in an actual network environment; the preliminary analysis module is used for performing preliminary analysis on the real-time running data to obtain preliminary analysis results; The first uploading module is configured to upload the real-time operation data and the preliminary analysis result to a network cloud decision platform if it is determined according to the preliminary analysis result that the physical layer device group is at risk of failure. The second receiving module is configured to receive a failure solution issued by the network cloud decision platform. The testing module is configured to simulate test the failure solution by using a digital twin topology module to obtain a test result. The second uploading module is configured to upload the test result to the network cloud decision platform for the network cloud decision platform to determine the feasibility of the failure solution. The digital twin topology module is a digital twin of an actual network environment of at least part of devices of the physical layer device group.
8. A network simulation apparatus based on digital twinning, characterized by, The network cloud decision platform comprises a third receiving module, a root cause analysis module, a solution generation module, a first issuing module, a fourth receiving module, a feasibility determination module, and a second issuing module. The third receiving module is configured to receive real-time operation data of a physical layer device group uploaded by any edge computing node and a preliminary analysis result of the real-time operation data, the preliminary analysis result indicating that the physical layer device group is at risk of failure. The root cause analysis module is configured to perform root cause analysis according to the real-time operation data and the preliminary analysis result. The solution generation module is configured to generate a failure solution according to a root cause analysis result. The first issuing module is configured to issue the failure solution to all edge computing nodes, for each edge computing node to simulate test the failure solution by using a digital twin topology module. The fourth receiving module is configured to receive test results of all edge computing nodes. The feasibility determination module is configured to determine the feasibility of the failure solution according to the test results. The second issuing module is configured to issue a control instruction to the physical layer device group according to the failure solution if the failure solution is feasible, for the physical layer device group to execute the control instruction to verify the effect of the failure solution. The digital twin topology module is a digital twin of an actual network environment of at least part of devices of the physical layer device group. 9.A network simulation system based on digital twinning, characterized in that, The network cloud decision platform comprises a third receiving module, a root cause analysis module, a solution generation module, a first issuing module, a fourth receiving module, a feasibility determination module, and a second issuing module. The physical layer device group is configured to collect real-time operation data and transmit the real-time operation data to a corresponding edge computing node, and receive and execute a control instruction issued by the network cloud decision platform. The edge computing node is configured to preliminarily analyze the real-time operation data to obtain a preliminary analysis result. The real-time operation data and the preliminary analysis result are uploaded to the network cloud decision platform if it is determined according to the preliminary analysis result that the physical layer device group is at risk of failure. The failure solution issued by the network cloud decision platform is simulated tested by using a digital twin topology module, and a test result is uploaded to the network cloud decision platform. The network cloud decision platform is configured to perform root cause analysis based on the real-time operation data and the preliminary analysis result uploaded by the edge computing nodes, and generate a fault solution; The fault solution is sent to all the edge computing nodes; In the case that the fault solution is feasible, control instructions are sent to the group of physical layer devices according to the fault solution; The digital twin topology module is a digital twin of the actual network environment of at least part of the devices in the group of physical layer devices.
10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein The processor executes the computer program to implement the network simulation method based on digital twin as claimed in any one of claims 1 to 4 or any one of claims 5 to 6.
11. A non-transitory computer-readable storage medium having stored thereon a computer program, wherein The computer program is executed by the processor to implement the network simulation method based on digital twin as claimed in any one of claims 1 to 4 or any one of claims 5 to 6.
12. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the network simulation method based on digital twin as claimed in any one of claims 1 to 4 or any one of claims 5 to 6.
Citation Information
Cited By
Cloud platform container scheduling method and device, electronic equipment and storage medium
CN121636175A
Internet of Things data acquisition and abnormity early warning method and system
CN121887795A