Digital twinning of network
By selecting probability distributions associated with network configuration and terminal device services, a digital twin model is generated, solving the problems of accuracy and configuration optimization in existing telecommunications networks, and realizing accurate simulation of network behavior and performance optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies struggle to create accurate digital twins in the telecommunications network field, failing to effectively reflect the interactions and mechanisms between access networks and user home networks, devices, and services. This makes it difficult to assess capacity and bandwidth requirements and to find the optimal network configuration.
By selecting probability distributions associated with network configuration and terminal device services, the time series of data volume is determined, and network data flow is simulated based on these values to generate an accurate digital twin model that takes into account factors such as network device behavior, service provision, home/terminal networks, and user interactions.
It enables the replication of real-world network behavior in the digital world, allowing for the evaluation of different network configurations and scenarios, network testing, generation of realistic datasets, optimization of network performance, and cost reduction.
Smart Images

Figure CN121644376A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Example embodiments relate to a system, method and computer program for providing a digital twin of a network. The network can be any access network and associated nodes. The digital twin is configured to provide information indicative of data flow through the network. In some examples, the digital twin is used in the context of a Digital Subscriber Line (DSL) or Passive Optical Network (PON) network. BACKGROUND
[0002] A simulator is a device, software program or system designed to replicate the behavior of a real-world system, process or environment. Simulators are widely used in numerous fields including aviation, engineering, healthcare and entertainment to train personnel, test designs, and improve performance. Simulators use mathematical models / algorithms to create a realistic representation of the system or process being simulated. By adjusting variables and inputs, a simulator can test different scenarios and predict outcomes. In this way, the simulation replicates what can happen to the system or object. However, simulators are not perfect representations of reality. They are created to mimic the behavior of a real-world system, process or environment as realistically as possible, but there are always limitations to the accuracy of the simulation and there are always some factors that cannot be captured in the simulation.
[0003] In contrast, a digital twin is a virtual representation of a physical object, system or process. It is a digital model that uses data and algorithms to simulate and predict the behavior and performance of the physical object or entity. In this way, the digital twin replicates (in a digital environment) the real-world process, thereby digitally replicating or mirroring what happens in the real world to the particular physical object, system or entity.
[0004] The concept of digital twins has become a way for various industries to improve efficiency, productivity and quality. For example, in manufacturing, digital twins can simulate (or represent) the entire production process from design to assembly, enabling engineers to optimize performance, predict maintenance needs and reduce downtime. In addition to manufacturing, digital twins are used in areas such as healthcare, transportation, and construction. In the healthcare field, digital twins can be used to simulate and analyze the behavior of organs or biological systems, helping doctors to develop personalized treatment plans. In the transportation field, digital twins can simulate traffic patterns, optimize routes, and improve safety.
[0005] Digital twins are considered in many ways to be an improvement over simulators because they are able to more accurately represent the physical system or process being modeled. Digital twins can simulate and predict the behavior of a physical system under different conditions, enabling engineers and operators to optimize performance, identify potential problems, and make informed decisions. This improved predictive capability helps prevent downtime, reduces maintenance costs, and improves the safety of the physical system / process. Another advantage of digital twins is that they can incorporate machine learning and other advanced analysis techniques. By analyzing large amounts of data, digital twins can identify patterns and make predictions that traditional simulators are unable to achieve or implement. While simulators still hold a place in many industries, the higher accuracy, predictive capability, and advanced analysis capabilities of digital twins are making them an increasingly important tool for optimizing performance, improving safety, and reducing the cost of real-world physical objects, systems, or entities.
[0006] In the field of telecommunications networks, having a reliable digital representation of a network in such a digital twin would bring significant advantages. However, creating such a digital representation can be quite challenging, especially to create a representation that accurately reflects the interactions between an access network and a customer's home (or terminal) network, devices and services, and their associated behaviors and mechanisms.
[0007] Therefore, it is desirable to provide a digital twin of a network to facilitate the evaluation of capacity and / or bandwidth requirements and optionally finding an optimal configuration of the network. SUMMARY
[0008] The scope of protection of embodiments of the application is defined by the independent claims. Embodiments and features that are not within the scope of the independent claims, if any, are to be interpreted as examples useful in understanding the various embodiments of the present application.
[0009] According to a first aspect, the present specification describes a system for implementing a digital twin of a network, the system comprising: a module for selecting one or more first values, wherein each first value represents a respective parameter of a configuration of the network and is selected from a probability distribution associated with the respective parameter of the configuration of the network; a module for selecting one or more second values, wherein each second value represents a respective parameter of a service to be executed on one or more terminal devices of the network, and is selected from a probability distribution associated with the respective parameter of the service to be executed, each service being associated with the transmission of one or more data packets through the network; a module for determining, based on the one or more selected second values, a time series representing an amount of data to be injected into the network in each of one or more time intervals; a module for determining, based on the one or more selected first values and the time series, a flow of the amount of data through the network; and a module for outputting information indicative of the flow of the amount of data.
[0010] According to a second aspect, the present specification describes a method for implementing digital twinning of a network, the method comprising: selecting one or more first values, wherein each first value represents a respective parameter of a configuration of the network and is selected from a probability distribution associated with the respective parameter of the configuration of the network; selecting one or more second values, wherein each second value represents a respective parameter of a service to be executed on one or more terminal devices of the network and is selected from a probability distribution associated with the respective parameter of the service to be executed, each service being associated with transmission of one or more data packets over the network; determining, based on the one or more selected second values, a time series representing an amount of data to be injected into the network at each of one or more time intervals; determining, based on the one or more selected first values and the time series, a flow of the amount of data over the network; and outputting information indicative of the flow of the amount of data.
[0011] According to a third aspect, the present specification describes a computer program comprising instructions for causing an apparatus to perform: selecting one or more first values, wherein each first value represents a respective parameter of a configuration of the network and is selected from a probability distribution associated with the respective parameter of the configuration of the network; selecting one or more second values, wherein each second value represents a respective parameter of a service to be executed on one or more terminal devices of the network and is selected from a probability distribution associated with the respective parameter of the service to be executed, each service being associated with transmission of one or more data packets over the network; determining, based on the one or more selected second values, a time series representing an amount of data to be injected into the network at each of one or more time intervals; determining, based on the one or more selected first values and the time series, a flow of the amount of data over the network; and outputting information indicative of the flow of the amount of data.
[0012] According to a fourth aspect, the present specification describes a computer- readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the following operations: selecting one or more first values, wherein each first value represents a respective parameter of a configuration of the network and is selected from a probability distribution associated with the respective parameter of the configuration of the network; selecting one or more second values, wherein each second value represents a respective parameter of a service to be performed on one or more terminal devices of the network and is selected from a probability distribution associated with the respective parameter of the service to be performed, each service being associated with transmission of one or more data packets over the network; determining, based on the one or more selected second values, a time series representing an amount of data to be injected into the network at each of one or more time intervals; determining, based on the one or more selected first values and the time series, a flow of the amount of data over the network; and outputting information indicative of the flow of the amount of data.
[0013] According to a fifth aspect, the present specification describes an apparatus comprising: at least one processor; and at least one memory including computer program code, the computer program code, when executed by the at least one processor, causing the apparatus to perform the following operations: selecting one or more first values, wherein each first value represents a respective parameter of a configuration of the network and is selected from a probability distribution associated with the respective parameter of the configuration of the network; selecting one or more second values, wherein each second value represents a respective parameter of a service to be performed on one or more terminal devices of the network and is selected from a probability distribution associated with the respective parameter of the service to be performed, each service being associated with transmission of one or more data packets over the network; determining, based on the one or more selected second values, a time series representing an amount of data to be injected into the network at each of one or more time intervals; determining, based on the one or more selected first values and the time series, a flow of the amount of data over the network; and outputting information indicative of the flow of the amount of data.
[0014] Example embodiments of each of these aspects will be hereinafter described.
[0015] The means for determining the flow of the amount of data over the network can further comprise: means for configuring a simulation of the network based on the one or more selected first values; means for injecting the time series into the configured simulation of the network; and means for determining the flow of the amount of data over the configured simulation of the network.
[0016] The module for determining the time series to be injected into the network based on the one or more selected second values may further include a module for determining the time series based on one or more of the following: the amount of bits associated with the one or more selected second values, the scheduling associated with the one or more selected second values, and the delay associated with the one or more selected second values.
[0017] The module for determining the flow of the data volume through the network based on the first value of the one or more selections and the time series may further include a module for determining the flow of the data volume based on one or more of the following: a capacity associated with the first value of the one or more selections, a buffer size associated with the first value of the one or more selections, and a drop policy associated with the first value of the one or more selections.
[0018] Optionally, the module for selecting the one or more first values includes a module for automatically selecting the one or more first values. Optionally, the module for selecting the one or more second values includes a module for automatically selecting the one or more second values.
[0019] In some implementations, a module is also provided for receiving one or more other first values of corresponding parameters representing the configuration of the network based on user input, wherein the stream determining the amount of data passing through the network is also based on the one or more received other first values.
[0020] In some implementations, a module is also provided for receiving one or more other second values of a corresponding parameter representing the service to be performed on the one or more terminal devices based on user input, wherein determining the time series to be injected into the network is also based on the one or more received other second values.
[0021] In some implementations, the network includes: one or more home gateways; one or more home networks; the one or more terminal devices located after the one or more home gateways and connected to the one or more gateways via the one or more home networks; an access node providing access to one or more service providers configured to provide the services to be performed on the one or more terminal devices; and an access network connecting the access node and the one or more home gateways.
[0022] Optionally, the time series is injected into the access node. Optionally, the module for outputting information indicating the flow of the data volume includes a module for outputting information indicating the flow of the data volume passing through the access node.
[0023] Optionally, the information indicating the flow of the data volume includes one or more of the following: the amount of data sent, the amount of data queued, the amount of data dropped, a quality of service indicator, a service slowdown indicator, a service latency indicator, a data queue status, or an impact qualification. The information indicating the flow of the data volume may be provided in the upstream and / or downstream directions.
[0024] Optionally, the network configuration parameters include one or more of the following:
[0025] The type, connectivity, or characteristics of the network node. Optionally, the characteristics of the network node include the number of gateways, the number of terminal devices, the type of terminal devices, or the network subscription. Additionally or alternatively, the characteristics of the network node include the technology of the network node. Optionally, the technology includes one of the following: Asymmetric Digital Subscriber Line (ADSL), Very High Speed Digital Subscriber Line (VDSL), GFast, Gigabit Passive Optical Network (GPON), 10 Gigabit Passive Optical Network (XGPON), or 10 Gigabit Symmetric Passive Optical Network (XGSPON).
[0026] Optionally, the parameters for the service to be executed on the one or more terminal devices include the number of services per terminal device. Additionally or alternatively, the parameters for the service to be executed on the one or more terminal devices include one or more of the following: the type, scheduling, or characteristics of the service. Optionally, the type of service may include: audio, browsing, file transfer, gaming, Internet of Things (IoT), Internet Protocol (IP) TV, video conferencing, IPv4, or video over the top.
[0027] In some implementations, each parameter of the network configuration is associated with one or more specific probability (or probability) distributions. In some implementations, each parameter of the service to be executed on the one or more terminal devices is associated with one or more specific probability (or probability) distributions.
[0028] Optionally, each probability distribution is determined based on data measured from real-world networks.
[0029] The module for selecting the one or more second values may include a module for selecting the one or more second values based on the one or more selected first values.
[0030] In some embodiments, modules are also provided for storing each probability distribution associated with corresponding parameters of the network configuration and each probability distribution associated with corresponding parameters of the service to be performed. In some embodiments, modules are also provided for accessing each probability distribution associated with corresponding parameters of the network configuration and for accessing each probability distribution associated with corresponding parameters of the service to be performed from a remote storage device.
[0031] The module for determining may further include: a module for injecting the data volume into the network in each of the one or more time intervals; and a module for processing the flow of each data volume injected through the network in each corresponding time interval to determine the flow of the data volume.
[0032] In some implementations, the network includes: one or more home (or user) gateways; the one or more terminal devices located after the one or more home (or user) gateways and connected to the one or more home (or user) gateways via the one or more home networks; an access node providing access to one or more service providers configured to provide the services to be performed on the one or more terminal devices; and an access network connecting the access node and the one or more home gateways. The access node may be referred to as a central node. The access network may be referred to as an access link.
[0033] In some embodiments, the module for selecting the one or more first values may include a module for selecting a plurality of first values. Additionally or alternatively, in some embodiments, the module for selecting the one or more second values may include a module for selecting a plurality of second values.
[0034] The selected one or more first values can instantiate the network configuration. The network configuration may include the network topology and one or more underlying network technologies, as well as any other network characteristics. The network topology includes the nodes (or cells) within the network and their interconnections. Each node in the network can be associated with a network value. The one or more first values can be associated with the network value. The network value may include one or more of the following: the node's capacity, the node's buffer size, and the node's drop policy.
[0035] The one or more second values may be associated with a service value. The service value may include one or more of the following: scheduling for each service, the type (or category) of the service, and the latency or round-trip time for each service. The type / category of the service may determine the pattern of data packets to be transmitted during the service. In some examples, the service value may also include the amount of data to be transmitted. In other examples, the amount of data to be transmitted over time may be determined based on the type, scheduling, and / or latency of the service (without needing to define the data amount through the service value).
[0036] This document also discloses a system comprising: a module for receiving one or more first values, each first value representing a parameter of a network configuration, the network including one or more terminal devices; a module for receiving one or more second values, each second value representing a parameter of a service to be performed on the one or more terminal devices, each service being associated with transmitting one or more data packets through the network; a module for determining a time series based on the one or more second values, the time series representing the amount of data to be injected into the network at each of one or more time intervals; a module for determining a flow of the amount of data through the network based on the one or more first values and the time series; and a module for outputting information indicating the flow of the amount of data.
[0037] Optionally, the system further includes a module for selecting each first value from a probability distribution associated with the parameters configured in the network. Optionally, the system further includes a module for selecting each second value from a probability distribution associated with the parameters of the service. In some examples, the selected value may be received from local or remote storage. In some examples, the selected value may be received in response to user input or user selection. The user input or user selection may be an input / selection of a seed value for automatically instantiating the digital twin.
[0038] This document also discloses a system comprising: a module for selecting a plurality of first values, each first value being selected from a corresponding probability distribution and representing a corresponding parameter of a network configuration, the network including one or more terminal devices; a module for selecting a plurality of second values, each second value being selected from a corresponding probability distribution and representing a corresponding parameter of a service to be performed on the one or more terminal devices, each service being associated with the transmission of one or more data packets through the network; a module for determining a time series based on the plurality of second values, the time series representing the amount of data to be injected into the network at each of one or more time intervals; a module for determining the flow of the data amount through the network based on the plurality of first values and the time series; and a module for outputting information indicating the flow of the data amount. Attached Figure Description
[0039] Exemplary embodiments will now be described by way of non-limiting example with reference to the accompanying drawings, in which:
[0040] Figure 1(a) is a schematic diagram of an example network to be implemented in digital twin, and Figure 1(b) is a schematic diagram of the pipeline-level representation of the network in Figure 1(a);
[0041] Figure 2 (a) is the first example probability distribution. Figure 2 (b) is the second example probability distribution;
[0042] Figure 3(a) is a schematic representation of an example operation of a system implementing a digital twin, Figure 3(b) is a schematic representation of an example pipeline-level operation of a digital twin, Figure 3(c) is a schematic representation of another example pipeline-level operation of a digital twin, and Figure 3(d) is a schematic block diagram of an example system implementing a digital twin (e.g., shown in Figure 3(a)).
[0043] Figure 4 (a) shows an example output of downstream traffic from a system implementing digital twins, at a first granularity, and Figure 4 (b) shows an example output of the upstream flow from the system using this first granularity;
[0044] Figure 5 (a) shows an example output of downstream traffic from a system implementing digital twins, at a second granularity (fineer than the first granularity), and Figure 5 (b) shows an example output of the upstream flow from the system using this first granularity;
[0045] Figure 6 (a) shows a schematic representation of an example network replicated in a digital twin, and Figure 6 (b) through 6(e) show example outputs of this digital twin;
[0046] Figure 7 (a) shows a schematic representation of another example network replicated in a digital twin, and Figure 7 (b) through 7(d) show example outputs from this digital twin;
[0047] Figure 8 This is an example process diagram illustrating the processing operations based on some examples of implementing digital twins of networks;
[0048] Figure 9 A schematic diagram of an apparatus that can be configured according to one or more exemplary embodiments of the process described herein; and
[0049] Figure 10 It is a planar diagram of a non-transient medium.
[0050] In the specification and drawings, similar reference numerals always refer to similar units. Detailed Implementation
[0051] Existing simulators are used to simulate network traffic response under given network topology characteristics (and / or associated pipe capacity). These simulators are based on, for example, data switching mechanisms. However, the inventors recognized that simply simulating packet switching in the pipes of the access network is insufficient when attempting to provide a digital twin of a network to facilitate the assessment of capacity and / or bandwidth requirements (and / or optionally find the optimal configuration of the network). Instead, the effects of different network device behaviors and mechanisms, service provision, the presence and behavior of devices on home / terminal networks and user interactions with these devices, capacity sharing among users, technical protocols, and so on, need to be considered.
[0052] The digital twins described in this paper aim to replicate the characteristics of an entire network, thereby replicating or generating data that closely approximates the expected responses of a real network. Specifically, digital twins aim to replicate different technologies, media, protocols, home / end-device devices, and user habits in the digital world to realistically represent real-world network behavior. The resulting digital twins can be used to evaluate different network configurations and scenarios, perform network tests, and generate realistic datasets (which can, for example, be used to train machine learning models).
[0053] In particular, the method described in this paper allows for the reproduction of real-world network conditions by generating a large number of simulations (each representing a different network scenario), thereby obtaining a variety of different situations to mimic what might be encountered in real-world networks. This method can also be used to achieve a single simulation of a specific network scenario. These results can be achieved without requiring users to define how traffic is generated on one side of the network and how it is consumed on the other side, or to specify links or connections between different network units. In this way, networks can be replicated or simulated without requiring users to input the various technical and technological aspects of the network. Therefore, improved copies of real-world networks can be provided in a simple and efficient manner using digital twins.
[0054] Referring to Figure 1(a), a network digital twin of the access network and home network is depicted. It replicates in the digital world the behavior of access network technologies (e.g., access nodes or optical line terminals (OLTs), optical network terminals (ONTs), etc.), the home environment (including home / user gateways, home / terminal network protocols, and the presence and behavior of devices connected to the home or terminal network), and the interactions of different users with network 100. These units can also be implemented at various scales to mimic the responses of real-world access networks.
[0055] As shown in Figure 1(a), network 100 includes access node 102, access network 104, one or more home or user gateways 108, one or more terminal or (home) networks 106 (which may be represented by end-user profiles registered on access node 102, for example, as shown in Figure 1(b)), and one or more terminal devices 110. Home gateway 108 connects the home network to access network 104. Home network 106 connects the terminal devices 110 within the home to their respective home gateways 108. In this art, access node 102 may also be referred to as a central node. Access network 104 may also be referred to as an access link.
[0056] It is possible to replicate various access network technologies, including but not limited to: Asymmetric Digital Subscriber Line (ADSL), Very High Speed Digital Subscriber Line (VDSL), GFast (Digital Subscriber Line Protocol standard), Gigabit Passive Optical Network (GPON), 10 Gigabit Passive Optical Network (or XGPON), and 10 Gigabit Symmetric Passive Optical Network (XGSPON). Any suitable gateway technology can be replicated. Appropriate access nodes for a given access network technology can be replicated.
[0057] As shown in Figure 1(a), each home network 106 is connected to a home or user gateway 108, which connects the home network to the access network 104. In some examples, the various home network technologies and protocols, including wireless and Ethernet connections, may be replicated. Various network subscriptions (e.g., bandwidth, upload and / or download speeds) for each home network, as well as user profiles or service level agreements (SLAs) for each home network 106, may also be replicated. In other examples, such as in the case of a shared access network like PON technology, the home network is not replicated; instead, the behavior of one or more end devices 110 behind the gateway is replicated without referencing the home network behavior.
[0058] As shown in Figure 1(a), one or more terminal devices 110 (home devices, device 1... device M) are connected to the terminal (or home) network 106 (i.e., one or more terminal devices are located after one or more home / user gateways 108). These terminal devices 110-1 to 110-M can include any suitable home user devices, including but not limited to: desktop computers, laptops, tablets or pads, digital enhanced cordless telecommunications (DECT) devices (e.g., landline phones), smartphones, Internet of Things (IoT) sensors or devices, tuners or set-top boxes, and televisions (TVs).
[0059] Network configuration can include network topology and underlying network technology, as well as any other network characteristics. Network topology includes the nodes (or cells) within the network and their interconnections. To replicate the network configuration in a digital twin, each cell of the network must be digitally replicated. Access network 104 is associated with access network technology (GPON, XGPON, etc.). This access network technology defines the transmission capacity (bit rate, throughput, etc.) used for the network. Access network 104 connects to access node 102 through an interface aligned (or adapted) to the network access technology. Access node 102 implements buffering and drop policies for access network 104.
[0060] Access network 104 and access node 102 digitally replicate capacity in one or more of the following aspects: defined / allocated capacity, buffer size, and drop policies, as discussed in more detail below with reference to Figure 1(b). At access node 102, mechanisms dependent on service level agreements (SLAs) and technology-dependent mechanisms have been implemented to match the way network 100 provides / shares capacity over multi-user links. Similarly, the capacity defined / allocated by the terminal or home network 106, and through the home gateway 108 interface (Wi-Fi 802.11a / g / n / ac / ax / … or Ethernet 10 / 100 / 1000 / … etc.), can also be digitally replicated within the capacity chain.
[0061] All these network units collectively define the capacity chain for each user and for the total shared capacity of the network. In addition to capacity, the network is also defined by the devices involved (110) for a given communication chain (e.g., a given internet subscription). The bandwidth requirements and responses of the various devices have been considered and digitally replicated, as discussed in more detail below.
[0062] To determine the capacity of nodes / cells in network 100, as well as the buffer size and drop policy for each node, a network configuration must first be selected. Values that can instantiate a network configuration are chosen. Each value represents the value of a corresponding parameter in the network configuration. Therefore, the entire network can be defined by one or more parameters that represent different aspects of the configuration of network 100 to be replicated. The network configuration may include network topology, one or more network technologies, and any other configuration parameters representing the physical network and its implementation. The combination of capacity, buffer size, and drop policy for each node in network 100 is referred to herein as a capacity chain.
[0063] The access network technology is an example of the configuration parameters of the network 100 to be replicated by the digital twin. The network configuration parameters may also include one or more of the types, connections, or characteristics of any network nodes. Examples of characteristics for each network node include: capacity, queue size, and technology type. Other examples of configuration parameters for the network to be replicated by the digital twin include, but are not limited to: number of gateways, number of terminal devices, terminal device types, number of local area network (LAN) ports, network subscription type for home networks, etc. Any suitable parameters can be used to represent the network configuration (and therefore, its physical and technical network characteristics).
[0064] These values are selected from the corresponding probability distributions. Each parameter of the network configuration can be associated with one or more specific probability (or probability) distributions. See below for further details. Figure 2 (a) and (b) discuss in more detail how to determine or select the probability values of these different parameters to provide the digital twin described herein. Once the values of the parameters are selected from these distributions, the capacity associated with one or more of the selected values, the buffer size associated with one or more of the selected values, and / or the discarding policy associated with one or more of the selected values can be used to generate or define the behavior of the digital twin, as described in more detail below with reference to Figure 1(b).
[0065] One or more services 112 can run on the terminal device, and each service can be associated with a service type and service schedule. Service types (or service categories) can include one or more of the following: audio, browsing, file transfer, gaming, Internet of Things (IoT), Internet Protocol Television (IPTV, also known as broadband television), video conferencing, Voice over Internet Protocol (VoIP), Internet Video Transport (VotT, also known as over-the-top content or OTT content), and / or iperf (a network performance measurement tool that can create data streams to measure one-way or two-way throughput between two ends), used for Transmission Control Protocol (TCP) or User Datagram Protocol (UDP), or for other protocols built on top of TCP or UDP (such as specific protocols implemented by VotT providers, sometimes proprietary protocols)—the methods described herein can embed traffic patterns and behaviors specific to any of these protocols. These services can be provided by a service server 114 located at the other end of access network 104 (from home network 106). Scheduling can also be associated with each service and the amount of data transmitted for uploading / downloading.
[0066] To replicate the behavior of network 100 in a digital twin, services are defined and implemented based on their required specifications and the associated traffic patterns that may arise when running on terminal device 110. Specifically, each service may be associated with one or more of the following: bit count (in each time interval), scheduling, and latency or round-trip time.
[0067] For example, a file transfer service (download) can be defined in terms of the amount of bytes to be exchanged (primarily downstream) and by attempting to match the maximum available bandwidth until the service ends (after the entire file is downloaded). In contrast, internet streaming video services (such as Netflix) require buffered byte transfers using the maximum available bandwidth at service initialization, but then periodically exchange and rebuffer data thereafter. All of these mechanisms are affected by network capacity chains and are also limited by the home network 106, network interfaces, and the devices 110 that initiate each service 112. Service servers 114 and traffic conditioning mechanisms (such as Enhanced Functional Port (ECP) mechanisms, PON aggregation mechanisms, etc.) also affect the defined / allocated capacity of network 100.
[0068] To determine the bit count, scheduling, and / or latency of each service, multiple services 112 to be run on the terminal device 110 can be selected first. Services 112 as a whole can be defined by one or more parameters.
[0069] Service type is one example of the parameters of the service to be run on the terminal device. Other examples of parameters for the service to be run on the terminal device include, but are not limited to: the number of services per terminal device, service scheduling, and service characteristics. Service characteristics may include one or more of the following: bit rate, amount of data to be transmitted, buffering policy, Quality of Service (QoS) requirements, Quality of Experience (QoE) requirements, Service Level Agreement (SLA), or user interaction with the service. Any suitable parameters can be used to represent the service on the terminal device (and therefore the expected data traffic on the network).
[0070] These values are selected from corresponding probability distributions. Each parameter of the service to be executed on one or more terminal devices can be associated with one or more specific probability (or probability) distributions. See below for reference. Figure 2 (a) and (b) discuss in more detail how to determine or select the probability values of these different service parameters. Once the parameter values are selected from these distributions, services (including associated service requirements and associated traffic pattern requirements) can be generated or defined using the amount of bits associated with one or more of the selected values, the scheduling associated with one or more of the selected values, and / or the latency associated with one or more of the selected values, and thus the behavior of the digital twin is defined, as described in more detail below with reference to Figure 1(b).
[0071] Figure 1(b) illustrates a network 100 represented by multiple communication pipes (or comPipes). Communication pipes are resources that allow the production and consumption of objects and enable the handling of shared resource availability. There are three levels of pipes:
[0072] 1. GW / packet pipes, which are pipes corresponding to an end-user profile (or SLA) and also to the bandwidth promised to the home gateway for that end-user (e.g., representing the pipe connecting the gateway to the end device and gateway 108 in home network 106). These pipes are "per-user" pipes.
[0073] 2. Access network pipe, corresponding to access network or access link 104 and access node 102. This pipe is a network-level pipe shared by multiple users.
[0074] 3. Home device conduits, corresponding to each device 110 connected to the home network 106 (e.g., conduits corresponding to terminal devices 110). These conduits are "per device" conduits.
[0075] Each communication channel is characterized by one or more of the following: a capacity to transmit a given amount of bits per sampling interval or iteration (this capacity is a configurable parameter that is configured based on the value selected for the parameters configured for the network); a buffer (or queue) of a given size (the queue size is a configurable parameter that is configured based on the value selected for the parameters configured for the network); and / or a drop policy.
[0076] Based on the selected values of the parameters used to represent the network configuration, a capacity chain can be defined for each node / cell in the network, including capacity, buffer size, and / or drop policy, thus defining a given network 100. Dropping traffic from a node occurs when the amount of traffic a network cell needs to process exceeds the transmission capacity plus the buffer size. Dropped traffic may also include data that was not transmitted due to insufficient capacity on the chain. The drop policy defines the option to reject or retransmit this unprocessed traffic. A drop bucket is an implementation concept used to store data traffic / data volume that will be retransmitted and to calculate the final amount of data to be dropped. As shown in Figure 1(b), this overall capacity chain (capacity, buffer size, drop policy) may be referred to herein as the "network value".
[0077] Network values are the outputs of the configuration or instantiation of network 100 (where the selected parameter values are the inputs to the configuration of network 100). Thus, the parameter values selected from the probability distribution define the network topology, including the selection of nodes and their interconnections, as well as the technologies used by the network. The network values associated with these parameter values define the final capacity of the network, including nodes / technologies instantiated using parameter values selected from the probability distribution. These network values also serve as inputs for subsequent operations of the digital twin.
[0078] The communication channel generates and consumes bit packets, which constitute a sequence of bits resulting from multiple services 112 performed on terminal device 110, where each packet represents a certain number of bits for a given service within a specific iteration or time interval. In other words, each service traffic sequence is segmented into bit packets, with one packet corresponding to a (potentially different) number of bits per sampling time interval or iteration. These packets may be characterized by: size or traffic volume; optionally, service priority; and / or a device ID linking service 112 to a specific device 110 connected to a specific home gateway 108. In some examples, the system implements a service prioritization mechanism, just as it would on a real IP network. Packets for higher-priority services are processed before packets for lower-priority services. For the same priority, the system may implement mechanisms such as fair queuing and round-robin. Priority can be another example of a parameter representing a service.
[0079] Based on the selected values of the parameters representing the services, a bit quantity, schedule (or timetable), and / or latency or round-trip time can be defined for each service (and optionally, service priority). By organizing all services, the system described herein can determine or define the traffic schedule for each service. As shown in Figure 1(b), this traffic schedule (bit quantity, schedule, latency) may be referred to herein as a "service value". The service value is the output for determining which services are running on the device (where the selected parameter values are the inputs for determining which services 112 are running on device 110). The service value is also the input for subsequent operations of the digital twin.
[0080] In particular, service values can be used to determine a time series representing the amount of data to be injected into the network at each of one or more time intervals during the operation of the digital twin, as discussed in more detail below with reference to Figure 3(a). This time series can represent the total or aggregated amount of data (in bytes) to be transmitted during the execution of service 112, and allows for a way of representing data flows without copying additional structured information (e.g., header information) of individual packets themselves. The series can represent the amount of data at any suitable time interval, where the time interval can depend on the granularity of the simulation.
[0081] Then, network values and time series can be used to determine the flow of data volume through network 100 (represented by packets on the right side of Figure 1(b), where packets are associated with each home or terminal device conduit), as discussed in more detail below with reference to Figure 3(a). This flow of data volume through the communication conduits of network 100 replicates the real-world behavior of packets on network 100 in the digital twin.
[0082] While network behavior can ultimately be defined in terms of capacity chaining and traffic scheduling (i.e., through the network and service values shown in Figure 1(b) and discussed above), accurately determining these network and service values so that network 100 can be replicated in a digital twin can be challenging. This is because each user's behavior may differ in terms of their connection usage, time spent at home, and habits of using service 112 on terminal device 110. Furthermore, in terms of network configuration, not every home environment includes the same number and / or type of devices 110; not every gateway 108 exhibits the same interface; not all access nodes 102 have the same users or user configurations, and so on. Moreover, regarding the use of service 112, each service may have different data usage patterns and different bit rates, thus requiring different data transmissions from server 114.
[0083] To replicate these diverse degrees of freedom, a statistical or probability distribution 200 is determined for each parameter of the configuration of the network 100 to be replicated and for each parameter representing the service 112 to be executed on the terminal device 110. Each probability distribution can be determined based on measured real-world data so that these distributions fit the real-world behavior of the network 100. The probability distributions can be parameterized to fit the corresponding real-world characteristics of different network configurations (topology, technology, etc.) and service uses. Figure 2 Examples of these probability distributions are shown in (a) and (b). These distributions provide a statistical representation of the possible values for each parameter, from which a value for each parameter can be selected based on probability. Then, as described above, the selected value for each parameter will be used to determine the network value and service value.
[0084] Figure 2 (a) An example distribution 200-1 is provided, representing the number of optical network terminals (ONTs) in a given implementation of the passive optical network (PON) access network 104. This distribution is a gamma distribution parameterized to fit the measurement data. The distribution shown in this example is the one used by the system when the network distribution was previously determined to be GPON. The network technology itself is determined based on its own distribution; in this example, the network technology has a normal distribution set to 80% GPON, 15% XGPON, and 5% XGSPON.
[0085] Figure 2 (b) provides another example distribution 200-2 for parameters representing the duration of intermediate peaks for the Internet transmission video service type used for service 112. This distribution is a gamma distribution parameterized to fit the measurement data.
[0086] It should be understood that these distributions are merely examples to illustrate the basic principles described herein. The probability distributions used by the system can be adjusted and modified over time, allowing the digital twin to evolve continuously. By adjusting the probability distributions, accuracy relative to real data can be improved, enabling the digital twin to reflect constantly changing real-world conditions.
[0087] A probability distribution 200 can be determined for each parameter to be represented in the digital twin. Therefore, multiple distributions 200-1…200-N can be determined and stored for use when operating the digital twin of network 100. Distribution 200 can be stored in / as part of the digital twin and / or accessed through the digital twin in other ways (e.g., from a remote storage device, such as a cloud server). When instantiating or configuring a scenario of the digital twin, or multiple scenarios together (e.g., a traffic generation activity), a specific value for each parameter can be selected based on the corresponding probability distribution for each parameter. These distributions can eliminate the need for the digital twin's users to have expertise in, for example, a specific type of network technology or knowledge of end-user behavior. Instead, multiple different types of network and user / service behaviors can be replicated or simulated probabilistically.
[0088] It is understandable that probability distributions can be independent of each other, or at least some distributions can depend on the choice of certain values for a given parameter. For example, Figure 2 The distribution shown in (b) can depend on the service type being Video over Internet Transport (VOTT); in response to selecting a value for the service type parameter indicating VOTT, the following can be used: Figure 2 (b) shows the distribution of intermediate peaks to select the parameter for the duration of the intermediate peaks. However, if a different service type parameter (e.g., VoIP) is selected, a different distribution will be used to select subsequent parameters.
[0089] This statistical approach allows for the replication of end-user behavior and activities, meaning that real-world home environments and scheduling can be digitally acquired for various types of users. For example, for IPTV users, a home environment containing devices including a television, laptop, and smartphone can be set up. In contrast, for non-IPTV users, the instantiation probability of a television might be higher than that of an internet-transmitted video service. These are stochastic processes generated under statistical guidance to match the distribution and responses of real-world networks (determined through measurements of real-world networks).
[0090] The use of these probability distributions 200 in initializing or configuring digital twins will now be discussed with reference to Figure 3(a).
[0091] Referring to Figure 3(a), a system for implementing digital twins of networks is described. As described herein, the system for implementing digital twins is capable of executing specific and macroscopic network scenarios (i.e., simulating networks from the single-user level to the entire network level), generating on-demand data (including large datasets containing various network scenarios and / or labeled datasets that can be used to train machine learning models), and matching real-world network conditions. The system also exhibits consistency, meaning that all its components, input data, and intermediate data (data generated by the components during the operation of the digital twin) interact in a logically / consistent manner, resulting in outputs that mimic what might occur in real-world networks. The digital twin can be used to output information indicating packet flows through the network (e.g., network traffic patterns).
[0092] The system includes a module 350 for selecting one or more first values. Each first value represents a corresponding parameter of the network configuration and is selected from a probability distribution 200A associated with the corresponding parameter of the network configuration. The system also includes a module 352 for selecting one or more second values. Each second value represents a corresponding parameter of a service 112 to be executed on one or more terminal devices 110 and is selected from a probability distribution 200B associated with the corresponding parameter of the service to be executed.
[0093] Probability distribution 200B may depend on the number and type of devices 110 represented by the first value selected by module 352 from probability distribution 200A. In other words, the selection of one or more second values depends on one or more of the selected first values. In this regard, the selection of the first value is shown as part of step 1, while the selection of the second value is shown as part of step 2. This dependency is partly due to the impact of network capacity on service demand and traffic patterns discussed above.
[0094] More specifically, in step 1, module 350 of the system selects a first value for determining or configuring the network configuration. These parameters may include the number and type of network units / nodes and their characteristics (pipeline capacity, queue / buffer size), end-user profiles and their associated characteristics (committed capacity or SLA, single buffer size), etc., as described above. Any suitable parameters may be included. The selected one or more first values may instantiate the network configuration and be associated with the network value. In step 2, module 352 selects a second value for determining or configuring service scheduling, i.e., which services 112 will be executed on which devices 110, when services start and stop, etc. One or more second values are associated with the service value. Any suitable parameters may be included. The second value for service scheduling is selected based on the number and type of devices 110, which are part of the configuration (selected or defined in step 1). Steps 1 and 2 involve selecting a value for each parameter from a dedicated probability distribution 200 for each parameter, as referenced above. Figure 2 As described in (a) and (b).
[0095] In some implementations, the selection of values in steps 1 and 2 (via modules 350 and 352) is fully automated. For example, module 350 for selecting one or more first values includes a module 350 for automatically selecting one or more first values. Additionally or alternatively, module 352 for selecting one or more second values includes a module 352 for automatically selecting one or more second values.
[0096] In some implementations, the module for selecting one or more first values includes a module for selecting multiple first values. Additionally or alternatively, the module for selecting one or more second values includes a module for selecting multiple second values. Optionally, if the selection is fully automated, multiple first values and / or second values can be automatically selected based on probability distribution 200.
[0097] This automatic selection can be viewed as a standalone mode of the digital twin, where the responses of the entire network can be replicated because all network characteristics (access, home, devices, services, and user behavior) are captured in the digital twin. In some specific implementations of the standalone mode, automatic selection is performed based on a seed value, which can be passed to the system implementing the digital twin, allowing the system to automatically make all decisions related to network configuration, activity scheduling, and traffic sequence generation based on that seed value. The seed value can be provided as part of user input or in response to user input. In particular, the system can randomly select first and second values (but according to and subject to a corresponding probability distribution 200 associated with various dynamic parameters of the system) to initialize the network. This is a probabilistic or statistical approach, where different seeds will lead to completely different (but still realistic) final outputs.
[0098] For ease of selection, the system may also include a module 340 for storing each probability distribution 200A associated with corresponding parameters of the network configuration and each probability distribution 200B associated with corresponding parameters of the service to be performed (see Figure 3(d) discussed below). In other words, the probability distributions 200 may be stored locally as part of the system implementing the digital twin. Additionally or alternatively, the system may also include a module for accessing each probability distribution 200A associated with corresponding parameters of the network configuration and each probability distribution 200B associated with corresponding parameters of the service to be performed; the probability distributions 200 may be stored outside the system implementing the digital twin, for example, on a remote storage device. Thus, selection may include receiving one or more selected values from said remote storage. As described above, such selection and subsequent reception from the remote storage device can be performed automatically.
[0099] In some examples, at least a portion of the first and second values is received in response to user input. As mentioned above, user input may include a seed value. However, in other examples, the user can configure or enforce one or more parameters or settings of the digital twin by providing specific first and / or second values for certain parameters. As mentioned above, other values of the first and second values can be automatically selected. For example, the user can configure a subset of parameters, and then the values of the remaining parameters can be automatically selected at least in part based on the user configuration of that subset (e.g., due to the dependence of certain probability distributions on the values of the parameters as described above). Thus, the user can provide their own settings for one or more aspects of network configuration and / or activity scheduling by providing user input for values of one or more parameters.
[0100] This can be viewed as an on-demand model of digital twins, where the response of the entire network, initialized by user input values for certain parameters, can be determined. Upon receiving these additional values, the system can bypass the selection of one or more first and / or second values (i.e., it will not select values for parameters specified by the user).
[0101] The system also includes a module 354 for determining a time series 360 based on one or more selected second values, the time series representing the amount of data to be injected into the network at each of the one or more time intervals. Thus, the time series represents the amount of data per unit time (wherein the unit time can be selected based on the granularity of the simulation). Optionally, when a user provides one or more second values to be used with automatically selected second values, the determination of the time series 360 is also based on these other provided second values.
[0102] To determine time series 360, module 354 uses the second value selected in step 2 to determine service scheduling, i.e., which services will be executed on which devices, when they will start and stop, etc. Specifically, module 354 can use the second value to determine service values, where the service values include one or more of the following: the amount of bits associated with one or more selected second values, the scheduling associated with one or more selected second values, and the latency associated with one or more selected second values. These service values can be used to schedule services.
[0103] Once a service is scheduled, module 354 can use the service value to define the traffic for each activity, i.e., define the upstream and downstream data streams or data volumes, which consist of bit packets to be submitted to the network in each sampling interval. In other words, module 354 uses the service value to organize all the different packets for each service 112 to run on device 100 of network 100 and determines the total number of bit packets passing through the network in each time interval. Once the packets are organized / aggregated, the total amount of data (in bytes) to be injected into the network in each time interval can be determined; this is time series 360. Therefore, time series 360 can be viewed as a traffic sequence representing the total amount of data to be input into the network per unit time, without replicating packet-specific information (such as packet structure information or other packet content). The time series can be determined based on one or more of the following: the number of bits associated with one or more selected second values, the scheduling associated with one or more selected second values, and the latency associated with one or more selected second values.
[0104] When creating these traffic sequences, the system considers the type and characteristics of each service 112 to generate traffic sequences that include patterns and effects specific to that service type. These characteristics may include the actual range of the nominal bit rate, how the total amount of data to be transmitted is distributed over the duration of service 112, initial and intermediate buffers, quality control, buffering strategies, intermediate peaks, possible end-user interactions (pause, screen sharing, video on / off, etc.), etc. In this way, module 354 of the system implementing the digital twin controls the management of service and traffic usage, including determining the scheduling of service 112 and mimicking the decisions made by end users (humans and / or devices 110) when performing service 112. Module 354 also controls the generation of traffic sequences corresponding to the type of service 112 (file transfer, VOTT, video conferencing, etc.), their scheduling, and the characteristics of the service type.
[0105] The system also includes a module 356 for determining the flow of data volume (e.g., data volume stream) through network 100 based on one or more selected first values and time series 360. Optionally, when a user provides one or more first values in conjunction with automatically selected first values, determining the flow of data volume through the network is also based on these other provided first values. In particular, module 356 can use the first values to determine network values, wherein the network values include one or more of the following: capacity associated with one or more selected first values, buffer size associated with one or more selected first values, and drop policy associated with one or more selected first values. In other words, each network node, connection, and network technology instantiated using the selected first values has associated network values (node capacity, buffer size for the node, drop policy for the node). These network values, along with the time series 360 of the packets, can then be used to determine the flow of data volume through the entire network 100 by defining an overall capacity chain.
[0106] In some examples, module 356 for determining the flow of data through network 100 includes a module for configuring a simulation of the network based on one or more selected first values. In some examples, the module for configuring the simulation of the network based on one or more selected first values also includes a module for configuring the simulation of the network based on network values associated with the selected first values. In this way, the digital twin of the network is configured based on parameters representing the network configuration.
[0107] As part of this configuration, module 356 uses the first value selected in step 1 and determines the configuration of network 100 represented by network values: the number and type of network cells / nodes and their characteristics (pipeline capacity, queue / buffer size), end-user profiles and their associated characteristics (committed capacity or SLA, single buffer size), etc. In other words, module 356 controls the configuration and management of network 100, including determining the network configuration based on the network values associated with the selected first value, through the generation, interconnection, and sizing of multiple nodes of network 100 and their key characteristics (capacity, buffer size, etc.).
[0108] Module 356 may also include a module for injecting the time series 360 of data packets into a configured simulation of the network. In this example, the time series 360 of data packets is configured to be injected into access node 102 (as shown in step 3). Module 356 also includes a module for determining the flow of the configured simulated data volume through the network. Once services 112 are scheduled and the flow sequence is defined as the time series 360 of data packets, module 356 can submit the amount of data in bit packets representing each service or activity sequence to access node 102 of the simulated network 100 based on their scheduling. Based on the time series 360, packets are injected at sampling intervals. Module 356 determines the flow of injected data by managing the transmission of the amount of data through all consecutive nodes and pipes constituting the network, while taking into account the mixture and temporal evolution conditions of services 112 at each sampling interval (e.g., service priority, congestion of one or more nodes, queuing delay, and policies or protocols that mean that whole or part of the data packets are delayed, dropped, or retransmitted). In this way, the overall flow of packets can be simulated or replicated without modeling the behavior of individual packets.
[0109] Examples of this injection and data flow through the network are further discussed with reference to Figures 3(b) and (c).
[0110] Figure 3(b) illustrates the communication pipeline of Figure 1(b) and the time series 360 of the data packets (as defined above). At each sampling interval n (from 1 to T, where T is the duration of the simulation or test scenario), an aggregated amount of data representing all packets of each service scheduled in that iteration is submitted to the communication pipeline. In this example, the time series is submitted to the pipeline corresponding to home gateway 108, behind which service 112 runs. The pipeline corresponding to the home (i.e., the end user) collects the aggregation of all data from the home network and all services running on the end devices behind home gateway 108. After any possible drops at previous pipeline levels, the pipeline corresponding to access link 104 collects the aggregation of all data from all services 112 of all end-user devices 110 connected to that link. When the data flow arrives at access link 104, the time series 360 is injected into access node 102. In some examples, this could be a measurement of the amount of data flowing out of the access node (e.g., a measurement or count of the amount of traffic transmitted or dropped downstream and / or upstream at access node 102) output by a digital twin. In other examples, the output (representing a measurement or count of the amount of traffic) can be obtained at any appropriate point on the network.
[0111] If, within a given iteration (i.e., time interval), the amount of aggregated data entering a given pipeline exceeds its capacity (defined by the network value), the excess data will enter the corresponding queue or buffer (defined by the network value) and will be processed in the next iteration (i.e., re-injected into the pipeline's inlet). If the queue is full, the remaining data will enter the corresponding drop bucket (defined by the network value) and will certainly be lost. Thus, the capacity chain defined by the network value defines the data flow between pipelines (which is injected according to the input time series 360). In other words, packet transmission, queuing, and dropping can be replicated within the data volume without requiring separate modeling of the behavior of each packet.
[0112] The injection and subsequent processing of data across multiple nodes (pipelines) of network 100 can be controlled by discrete event processors (which handle event processing and time management in real-time or non-real-time). In other words, the discrete event processor iteratively manages the next volume of data injected at the system's entry point for all active services. The discrete event processor is an example of module 356, but any suitable module can be used.
[0113] In other implementations, as illustrated with reference to Figure 3(c), the system implementing the digital twin can implement or embed an optional service rate control mechanism. The service rate control mechanism can provide rate scaling and adaptive pipeline capacity, and is an optional process that can be activated on top of the core communication pipeline structure and event handling described in Figure 3(b). In the example implementation of Figure 3(c), instead of injecting data volume into the pipeline that regulates the traffic volume of each home gateway, the data volume is first submitted to a pipeline dedicated to each service (shown as the "rate control pipeline"). These pipelines are used for traffic scaling, i.e., intermediate buffering when the service speed slows down, and for re-injecting bits that might be dropped later in the link, thereby replicating the retransmission of the controlled transport protocol. Only rate-controlled services are injected into the rate control pipeline; non-rate-controlled services are injected into the gateway pipeline, as shown in Figure 3(b).
[0114] The service rate control mechanism assesses at each iteration or interval (and for each ongoing controlled service) whether a rate adjustment is needed to address congestion in the previous iteration. For example, it compares the amount of data buffered in the queue and / or the amount of data dropped in all subsequent pipelines to the expected transmission volume of the service. If, in a given iteration, the full amount of data committed by a given controlled service has not yet been fully transmitted to the home gateway (e.g., when service experiences latency and / or packet drop), the service can be slowed down, and its bit rate reduced in the next iteration. The rate reduction can be determined based on the amount of data buffered and / or dropped (in bytes), and / or by applying a conditioning mechanism similar to the actual control mechanism. In highly congested networks, the rate reduction may be larger than in less congested networks. In other words, the magnitude of the rate reduction may depend on the level of congestion within the network. Optionally, in some examples, in a given iteration, portions of the data volume for a given service that have been dropped at any pipeline layer are re-injected into the correct service control pipeline, resulting in the retransmission of this "lost" data.
[0115] Referring further to Figure 3(a), module 356 therefore controls the time series 360 of network 100 (representing the amount of data aggregated on the traffic sequence of each of the multiple services 112 executed on device 110) through the processing of multiple nodes (pipelines) of network 100, taking into account the mix of services 112 and network capacity (queuing, bottlenecks, packet loss / dropping, etc.), and determines the final data transmission for each service 112 on a per-time interval basis, and can mimic the workings of broadband communication networks and replicate all their potential effects (congestion, latency, service control, packet loss, etc.).
[0116] The system also includes a module 358 for outputting information indicating the flow of data volume. The specific information to be output and the granularity level of that output can be determined or selected by the user, or it can be automatically configured by module 358 based on the first and second input values. The granularity of the output can be the same as the granularity of the input, for example, representing the time interval of the injected data volume flow, thus representing the time interval of the measurement or counting of the output data volume flow. Step 4 illustrates an example of the output information.
[0117] Module 358 can be configured to output information to a remote system or device, such as by sending information via any suitable wireless or wired connection. In some examples, module 358 may provide or report output information to a user that details the flow of data volume. Additionally or alternatively, module 358 may be configured to present information to a user, for example, by displaying the information on a monitor. By outputting information indicating the flow of data volume, module 358 helps to report network usage on multiple different network nodes or pipes. For example, output may be provided on one, more, or each communication pipe in Figure 3(b) or 3(c).
[0118] The information output by module 358 may include any suitable information indicating the amount of data flowing per time interval. In one implementation, the output information includes one or more data traffic patterns. Traffic data patterns may indicate the amount of data sent, queued, and / or dropped per time interval. For example, the output information may include data detailing the amount of data sent, queued, or dropped for each service type in each sampling interval, and / or packet latency, and / or information regarding the impact on the transmission of each service. In some examples, the output information may include information about the mutual influence of services 112, including services in the same and reverse flows (i.e., upstream affects downstream packets, and vice versa). In some examples, the output information may include information about initial buffer rates, buffer periods, etc. In some specific examples, a complete traffic dataset (number of bytes exchanged upstream / downstream, number of bytes dropped upstream / downstream, number of bytes queued upstream / downstream, etc.) may be output. Optionally, the output information may be output at the individual user level, at the network level, or at different granularity levels (as needed).
[0119] In some examples, information indicating the flow of data volume may include one or more of the following: data volume sent, data volume queued, data volume dropped, quality of service indication, service slowdown indication, service latency indication, data queue status, and impact eligibility. Impact eligibility may indicate the cause of latency and / or quality of service degradation (i.e., whether it is due to the user using access link 104 in accordance with their service level agreement (SLA), and if so, whether it is due to traffic increases or decreases, or whether it is due to other users sharing the same access link, resulting in access link, downlink, or uplink congestion). Impact eligibility may also indicate the type of slowdown, i.e., data volume starting to queue or data volume being dropped. Information indicating the flow of data volume may be provided in the upstream and / or downstream directions.
[0120] In a specific example, module 358 can monitor the flow of data volume on the network for each service 112 within each sampling interval to obtain output data traffic patterns. For example, module 358 can count the amount of data transmitted through the corresponding communication channel, stored in the corresponding communication channel queue, or dropped in the corresponding communication channel, and then report that count to a usage counter. This count can be aggregated across services or a separate count can be maintained for each service. The count can be aggregated according to a determined level of output granularity. Throughout the analysis, additional information regarding the impact on the provision of other services 112 (e.g., network congestion, packet queuing, the impact of service flow on other service flows due to service priorities, etc.) can also be recorded.
[0121] In one specific implementation, referring to step 4 in FIG3, each unit of data volume (e.g., each data byte) flowing through the pipeline implementing access node 102 is counted and recorded in a dedicated usage counter, and reported as part of the output at the end of the simulation or test scenario. In this example, module 358 for outputting information indicating the flow of data volume includes module 358 for outputting information indicating the flow of data volume through access node 102. For example, module 358 may count the amount of data transmitted through the simulated access network 104, the amount of data stored in the queue at the simulated access node 102, or the amount of data dropped at the simulated access node 102, and then report the count to the usage counter. However, the output can be performed at any suitable part of network 100.
[0122] In some examples, the output information also includes the amount of data dropped at each pipe level. In some examples, the output information includes a record of data traffic successfully transmitted through all pipes corresponding to each home gateway 108. The output information can provide an indication of the amount of traffic (i.e., the number of data bytes) at each sampling interval. In some other implementations, this amount can be aggregated over N sampling intervals, depending on the configuration of the digital twin. In some other examples, the output information may include the number of bytes exchanged (i.e., data usage) per gateway 108 and / or per service 112. In some other examples, the number of bytes exchanged may be considered at the user profile level (e.g., for monitoring end users) and / or at the access link level (e.g., to determine network congestion). Optionally, the output information may additionally or alternatively include quality of service or quality of experience information.
[0123] Referring further to step 5, the output information may include corresponding datasets, each representing a different simulation performed using the digital twin. These datasets can be used, for example, to evaluate different network configurations and scenarios, perform tests on the network, and train machine learning models, etc. In other words, by repeating the process described with reference to steps 1 to 5, but using different initial conditions (i.e., utilizing different first and second values selected from probability distribution 200), the digital twin described herein can create datasets including hundreds of thousands or even millions of different combinations of services 112 and service usage sequences, covering a variety of different network topologies and end-user profiles. In other words, the output of module 358 may include information about various combinations of services, end-users, service schedules, and interval granularity. Therefore, by using the digital twin, a method can be provided to represent the network at both the individual user-specific level and the macro-network level.
[0124] The output information can be configured as labeled data, allowing its use in the training of one or more machine learning models. In this way, the digital twins described herein enable the creation of relevant data and datasets suitable for training and evaluating machine learning models or AI-based features. This is particularly advantageous when the machine learning models or AI-based features are implemented for purposes requiring the matching of real-world traffic trajectories with real-world access network behavior (whether macroscopically (network level) or concretely (per user, per PON level)). Thus, data for training the aforementioned models / features can be provided faster and more efficiently compared to training with measurement data.
[0125] However, the output information can be used for purposes other than generating datasets. In another example, the output information can be used for capacity planning when introducing new users or new topologies to the network, and / or to assess the impact of upgrading network technologies. In particular, by selecting first and second values from corresponding probability distributions, digital twins can perform statistical analysis on different network topologies and / or different access link technologies to understand the impact of topology and / or technology on throughput efficiency, congestion, and dropped traffic. In some examples, this statistical analysis can be used to find optimal configurations, whether at the individual user level or on average across a broader network scale. Therefore, more efficient networks can be designed or configured.
[0126] Figure 3(d) further illustrates the system described above for implementing a digital twin of a network. Specifically, a system 300 for implementing a digital twin of network 100 is shown. The modules of system 300 may be provided separately, and / or, some or all of the modules may be combined into one or more software or hardware modules as needed.
[0127] System 300 includes a module 350 for selecting one or more first values 344, wherein each first value represents a corresponding parameter of the network configuration and is selected from a probability distribution 200A associated with the corresponding parameter of the network configuration. One or more first values may be associated with network values. Network values may include one or more of the following: capacity, buffer size, and drop policy.
[0128] The system includes a module 354 for selecting one or more second values 346, wherein each second value represents a corresponding parameter of a service to be performed on one or more terminal devices, and is selected from a probability distribution 200B associated with the corresponding parameter of the service to be performed, each service being associated with the transmission of one or more data packets over the network. The one or more second values may be associated with a service value. A service value may include one or more of the following: bit count, scheduling, and latency or round-trip time.
[0129] The system may optionally include a module 340 for storing each probability distribution 200A associated with corresponding parameters of network configuration and each probability distribution 200B associated with corresponding parameters of the service to be performed. Additionally or alternatively, the system may also optionally include a module 342 for accessing each probability distribution 200A associated with corresponding parameters of network configuration and each probability distribution 200B associated with corresponding parameters of the service to be performed.
[0130] System 300 further includes a module 354 for determining a time series representing the amount of data to be injected into the network in each of the one or more time intervals, based on one or more selected second values. System 300 also includes a module 356 for determining the flow of the data volume through the network based on one or more selected first values and the time series. The determining module 356 may further include: a module for injecting the data volume into the network in each of the one or more time intervals; and a module for processing the flow of each injected data volume through the network in each corresponding time interval to determine the (total) flow of the data volume. The system also includes a module 358 for outputting information 348 indicating the flow of the data volume.
[0131] The following will refer to Figure 4 (a) through 7(d) discuss the information output by module 358 in more detail, as exemplified by 348.
[0132] Figure 4 (a) and (b) illustrate example traffic patterns (e.g., the amount of data per interval). Specifically, downstream and upstream traffic trajectories of exchanged bytes are shown, aggregated at 10-second intervals (e.g., 10-second granularity); this traffic includes packets originating from / from services initiated by end device 110, and the selected network technology is PON. The downstream traffic trajectories shown... Figure 4 (a) and showing the upstream flow trajectory Figure 4 In example (b), service 112 includes: file transfer, video conferencing, internet video streaming, and browsing. Figure 5 (a) and (b) show the relationship with Figure 4 (a) and (b) have the same traffic pattern, but the output is aggregated at 1-second intervals (e.g., 1-second granularity). Figure 4 (a) Same Figure 5 (a) shows the downstream flow trajectory. Figure 5 (b) shows the upstream flow trajectory; Figure 4 (a) and (b) and Figure 5 The network conditions and service types remain unchanged between examples (a) and (b).
[0133] Figure 6 (b) through (e) illustrate example traffic patterns of an exemplary network, such as... Figure 6 As shown in (a). The network includes access node 102, access network 104 using GPON technology, home / user gateway 108, and terminal / home network 106 including an Ethernet connection to a single terminal device 110-1 running a file transfer service. Users have an SLA or profile with download / upload speeds of 500 / 250 Mbps, and the access network capacity is 2.5 Gbps. Figure 6(b) illustrates an example traffic pattern for file transfer service 112 with a round-trip time of 100 milliseconds. In this instance, all data (e.g., all packets) is sent because data usage never reaches the 500 Mbps limit. In contrast, Figure 6 (c) illustrates an example traffic pattern for file transfer service 112 with a round-trip time of 20 milliseconds. The dashed line represents 500 Mbps of data usage. In this example, from Figure 6 (d) It can be seen that some data (e.g., packets) is discarded, while some data is queued, such as... Figure 6 As shown in (e).
[0134] Figure 7 (b) through (d) illustrate example traffic patterns of an exemplary network, such as... Figure 7 As shown in (a). The network includes access node 102, access network 104 using GPON technology, home / user gateway 108, and terminal / home network 106, which includes an Ethernet connection to a single terminal device 110-1 running Video over Internet Transport (VOTT) service 112. In one example, the video bitrate of service 112 is 1050 Kbps, and in another example, it is 2350 Kbps. Users have an SLA or profile with download / upload speeds of 500 / 250 Mbps in one example and 100 / 50 Mbps in another example. The access network has a capacity of 2.5 Gbps. Figure 7 (b) shows an example initial buffering rate for a video bitrate of 1050 Kbps for VOTT service 112. In contrast, Figure 7 (c) shows an example initial buffering mode for VOTT service 112 at a video bitrate of 2350 Kbps. In both cases, the user profile speed is 500 Mbps. It can be seen that... Figure 7 The initial buffer rate in (c) is higher than Figure 7 The initial buffer rate in (b). In another comparative example, Figure 7 (d) shows an example initial buffering mode for VOTT service 112 at a video bitrate of 2350 Kbps, but the user profile speed is 100 Mbps. Because the buffering rate is limited by the user profile, Figure 7 Initial buffer ratio in (d) Figure 7 (c) is longer.
[0135] Figure 8 This is a flowchart illustrating the operations used to implement a digital twin of a network. These operations can be performed by the aforementioned system 300 and modules.
[0136] The first operation 1310 may include selecting one or more first values, where each first value represents a corresponding parameter of the network configuration and is selected from a probability distribution associated with the corresponding parameter of the network configuration. The one or more first values may be associated with network values. Network values may include one or more of the following: capacity, buffer size, and drop policy.
[0137] The second operation 1320 may include selecting one or more second values, wherein each second value represents a corresponding parameter of a service to be performed on one or more terminal devices, and is selected from a probability distribution associated with the corresponding parameter of the service to be performed, each service being associated with the transmission of one or more packets over the network. The one or more second values may be associated with a service value. A service value may include one or more of the following: bit count, scheduling, and latency or round-trip time.
[0138] The third operation 1330 may include determining a time series representing the amount of data to be injected into the network in each of one or more time intervals based on one or more selected second values. The fourth operation 1340 may include determining the flow of the data volume through the network based on one or more selected first values and the time series. The fifth operation 1350 may include outputting information indicating the flow of the (determined) data volume.
[0139] Optionally, the determination operation 1340 may include injecting a data volume into the network at each of one or more time intervals, and then processing the flow of each injected data volume through the network at each corresponding time interval to determine the total flow of the data volume. Optionally, the processing operation may include processing the data volume flow according to one or more first values (and associated network values); in other words, processing may include processing the data volume flow according to a capacity chain defined by network values (capacity, buffer size, drop policy) for the network, wherein the network configuration is defined itself according to one or more selected first values.
[0140] Figure 9Modules according to some example embodiments are shown, which may include system 300 (and any modules described in association with the system). The device may be configured to perform the operations described herein, such as those described with reference to any disclosed process. The device includes at least one processor 900 and at least one memory 901 directly or closely connected to the processor. Memory 901 includes at least one random access memory (RAM) 901a and at least one read-only memory (ROM) 901b. Computer program code (software) 905 is stored in ROM 901b. The device may be connected to a transmitter (TX) and a receiver (RX). The device may optionally be connected to a user interface (UI) for indicating the device and / or for outputting data. At least one processor 900, at least one memory 901, and computer program code 905 are arranged to cause the device to at least perform a method according to any of the foregoing processes, such as, in combination with... Figure 8 The flowchart and related features are disclosed herein. At least one memory 901 may also include the digital twin described herein, and optionally include a stored probability distribution 200 (e.g., the module 340 for storage may include at least one memory 901).
[0141] Figure 10 A non-transitory medium 1000 according to some embodiments is illustrated. The non-transitory medium 1000 is a computer-readable storage medium. It can be, for example, a CD, DVD, USB flash drive, Blu-ray disc, etc. The non-transitory medium 1000 stores computer program code, methods for causing a device to perform any of the aforementioned processes, for example, as disclosed in the flowcharts and their related features. The memory can be volatile or non-volatile. It can be, for example, RAM, SRAM, flash memory, FPGA block RAM, DCD, CD, USB flash drive, and Blu-ray disc.
[0142] Unless otherwise stated or the context clearly specifies, "two entities different" means that they perform different functions. This does not necessarily mean that they are based on different hardware. That is, each entity described in this specification may be based on different hardware, or some or all of the entities may be based on the same hardware. This does not necessarily mean that they are based on different software. That is, each entity described in this specification may be based on different software, or some or all of the entities may be based on the same software. Each entity described in this specification can be implemented in the cloud.
[0143] Implementations of any of the above-described blocks, devices, systems, technologies, or methods include (but are not limited to) implementations of hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof. Some embodiments can be implemented in the cloud.
[0144] It should be understood that the above description represents what is currently considered a preferred embodiment. However, it should be noted that the description of the preferred embodiment is given by way of example only, and various modifications can be made without departing from the scope defined by the appended claims.
Claims
1. A system implementing digital twinning of a network, the system comprising: a module for selecting one or more first values, wherein each first value represents a respective parameter of a configuration of the network and is selected from a probability distribution associated with the respective parameter of the configuration of the network; a module for selecting one or more second values, wherein each second value represents a respective parameter of a service to be executed on one or more end devices of the network and is selected from a probability distribution associated with the respective parameter of the service to be executed, each service being associated with the transmission of one or more data packets over the network; a module for determining, based on the one or more selected second values, a time series representing an amount of data to be injected into the network in each of one or more time intervals; a module for determining, based on the one or more selected first values and the time series, a flow of the amount of data through the network; and a module for outputting information indicative of the flow of the amount of data. The module for determining the flow of the amount of data through the network comprises:
2. The system of claim 1, wherein, a module for configuring a simulation of the network based on the one or more selected first values; a module for injecting the time series into the configured simulation of the network; and a module for determining the flow of the amount of data through the configured simulation of the network.
3. The system of claim 1 or 2, wherein: the module for determining, based on the one or more selected second values, a time series to be injected into the network comprises a module for determining the time series based on one or more of: an amount of bits associated with the one or more selected second values, a schedule associated with the one or more selected second values, and a delay associated with the one or more selected second values; and / or the module for determining, based on the one or more selected first values and the time series, a flow of the amount of data through the network can further comprise a module for determining the flow of the amount of data based on one or more of: a capacity associated with the one or more selected first values, a buffer size associated with the one or more selected first values, and a dropping policy associated with the one or more selected first values.
4. The system of any preceding claim, wherein: the module for selecting the one or more first values comprises a module for automatically selecting the one or more first values; and / or the module for selecting the one or more second values comprises a module for automatically selecting the one or more second values.
5. The system of any preceding claim, further comprising: a module for receiving, based on user input, one or more further first values representing respective parameters of the configuration of the network, wherein determining the flow of the amount of data through the network is further based on the one or more received further first values; and / or a module for receiving, based on user input, one or more other second values representative of respective parameters of the services to be executed on the one or more terminal devices, wherein determining the time series to be injected into the network is further based on the one or more received other second values.
6. The system of any of the preceding claims, wherein, The network comprises: one or more home gateways; one or more home networks; the one or more terminal devices, which are located behind the one or more home gateways and are connected to the one or more gateways through the one or more home networks; an access node providing access to one or more service providers, which are configured to provide the services to be executed on the one or more terminal devices; and an access network connected between the access node and the one or more home gateways.
7. The system of claim 6, wherein: the time series is injected into the access node; and / or the module for outputting information indicative of the flow of the amount of data comprises a module for outputting information indicative of the flow of the amount of data through the access node.
8. The system of any of the preceding claims, wherein, The information indicative of the flow of the amount of data comprises one or more of: an amount of data transmitted, an amount of data queued, an amount of data dropped, a quality of service indication, a service slowdown indication, a service latency indication, a data queue status, or an impact eligibility.
9. The system of any of the preceding claims, wherein, The configured parameters of the network comprise one or more of: a type, a connection, or a characteristic of a network node, optionally wherein the characteristic of the network node comprises a technology of the network node, the technology comprising one or more of: Asymmetric Digital Subscriber Line (ADSL), Very High Digital Subscriber Line (VDSL), GFast, Gigabit-capable Passive Optical Network (GPON), 10 Gigabit-capable Passive Optical Network (XGPON), or 10 Gigabit-capable Symmetric Passive Optical Network (XGSPON); and / or a number of gateways, a number of terminal devices, a type of terminal device, or a network subscription.
10. The system of any of the preceding claims, wherein, The parameters of the services to be executed on the one or more terminal devices comprise one or more of: a type, a schedule, or a characteristic of the service, optionally wherein the type of the service comprises: audio, browsing, file transfer, gaming, Internet of Things, Internet Protocol Television, video conferencing, Voice over Internet Protocol, or Internet video transmission; and / or a number of services on each terminal device.
11. The system of any of the preceding claims, wherein, Each probability distribution is determined based on data measured from a real-world network.
12. The system of any of the preceding claims, wherein, The module for selecting the one or more second values comprises a module for selecting the one or more second values as a function of the one or more selected first values.
13. The system of any one of the preceding claims, further comprising: a module for storing each probability distribution associated with respective parameters of a configuration of the network and each probability distribution associated with respective parameters of the services to be executed; and / or a module for selecting the one or more second values comprises a module for selecting the one or more second values as a function of the one or more selected first values. a module for accessing each probability distribution associated with a respective parameter of a configuration of the network and accessing each probability distribution associated with a respective parameter of a service to be executed from a remote storage device.
14. A method for implementing digital twinning of a network, the method comprising: selecting one or more first values, wherein each first value represents a respective parameter of a configuration of the network and is selected from a probability distribution associated with a respective parameter of a configuration of the network; selecting one or more second values, wherein each second value represents a respective parameter of a service to be executed on one or more terminal devices of the network and is selected from a probability distribution associated with a respective parameter of a service to be executed, each service being associated with transmission of one or more data packets through the network; determining, based on the one or more selected second values, a time series representing an amount of data to be injected into the network at each of one or more time intervals; determining, based on the one or more selected first values and the time series, a flow of the amount of data through the network; and outputting information indicative of the flow of the amount of data.
15. A computer program comprising instructions for causing an apparatus to perform: selecting one or more first values, wherein each first value represents a respective parameter of a configuration of the network and is selected from a probability distribution associated with a respective parameter of a configuration of the network; selecting one or more second values, wherein each second value represents a respective parameter of a service to be executed on one or more terminal devices of the network and is selected from a probability distribution associated with a respective parameter of a service to be executed, each service being associated with transmission of one or more data packets through the network; determining, based on the one or more selected second values, a time series representing an amount of data to be injected into the network at each of one or more time intervals; determining, based on the one or more selected first values and the time series, a flow of the amount of data through the network; and outputting information indicative of the flow of the amount of data.