Control device, DCI gateway system, DCI gateway, setting method, and program
The DCI gateway system addresses reliability issues in DCI systems by implementing redundancy and packet recovery mechanisms, ensuring high reliability and low latency data transfer between DCI clusters.
Patent Information
- Application Number
- PCT/JP2024/026340
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-01-29
AI Technical Summary
Current DCI systems lack technologies for high reliability in path repair and packet recovery, particularly in geographically distant clusters, leading to increased propagation delays and communication interruptions during failures, with existing protocols assuming a lossless network and limited redundancy.
A control device and DCI gateway system that includes a route selection unit, delay calculation unit, resource calculation unit, and determination unit to configure a DCI gateway connecting a DCI cluster to an All Photonics Network, employing redundancy, flow control, and lost packet recovery to enhance reliability.
The system achieves high reliability in DCI by minimizing communication overhead and ensuring low latency data transfer between DCI clusters through redundant configurations and efficient packet recovery, reducing packet loss and communication interruptions.
Smart Images

Figure JP2024026340_29012026_PF_FP_ABST
Abstract
Description
Control device, DCI gateway system, DCI gateway, setting method and program
[0001] The present disclosure relates to a control device, a DCI gateway system, a DCI gateway, a setting method, and a program.
[0002] Advances in AI technology and new use cases such as autonomous driving are placing increasing demands on communication and computing power. It is expected that systems that guarantee high speed, low latency, and low jitter will be flexibly built for real-time AI analysis, including autonomous driving, and for e-sports.
[0003] In response to this, the Innovative Optical and Wireless Network (IOWN) Global Forum has proposed an architecture called Data Centric Infrastructure (DCI), which aims to achieve the above requirements (Non-Patent Document 1). DCI is composed of a DCI cluster, a DCI GW, and an APN (All Photonics Network). In addition, development of a DCI system (deterministic delay) that guarantees the delay from data transfer to application processing has begun (Non-Patent Document 2).
[0004] Today's datacenters are built using Clos Networks (Non-Patent Document 3), and DCI clusters are expected to adopt a similar architecture. Each rack is built using Compute Express Link (CXL) or similar protocols (Non-Patent Document 4), and communication between racks and between DCI clusters is performed using Remote Direct Memory Access (RDMA). Protocols such as RoCEv2 are used for long-distance RDMA communication (Non-Patent Document 5).
[0005] IWON Global Forum, "IDC Product Concept Paper", [online], Internet <https: / / iowngf.org / wp-content / uploads / formidable / 21 / IOWN-GF-RD-DCI_PCP-1.1.pdf>, "Research and Development of Highly Efficient Deterministic Delay Computing Infrastructure Technology", [online], Internet <https: / / www.meti.go.jp / policy / mono_info_service / joho / post5g / pdf / 240130_theme_01.pdf>, "Use of BGP for Routing in Large-Scale Data Centers", [online], Internet <https: / / www.rfc-editor.org / rfc / rfc7938>, "CXL3.1 Specification Now Available", [online], Internet <https: / / computeexpresslink.org / > "RDMA Acceleration Methods in Long-Distance Optical Networks," [online], Internet <https: / / ken.ieice.org / ken / paper / 20211007tCfr / > "Open All-Photonic Network Functional Architecture," [online], Internet <https: / / iowngf.org / wp-content / uploads / formidable / 21 / IOWN-GF-RD-Open_APN_Functional_Architecture-2.0.pdf> "ITU-T G.873.1: Optical transport network: Linear protection, 2017," [online], Internet <https: / / www.itu.int / rec / T-REC-G.873.1-201710-I / en> "IEEE 802.3ad: Link Aggregation Control Protocol, 2000#," [online], Internet <https: / / www.ieee802.org / 3 / hssg / public / apr07 / frazier_01_0407.pdf>“RFC5798: Virtual Router Redundancy Protocol (VRRP), 2010”, [online], Internet <https: / / www.rfc-editor.org / rfc / rfc5798.txt> “RFC4271: A Border Gateway Protocol 4 (BGP-4), 1994”, [online], Internet <https: / / www.rfc-editor.org / rfc / rfc4271.txt> “SMPTE ST 2022-7: STANDARD Seemless Protection Switching of RTP Datagrams, 2019.” [online], Internet <https: / / pub.smpte.org / latest / st2022-7 / st2022-7-2019.pdf> “SMPTE ST 2022-5: Forward Error Correction for Transport of High Bit Rate Media Signals over IP Networks (HBRMT), 2013, [online], Internet <https: / / pub.smpte.org / pub / st2022-5 / st2022-5-2013.pdf>, "RFC3168: The Addition of Explicit Congestion Notification (ECN) to IP", [online], Internet <https: / / www.rfc-editor.org / rfc / rfc3168.txt>, "Tactile Remote Robot Control Using Low-Latency Transport and Precision Bilateral Control Technologies", NTT Journal, [online], Internet <https: / / journal.ntt.co.jp / article / 26175>.
[0006] To build applications with deterministic latency using DCI, high-reliability technologies are important in addition to high-speed technologies. However, currently, technologies for path repair and packet recovery in each component have not been considered. DCI also works with geographically distant DCI clusters, which increases propagation delays and increases switching times when a failure occurs. For example, in the case of a path failure within a datacenter, switching can be achieved in a few milliseconds by monitoring the link. However, the target path switching time for transmission equipment used within an APN is 50 msec or less, resulting in the loss of many packets during the switching process (Non-Patent Documents 6 and 7).
[0007] Even within a DCI cluster, failures can occur due to buffer overflow caused by processing delays or congestion, fiber breakage, or equipment failure. In the case of a link failure, switching occurs with a communication interruption of several milliseconds using Link Aggregation (Non-Patent Document 8). On the other hand, in the IP layer, route switching is performed using VRRP or BGP, resulting in communication interruptions of several seconds (Non-Patent Documents 9 and 10). If the sender is in a remote location, an additional propagation delay is required to notify the communication interruption.
[0008] To reduce communication overhead, communication between DCI clusters requires memory-to-memory data transfer, such as RDMA, which does not involve the OS or CPU. However, the RDMA protocol presupposes a lossless network, and high reliability, such as redundancy, is only defined for Ethernet and IP (Non-Patent Documents 11, 12, 13).
[0009] Therefore, in a system that connects DCI clusters in remote locations, it is necessary to build a DCI GW that takes into account reliability requirements.In addition, a DCI GW that can achieve redundancy in data transfer between memories between DCI clusters is also required.
[0010] The present disclosure has been made in consideration of the above circumstances, and an object of the present disclosure is to provide a technology that realizes high reliability in DCI.
[0011] In order to achieve the above object, one aspect of the present disclosure is a control device that configures a DCI gateway that connects a Data Centric Infrastructure (DCI) cluster to an All Photonics Network (APN), the control device including: a route selection unit that uses a route database to select multiple route candidates for a communication route between the DCI cluster and a communication device that communicates with the DCI cluster, and selects at least one high-reliability method according to user requirements; a delay calculation unit that calculates a delay time for each route candidate for each high-reliability method using a transmission time of the communication route and a processing time of the DCI gateway; a resource calculation unit that calculates resources required for each route candidate for each high-reliability method based on the delay time calculated by the delay calculation unit; a determination unit that determines, for each route candidate, one of the multiple route candidates and a high-reliability method based on the delay time and resources calculated for each high-reliability method; and a configuration unit that configures the DCI gateway and the APN using the determined communication route and high-reliability method.
[0012] One aspect of the present disclosure is a DCI gateway system comprising: a first DCI (Data Centric Infrastructure) gateway connected to a first APN (All Photonics Network); and a second DCI gateway connected to a second APN, wherein the first DCI gateway comprises a receiving unit that receives data transmitted from a communication device via the first APN and a processing unit that writes the data to a memory of a DCI cluster, and the second DCI gateway comprises a receiving unit that receives the same data as the data transmitted from the communication device via the second APN, and a processing unit that reads the data written to the memory of the DCI cluster and, when packet loss is detected, writes a packet corresponding to the packet lost from the received data to the memory.
[0013] One aspect of the present disclosure is a DCI gateway, which is a DCI (Data Centric Infrastructure) gateway connected to a first APN (All Photonics Network) and a second APN, and includes a first receiver that receives data transmitted from a communication device via the first APN, a second receiver that receives the same data as the data transmitted from the communication device via the second APN, and a processing unit that uses the data received by the first receiver and the data received by the second receiver to generate transmission data to be transmitted to a DCI cluster, and transmits the transmission data to the DCI cluster.
[0014] One aspect of the present disclosure is a configuration method performed by a control device to configure a DCI gateway that connects a Data Centric Infrastructure (DCI) cluster to an All Photonics Network (APN), the method selecting, using a route database, multiple route candidates for a communication route between the DCI cluster and a communication device that communicates with the DCI cluster, and selecting at least one high-reliability method according to user requirements, calculating a delay time for each route candidate for each high-reliability method using a transmission time of the communication route and a processing time of the DCI gateway, calculating resources required for each route candidate for each high-reliability method based on the calculated delay time, determining a communication route and a high-reliability method from among the multiple route candidates based on the delay time and resources calculated for each high-reliability method, and configuring the DCI gateway and the APN using the determined communication route and high-reliability method.
[0015] One aspect of the present disclosure is a program that causes a computer to function as the control device.
[0016] According to the present disclosure, it is possible to provide a technique for achieving high reliability in DCI.
[0017] Fig. 1 shows an example of a DCI configuration. Fig. 2 is a diagram showing an example of the configuration of a DCI cluster, DCI GW, and control device of this embodiment. Fig. 3 is a schematic diagram showing a redundant configuration using redundancy of DCI GWs. Fig. 4 is a schematic diagram showing another redundant configuration using redundancy of DCI GWs. Fig. 5 is a schematic diagram showing a redundant configuration using redundancy of DCI GW ports. Fig. 6 is a schematic diagram showing an example of a configuration for recovering lost packets. Fig. 7 is an example of a hardware configuration.
[0018] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0019] 1 shows an example of the configuration of a DCI (Data Centric Infrastructure). The DCI is configured using a DCI cluster 2, a DCI GW 3 (DCI gateway), and an APN (All Photonics Network) 9.
[0020] A data center 1 is configured with a DCI cluster 2 and at least one DCI GW 3. The DCI cluster 2 is constructed from a large number of hosts and HW accelerators (GPUs, FPGAs, etc.), and a network with guaranteed quality in terms of latency, jitter, bandwidth, etc. The DCI cluster 2 allows for flexible combination of HW accelerators, which was not possible with conventional CPU-centric architectures, and also makes it possible to guarantee application processing times.
[0021] DCI GW3 functions as a termination function for APN9 and as a gateway for the network of Data Center 1. If DCI Cluster 2 is constructed using an IP-based Clos network, it will have the termination function for APN9 and the function of a router, but in the future, routers will be eliminated and the network will be constructed entirely using optical fiber and switches.
[0022] The DCI GW 3 connects an in-center DCI cluster 2 located in the same data center with an out-of-center DCI cluster 2 located in another data center via an APN 9. The DCI GW 3 also connects the in-center DCI cluster 2 with a user system 7 via the APN 9. The user system 7 is installed at a user site together with the DCI GW 3.
[0023] The DCI architecture aims to link multiple geographically separated DCI clusters 2, and an accelerator has been proposed to transfer RoCEv2 data remotely at high speed (Non-Patent Document 5). Such acceleration and protocol conversion functions are built into the DCI GW 3.
[0024] APN9 occupies optical wavelengths according to the application, and because the relay network is entirely optical, packets can be transferred without processing delays such as buffering and routing.
[0025] 2 is a diagram showing the configuration of the DCI cluster 2, DCI GW 3, and control device (controller) 5 of this embodiment. The control device 5 may be disposed in each data center 1, or one control device 5 may be disposed for multiple data centers 1.
[0026] <DCI Cluster> The DCI cluster 2 comprises a computational resource group 21, a time synchronization unit 22, a delay measurement unit 23, and a resource management unit 24. The computational resource group 21 executes applications using multiple CPUs (e.g., servers). The time synchronization unit 22 achieves highly accurate time synchronization with the DCI GW 3, control device 5, etc. The delay measurement unit 23 measures the processing time (processing delay time) of the DCI cluster 2. The resource management unit 24 manages resources such as the memory and CPU of the DCI cluster 2.
[0027] <DCI GW> The DCI GW 3 includes a port 31 (internal interface), a buffer 32 (transmission / reception buffer), a protocol processing unit 33, a buffer 38, a port 39 (external interface), a time synchronization unit 42, a delay measurement unit 43, and a resource management unit 44. The time synchronization unit 42 achieves highly accurate time synchronization in order to measure delay times with high precision. The delay measurement unit 43 measures the processing time (processing delay time) of the DCI GW 3. The resource management unit 44 manages resources such as the memory, buffers 32 and 38, and CPU of the DCI GW 3.
[0028] The port 31 is a connection interface with the DCI cluster 2. The port 39 is a connection interface with the APN 9. The buffers 32 and 38 are memories for temporarily storing data when transmitting and receiving data.
[0029] The protocol processing unit 33 includes a DC protocol processing unit 34 , a DC external protocol processing unit 37 , a high functionality unit 35 , and a high reliability unit 36 .
[0030] The DC protocol processing unit 34 processes (converts) protocols (such as CXL or RoCEv2) within the DCI cluster 2. The DC protocol processing unit 34 speeds up communication between DCI GWs 3 by sending pseudo ACKs of RDMA (Remote Direct Memory Access) (Non-Patent Document 14). The DC external protocol processing unit 37 processes communication protocols via the APN 9. For example, when the high reliability unit 36 uses SMPTE ST 2022-7, it inserts or deletes an RTP header.
[0031] The high performance unit 35 performs time synchronization, time stamp assignment for measuring delays in application processing and communication processing, and tunneling such as RTP.
[0032] The high reliability unit 36 performs communication redundancy, flow control to prevent buffer overflow on the receiving side, recovery of lost packets, etc. The type of high reliability method that the high reliability unit 36 uses to achieve a confirmed delay is determined by settings from the control device 5.
[0033] 1. Redundancy Redundancy of communication with the DCI cluster 2, such as CXL, is only considered at the IP or transmission layer. For this reason, this embodiment employs the following configuration (architecture) to achieve redundancy of the DCI GW 3 device or redundancy of the DCI GW 3 port.
[0034] (1) Redundancy of DCI GW 3 (Part 1) Fig. 3 is a schematic diagram showing a redundancy configuration using redundancy of DCI GW 3. A transmitter 6 on the sending side transmits the same data to multiple DCI GWs 3 via an APN 9. Note that the same data may have the same payload but different headers. The transmitter 6 is, for example, a DCI cluster 2 in another data center 1 or a user system 7. The data center 1 in the illustrated example includes a primary DCI GW 3A and a secondary DCI GW 3B.
[0035] The high reliability unit 36 of the DCI GW 3A receives data transmitted from the transmitter 6 via the APN 9A and writes the data to a memory included in the DCI GW 3A. The transmitter 6 may assign an ID to the data before transmitting it. The ID may include a timestamp, version information, a sequence number, and the like. This allows the DCI GW 3A and DCI cluster 2 to detect the latest data, missing data, and the like using the high reliability unit 36. Similarly, the high reliability unit 36 of the DCI GW 3B receives data transmitted from the transmitter 6 via the APN 9B and writes the data to a memory included in the DCI GW 3B.
[0036] The DCI cluster 2 receives data transmitted from the transmitter 6 by reading data written in the memory of each DCI GW 3A, 3B. That is, the DCI cluster 2 polls the DCI GWs 3A, 3B for data. The DCI cluster 2 may use an ID to avoid redundantly reading data that has been written in multiple DCI GWs 3A, 3B. For example, the DCI cluster 2 may read the memory of the primary DCI GW 3A and use the ID to check whether the data in that memory is the latest or whether there are any missing data. If the data is the latest or there are no missing data, the DCI cluster 2 may omit reading the memory of the secondary DCI GW 3B. Data acquisition from the DCI cluster 2 to each DCI GW 3A, 3B is communication within the same data center 1, so low latency data acquisition is possible.
[0037] When transmitting data from the DCI cluster 2 to the transmitter 6, the DCI cluster 2 writes the data to be transmitted to the memory of each DCI GW 3A, 3B. This allows the same data to be transmitted to the transmitter 6 via two APNs 9A, 9B. The two APNs 9A, 9B may be networks provided by the same telecommunications carrier or networks provided by different telecommunications carriers. Data (packets) may be replicated within the APN 9 in the transmitter 6. Using a splitter or the like within the APN 9 to replicate data can reduce the computational resources of the transmitter 6. The same applies to the redundant configuration described below.
[0038] (2) Redundancy of DCI GW 3 (Part 2) Fig. 4 is a schematic diagram showing another redundancy configuration using redundancy of the DCI GW 3. In the configuration shown in Fig. 4, multiple DCI GWs 3A and 3B write data transmitted from the transmitter 6 to the memory (shared memory) of the DCI cluster 2. The example shown in the figure includes a primary DCI GW 3A and a secondary DCI GW 3B.
[0039] A transmitter 6 on the sending side transmits the same data to a plurality of DCI GWs 3A and 3B via APNs 9A and 9B, respectively. Note that the same data may have the same payload but different headers.
[0040] The high reliability unit 36 of each DCI GW 3A, 3B writes the received data to a write memory included in the DCI cluster 2. Each DCI GW 3A, 3B may write the received data to a different area (a primary memory area, a secondary memory area) of the write memory of the DCI cluster 2. In this case, the application of the DCI cluster 2 may read data from either memory area (for example, the primary memory area). Note that, to avoid data contention when each DCI GW 3A, 3B writes data to a memory area of the DCI cluster 2, the high reliability unit 36 of each DCI GW 3A, 3B may use the time synchronization unit 42 to set a time period for writing to the memory area that is different from that of the other DCI GW. Furthermore, each DCI GW 3A, 3B may write different parts (for example, the first half and the second half) of the same data transmitted from the transmitter 6 to the DCI cluster 2, thereby distributing the write destination memories.
[0041] Alternatively, the primary DCI GW 3A may always write data to the memory of the DCI cluster 2, and the secondary DCI GW 3B may read the data written to the memory of the DCI cluster 2 and write the missing data to the memory only if it detects missing data (packet loss). That is, the DCI GW system shown in Fig. 4 includes a DCI GW 3A (first DCI GW) connected to an APN 9A (first APN) and a DCI GW 3B (second DCI GW) connected to an APN 9B. The DCI GW 3A includes a port 39 (receiving unit) that receives data transmitted from a transmitter 6 (communication device) via the APN 9A, and a protocol processing unit 33 (processing unit) that writes the data to the memory of the DCI cluster 2. The DCI GW 3B (second DCI GW) includes a port 39 (receiving unit) that receives the same data as the data transmitted from the transmitter 6 via the APN 9B (second APN), and a protocol processing unit 33 (processing unit) that reads data written to the memory of the DCI cluster 2 and, if packet loss is detected, writes a packet corresponding to the packet that was lost from the received data to the memory.
[0042] (3) Port Redundancy Figure 5 is a schematic diagram showing a redundant configuration using redundancy of the port 39 of the DCI GW 3. In the configuration shown in Figure 5, a redundant configuration similar to SMPTE ST 2022-7 is realized by making the port 39 redundant within one DCI GW 3. Since SMPTE ST 2022-7 uses RTP, the reliability enhancement unit 36 adds or deletes a protocol header of the transmitted data.
[0043] The illustrated DCI GW 3 includes multiple ports 39A and 39B. Port 39A receives data transmitted from the transmitter 6 via APN 9A. Port 39B receives the same data transmitted from the transmitter 6 via APN 9B. The reconfiguration unit 361 of the high reliability unit 36 reconfigures the data received by ports 39A and 39B into a single piece of data and writes the reconfigured data to the memory of the DCI cluster 2. Note that the maximum transmission delay among the multiple APNs 9A and 9B is used as the transmission delay. Therefore, the transmission delay from the transmitter 6 is calculated as the protocol processing time of the DCI GW 3 + max (transmission delay of APN 9A, transmission delay of APN 9B). Therefore, the buffer 38 connected to each port must reserve a buffer for the difference in transmission delay between the connected APNs 9A and 9B (i.e., the difference between the transmission delay of APN 9A and the transmission delay of APN 9B). The buffer may be maintained on the APN side.
[0044] As such, the DCI GW 3 shown in FIG. 5 is a DCI GW connected to an APN 9A (first APN) and an APN 9B (second APN), and includes a port 39A (first receiving unit) that receives data transmitted from a transmitter 6 (communication device) via the APN 9B, a port 39 (second receiving unit) that receives the same data as the data transmitted from the transmitter 6 via the APN 9B, and a protocol processing unit 33 that uses the data received by port 39A and the data received by port 39B to generate transmission data to be transmitted to the DCI cluster 2, and transmits the transmission data to the DCI cluster 2.
[0045] 2. Flow Control The high reliability unit 36 of the DCI GW 3 may send a congestion notification to the transmitter 6 to prevent overflow of the buffers 32, 38. This controls the flow rate of data transmitted from the transmitter 6 and prevents packet loss due to buffer overflow. For example, the high reliability unit 36 may monitor the free space in the buffers 32, 38, and send a congestion notification to the transmitter 6 when the free space falls below a predetermined capacity. This allows the transmitter 6 to transmit data (packets) taking into account the buffers 32, 38 of the DCI GW 3 on the receiving side, regardless of the flow rate of data to be transmitted.
[0046] In this way, the flow rate control in this embodiment may be performed using information notified from the DCI GW 3 on the receiving side, such as Explicit Congestion Notification (Non-Patent Document 13). That is, the high reliability unit 36 monitors the free space in the buffers 32 and 38, and prevents packet loss due to buffer overflow by controlling the flow rate that the transmitter 6 sends to the APN 9. Furthermore, the high reliability unit 36 may perform priority control for each communication flow, thereby preferentially transferring traffic with high delay requirements or importance.
[0047] 3. Lost Packet Recovery Figure 6 is a schematic diagram showing an example of a configuration for recovering lost packets. In the configuration shown in Figure 6, high reliability is achieved by recovering lost packets using a single DCI GW 3.
[0048] The high reliability unit 36 may recover lost packets using FEC (Forward Error Correction) and / or a retransmission function. The high reliability unit 36 may also discard data that could not be processed within a specified delay time without transmitting it to the DCI cluster 2. This prevents the sending of unnecessary data, and reduces the transmission and processing of unnecessary data.
[0049] The reliability enhancement unit 36 may also use different lost packet recovery methods for communication with the DCI cluster 2 and communication with the transmitter 6 via the APN 9. In the example shown, lost packets are recovered by retransmission in communication with the DCI cluster 2, and lost packets are recovered by FEC in communication with the transmitter 6. In this case, error correction packets are added to the data transmitted from the transmitter 6. The reliability enhancement unit 36 detects packet loss using the error correction packets and restores the lost packets.
[0050] When FEC is used, the communication bandwidth used increases, so the reliability enhancement unit 36 may perform control to reduce the bandwidth of port 31 on the DCI cluster 2 side. Also, a buffer for FEC needs to be reserved in buffer 38 on the port 39 side. Also, when recovering lost packets by retransmission in communication with DCI cluster 2, a buffer for retransmission needs to be reserved in buffer 32 on the port 31 side. In the example shown in the figure, there is no need to implement special functions in DCI cluster 2, and since packets are not retransmitted via APN 9, there is the effect of suppressing increases in retransmission delays.
[0051] <Control Device> The control device 5 sets the configuration of the DCI GW 3 that connects the DCI cluster 2 to the APN 9. The control device 5 shown in the figure includes a requirement reception unit 51, a route selection unit 52, a delay acquisition unit 53, a delay calculation unit 54, a resource calculation unit 55, a determination unit 56, a setting unit 57, a route DB (route database) 58, and a resource management unit 59.
[0052] The requirement receiving unit 51 receives requirements for improving reliability input by the user. For example, the following requirements may be given:
[0053] -Whether or not to make the route redundant.
[0054] -Whether or not to make DCI GW3 redundant.
[0055] - Whether or not to allow packet retransmission.
[0056] - Whether or not to consider route disconnection. If so, whether or not to allow packet loss (50 ms communication interruption) when switching routes.
[0057] - Whether geographical distribution and other physical conditions should be taken into consideration (for example, whether overhead lines are permitted or not).
[0058] The route selection unit 52 uses a route database to select multiple route candidates for the communication route between the DCI cluster 2 and the transmitter 6 (communication device) that communicates with the DCI cluster 2, and also selects at least one high-reliability method according to user requirements.
[0059] Specifically, the route selection unit 52 references information about available routes for each APN 9 stored in the route DB 58 and extracts possible communication routes between the DCI cluster 2 and the transmitter 6 as route candidates. The route selection unit 52 also selects at least one high reliability method (a communication method related to high reliability) according to the requirements. The high reliability method may include at least one of the above-mentioned communication route redundancy, DCI GW 3 redundancy, flow control, and lost packet recovery.
[0060] The delay calculation unit 54 calculates the delay time of each route candidate for each high reliability method using the transmission time of the communication route and the protocol processing time of the DCI GW 3. The delay calculation unit 54 may add the application processing time of the DCI cluster 2 to the delay time. The delay calculation unit 54 may use the delay acquisition unit 53 to acquire the transmission time and processing time measured by the DCI GW 3, APN 9, and DCI cluster 2, or may calculate the transmission time using the route DB 58.
[0061] The delay calculation unit 54 also calculates the delay time for each candidate route for each high-reliability method. For example, if packet retransmission is permitted when recovering lost packets, the transmission delay time for one round trip per packet is added to the delay time. If FEC is used, the processing time required for packet recovery is added to the delay time in addition to the increased communication charges due to FEC. Note that the transmission delay time for one round trip per packet and the processing time for packet recovery may be measured or calculated. For each high-reliability method, the delay calculation unit 54 calculates the worst-case (or 99th percentile, etc.) delay time by taking into account the time required for that high-reliability method. Furthermore, in the case of a high-reliability method that adds redundancy to communication paths, the transmission delay time of the slowest communication path is used.
[0062] The resource calculation unit 55 calculates the resources required for each route candidate for each high reliability method based on the delay time calculated by the delay calculation unit. The resources include the buffer capacity of the buffers 32 and 38 of the DCI GW 3. Specifically, the resource calculation unit 55 calculates resources such as the buffer capacity required for the DCI GW 3 based on the transmission delay calculated by the delay calculation unit 54, the protocol processing of the DCI GW 3, and the delay time due to application processing by the DCI cluster 2. Note that, in the case of a communication method that provides route redundancy, the resource calculation unit 55 calculates the buffer resource for the transmission delay, taking into account the delay difference between the redundant routes. The resource calculation unit 55 also calculates the required bandwidth in a similar manner.
[0063] For example, in a high-reliability system that uses redundant communication paths, the buffer capacity must be sufficient to accommodate the amount of data transferred over the time difference in delay between each communication path. Also, in a high-reliability system designed to allow only one retransmission for packet loss, the buffer capacity must be sufficient to accommodate the amount of data transferred over the propagation delay of one round trip.
[0064] The resource calculation unit 55 may refer to the resource management unit 59 (resource database) and delete route candidates of high reliability methods that cannot be realized by securing the calculated resources. Specifically, the resource calculation unit 55 may compare the resources calculated for each high reliability method of each route candidate with the current resource status managed by the resource management unit 59, determine the feasibility of each high reliability method of each route candidate, and delete route candidates of high reliability methods that cannot be realized due to a lack of resources.
[0065] The determination unit 56 determines a communication route and a high reliability method from among the multiple route candidates based on the delay time and resources calculated for each high reliability method for each route candidate. Specifically, the determination unit 56 determines an optimal communication route and high reliability method based on the calculated delay time and resources (including buffer capacity). Note that the determination unit 56 may also take into account costs such as communication charges and electricity charges for the DCI GW 3 when determining the communication route and high reliability method. Alternatively, the determination unit 56 may present multiple route candidates and high reliability methods to the user and allow the user to select one.
[0066] The setting unit 57 sets the DCI GW 3 and the APN 9 using the determined communication path and high reliability method. Specifically, the DCI cluster 2, the APN 9, and the DCI GW 3 are actually constructed based on the communication path, high reliability method, and resources determined by the determination unit 56.
[0067] The control device 5 may include a time synchronization unit (not shown) that performs highly accurate time synchronization in order to measure the delay time with high accuracy.
[0068] The control device 5 of the present embodiment described above is a control device 5 that sets a DCI GW 3 that connects a DCI cluster 2 to an APN 9, and includes: a route selection unit 52 that uses a route DB 58 to select multiple route candidates for a communication route to a transmitter 6 (communication device) that communicates with the DCI cluster 2, and selects at least one high reliability method according to user requirements; a delay calculation unit 54 that calculates a delay time for each route candidate for each high reliability method using a transmission time of the communication route and a processing time of the DCI GW 3; a resource calculation unit 55 that calculates, for each high reliability method, resources required for each route candidate based on the delay time calculated by the delay calculation unit 54; a determination unit 56 that determines, for each route candidate, one communication route and high reliability method from among the multiple route candidates based on the delay time and resources calculated for each high reliability method; and a setting unit that sets the DCI GW 3 and the APN 9 using the determined communication route and high reliability method.
[0069] As a result, in this embodiment, it is possible to automatically build a system that satisfies given requirements, such as high reliability and delay. That is, in a system in which remote DCI clusters 2 are linked together, it is possible to take into account reliability requirements and determine the high reliability of the entire DCI system in an integrated manner, and build the system. Specifically, it is possible to calculate the required resources (buffers) and delay time based on the communication path, the high reliability method for the DCI GW 3, and the like, determine the high reliability method for the DCI GW 3 that is optimal for the requirements requested by the user, and secure and configure the resources for the DCI GW 3.
[0070] Furthermore, the DCI GW 3 of this embodiment can realize redundancy between DCI clusters 2 via the APN 9 by using the high reliability unit 36 of the protocol processing unit 33. Specifically, the DCI GW 3 of this embodiment can realize redundancy in data transfer between memories between DCI clusters 2.
[0071] The control device 5 and DCI GW 3 described above can use, for example, a general-purpose computer system as shown in Fig. 7. The computer system shown in the figure includes a CPU (Central Processing Unit, processor) 901, a memory 902, a storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), a communication device 904, an input device 905, and an output device 906. The memory 902 and the storage 903 are storage devices. In this computer system, the CPU 901 executes a predetermined program loaded on the memory 902, thereby realizing each function of the control device 5 and the DCI GW 3.
[0072] The control device 5 and the DCI GW 3 may be implemented on a single computer or multiple computers. The control device 5 and the DCI GW 3 may be virtual machines implemented on a computer. The programs for the control device 5 and the DCI GW 3 may be stored on a computer-readable recording medium such as a HDD, SSD, USB (Universal Serial Bus) memory, CD (Compact Disc), or DVD (Digital Versatile Disc), or may be distributed via a network. The computer-readable recording medium is, for example, a non-transitory recording medium.
[0073] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the present disclosure.
[0074] DESCRIPTION OF SYMBOLS 1: Data center 2: DCI cluster 21: Computational resource group 22, 42: Time synchronization unit 23, 43: Delay measurement unit 24, 44: Resource management unit 3: DCI GW 31, 39: Port (internal interface, external interface) 32, 38: Buffer (transmission / reception buffer) 33: Protocol processing unit 34: DC internal protocol processing unit 35: High functionality unit 36: High reliability unit 37: DC external protocol processing unit 5: Control device (controller) 51: Requirements reception unit 52: Route selection unit 53: Delay acquisition unit 54: Delay calculation unit 55: Resource calculation unit 56: Determination unit 57: Setting unit 58: Route DB 59: Resource management unit 6: Transmitter 7: User system
Claims
1. A control device that configures a DCI gateway that connects a DCI (Data Centric Infrastructure) cluster to an APN (All Photonics Network), comprising: a route selection unit that uses a route database to select multiple route candidates for a communication route between the DCI cluster and a communication device that communicates with the DCI cluster, and selects at least one high-reliability method according to user requirements; a delay calculation unit that calculates the delay time of each route candidate for each high-reliability method using the transmission time of the communication route and the processing time of the DCI gateway; a resource calculation unit that calculates the resources required for each route candidate for each high-reliability method based on the delay time calculated by the delay calculation unit; a determination unit that determines one of the communication routes and high-reliability method from among the multiple route candidates for each route candidate based on the delay time and resources calculated for each high-reliability method; and a configuration unit that configures the DCI gateway and the APN using the determined communication route and high-reliability method.
2. The control device according to claim 1, wherein the resource calculation unit refers to a resource database and eliminates route candidates of high reliability methods that cannot secure the calculated resources.
3. The control device according to claim 1, wherein the high reliability method includes at least one of communication path redundancy, DCI gateway redundancy, flow control, and lost packet recovery.
4. A DCI gateway system comprising: a first DCI (Data Centric Infrastructure) gateway connected to a first APN (All Photonics Network); and a second DCI gateway connected to a second APN, wherein the first DCI gateway comprises: a receiving unit that receives data transmitted from a communication device via the first APN; and a processing unit that writes the data to a memory of a DCI cluster, and the second DCI gateway comprises: a receiving unit that receives the same data as the data transmitted from the communication device via the second APN; and a processing unit that reads the data written to the memory of the DCI cluster and, when packet loss is detected, writes a packet corresponding to the packet lost from the received data to the memory.
5. A DCI (Data Centric Infrastructure) gateway connected to a first APN (All Photonics Network) and a second APN, comprising: a first receiving unit that receives data transmitted from a communication device via the first APN; a second receiving unit that receives the same data as the data transmitted from the communication device via the second APN; and a processing unit that uses the data received by the first receiving unit and the data received by the second receiving unit to generate transmission data to be transmitted to a DCI cluster, and transmits the transmission data to the DCI cluster.
6. A configuration method performed by a control device to configure a DCI gateway connecting a DCI (Data Centric Infrastructure) cluster to an APN (All Photonics Network), comprising: selecting, using a route database, multiple route candidates for a communication route between the DCI cluster and a communication device that communicates with the DCI cluster, and selecting at least one high-reliability method according to user requirements; calculating the delay time for each route candidate for each high-reliability method using the transmission time of the communication route and the processing time of the DCI gateway; calculating, for each high-reliability method, the resources required for each route candidate based on the calculated delay time; determining, for each route candidate, one of the multiple route candidates and a high-reliability method based on the delay time and resources calculated for each high-reliability method; and configuring the DCI gateway and the APN using the determined communication route and high-reliability method.
7. A program that causes a computer to function as the control device according to any one of claims 1 to 3.
Citation Information
Patent Citations
Transmission device, communication control method, and communication system
JP2015097349A
Management device, communication control system, and communication control method
JP2023044932A