Computer System and Congestion Avoidance Method
The timing adjustment mechanism addresses network congestion in public clouds by individually delaying retransmissions, enhancing communication reliability and reducing delays in shared network systems.
Patent Information
- Application Number
- JP2023001858
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2043-01-10
AI Technical Summary
In public cloud environments, simultaneous retransmissions of lost packets due to OS-driven timeout mechanisms can lead to network congestion and prolonged communication delays, affecting application performance and reliability, especially in critical applications like life-and-death monitoring.
A timing adjustment mechanism that calculates and applies individualized delay times for retransmissions based on packet characteristics, preventing simultaneous retransmission congestion by varying the transmission timing of packets.
Effectively avoids network congestion and minimizes communication delays by dynamically adjusting retransmission times, ensuring stable and reliable communication even in shared network environments.
Smart Images

Figure 0007712306000001 
Figure 0007712306000002 
Figure 0007712306000003
Abstract
Description
Technical Field
[0001] The present invention relates to a computer system and a congestion avoidance method.
Background Art
[0002] Cloud computing is an attractive means that can provide IT services without having hardware assets and can obtain computing power with high flexibility against loads. In recent years, attempts have been made to implement business software that requires high reliability on the cloud, and it is expected that the use of the cloud will continue to spread in the future.
[0003] In a public cloud that can be used by various people, since different users are assigned to and share the same hardware, ensuring performance and reliability becomes an issue. Regarding resources (such as CPUs and memories) enclosed within a computer, separation between users is achieved by utilizing the hardware support functions, but for networks, although bandwidth restrictions and the like are imposed, there are parts in the generally used Ethernet technology where the influence on others cannot be avoided. Network congestion is one of the unavoidable influences, and it is known as a phenomenon in which some communications become impossible when the communication volume of users concentrates beyond the bandwidth that can be provided by the shared communication path.
[0004] In such congestion, it is known that loss of transmitted packets occurs. Since the receiving side cannot know that a packet has been lost, a timeout is set on the transmitting side in case there is no response from the receiving side for a certain period of time, and a procedure for retransmitting the packet after the timeout period has elapsed is defined in the standard specification. Such a retransmission operation is realized by driver software incorporated in the Operating System. However, in usage forms such as video streaming where delay is important, relying on the retransmission function of the OS may result in a timeout until retransmission being too long and affecting the provision of the function. In such cases, as described in Patent Document 1, by managing the timeout on the application side and actively retransmitting at a timing earlier than the retransmission function of the OS, the impact of packet loss can be mitigated. Packet loss due to congestion is addressed by appropriately implementing retransmission as described above, but there are also issues in the public cloud that cannot be solved by simple retransmission.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] Retransmission as a countermeasure against packet loss during convergence is an effective means. However, when convergence occurs in a communication path shared by a large number of transmission and reception entities (instances) such as a public cloud, a phenomenon may occur where consecutive retransmissions fail. This occurs when the retransmission operations of the instances are all operated according to the same rules. When packet loss occurs in multiple instances and retransmissions are started all at once after the timeout period has elapsed due to the retransmission function of the OS, a large number of retransmission packets will concentrate on a specific switch all at once, causing convergence again and resulting in packet loss. This phenomenon will occur repeatedly among the instances operating the retransmission operation in the same procedure, and as a result, a phenomenon will occur where an unexpectedly long time is required until retransmission is successful. From the perspective of the application, since the retransmission operation by the OS is hidden, it is observed from the application that the communication delay time suddenly extends abnormally. During this period, since the communication path is determined to be operating normally, abnormal communication delays during normal times may have an unexpected impact on the application, and in some cases, may lead to misdetection of failures.
[0007] Changing the timeout until retransmission for each instance can reduce the probability of being involved in the event. However, changing the operation of the OS's retransmission function affects the entire system, making it difficult to evaluate the impact. In addition, making modifications to change the OS's retransmission operation is a difficult option in terms of maintenance costs. For these reasons, in a computer system that is an instance connected to a public network, it has become an important issue to efficiently avoid convergence without making changes to the OS or applications.
Means for Solving the Problem
[0008] To achieve the above object, one of the typical computer systems of the present invention is connected to a network including a network switch, composed of a plurality of computers, and in a computer system that recovers the transfer operation of the lost data packet by a retransmission operation when the data packet is lost on the network, it includes a computer, software operating on the computer, and a timing adjustment mechanism existing between the computer and the network. The timing adjustment mechanism calculates a delay time for delaying the transmission of the data packet based on the characteristics of the data packet transmitted from the software, and delays the transmission of the data packet by the calculated delay time. Also, one of the typical congestion avoidance methods of the present invention is a congestion avoidance method for a computer system connected to a network including a network switch, composed of a plurality of computers, and recovering the transfer operation of the lost data packet by a retransmission operation when the data packet is lost on the network. A timing adjustment mechanism existing between the computer on which the software operates and the network acquires the characteristics of the data packet transmitted from the software, calculates a delay time for delaying the transmission of the data packet based on the characteristics of the data packet, and delays the transmission of the data packet by the calculated delay time.
Advantages of the Invention
[0009] According to the present invention, congestion can be efficiently avoided. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Best Mode for Carrying Out the Invention
[0011] The present invention is implemented in a network system composed of two or more computer systems and a network including network switches connecting them. FIG. 2 is a schematic configuration diagram of hardware. There are a plurality of servers 300 composed of a central processing unit 310, a main memory device 320, and a network interface 330. A network switch 410 composed of a network interface 330, a switch module 420, and a switch controller 430 is connected to the server by a network cable to form a network system. In principle, the present invention is assumed to be executed as software arranged in the main memory device 320 by the central processing unit 310 installed in the server. However, dedicated hardware performing the same operation may be separately installed in the server or the network switch. Alternatively, it may be implemented by being executed as firmware on the switch controller 430 of the network switch 410.
[0012] FIG. 3 is a diagram for explaining a phenomenon in which packet loss occurs due to congestion assumed in the present invention. Although there is an input / output buffer for each network interface 330, when packets 500 arrive at the output buffer from another network interface 330 at a rate faster than the rate at which data is transferred from the output buffer to another network switch 410 or server 300, and this situation continues, the output buffer becomes full and a situation occurs where packets 500 cannot be added to the output buffer. When the output buffer of a specific network interface 330 becomes full and packets 500 cannot be added, in many cases, the packets 500 that were attempted to be added are discarded and the packets 500 are lost. Such a situation is called congestion, and communication cannot be performed normally. The congestion state is resolved as the transfer of the packets 500 in the output buffer progresses over time and free space becomes available in the buffer, but the packets 500 that arrive before the resolution are lost.
[0013] The packets 500 lost in this way are detected by the fact that a normal response is not returned from the destination within a certain time according to the timeout setting of the source, and the standard operation is that the same packets are retransmitted and thus recovered. Therefore, even if congestion occurs, measures are taken so that communication does not immediately fail.
[0014] Figure 4 is a schematic diagram of a situation where a large number of packets 500 are being sent from a large number of servers 300 (described as Node in the figure) to a certain network switch 410. When the packets 500 sent from each server 300 overlap at a certain point, the load on the network switch 410, that is, the operating rate of the output buffer, may increase to the limit and enter a congestion state. At this time, multiple packets 500 may be lost at once. The lost packets 500 are recovered by each server 300 retransmitting them after a certain period of time. However, if the retransmission time is set the same for all of nodes A / B / C (300a, 300b, 300c) due to circumstances such as all using the same type of operating system, the packets 500 will be retransmitted at the same timing. At this time, depending on the state of the network, re-congestion may occur due to the retransmission, and the phenomenon of the packets 500 being lost again may occur. During the repeated occurrence of this phenomenon several times, there may be a case where, like node C (300c) in Figure 4, a deviation occurs in the time until retransmission, and it is possible to avoid being involved in a simultaneous retransmission. Also, there may be a case where some of the retransmitted packets 500 reach without being lost even when the traffic from other nodes 300 decreases and congestion occurs. In this way, the congestion caused by retransmission gradually subsides, but there may also be a case where, like node A (300a), it is repeatedly involved in re-congestion and continuously fails to retransmit. Such an event where retransmission continuously fails and retransmission is repeated many times will not be recognized by the application 100 that is communicating and will simply be observed as if the communication delay time has increased. In an application 100 where the communication delay time has a fatal impact, the occurrence of such an event may have a fatal impact on the continued operation of the application 100. For example, in the case of life-and-death monitoring communication, there may be a case where, even though the partner server is operating, it is recognized as a failure due to the long delay in the life-and-death confirmation communication. Such problems cannot be addressed by a simple retransmission operation.
[0015] To address abnormal delays in communication due to continuous retransmission failures, adjusting the transmission timing is an effective means. Since the simultaneous retransmission by a large number of servers 300 during retransmission causes re-convergence, adjusting the retransmission time prevents the sudden concentration of packets 500 in the network switch 410 and suppresses re-convergence.
[0016] FIG. 5 is a conceptual diagram for adjusting the transmission delay time. In addition to the retransmission delay caused by the default operation of a normal OS or the like, a delay for avoiding convergence and a random delay are added to prevent the concentration of packets 500, thereby adjusting the transmission interval of packets. Since the appropriate convergence avoidance delay time varies depending on the network state, it is desirable to adjust it dynamically. For example, when retransmissions fail continuously more than a threshold number of times, the convergence avoidance time is extended, and when the number of successful retransmissions reaches or exceeds the threshold, the convergence avoidance time is decreased for adjustment. By using the timing adjustment system 1000 that adjusts the transmission timing of such packets 500, continuous retransmission failures and the accompanying abnormal increase in communication delay can be prevented.
[0017] In the following description, there may be one or more timing adjustment systems for each server 300 or instance. An individual timing adjustment system 1000 may be applied to each application 100, or multiple applications 100 may share it. Also, in the following description, when explaining without distinguishing between elements of the same type, the common part of the reference signs may be used, and when distinguishing between elements of the same type, individual reference signs may be used.
Example
[0018] (Basic Example) FIG. 1 shows an overview of the flow of packets containing data in an embodiment of the present invention. In order to send the data generated by application 100a to application 100b operating elsewhere, packets containing the data are generated, and the operating system 200 is instructed to send them using system calls or the like. The operating system 200 sends the packet 500 toward the network 400 in order to send the packet to the destination according to the instruction, but by passing through the timing adjustment system 1000 for adjusting the transmission delay on the transmission path, the continuous retransmission failure due to re-convergence is suppressed. The timing adjustment system 1000 may operate as a part of the function of the operating system 200, or may operate as a dedicated device physically separated from the server 300 as a dedicated device.
[0019] Figure 6 shows the configuration of the timing adjustment system 1000. The timing adjustment system 1000 is composed of the following components. The retransmission confirmation mechanism 1100 determines whether the received packet 500 is a retransmission packet based on the information in the session state table 1200, and issues an instruction to the packet switch 1300 to distribute it to the path for adjusting the transmission timing if it is a retransmission packet. The session state table 1200 provides the retransmission confirmation mechanism 1100 with information regarding the state of the session to which the packet was transmitted from the information stored in the header of the packet. The packet switch 1300 controls whether to pass the packet through the path as it is or to pass it through the path for adjusting the transmission timing based on the signal from the retransmission confirmation mechanism 1100. The delay calculator 1400 queries the delay setting table 1700 using the header information of the packet 500, calculates how much to delay the transmission of the packet 500, and notifies the delay buffer 1800 of the delay time. The delay buffer 1800 accumulates the received packet 500 based on the delay time of the packet notified by the delay calculator 1400 and transmits it after the delay time has elapsed. The delay adjuster 1500 queries the delay history table 1600 and the delay setting table 1700 based on the header information of the packet, adjusts the current delay setting based on the history information, and sets the new delay setting in the delay setting table 1700.
[0020] Figure 7 shows the operation flowchart of the timing adjustment system 1000. The flowchart shown in Figure 7 is activated by receiving a packet or by a delay timer. When a packet is received, the process starts from step 10000, and when it is activated by a delay timer, the process starts from step 10100.
[0021] Step 10000: The timing adjustment system 1000 sends the received packet to the retransmission confirmation mechanism 1100 and proceeds to step 10010. Step 10010: The retransmission confirmation mechanism 1100 checks whether the packet has been retransmitted. If it has been retransmitted, proceed to step 10020. If it has not been retransmitted, proceed to step 10120. Step 10020: The packet switch 1300 sends the packet to the delay calculator 1400. Then, proceed to step 10030. Step 10030: The delay calculator 1400 calculates the delay time for the sent packet, sets the calculated delay time for the packet, and registers it in the delay buffer 1800. Then, proceed to step 10040. Step 10040: The packet switch 1300 sends the packet to the delay adjuster 1500 to perform delay adjustment. Then, proceed to step 10200.
[0022] Step 10120: The timing adjustment system 1000 transmits the packet. Then, proceed to step 10200. Step 10200: The timing adjustment system 1000 determines whether one or more packets are registered in the delay buffer 1800. If not, return to the standby state. If so, proceed to step 10210. Step 10210: The timing adjustment system 1000 searches for the packet with the shortest waiting time among the packets registered in the delay buffer 1800 and proceeds to step 10220. Step 10220: The timing adjustment system 1000 sets a delay timer for the packet with the shortest waiting time and returns to the standby state.
[0023] Step 10100: If the process is activated by the delay timer, the timing adjustment system 1000 extracts the packet that has waited for the set delay time from the delay buffer and proceeds to step 10120.
[0024] Figure 8 shows the configuration of the retransmission confirmation mechanism 1100. The retransmission confirmation mechanism is composed of a header separation unit 1110 that extracts the header part from the packet 500, and a comparator 1120 that compares the final sequence number obtained from the session state table 1200 with the sequence number of the packet. When the final sequence number is equal to or greater than the sequence number included in the packet, the comparator 1120 determines that the packet has been retransmitted and sends a signal indicating that the packet is the one for which transmission delay should be inserted. Due to the limitation of the storage data size, the sequence number that can be stored in the packet 500 may return to zero after one cycle. To address this, the data size for storing the final sequence number in the session state table 1200 may be ensured to be larger than the data size of the packet. When detecting that one cycle has occurred, the upper digits higher than the data size of the packet may be added for identification.
[0025] Figure 9 shows the operation flowchart of the retransmission confirmation mechanism 1100. When the retransmission confirmation mechanism 1100 receives a packet, it sequentially executes the processes from step 11000 to step 11050. Step 11000: The header separation unit 1110 of the retransmission confirmation mechanism 1100 extracts header information (destination IP address, source port number, destination port number) from the packet and sends it to the session state table. Then, it proceeds to step 11010. Step 11010: The retransmission confirmation mechanism 1100 receives the sequence number at the time of the last transmission from the session state table. Then, it proceeds to step 11020. Step 11020: The comparator 1120 compares the sequence number of the last transmission (SeqL) with the sequence number of the received packet (SeqC). Then, it proceeds to step 11030.
[0026] Step 11030: As a result of the comparison by the comparator 1120, if SeqC < SeqL, it is determined as retransmission and proceeds to step 11040. If SeqC ≥ SeqL, it proceeds to step 11050. Step 11040: The retransmission confirmation mechanism 1100 sends a signal to delay the packet and returns to the standby state. Step 11050: The retransmission confirmation mechanism 1100 sends a signal to pass the packet without delay and returns to the standby state.
[0027] Figure 10 shows the configuration of the delay calculator 1400. The delay calculator 1400 is composed of a header separation unit 1110, a random number generator 1410, a multiplier that multiplies the random delay width, and an adder that adds the congestion avoidance delay. The header information extracted by the header separation unit 1110 is sent to the delay setting table 1700, and the setting values of the random delay width and the congestion avoidance delay time are received. The random delay width is multiplied by the random number generated by the random number generator 1410 to calculate the random delay time. The random delay time and the congestion avoidance delay time are added to output the transmission delay time of the packet to the outside.
[0028] Figure 11 shows the operation flowchart of the delay calculator 1400. When the delay calculator 1400 receives a packet, it sequentially executes the processes of steps 12000 to 12050. Step 12000: The header separation unit 1110 of the delay calculator 1400 extracts the header information (destination IP address, source port number, destination port number) from the packet and sends it to the delay setting table 1700. Then, it proceeds to step 12010. Step 12010: The delay calculator 1400 receives the random delay width and the congestion avoidance delay value from the delay setting table 1700. Then, it proceeds to step 12020. Step 12020: The delay calculator 1400 obtains a random number from the random number generator 1410. Then, it proceeds to step 12030. Step 12030: The delay calculator 1400 multiplies the random number and the random delay width to calculate the random delay value. Then, it proceeds to step 12040. Step 12040: The delay calculator 1400 adds the congestion avoidance delay value to the random delay value to calculate the packet delay time. Then, it proceeds to step 12050. Step 12050: The delay computer 1400 sends out the packet delay time and returns to the standby state.
[0029] Figure 12 shows the configuration of the delay adjuster 1500. The delay adjuster 1500 is composed of a header separation unit 1110, a block for adjusting the congestion avoidance time in the increasing direction, a block for adjusting in the decreasing direction, and a block for sending a request to update the delay setting table 1700 by combining the adjusted congestion avoidance time with the packet header information. The block for adjusting the congestion avoidance time of the delay adjuster 1500 in the increasing direction includes a threshold 1510 for increasing the congestion avoidance time with respect to the number of packet retransmissions, a counter for counting the number of times the retransmission count condition in the increasing direction is satisfied, a threshold 1520 for making a determination to actually increase with respect to the retransmission count condition in the increasing direction, and an increase width 1530. The block for adjusting the congestion avoidance time of the delay adjuster 1500 in the decreasing direction includes a threshold 1550 for decreasing the congestion avoidance time with respect to the number of packet retransmissions, a counter for counting the number of times the retransmission count condition in the decreasing direction is satisfied, a threshold 1560 for making a determination to actually decrease with respect to the retransmission count condition in the decreasing direction, and a decrease width 1570.
[0030] Figure 13 shows the operation flowchart of the delay adjuster 1500. When the delay adjuster 1500 receives a packet, it sequentially executes the processes of steps 13000 to 13240. Step 13000: The header separation unit 1110 of the delay adjuster 1500 extracts header information (destination IP address, source port number, destination port number, sequence number) from the packet and sends it to the delay history table 1600. Then, it proceeds to step 13010. Step 13010: The delay adjuster 1500 receives the number of delays from the delay history table 1600. Then, it proceeds to step 13020. Step 13020: The header separation unit 1110 of the delay adjuster 1500 extracts header information (destination IP address, source port number, destination port number) from the packet and sends it to the query port of the delay setting table 1700. Then, it proceeds to step 13030. Step 13030: The delay adjuster 1500 receives the convergence avoidance delay value from the delay setting table 1700. Then, it proceeds to step 13100.
[0031] Step 13100: The delay adjuster 1500 determines whether the number of delays is greater than or equal to the retransmission threshold in the increasing direction. If the number of delays is greater than or equal to the retransmission threshold in the increasing direction, it proceeds to step 13110. If the number of delays is less than the retransmission threshold in the increasing direction, it proceeds to step 13200. Step 13110: The delay adjuster 1500 increments the increasing direction counter by 1. Then, it proceeds to step 13120. Step 13120: The delay adjuster 1500 determines whether the increasing direction counter is greater than or equal to the threshold. If the increasing direction counter is greater than or equal to the threshold, it proceeds to step 13130. If the increasing direction counter is less than the threshold, it returns to the standby state.
[0032] Step 13130: The delay adjuster 1500 resets the increasing direction counter to 0. Then, it proceeds to step 13140. Step 13140: The delay adjuster 1500 adds the convergence avoidance delay increase width value to the current convergence avoidance delay value to calculate a new convergence avoidance delay value. Then, it proceeds to step 13150. Step 13150: The delay adjuster 1500 sends the header information (destination IP address, source port number, destination port number) extracted from the packet and the new convergence avoidance delay value to the update port of the delay setting table 1700. Then, it returns to the standby state.
[0033] Step 13200: The delay adjuster 1500 determines whether the number of delays is greater than or equal to the retransmission threshold in the decreasing direction. If the number of delays is greater than or equal to the retransmission threshold in the decreasing direction, it proceeds to step 13210. If the number of delays is less than the retransmission threshold in the decreasing direction, it returns to the standby state. Step 13210: The delay adjuster 1500 increments the decreasing direction counter by 1. Then, it proceeds to step 13220. Step 13220: The delay adjuster 1500 determines whether the decrement counter is greater than or equal to the threshold value. If the decrement counter is greater than or equal to the threshold value, proceed to step 13230. If the decrement counter is less than the threshold value, return to the standby state.
[0034] Step 13230: The delay adjuster 1500 resets the decrement counter to 0. Then, proceed to step 13240. Step 13240: The delay adjuster 1500 adds the convergence avoidance delay reduction width value to the current convergence avoidance delay value to calculate a new convergence avoidance delay value. Then, proceed to step 13150.
[0035] Figure 14 shows the configuration of the session state table 1200. The session state table stores combinations of the destination IP address, source port number, destination port number, and the final sequence number. It receives a combination of the destination IP address, source port number, destination port number, and sequence number from the outside, searches for a row where the destination IP address, source port number, and destination port number match, and updates the final sequence number if the final sequence number in the matching row is smaller than the received sequence number. If the final sequence number is larger, it does not update. Then, it outputs the final sequence number of the matching row.
[0036] Figure 15 shows the operation flowchart of the session state table 1200. When the session state table 1200 receives a query, it sequentially executes the processes of steps 14000 to 14060. Step 14000: The session state table 1200 searches for an entry that matches the destination IP address, source port number, and destination port number included in the query. Then, proceed to step 14010. Step 14010: The session state table 1200 determines whether an entry that matches the search condition exists. If an entry that matches the search condition exists, proceed to step 14020. If an entry that matches the search condition does not exist, proceed to step 14050.
[0037] Step 14020: The session state table 1200 compares the sequence number of the discovered entry with the sequence number included in the query. Then, it proceeds to step 14030. Step 14030: The session state table 1200 determines whether the sequence number of the query is greater than the sequence number of the entry. If the sequence number of the query is greater than the sequence number of the entry, it proceeds to step 14040. If the sequence number of the query is not greater than the sequence number of the entry, it proceeds to step 14060. Step 14040: The session state table 1200 updates the sequence number of the entry with the sequence number of the query. Then, it proceeds to step 14060.
[0038] Step 14050: The session state table 1200 registers the destination IP address, source port number, destination port number, and sequence number included in the query in the table. Then, it proceeds to step 14060. Step 14060: The session state table 1200 outputs the sequence number included in the query as the last sequence number. Then, it returns to the standby state.
[0039] FIG. 16 shows the configuration of the delay setting table 1700. The delay setting table 1700 stores the destination IP address, source port number, destination port number, congestion avoidance delay time, and random delay width. When receiving the destination IP address, source port number, and destination port number from the outside through the query port, it searches for the information that matches the destination IP address, source port number, and destination port number from the stored information, and outputs the congestion avoidance delay time and random delay width of the matching row. If no matching row is found, it outputs the default congestion avoidance delay time and random delay width. When receiving the combination of the destination IP address, source port number, destination port number, and congestion avoidance delay time from the outside through the update port, it searches for the information that matches the destination IP address, source port number, and destination port number from the stored information, and updates the congestion avoidance delay time of the matching row. If no matching row is found, it adds a new row in combination with the default random delay width.
[0040] FIG. 17 shows the operation flowchart of the delay setting table 1700. When the delay setting table 1700 receives a query, it sequentially executes the processes from step 15000 to step 15030. When the delay setting table 1700 receives an update instruction, it sequentially executes the processes from step 15100 to step 15130.
[0041] Step 15000: The delay setting table 1700 searches for an entry that matches the destination IP address, source port number, and destination port number included in the query. Then, it proceeds to step 15010. Step 15010: The delay setting table 1700 determines whether there is an entry that meets the search criteria. If there is an entry that meets the search criteria, it proceeds to step 15020. If there is no entry that meets the search criteria, it proceeds to step 15030. Step 15020: The congestion avoidance delay setting table 1700 outputs the congestion avoidance delay value and the random delay width value included in the discovered entry. Then, it returns to the standby state. Step 15030: The congestion avoidance delay setting table 1700 outputs the system default congestion avoidance delay value and the random delay width value. Then, it returns to the standby state.
[0042] Step 15100: The congestion avoidance delay setting table 1700 searches for an entry that matches the destination IP address, source port number, and destination port number included in the update instruction. Then, it proceeds to step 15110. Step 15110: The congestion avoidance delay setting table 1700 determines whether there is an entry that matches the search condition. If there is an entry that matches the search condition, it proceeds to step 15120. If there is no entry that matches the search condition, it proceeds to step 15130. Step 15120: The congestion avoidance delay setting table 1700 updates the congestion avoidance delay value included in the discovered entry with the value included in the update instruction. Then, it returns to the standby state. Step 15130: The congestion avoidance delay setting table 1700 creates a new entry by adding the default random delay width value to the destination IP address, source port number, destination port number, and congestion avoidance delay value included in the update instruction, and registers it in the table. Then, it returns to the standby state.
[0043] Figure 18 shows the configuration of the delay history table 1600. The delay history table 1600 stores multiple rows with the destination IP address, source port number, destination port number, sequence number, and retransmission count counter as one row. When receiving the destination IP address, source port number, destination port number, and sequence number from the outside, it searches for a row that matches the information, adds 1 to the value of the retransmission count counter of the matching row, and outputs the added value. If there is no matching row, it newly adds a row with the retransmission count counter set to 1 to the received information.
[0044] Figure 19 shows the operation flowchart of the delay history table 1600. When the delay history table 1600 receives a Query, it starts processing from step 16000. Step 16000: The delay history table 1600 searches for an entry that matches the destination IP address, source port number, destination port number, and sequence number included in the query. Then, it proceeds to step 16010. Step 16010: The delay history table 1600 determines whether there is an entry that matches the search condition. If there is no entry that matches the search condition, it proceeds to step 16020. If there is no entry that matches the search condition, it proceeds to step 16100.
[0045] Step 16020: The delay history table 1600 increments the retransmission count (Count) of the discovered entry by 1 and outputs the incremented count as the retransmission count. Then, it returns to the standby state. Step 16030: The delay history table 1600 creates an entry with the retransmission count set to 1 from the destination IP address, source port number, destination port number, and sequence number included in the query, registers it in the table. Then, it proceeds to step 16110. Step 16110: The delay history table 1600 outputs 1 as the retransmission count. Then, it returns to the standby state.
[0046] Figure 20 shows an example of the setting screen of the delay setting table 1700. The congestion avoidance delay time and random delay width for each destination can be changed, for example, from a GUI (Graphical User Interface). Also, when there is no row for the destination, it can be newly set. The congestion avoidance delay time is automatically adjusted, but the setting can also be changed externally through such an interface. In this example, an example of a GUI is shown, but this setting may also be changed through interfaces such as a CUI (Commandline User Interface) or a REST (REpresentational State Transfer) API.
[0047] FIG. 21 shows an example of a setting screen of the delay adjuster 1500. Regarding the block for adjusting the convergence avoidance time in the increasing direction, a threshold 1510 for increasing the convergence avoidance time with respect to the number of retransmissions of a packet, a threshold 1520 for making a determination to actually increase with respect to the increasing direction retransmission count condition, and an increase width 1530 can be set. Regarding the block for adjusting the convergence avoidance time in the decreasing direction, a threshold 1550 for decreasing the convergence avoidance time with respect to the number of retransmissions of a packet, a threshold 1560 for making a determination to actually decrease with respect to the decreasing direction retransmission count condition, and a decrease width 1570 can be set. Although an example of a GUI is shown in this example, this setting may be changed through an interface such as a CUI (Commandline User Interface) or a REST (REpresentational State Transfer) API.
Example
[0048] (Example of obtaining the convergence avoidance time from a network switch) In Example 2, the adjustment of the convergence avoidance time is shown when the average convergence time is obtained from the network 400, rather than the automatic adjustment by observing the retransmission execution state shown in Example 1.
[0049] FIG. 22 shows an overview of the flow of packets including the data in Example 2. The difference from Example 1 is that a convergence time calculator 440, which is a mechanism for calculating the average convergence time, is added to the network 400, and in response to an inquiry from the timing adjustment system 1000, the average convergence time observed in the network 400 is notified, and based on this, the timing adjustment system 1000 performs an operation of delaying the transmission of packets.
[0050] Figure 23 is a schematic diagram of the hardware configuration in Example 2. As a difference from Example 1, a congestion time calculator 440 is added to the network switch 410. The congestion time calculator 440 monitors the operating status of the network switch 410, measures the congestion time for each network interface 330, and calculates the average congestion time by averaging them overall.
[0051] Figure 24 is a schematic diagram of congestion time measurement. When a packet arrives at a certain network interface 330, the transmission buffer of the network interface is full, and when packet loss starts to occur, the time until there is space in the transmission buffer is defined as the congestion time. Different congestion times are measured for each transmission and reception port of the network interface, so the average value is calculated for the entire switch to be able to respond to inquiries from the timing adjustment system.
[0052] Figure 25 is an operation flowchart of the congestion time calculator 440. When the congestion time calculator 440 detects that port queue addition is not possible, it starts the processing from step 20000. Step 20000: The congestion time calculator 440 refers to the port congestion flag to determine whether the port is in a congested state. If the port is in a congested state, it returns to the standby state. If the port is not in a congested state, it proceeds to step 20010. Step 20010: The congestion time calculator 440 sets the port congestion flag and proceeds to step 20020. Step 20020: The congestion time calculator 440 starts the port congestion time measurement timer. Then, it returns to the standby state.
[0053] When the congestion time calculator 440 detects port queue addition, it starts the processing from step 20100. Step 20100: The congestion time calculator 440 refers to the port congestion flag to determine whether the port is in a congested state. If the port is in a congested state, it proceeds to step 20110. If the port is not in a congested state, it returns to the standby state. Step 20110: The convergence time calculator 440 resets the port convergence flag and proceeds to step 20120. Step 20120: The convergence time calculator 440 stops the port convergence time measurement timer and calculates the convergence time T curr . Then, it proceeds to step 20130. Step 20130: The convergence time calculator 440 calculates a new average convergence time T' avg from the previous average convergence time T avg and the current convergence time T. T avg= αT curr + (1 - α)T avg Then, it proceeds to step 20140. Step 20140: The convergence time calculator 440 calculates and updates the average convergence time of all ports. Then, it returns to the standby state.
[0054] When the convergence time calculator 440 receives a convergence time inquiry, it starts the processing of step 20200. Step 20200: It transmits the average convergence time of all ports to the inquirer. Then, it returns to the standby state.
[0055] Figure 26 is a configuration diagram of the timing adjustment system 1000 in Example 2. The difference from Example 1 is that there is no delay adjuster 1500, and the delay calculator 1400 is changed to inquire about the convergence avoidance time from the convergence time calculator 440. The operations of other components are the same as in Example 1. Also, regarding the operation flowchart of the timing adjustment system 1000, there is no change compared to Example 1 because only the internal operation of the delay calculator 1400 changes.
[0056] Figure 27 is a configuration diagram of the delay calculator 1400 in Example 2. The difference from Example 1 is that, regarding the convergence avoidance time of the parameter obtained from the delay setting table 1700 in Example 1, in Example 2, it is changed to be acquired from the convergence time calculator 440. Other operations are the same as in Example 1.
[0057] Figure 28 is an operation flowchart of the delay calculator 1400 in Embodiment 2. When the delay calculator 1400 in the embodiment receives a packet, it starts processing from step 12000. Step 12000: The header separation unit 1110 of the delay calculator 1400 extracts header information (destination IP address, source port number, destination port number) from the packet and sends it to the delay setting table 1700. Then, it proceeds to step 21010. Step 21010: The delay calculator 1400 receives a random delay width from the delay setting table 1700. Then, it proceeds to step 21020. Step 21020: The delay calculator 1400 queries the congestion time calculator 440 for the average congestion time. Then, it proceeds to step 21030. Step 21030: The delay calculator 1400 receives the average congestion time from the congestion time calculator 440. Then, it proceeds to step 12020. The subsequent processing is the same as that in Embodiment 1, so the description is omitted.
Embodiment
[0058] (Embodiment of statistically calculating the default value of the congestion avoidance time) In Embodiment 3, a configuration is shown in which the default value before the automatic adjustment of the congestion avoidance time is calculated from the previous congestion avoidance time setting state and the statistical value of the number of consecutive retransmission occurrences at that time. In Embodiment 1, the default value was a fixed value regardless of the destination of the packet, but in this embodiment, by calculating a default value of the congestion avoidance time sufficient to suppress the occurrence of consecutive retransmissions from the past statistical information for each destination, the occurrence of early consecutive retransmissions is suppressed.
[0059] FIG. 29 is a configuration diagram of the timing adjustment system 1000 in Embodiment 3. As a difference from Embodiment 1, a delay statistic calculator 2000 and a delay statistic table 2100 are added. The delay statistic calculator 2000 records in the delay statistic table 2100 the relationship between the congestion avoidance delay time set for a certain packet and the maximum number of retransmissions. The delay statistic calculator 2000 receives an inquiry about the default congestion avoidance time from the delay setting table 1700, calculates the minimum congestion avoidance time with the smallest maximum number of retransmissions from the records in the delay statistic table 2100, and returns it to the delay setting table 1700.
[0060] FIG. 30 is an operation flowchart of the timing adjustment system 1000 in Embodiment 3. The timing adjustment system 1000 in Embodiment 3 is different from Embodiment 1 in that after step 10040, steps 30000 and 30010 are executed and then it proceeds to step 10200. Since the rest is the same as in Embodiment 1, the description is omitted. Step 30000: The timing adjustment system 1000 sends a packet to the delay statistic calculator 2000. Then, it proceeds to step 30010. Step 30010: The delay statistic calculator 2000 updates the delay statistic table 2100 and transmits a default delay value to the delay setting table 1700. Then, it proceeds to step 10200.
[0061] FIG. 31 is a configuration diagram of the delay statistics computer 2000 in Embodiment 3. The delay statistics computer 2000 includes a header separation unit 1110, a delay statistics table update block, and a default congestion delay time calculation unit 2010. The delay statistics table update block includes a threshold 2020 for the number of retransmissions of a packet, a comparator that compares the threshold with the number of retransmissions of the packet, and a merge unit that combines the congestion avoidance delay time, the destination IP address, the source port number, and the destination port number included in the packet based on the comparison result. The default congestion delay time calculation unit 2010 receives a plurality of pairs of a congestion avoidance delay time and the number of retransmission occurrences obtained by sending the destination IP address, the source port number, and the destination port number extracted from the packet to the query port of the delay statistics table 2100, selects a combination in which the number of retransmission occurrences is the minimum and the congestion avoidance delay time is the minimum from among them, and outputs the congestion avoidance delay time as a default congestion avoidance delay time to the delay setting table 1700.
[0062] FIG. 32 is an operation flowchart of the delay statistics computer 2000 in Embodiment 3. When the delay statistics computer 2000 receives a packet, it starts processing from step 31000. Step 31000: The header separation unit 1110 of the delay statistics computer 2000 extracts header information (destination IP address, source port number, destination port number, sequence number) from the packet and sends it to the delay history table 1600. Then, it proceeds to step 31010. Step 31010: The delay statistics computer 2000 receives the number of delays from the delay history table 1600. Then, it proceeds to step 31020. Step 31020: The header separation unit 1110 of the delay statistics computer 2000 extracts header information (destination IP address, source port number, destination port number) from the packet and sends it to the query port of the delay setting table 1700. Then, it proceeds to step 31030. Step 31030: The delay statistics computer 2000 receives the congestion avoidance delay value from the delay setting table 1700. Then, it proceeds to step 31040.
[0063] Step 31040: The delay statistics calculator 2000 determines whether the number of delays is equal to or greater than a threshold value. If it is equal to or greater than the threshold value, proceed to step 31050. If it is less than the threshold value, proceed to step 31100. Step 31050: The header separation unit 1110 of the delay statistics calculator 2000 extracts header information (destination IP address, source port number, destination port number) from the packet, combines it with the congestion avoidance delay value, and sends it to the increment port of the delay statistics table 2100. Then, proceed to step 31100.
[0064] Step 31100: The header separation unit 1110 of the delay statistics calculator 2000 extracts header information (destination IP address, source port number, destination port number) from the packet and sends it to the Query port of the delay statistics table 2100. Then, proceed to step 31110. Step 31110: The default congestion delay time calculation unit 2010 of the delay statistics calculator 2000 receives a set of pairs of congestion avoidance time and retransmission count from the delay statistics table 2100. Then, proceed to step 31120.
[0065] Step 31120: The default congestion delay time calculation unit 2010 determines whether there is one or more pairs of delay avoidance time and retransmission count. If there is one or more pairs, proceed to step 31130. If not, proceed to step 31140. Step 31130: The default congestion delay time calculation unit 2010 searches for the pair with the minimum retransmission count and outputs the congestion avoidance time of that pair as the default congestion avoidance delay time. Then, return to the standby state. Step 31140: The default congestion delay time calculation unit 2010 outputs the system default congestion avoidance delay time. Then, return to the standby state.
[0066] FIG. 33 is a configuration diagram of the delay statistics table 2100. The delay statistics table holds a plurality of rows each consisting of a destination IP address, a source port number, a destination port number, a congestion avoidance delay time, and a retransmission count. When a query including a destination IP address, a source port number, and a destination port number comes to the query port, it searches for a row in which the destination IP address, the source port number, and the destination port number match, and outputs a pair of the congestion avoidance delay time and the retransmission count for all the rows that meet the conditions. When an increment instruction including a destination IP address, a source port number, a destination port number, and a congestion avoidance delay time comes to the increment port, it searches for a row in which the destination IP address, the source port number, the destination port number, and the congestion avoidance delay time match, and increments the retransmission count of the matching row by 1. If there is no row that meets the conditions, it newly adds a row with the destination IP address, the source port number, the destination port number, the congestion avoidance delay time, and the retransmission count set to 1.
[0067] FIG. 34 is an operation flowchart of the delay statistics table 2100. When the delay statistics table 2100 receives a query, it starts processing from step 32000. Step 32000: The delay statistics table 2100 searches for an entry that matches the destination IP address, the source port number, and the destination port number included in the query. Then it proceeds to step 32010. Step 32010: The delay statistics table 2100 determines whether there is an entry that meets the search conditions. If there is an entry that meets the search conditions, it proceeds to step 32020. If there is no entry that meets the search conditions, it proceeds to step 32030.
[0068] Step 32020: The delay statistics table 2100 extracts and outputs a pair of the congestion avoidance time and the count from the entry that meets the conditions. If there are multiple matches, it outputs all the matching entries. Then it returns to the standby state. Step 32030: The delay statistics table 2100 outputs a signal representing an empty set. Then, it returns to the standby state.
[0069] When the delay statistics table 2100 receives an increment instruction, it starts the processing from step 32100. Step 32100: The delay statistics table 2100 searches for an entry that matches the destination IP address, source port number, destination port number, and congestion avoidance delay time included in the increment instruction. Then, it proceeds to step 32110. Step 32110: The delay statistics table 2100 determines whether there is an entry that matches the search condition. If there is an entry that matches the search condition, it proceeds to step 32120. If there is no entry that matches the search condition, it proceeds to step 32130.
[0070] Step 32120: The delay statistics table 2100 increments the count of the entry that matched the search by 1 and updates the entry. Then, it returns to the standby state. Step 32030: The delay statistics table 2100 creates a new entry with a count of 1 using the destination IP address, source port number, destination port number, and congestion avoidance delay time included in the increment instruction, and adds it to the table. Then, it returns to the standby state.
[0071] FIG. 35 shows a configuration diagram of the delay setting table 1700 in Example 3. The difference from Example 1 is that a path for obtaining from the delay statistics calculator 2000 is added for the default congestion avoidance delay time to be referred to when there is no row in the table that matches the conditions of the query. Except for the operation of referring to the default value of the congestion avoidance delay time, the operation is the same as that in Example 1.
[0072] FIG. 36 is an operation flowchart of the delay setting table 1700 in Example 3. In this flowchart, as a result of the determination in step 15010, when there is no entry that matches the conditions of the query, it proceeds to step 33000. Step 33000: The delay setting table 1700 outputs the default delay time passed from the delay statistic computer 2000 as the congestion avoidance delay value, together with the random delay width value. Then, it returns to the standby state. Since other processes are the same as those of the delay setting table 1700 in the first embodiment, the description thereof is omitted.
[0073] As described above, the disclosed computer system is connected to a network including a network switch, is composed of a plurality of computers, and in a computer system that recovers the transfer operation of the lost data packet by retransmission when the data packet is lost on the network, it includes a computer (200), software (100) operating on the computer, and a timing adjustment mechanism (1000) existing between the computer and the network. The timing adjustment mechanism calculates a delay time for delaying the transmission of the data packet based on the characteristics of the data packet transmitted from the software, and delays the transmission of the data packet by the calculated delay time. By introducing the mechanism of the present invention, when congestion occurs, retransmission packets are transmitted at different delay times from a large number of instances, preventing them from being involved in congestion caused by the simultaneous transmission of retransmission packets. Also, even when the duration of the congestion changes depending on the operating conditions of other instances, the automatic adjustment mechanism automatically adjusts an appropriate delay time to avoid congestion, adjusts it to the minimum necessary delay time according to changes in the communication situation, and minimizes the impact caused by inserting a delay. Therefore, congestion can be efficiently avoided.
[0074] Also, the timing adjustment mechanism determines whether to delay the transmission of the data packet based on the characteristics of the data packet. Also, the timing adjustment mechanism calculates the delay time by adding a random value to a fixed time determined from the characteristics of the data packet. In this way, the disclosed computer system can easily obtain an appropriate delay time.
[0075] In addition, when the timing adjustment mechanism detects that, among the delay times, for a fixed time determined from the characteristics of the data packet, data packets with the same content are continuously retransmitted more than a first threshold number of times, the detection count of continuous retransmission for the destination of the data packet is added. When it is detected that the detection count of continuous retransmission is equal to or more than a second threshold number of times, the fixed time of the delay time set for transmission to the destination is increased by a set width. When it is detected that data packets with the same content are retransmitted less than a third threshold number of times, the detection count of non - continuous retransmission for the destination of the data packet is added. When it is detected that the detection count of non - continuous retransmission is equal to or more than a fourth threshold number of times, the fixed time of the delay time set for transmission to the destination is decreased by a set width. By doing so, the fixed time is automatically adjusted according to the state of the network. Therefore, it is possible to obtain a delay time that conforms to the state of the network and efficiently avoid congestion.
[0076] In addition, the network is a publicly available network shared by a plurality of computers. Even in a publicly available network, without having information about what other instances exist, the disclosed computer system can achieve congestion avoidance.
[0077] In addition, the timing adjustment mechanism may be configured to receive a statistical value related to the congestion occurrence time from the network switch and calculate the time to delay transmission from the statistical value. In this configuration, by obtaining information about congestion from the network switch, the computer system can avoid congestion with high precision. Note that the configuration of obtaining information about congestion from the network switch and calculating the delay time is preferably selectively applied to a specific type of instance. This is to avoid a situation where a plurality of instances that have obtained information from the network switch calculate similar delay times and congestion occurs repeatedly. The specific type is, for example, an instance responsible for life - death monitoring.
[0078] Further, the timing adjustment mechanism records the correspondence relationship between the delay time and the number of times of consecutive retransmission, and from the correspondence relationship, sets the delay time when the number of times of consecutive retransmission is the smallest as the default value of the delay time, and calculates the delay time to be applied to the data packet based on the default value. With this configuration, an appropriate delay time can be obtained quickly.
[0079] As described above, several embodiments have been described, but these are examples for explaining the present invention, and the scope of the present invention is not limited only to these embodiments. The present invention can be implemented in various other forms. For example, although an embodiment of providing a through path for a packet that is not retransmitted has been described, instead of providing a separate path for a packet that is not retransmitted, it may be implemented as a configuration in which the delay time is set to be extremely small (for example, 0).
Explanation of Reference Numerals
[0080] 1000: Timing adjustment system 1100: Retransmission confirmation mechanism 1200: Session state table 1300: Packet switch 1400: Delay calculator 1500: Delay adjuster 1600: Delay history table 1700: Delay setting table 1800: Delay buffer 2000: Delay statistical calculator 2100: Delay statistical table
Claims
1. In a computer system that is connected to a network including a network switch, is composed of a plurality of computers, and, when a data packet is lost on the network, recovers the transfer operation of the lost data packet by a retransmission operation, a computer; software operating on the computer; and a timing adjustment mechanism existing between the computer and the network, wherein the timing adjustment mechanism determines whether to delay the transmission of the data packet based on the characteristics of the data packet transmitted from the software, calculates a delay time for delaying the transmission of the data packet by adding a random value to a fixed time determined from the characteristics of the data packet, and delays the transmission of the data packet by the calculated delay time. A computer system characterized by the above.
2. In the computer system according to Claim 1, wherein the timing adjustment mechanism adds the detection count of continuous retransmission to the destination of the data packet when it detects that the data packet of the same content has been continuously retransmitted more than a first threshold number of times for the fixed time among the delay times, when it detects that the detection count of continuous retransmission is equal to or more than a second threshold number of times, increases the fixed time of the delay time set for transmission to the destination by a set width, adds the detection count of non - continuous retransmission to the destination of the data packet when it detects that the data packet of the same content has been retransmitted less than a third threshold number of times, and when it detects that the detection count of non - continuous retransmission is equal to or more than a fourth threshold number of times, decreases the fixed time of the delay time set for transmission to the destination by a set width, thereby automatically adjusting the fixed time according to the state of the network. A computer system characterized by the above.
3. In the computer system according to Claim 1, wherein the network is a publicly shared network shared by a plurality of computers. A computer system characterized by the above.
4. In a computer system that is connected to a network including a network switch, is composed of a plurality of computers, and, when a data packet is lost on the network, recovers the transfer operation of the lost data packet by a retransmission operation, a computer; software operating on the computer; and a timing adjustment mechanism existing between the computer and the network, The timing adjustment mechanism is based on the characteristics of the data packet transmitted from the software, calculates a delay time for delaying the transmission of the data packet, and delays the transmission of the data packet by the calculated delay time, The timing adjustment mechanism is receives a statistical value regarding the congestion occurrence time from the network switch, and calculates the time for delaying the transmission from the statistical value A computer system characterized by the above.
5. In the computer system according to claim 1, the timing adjustment mechanism records the correspondence relationship between the delay time and the number of consecutive retransmissions, from the correspondence relationship, sets the delay time when the number of consecutive retransmissions is the smallest as the default value of the delay time, and calculates the delay time to be applied to the data packet based on the default value A computer system characterized by the above.
6. A method for avoiding congestion in a computer system that is connected to a network including a network switch, is composed of a plurality of computers, and when a data packet is lost on the network, recovers the transfer operation of the lost data packet by a retransmission operation, a timing adjustment mechanism existing between the computer on which the software operates and the network acquires the characteristics of the data packet transmitted from the software, based on the characteristics of the data packet, determines whether to delay the transmission of the data packet, adds a random value to a fixed time determined from the characteristics of the data packet to calculate a delay time for delaying the transmission of the data packet, and delays the transmission of the data packet by the calculated delay time, A method for avoiding congestion characterized by the above.
7. A method for avoiding congestion in a computer system that is connected to a network including a network switch, is composed of a plurality of computers, and when a data packet is lost on the network, recovers the transfer operation of the lost data packet by a retransmission operation, a timing adjustment mechanism existing between the computer on which the software operates and the network acquires the characteristics of the data packet transmitted from the software, based on the characteristics of the data packet, calculates a delay time for delaying the transmission of the data packet, and delays the transmission of the data packet by the calculated delay time, The timing adjustment mechanism is Receiving a statistical value regarding the congestion occurrence time from the network switch, Calculating the time for delaying transmission from the statistical value A congestion avoidance method characterized by the above.
Citation Information
Patent Citations
Early packet loss detection and feedback
JP2016515775A
Data rate adjuster using transport latency
US20020172153A1
Method and apparatus for timeout reduction and improved wireless network performance by delay injection
US20060104313A1
Redundancy processing device, information processing device, redundancy system, method, and storage medium
WO2017094235A1