High access rate data object transfer

By detecting storage node congestion in distributed storage systems, marking high access rate data objects and transmitting across nodes, the storage node-level congestion problem is solved, and a more uniform data object access rate and reduced storage network delay is achieved.

CN120029534APending Publication Date: 2025-05-23FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411462245.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-10-18
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In distributed storage systems, congestion may occur at both the storage node and the network level, resulting in reduced processing delays and performance. Especially when data object access rates change rapidly, it is difficult for the prior art to effectively solve storage node-level congestion.

Method used

By detecting the congestion status at the storage nodes in the storage network, access rate data of the data objects is obtained, high access rate data objects are marked, and transmission paths between storage nodes are calculated, along which high access rate data objects are transmitted from the congested storage node to the non-congested storage node.

Benefits of technology

The effect of reducing congestion at both the storage node level and the network level is achieved. By relocating high-access rate data objects between different storage nodes, the access rate of data objects is ensured to be more uniform and the overall delay of the storage network is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029534A_ABST
    Figure CN120029534A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to high access rate data object transmission. A computing system includes one or more processing devices configured to detect a congestion condition occurring at a first storage node located in a storage network of a distributed storage system. The one or more processing devices are further configured to obtain respective first access rate data for the first plurality of data objects stored at the first storage node. Based at least in part on the first access rate data, the one or more processing devices are further configured to mark the first data object as a high access rate data object. The one or more processing devices are further configured to compute a transmission path between the first storage node and a second storage node in the storage network. The one or more processing devices are further configured to transmit the high access rate data object from the first storage node to the second storage node along a transmission path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and more specifically, to a high access rate data object transmission. Background Art

[0002] Computing devices called storage nodes are used for cloud storage of data. These storage nodes are included in a distributed storage system, in which multiple storage nodes are networked together. For example, the storage nodes can be located in a data center. In order to upload and download data stored at the distributed storage system, client devices communicate with the storage nodes through a storage network connecting the storage nodes.

[0003] When client devices access data stored in a distributed storage system, the distributed storage system sometimes experiences congestion. Congestion occurs when a large number of requests are received at a specific component of the distributed storage system, resulting in processing delays. This congestion may occur during network transmission or at a storage node endpoint. Summary of the invention

[0004] To address these issues, according to one aspect of the present disclosure, a computing system is provided. The computing system includes one or more processing devices configured to detect a congestion condition occurring at a first storage node in a storage network located in a distributed storage system. In response to detecting the congestion condition, the one or more processing devices are also configured to obtain corresponding first access rate data for a first plurality of data objects stored at the first storage node. Based at least in part on the first access rate data, the one or more processing devices are also configured to mark a first data object in the first plurality of data objects as a high access rate data object. In response to marking the high access rate data object, the one or more processing devices are also configured to calculate a transmission path between the first storage node and a second storage node in the storage network. The one or more processing devices are also configured to transmit the high access rate data object from the first storage node to the second storage node along the transmission path.

[0005] This summary is provided to introduce in a simplified form a selection of concepts that will be further described in the following detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to implementations that address any or all of the disadvantages noted in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1A A storage network is schematically illustrated according to an example embodiment within which one or more data objects may be relocated.

[0007] Figure 1B Schematically shows the Figure 1A An example storage network of an example when congestion occurs at a storage node.

[0008] Figure 2 Schematically shows the Figure 1A An example computing system is configured to detect a congestion condition and identify high access rate data objects.

[0009] Figure 3A Schematically shows the Figure 2 An example of a computing system, a first storage node, and a second storage node when transferring a data object between storage nodes.

[0010] Figure 3B Schematically shows the Figure 2 The example of a computing system, a first storage node, and a second storage node when the first storage node and the second storage node store copies of a high access rate data object.

[0011] Figure 3C Schematically shows the Figure 2 The example of a computing system, a first storage node, and a second storage node when a high access rate data object is returned from a second storage node to a first storage node.

[0012] Figure 4 Schematically shows the Figure 2 An example computing system when using storage node performance data to detect congestion conditions.

[0013] Figure 5 Schematically shows the Figure 2 An example of a computing system when calculating a transmission path between a first storage node and a second storage node.

[0014] Figure 6 Schematically shows the Figure 2 An example computing system when computing a performance simulation of a second storage node.

[0015] Figure 7 Schematically shows the Figure 2 An example computing system is provided when using additional input to identify high access rate data objects.

[0016] Fig. 8A Shown according to Figure 1A A flowchart of an example method for use with a computing system to execute a scheduler and a controller for a storage network of a distributed storage system.

[0017] Figure 8B-8G shows that in some examples, Fig. 8A Additional steps of the method.

[0018] Fig. 9 Shows that you can instantiate Figure 1A A schematic diagram of an example computing environment for a computing system of FIG. DETAILED DESCRIPTION

[0019] Various techniques have been previously developed to alleviate congestion that occurs at the network level. Routing algorithms have been used to direct network traffic along different paths through a network in order to prevent or alleviate network congestion. For example, multipath transmission can be used to reduce the variation in the amount of network traffic through different paths.

[0020] In addition to these methods for reducing network-level congestion, techniques have been developed to improve the efficiency of storage nodes. For example, key-value engines for distributed storage systems have been developed in a manner designed to improve the efficiency of specific database operations. At the hardware level, different types of memory devices (e.g., solid-state drives (SSDs), hard disk drives (HDDs), and magnetic storage) are used to store data, depending on the expected access rate of the data.

[0021] According to previous methods of reducing congestion, technologies based on rerouting traffic at the network level do not address congestion that occurs at the storage node level, and storage node-level technologies do not address congestion that occurs at the network level. In existing storage networks, storage and networking each have their own control plane and data plane, which provide programmability for developers to define and implement their resource management and scheduling policies. Using these control planes and data planes, storage and networking are managed separately in existing data centers. This separate management sometimes reduces end-to-end performance due to a lack of coordination between storage and networking.

[0022] In some applications, such as video streaming, the rate at which data objects are accessed by client devices may vary widely over time. For example, a video that is infrequently accessed may "go viral" and have its access rate suddenly increase. As a result, the storage nodes storing the video may experience congestion. These rapid changes in access rates may occur unpredictably and, therefore, may be difficult to address using existing storage node efficiency techniques.

[0023] To address the above challenges, the following methods are provided to reduce storage node-level congestion. Using the following techniques, data objects with high access rates are relocated to different storage nodes. This relocation allows the storage network to achieve a more even distribution of traffic at the storage nodes, thereby reducing congestion.

[0024] Figure 1AAn example storage network 10 in which a data object relocation method may be used is schematically illustrated. The storage network 10 includes a plurality of storage nodes 12, each storing a plurality of data objects 14. The storage nodes 12 are physical server computing devices at which the data objects 14 are stored in memory devices. The storage network 10 also includes networking hardware 16, which includes a plurality of routers 17 and a network interface controller (NIC) 18. Via the networking hardware 16, one or more storage nodes 12 are configured to communicate with a plurality of client devices 40 over a computer network 11 such as the Internet. The client devices 40 execute corresponding applications 42 that communicate with the storage nodes 12 to upload and / or download the data objects 14. It should be understood that a content distribution network 11A separate from the storage network 10 may be used to cache copies of the data objects 14 on servers within the content distribution network 11A, which are closer to the client devices 40 on the network 11, thereby accelerating the delivery of the data objects 14 to the client devices 40 and reducing network congestion. Although multiple copies of data object 14 may be cached within content distribution network 11A, typically only one copy of data object 14 is stored within storage network 10 to save storage space, but in some cases two copies of a data object may be stored in the storage network, such as during migration of a data object from one location to another. An archived copy of data object 14 may also be stored in an archive location accessible by the storage network. In other examples, multiple copies of data object 14 may be stored in storage network 10.

[0025] The storage network 10 also includes a computing system 30 at which the scheduler 20 and the controller 22 are configured to be executed. As discussed in further detail below, the scheduler 20 is configured to monitor the performance of the network path and determine whether congestion occurs. In addition, the scheduler 20 is configured to calculate a prediction of future network performance. The controller 22 is configured to perform the relocation of the data objects 14 between the storage nodes 12 as described below.

[0026] Figure 1B An example storage network 10 is schematically illustrated when congestion occurs at a storage node. Figure 1BThe example storage network 10 includes storage nodes 12A, 12B, 12C, 12D, and 12E. The example storage network 10 also includes routers 17A, 17B, 17C, 17D, 17E, and 17F, which connect the storage nodes to applications 42A and 42B executed at corresponding client devices 40. Router 17 is arranged in an upper layer, which includes routers 17A and 17B, and in a lower layer, which includes routers 17C, 17D, 17E, and 17F. The scheduler 20 and the controller 22 are configured to interface with the routers 17A and 17B of the upper layer, and the routers 17A and 17B of the upper layer connect the routers of the lower layer to the applications 42A and 42B. Figure 1B Also shown are connections 24 between components included in the storage network 10.

[0027] exist Figure 1B In the example storage network 10 of FIG. 1 , storage nodes 12A and 12B both store a plurality of high access rate data objects 50, while storage nodes 12C, 12D, and 12E store only low access rate data objects 52. Therefore, the high access rate data objects 50 are unevenly distributed among the storage nodes 12. This uneven distribution causes overloading of routers 17A, 17C, and 17D, as well as overloading of connections connected to those routers. Figure 1B In the example of FIG. 1 , the overload connection 26 is indicated by a dashed line. Thus, Figure 1B It is shown that congestion at a storage node 12 may lead to congestion in a larger portion of the storage network 10 .

[0028] To alleviate congestion at routers 17A, 17C, and 17D, and at overloaded connections 26 associated with those routers, controller 22 may be configured to use conventional network routing techniques to establish alternative connections 28. These alternative connections 28 redirect portions of network traffic from router 17A to router 17B and from router 17D to router 17E. However, because access to high access rate data objects 50 is bottlenecked at storage nodes 12A and 12B rather than at any router 17, connections 26 to and from storage nodes 12A and 12B remain overloaded.

[0029] Figure 2 According to an example, it is schematically shown Figure 1A 1. The computing system 30 includes one or more processing devices 32 and one or more memory devices 34. The computing system 30 may be implemented at a single physical computing device or across multiple networked computing devices.

[0030] To address the congestion problem described above, one or more processing devices 32 are configured to detect a congestion condition 44 occurring at a first storage node 12A located in the storage network 10. For example, the congestion condition 44 may be detected based on latency data associated with the first storage node 12A, as discussed in further detail below. Therefore, the one or more processing devices 32 are configured to determine that congestion has occurred at the first storage node 12A.

[0031] Figure 2 The computing system 30 is shown when one or more processing devices 32 are configured to identify high access rate data objects 50 and low access rate data objects 52. The identification can be performed at the scheduler 20 in response to detecting a congestion condition 44. At the scheduler 20, the one or more processing devices 32 are configured to obtain access rate data 60. The access rate data 60 includes a plurality of access rates 62 and a plurality of derivatives of the access rates 64. The access rate data 60 includes corresponding first access rate data 60A for a first plurality of data objects 14 stored at a first storage node 12A. Figure 2 As shown in the example of , the first access rate data 60A includes a plurality of access rates 62A for those data objects 14 and derivatives 64A of the access rates.

[0032] Based at least in part on the first access rate data 60A, the one or more processing devices 32 are also configured to mark the first data object 14A in the first plurality of data objects 14 as a high access rate data object 50. In some examples, the one or more processing devices 32 are configured to mark the first data object 14A as a high access rate data object 50 in response to determining that the access rate 62A of the first data object 14A is higher than the first predefined access rate threshold 70. The detection and marking can occur in the storage network control plane. The mark on the first data object 14A can be metadata stored in a table accessible to the storage network control plane, or as metadata in the data object itself, as some examples. Additionally or alternatively, the one or more processing devices 32 can be configured to mark the first data object 14A as a high access rate data object 50 in response to determining that the derivative 64A of the access rate of the first data object 14A is higher than the first predefined access rate derivative threshold 71. Thus, the first data object 14A is identified as having a high access rate 62A or a rapidly increasing access rate 62A compared to the corresponding baseline values ​​of the access rate 62A and the derivative of the access rate 64A.

[0033] In some examples, the one or more processing devices 32 may be further configured to determine that the access rate 62A of the first data object 14A is above a second predefined access rate threshold 72, which is higher than the first predefined access rate threshold 70. Additionally or alternatively, the one or more processing devices 32 may determine that the derivative 64A of the access rate of the first data object 14A is above a second predefined access rate derivative threshold 73, which is higher than the first predefined access rate derivative threshold 71. Thus, the one or more processing devices 32 are configured to classify the first data object 14A as being at an access rate level that is higher than the access rate level defined by the first threshold.

[0034] Now turn to Figure 3A , the one or more processing devices 32 are also configured to calculate a transmission path 54 between the first storage node 12A and the second storage node 12B in the storage network 10. The first storage node 12A and the second storage node 12B may be located in the same data center. The one or more processing devices 32 may be configured to calculate the transmission path 54 at the scheduler 20. The transmission path 54 is selected from a plurality of possible network paths 56 that connect the first storage node 12A to the second storage node 12B. Each of the network paths 56 includes a plurality of connections 24 between components of the storage network 10.

[0035] After calculating the transmission path 54, the one or more processing devices 32 are further configured to transmit the high access rate data object 50 from the first storage node 12A to the second storage node 12B along the transmission path 54. Figure 3A In the example of , the controller 22 is configured to generate a transmission instruction 58 for the first storage node 12A to transmit the high access rate data object 50 to the second storage node 12B, and the computing system 30 is configured to transmit the transmission instruction 58 to the first storage node 12A. Therefore, the high access rate data object 50 is unloaded from the first storage node 12A, which can reduce the latency of communication with the first storage node 12A. In some examples, the controller 22 is configured to perform multi-path transmission by transmitting portions of the high access rate data object 50 along different corresponding transmission paths 54, thereby avoiding transmission path congestion when transmitting large data objects.

[0036] Back to Figure 2In an example of, the one or more processing devices 32 may be further configured to obtain second access rate data 60B for a second plurality of data objects 14 stored at a second storage node 12B. The second access rate data 60B includes respective access rates 62B and derivatives of the access rates 64B for the second plurality of data objects 14. Based at least in part on the second access rate data 60B, the one or more processing devices 32 may be further configured to mark a second data object 14B in the second plurality of data objects 14 as a low access rate data object 52. In response to determining that the access rate 62B of the second data object 14B is below a third predefined access rate threshold 74 or the derivative 64B of the access rate is below a third predefined access rate derivative threshold 75, the second data object 14B may be marked as a low access rate data object 52.

[0037] like Figure 3A As shown in , in response to transmitting the high access rate data object 50 from the first storage node 12A to the second storage node 12B, and marking the second data object 14B as a low access rate data object 52, in some examples, the one or more processing devices 32 are further configured to transmit the low access rate data object 52 to the first storage node 12A along the transmission path 54. Therefore, the one or more processing devices 32 are configured to redistribute the data object 14 between the first storage node 12A and the second storage node 12B. Similar to the high access rate data object 50, in some examples, the low access rate data object 52 can also be transmitted via multi-path transmission.

[0038] In some examples, when determining whether to transfer the second data object 14B to the first storage node 12A, data related to storage and memory usage at the first storage node 12A may be used at the scheduler 20. As an additional criterion for transferring the second data object 14B to the first storage node 12A, the one or more processing devices 32 may be further configured to determine that the first storage node 12A has sufficient storage capacity 76 to store the second data object 14B, and / or that the first storage node 12A also has sufficient memory write bandwidth 77 to write the second data object 14B to the first storage node 12A.

[0039] Figure 3B Another example of a high access rate data object 50 being transferred from a first storage node 12A to a second storage node 12B is schematically shown. Figure 3BIn the example of , the one or more processing devices 32 are further configured to replicate the high access rate data object 50. The one or more processing devices 32 are further configured to transmit the copy of the high access rate data object 50 to the second storage node 12B, so that two or more copies of the high access rate data object 50 are stored and accessible simultaneously within the storage network 10. Figure 3B In the example of FIG. 5 , by enabling high access rate data objects 50 to be accessed from two different storage nodes 12 , congestion at the first storage node 12A is reduced.

[0040] In some examples, the one or more processing devices 32 may be configured to replicate the high access rate data object 50 and transmit the copy in response to determining that the access rate 62A of the first data object 14A is above a second predefined access rate threshold 72, wherein the second predefined access rate threshold 72 is higher than the first predefined access rate threshold 70. Additionally or alternatively, the one or more processing devices 32 may be configured to replicate the high access rate data object 50 and transmit the copy in response to determining that the derivative 64A of the access rate of the first data object 14A is above a second predefined access rate derivative threshold 73, wherein the second predefined access rate derivative threshold 73 is higher than the first predefined access rate derivative threshold 71. Thus, the one or more processing devices 32 may replicate the high access rate data object 50 under the condition that the access rate 62A or the derivative 64A of the access rate is high enough that it exceeds the second threshold as well as the first threshold.

[0041] In some examples, such as Figure 3C As shown in , the one or more processing devices 32 are also configured to detect a congestion condition 44 occurring at the second storage node 12B after the high access rate data object 50 is transmitted from the first storage node 12A to the second storage node 12B. In response to detecting the congestion condition occurring at the second storage node, the one or more processing devices 32 may be further configured to return the high access rate data object 50 to the first storage node 12A. Therefore, in examples where the transmission does not alleviate the congestion condition 44, the one or more processing devices 32 may be configured to roll back the transmission of the high access rate data object 50 to the second storage node 12B. In some such examples, a copy of the high access rate data object 50 is transmitted to the first storage node 12A so that, as shown in FIG. Figure 3B As shown in FIG. 1 , the first storage node 12A and the second storage node 12B simultaneously store copies of the high access rate data object 50 .

[0042] Figure 4 The computing system 30 is schematically illustrated when a congestion condition 44 is detected, as described above. Figure 4As shown in the example, the identification of the congestion condition 44 can be at least partially based on one or more corresponding storage node weights 80 calculated for the storage node 12 at the scheduler 20. In Figure 4 the example, one or more processing devices 32 are configured to calculate the storage node weights 80, and the storage node weights 80 include a first storage node weight 80A of the first storage node 12A and a second storage node weight 80B of the second storage node 12B. In addition to detecting the congestion condition 44 in Figure 4 the example, one or more processing devices 32 are further configured to identify a non-congestion condition 46 at the second storage node 12B. The non-congestion condition 46 can indicate that the second storage node 12B is a qualified recipient of the high access rate data object 50.

[0043] In Figure 4 the example, one or more processing devices 32 are further configured to calculate the storage node weights 80 using the storage node performance data 90 associated with the plurality of storage nodes 12. Example storage node performance data 90 includes corresponding write latency data 92, read latency data 94, and write ratio 96 for the storage nodes. The write latency data 92 indicates the latency of the write operation at the storage node 12, and the read latency data 94 indicates the latency of the read operation at the storage node 12. The write ratio 96 indicates the ratio of the write operation to the read operation, which varies according to the workload of the storage node 12. Figure 4 the example shows first storage node performance data 90A associated with the first storage node 12A and including write latency data 92A, read latency data 94A, and write ratio 96A. Figure 4 the example also shows second storage node performance data 90B associated with the second storage node 12B and including write latency data 92B, read latency data 94B, and write ratio 96B.

[0044] In one example of the calculation of the storage node weights 80, one or more processing devices 32 are configured to calculate each storage node weight 80 as follows:

[0045] weight = write_latency_weight * write_ratio + read_latency_weight * (1 - write_ratio)

[0046] Therefore, in the above example, the storage node weight 80 is the average latency. In the above equation, write_latency_weight and read_latency_weight can be selected from a set of write latency intervals and a set of read latency intervals corresponding to different write latency levels and read latency levels, respectively (for example, intervals indicating low, medium and high write latency, respectively, and intervals indicating low, medium and high read latency, respectively). In some examples, the total storage node weight 80 across all storage nodes 12 can be normalized to 1, so that the corresponding storage node weight 80 of each storage node 12 is expressed relative to the storage node weights 80 of other storage nodes 12 in the storage network 10.

[0047] The one or more processing devices 32 may be further configured to compare the storage node weight 80 to a storage node weight threshold 82 to determine whether congestion has occurred at the storage node 12. Figure 4 In the example of , the one or more processing devices 32 determine that the first storage node weight 80A is below the storage node weight threshold 82 (in this example, a lower weight corresponds to more congestion), and therefore the first storage node 12A has a congested condition 44. In contrast, the one or more processing devices 32 determine that the second storage node weight 80B is above the storage node weight threshold 82, and therefore the second storage node 12B has a non-congested condition 46.

[0048] Figure 5 The computing system 30 is schematically shown when computing the transmission path 54 at the scheduler 20. Figure 5 In the example of , the one or more processing devices 32 are also configured to obtain path congestion data 100 associated with a plurality of network paths 56 within the storage network 10. Each of these network paths 56 has a storage node 12 as an endpoint. The path congestion data 100 may, for example, include a plurality of round trip times (RTTs) 102 of corresponding probe packets 108 transmitted along the plurality of network paths 56. The path congestion data 100 may also include a corresponding bandwidth 104 of the network path 56.

[0049] Based at least in part on the path congestion data 100, the one or more processing devices 32 are also configured to calculate a plurality of network path weights 106 associated with the corresponding plurality of network paths 56 between the storage nodes 12. Similar to the storage node weights 80, the network path weights 106 may each be selected from a set of predefined values ​​associated with intervals corresponding to amounts of latency (e.g., low RTT, medium RTT, and high RTT 102).

[0050] One or more processing devices 32 may also be configured to calculate a plurality of combined weights 110 based at least in part on network path weights 106 and storage node weights 80 of storage nodes 12 located at respective endpoints of network paths 56. In some examples, each of the combined weights 110 is the sum of the network path weight 106 of network path 56 and the storage node weights 80 of the two storage nodes 12 located at the endpoints of that network path 56. In other examples, the combined weights 110 may be calculated as a weighted sum of the network path weights 106 and the storage node weights 80.

[0051] One or more processing devices 32 are also configured to select a candidate transmission path pool 112. The candidate transmission path pool 112 includes the top N of the network paths 56 with the highest combined weights among the plurality of network paths 56, where N is a predetermined pool size. In Figure 5 an example, one or more processing devices 32 are also configured to randomly select a transmission path 54 from among the plurality of network paths 56 included in the candidate transmission path pool 112. The scheduler 20 is also configured to convey the selected transmission path 54 to the controller 22, where one or more processing devices 32 are also configured to generate a transmission instruction 58 to transmit a high access rate data object 50 along the transmission path 54.

[0052] In some examples, prior to transmitting the high access rate data object 50, one or more processing devices 32 are also configured to calculate a performance simulation 120 of a second storage node 12B, as Figure 6 shown in an example. In the performance simulation 120, the scheduler 20 is configured to simulate the second storage node 12B when the second storage node 12B stores the high access rate data object 50. One or more processing devices 32 are configured to calculate a predicted access rate 124 of the high access rate data object 50 at the simulated second storage node 122.

[0053] The predicted access rate 124 may be calculated, for example, at a storage network simulation machine learning model 126. In Figure 6 an example, the input to the storage network simulation machine learning model 126 includes simulated storage location data 128, which indicates data objects 14 stored at respective storage nodes 12 in the performance simulation 120. The storage of the high access rate data object 50 at the simulated second storage node 122 is indicated in the simulated storage location data 128. The input to the storage network simulation machine learning model 126 may also include simulated workload data 130, which indicates simulated read and write operations performed at the storage nodes 12. Additionally, the input to the storage network simulation machine learning model 126 may also include a simulated network topology 132 of the storage network 10.

[0054] At the storage network simulation machine learning model 126, the one or more processing devices 32 are configured to calculate predicted performance data 134. The predicted performance data 134 may include predicted write latency data 136, predicted read latency data 138, and predicted write ratio data 140. The predicted performance data 134 is calculated for the simulated second storage node 122, and in some examples may also be calculated for one or more other simulated storage nodes.

[0055] The one or more processing devices 32 may be further configured to determine that the congestion condition 44 did not occur at the simulated second storage node 122 in the performance simulation 120 based at least in part on the predicted performance data 134. Therefore, the one or more processing devices 32 determine that the simulated second storage node 122 has a non-congested condition 46. In response to determining that the congestion condition 44 did not occur at the simulated second storage node 122, the one or more processing devices 32 may be further configured to transfer the high access rate data object 50 to the second storage node 12B. Figure 6 The example shows that in response to making this determination, the scheduler 20 outputs a transfer instruction 58 to the controller 22. Therefore, before transferring the high access rate data object 50 to the second storage node 12B, the one or more processing devices 32 are configured to test whether congestion still occurs after the transfer.

[0056] Figure 7 Schematically shows that Figure 2 Additional inputs not shown in FIG. 5 are used by the computing system 30 when identifying high access rate data objects 50 and low access rate data objects 52. Figure 7 In the example of , the one or more processing devices 32 are also configured to receive priority metadata 150 associated with the data object 14 stored at the storage node 12 . Figure 7 Priority metadata 150A associated with a first data object 14A and priority metadata 150B associated with a second data object 14B are shown. The priority metadata 150 for a data object 14 may, for example, indicate an expected access rate for the data object 14. A target latency level may also be indicated in the priority metadata 150. For example, log data may have priority metadata 150 indicating a short target write latency. The priority metadata 150 for log data may also indicate that the log data has a low expected access rate because the log data is not typically read after its initial write except when recovering from a crash associated with the log data at the computing device.

[0057] The one or more processing devices 32 may be configured to identify the first data object 14A as a high access rate data object 50 based at least in part on the priority metadata 150A of the first data object 14A. Figure 7In the example of , the priority metadata 150A is used to select or modify the predefined access rate thresholds 70 and 72 and the predefined access rate derivative thresholds 71 ​​and 73 for the first data object 14A. The one or more processing devices 32 may be further configured to use the priority metadata 150B to modify the third predefined access rate threshold 74 and the third predefined access rate derivative threshold 75 for identifying the second data object 14B as the low access rate data object 52.

[0058] The one or more processing devices 32 may be further configured to receive the corresponding storage age 152 of the storage node 12. Figure 7 In an example of the present invention, one or more processing devices 32 are configured to receive a storage age 152A of a first storage node 12A and a storage age 152B of a second storage node 12B. The storage age 152 of the storage node 12 may indicate an estimated remaining hardware life of a storage device included in the storage node 12. For example, the storage age 152 of a solid state drive (SSD) may be a program erase count (PEC) of the SSD. In some examples, the storage age 152 of the storage node 12 may be calculated as an average PEC of the SSDs included in the storage node 12.

[0059] The one or more processing devices 32 may be configured to identify the first data object 14A as a high access rate data object 50 based at least in part on the storage age 152A of the first storage node 12A. Figure 7 As shown in , the predefined access rate thresholds 70 and 72 and the predefined access rate derivative thresholds 71 ​​and 73 may decrease as the storage age 152A increases. Therefore, to avoid hardware failures during high traffic at the first storage node 12A, the threshold at which the first data object 14A is transferred to the second storage node 12B may be lowered as the memory device of the first storage node 12A ages.

[0060] The one or more processing devices 32 may also be configured to use the storage age 152B of the second data object 14B to identify the second data object 14B as a low access rate data object 52. For example, to avoid transferring a high access rate data object 50 to an SSD memory device having a high storage age 152B, the one or more processing devices 32 may have a storage age threshold 158 for the storage age 152B. As an additional criterion for transferring the high access rate data object 50 to the second storage node 12B, the one or more processing devices 32 may be configured to determine that the storage age 152B is below the storage age threshold 158.

[0061] Fig. 8AA flow chart of a method 200 for use with a computing system is shown. The computing system executing the method 200 is configured to execute a scheduler and a controller for a storage network of a distributed storage system. The storage network includes a plurality of storage nodes storing data objects.

[0062] At step 202, method 200 includes detecting a congestion condition occurring at a first storage node in a storage network located in a distributed storage system. The congestion condition is a condition in which latency at the first storage node is increased due to high traffic. At step 204, in response to detecting the congestion condition, method 200 also includes obtaining corresponding first access rate data for a first plurality of data objects stored at the first storage node. The first access rate data is time series data that indicates, for each of the first plurality of data objects, a frequency at which those data objects are read from a storage device.

[0063] At step 206, method 200 also includes marking a first data object in the first plurality of data objects as a high access rate data object based at least in part on the first access rate data. In some examples, performing step 206 includes performing step 208. At step 208, method 200 may also include determining that an access rate of the first data object is above a first predefined access rate threshold, or that a derivative of the access rate of the first data object is above a first predefined access rate derivative threshold. In response to making any of the above determinations, the first data object may be marked as a high access rate data object.

[0064] In some examples, other attributes of the first data object may also be considered when determining whether to mark the first data object as a high access rate data object. For example, based on priority metadata of the first data object, the first data object may be at least partially identified as a high access rate data object. Additionally or alternatively, when determining whether to identify the first data object as a high access rate data object, the storage age of the first storage node may be used.

[0065] In step 210, in response to marking the high access rate data object, method 200 may also include calculating a transmission path between the first storage node and a second storage node in the storage network. In step 212, method 200 also includes transmitting the high access rate data object from the first storage node to the second storage node along the transmission path. Therefore, when congestion occurs at the first storage node, the high access rate data object is unloaded from the first storage node to the second storage node to reduce congestion. In some examples, multi-path transmission of the high access rate data object is performed.

[0066] Figure 8BAdditional steps of method 200 that are performed in some examples are shown. When the first data object is marked at step 206, step 214 of method 200 may be performed. At step 214, method 200 may also include determining that an access rate of the first data object is above a second predefined access rate threshold, the second predefined access rate threshold being higher than the first predefined access rate threshold. Additionally or alternatively, step 214 may include determining that a derivative of the access rate of the first data object is above a second predefined access rate threshold, the second predefined access rate threshold being higher than the first predefined access rate threshold.

[0067] When the high access rate data object is transmitted at step 212, steps 216 and 218 may be performed. At step 216, method 200 may also include replicating the high access rate data object. At step 218, method 200 may also include transmitting a copy of the high access rate data object to a second storage node, such that two or more copies of the high access rate data object are simultaneously stored and accessible within the storage network. Thus, the storage network may be configured to store additional copies of data objects having very high access rates or whose access rates are increased.

[0068] Figure 8C Additional steps of method 200 that may be performed in order to detect a congestion condition at step 202 are shown. At step 220, method 200 may also include obtaining storage node performance data associated with a plurality of storage nodes included in the storage network, respectively. The plurality of storage nodes include a first storage node and a second storage node. The storage node performance data may, for example, include corresponding write latency, read latency, and write ratio of the storage nodes.

[0069] In step 222, method 200 may also include calculating a plurality of storage node weights associated with the storage node based at least in part on the storage node performance data. For example, the storage node weight may be an average latency calculated from the write latency, read latency, and write ratio of the storage node. In step 224, method 200 may also include detecting a congestion condition occurring at a first storage node by at least in part comparing the storage node weight to a storage node weight threshold. When the storage node weight of the storage node exceeds the storage node weight threshold, the scheduler may indicate that a congestion condition has occurred at the storage node.

[0070] Fig.8D Shows that in the implementation Figure 8C In some examples, the method 200 may include additional steps that can be performed in the example of the steps of . At step 226, the method 200 may also include obtaining path congestion data associated with multiple network paths within the storage network. In some examples, the path congestion data may include multiple round trip times (RTTs) of the probe packets transmitted along the multiple network paths.

[0071] In step 228, based at least in part on the path congestion data, method 200 may further include calculating a plurality of network path weights associated with a corresponding plurality of network paths between storage nodes. These network weights may be calculated based at least in part on the RTT of the probe packets. When calculating the network path weights, network path bandwidth data may also be used in some examples. In step 230, based at least in part on the storage node weights and the network path weights, method 200 may further include selecting a transmission path along which high access rate data objects are transmitted. For example, a transmission path with the lowest total weight may be selected. In some examples, multiple transmission paths are selected, and multipath transmission is performed.

[0072] Fig. 8E Additional steps of method 200 performed in some examples are shown. At step 232, method 200 may also include obtaining second access rate data for a second plurality of data objects stored at a second storage node. Based at least in part on the second access rate data, method 200 may also include marking a second data object in the second plurality of data objects as a low access rate data object at step 234. At step 236, method 200 may also include transmitting the low access rate data object to the first storage node along the transmission path in response to transmitting the high access rate data object from the first storage node to the second storage node. Thus, the low access rate data object may replace the high access rate data object at the first storage node, thereby filling storage capacity that may otherwise be unoccupied.

[0073] Fig.8F Additional steps of method 200 performed in some examples are shown. In step 238, before transmitting the high access rate data object to the second storage node, method 200 may also include calculating a performance simulation of the second storage node. In the performance simulation, the second storage node is simulated as storing the high access rate data object. In some examples, the performance simulation may include a storage network simulation machine learning model that receives simulated storage location data, simulated workload data, and simulated network topology as input. In such an example, the storage network simulation machine learning model may be configured to output predicted performance data of the storage node.

[0074] At step 240, method 200 may further include determining that a congestion condition does not occur at the second storage node in the performance simulation. The weight-based method discussed above may be used to determine whether a congestion condition occurs at the simulated second storage node. At step 242, method 200 may further include transmitting the high access rate data object to the second storage node in response to determining that a congestion condition does not occur at the second storage node in the performance simulation. Thus, the scheduler may predict the performance of the second storage node to determine whether transmitting the high access rate data object to the second storage node will reduce congestion.

[0075] Figure 8G Additional steps of method 200 performed in some examples are shown. In step 244, method 200 may also include detecting a congestion condition occurring at the second storage node after the high access rate data object is transmitted to the second storage node. In step 246, method 200 may also include returning the high access rate data object to the first storage node in response to detecting the congestion condition occurring at the second storage node. Therefore, when transmitting the high access rate data object to the second storage node still causes storage node-level congestion, the high access rate data object may be returned. In some examples, when the high access rate data object is returned to the first storage node, copies of the high access rate data object are stored simultaneously at the first storage node and the second storage node.

[0076] Using the techniques discussed above, data objects can be relocated between storage nodes in a storage network in order to alleviate storage node level congestion. This relocation is performed in a manner that jointly considers the properties of the storage nodes and the network paths at the scheduler when calculating transfer instructions. Therefore, data object relocation can be performed in a manner that avoids congestion at both the storage node level and the network level. Using the devices and methods discussed above, the quality of service at the storage network can be robust to rapid changes in demand for specific data objects. Therefore, client devices can access high-traffic data objects reliably and with low latency.

[0077] In some embodiments, the methods and processes described herein may be bound to a computing system of one or more computing devices. In particular, such methods and processes may be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.

[0078] Fig. 9 A non-limiting embodiment of a computing system 300 is schematically shown that can implement one or more of the above methods and processes. The computing system 300 is shown in simplified form. The computing system 300 can embody the above description and Figure 1A. The computing system 300 may take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, video gaming devices, mobile computing devices, mobile communication devices (e.g., smart phones), and / or other computing devices, as well as wearable computing devices (such as smart watches and head-mounted augmented reality devices).

[0079] The computing system 300 includes a logic processor 302, a volatile memory 304, and a non-volatile storage device 306. The computing system 300 may optionally include a display subsystem 308, an input subsystem 310, a communication subsystem 312, and / or Fig. 9 Other components not shown.

[0080] Logical processor 302 includes one or more physical devices configured to execute instructions. For example, a logical processor may be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, implement technical effects, or otherwise obtain desired results.

[0081] The logical processor may include one or more physical processors configured to execute software instructions. Additionally or alternatively, the logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processor of the logical processor 302 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel and / or distributed processing. The various components of the logical processor may optionally be distributed in two or more separate devices, which may be remotely located and / or configured for coordinated processing. Various aspects of the logical processor may be virtualized and executed by a remotely accessible networked computing device configured in a cloud computing configuration. In this case, these virtualized aspects are run on different physical logical processors of various different machines.

[0082] The non-volatile storage device 306 includes one or more physical devices configured to store instructions executable by a logical processor to implement the methods and processes described herein. When such methods and processes are implemented, the state of the non-volatile storage device 306 can be transformed—for example, to store different data.

[0083] The non-volatile storage device 306 may include a removable and / or built-in physical device. The non-volatile storage device 306 may include an optical storage device (e.g., CD, DVD, HD-DVD, Blu-ray disc, etc.), a semiconductor memory (e.g., ROM, EPROM, EEPROM, flash memory, etc.), and / or a magnetic storage device (e.g., a hard disk drive, a floppy disk drive, a tape drive, MRAM, etc.) or other mass storage device technology. The non-volatile storage device 306 may include a non-volatile, dynamic, static, read / write, read-only, sequential access, location addressable, file addressable, and / or content addressable device. It should be understood that the non-volatile storage device 306 is configured to retain instructions even when power is cut off to the non-volatile storage device 306.

[0084] The volatile memory 304 may include physical devices including random access memory. The volatile memory 304 is typically used by the logical processor 302 to temporarily store information during processing of software instructions. It should be understood that when power is cut off to the volatile memory 304, the volatile memory 304 typically does not continue to store instructions.

[0085] Aspects of the logic processor 302, volatile memory 304, and non-volatile storage device 306 may be integrated together into one or more hardware logic components. Such hardware logic components may include, for example, field programmable gate arrays (FPGAs), programmed and application specific integrated circuits (PASIC / ASIC), programmed and application specific standard products (PSSP / ASSP), systems on chips (SOCs), and complex programmable logic devices (CPLDs).

[0086] The terms "module," "program," and "engine" may be used to describe aspects of the computing system 300 that are typically implemented in software by a processor to use portions of volatile memory to perform a specific function that involves a transformation process that specifically configures the processor to perform the function. Thus, a module, program, or engine may be instantiated by a logical processor 302 executing instructions stored by a non-volatile storage device 306, using portions of volatile memory 304. It should be understood that different modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Likewise, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module," "program," and "engine" may include individual or grouped executable files, data files, libraries, drivers, scripts, database records, and the like.

[0087] When included, the display subsystem 308 can be used to present a visual representation of the data stored by the non-volatile storage device 306. The visual representation can take the form of a graphical user interface (GUI). When the methods and processes described herein change the data stored by the non-volatile storage device and thus transform the state of the non-volatile storage device, the state of the display subsystem 308 can also be transformed to visually represent the changes in the underlying data. The display subsystem 308 can include one or more display devices utilizing almost any type of technology. Such a display device can be combined with the logical processor 302, the volatile memory 304, and / or the non-volatile storage device 306 in a shared housing, or such a display device can be a peripheral display device.

[0088] When included, the input subsystem 310 may include or interface with one or more user input devices, such as a keyboard, mouse, touch screen, or game controller. In some embodiments, the input subsystem may include or interface with selected natural user input (NUI) components. Such components may be integrated or peripheral, and the conversion and / or processing of input actions may be handled on-board or off-board. Example NUI components may include microphones for voice and / or speech recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; head trackers, eye trackers, accelerometers, and / or gyroscopes for motion detection and / or intent recognition; and electric field sensing components for assessing brain activity; and / or any other suitable sensors.

[0089] When included, the communication subsystem 312 can be configured to communicatively couple the various computing devices described herein to each other and to communicatively couple to other devices. The communication subsystem 312 can include wired and / or wireless communication devices compatible with one or more different communication protocols. As non-limiting examples, the communication subsystem can be configured for communication via a wireless telephone network or a wired or wireless local or wide area network. In some embodiments, the communication subsystem can allow the computing system 300 to send and / or receive messages from other devices via a network such as the Internet.

[0090] The following paragraphs provide additional descriptions of the subject matter of the present disclosure. According to one aspect of the present disclosure, a computing system is provided, including one or more processing devices, the one or more processing devices being configured to detect a congestion condition occurring at a first storage node in a storage network located in a distributed storage system. In response to detecting the congestion condition, the one or more processing devices are also configured to obtain corresponding first access rate data for a first plurality of data objects stored at the first storage node. Based at least in part on the first access rate data, the one or more processing devices are also configured to mark a first data object in the first plurality of data objects as a high access rate data object. In response to marking the high access rate data object, the one or more processing devices are also configured to calculate a transmission path between the first storage node and a second storage node in the storage network. The one or more processing devices are also configured to transmit the high access rate data object from the first storage node to the second storage node along the transmission path. The above features may have a technical effect of transferring high traffic data objects away from a storage node when the storage node experiences congestion.

[0091] According to this aspect, one or more processing devices may be configured to, in response to determining that the access rate of the first data object is higher than a first predefined access rate threshold or the derivative of the access rate of the first data object is higher than the first predefined access rate derivative threshold, mark the first data object as a high access rate data object. The above features may have the technical effect of identifying a high access rate data object as a data object with high traffic or rapidly increasing traffic.

[0092] According to this aspect, in order to transfer a high access rate data object from a first storage node to a second storage node, one or more processing devices may also be configured to copy the high access rate data object and transfer the copy of the high access rate data object to the second storage node, so that two or more copies of the high access rate data object are simultaneously stored and accessible within the storage network. The above features may have the technical effect of reducing congestion at a storage node by making additional copies of the high access rate data object accessible at different storage nodes.

[0093] According to this aspect, one or more processing devices may be configured to, in response to determining that an access rate of the first data object is above a second predefined access rate threshold that is higher than the first predefined access rate threshold or a derivative of the access rate of the first data object is above a second predefined access rate derivative threshold that is higher than the first predefined access rate derivative threshold, replicate the high access rate data object and transmit the copy. The above features may have the technical effect of replicating the high access rate data object in the event of very high traffic or very high increase in traffic.

[0094] According to this aspect, one or more processing devices may also be configured to calculate a performance simulation of the second storage node before transmitting the high access rate data object. In the performance simulation, the second storage node may store the high access rate data object. One or more processing devices may also be configured to determine that no congestion condition occurs at the second storage node in the performance simulation. One or more processing devices may also be configured to transmit the high access rate data object to the second storage node in response to determining that no congestion condition occurs at the second storage node in the performance simulation. The above-mentioned feature may have a technical effect of testing whether transmitting a high access rate data object will alleviate storage node congestion before transmitting the high access rate data object.

[0095] According to this aspect, one or more processing devices may also be configured to obtain storage node performance data respectively associated with multiple storage nodes included in the storage network. The multiple storage nodes may include a first storage node and a second storage node. Based at least in part on the storage node performance data, the one or more processing devices may also be configured to calculate multiple storage node weights associated with the storage nodes. The one or more processing devices may also be configured to detect a congestion condition occurring at the first storage node at least in part by comparing the storage node weight with a storage node weight threshold. The above features may have a technical effect of identifying when congestion occurs at the first storage node.

[0096] According to this aspect, one or more processing devices may also be configured to obtain path congestion data associated with multiple network paths within the storage network. The path congestion data may include multiple round trip times (RTTs) of probe packets transmitted along the multiple network paths. The above features may have a technical effect of identifying congestion in the network path.

[0097] According to this aspect, one or more processing devices may be configured to calculate a transmission path at least in part by calculating a plurality of network path weights associated with a corresponding plurality of network paths between storage nodes based at least in part on the path congestion data. Calculating the transmission path may also include selecting a transmission path along which high access rate data objects are transmitted based at least in part on the storage node weights and the network path weights. The above features may have the technical effect of selecting a transmission path that avoids congestion at both the storage nodes and the network paths.

[0098] According to this aspect, one or more processing devices may also be configured to obtain second access rate data for a second plurality of data objects stored at a second storage node. Based at least in part on the second access rate data, one or more processing devices may also be configured to mark a second data object in the second plurality of data objects as a low access rate data object. In response to transmitting the high access rate data object from the first storage node to the second storage node, the one or more processing devices may also be configured to transmit the low access rate data object to the first storage node along the transmission path. The above features may have a technical effect of more efficiently allocating storage space between the first storage node and the second storage node.

[0099] According to this aspect, one or more processing devices can be configured to identify the first data object as a high access rate data object based at least in part on priority metadata of the first data object. The above features can have a technical effect of using the priority level indicated in the priority metadata to set a threshold for determining that a data object is a high access rate data object.

[0100] According to this aspect, one or more processing devices may be configured to identify the first data object as a high access rate data object based at least in part on the storage age of the first storage node. The above features may have a technical effect of avoiding storage device failures at the first storage node under high traffic conditions.

[0101] According to this aspect, after transmitting the high access rate data object, the one or more processing devices may also be configured to detect a congestion condition occurring at the second storage node. In response to detecting the congestion condition occurring at the second storage node, the one or more processing devices may also be configured to return the high access rate data object to the first storage node. The above feature may have a technical effect of rolling back the data object transmission when storage node-level congestion still occurs after transmitting the high access rate data object.

[0102] According to another aspect of the present disclosure, a method for use with a computing system is provided. The method may include detecting a congestion condition occurring at a first storage node in a storage network located in a distributed storage system. In response to detecting the congestion condition, the method may also include obtaining corresponding first access rate data for a first plurality of data objects stored at the first storage node. Based at least in part on the first access rate data, the method may also include marking a first data object in the first plurality of data objects as a high access rate data object. In response to marking the high access rate data object, the method may also include calculating a transmission path between the first storage node and a second storage node in the storage network. The method may also include transmitting the high access rate data object from the first storage node to the second storage node along the transmission path. The above-mentioned features may have the technical effect of transferring high-traffic data objects away from a storage node when the storage node experiences congestion.

[0103] According to this aspect, in response to determining that the access rate of the first data object is higher than a first predefined access rate threshold or the derivative of the access rate of the first data object is higher than the first predefined access rate derivative threshold, the first data object can be marked as a high access rate data object. The above feature can have a technical effect of identifying a high access rate data object as a data object with high traffic or rapidly increasing traffic.

[0104] According to this aspect, transmitting the high access rate data object from the first storage node to the second storage node may include copying the high access rate data object. Transmitting the high access rate data object may also include transmitting a copy of the high access rate data object to the second storage node, so that two or more copies of the high access rate data object are simultaneously stored and accessible within the storage network. The above features may have the technical effect of reducing congestion at the storage node by making additional copies of the high access rate data object accessible at different storage nodes.

[0105] According to this aspect, the method may also include obtaining storage node performance data respectively associated with a plurality of storage nodes included in the storage network. The plurality of storage nodes may include a first storage node and a second storage node. Based at least in part on the storage node performance data, the method may also include calculating a plurality of storage node weights associated with the storage nodes. The method may also include detecting a congestion condition occurring at the first storage node at least in part by comparing the storage node weight with a storage node weight threshold. The above features may have a technical effect of identifying when congestion occurs at the first storage node.

[0106] According to this aspect, calculating the transmission path may include obtaining path congestion data associated with a plurality of network paths within the storage network. Based at least in part on the path congestion data, calculating the transmission path may also include calculating a plurality of network path weights associated with the respective plurality of network paths between the storage nodes. Based at least in part on the storage node weights and the network path weights, calculating the transmission path may also include selecting a transmission path along which the high access rate data object is transmitted. The above features may have the technical effect of selecting a transmission path that avoids congestion at both the storage nodes and the network paths.

[0107] According to this aspect, the method may also include obtaining second access rate data for a second plurality of data objects stored at the second storage node. Based at least in part on the second access rate data, the method may also include marking a second data object in the second plurality of data objects as a low access rate data object. In response to transmitting the high access rate data object from the first storage node to the second storage node, the method may also include transmitting the low access rate data object to the first storage node along the transmission path. The above features may have a technical effect of more efficiently allocating storage space between the first storage node and the second storage node.

[0108] According to this aspect, based on the priority metadata of the first data object and / or the storage age of the first storage node, the first data object is identified as a high access rate data object. The above feature can have the technical effect of using the priority level indicated in the priority metadata to set a threshold for the data object to be determined as a high access rate data object. Additionally or alternatively, the above feature can have the technical effect of avoiding storage device failure at the first storage node in the case of high traffic.

[0109] According to another aspect of the present disclosure, a computing system is provided, including one or more processing devices, the one or more processing devices being configured to detect a congestion condition occurring at a first storage node in a storage network located in a distributed storage system. In response to detecting the congestion condition, the one or more processing devices are further configured to obtain corresponding access rate data for a first plurality of data objects stored at the first storage node and a second plurality of data objects stored at a second storage node in the storage network. Based at least in part on the access rate data, the one or more processing devices are further configured to mark a first data object in the first plurality of data objects as a high access rate data object, and mark a second data object in the second plurality of data objects as a low access rate data object. In response to marking the high access rate data object and the low access rate data object, the one or more processing devices are further configured to transfer the high access rate data object from the first storage node to the second storage node. The one or more processing devices are further configured to transfer the low access rate data object from the second storage node to the first storage node.

[0110] As used herein, "and / or" is defined as inclusive OR∨ as specified by the following truth table:

[0111] A B A∨B real real real real Fake real Fake real real Fake Fake Fake

[0112] It should be understood that the configuration and / or method described herein is exemplary in nature, and these specific embodiments or examples should not be considered restrictive, because many variations are possible. The specific routine or method described herein can represent one or more of any number of processing strategies. Therefore, the various actions shown and / or described can be performed in the order shown and / or described, in other orders, in parallel or omitted. Likewise, the order of the above process can be changed.

[0113] The subject matter of the present disclosure includes all novel and nonobvious combinations and subcombinations of the various processes, systems and configurations, and other features, functions, acts, and / or properties disclosed herein, as well as any and all equivalents.

Claims

1. A computing system comprising: One or more processing devices configured to: Detecting a congestion condition occurring at a first storage node in a storage network of a distributed storage system; In response to detecting the congestion condition, obtaining corresponding first access rate data for a first plurality of data objects stored at the first storage node; Based at least in part on the first access rate data, marking a first data object in the first plurality of data objects as a high access rate data object; In response to marking the high access rate data object, calculating a transmission path between the first storage node and a second storage node in the storage network; as well as The high access rate data object is transmitted from the first storage node to the second storage node along the transmission path.

2. The computing system of claim 1 , wherein the one or more processing devices are configured to mark the first data object as the high access rate data object in response to determining: The access rate of the first data object is higher than a first predefined access rate threshold; or A derivative of the access rate of the first data object is above a first predefined access rate derivative threshold.

3. The computing system of claim 2, wherein in order to transfer the high access rate data object from the first storage node to the second storage node, the one or more processing devices are further configured to: replicating the high access rate data object; and The copy of the high access rate data object is transmitted to the second storage node so that two or more copies of the high access rate data object are simultaneously stored and accessible within the storage network.

4. The computing system of claim 3, wherein the one or more processing devices are configured to replicate the high access rate data object and transmit the replica in response to determining: The access rate of the first data object is higher than a second predefined access rate threshold, and the second predefined access rate threshold is higher than the first predefined access rate threshold; or The derivative of the access rate of the first data object is above a second predefined access rate derivative threshold, and the second predefined access rate derivative threshold is higher than the first predefined access rate derivative threshold.

5. The computing system of claim 1 , wherein the one or more processing devices are further configured to: Before transmitting the high access rate data object, calculating a performance simulation of the second storage node, wherein in the performance simulation, the second storage node stores the high access rate data object; determining that the congestion condition does not occur at the second storage node in the performance simulation; as well as In response to determining that the congestion condition does not occur at the second storage node in the performance simulation, the high access rate data object is transmitted to the second storage node.

6. The computing system of claim 1 , wherein the one or more processing devices are further configured to: Obtaining storage node performance data respectively associated with a plurality of storage nodes included in the storage network, wherein the plurality of storage nodes include the first storage node and the second storage node; calculating a plurality of storage node weights associated with the storage node based at least in part on the storage node performance data; as well as The congestion condition occurring at the first storage node is detected at least in part by comparing the storage node weight to a storage node weight threshold.

7. The computing system of claim 6, wherein: The one or more processing devices are further configured to: obtain path congestion data associated with a plurality of network paths within the storage network; and The path congestion data includes a plurality of round trip times (RTTs) of probe packets transmitted along the plurality of network paths.

8. The computing system of claim 7, wherein the one or more processing devices are configured to compute the transmission path at least in part by: calculating a plurality of network path weights associated with a corresponding plurality of network paths between the storage nodes based at least in part on the path congestion data; and The transmission path along which the high access rate data object is transmitted is selected based at least in part on the storage node weights and the network path weights.

9. The computing system of claim 1 , wherein the one or more processing devices are further configured to: obtaining second access rate data for a second plurality of data objects stored at the second storage node; marking a second data object in the second plurality of data objects as a low access rate data object based at least in part on the second access rate data; as well as In response to transmitting the high access rate data object from the first storage node to the second storage node, the low access rate data object is transmitted to the first storage node along the transmission path.

10. The computing system of claim 1, wherein the one or more processing devices are configured to identify the first data object as the high access rate data object based at least in part on priority metadata of the first data object.

11. The computing system of claim 1, wherein the one or more processing devices are configured to: identify the first data object as the high access rate data object based at least in part on a storage age of the first storage node.

12. The computing system of claim 1 , wherein after transmitting the high access rate data object, the one or more processing devices are further configured to: detecting a congestion condition occurring at the second storage node; and In response to detecting the congestion condition occurring at the second storage node, the high access rate data object is returned to the first storage node.

13. A method for use with a computing system, the method comprising: Detecting a congestion condition occurring at a first storage node in a storage network of a distributed storage system; In response to detecting the congestion condition, obtaining corresponding first access rate data for a first plurality of data objects stored at the first storage node; Based at least in part on the first access rate data, marking a first data object in the first plurality of data objects as a high access rate data object; In response to marking the high access rate data object, calculating a transmission path between the first storage node and a second storage node in the storage network; as well as The high access rate data object is transmitted from the first storage node to the second storage node along the transmission path.

14. The method of claim 13, wherein marking the first data object as the high access rate data object is responsive to determining: The access rate of the first data object is higher than a first predefined access rate threshold; or A derivative of the access rate of the first data object is above a first predefined access rate derivative threshold.

15. The method according to claim 13, wherein transferring the high access rate data object from the first storage node to the second storage node comprises: Copying the high access rate data object; as well as The copy of the high access rate data object is transmitted to the second storage node so that two or more copies of the high access rate data object are simultaneously stored and accessible within the storage network.

16. The method according to claim 13, further comprising: Obtaining storage node performance data respectively associated with a plurality of storage nodes included in the storage network, wherein the plurality of storage nodes include the first storage node and the second storage node; calculating a plurality of storage node weights associated with the storage nodes based at least in part on storing the node performance data; as well as The congestion condition occurring at the first storage node is detected at least in part by comparing the storage node weight to a storage node weight threshold.

17. The method of claim 16, wherein calculating the transmission path comprises: obtaining path congestion data associated with a plurality of network paths within the storage network; calculating a plurality of network path weights associated with a corresponding plurality of network paths between the storage nodes based at least in part on the path congestion data; as well as The transmission path along which the high access rate data object is transmitted is selected based at least in part on the storage node weights and the network path weights.

18. The method according to claim 13, further comprising: obtaining second access rate data for a second plurality of data objects stored at the second storage node; marking a second data object in the second plurality of data objects as a low access rate data object based at least in part on the second access rate data; as well as In response to transmitting the high access rate data object from the first storage node to the second storage node, the low access rate data object is transmitted to the first storage node along the transmission path.

19. The method of claim 13, wherein the first data object is identified as the high access rate data object based on at least one of the following: priority metadata of the first data object; and / or The storage age of the first storage node.

20. A computing system comprising: One or more processing devices configured to: Detecting a congestion condition occurring at a first storage node in a storage network of a distributed storage system; In response to detecting the congestion condition, obtaining corresponding access rate data for: a first plurality of data objects stored at the first storage node; as well as a second plurality of data objects stored at a second storage node in the storage network; Based at least in part on the access rate data: marking a first data object among the first plurality of data objects as a high access rate data object; as well as marking a second data object in the second plurality of data objects as a low access rate data object; as well as In response to marking the high access rate data object and the low access rate data object: Transmitting the high access rate data object from the first storage node to the second storage node; as well as The low access rate data object is transferred from the second storage node to the first storage node.