Path selection method, data pulling method and wide area storage service cluster system
By determining the shortest path and the sub-short path in the wide-area storage service cluster and transmitting data requests at the same time, the problem of low data pulling efficiency in the prior art is solved, and more efficient data transmission is achieved.
Patent Information
- Application Number
- CN202411969986.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In the prior art, when data is transmitted through the shortest network access path, there is a problem of low data pulling efficiency.
By determining the shortest path and the sub-short path in the wide-area storage service cluster, the path weight is calculated using the node path map and node attribute information, and the sub-short path is determined based on the preset weight difference threshold, the method of transmitting data requests at the same time as the shortest path and the sub-short path are realized.
Improves data pulling efficiency and avoids the inefficiency problem when a path fails when a request fails.
Smart Images

Figure CN119945967A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer network technology, and in particular to a path selection method, a data pulling method, and a wide area storage service cluster system. Background Art
[0002] In the application of computer cloud storage, the infrastructure for data storage and access is one of the key factors in user experience. When users upload or download files to the cloud via the Internet, the speed, stability and security of data transmission directly determine the user's overall perception of the service. Therefore, how to ensure an efficient and secure data interaction experience has become a key issue of concern.
[0003] At present, the way to pull the data required by users through a wide area storage network is often to obtain the shortest network access path based on the weight calculation between the node networks, and then pull data from the cache of the target node through the shortest network access path based on the data pull request issued by the user, so as to quickly obtain the data required by the user.
[0004] However, when pulling data through the above path selection method, there is a problem of low data pulling efficiency. Summary of the invention
[0005] The embodiments of the present application provide a path selection method, a data pulling method and a wide area storage service cluster system to solve the problem of low data pulling efficiency.
[0006] In a first aspect, an embodiment of the present application provides a path selection method, including:
[0007] Determine the target node corresponding to the data pull request in the wide area storage service cluster;
[0008] Determine multiple target paths to the target node based on the node path graph in the wide area storage service cluster;
[0009] Determine path weights of multiple target paths based on attribute information of nodes in the wide area storage service cluster;
[0010] According to the path weight, determine the shortest path among multiple target paths;
[0011] According to the path weight, the shortest path and the preset weight difference threshold, the second shortest path is determined, and both the shortest path and the second shortest path are used to transmit the pull data request to the target node.
[0012] In a possible implementation, determining the second shortest path according to the path weight, the shortest path, and a preset weight difference threshold includes:
[0013] Determine the weight differences between the multiple target paths and the shortest path according to the path weight and the path weight of the shortest path;
[0014] The second shortest path is determined according to the weight difference and a preset weight difference threshold, where the second shortest path is a path whose weight difference with the shortest path is not greater than the preset weight difference threshold.
[0015] In a possible implementation, determining path weights of multiple target paths according to attribute information of nodes in the wide area storage service cluster includes:
[0016] Determine the node weight according to the attribute information of the nodes in the wide area storage service cluster;
[0017] According to the node weights, the path weights of multiple target paths are determined.
[0018] In a possible implementation, determining the node weight according to the attribute information of the node in the wide area storage service cluster includes:
[0019] Determine a first weight of the node according to the network delay information in the attribute information, determine a second weight of the node according to the bandwidth capacity information in the attribute information, and determine a third weight of the node according to the cost information in the attribute information;
[0020] The node weight is determined according to the node first weight, the node second weight and the node third weight, wherein the node first weight and the node third weight are in inverse proportion to the node weight, and the node second weight is in direct proportion to the node weight.
[0021] In a second aspect, an embodiment of the present application provides a data pulling method, including:
[0022] In response to receiving a pull data request, determining a shortest path and a second shortest path for transmitting the pull data request, wherein the shortest path and the second shortest path are obtained according to the path selection method of the first aspect;
[0023] According to the shortest path and the second shortest path, the data pull request is transmitted to the target node;
[0024] Receive the returned target data, which is the fastest returned required data among the shortest path and the second shortest path.
[0025] In a possible implementation, after receiving the returned target data, the method further includes:
[0026] The adjacent intermediate nodes in the shortest path and the second shortest path cache the demand data to obtain the cached demand data. The adjacent intermediate nodes are intermediate nodes connected to the storage service node. The intermediate nodes are nodes between the storage service node and the target node.
[0027] If the shortest path and the second shortest path receive a new data pull request, and the target of the new data pull request is the demand data, the storage service node receives the demand data of the returned target cache, and the demand data of the target cache is the demand data of the cache that returns the fastest among the adjacent intermediate nodes of the shortest path and the adjacent intermediate nodes of the second shortest path.
[0028] In a possible implementation manner, after the demand data is cached at adjacent intermediate nodes in the shortest path and the second shortest path, and the cached demand data is obtained, the method further includes:
[0029] If the shortest path and the second shortest path do not receive a new request to pull data, or the target of the new request to pull data is not the required data, then the cache retention time of the adjacent intermediate node is obtained;
[0030] If the cache retention time is longer than the preset cache retention time threshold, the cached demand data is deleted.
[0031] In the third aspect, an embodiment of the present application provides a wide-area storage service cluster system, including a central service node and storage service nodes corresponding to multiple regions. The central service node is used to periodically detect attribute information and reachability information of multiple storage service nodes, and update the node path map according to the reachability information. The storage service node is used to execute the path selection method of the first aspect above.
[0032] In a fourth aspect, an embodiment of the present application provides a path selection device, including:
[0033] A first determination module is used to determine a target node corresponding to a data pull request in the wide area storage service cluster;
[0034] A second determination module is used to determine multiple target paths to the target node according to the node path graph in the wide area storage service cluster;
[0035] A third determination module is used to determine the path weights of multiple target paths according to the attribute information of the nodes in the wide area storage service cluster;
[0036] A fourth determination module, used to determine the shortest path among multiple target paths according to the path weight;
[0037] The fifth determination module is used to determine the second shortest path according to the path weight, the shortest path and a preset weight difference threshold, and both the shortest path and the second shortest path are used to transmit a request to pull data to the target node.
[0038] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0039] Memory stores computer-executable instructions;
[0040] The processor executes the computer-executable instructions stored in the memory to implement the above-mentioned first aspect and / or various possible implementations of the first aspect.
[0041] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.
[0042] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0043] The path selection method, data pulling method and wide area storage service cluster system provided in the embodiments of the present application determine the target node corresponding to the data pulling request in the wide area storage service cluster; determine multiple target paths to the target node according to the node path diagram in the wide area storage service cluster; determine the path weights of the multiple target paths according to the attribute information of the nodes in the wide area storage service cluster; determine the shortest path among the multiple target paths according to the path weight; determine the second shortest path according to the path weight, the shortest path and a preset weight difference threshold, and both the shortest path and the second shortest path are used as a means to transmit the data pulling request to the target node, so that after the shortest path is determined, the difference in weight between other paths and the shortest path is calculated, and the second shortest path with a weight close to the shortest path is determined according to the weight difference and the preset weight difference threshold, and the data pulling request is transmitted simultaneously through the shortest path and the second shortest path, so as to achieve the effect of improving data pulling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] Figure 1 Schematic diagram of the scenario for path selection provided for this application;
[0046] Figure 2 A schematic diagram of a flow chart of a path selection method provided in an embodiment of the present application;
[0047] Figure 3 A schematic diagram of determining the shortest path and the second shortest path provided in an embodiment of the present application;
[0048] Figure 4 A flowchart of another path selection method provided in an embodiment of the present application;
[0049] Figure 5A schematic diagram of a data extraction method according to an embodiment of the present invention;
[0050] Figure 6 A schematic diagram of an update node path diagram provided in an embodiment of the present application;
[0051] Figure 7 A schematic diagram of the structure of a path selection device provided in an embodiment of the present application;
[0052] Figure 8 A schematic diagram of the structure of a data extraction device provided in an embodiment of the present application;
[0053] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0054] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0055] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0056] First, the terms involved in this application are explained:
[0057] Wide-area storage service cluster: refers to a storage system consisting of storage service nodes distributed in multiple geographical locations. These nodes are connected through a high-speed network to jointly provide unified, highly available data storage services.
[0058] Storage service node: refers to the physical or virtual server responsible for actual data storage, management and processing in the wide area storage service cluster. Each node is an independent working unit in the corresponding region, working together to provide unified data storage services.
[0059] Figure 1 A schematic diagram of the path selection scenario provided for this application, such as Figure 1As shown, the specific application scenario of the present application is that user A-1 in region A creates a bucket bucket-movie (not shown in the figure) and specifies the index location as region A; user A-2 in region A writes object movie 1 to bucket bucket-movie, and the object will be written to the BOSS data storage space of region A; user B in region B writes object movie 2 to bucket bucket-movie, and the object will be written to the BOSS data storage space of region B; when the user in region A requests object movie 2, the BOSS service in region A will pull data from region B through the cloud-to-cloud highway, through the path region A-region B, or the path region A-region C-region B, and return it to the user.
[0060] In combination with the above scenario, it can be seen that in the prior art, the method of pulling data by selecting the shortest network access path to transmit the data pull request has a technical problem of low data pulling efficiency because the performance of storage clusters in different regions varies and the response time to the request is different. Only the shortest network access path is selected.
[0061] The path selection method, data pulling method and wide area storage service cluster system provided in the present application determine the shortest path through the node path graph and the attribute information of the node, and then calculate the difference in weights between other paths and the shortest path, and determine the second shortest path with a weight close to the shortest path based on the difference in weights and a preset weight difference threshold, and solve the technical problem of low data pulling efficiency by simultaneously transmitting the shortest path and the second shortest path to pull data through technical means.
[0062] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0063] Figure 2 The flowchart of the path selection method provided in the embodiment of the present application is shown in FIG. The execution subject of the present method can be a server or other servers, and this embodiment is not particularly limited here. Figure 2 As shown, the method may include:
[0064] S201. Determine a target node corresponding to a data pull request in a wide area storage service cluster.
[0065] Among them, pulling data request can refer to the client or application actively initiating a request to the wide area storage service cluster to obtain specific data resources. The pulling data request can include required data information and target node information. The required data information can be the description information of the specific data to be obtained, and the target node information can be the information specifying which node to pull data from.
[0066] The method for determining a target node corresponding to a data pull request in a wide area storage service cluster may include determining a target node in storage service nodes of the wide area storage service cluster according to target node information in the data pull request.
[0067] S202: Determine multiple target paths to target nodes according to a node path graph in the wide area storage service cluster.
[0068] The node path diagram may refer to a diagram used to visually represent the nodes in the wide area storage service cluster and the connection relationships between them.
[0069] The target path may refer to all connection paths between the storage service node that issues the data pull request and the target node.
[0070] The method for determining multiple target paths to target nodes may include taking the storage service node that issues a request to pull data as the starting node, and determining all possible connection methods between the starting node and the target node in a node path graph, including direct connection and indirect connection through other storage service nodes as intermediate nodes.
[0071] S203: Determine path weights of multiple target paths according to attribute information of nodes in the wide area storage service cluster.
[0072] The node attribute information may refer to various performance-related parameters of the node, such as bandwidth, delay, packet loss rate, etc.
[0073] Path weight can refer to a comprehensive indicator that measures the superiority of a path relative to other paths, and can be calculated based on the attribute information of multiple nodes on the path.
[0074] In the embodiment of the present application, the method for determining the path weights of multiple target paths according to the attribute information of the nodes in the wide area storage service cluster may include:
[0075] Determine the node weight according to the attribute information of the nodes in the wide area storage service cluster;
[0076] According to the node weights, the path weights of multiple target paths are determined.
[0077] Among them, node weight can refer to a numerical indicator used to quantify and evaluate the relative importance, performance or fitness of each node in the wide area storage service cluster, which comprehensively reflects various attribute information of the node.
[0078] The method of determining the path weights of multiple target paths based on node weights may include summing the weights of all nodes on the path, or pre-defining a weight calculation model to assign an initial weight to each node on the path, and then combining the node weights to comprehensively obtain the path weights of multiple target paths.
[0079] In the embodiment of the present application, according to the attribute information of the nodes in the wide area storage service cluster, the method for determining the node weight may include:
[0080] Determine a first weight of the node according to the network delay information in the attribute information, determine a second weight of the node according to the bandwidth capacity information in the attribute information, and determine a third weight of the node according to the cost information in the attribute information;
[0081] The node weight is determined according to the node first weight, the node second weight and the node third weight, wherein the node first weight and the node third weight are in inverse proportion to the node weight, and the node second weight is in direct proportion to the node weight.
[0082] Among them, the first weight of the node, the second weight of the node and the third weight of the node can respectively refer to the weights calculated based on the network delay information, bandwidth capacity information and cost information. The calculation method can include first standardizing the information, and then determining the first, second and third weights of the node according to a preset weight mapping table.
[0083] The method for determining the node weight may include weighted summing the first, second and third weights of the node according to a certain ratio, wherein the lower the network delay and cost and the higher the bandwidth capacity, the higher the relative importance of the node should be. Therefore, the first weight of the node and the third weight of the node can be set to be inversely proportional to the node weight, and the second weight of the node can be set to be in direct proportion to the node weight.
[0084] In other embodiments, the first node weight and the third node weight may be set in direct proportion to the node weight, and the second node weight may be set in inverse proportion to the node weight. In this case, the lower the node weight, the higher the relative importance of the node.
[0085] S204: Determine the shortest path among multiple target paths according to the path weights.
[0086] The shortest path may refer to a path with the best comprehensive performance among multiple target paths.
[0087] The method for determining the shortest path may include first obtaining the path weights of multiple target paths based on the path weights. If the weight of the node in the path is higher, the relative importance of the node is higher, and the target path with the highest path weight is selected as the shortest path. If the weight of the node in the path is lower, the relative importance of the node is higher, and the target path with the lowest path weight is selected as the shortest path.
[0088] S205: Determine the next shortest path according to the path weight, the shortest path and a preset weight difference threshold, and both the shortest path and the next shortest path are used to transmit a request to pull data to the target node.
[0089] The second shortest path may refer to a path whose weight difference with the shortest path is within a preset range, and there may be multiple second shortest paths.
[0090] Among them, in the embodiment of the present application, the method for determining the second shortest path according to the path weight, the shortest path and the preset weight difference threshold may include:
[0091] Determine the weight differences between the multiple target paths and the shortest path according to the path weight and the path weight of the shortest path;
[0092] The second shortest path is determined according to the weight difference and a preset weight difference threshold, where the second shortest path is a path whose weight difference with the shortest path is not greater than the preset weight difference threshold.
[0093] The weight difference may refer to the absolute value of the weight difference between the target path excluding the shortest path and the shortest path. The smaller the weight difference, the smaller the gap between the target path and the shortest path. The larger the weight difference, the larger the gap between the target path and the shortest path.
[0094] The method for determining the weight difference may include subtracting the weights of the multiple target paths from the weight of the shortest path, and then calculating the absolute value to obtain the weight difference.
[0095] The method for determining the second shortest path may include comparing the weight differences between the multiple target paths and the shortest path with a preset weight difference threshold, and if the weight difference is not greater than the weight difference threshold, determining that the target path is the second shortest path, for example Figure 3 , node 1 needs to pull data from node 3, and calculate the shortest path ranking according to the Dijkstra algorithm:
[0096] First place: Node 1->Node 2->Node 3, weight 3+5=8;
[0097] Second place: Node 1->Node 4->Node 3, weight 6+3=9;
[0098] Third place: Node 1->Node 3, weight 20;
[0099] The first path is determined to be the shortest path, and the weight difference between the second and third paths is calculated: |8-9|=1, |8-20|=12;
[0100] The weight differences are 1 and 12 respectively. The weight difference of the second place is less than the preset weight difference threshold of 5, and the weight difference of the third place is greater than the preset weight difference threshold of 5, so the path of the second place is determined to be the second shortest path.
[0101] Therefore, node 1 sends a data pull request to node 2, the next node of the shortest path, and node 4, the next node of the second shortest path at the same time.
[0102] The path selection method provided in the embodiment of the present application can determine the shortest path through the node path graph and the attribute information of the node, and then calculate the difference in weights between other paths and the shortest path. According to the difference in weights and a preset weight difference threshold, a second shortest path with a weight close to the shortest path is determined, and the technical means of simultaneously transmitting the data pulling request through the shortest path and the second shortest path can be used to achieve the effect of improving data pulling efficiency. It can also avoid the problem of inefficiency caused by re-requesting when a request fails due to a sudden failure of a path.
[0103] Figure 4 A flow chart of another path selection method provided in an embodiment of the present application is shown as follows: Figure 4 As shown, the method includes:
[0104] S401: Determine a target node corresponding to a data pull request in a wide area storage service cluster.
[0105] Among them, the data pull request may refer to the client or application actively initiating a request to the wide area storage service cluster to obtain specific data resources. The data pull request may include the required data information and the target node information.
[0106] S402: Determine the shortest path to the target node according to the undirected weighted graph in the wide area storage service cluster.
[0107] Among them, an undirected weighted graph may refer to a structural graph consisting of nodes and edges connecting these nodes, each edge having a weight associated therewith, representing the cost, distance or other metric of the edge, and in particular, in an undirected weighted graph, the edges are non-directional.
[0108] The method for determining the shortest path between target nodes may include first determining all paths that can reach the target node in the wide area storage service cluster based on an undirected weighted graph, then calculating the weight of each path based on the weighted edges between the nodes in the path, and determining the shortest path among them based on the weight of each path.
[0109] S403: Determine the second shortest path according to the undirected weighted graph, the shortest path and a preset weight difference threshold, and both the shortest path and the second shortest path are used to transmit a request to pull data to a target node.
[0110] Among them, the method for determining the second shortest path may include comparing the weight differences between multiple target paths and the shortest path with a preset weight difference threshold, and if the weight difference is not greater than the weight difference threshold, determining that the target path is the second shortest path.
[0111] Another path selection method provided in an embodiment of the present application can determine the shortest path based on an undirected weighted graph, and then calculate the difference in weights between other paths and the shortest path, and determine the second shortest path with a weight close to the shortest path based on the difference in weights and a preset weight difference threshold, and achieve the effect of improving data pulling efficiency by simultaneously transmitting data pulling requests through the shortest path and the second shortest path.
[0112] Figure 5 A flow chart of a data extraction method provided in an embodiment of the present application is shown as follows: Figure 5 As shown, the method includes:
[0113] S501 . In response to receiving a request to pull data, determine a shortest path and a second shortest path for transmitting the request to pull data, where the shortest path and the second shortest path are obtained according to the above path selection method.
[0114] The shortest path may refer to a path with the best comprehensive performance among multiple target paths.
[0115] The second shortest path may refer to a path whose weight difference with the shortest path is within a preset range, and there may be multiple second shortest paths.
[0116] S502: Transmit a request to pull data to a target node according to the shortest path and the second shortest path.
[0117] The target node may refer to a storage service node whose information needs to be pulled.
[0118] S503: Receive returned target data, where the target data is the fastest returned demand data among the shortest path and the second shortest path.
[0119] Among them, the shortest path and the second shortest path both pull the required data from the target node to the storage service node that issues the data pull request based on the data pull request. The storage service node that issues the data pull request receives the required data returned first and uses it as the target data. The required data returned later is discarded.
[0120] In this embodiment of the present application, after receiving the returned target data, the method may further include:
[0121] The adjacent intermediate nodes in the shortest path and the second shortest path cache the demand data to obtain the cached demand data. The adjacent intermediate nodes are intermediate nodes connected to the storage service node. The intermediate nodes are nodes between the storage service node and the target node.
[0122] If the shortest path and the second shortest path receive a new data pull request, and the target of the new data pull request is the demand data, the storage service node receives the demand data of the returned target cache, and the demand data of the target cache is the demand data of the cache that returns the fastest among the adjacent intermediate nodes of the shortest path and the adjacent intermediate nodes of the second shortest path.
[0123] Among them, when the shortest path and the second shortest path pull the required data according to the data pull request, the required data can be cached at the adjacent intermediate node. The adjacent intermediate node is a node in the intermediate node that is connected to the storage service node that issues the data pull request. If the storage service node issues the same data pull request again, after selecting the shortest path and the second shortest path, when the data pull request is transmitted according to the shortest path and the second shortest path, it passes through the adjacent intermediate nodes of the shortest path and the second shortest path. At this time, the adjacent intermediate node can directly return the previously cached required data, compared with transmitting to the target node and then pulling the required data and then returning it, it saves time and improves efficiency.
[0124] In other embodiments, the nodes for caching the required data can be set according to the specific situation. The caching function can be turned on for the intermediate nodes with larger memory, and the caching function can be turned off for the intermediate nodes with smaller memory. On the basis of ensuring the working capacity of the nodes, the efficiency of data pulling is improved.
[0125] When the adjacent intermediate nodes of the shortest path and the adjacent intermediate nodes of the second shortest path return the cached demand data, the storage service node that issues the data pull request receives the cached demand data returned first and uses it as the target cached demand data, and discards the cached demand data returned later.
[0126] In the embodiment of the present application, after the demand data is cached at adjacent intermediate nodes in the shortest path and the second shortest path, and the cached demand data is obtained, the method may further include:
[0127] If the shortest path and the second shortest path do not receive a new request to pull data, or the target of the new request to pull data is not the required data, then the cache retention time of the adjacent intermediate node is obtained;
[0128] If the cache retention time is longer than the preset cache retention time threshold, the cached demand data is deleted.
[0129] The cache retention time may refer to the cache time of the demand data at the adjacent intermediate node.
[0130] When the cache retention time is greater than the preset cache retention time threshold, it means that the demand data has cooled down. In order to save node space, the cooled cache demand data can be deleted. The cache retention time threshold can be set according to the specific situation and is not specifically limited here.
[0131] A data pulling method provided in an embodiment of the present application can cache required data when pulling data. When a new request to pull data arrives, the cached required data is returned, and the cached data that has not been pulled for a long time is deleted. On the basis of ensuring the working capacity of the node, the efficiency of data pulling is improved.
[0132] A wide-area storage service cluster system includes a central service node and storage service nodes corresponding to multiple regions. The central service node is used to periodically detect the attribute information and reachability information of multiple storage service nodes, and update the node path map according to the reachability information. The storage service node is used to execute the above-mentioned path selection method.
[0133] Among them, the storage service node may include a BOSS service unit, and the central service node can be connected to the BOSS service units in multiple storage service nodes respectively, and push the periodically detected attribute information and reachability information, as well as the node path map to the BOSS service unit in the storage service node.
[0134] Reachability information may refer to data describing the status of each storage service node itself, as well as data describing the connection status between each storage service node.
[0135] For example Figure 6 , the method of updating the node path graph according to the reachability information may include updating the state of each node itself and the connection state between the nodes in the node path graph. When a network failure is detected between node 2 and node 3, the connection state between node 2 and node 3 is updated in the node path graph as unreachable, and the connection between node 2 and node 3 is disconnected to avoid the faulty connection from affecting the process of pulling data.
[0136] In other embodiments, the central service node may also update the undirected weighted graph according to the reachability information and the attribute information. Specifically, the nodes and edges in the undirected weighted graph may be updated according to the reachability information, and the weights of the edges in the undirected weighted graph may be updated according to the attribute information.
[0137] The storage service node can determine the target path and the weight of each target path based on the attribute information and the node path graph, and then execute the above path selection method.
[0138] Figure 7 This is a schematic diagram of the structure of the path selection device provided in the embodiment of the present application. Figure 7 As shown, the path selection device 70 includes: a first determination module 701, a second determination module 702, a third determination module 703, a fourth determination module 704 and a fifth determination module 705. Among them:
[0139] The first determination module 701 is used to determine the target node corresponding to the data pull request in the wide area storage service cluster;
[0140] A second determination module 702 is used to determine multiple target paths to the target node according to the node path graph in the wide area storage service cluster;
[0141] The third determination module 703 is used to determine the path weights of multiple target paths according to the attribute information of the nodes in the wide area storage service cluster;
[0142] A fourth determination module 704, configured to determine the shortest path among multiple target paths according to the path weights;
[0143] The fifth determination module 705 is used to determine the second shortest path according to the path weight, the shortest path and a preset weight difference threshold, and both the shortest path and the second shortest path are used to transmit a request to pull data to the target node.
[0144] In the embodiment of the present application, the third determining module 703 may also be used to:
[0145] Determine the node weight according to the attribute information of the nodes in the wide area storage service cluster;
[0146] According to the node weights, the path weights of multiple target paths are determined.
[0147] In the embodiment of the present application, the third determining module 703 may also be used to:
[0148] Determine a first weight of the node according to the network delay information in the attribute information, determine a second weight of the node according to the bandwidth capacity information in the attribute information, and determine a third weight of the node according to the cost information in the attribute information;
[0149] The node weight is determined according to the node first weight, the node second weight and the node third weight, wherein the node first weight and the node third weight are in inverse proportion to the node weight, and the node second weight is in direct proportion to the node weight.
[0150] In the embodiment of the present application, the fifth determining module 705 may also be used to:
[0151] Determine the weight differences between the multiple target paths and the shortest path according to the path weight and the path weight of the shortest path;
[0152] The second shortest path is determined according to the weight difference and a preset weight difference threshold, where the second shortest path is a path whose weight difference with the shortest path is not greater than the preset weight difference threshold.
[0153] As can be seen from the above, the path selection device of the embodiment of the present application consists of a first determination module 701, which is used to determine the target node corresponding to the data pull request in the wide area storage service cluster; a second determination module 702, which is used to determine multiple target paths between the target nodes based on the node path map in the wide area storage service cluster; a third determination module 703, which is used to determine the path weights of multiple target paths based on the attribute information of the nodes in the wide area storage service cluster; a fourth determination module 704, which is used to determine the shortest path among multiple target paths based on the path weight; and a fifth determination module 705, which is used to determine the second shortest path based on the path weight, the shortest path and a preset weight difference threshold, and both the shortest path and the second shortest path are used to transmit the data pull request to the target node. Therefore, the embodiment of the present application can determine the shortest path through the node path graph and the node attribute information, and then calculate the difference in weights between other paths and the shortest path. According to the difference in weights and a preset weight difference threshold, the second shortest path with a weight close to the shortest path is determined, and the technical means of simultaneously transmitting the data pulling request through the shortest path and the second shortest path can be used to achieve the effect of improving data pulling efficiency. It can also avoid the problem of inefficiency caused by re-requesting when a request fails due to a sudden failure of a path.
[0154] Figure 8 This is a schematic diagram of the structure of the data extraction device provided in the embodiment of the present application. Figure 8 As shown, the data pulling device 80 includes: a response module 801, a transmission module 802 and a receiving module 803. Among them:
[0155] A response module 801, in response to receiving a pull data request, determines a shortest path and a second shortest path for transmitting the pull data request, wherein the shortest path and the second shortest path are obtained according to the above-mentioned path selection method;
[0156] The transmission module 802 transmits the pull data request to the target node according to the shortest path and the second shortest path;
[0157] The receiving module 803 is used to receive the returned target data, where the target data is the fastest returned demand data among the shortest path and the second shortest path.
[0158] In the embodiment of the present application, the receiving module 803 may also be used for:
[0159] The adjacent intermediate nodes in the shortest path and the second shortest path cache the demand data to obtain the cached demand data. The adjacent intermediate nodes are intermediate nodes connected to the storage service node. The intermediate nodes are nodes between the storage service node and the target node.
[0160] If the shortest path and the second shortest path receive a new data pull request, and the target of the new data pull request is the demand data, the storage service node receives the demand data of the returned target cache, and the demand data of the target cache is the demand data of the cache that returns the fastest among the adjacent intermediate nodes of the shortest path and the adjacent intermediate nodes of the second shortest path.
[0161] In the embodiment of the present application, the receiving module 803 may also be used for:
[0162] If the shortest path and the second shortest path do not receive a new request to pull data, or the target of the new request to pull data is not the required data, then the cache retention time of the adjacent intermediate node is obtained;
[0163] If the cache retention time is longer than the preset cache retention time threshold, the cached demand data is deleted.
[0164] As can be seen from the above, the path selection device of the embodiment of the present application is composed of a response module 801, which responds to receiving a request to pull data and determines the shortest path and the second shortest path for transmitting the request to pull data, and the shortest path and the second shortest path are obtained according to the above path selection method; a transmission module 802, which transmits the request to pull data to the target node according to the shortest path and the second shortest path; a receiving module 803, which is used to receive the returned target data, and the target data is the fastest returned demand data in the shortest path and the second shortest path. Therefore, the embodiment of the present application can cache the demand data when pulling data, and when a new request to pull data arrives, the cached demand data is returned, and the cached data that has not been pulled for a long time is deleted, so as to improve the efficiency of data pulling on the basis of ensuring the working capacity of the node.
[0165] Fig. 9 This is a schematic diagram of the structure of the electronic device provided in this application. Fig. 9 As shown, the electronic device 90 provided in this embodiment includes: at least one processor 901 and a memory 902. Optionally, the device 90 further includes a communication component 903. The processor 901, the memory 902 and the communication component 903 are connected via a bus 904.
[0166] In a specific implementation process, at least one processor 901 executes the computer execution instructions stored in the memory 902, so that at least one processor 901 executes the above method.
[0167] The specific implementation process of the processor 901 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.
[0168] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the invention may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0169] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0170] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0171] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0172] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.
[0173] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.
[0174] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0175] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0176] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0177] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0178] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0179] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0180] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A path selection method, characterized in that: A storage service node applied to a wide-area storage service cluster, wherein the wide-area storage service cluster includes storage service nodes corresponding to a plurality of regions, wherein the storage service node stores data of the corresponding region, and the path selection method includes: Determine a target node corresponding to the data pull request in the wide area storage service cluster; Determine multiple target paths to the target node according to the node path graph in the wide area storage service cluster; Determining path weights of the plurality of target paths according to attribute information of nodes in the wide area storage service cluster; Determining the shortest path among the multiple target paths according to the path weights; A second shortest path is determined according to the path weight, the shortest path and a preset weight difference threshold, and both the shortest path and the second shortest path are used to transmit the data pull request to the target node.
2. The method according to claim 1, characterized in that The determining the second shortest path according to the path weight, the shortest path and a preset weight difference threshold comprises: Determining weight differences between the multiple target paths and the shortest path according to the path weight and the path weight of the shortest path; The second shortest path is determined according to the weight difference and the preset weight difference threshold, wherein the second shortest path is a path whose weight difference with the shortest path is not greater than the preset weight difference threshold.
3. The method according to claim 1 or 2, characterized in that: The determining the path weights of the plurality of target paths according to the attribute information of the nodes in the wide area storage service cluster includes: Determining node weights according to attribute information of nodes in the wide area storage service cluster; The path weights of the multiple target paths are determined according to the node weights.
4. The method according to claim 3, characterized in that The determining the node weight according to the attribute information of the nodes in the wide area storage service cluster includes: Determine a first node weight according to the network delay information in the attribute information, determine a second node weight according to the bandwidth capacity information in the attribute information, and determine a third node weight according to the cost information in the attribute information; The node weight is determined according to the node first weight, the node second weight and the node third weight, wherein the node first weight and the node third weight are inversely proportional to the node weight, and the node second weight is in direct proportion to the node weight.
5. A data extraction method, characterized in that: The storage service nodes used in the wide-area storage service cluster include: In response to receiving a pull data request, determining a shortest path and a second shortest path for transmitting the pull data request, wherein the shortest path and the second shortest path are obtained according to the path selection method according to any one of claims 1 to 4; Transmitting the data pull request to a target node according to the shortest path and the second shortest path; The returned target data is received, where the target data is the fastest returned required data among the shortest path and the second shortest path.
6. The method according to claim 5, characterized in that After receiving the returned target data, the method further includes: The adjacent intermediate nodes in the shortest path and the second shortest path cache the demand data to obtain the cached demand data, the adjacent intermediate nodes are intermediate nodes connected to the storage service node, and the intermediate nodes are nodes between the storage service node and the target node; If the shortest path and the second shortest path receive a new data pull request, and the target of the new data pull request is the demand data, the storage service node receives the demand data of the returned target cache, and the demand data of the target cache is the fastest returned cache demand data among the adjacent intermediate nodes of the shortest path and the adjacent intermediate nodes of the second shortest path.
7. The method according to claim 6, characterized in that After the adjacent intermediate nodes in the shortest path and the second shortest path cache the demand data and obtain the cached demand data, the method further includes: If the shortest path and the second shortest path do not receive a new request to pull data, or the target of the new request to pull data is not the required data, then obtaining the cache retention time of the adjacent intermediate node; If the cache retention time is longer than a preset cache retention time threshold, the cached demand data is deleted.
8. A wide area storage service cluster system, characterized in that: It includes a central service node and storage service nodes corresponding to multiple regions. The central service node is used to periodically detect the attribute information and reachability information of the multiple storage service nodes, and update the node path map according to the reachability information. The storage service node is used to execute the path selection method described in any one of claims 1 to 4.
9. A path selection device, characterized in that: include: A first determination module is used to determine a target node corresponding to a data pull request in the wide area storage service cluster; A second determination module is used to determine multiple target paths to the target node according to a node path graph in the wide area storage service cluster; A third determination module, configured to determine path weights of the plurality of target paths according to attribute information of nodes in the wide area storage service cluster; a fourth determination module, configured to determine the shortest path among the plurality of target paths according to the path weights; The fifth determination module is used to determine the second shortest path according to the path weight, the shortest path and a preset weight difference threshold, and the shortest path and the second shortest path are both used to transmit the data pull request to the target node.
10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 4.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 4 when executed by a processor.
12. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 4 when being executed by a processor.
Citation Information
Patent Citations
Network path determination method and device, computer equipment and storage medium
CN117354229A
Policy based path management
US20190238412A1
Server-side path selection in multi virtual server environment
US20200177499A1
Bandwidth constraint for multipath segment routing
US20220103458A1
Multi-channel packet forwarding method and device
WO2017140112A1