Manipulating a distributed agreement protocol to identify a desired set of storage units
Patent Information
- Application Number
- DE112017002144
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-07-12
- Filing Date
- 2017-07-04
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2037-07-04
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND OF THE INVENTION TECHNICAL FIELD OF THE INVENTION
[0001] The invention relates generally to computer networks and, more particularly, to scattering error-encoded data. DESCRIPTION OF THE RELEVANT PRIOR ART
[0002] The document US 2010 / 0 287 200 A1 describes a user device including a browser module, a DSN interface to a local or external DSN memory, and a DS processing module coupled to the DSN interface for storing and retrieving the data object from the DSN memory, wherein the data object is divided into a plurality of data segments, and wherein each of the plurality of data segments is stored in the DSN memory as a plurality of encoded data slices generated based on an error coding distribution function.
[0003] The document US 2015 / 0 378 625 A1 describes a method that identifies a change in the DSN storage of a dispersed storage network (DSN), generates updated storage evaluation results, identifies an updated set of storage units, and sends at least one data migration request to at least one storage unit of the updated set of storage units.
[0004] Computing units are known for transmitting, processing, and / or storing data. Such computing units range from wireless smartphones, laptops, tablets, PCs, workstations, and video game devices to data centers that support millions of web searches, stock trades, or online purchases every day. Generally, a computing unit includes a central processing unit (CPU), a memory system, user input / output interfaces, peripheral interfaces, and an interconnecting bus structure.
[0005] As is further known, by using "cloud computing," a computer can effectively extend its CPU to perform one or more computing functions for the computer (e.g., a service, application, algorithm, arithmetic logic function, etc.). Furthermore, cloud computing for large-scale services, applications, and / or functions can be executed across multiple cloud computing resources in a distributed manner to improve the response time for completing the service, application, and / or function. For example, Hadoop is an open-source software framework that supports distributed applications by enabling application execution by thousands of computers.
[0006] In addition to cloud computing, a computer can use "cloud storage" as part of its memory storage system. As is well known, cloud storage allows a user to store data, applications, etc., via their computer on an internet storage system. The internet storage system may include a RAID (Redundant Array of Independent Disks) system and / or a distributed storage system that uses an error-correcting scheme to encode data for storage.
[0007] In a distributed storage system containing multiple storage units, there are cases where it is more efficient, faster, and / or more reliable for a computing unit to write to a particular group of storage units than to other groups of storage units. In distributed storage systems using a resilience-balancing storage protocol, selection of specific groups of storage units is typically not permitted. Thus, a computing unit is assigned a group of storage units to write to that may not be the most efficient, fastest, or most reliable for the computing unit. SUMMARY OF THE INVENTION
[0008] According to one aspect, a method is provided comprising: obtaining a plurality of groups of encoded data chunks for storage in a memory of the DSN from a computing device of a distributed storage network (DSN), wherein the memory of the DSN includes a plurality of pools of storage devices located across a geographical area, wherein a data object is divided into a plurality of data segments, and wherein the plurality of data segments is a distributed storage error encoded into the plurality of groups of encoded data chunks; identifying, by the computing device, a desired group of storage devices within the plurality of pools of storage devices to store the plurality of groups of encoded data chunks;generating, by the computing device, a particular source name based on the desired group of storage units and a distributed agreement protocol (DAP), wherein the DAP identifies a group of storage units from the plurality of pools of storage units based on a slice identifier and a plurality of storage allocation coefficients, and wherein, when a device in the DSN executes the DAP, the device uses the particular source name as the slice identifier to identify the desired group of storage units; generating, by the computing device, a plurality of sets of slice names for the plurality of sets of encoded data slices, wherein the plurality of sets of slice names includes the particular source name;and sending a plurality of groups of write requests by the data processing unit to the desired group of storage units with respect to the plurality of groups of encoded data slices and according to the plurality of groups of slice names;
[0009] According to another aspect, there is provided a computer-readable memory device comprising: a first memory element storing operational instructions that, when executed by a computing device of a distributed storage network (DSN), cause the computing device to: obtain a plurality of groups of encoded data chunks for storage in a memory of the DSN, wherein the memory of the DSN includes a plurality of pools of memory devices located across a geographical area, wherein a data object is divided into a plurality of data segments, and wherein the plurality of data segments are distributed memory errors encoded into the plurality of groups of encoded data chunks;and identifying a desired group of storage units within the plurality of pools of storage units for storing the plurality of groups of encoded data slices; a second memory element storing operational instructions that, when executed by the computing device, cause the computing device to: generate a particular source name based on the desired group of storage units and a distributed agreement protocol (DAP), wherein the DAP identifies a group of storage units from the plurality of pools of storage units based on a slice identifier and a plurality of memory allocation coefficients, and wherein, when a device in the DSN executes the DAP, the device uses the particular source name as the slice identifier to identify the desired group of storage units;and generating a plurality of groups of snippet names for the plurality of groups of encoded data snippets, the plurality of groups of snippet names including the determined source name; and a third memory element storing operational instructions that, when executed by the computing device, cause the computing device to: send a plurality of groups of write requests to the desired group of storage devices with respect to the plurality of groups of encoded data snippets and according to the plurality of groups of snippet names;
[0010] According to another aspect, a computing unit of a distributed storage network (DSN) is provided, the computing unit comprising: an interface; a memory; and a processing module operably connected to the memory and the interface, the processing module operable to: obtain a plurality of groups of encoded data chunks for storage in a memory of the DSN, wherein the memory of the DSN includes a plurality of pools of storage units positioned across a geographical area, wherein a data object is divided into a plurality of data segments, and wherein the plurality of data segments are scattered memory errors encoded into the plurality of groups of encoded data chunks;Identifying a desired group of storage units within the plurality of pools of storage units for storing the plurality of groups of encoded data snippets; Generating a particular source name based on the desired group of storage units and a distributed agreement protocol (DAP), wherein the DAP identifies a group of storage units from the plurality of pools of storage units based on a snippet identifier and a plurality of memory allocation coefficients, and wherein, when a unit in the DSN executes the DAP, the unit uses the particular source name as the snippet identifier to identify the desired group of storage units; Generating a plurality of groups of snippet names for the plurality of groups of encoded data snippets, wherein the plurality of groups of snippet names includes the particular source name;and sending a plurality of groups of write requests to the desired group of storage units with respect to the plurality of groups of encoded data slices and according to the plurality of groups of slice names; BRIEF DESCRIPTION OF THE DIFFERENT VIEWS OF THE DRAWING(S)
[0011] Preferred embodiments of the present invention are described below by way of example only and with reference to the following drawings: Fig. 1 is a schematic block diagram of an embodiment of a distributed storage network (DSN) according to a preferred embodiment of the present invention; Fig. 2 is a schematic block diagram of an embodiment of a computing core according to a preferred embodiment of the present invention; Fig. 3 is a schematic block diagram of an example of scattered memory error encoding of data according to a preferred embodiment of the present invention; Fig. 4 is a schematic block diagram of a general example of an error encoding function according to a preferred embodiment of the present invention; Fig. 5 is a schematic block diagram of a specific example of an error encoding function according to a preferred embodiment of the present invention; Fig. 6 is a schematic block diagram of an example slice name of an encoded data slice (EDS) according to a preferred embodiment of the present invention; Fig. 7 is a schematic block diagram of an example of scattered memory error decoding of data according to a preferred embodiment of the present invention; Fig. 8 is a schematic block diagram of a general example of an error decoding function according to a preferred embodiment of the present invention; Fig. 9 is a schematic block diagram of an embodiment of a decentralized or distributed agreement protocol (DAP) according to a preferred embodiment of the present invention; Fig. 10 is a schematic block diagram of an example of creating multiple groups of cutouts according to a preferred embodiment of the present invention; Fig. 11 is a schematic block diagram of an example of storage vaults according to a preferred embodiment of the present invention; and Fig. 12 is a logic diagram of an example method for manipulating a DAP to identify a desired group of storage units in accordance with a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0012] Fig. 1 is a schematic block diagram of one embodiment of a distributed storage network (DSN) 10 including a plurality of computing devices 12-16, a management device 18, an integrity processing device 20, and a DSN memory 22. The components of the DSN 10 are interconnected via a network 24, which may include one or more wireless and / or wired communication systems; one or more private intranet systems and / or public internet systems; and / or one or more local area networks (LANs) and / or wide area networks (WANs).
[0013] The DSN memory 22 includes a plurality of storage units 36, which may be located at geographically different locations (e.g., one in Chicago, one in Milwaukee, etc.), at a common location, or a combination thereof. For example, if the DSN memory 22 includes eight storage units 36, each storage unit is positioned at a different location. In another example, if the DSN memory 22 includes eight storage units 36, all eight storage units are located at the same location.In yet another example, when DSN memory 22 includes eight memory units 36, a first pair of memory units is located at a first common location, a second pair of memory units is located at a second common location, a third pair of memory units is located at a third common location, and a fourth pair of memory units is located at a fourth common location. Note that a DSN memory 22 may include more or fewer than eight memory units 36. Further, note that each memory unit 36 may comprise a compute core (as shown in FIG. Fig. 2 or components thereof) and a plurality of memory units for storing scattered error-encoded data.
[0014] Each of the computing devices 12 through 16, the processing device 18, and the integrity processing device 20 includes a computing core 26 that includes network interfaces 30 through 33. The computing devices 12 through 16 may each be a portable computing device and / or a fixed computing device. A portable computing device may be a social networking device, a gaming device, a mobile phone, a smartphone, a digital assistant, a digital music player, a digital video player, a laptop computer, a handheld computer, a tablet, a video game controller, and / or any other portable device that includes a computing core.A fixed computing device may be a computer (PC), a computer server, a cable set-top box, a satellite receiver, a television, a printer, a fax machine, entertainment electronics, a video game console, and / or any type of personal or business computing equipment. Note that both the management device 18 and the integrity device 20 may be separate computing devices, a common computing device, and / or integrated with one or more of the computing devices 12-16 and / or one or more of the storage devices 36.
[0015] Each interface 30, 32, and 33 includes software and hardware for indirectly and / or directly supporting one or more communication connections over the network 24. For example, the interface 30 supports a communication connection (e.g., wired, wireless, direct, over a LAN, over the network 24, etc.) between computing devices 14 and 16. As another example, the interface 32 supports communication connections (e.g., a wired connection, a wireless connection, a LAN connection, and / or any other type of connection to / from the network 24) between the computing devices 12 and 16 and the DSN memory 22. As yet another example, the interface 33 supports a communication connection to the network 24 for both the management device 18 and the integrity processing device 20.
[0016] The data processing units 12 and 16 include a distributed memory (DS) client module 34 that enables the data processing unit to perform distributed memory error encoding and decoding of data (e.g., data 40), as described below with reference to one or more of the Fig. 3 to 8. In this exemplary embodiment, the computing unit 16 operates as a distributed storage processing agent for the computing unit 14. In this role, the computing unit 16 performs distributed storage error encoding and decoding of data for the computing unit 14. By using distributed storage error encoding and decoding, the DSN 10 tolerates a significant number of storage unit errors (the number of errors is based on parameters of the distributed storage error encoding function) without data loss and without the need for redundant or backup copies of the data. Furthermore, the DSN 10 stores data for an indefinite period of time without data loss and in a secure manner (e.g., the system is highly resilient to unauthorized attempts to access the data).
[0017] During operation, management unit 18 performs DS management services. For example, management unit 18 creates distributed data storage parameters (e.g., vault creation, distributed storage parameters, security parameters, billing information, user profile information, etc.) individually or as part of a group of user units for computing units 12 through 14. As a specific example, management unit 18 coordinates the creation of a vault (e.g., a virtual memory block associated with a portion of a DSN aggregate namespace) in DSN memory 22 for a user unit, a group of units, or for public access, and creates per-vault distributed storage (DS) error encoding parameters for a vault.The management unit 18 facilitates the storage of DS error encoding parameters for each vault by updating registry information for the DSN 10, where the registry information may be stored in the DSN memory 22, a data processing unit 12 to 16, the management unit 18 and / or the integrity processing unit 20.
[0018] The management unit 18 creates and stores user profile information (e.g., an access control list (ACL)) in a local memory and / or in a memory of the DSN memory 22. The user profile information contains authentication information, permissions, and / or the security parameters. The security parameters may include an encryption / decryption scheme, one or more encryption keys, a key generation scheme, and / or a data encoding / decoding scheme.
[0019] Management unit 18 creates billing information for a specific user, user group, vault access, public vault access, etc. For example, management unit 18 tracks how often a user accesses a non-public vault and / or public vaults, which can be used to generate billing information per access. In another example, management unit 18 tracks the amount of data stored and / or accessed by a user unit and / or user group, which can be used to generate billing information by data volume.
[0020] In another example, the management unit 18 performs network operations, network management, and / or network maintenance. Network operations include authenticating user data allocation requests (e.g., read and / or write requests), managing vault creations, creating authentication data for user units, adding / deleting components (e.g., user units, storage units, and / or compute units with a DS-Client module 34) to / from the distributed DSN 10, and / or creating authentication data for the storage units 36. Network management includes monitoring devices and / or units for errors, managing vault information, determining device and / or unit activation states, determining device and / or unit loads, and / or determining any other system-level operation that affects the performance level of the DSN 10.Network maintenance includes simplifying the replacement, updating, repairing, and / or expanding of a unit and / or a unit of the DSN10.
[0021] The integrity processing unit 20 performs reconstruction of "invalid" or missing encoded data snippets. At a higher level, the integrity processing unit 20 performs reconstruction by periodically attempting to retrieve / list encoded data snippets and / or snippet names of the encoded data snippets from the DSN memory 22. Retrieved encoded snippets are checked for errors due to data corruption, outdated versions, etc. If a snippet contains an error, it is marked as an "invalid" snippet. Encoded data snippets that were not received and / or not listed are marked as missing snippets. Invalid and / or missing snippets are then reconstructed using other retrieved encoded data snippets that are determined to be good snippets to generate reconstructed snippets.The reconstructed sections are stored in DSN memory 22.
[0022] Fig. 2 is a schematic block diagram of one embodiment of a computing core 26 including a processing module 50, a memory controller 52, a main memory 54, a video graphics processing unit 55, an input / output (I / O) controller 56, a peripheral component interconnect (PCI) interface 58, an I / O interface module 60, at least one I / O device interface module 62, a basic input / output system (BIOS) read-only memory (ROM) 64, and one or more memory interface modules. The one or more memory interface modules include one or more of a Universal Serial Bus (USB) interface module 66, a Host Bus Adapter (HBA) interface module 68, a network interface module 70, a flash interface module 72, a hard disk interface module 74, and a DSN interface module 76.
[0023] The DSN interface module 76 operates to emulate a conventional operating system (OS) file system interface (e.g., network file system (NFS), flash file system (FFS), disk file system (DFS), file transfer protocol (FTP), web-based distributed authoring and versioning (WebDAV), etc.) and / or a block memory interface (e.g., small computer system interface (SCSI), Internet Small Computer System Interface (iSCSI), etc.). The DSN interface module 76 and / or the network interface module 70 may be implemented as one or more of the interfaces 30 through 33 of Fig. 1. Note that the I / O device interface module 62 and / or the memory interface modules 66 through 76 may be referred to collectively or individually as I / O ports.
[0024] Fig. 3 is a schematic block diagram of an example of scattered memory error encoding of data. When data of a data processing unit 12 or 16 needs to be stored, it performs scattered memory error encoding of the data according to a scattered memory error encoding process based on scattered memory error encoding parameters. The scattered memory error encoding parameters include an encoding function (e.g., information scattering algorithm, Reed-Solomon, Cauchy Reed-Solomon, systematic encoding, unsystematic encoding, online codes, etc.), a data segmentation protocol (e.g., data segment size, fixed, variable, etc.), and encoding values per data segment. The encoding values per data segment include a sum or a pillar width, a number (T) of encoded data segments per encoding of a data segment (i.e.,in a group of encoded data snippets); a decoding threshold number (D) of encoded data snippets of a group of encoded data snippets required to recover the data segment; a read threshold number (R) of encoded data snippets to specify a number of encoded data snippets per group to be read from the memory to decode the data segment; and / or a write threshold number (W) to specify a number of encoded data snippets per group that must be correctly stored before it can be assumed that the encoded data segment has been properly stored. The scattered memory error encoding parameters may further include snippet information (e.g., the number of encoded data snippets created for each data segment) and / or snippet security information (e.g.,per encryption, compression, integrity checksum, etc. per encoded data section).
[0025] In the present example, Cauchy Reed-Solomon was chosen as the encoding function (a general example is given in Fig. 4, and a specific example is shown in Fig. 5); the data segmentation protocol is used to divide the data object into fixed-size data segments; and the encoding values per data segment include: a pillar width of 5, a decoding threshold of 3, a read threshold of 4, and a write threshold of 4. According to the data segmentation protocol, the data processing unit 12 or 16 divides the data (e.g., a file (e.g., text, video, audio, etc.), a data object, or other data arrangement) into a plurality of fixed-size data segments (e.g., 1 to Y of a fixed size in the range of kilobytes to terabytes or higher). The number of data segments created depends on the size of the data and the data segmentation protocol.
[0026] The data processing unit 12 or 16 then performs memory error encoding of a data segment using the selected encoding function (e.g., Cauchy Reed-Solomon) to generate a group of encoded data segments. Fig. Figure 4 illustrates a general Cauchy Reed-Solomon encoding function that includes an encoding matrix (EM), a data matrix (DM), and an encoded matrix (CM). The size of the encoding matrix (EM) depends on the pillar width number (T) and the decoding threshold number (D) of selected encoding values per data segment. To generate the data matrix (DM), the data segment is divided into a plurality of data blocks, and the data blocks are arranged in a number D of rows with Z data blocks per row. Note that Z is a function of the number of data blocks created from the data segment and the decoding threshold number (D). The encoded matrix is generated by multiplying the data matrix by the encoded matrix.
[0027] Fig. Figure 5 illustrates a specific example of a Cauchy Reed-Solomon encoding with a pillar number (T) of five and a decoding threshold number of three. In this example, a first data segment is divided into twelve data blocks (D1 to D12). The encoded matrix contains five rows of encoded data blocks, where the first row from X11 to X14 corresponds to a first encoded data segment (EDS 1_1), the second row from X21 to X24 corresponds to a second encoded data segment (EDS 2_1), the third row from X31 to X34 corresponds to a third encoded data segment (EDS 3_1), the fourth row from X41 to X44 corresponds to a fourth encoded data segment (EDS 4_1), and the fifth row from X51 to X54 corresponds to a fifth encoded data segment (EDS 5_1). Please note that the second number of the EDS designation corresponds to the data segment number.
[0028] Referring again to the explanation of Fig. 3, the data processing unit also creates a snippet name (SN) for each encoded data snippet (EDS) in the group of encoded data snippets. A typical format for a snippet name 78 is shown in Fig. 6. As shown, the slice name (SN) 78 includes a pillar number of the encoded data slice (e.g., one from 1 to T), a data segment number (e.g., one from 1 to Y), a vault identifier, a data object identifier (ID), and may further include revision-level information of the encoded data slices. The slice name functions, at least in part, as a DSN address for the encoded data slice for storage and retrieval from the DSN memory 22.
[0029] As a result of the encoding, the data processing unit 12 or 16 generates a plurality of groups of encoded data snippets, which are provided to the storage units with their respective snippet names for storage. As shown, the first group of encoded data snippets includes EDS 1_1 to EDS 5_1, and the first group of snippet names includes SN 1_1 to SN 5_1, and the last group of encoded data snippets includes EDS 1_Y to EDS 5_Y, and the last group of snippet names includes SN 1_Y to SN 5_Y.
[0030] Fig. 7 is a schematic block diagram of an example of scattered memory error decoding of a data object used in the example of Fig. 4 scattered memory error encoded and stored. In this example, the computing device 12 or 16 retrieves from the storage devices at least the decoding threshold number of encoded data slices per data segment. In a specific example, the computing device retrieves a reading threshold number of encoded data slices.
[0031] To recover a data segment from a decoding threshold number of encoded data segments, the data processing unit uses a decoding function as in Fig. 8. As shown, the decoding function is essentially an inversion of the encoding function of Fig. 4. The encoded matrix contains a decoding threshold number of rows (e.g., three in this example), and the decoding matrix is an inverse of the encoding matrix containing the corresponding rows of the encoded matrix. For example, if the encoded matrix contains rows 1, 2, and 4, the encoding matrix is restricted to rows 1, 2, and 4 and then inverted to generate the decoding matrix.
[0032] Fig. Figure 9 is a schematic block diagram of one embodiment of a decentralized or distributed agreement protocol (DAP) 80, which may be implemented by a computing device, a storage device, and / or any other device or device of the DSN, for determining where to store encoded data snippets or where to find stored encoded data snippets. The DAP 80 includes a plurality of functional evaluation modules 81. Each of the functional evaluation modules 81 includes a deterministic function 83, a normalization function 85, and an evaluation function 87.
[0033] Each functional evaluation module 81 receives as inputs a slice identifier 82 and storage pool (SP) coefficients (e.g., a first functional evaluation module 81-1 receives SP-1 coefficients "a" and b). Based on the inputs, where the SP coefficients are different for each functional evaluation module 81, each functional evaluation module 81 generates a unique score 93 (e.g., an alphanumeric value, a numeric value, etc.). The ranking function 84 receives the unique scores 93 and sorts them based on an ordering function (e.g., from highest to lowest, from lowest to highest, alphabetically, etc.) and then selects one as a selected storage pool 86. Note that a storage pool contains one or more groups of storage units.
[0034] It should also be noted that the snippet identifier corresponds to a snippet name or to common attributes of a group of snippet names. For example, snippet identifier 82 for a group of encoded data snippets specifies a data segment number, a vault ID, and a data object ID, but leaves the pillar number open. In another example, snippet identifier 82 specifies a range of snippet names (e.g., 0000 0000 to FFFF FFFF).
[0035] In a specific example, the first functional evaluation module 81-1 receives the slice identifier 82 and SP coefficients for storage pool 1 of the DSN. The SP coefficients include a first coefficient (e.g., "a") and a second coefficient (e.g., "b"). The first coefficient is, for example, a unique identifier for the corresponding storage pool (e.g., the ID SP #1 for the coefficient "a" of SP 1), and the second coefficient is a weighting factor for the storage pool. The weighting factors are derived to permanently ensure that data is stored in the storage pools in a fair and distributed manner based on the capacities of the storage units in the storage pools.
[0036] For example, the weighting factor contains an arbitrary bias that adjusts a proportion of selections to an associated location, such that the probability that a source name is assigned to that location is equal to the location weight divided by the sum of all location weights for all comparison locations (e.g., locations corresponding to storage units). In a specific example, each storage pool is associated with a location weighting factor based on storage capacity, such that storage pools with more storage capacity have a higher location weighting factor than storage pools with less storage capacity.
[0037] The deterministic function 83, which may be a hashing function, a hash-based message authentication code function, a cyclic redundancy code function, a hashing module of a number of locations, consistent hashing, rendezvous hashing, and / or a sponge function, performs a deterministic function on a combination and / or concatenation (e.g., adding, appending, interleaving) of the slice identifier 82 and the first SP coefficient (e.g., SU 1 coefficient "a") to produce an intermediate result 89.
[0038] The normalization function 85 normalizes the intermediate result 89 to produce a normalized intermediate result 91. For example, the normalization function 85 divides the intermediate result 89 by a number of possible output permutations of the deterministic function 83 to produce the normalized intermediate result. For example, if the intermediate result is 4,325 (decimal) and the number of possible output permutations is 10,000, the normalized result is 0.4325.
[0039] The evaluation function 87 performs a mathematical function on the normalized result 91 to generate the score 93. The mathematical function can be division, multiplication, addition, subtraction, a combination thereof, and / or any mathematical operation. For example, the evaluation function divides the second SP coefficient (e.g., SP 1 coefficient "b") by the negative logarithm of the normalized result (e.g., e y= x and / or In(x) = y). For example, if the second SP coefficient is 17.5, and the negative logarithm of the normalized result is 1.5411 (e.g., e (0.4235) ), the score is 11.3555.
[0040] The ranking function 84 receives the scores 93 from each of the feature evaluation modules 81 and sorts them to generate a ranking of the memory pools. For example, if the sorting is from highest to lowest and there are five memory units in the DSN, the ranking function evaluates the scores for five memory units to place them in an ordered sequence. From the ranking, the ranking module 84 selects one of the memory pools 86 that is the destination for a group of encoded data slices.
[0041] The DAP 80 can also be used to identify a group of storage units, a single storage unit, and / or a working memory unit within the storage unit. To achieve different output results, the coefficients are modified according to the desired location information. The DAP 90 can also output the ordered sequence of scores.
[0042] Fig. 10 is a schematic block diagram of an example of creating multiple groups of snippets. Each plurality of groups of encoded data snippets (EDS) corresponds to the encoding of a data object, a portion of a data object, or multiple data objects, where a data object is one or more of a file, text, data, digital information, etc. The visually highlighted plurality of encoded data snippets corresponds, for example, to a data object with a data identifier of "a2."
[0043] Each encoded data slice from each group of encoded data slices is uniquely identified by its slice name, which is also used as at least part of the DSN address for storing the encoded data slice. As shown, an EDS group contains EDS 1_1_1_a1 through EDS 5_1_1_a1. The EDS number includes a pillar number, data segment number, vault ID, and data object ID. Thus, for EDS 1_1_1_a1, it is the first EDS of a first data segment of a data object "a1" and is to be stored in or is stored in vault 1. Note that vaults are a logical memory container supported by the storage units of the DSN. A vault can be assigned to one or more user compute units.
[0044] As further shown, a further plurality of groups of encoded data segments are stored in Vault 2 for the data object "b2." There are Y EDS groups, where Y corresponds to the number of data segments created by segmenting the data object. The last EDS group of the data object "b1" contains EDS 1_Y_2_b1 through EDS 5_Y_2_b1. Thus, for EDS 1_Y_2_b1, it is the first EDS of the last data segment "Y" of a data object "b1" and is to be stored or is stored in Vault 2.
[0045] Fig. 11 is a schematic block diagram of an example of a plurality of storage pools (e.g., pool 1 to pool n) supporting one or more storage vaults. Each storage pool (e.g., 1 to n) contains one or more groups of storage units, where the number of storage units in a group of storage units corresponds to the pillar width number of groups of encoded
[0046] corresponds to the data segments it stores. For example, storage pool 1 and storage pool 2 each contain seven storage units, and storage pool n contains twelve storage units. Note that a storage pool may have more or fewer storage units than illustrated, and the number of storage units may vary from storage pool to storage pool.
[0047] In this example, storage pools 1 through n support three vaults (Vault 1, Vault 2, and Vault 3). Vaults 1 and 2 use five of the storage units and span multiple storage pools. Vault 3 uses seven of the storage units and is located only in storage pool "n." The number of storage units in a vault corresponds to the pillar width number, which in this example is five for Vaults 1 and 2 and seven for Vault 3. Note that a storage pool can have rows of storage units, where SU #1 represents a group of storage units, each corresponding to a first pillar number; SU #2 represents a second plurality of storage units, each corresponding to a second pillar number; and so on. For example, the field labeled Storage Unit SU #1 of storage pool 1 is representative of a plurality of storage units.
[0048] Each unit (e.g., the data processing units 12 to 16, the management unit 18, the integrity processing unit 29, the storage unit) of the DSN may contain the distributed agreement protocol 80 as described in Fig. 9. The DAP 80 uses snippet identifiers (e.g., the snippet name and / or one or more common elements thereof (e.g., the pillar number, the data segment number, the vault ID, and / or the data object ID)) to identify a group or pool of storage units for one or more groups of encoded data snippets. With respect to the three pluralities of groups of encoded data snippets (EDS) of Fig. 11, the DAP 80 distributes the groups of encoded data segments approximately evenly across the DSN memory (e.g., across the various memory units).
[0049] A computing device may specify a group of storage devices as a destination for storing its encoded data fragments because, from the perspective of computing devices, the target group of storage devices can be written to more efficiently, reliably, and / or quickly than a group allocated via the DAP. As described in more detail with reference to Fig. As described in Figure 12, the computing unit can manipulate the DAP so that the DAP allocates the desired group of storage units. To do this, the computing unit creates a specific slice identifier (e.g., a specific source name) instead of one containing one or more randomly generated components (e.g., a data object ID).
[0050] If many data processing units manipulate the DAP over time to obtain the desired group of memory units, memory asymmetry is likely to occur. The memory asymmetry (e.g., DAP asymmetry) can be corrected, at least to some extent, by the memory units. For example, the memory unit initiates a change in the DAP coefficients (e.g., memory allocation coefficients), causing some of the encoded data segments stored in the desired groups of memory units to be stored in other memory units. This will be discussed in more detail with reference to Fig. 12 explained.
[0051] Fig. 12 is a logic diagram of an example method for manipulating a distributed agreement protocol (DAP) to identify a desired group of storage devices. The method begins with step 100, in which a computing device of a distributed storage network obtains (e.g., receives, creates, etc.) a plurality of encoded data snippets for storage in the memory of the DSN (e.g., the DSN memory contains a plurality of pools of storage devices (e.g., a pool contains one or more groups of storage devices) located across a geographic area.
[0052] The method continues with step 102, where the computing device identifies a desired group of storage devices within the plurality of pools of storage devices to store the plurality of groups of encoded data snippets. For example, the computing device identifies the desired group of storage devices within the plurality of pools of storage devices with a desired write speed (e.g., data transfer rate and write latency are better than a group of storage devices identified by the DAP). In another example, the computing device identifies the desired group of storage devices as a group with a desired efficiency (e.g., that meets and exceeds a write threshold in time for a majority of the groups of encoded stored data snippets).In yet another example, the computing device identifies the desired group of storage devices as the group having a desired reliability (e.g., consistency with at least a write threshold number of storage devices available for storing encoded data snippets).
[0053] Identifying how the computing device locates the desired group of storage devices can be accomplished in many ways. For example, the computing device performs a search function (e.g., it performs a search function to identify the group). In another example, the computing devices initiate a query that tests the desired write speed, efficiency, and / or reliability and receives a response. In another example, the computing device accesses a historical data set that tracks the write speed, efficiency, and / or reliability of the computing device's data when writing to groups of storage devices of the DSN. From the historical data set, the computing device selects the group of storage devices with the desired write characteristics (e.g.,desired write speed, desired efficiency, and / or desired reliability). In yet another example, the computing device accesses a lookup table to identify the group.
[0054] The method continues with step 104, in which the computing device generates a particular source name for the DAP (e.g., a particular data object ID combined with a vault ID) to identify a desired group of storage devices. For example, the computing device generates a particular data object ID based on the DAP so that the DAP identifies the desired group of storage devices from the plurality of pools of storage devices. The computing device generates the particular source name by generating a particular data object identifier and combining it with a vault identifier and / or revision-level information. Within the DSN, devices (e.g., the computing device, other computing devices, the storage devices, management device, integrity device, etc.) use the DSN to identify the desired group of storage devices.) the particular source name as the slice identifier when executing the DAP to identify the desired group of storage units as the storage units storing the plurality of groups of encoded data slices.
[0055] The method continues with step 106, in which the computing device generates a plurality of groups of slice names for the plurality of groups of encoded data slices. The computing device generates a slice name by combining the determined source name with a pillar number and a data segment number. The method continues with step 108, in which the computing device sends a plurality of groups of write requests to the desired group of storage devices corresponding to the plurality of groups of slice names with respect to the plurality of groups of encoded data slices.
[0056] The method continues with step 110, where the DSN's memory units determine whether a DAP imbalance exists. If the memory units determine that there is no DAP, the method loops back to step 100.
[0057] If a DAP asymmetry exists, the method continues with step 112, where the storage units initiate an adjustment of the allocation coefficients. For example, one of the storage units adjusts the memory allocation coefficients. In another example, a storage unit requests that the management unit adjust the memory allocation coefficients. The storage unit or management unit adjusts the memory allocation coefficients (e.g., coefficient b for one or more functional evaluation modules of Fig.9) such that one or more encoded data segments are transferred from storage units of the desired group of storage units to other storage units in the DSN.
[0058] The method continues with step 114, where the storage units execute the DAP using the determined source name and the adjusted storage allocation coefficients to identify one or more encoded data segments to be transferred to one or more other storage units. The method continues with step 116, where the storage units of the desired group of storage units transfer the one or more encoded data segments to the one or more other storage units.
[0059] It should be noted that terminologies as they may be used herein, such as bitstream, stream, signal sequence, etc. (or their equivalents) have been used interchangeably to describe digital information whose content corresponds to any number of desired types (e.g., data, video, voice, audio, etc., all of which may be referred to generically as "data").
[0060] The terms "substantially" and "approximately," as used herein, provide an industry-accepted tolerance for the respective term and / or a relativity between elements. Such industry-accepted tolerance ranges from less than one percent to fifty percent and corresponds to, but is not limited to, component values, integrated circuit process variations, temperature variations, rise and fall times, and / or thermal noise. Such relativity between elements ranges from a difference of a few percent to size differences.
Claims
[1] Method comprising: obtaining (100) a plurality of groups of encoded data segments for storage in a working memory (22) of the DSN from a data processing unit (12, 16) of a distributed storage network, DSN, (10), wherein the working memory (22) of the DSN includes a plurality of pools of storage units (36) positioned across a geographical area, wherein a data object is divided into a plurality of data segments, and wherein the plurality of data segments is a distributed storage error encoded into the plurality of groups of encoded data segments; identifying (102) a desired group of storage units (36) within the plurality of pools of storage units for storing the plurality of groups of encoded data segments by the data processing unit (12, 16); generating (104) by the data processing unit (12, 16) a particular source name based on the desired group of storage units (36) and a distributed agreement protocol, DAP, (80), wherein the DAP (80) identifies a group of storage units (36) from the plurality of pools of storage units (36) based on a slice identifier and a plurality of storage allocation coefficients, and wherein, when a unit in the DSN (10) executes the DAP (80), the unit uses the particular source name as the slice identifier to identify the desired group of storage units (36); generating (106) a plurality of groups of clipping names (78) for the plurality of groups of encoded data clippings by the data processing unit (12, 16), the plurality of groups of clipping names (78) containing the determined source name; and sending (108) a plurality of groups of write requests by the data processing unit (12, 16) to the desired group of storage units (36) with respect to the plurality of groups of encoded data slices and according to the plurality of groups of slice names (78). [2] The method of claim 1, wherein identifying (102) the desired group of storage units (36) comprises one or more of: identifying a group of storage units (36) within the plurality of pools of storage units having a desired write speed relative to the data processing unit (12,16); identifying the group of storage units (36) within the plurality of pools of storage units having a desired efficiency with respect to the data processing unit (12, 16); and identifying the group of storage units (36) within the plurality of pools of storage units having a desired reliability with respect to the data processing unit (12,16). [3] The method of claim 2, wherein identifying (102) the desired group of storage units (36) comprises one or more of: performing a search; initiating a query; receiving a query response; accessing a historical data set; accessing a table; and receiving a list of a desired group of storage units (36). [4] The method of claim 1, wherein generating (104) the particular source name comprises: generating a specific data object identifier; and combining the particular data object identifier with one or more of a vault identifier and revision-level information to produce the particular source name. [5] The method of claim 1, wherein generating (106) a clipping name (78) of the plurality of groups of clipping names (78) comprises: combining the determined source name with a pillar number and a data segment number to generate the cutout name (78). [6] The method of claim 1, wherein the DAP (80) comprises: a plurality of functions operable to generate one or more unique scores based on a plurality of memory pool coefficients and one or more slice identifiers corresponding to the plurality of groups of encoded data slices; and a ranking function (84) that processes the one or more unique scores to identify a selected memory pool (86) for storing the plurality of groups of encoded data snippets. [7] The method of claim 1, further comprising: determining (110) by at least some storage units (36) of the plurality of pools of storage units (36) a DAP asymmetry originating from data processing units (12, 16) that create source names having particular data object identifiers instead of randomly generating data object identifiers; adjusting (112) by the at least some memory units (36) one or more memory allocation coefficients of the plurality of memory allocation coefficients to generate an adjusted plurality of memory allocation coefficients; executing (114) the DAP (80) by the at least some storage units (36) using the determined source name and the adjusted plurality of storage allocation coefficients to identify one or more encoded data chunks from the plurality of groups of encoded data chunks to be transferred to one or more other storage units within the plurality of pools of storage units; and transferring (116) the one or more encoded data portions to the one or more other storage units (36) by one or more storage units of the desired group of storage units. [8] Computer-readable memory device comprising: a first memory element storing operational instructions that, when executed by a computing device (12, 16) of a distributed storage network, DSN, (10), cause the computing device (12, 16) to: Obtaining (100) a plurality of groups of encoded data segments for storage in a working memory (22) of the DSN, wherein the working memory of the DSN (22) includes a plurality of pools of storage units (36) positioned across a geographical area, wherein a data object is divided into a plurality of data segments, and wherein the plurality of data segments is a scattered memory error encoded into the plurality of groups of encoded data segments; identifying (102) a desired group of storage units (36) within the plurality of pools of storage units for storing the plurality of groups of encoded data snippets; a second memory element storing operational instructions that, when executed by the data processing unit (12, 16), cause the data processing unit (12, 16) to: Generating (104) a particular source name based on the desired group of storage units (36) and a distributed agreement protocol, DAP (80), wherein the DAP (80) identifies a group of storage units from the plurality of pools of storage units based on a slice identifier and a plurality of storage allocation coefficients, and wherein, when a unit in the DSN (10) executes the DAP (80), the unit uses the particular source name as the slice identifier, to identify the desired group of storage units; and generating (106) a plurality of groups of snippet names (78) for the plurality of groups of encoded data snippets, the plurality of groups of snippet names (78) including the determined source name; and a third memory element storing operational instructions that, when executed by the data processing unit (12, 16), cause the data processing unit (12, 16) to: Sending (108) a plurality of groups of write requests to the desired group of storage units (36) with respect to the plurality of groups of encoded data slices and according to the plurality of groups of slice names (78). [9] The computer-readable memory device of claim 8, wherein the first memory element further stores operational instructions that, when executed by the computing device (12, 16), cause the computing device (12, 16) to identify (102) the desired group of memory devices (36) by one or more of: identifying a group of storage units (36) within the plurality of pools of storage units having a desired write speed relative to the data processing unit (12,16); identifying the group of storage units (36) within the plurality of pools of storage units having a desired efficiency with respect to the data processing unit (12, 16); and identifying the group of storage units (36) within the plurality of pools of storage units having a desired reliability with respect to the data processing unit (12,16). [10] The computer-readable memory device of claim 9, wherein the first memory element further stores operational instructions that, when executed by the computing device (12, 16), cause the computing device (12, 16) to identify (102) the desired group of memory devices (36) by one or more of: performing a search; initiating a query; receiving a query response; accessing a historical data set; accessing a table; and receiving a list of a desired group of storage units (36). [11] The computer-readable memory device of claim 8, wherein the second memory element further stores operational instructions that, when executed by the computing device (12, 16), cause the computing device (12, 16) to generate (104) the particular source name by: generating a specific data object identifier; and combining the particular data object identifier with one or more of a vault identifier and revision-level information to produce the particular source name. [12] The computer-readable memory device of claim 8, wherein the second memory element further stores operational instructions that, when executed by the computing device (12, 16), cause the computing device (12, 16) to generate (106) a clipping name (78) of the plurality of groups of clipping names (78) by: combining the determined source name with a pillar number and a data segment number to generate the cutout name (78). [13] The computer-readable memory device of claim 8, wherein the second memory element stores operational instructions that, when executed by the computing device (12, 16), cause the computing device (12, 16) to use the DAP (80) by performing: a plurality of functions operable to generate one or more unique scores based on a plurality of memory pool coefficients and one or more slice identifiers corresponding to the plurality of groups of encoded data slices; and a ranking function (84) that processes the one or more unique scores to identify a selected memory pool (86) for storing the plurality of groups of encoded data snippets. [14] A computer-readable memory device according to claim 8, further comprising: a fourth memory element storing operational instructions that, when respectively executed by at least some memory units (36) of the plurality of pools of memory units (36), cause the at least some memory units to: Determining (110) a DAP asymmetry originating from data processing units (12,16) that create source names that have specific data object identifiers instead of randomly generating data object identifiers; adjusting (112) one or more memory allocation coefficients of the plurality of memory allocation coefficients to generate an adjusted plurality of memory allocation coefficients; a fifth memory element storing operational instructions that, when executed by the at least some memory units (36), cause the at least some memory units to: Executing (114) the DAP (80) using the determined source name and the adjusted plurality of memory allocation coefficients to identify one or more encoded data slices from the plurality of groups of encoded data slices to be transferred to one or more other memory units (36) within the plurality of pools of memory units; and Transferring (116) the one or more encoded data portions to the one or more other storage units (36) of the at least some storage units by one or more storage units (36) of the desired group of storage units. [15] Data processing unit (12,16) of a distributed storage network (DSN), (10), wherein the data processing unit (12,16) comprises: an interface (30,32,33): a working memory (54); and a processing module (50) operatively connected to the memory (54) and the interface (30, 32, 33), the processing module (50) being operable to: Obtaining (100) a plurality of groups of encoded data segments for storage in a working memory (22) of the DSN, wherein the working memory (22) of the DSN includes a plurality of pools of storage units (36) positioned across a geographical area, wherein a data object is divided into a plurality of data segments, and wherein the plurality of data segments is a scattered memory error encoded into the plurality of groups of encoded data segments; Identifying (102) a desired group of storage units (36) within the plurality of pools of storage units (36) for storing the plurality of groups of encoded data snippets; Generating (104) a particular source name based on the desired group of storage units (36) and a distributed agreement protocol, DAP, (80), wherein the DAP (80) identifies a group of storage units (36) from the plurality of pools of storage units (36) based on a slice identifier and a plurality of storage allocation coefficients, and wherein, when a unit in the DSN (10) executes the DAP (80), the unit uses the particular source name as the slice identifier to identify the desired group of storage units (36); generating (106) a plurality of groups of snippet names (78) for the plurality of groups of encoded data snippets, the plurality of groups of snippet names (78) including the determined source name; and Sending (108) a plurality of groups of write requests to the desired group of storage units (36) with respect to the plurality of groups of encoded data slices and according to the plurality of groups of slice names (78). [16] The data processing unit (12, 16) of claim 15, wherein the processing module (50) is further operable to identify (102) the desired group of storage units (36) by one or more of: identifying a group of storage units (36) within the plurality of pools of storage units having a desired write speed relative to the data processing unit (12,16); identifying the group of storage units (36) within the plurality of pools of storage units having a desired efficiency with respect to the data processing unit (12, 16); and identifying the group of storage units (36) within the plurality of pools of storage units having a desired reliability with respect to the data processing unit (12,16). [17] The data processing unit (12, 16) of claim 16, wherein the processing module (50) is further operable to identify (102) the desired group of storage units (36) by one or more of: performing a search; initiating a query; receiving a query response; accessing a historical data set; accessing a table; and receiving a list of a desired group of storage units (36). [18] The data processing unit (12, 16) of claim 15, wherein the processing module (50) is further operable to generate (104) the determined source name by: generating a specific data object identifier; and combining the particular data object identifier with one or more of a vault identifier and revision-level information to produce the particular source name. [19] The data processing unit (12, 16) of claim 15, wherein the processing module (50) is further operable to generate (106) a clipping name (78) of the plurality of groups of clipping names (78) by: combining the determined source name with a pillar number and a data segment number to generate the cutout name (78). [20] The data processing unit (12, 16) of claim 15, wherein the processing module (50) is further operable to use the DAP (80) by performing: a plurality of functions operable to generate one or more unique scores based on a plurality of memory pool coefficients and one or more slice identifiers corresponding to the plurality of groups of encoded data slices; and a ranking function (84) that processes the one or more unique scores to identify a selected memory pool (86) for storing the plurality of groups of encoded data snippets. [21] Data processing unit (12,16) according to claim 15, further operable to: determining (110) by at least some storage units (36) of the plurality of pools of storage units from a DAP asymmetry originating from data processing units (12, 16) that create source names having particular data object identifiers instead of randomly generating data object identifiers; adjusting (112) by the at least some memory units (36) one or more memory allocation coefficients of the plurality of memory allocation coefficients to generate an adjusted plurality of memory allocation coefficients; Executing (114) the DAP (80) by the at least some storage units (36) using the determined source name and the adjusted plurality of storage allocation coefficients to identify one or more encoded data chunks from the plurality of groups of encoded data chunks to be transferred to one or more other storage units within the plurality of pools of storage units; and Transferring (116) the one or more encoded data portions to the one or more other storage units (36) by one or more storage units of the desired group of storage units.
Citation Information
Patent Citations
System and method for accessing a data object stored in a distributed storage network
US20100287200A1
Migrating encoded data slices in a dispersed storage network
US20150378625A1