Encoding for data recovery in a storage system - Patents.com

The method addresses the challenge of sector error recovery in digital data storage systems by using redundant data and linear combinations of information payloads to recover missing sectors, achieving reliable and efficient data recovery.

JP7671759B2Active Publication Date: 2025-05-02MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022537460
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-09
Filing Date
2020-12-14
Publication Date
2025-05-02
Estimated Expiration
2040-12-14

AI Technical Summary

Technical Problem

Existing digital data storage systems face challenges in efficiently recovering from sector errors, especially in large-scale systems where redundancy methods are not effectively utilized across multiple sectors or media fragments.

Method used

A computer-implemented method that uses redundant data stored on a storage medium, where each information sector is part of a group with multiple redundant codes. These redundant codes are linear sums of information payloads from different sectors, weighted by coefficients. The method determines a square matrix E and its inverse D to recover missing information sectors by decoding the redundant codes.

Benefits of technology

The method provides reliable data recovery for applications like archive storage, reduces memory overhead, and ensures quantifiable reliability assurance by effectively utilizing redundancy across multiple sectors and media fragments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007671759000003
    Figure 0007671759000003
  • Figure 0007671759000004
    Figure 0007671759000004
  • Figure 0007671759000005
    Figure 0007671759000005
Patent Text Reader

Abstract

A method for reading from a storage medium to recover a group of information sectors, each containing a respective information payload. The medium stores redundant data including a plurality of distinct redundancy codes for the group, each code being a linear sum of terms, each term being an information payload from a different one of the information sectors in the group weighted by a respective coefficient of a set of coefficients for the redundancy code. After the redundant data has been stored on the medium, the method includes: identifying a set of k' information sectors to be recovered, selecting k' of the redundancy codes, determining a square matrix E of the k' information sectors by the k' sets of coefficients of the selected codes, determining a matrix D that is the inverse of E, and recovering the k' information payloads from the inverse matrix D.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background

[0001] Most large scale practical digital data storage media systems use redundancy to help correct from common types of errors. These can be within individual sectors of data, across multiple sectors of data within a single medium, or across multiple pieces of media. For example, within a single sector, it is common for hard drives to use LDPC (low density parity check) as an error correction method for small bit errors. Across media, it is common to use RAID (redundant array of independent disks) to protect against individual disk failures.

[0002]

[0002] Methods and systems for recovering from sector errors within a single medium tend to be less frequently used and more primitive. For example, on disk drives, it is common to reserve a small portion of the medium for spares and use one of these locations as an alternative if a pre-allocated area of ​​a particular sector is deemed to be failing. This can be problematic if the spare area is too large or too small because it must be selected in advance, and it also introduces unpredictability in access latency. In another example, on tape-based systems, it is common for read-after-writes to check whether the write was successful, and if not, an additional copy of the sector is written. This can be problematic because it creates unpredictable capacity in the tape.

[0003]

[0003] More complex and efficient schemes are impractical in typical storage systems because they must deal with cases where data is rewritten or where the total data to be written within the entire medium is not simultaneously available. Moreover, they are usually made unnecessary by the need for protection across multiple pieces of media, due to, for example, the relatively high failure rates of disk drives and tapes.

[0004]

[0004] A WORM (write-once-read-many) storage system is a storage format in which all data is written in one operation at once. Classical optical media such as CDs and DVDs are both WORM and write all data in one operation, but because of the need to provide data at a fixed rate when read, for example to an audio or video playback device, and / or the desire to maintain very low-cost playback devices in consumer scenarios, they do not tend to use complex, media-wide redundancy systems.

[0005]

[0005] Another known type of optical WORM storage uses quartz glass as a storage medium. Information is inscribed on the structure with the help of an ultrafast laser (usually a femtosecond laser). Such a laser has the ability to direct a large amount of energy into a very confined space, changing the structure of the glass in that area in a controlled and permanent way, and thus storing information there. Some such systems allow the storage of data across three dimensions of the medium, in which case the location of a given bit or symbol can be called a voxel. Reading then works by using polarization-sensitive microscopy to shine light into a particular part of the glass and infer the data written in that area by measuring certain properties of the light observed. Summary of the Invention [Means for solving the problem]

[0006] overview

[0006] According to one aspect disclosed herein, a computer-implemented method is provided for reading from a storage medium to recover a group of information sectors, each information sector including a respective information payload. The storage medium stores redundancy data including a plurality of distinct redundancy codes for the group, each redundancy code being a linear sum of terms, each term in the sum being an information payload from a different one of the information sectors in the group weighted by a respective coefficient of a set of coefficients for the redundancy code. The method includes the steps of: identifying a set of k' information sectors from the group, whose respective information payloads are to be recovered based on the redundancy data after the redundancy data has already been stored on the storage medium; selecting k' of the redundancy codes; determining a square matrix E, each matrix column including a respective coefficient of a different one of the k' information sectors and each matrix row including a set of coefficients of a different one of the k' redundancy codes, or vice versa; determining a matrix D that is the inverse of E; and determining v i =Σ j (d i,j ·r j ), where v i is the information payload, i is an index indicating each information sector, and j is the length of each redundant code r j is the index that indicates i,j is a matrix element of D, and Σ j This involves performing a decoding process that includes recovering, where k' is the sum over the k' redundancy codes, a calculation being performed for each i of the k' information sectors.

[0007]

[0007] This Summary is provided to convey in a simplified form a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Nor is the claimed subject matter limited to implementations that solve any or all of the disadvantages discussed herein.

[0008] BRIEF DESCRIPTION OF THE DRAWINGS

[0008] To assist in understanding embodiments of the present disclosure and to show how such embodiments may be put into effect, reference is made by way of example only to the accompanying drawings, in which: [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a schematic diagram of a data storage scheme. [Diagram 2]

[0010] FIG. 2 is a flow chart of a method for storing data on a medium that includes redundant data. [Diagram 3]

[0011] FIG. 3 is a flow chart of a method for reading data from a medium including recovery based on redundant data. [Figure 4]

[0012] FIG. 4 illustrates generally a scheme that uses both rows and columns of redundant data. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Detailed Description of the Embodiments

[0013] The following describes network storage and recovery methods that have particular, but not limiting, applicability to WORM (write once read many) storage systems, such as fused silica, where all data is written simultaneously in one operation and the media is extremely durable. This enables superior in-media redundancy error recovery methods and systems. Although embodiments may be described with respect to fused silica media, the methods of the present disclosure may also be applied to any WORM media where all data to be written is available to be processed as it is written.

[0011]

[0014] Even more generally, the disclosed method may be used with any media, including more conventional optical disks, magnetic media, or electronic storage media. In practice, however, they tend to be used for real-time or on-the-fly reading and writing of small portions of data. For example, classical optical media such as CDs and DVDs are both WORM, writing all data in one operation, but they do not tend to use complex, media-wide redundancy systems due to the need to provide data at a fixed rate when read, for example to an audio or video playback device, and / or the desire to maintain very low-cost playback devices in consumer scenarios. The disclosed method has particular, but not limited, applicability to long-term storage systems, such as those in archival storage systems, where it is acceptable to simultaneously write relatively large amounts of data and, if recovery is required, to simultaneously read relatively large amounts of data. The disclosed method could be used, for example, in quartz glass data repositories and / or in cloud-based archival systems.

[0012]

[0015] The encoding method divides the data to be stored into sectors, which may be called content sectors or information sectors. The information sectors are used to build random linear combinations of the information in those sectors. These combinations may be called recovery sectors or redundant sectors. Both the information sectors and the redundant sectors may be stored on the storage medium. Alternatively, the redundant information could be communicated to the reader via a separate medium. In either case, during the reading process, a method is applied to determine which sectors have been read correctly and which are missing. The correctly read redundant sectors are then used to recover the missing information sectors by inverting the random linear combinations. This process is similar to inverting a system of linear equations.

[0013]

[0016] As an optional optimization, the information sectors can be divided into groups and the redundancy encoding operation can be applied to each group independently. This approach reduces the computational overhead of the above encoding while minimizing the degradation of error correction performance. Alternatively, all sectors on the entire medium can be encoded as one group.

[0014]

[0017] The set of coefficients used for the linear combination can be predefined. As another optional extension that increases the flexibility of the code, the coefficients can be stored together with the encoded sector. However, this would have a relatively large storage cost. Therefore, instead, to reduce the overhead, the coefficients can be generated from a deterministic pseudorandom process. Therefore, it is sufficient to store the seed that initializes the random process and the exponent of the random process that created the coefficients.

[0015]

[0018] Encoding is a common technique for providing reliability in both storage and networking applications. Codes differ in the way they combine original content to construct the encoded information. The differences are aimed at trading off reliability performance, coding overhead, computational complexity of encoding and decoding, and other factors. The method of the present disclosure is similar to the network coding technique disclosed in U.S. Pat. No. 7,756,051 for encoding data to be transmitted over a network for the purpose of content distribution, as opposed to storage. However, existing network coding techniques are "unsystematic." That is, the payload data itself is encoded by the same encoding scheme that is adding redundancy, and decoding is always required as a matter of course to extract the payload information. In contrast, in the method of the present disclosure, the encoding is "systematic." That is, the payload remains unencoded by applying the encoding method to add redundancy (even if the payload happens to be encoded by some other orthogonal lower layer encoding scheme, such as for compression and / or encryption purposes, rather than redundancy). Instead, the redundant data is separate to the payload information, e.g., appended in separate redundant sectors or even communicated over a separate medium. In the systematic case, the corresponding decoding method is only needed if an error needs to be recovered in the payload information, and not necessarily at every read.

[0016]

[0019] Compared to other coding techniques, network coding offers very good reliability performance for a wide range of failure scenarios (i.e., it can be decoded with high probability with modest storage overhead) and also allows the construction of an arbitrary number of encoded sectors. To provide these benefits, it may sacrifice computational performance. This disclosure extends the scope of existing network-coded encoding schemes to include storage applications.

[0017]

[0020] Therefore, the disclosed approach to constructing error-correcting codes can provide reliable data recovery for applications such as archival storage and provide quantifiable reliability guarantees. In embodiments, it also reduces storage overhead compared to existing systems.

[0018]

[0021] In embodiments, nearby storage elements (each storing an individual bit of data or other such basic symbol) may be grouped together to define a sector. Sectors typically store tens of kilobytes of data and are usually the same size. As with all storage systems, due to imperfections in writing and reading, elements within a sector may be written or read with errors. To address that problem, redundancy information is stored within each sector that allows for detection of errors within the sector and, typically, correction of those errors. In addition, an integrity check in the form of a hash or checksum may be stored to determine whether the correction was successful. There are many known algorithms for determining integrity. A portion of the available bytes within a sector may be reserved to store integrity check information. After reading all elements within a sector and checking the integrity, it is then possible to determine whether the sector was read correctly as a whole (with a very high probability) or whether errors exist within the sector. The decoding method of the present disclosure may then recover the sector containing the error.

[0019]

[0022] The approach involves generating (and in embodiments storing within the media) specially crafted codes that can be used to recover erroneous sectors. In embodiments, these redundant codes may be stored in their own redundant sectors on the media. Some sectors may therefore be reserved within the media to store special redundant sectors instead of user data. To recover missing user sectors, the redundant sectors are combined with other user sectors that were read correctly.

[0020]

[0023] In one implementation, any redundant sector can be used to recover any user sector. In other words, all user data stored across the medium can be used to build all redundant sectors, so that every redundant sector can help recover every user sector. Although this approach is possible, in practice, it can be computationally impractical. Therefore, as an optimization, user sectors can instead be divided into groups, and the encoding method builds group-specific redundant sectors. In this case, recovery is performed only within a given group, combining user sectors from a single group with redundant sectors from the same group to recover missing user sectors (of that group). The size of those groups (number of users and redundant sectors) is a design parameter that depends on the expected sector read failure rate, the desired reliability target, and other practical concerns.

[0021]

[0024] Several exemplary embodiments of the techniques of this disclosure will now be described in more detail with reference to FIGS.

[0022]

[0025] FIG. 1 shows a scheme for arranging data in sectors on a storage medium. The data is stored by a storage computer. The storage medium may be an integral part of the storage computer or may be an external or removable medium. The storage computer may include one or more computer units (in one or more housings) in one or more locations. The data may be read by a reading computer and errors may be recovered. The reading computer may be the same computer as the storage computer, or a different computer, or may share one or parts of the same units when implemented across multiple computer units. In the case of multiple computer units, suitable distributed computing techniques will be familiar per se to those skilled in the art. The or each computer unit of the storage and reading computer includes one or more processors on which the storage and reading method of the present disclosure is respectively executed. The or each processor may include one or more cores. Examples of such processors include general purpose CPUs (central processing units); as well as dedicated memory processors, or co-processors, accelerator processors, or application-specific processors, such as repurposed GPUs (graphics processors), DSPs (digital signal processors), or cryptoprocessors, etc.

[0023]

[0026] In embodiments, the storage medium is a quartz glass-based storage medium. However, this is not a limitation. More generally, the storage medium can be any WORM (write once read many) medium, or a ROM (read only memory), or even a rewritable (but non-volatile) memory. The difference between ROM and WORM is that ROM is written at the factory, whereas WORM can be written once in the field by the end user (consumer). The storage medium can be an optical medium such as quartz glass or an optical disk (e.g., CD or DVD); a magnetic medium such as a hard disk drive, a floppy disk, or a magnetic tape; or an electronic medium, for example, a hardwired ROM, an EPROM (erasable programmable ROM), an EEPROM (electrically erasable programmable ROM), or a flash memory, or a solid state drive. Other types of storage media may be familiar to those skilled in the art, and most broadly, the methods of the present disclosure are applicable to any type of computer-readable medium for storing data.

[0024]

[0027] In embodiments, the storage medium may include a single contiguous piece of writeable material, such as a single piece of fused silica glass, or a single magnetic disk, etc. Alternatively, the medium may include more than one piece of writeable material, such as multiple pieces of glass, or multiple magnetic disks, etc. In the latter case, the disclosed schemes can provide redundancy across multiple separate pieces of writeable material ("platters") in addition to redundancy of multiple sectors within a platter. This application space may be particularly important in archival storage, for example.

[0025]

[0028] Whatever form the medium takes, the information to be stored is stored in a group S of information sectors. i , i=1...n. In the notation used herein, i denotes each information sector S in the group. iwhere i denotes an index of n sectors of information, and a plurality of n sectors of information are present in the group. For notational purposes, the index i is represented as ranging from 1 to n to indicate that it can take a value indicating any of the n sectors of information. This does not necessarily imply that the actual digital parameter used by the reading or writing computer to index the sectors ranges numerically from 1 to n (e.g., in practice, this could range from 0 to n-1 in binary). It will be understood that this is merely a convenient mathematical representation or notation.

[0026]

[0029] The group of n information sectors could be all the information sectors on a medium (e.g., glass, disk, or tape, etc.) or it could be only one of multiple groups of information sectors on the medium. In the latter case, the following scheme could be applied to only one of the groups, or to each of some or all of the groups independently. The following describes the storage and encoding within a given group.

[0027]

[0030] Each information sector S i Each information sector S i are stored in different respective physical sectors of the storage medium. In embodiments, each of the physical sectors may be the same size as each other. In embodiments, the information sectors stored on the physical sectors may each be the same size as each other. In embodiments, the physical sectors in a group may be adjacent to each other or may be interleaved with sectors of one or more other groups. Each physical sector includes multiple storage elements, each for storing an individual basic symbol (e.g., bit) of data. For example, in a 3D fused silica medium, these would be individual voxels. In embodiments, the storage elements in a given physical sector may form a contiguous series or array of elements, e.g., a contiguous rectangle or cuboid of voxels. However, this is not required and the sectors could simply be logical sectors that are not tied to the underlying physical layout of the physical storage elements.

[0028]

[0031] Each information sector S i is at least the information payload v i , i.e., data content (user data). This is the actual data that the user wants to store. In embodiments, each information payload contains a vector of component information values ​​and can be treated as a vector of component information values. Thus, an information payload v i may also be referred to herein as an information vector.

[0029]

[0032] The reading computer reads each information sector S i Information payload of v i We also need some mechanism for detecting whether each error-detecting code z i is the information payload v i In embodiments, this may be, for example, associated with each of the payloads v i and its respective information payload v i Together with each individual information sector S i However, alternatively, it could in principle be stored elsewhere on the same medium, or even elsewhere on a different medium of the reading computer. Wherever it is stored, the error recovery code z i When the reading computer reads the payload and the error detection code, each information payload v i A small piece of redundant data, such as a parity bit or checksum, that allows detection (but not necessarily correction) of errors in the information payload v i This could include errors that occur when initially writing the data to the media, or errors that occur due to degradation on the media in the time between writing and reading, or read errors that occur during the reading process, or a combination of any two or more of these.

[0030]

[0033] Information Sector S iand error detection code z i In addition, multiple redundant codes r j , j=1...k, are also calculated. These are the errors that occur when reading the information payload v if an error of any kind (writing, degradation or reading) is detected during reading, for example based on an error detection code or even a total failure to read at all. i In embodiments, each information payload includes a vector of constituent elements and can be treated as a vector of constituent elements. Hence, the recovery code r j may also be referred to herein as a redundant vector.

[0031]

[0034] Error-detecting code z i Note that, although redundancy data is also included for the purposes of error recovery, they are not included and, in embodiments, may not enable error recovery or, at best, may only enable error recovery (if used in this way) for information experiencing simpler or more limited errors. j may also be referred to as an error recovery code.

[0032]

[0035] In the notation used herein, j is the information sector S i Redundancy code (recovery code) r for a given group of j , where j denotes the index of each of the k redundant codes for the group. For notational purposes, the index j is represented as ranging from 1 to k to indicate that it can take on values ​​that indicate any of the k redundant codes. This does not necessarily mean that the redundant code r j This does not imply that the actual digital parameter used by a reading or writing computer to index k ranges numerically from 1 to k (e.g., in reality it could range from 0 to k-1 in binary), and it will be understood that this is merely a convenient mathematical representation or notation.

[0033]

[0036] In various embodiments, the redundant code r j The set of Information Sector S i In some such embodiments, each code is stored in a different respective redundant sector R on the same storage medium, which may be a separate physical sector of the storage medium. i In some embodiments, each redundant code r j is each information vector v i or may be the same size as the information sector. In various embodiments, each redundant sector R i Each information sector S i Alternatively, each redundant code r j The size of the information vector v i Or Information Sector S i , and / or there may be more than one code stored per redundant sector R. In a further alternative, the redundant codes are not stored in separate sectors, but rather in the information sector S. i It is not excluded that the assets could be distributed among

[0034]

[0037] Information Sector S i For a given group of , there must necessarily be as many redundant codes r as there are information sectors in the group. j (k is not necessarily equal to n), and there exists a redundant code r for a given group. j and Information Sector S i Note that there is no one-to-one correspondence between r and r. Rather, as will be explained in more detail shortly, j is the number of n information sectors in the group S i At least the information payload v i In general, the total number of redundant codes k may be less than or equal to the number of data items n, or in embodiments, the number of redundant codes k may be greater than the number of data items n.

[0035]

[0038] In addition, the redundant code r j The coefficient c used to calculatej,i (see below) may be stored on the medium together with the code itself, for example in a redundant sector R. In this case, the reading computer reads the coefficients from the medium and uses them to derive the redundant code r j Alternatively, the coefficients may be determined according to a predetermined deterministic process, such as a pseudorandom process, and only an indication of the process used may be stored in the medium, for example, again in the redundant sector R. In this case, the reading computer reads the indications from the medium and uses them to determine the process that was used to determine the coefficients, which is used to determine the coefficients themselves, and uses these coefficients to calculate the redundant code r j The instructions may, for example, include a seed for the pseudorandom process. Optionally, it may also include an indication of which of a number of available algorithmic forms was used for the given process. Alternatively, the type of algorithm could be assumed by the reading computer.

[0036]

[0039] In yet a further variant, the redundant code r j , or coefficient c j,i Neither the codes, coefficients, and / or any instructions of the process for determining the coefficients need to be written onto the storage medium itself. Instead, any of these could be communicated to the reading computer via a separate medium (e.g., a communication channel). For example, they could be publicly available or could be sent specifically to the reading computer, for example, over a network or on a separate storage medium such as a dongle. If the storage and reading computers are the same computer, communication could simply involve storing the codes, coefficients, and / or instructions locally on the computer in question, for example, on a local hard drive.

[0037]

[0040] FIG. 2 shows a method for storing data in the format described with respect to FIG. 1 and storing a redundant code r jFIG. 1 shows an encoding method that can be applied in a storage computer to determine the sector S. The method is performed by software stored in a memory and executed on at least one processor of the storage computer. The memory can be, for example, a storage computer for determining the sector S. i The present invention may include one or more memory units, which could include any of the types of media described above with respect to the storage medium on which the program is stored, and / or different types, such as RAM (random access memory), and which could be the same as or separate from the storage medium, or a combination thereof.

[0038]

[0041] In step 210, the method comprises determining n information sectors S to be written on a storage medium (e.g., quartz glass). i This step determines the group of i=1...n for at least sector S i Each of the n payloads v i This is the user information (i.e., content) that the method uses to store and protect with redundancy codes. Step 210 also includes determining sector S i Each error-detecting code z i This may include generating a checksum byte or bytes (e.g., one or more checksum bytes). Alternatively, these could be stored elsewhere on the medium, or in principle even on a different medium, although this would make error detection slower.

[0039]

[0042] Next, steps 220-230 generate redundancy codes, which are to be stored, for example, in a redundant sector R. The method converts the information bytes into k redundancy codes or code words r j , j=1...k (e.g. a codeword can consist of 2 bytes). These can be described as random linear codes, or linearly independent codes, for reasons that will be explained shortly.

[0040]

[0043] In step 220, the method selects coefficients for performing the linear combination. i To accommodate the possibility of errors in more than one of the redundant codes r1, r2, ..., r k Therefore, in step 220, the method generates the redundancy code r for j=1...k to be generated. j The coefficients for each of j,i Each set (each with its own code r j for i=1...n information sector S i There are n nonzero coefficients c, one for each j,i These coefficients may be selected randomly, for example. Together, the set includes coefficients c j,i The k-by-n matrix C of 1...k, where j=1...k, and i=1...n.

number

[0041]

[0044] Each row is made up of k redundant codes r j , j=l...k. The k sets must be linearly independent of each other, or the probability that they would be so when chosen randomly must be within some tolerance threshold. The process for selecting these coefficients will be described in more detail shortly.

[0042]

[0045] Each column in the matrix C represents n information sectors S i , i=1...n respectively.

[0043]

[0046] In step 230, the method finds a redundant code r for a given j. j Calculate each of as a linear sum of: r=c1·v1+c2·v2+…+c n ·v n (2) In other words: r1=c 1,1 v1+c 1,2 ·v2+…+c 1,n ·v n r2=c 2,1 v1+c 2,2 ·v2+…+c 2,n ·v n … r k =c k,1 v1+c k,2 ·v2+…+c k,n ·v n (3) Or: r j =c j,1 v1+c j,2 ·v2+…+c j,n ·v n ,j=l...k (3a)

[0044]

[0047] Each redundant code r j is the n information sectors in the group being encoded, S i Each term is a sum of n terms, one for each of the information sectors S i Take the multiplicand from and calculate the respective coefficient c for that information sector. j,i As noted above, in embodiments, the multiplicand of each term is the information payload v i Error-detecting code z i (e.g., checksum bytes) may be ignored in the construction and recovery of the redundancy code, which is used to check for the presence of errors within a sector, but is not included in the cross-sector redundancy code r described herein. j However, in an alternative embodiment, each error detection code z i Sector S i It is not excluded that additional data from could be included in the respective multiplicands of each term.

[0045]

[0048] In embodiments, each information payload v in the groupi contains and can be treated as a vector of information values ​​(elements), e.g., a vector of individual bits or bytes. Similarly, each redundant code r j may contain and be treated as a vector of redundant elements. Therefore, the information payload and the redundant code may be referred to as an information vector and a redundant vector, respectively. However, this is not limiting, and in other variations described below, the information payload v i and / or redundant code r j It will be appreciated that each of could be treated as a single scalar value.

[0046]

[0049] In a preferred embodiment, the information payload v i and redundant code r j Each of is a vector with coefficients c j,i Each of is a scalar.

[0047]

[0050] Observe that the above addition and multiplication operations are performed in a finite field. For example, if the coefficients and codewords each consist of 2 bytes (16 bits), then the appropriate field for the operations is, for example, the Galois field GF(2 16 ). Otherwise the notation follows standard algebraic rules: the constants c1, c2, ..., c n , the corresponding vectors v1, v2, ..., v n are multiplied, and vector addition is element-wise.

[0048]

[0051] Also, again, note that in general, k can be less than, equal to, or greater than n, depending on the implementation.

[0049]

[0052] Typically, multiple information vectors v in a group i Since we would expect errors in r, the method generates multiple redundant vectors, r, r, ..., r for a given group. k As mentioned above, the corresponding coefficients are stored in the redundant vector r j For c j,1 , cj,2 , ..., c j,n Here, j is 1 to k. For all c j,i The value of is nonzero. Furthermore, for a different set of coefficients, c j =[c j,1 ,c j,2 ,...,c j,n ] are linearly independent of each other, where each set is a redundant code r j , i.e., one row of the matrix in equation (1) corresponding to one stage of equation (3). Thus, the redundant code r j=1 The set of coefficients c1=[c 1,1 ,c 1,2 ,...,c 1,n ] is another set of coefficients c2 = [c 2,1 ,c 2,2 , ..., c 2,n ]...c k =[c k,1 ,c k,2 ,...,c k,n ] etc. (Note that the vectors here are different kinds of vectors that encode user or redundant data - where vector c j is the information vector v i=1...n (Denotes a set of coefficients, one for each of

[0050]

[0053] "Linearly independent" means that one set cannot be created from a linear combination of the other sets. That is, for any given set c used for a given group of information sectors (i.e., among the sets of coefficients used in Equation 3), j=a =[c a,1 ,c a,2 ,...,c a,n ], c j=a =(β1·c1)+…+(β a-1 ·c a-1 )+(β a+1 ·c a+1 )+…(β k ·c k ), where c j [cj,i , c j,2 ,..., c j,n . This condition is equivalent to saying that each set of coefficients adds new redundant information to the redundant code. If this condition is not satisfied for one of the sets of coefficients, then the corresponding redundant code generated from the coefficients of that set does not add new redundant information to the code, and therefore, the information payload v i that can be corrected is one less than the existing redundant codes.

[0051]

[0054] Assuming that the linear independence condition is satisfied, at this time, using k redundant vectors, any k missing user sectors S i from the user information v i can be recovered. Similarly, if there are k' < k missing user sectors, any k' redundant vectors can be used to reconstruct the missing user sectors (with the help of the n - k' user sectors that were read correctly without error). The condition that the set of coefficients c j is linearly independent is equivalent to the matrix of Equation (1) having a row rank of k.

[0052]

[0055] To achieve linear independence, the set of coefficients can be selected according to a pseudo-random process. This does not strictly ensure linear independence by itself. However, it would mean that the probability that the set is linearly independent is within some threshold. In some embodiments, the method is such that the set is likely to be linearly independent within an acceptable threshold probability, and if, ultimately, they are not linearly independent, the result is acceptable (i.e., one less incorrect information sector S in the group of n sectors iThis may involve simply selecting a set of coefficients pseudo-randomly and not checking the linear independence condition, assuming that the linear independence condition can be corrected. In other words, simply generate the coefficients randomly and hope for the best. However, alternatively, the set of coefficients may be selected using a selection process that ensures that the set is linearly independent. For example, this may involve selecting them pseudo-randomly and then checking that the selected set is linearly independent, and if not, reselecting one, some, or all of the coefficients using a pseudo-random number generator until the linear independence condition is met. Another possibility is to use a pre-designed set of coefficients.

[0053]

[0056] In step 240, the information sector S i to the storage medium. Redundant sectors R may also be written to the medium, which in embodiments may also contain coefficients, or a seed for the process used to generate the coefficients (see below). Alternatively, some or all of this redundant data could be read separately and communicated to the computer. Also, FIG. 2 is given by way of example only, and other variations of the method may include writing information sectors S before steps 220 or 230. i It will be appreciated that the order of the steps is not important, except to the extent that dependencies exist in the information that is generated.

[0054]

[0057] Referring again to FIG. 1 and equation (3), the redundant code r j Note that the encoding scheme for encoding v is a systematic encoding scheme. That is, it encodes the payload information v i itself is left unconverted, and redundant information r j is the payload information v iIn contrast, in a non-systematic encoding scheme, such as that previously used in network coding as disclosed in U.S. Pat. No. 7,756,051, the encoding mathematically transforms the payload itself. In other words, the redundancy is spread over the entire information sector or packet. In the non-systematic case, as previously used in network coding, everything received at the receiver is encoded and decoding is required to read all the packets. In contrast, in the systematic case used for storage as disclosed herein, decoding is only required for erroneous sectors and the computational overhead scales with the number of lost sectors.

[0055]

[0058] In organizational cases, the information payload v i could be plain user data or could have been transformed by some lower layer of encoding (such as for media compression and / or encryption), but in either case it is not transformed by an encoding method or scheme at the layer that adds redundancy encoding, i.e., generates the redundant codes / vectors above. Also, the redundant codes are not stored as payload within the same or partially coincident physical storage elements of the media (e.g., they are separate voxels).

[0056]

[0059] However, it should be noted that the scope of the present disclosure is not limited to the systematic case. It is not excluded that in alternative embodiments, the non-systematic case could also be used for storage. In the non-systematic case, this means that only redundant codes will be stored on the storage medium, and "raw" information sectors will not be stored on the medium. In this case, as long as at least n codes are read correctly, the information sector will be fully recovered from the codes.

[0057]

[0060] In either the systematic or non-systematic case, the redundancy vector r j=1...kIncreasing the number k of vectors increases the probability of successful recovery of user data (information sectors). However, it also increases the computational cost (to create redundant vectors). The choice of k is a system parameter that can be specified at design time.

[0058]

[0061] Given the number of information sectors per group (n) and the number of redundant vectors (k), the coefficient c j ensure that the sets of c are linearly independent (or within an acceptable probability) i,j It is desirable to choose an appropriate value for

[0059]

[0062] sign r j and coefficient c j,i to a reading computer that will perform the decoding. One way is to define the coefficient values ​​and make them known to a process that generates redundant vectors and a process that restores missing vectors. This approach would require defining n·k coefficients and would effectively require a large amount of space to store those values.

[0060]

[0063] An alternative is to store the coefficients together with the redundancy vector on the storage medium, for example in a redundant sector R. In this approach, some space is saved for the value c j,1 , c j,2 , ..., c k,n To store j , which would require extra overhead per vector.

[0061]

[0064] A third alternative is to design a deterministic process that generates the coefficients and has a short description. One implementation is to use well-known algorithms for generating pseudorandom numbers. In this approach, the designer would define an algorithm and an initial seed that generates the pseudorandom sequence (and thus the coefficients). The designer would preferably check that the generated random numbers satisfy the linear independence assumption, for example by checking that the first n·k codewords define a matrix with the form (1) and that this matrix has row rank k. This check only needs to be done once per seed. In this approach, no extra information needs to be stored in each sector. However, it would be necessary to be able to map a physical location to a particular sector, i.e., to be able to identify the location of each information or redundant sector on the glass. For each read of a sector, the decoding method on the reading computer would need to be able to derive the group to which this sector belongs and whether it is the i-th information sector or the i-th redundant sector. This is possible by defining locations within the medium into groups and sectors.

[0062]

[0065] In a further variation of any of the above, the redundancy codes, coefficients, and / or seeds (or other such indicators of a deterministic process) could be communicated to the reading computer via a separate medium, for example on an attached dongle or over a network communications channel.

[0063]

[0066] A further point to note is the error detection code (e.g., checksum) that is added to each sector. i In various embodiments, a sector stores information and redundant code words, as well as additional information z that can be used to detect errors in the sector. iIn such embodiments, these checksums are preferably not used to create the redundancy code. Instead, in some such embodiments, the bytes of the redundant sector R may be calculated first and then used to calculate the checksum for the corresponding redundant sector.

[0064]

[0067] Figure 3 shows the redundant code r j FIG. 1 shows a decoding method that can be applied in a storage computer for reading data from a medium as described with reference to FIG. 1, including recovering erroneous sectors using the method. The method is performed by software stored in a memory and executed on at least one processor of a reading computer. The memory can, for example, store information sectors S i The storage medium on which the program code is stored may include any of the types of media described above, and / or different types, such as RAM (random access memory), and may include one or more memory units, which may be the same as or separate from the storage medium, or a combination thereof.

[0065]

[0068] In step 310, the method begins with the process of reading a sector from the storage medium (e.g., glass). i After reading (or actually R), the method reads each error-detecting code z i (e.g., a checksum) to determine if the sector is error-free. The process also knows which group the sector belongs to, and the location of the sector within the group (whether it is the ith information or a redundant sector). Alternatively, or in addition, errors could be detected upon complete failure to read the sector. i In the case of error detection using z, each bit or symbol in the sector is successfully read with a correct intended value, but each error detection code z iBased on the redundant information within, it is detected that at least one of those values is incorrect. In contrast, in the case of a read failure, a reliable value cannot be read from at least one of the bits or symbols within the sector.

[0066]

[0069] In either case, after the read, the method has k' < k information sectors S (for group i = 1...n) i that were not read correctly, and has at least succeeded in reading k' redundant codes (therefore, n - k' information sectors S i were read correctly). Next, the recovery process operates as follows.

[0067]

[0070] In step 320, the method subtracts the correct information sectors from each of the redundant codes. Assume that the information vector v j was read correctly for the n - k' values of i. This means that the method updates each redundant vector r j as follows: r j ← r j - c j,i · v i (4)

[0068]

[0071] Here, j now represents the k' redundant codes r used in the recovery jFor convenience of notation, this may be denoted as j=1...k'. It will be understood that this is again merely a convenient mathematical notation and does not limit the form taken by the actual digital parameters used to refer to the code on the reading computer. Also, strictly speaking, this is not necessarily the same sequence of j values ​​used to enumerate the k redundant codes during encoding. As a matter of notation, the new exponent could alternatively be labeled j', but in the following the simpler notation of j is adopted. In any case, the notation is not intended to imply that the k' redundant codes used in decoding are the first (lowest) k' exponent codes in the sequence of k redundant codes as indexed during encoding (it can be any k' of them, not necessarily the first k' in the sequence as indexed during encoding).

[0069]

[0072] From this point onwards, j denotes the redundancy code updated according to step 320, and j denotes the k′ updated redundancy codes r used in the recovery. j It refers to an index that indexes between.

[0070]

[0073] In step 330, the method comprises: i 3. Update the matrix in equation (1) by removing columns corresponding to k' by k' matrices (determined to have been read correctly in step 310). Doing so transforms matrix (1) into a k' by k' matrix. This reduced matrix may now be labeled E. In the reduction process, the correspondence of the original block's exponents to the updated exponents is also stored.

[0071]

[0074] In step 340, the method further comprises: -1 Invert this reduced matrix to generate the inverse matrix E -1 E.E. -1 = 1.

[0072]

[0075] In step 350, the method recovers the missing information sectors by performing the following operations:

number

[0073]

[0076] Here, this is the i-th erroneous sector S i The sum is the sum over all j of the k′ redundant codes used in the recovery for the information sectors S that were determined to be erroneous, incorrect, or not successfully read in step 310. i is carried out individually for each i in

[0074]

[0077] Finally, v i are remapped to the appropriate missing sector using the map from step 330 above.

[0075]

[0078] The redundant code r used in the above decoding method j and coefficient c j,i The value of is determined by any of the means described above: i The read data may be included on the storage medium itself (e.g., on one or more redundant sectors R) along with the read data, or may be communicated via a separate medium, or may be communicated to a process on a reading computer by any combination of these approaches.

[0076]

[0079] It should be noted that the above has been described for the systematic case where the n information sectors themselves are stored on the storage medium and the recovery is performed only for the minimum necessary to recover the lost or erroneous information sectors (i.e. the number of recovery codes k' used in the recovery is equal to the number of information sectors to be recovered). The codes that are present and are not erroneous are simply read directly from the medium. However, in principle the method could be used for any k' x k' matrix, where k' is the number of redundancy codes used in the decoding and is also equal to the number of information sectors desired to be read by the decoding method. In the ultimate conclusion of this, the method can even be applied to the non-systematic case where only redundancy codes (no information sectors) are stored on the medium and all desired information sectors are recovered from the codes. In this case there is no step 320 (equation 4) of updating the codes based on correctly read information sectors and the information vector v i is fully recovered from k' redundant symbols, where in this case k' is not the number of missing or incorrect symbols, but simply the number of information vectors to be recovered, and is the r j is simply the code for j=...k' used in decoding.

[0077]

[0080] There are two parameters to consider in the encoding scheme that will affect decoding.

[0078]

[0081] One is the ratio of k (the number of redundant codes provided) to n (the number of information sectors to be encoded), which will affect the possibility of constructing a valid encoding if the coefficients are chosen pseudo-randomly without actively ensuring that they are linearly independent.

[0079]

[0082] In any computer, a given value must be represented within a finite field (also called a Galois field) of some size L, and in various embodiments, addition is performed in a wrap-around (modulo) fashion within that field. For example, if the size L of the field is 8 bits, then adding 1 after 255 returns to the lowest value of 0.

[0080]

[0083] Let n be the number of information blocks. When a new code is constructed, a random number c i is selected to be combined with v i (see also Equation 2). The number c i can be any number within the field excluding 0. Thus, there are L = 2 i - 1 ways to choose c. There are L 16 ways to choose the n values of c. i n

[0081]

[0084] If k (< n) encodings have already been selected and exist, the following gives the probability that a randomly generated encoding (as described above) depends on the existing encodings (i.e., the new encoding is not good because it does not have the property of linear independence).

[0082]

[0085] The existing encoding can generate (i.e., span) (L + 1) k encodings. This is because the existing encodings r1, r2,..., r k can be linearly combined as d1·r1 +... + d k ·r k ·r k for random coefficients d1,..., d i from the field. Since the value of d k can take the value 0, the number of combinations of linear encodings is (L + 1).

[0083]

[0086] Therefore, the probability that a randomly generated coefficient vector (i.e., a code) is one of those spanned by an existing encoding is (L+1) k / L n ≒L k-n =1 / (2 16 -1) n-k As k approaches n, the probability of a bad selection increases. However, in embodiments, k may be less than 10%-20% of n, and n may be in the thousands. Therefore, the probability of a bad selection is very small.

[0084]

[0087] Note that unlike other redundancy encoding and decoding schemes used for storage, n is not limited by the field size L in our scheme. In other existing redundancy schemes, n+k is limited to be smaller than L. However, in our scheme, n+k can be larger than L. The number of encodings that can be generated is limited by n. Random construction of codes works better for larger fields (see the denominator in the above equation), but the effect of n is still more significant.

[0085]

[0088] Another parameter to consider is the group size. In particular, there are benefits to using a larger group size (larger than n) as it improves the chances of recovery.

[0086]

[0089] In a group having n information sectors and k associated redundancy codes, there will be a maximum number of errors that can be tolerated (at least n of the n information sectors and k codes must be successfully read). For example, if there are 8 information sectors and 2 associated redundancy codes, then the system will be able to tolerate a maximum of 2 errors in 10 sectors including 8 information sectors and 2 redundant sectors and still recover the entire group. That is, the total number of successfully read information sectors and codes is at least the number n of information sectors originally written on the medium in the group.

[0087]

[0090] However, errors in the information sectors stored on the storage medium are random. With a finite group size, there is always a chance that one will be unlucky enough to have more erroneous sectors than the number for which the redundancy code is designed. With a small group size, this chance can be quite large. For example, in the above example, statistical variations in the number and distribution of errors can easily lead to a given group having three or more erroneous information sectors by chance, and therefore being unable to recover. The larger the group size, the smaller the chance of being unlucky than planned. That is, as n approaches infinity, the probability that a group will be unrecoverable approaches the theoretical statistical value for a given number of associated redundancy codes.

[0088]

[0091] Let us again assume that n is the number of information blocks per group and that k redundancy blocks are generated per group. Let us also assume that p is the probability of reading an (information or encoding) block correctly. On average, we would expect that there will be (1-p)·(n+k) failures. As long as the number of failed blocks is less than k, then (due to the assumption of linear independence in the construction of the encoded blocks) it should be possible to reconstruct the missing blocks. If more than k failures are observed, then the decoder will not be able to recover at least one block of the group.

[0089]

[0092] The probability of failure is P fail ≦exp[-(n+k)·D(n / (n+k)||p)], where D[a||p]=a·log(a / p)+(1-a)·log((1-a) / (1-p)). (This can be derived from the tail bound of the binomial distribution, assuming k / n>1-p. The D(a||p) function is the relative entropy.) Observe that the failure probability drops exponentially with the group size n. Hence, larger n is much better.

[0090]

[0093] But there is also a D(.||.) term. This requires that for a reasonable amount of overhead compared to the expected failure probability (1-p), n must be on the order of a few thousand to ensure a very low failure probability. In memory, it is desirable to have a very low failure probability (P fail is 10 for as large an x ​​as possible. -x Therefore, in embodiments, very large group sizes may be desired.

[0091]

[0094] In some embodiments, the group size could even be an entire sector of information. In some embodiments, the group could span multiple pieces of writable material (e.g., multiple pieces of glass).

[0092]

[0095] A further optional optimization is now described with reference to Figure 4. This provides a smaller group of schemes for faster recovery and / or progressive use of the schemes.

[0093]

[0096] The recovery process outlined above is i (i=1...n) over the entire group. This means that to recover a sector, the method allows recovery of data if it successfully reads any n sectors out of a total of n data sectors and k redundant sectors (or more generally, if it successfully reads any n information vectors and / or redundancy codes out of n information vectors and k redundancy codes). This approach works best when the goal of the read process is to recover the entire contents stored in the storage medium (e.g., glass). This may be acceptable for applications such as archival storage, where reads are only required occasionally and slow recovery times are acceptable. However, in some scenarios, it may be desirable to correctly read only a subset of the data, and do so more quickly.

[0094]

[0097] In order to reduce the effort of recovering a partial set of sectors, the information sectors may be organized in a matrix format as shown by way of example in FIG. 4 (the information sectors are indicated by boxes). Redundant sectors are arranged row by row (redundant sector r x,i,j ), and for each column (redundant sector r y,i,j ) is calculated. For example, r x,1,1 ~r x,1,k’ v 1,1 ~v 1,n can be constructed using the same methodology as before, with input r y,1,1 ~r y,k’,1 For input v 1,1 , v 2,1 , ..., v n,1 Note that in the illustrated example there are m rows and m columns (a square matrix), but more generally there could be a different number of rows than columns.

[0095]

[0098] The use of smaller groups allows for faster recovery. It also allows for the incremental use of the scheme in encoding. For example, one column can be written with its redundancy at a time, and much later, after many columns have been written, row redundancy could also be written to provide additional redundancy protection. The row redundancy could even be chosen based on the error rate actually observed on previous columns.

[0096]

[0099] Information vector v i,j To recover, we can use the redundant vector from either row i or column j. The number of information vectors above is m m = m 2 Observe that the number of redundant vectors is 2 m k'.

[0097]

[0100] Before detailing the process of recovering sectors for such an embodiment, a few observations are made. To compare this approach with the base method in the previous section (Figures 1-3), n=m 2Or consider the case m=√n. The base scheme would require an effort of reading n to n+k sectors to recover the missing sector. The schemes in this section could recover using row or column redundant sectors with an effort on the order of √n+k'. The example described with respect to FIG. 4 uses a square m×m matrix: the number of rows is equal to the number of columns (m). It assumes the same number of redundant sectors for rows and columns (k'). Alternative implementations can use different dimensions for the rows and columns and can also appropriately adjust the number of redundant sectors per row or column depending on the size of the rows or columns as well as other parameters of the system. In the above, a linear coding was used to construct both the row and column redundancy vectors. Other methods could have been used to construct either the row redundancy sectors or the columns (or indeed both). That is, a linear code could be used for both to recover missing blocks from the rows or from the columns, or a linear code could be used to recover missing blocks row-wise and a different redundancy coding scheme to recover blocks column-wise, or vice versa.

[0098]

[0101] For the arrangement of Figure 4, the process of reading a sector is as follows: first try to read it directly. If that succeeds, then stop. The cost of the read is the cost of reading one sector. If that fails, then try to recover from the row (or column). Then the cost per sector read is m+1 to m+k (depending on how many sectors are in error). If that fails too, then try to recover using column redundancy (or row redundancy). This adds an additional cost, which is the reading of m+1 to m+k sectors.

[0099]

[0102] If recovery using row and column redundancy fails, recovery continues in an iterative process. Observe that since recovery using both row and column redundancy has failed, there are therefore multiple sectors that are erroneous. Create a list of missing sectors and select one of those sectors, say v i,j’ and try to recover it using the information in column j' (v i,j’ is the target sector v i,j (Observe that j' is in the same row as sector v i,j’ (where v i,j’ is missing). If that succeeds, then the original sector v i,j Try to recover the missing sector or repeat with another missing sector. i,j’ If it is not possible to recover j' then add all missing sectors from column j' to the list and repeat with another sector.

[0100]

[0103] This process is i,j Either it has read all information sectors v and all redundant sectors r and d and still has enough sectors to allow recovery of v i,j If it is not possible to make progress towards recovering the , it terminates. In the latter case, recovery fails.

[0101]

[0104] We show that the performance of the tabular array is comparable to the baseline method in the previous section in the worst case, i.e., the process 2 Note that this would require reading 2·(√n+k') sectors plus all redundant sectors, but in the general case the erroneous sector should be decoded after reading 2·(√n+k') sectors.

[0102]

[0105] As another optional optimization, in addition to or independent of that described with respect to FIG. 4, information storage from different recovery encoding groups may be physically interleaved with each other on the storage medium (i.e., their actual physical locations are spatially interleaved).

[0103]

[0106] Imperfections in the writing and reading process may cause errors in reading a sector correctly. These errors may occur independently of each other; that is, the failure probability of a sector is the same for all sectors, and a failure in one sector does not change the probability of failure in any other (adjacent) sector. Errors may also be spatially correlated, for example, when an imperfection in reading or writing affects many "neighboring" sectors, i.e., sectors close to each other in physical space. The approach described with respect to Figures 1-4 can deal with both types of errors. However, the decoding performance depends on the number of errors in a group: the decoding cost increases with the number of erroneous sectors. It may therefore be desirable to avoid situations where an error affects many sectors from the same group, and instead, it may be preferable to spread the errors evenly among the various groups.

[0104]

[0107] Correlated errors can be expected to affect sectors that are physically close together in the physical space of the material. To destroy those correlations, sectors from different groups can be interleaved, preferably maximizing the spatial distance of sectors from the same group. The exact layout depends on the physical properties of the media.

[0105]

[0108] Although the example of FIG. 4 has been described with groups arranged in rows and columns, the same principle can be applied more generally with any arrangement of information sectors arranged in partially matching storage partial sets (partially matching groups of information sectors). Redundancy codes are generated for each partial set (group). That is, the redundancy codes can be used to recover information sectors of the same partial set, but cannot be directly used to recover information sectors that are not part of the same partial set. A partial set can have partially matching information sectors, for example, partial set A can include information sectors s2, s3, s7, and s8, and partial set B can include s1, s2, s3, and s6. The decoding process works iteratively by identifying the partial sets that can be recovered, and then recovering more partial sets using (recovered) information sectors from these partial sets.

[0106]

[0109] As an example, the redundant code r A is generated for subset A, and r for subset B. B Similarly, information sectors s2 and s6 are lost (all other s i , r A , and r B Suppose that s1 is received correctly and s2 is received correctly. Since there are two missing sectors from subset B and only one redundant sector for B, it is not possible to recover subset B. However, there is sufficient redundancy to recover s2 using subset A. It is then possible to recover s6 using subset B with the reconstructed s2. Note that recovery will be possible even if s2 and s3 are lost: even if neither subset can be used alone for recovery, r A and r B Simplify (using the same process as in Eq. 4) to depend only on s2 and s3, and A and r B Rewrite {tilde over (s2)} to depend only on s2 and s3, and then observe that we can recover s2 and s3 using those relations.

[0107]

[0110] The motivation behind creating the subsets is to be able to recover missing information sectors using fewer sectors (i.e., locally). This speeds up the decoding stage, for example, by decoding smaller groups and by taking advantage of the placement of sectors from the same subset in nearby locations in the storage medium. (The execution time benefit comes at the cost of losing some coding efficiency.)

[0108]

[0111] The case of a product code (i.e., when we divide the information sector into columns and rows) is a special case of the above scheme in which the subsets defined by columns (rows) do not share any information sectors, and every subset defined by columns has exactly one information sector shared with every subset defined by rows.

[0109]

[0112] To improve reliability, some storage systems place replicas of content in geographically separate sites (data centers), e.g., sites in different continents. Existing storage systems typically store identical copies of both information and redundant sectors in those sites. As a further extension, the above method can be used to store a different set of redundant codes at each site. In other words, a unique redundant code for each site is generated, for example, by creating a coefficient matrix with a unique random seed for each site to generate the coefficients. In case of severe errors at multiple sites, where correctly read sectors in each site are not enough to recover the original content, the redundant codes from multiple sites can be combined, increasing the probability of successful recovery.

[0110]

[0113] This scheme is equivalent to the following: Given s sites, then s k redundant codes are generated, with k redundant codes assigned to each site. The storage overhead per site is the same as in the process described above. However, if combining of redundant codes is allowed across sites, the probability of successful recovery is equivalent to using s k redundant codes.

[0111]

[0114] It will be understood that the above-described embodiments have been described by way of example only.

[0112]

[0115] More generally, according to one aspect disclosed herein, there is provided a computer-implemented method of reading from a storage medium to recover a group of information sectors, each information sector including a respective information payload, the storage medium storing redundancy data including a plurality of distinct redundancy codes for the group, each redundancy code being a linear sum of terms, each term in the sum being an information payload from a different one of the information sectors in the group weighted by a respective coefficient of a set of coefficients for the redundancy code, the method including, after the redundancy data has already been stored on the storage medium: - identifying from said group a set of k' information sectors whose respective information payloads are to be recovered based on redundant data; - selecting k' of the redundant codes; - determining a square matrix E, each of whose columns contains a respective coefficient of a different one of the k' information sectors and each of whose rows contains a set of coefficients of a different one of the k' redundancy codes, or vice versa; - determining a matrix D which is the inverse of E; and -v i =Σ j (d i,j ·r j ), where v i is the information payload, i is an index indicating each information sector, and j is the length of each redundant code r jis the index that indicates i,j is a matrix element of D, and Σ j is the sum over k' redundant codes, and the calculation is performed for each i of the k' information sectors. A computer-implemented method is provided that includes performing a decoding process that includes:

[0113]

[0116] In embodiments, some or all of the information sectors are also stored on the storage medium. This is the systematic case. Alternatively, the method could be used in the non-systematic case, where only the redundancy code is stored on the medium, and not the information sectors, and the information payload is entirely recovered from the redundancy code.

[0114]

[0117] In various embodiments, identifying which of the groups of information sectors is missing or erroneous, does not exist on the storage medium, or contains an information payload v that contains an error. i , which are present and not erroneous, and which information payloads v are found to be present on the storage medium and do not contain errors. i , where the k′ information sectors include missing and / or erroneous information sectors of the group.

[0115]

[0118] In various embodiments, the k' information sectors may be only the missing and / or erroneous sectors of the group. Alternatively, it is not excluded that k' may be larger than strictly necessary, for example, even if there are fewer than k missing and erroneous sectors, it may be up to the same number of codes k stored on the storage medium for the group. However, this would make the recovery more computationally intensive than necessary.

[0116]

[0119] In embodiments, the method comprises, prior to said recovery, performing r over all i of the non-erroneous information sectors in the group found to be stored on the storage medium. j ←rj -(c j,i ·v i updating each of the k' redundant codes by performing j,i is a coefficient corresponding to the i-th information sector and the j-th redundancy code, j is the updated redundant code.

[0117]

[0120] In embodiments, the set of coefficients may not be stored on a storage medium; instead, the method comprises: - reading, from a storage medium, instructions of a predetermined deterministic process for determining the set of coefficients and, based thereon, determining the set of coefficients using said process; or - receiving the set of coefficients via a separate medium, or - receiving, via a separate medium, indications of a predetermined deterministic process for determining the set of coefficients and, based thereon, determining the set of coefficients using said process; may include.

[0118]

[0121] The separate medium could be another digital or computer readable storage medium (e.g., a dongle), or a network (e.g., the Internet). As another alternative, the separate medium could include paper or print media, or even another form of media such as audio. In either case, receiving could include receiving the coefficients in a communication specifically addressed or transmitted to the reading computer, or alternatively, via publication. For example, the coefficients could be published online.

[0119]

[0122] In embodiments, the sets of coefficients for different redundancy codes may be linearly independent with respect to each other.

[0120]

[0123] In embodiments, the method may include an initial stage of storing redundancy code on a storage medium prior to the decoding process.

[0121]

[0124] In embodiments, the initial storage stage may include selecting the coefficients according to a process that ensures that the sets of coefficients for different redundancy codes are linearly independent with respect to each other.

[0122]

[0125] In embodiments, the initial storage stage may include pseudo-randomly selecting the coefficients.

[0123]

[0126] In embodiments, the group of information sectors may be all information sectors on the storage medium.

[0124]

[0127] In an alternative embodiment, the group of information sectors may be one of a plurality of groups of information sectors on a storage medium, and the method may be applied to each of the groups individually.

[0125]

[0128] In some such embodiments, information sectors from different ones of the groups may be physically interleaved on the storage medium.

[0126]

[0129] Said group of information sectors may be one of a first group of information sectors stored on the storage medium and a second group of information sectors stored on the storage medium that partially match the first one, including some, but not all, of the same information sectors, each of the first and second groups being associated with a respective set of redundancy codes. In such an embodiment, the method may include, when it is not possible to first recover the information payload of all information sectors in the first group based on the respective set of redundancy codes associated with the first group, recovering the information payload of the information sectors in the second group based on the respective set of redundancy codes associated with the second group, thus recovering at least one of the payloads of the information sectors that partially match the first group, and then recovering the first group based on the redundancy sectors associated with the first group and at least one recovered information sector in the first group.

[0127]

[0130] The other of the first and second groups of information sectors could be encoded according to the same redundancy scheme or a different redundancy scheme.

[0128]

[0131] In an exemplary application of the various techniques disclosed herein, the storage medium may include a fused silica storage medium.

[0129]

[0132] In embodiments, the storage medium may include multiple separate pieces of writeable material, and the information sectors and / or redundancy code span the multiple pieces of writeable material.

[0130]

[0133] For example, a storage medium could include multiple separate platters, or even multiple storage units contained within separate housings. In embodiments, the separate pieces of material are of the same media type, although it is not excluded that they could instead include different media types, such as glass and magnetic, etc.

[0131]

[0134] Some or all of the information sectors may be duplicated across each piece of material (copies of the same information on each piece), or some of the information codes may be stored only on one piece while other of the information sectors are stored only on another piece. If redundancy codes are stored on the medium, then some or all of the redundancy codes may be duplicated across each piece of material (copies of the same code on each piece), or some of the information codes may be stored on one piece while other of the information sectors are stored on other pieces.

[0132]

[0135] In some embodiments, the material pieces could even be distributed across multiple different data centers at multiple different geographic sites.

[0133]

[0136] Each information sector may further include a respective error detection code. In embodiments, the multiplicand may be only the respective information payload and not the respective error detection code. Alternatively, it is not excluded that the multiplicand includes both the respective information payload and the respective error detection code.

[0134]

[0137] The errors may be errors that occur in the storage of one or more information payloads on the storage medium when written, or errors that occur due to degradation after writing but before being read. The errors may be detected when reading based on the error detection code. Once detected, the errors may be recovered based on the redundancy code.

[0135]

[0138] In embodiments, each of said information payloads is a vector of information elements and each of said coefficients is a scalar, and said multiplication comprises an element-wise multiplication of each information element of the respective information payload with a respective scalar coefficient for the respective information payload.

[0136]

[0139] In embodiments, the sets of coefficients may be selected according to a process that ensures that the sets are linearly independent of one another. Alternatively, the coefficients may be selected according to a pseudo-random process (and thus inherently have some certainty that the sets are linearly independent of one another).

[0137]

[0140] In embodiments, the storage medium may be a glass-based storage medium, such as a fused silica storage medium. Alternatively, the storage medium could be another form of optical storage medium, such as an optical disk, or a magnetic storage medium, such as a magnetic disk or tape, or an electronic storage medium, such as an EEPROM or flash memory.

[0138]

[0141] In embodiments, the storage medium may be a write-once, read-many (WORM) storage medium.

[0139]

[0142] In embodiments, the method may be used for archival storage.

[0140]

[0143] In embodiments, the information payload may include cleartext user data. Alternatively, the information payload may be encoded by a lower layer encoding scheme.

[0141]

[0144] In embodiments, the redundancy code may be stored on the storage medium, for example in one or more redundant sectors separate from the information sectors. Alternatively, the redundancy code may not be stored on the medium.

[0142]

[0145] For example, this could include publishing or communicating a code or instructions of a predetermined process to one or more designated parties via a communications channel separate from the storage medium. The predetermined process could include, for example, a deterministic pseudo-random process, and the instructions could include at least a seed of the pseudo-random process.

[0143]

[0146] The information sectors of a group could be interleaved with those of one or more other groups. Alternatively, the information sectors of said group could be physically adjacent on the storage medium.

[0144]

[0147] The detection of errors may be based on a respective error detection code.

[0145]

[0148] The method may also include directly reading the non-erroneous payload values ​​without requiring recovery.

[0146]

[0149] In embodiments, each of the elements of the inverse matrix D is a scalar and each of the redundant codes is a vector. i =Σ j (d i,j ·r j The product "·" in ') can be element-wise multiplication.

[0147]

[0150] According to another aspect disclosed herein, there is provided a computer program comprising code embodied on a computer readable storage and configured to perform the method of any embodiment disclosed herein when executed on one or more processing devices.

[0148]

[0151] According to another aspect, there is provided a computer system comprising a memory including one or more memory units and a processing device including one or more processing devices, the memory storing code configured to execute on the processing device, the code configured to perform a method according to any embodiment disclosed herein.

[0149]

[0152] Given the disclosure herein, other variations or uses of the techniques of the present disclosure may become apparent to those of ordinary skill in the art. The scope of the present disclosure is not limited by the above-described embodiments, but only by the appended claims.

Claims

1. 1. A computer-implemented method for reading from a storage medium to recover a group of information sectors, each sector containing a respective information payload, comprising: the storage medium stores redundancy data including distinct redundancy codes for the group, each redundancy code being a linear sum of terms, each term in the sum being the information payload from a different one of the information sectors in the group weighted by a respective coefficient for the redundancy code, the coefficients being pseudo-random coefficients determined according to a pseudo-random process; the storage medium does not store the pseudorandom coefficients, but instead stores instructions of a deterministic process for determining the pseudorandom coefficients, the pseudorandom coefficients being determined based on instructions of the deterministic process read from the storage medium; The method further comprising: identifying k′ information sectors from said group; selecting k′ of the redundant codes; updating each of the k' redundancy codes by performing r j ←r j -(c j,i ·v i ) over all i of the non-erroneous information sectors in the group found to be stored on the storage medium, where v i is the information payload, i is an index indicating each of the information sectors, j is an index indicating each of the k' selected redundancy codes r j , and c j,i is the coefficient corresponding to the i th information sector and the j th redundancy code; determining a square matrix E, each of whose columns includes the respective coefficients of a different one of the k' information sectors and each of whose rows includes the coefficients of a different one of the updated k' redundancy codes, or a square matrix E, each of whose rows includes the respective coefficients of a different one of the k' information sectors and each of whose columns includes the coefficients of a different one of the updated k' redundancy codes; determining a matrix D that is the inverse of E; and v i =Σ j (d i,j ・r j ), where d i,j is a matrix element of the matrix D, and Σ j is a sum over the updated k′ ​​redundancy codes, the calculation being performed for each i of the k′ information sectors; 4. A computer-implemented method comprising: performing a decryption process comprising:

2. The method of claim 1 , wherein some or all of the information sectors are also stored on the storage medium.

3. identifying the k′ information sectors; Which of the groups of information sectors is missing or erroneous, does not exist on the storage medium, or contains an erroneous information payload v i , which are present and not erroneous, and which information payloads v found to be present on the storage medium and which do not contain errors. i 3. The method of claim 2, comprising identifying which k' information sectors have the missing and / or erroneous information sectors of the group, wherein the k' information sectors include the missing and / or erroneous information sectors of the group.

4. The method of claim 3 , wherein the k′ information sectors are only the missing and / or erroneous sectors of the group.

5. The method according to any one of claims 1 to 4, wherein the pseudorandom coefficients for different redundancy codes are linearly independent with respect to each other.

6. 6. The method of claim 1, further comprising an initial stage of storing the redundancy codes on the storage medium prior to the decoding process, the initial stage including selecting the pseudorandom coefficients according to a process that ensures that the pseudorandom coefficients for different redundancy codes are linearly independent with respect to each other.

7. 7. The method according to claim 1, further comprising an initial stage of storing said redundancy code on said storage medium prior to said decoding process, said initial stage comprising pseudo-randomly said selection of said pseudo-random coefficients according to a pseudo-random process.

8. The method according to any one of claims 1 to 7, wherein said group of information sectors is all said information sectors on said storage medium.

9. The method according to any one of claims 1 to 8, wherein said group of information sectors is one of a plurality of groups of information sectors on said storage medium, and said method is applied to each of said groups individually.

10. The method according to any one of claims 1 to 9, wherein the information sectors from different ones of the groups are physically interleaved on the storage medium.

11. the group of information sectors being one of a first group of information sectors stored on the storage medium and a second group of information sectors stored on the storage medium that partially coincides with the first group, each of the first and second groups being associated with a respective redundancy code, and the method further comprising: when it is not possible to initially recover the information payloads of all of the information sectors in the first group based on the respective redundancy codes associated with the first group, recovering the information payloads of the information sectors in the second group based on the respective redundancy codes associated with the second group, thus recovering at least one of the information payloads of the information sectors that partially match the first group; thereafter, recovering the first group based on the redundancy code associated with the first group and the at least one recovered information payload of the information sector that partially matches the first group; The method according to any one of claims 1 to 10, comprising:

12. A method according to any preceding claim, wherein the storage medium comprises a plurality of separate pieces of writable material, and wherein information sectors and / or redundancy codes span the plurality of separate pieces of writable material.

13. A computer program embodied on a computer readable storage medium and comprising code configured to perform the method of any one of claims 1 to 12 when executed on a processing device.

Citation Information

Patent Citations

  • Parallel Reed-Solomon RAID (RS-RAID) architecture, devices, and methods

    JP2011504269A

  • Multiple node repair using high rate minimum storage regeneration erasure code

    US20180060169A1

  • Method and communication device for wireless optical communication

    WO2019154065A1