Secret sharing method, secret sharing system, and secret sharing program

By generating redundant fragmented data using the AMSSS method and recovering scattered object data using spare fragments, the problem of low memory efficiency in big data processing in existing technologies is solved, achieving a balance between high efficiency, confidentiality, and availability.

CN121399680APending Publication Date: 2026-01-23LIYING DEZHI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480042647.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-05-30
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing secret sharing methods suffer from reduced efficiency in distributed storage when dealing with large amounts of data, making it difficult to balance confidentiality and availability.

Method used

The AMSSS method is used to generate redundant fragmented data, and spare fragments are used to restore scattered object data during data recovery, ensuring successful reconstruction even when fragments are insufficient.

Benefits of technology

It improves the usability of secret sharing while maintaining high confidentiality, making it suitable for big data backup systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121399680A_ABST
    Figure CN121399680A_ABST
Patent Text Reader

Abstract

And the confidentiality and the availability of the secret sharing method are both realized. Comprises: a dispersion step in which a dispersion data generation unit (14) generates slice data on the basis of data to be dispersed; and a reconstruction step in which the data recovery unit (15) reconstructs the data to be distributed using the fragmented data. In the dispersion step, k slices and one or two spare slices are generated from the data to be dispersed, and are stored in a plurality of storage units (22), respectively. In the reconfiguration step, if the set of slices collected from each storage unit (22) satisfies the number of reconfiguration dispersion, the data to be dispersed is reconfigured using the collected set of slices. On the other hand, if the number of reconfiguration spreads is not satisfied, the spare slice is used when the data to be distributed is reconfigured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a secret sharing method (hereinafter, referred to as a secret sharing method) for secret dispersing secret information of a dispersed object data into a plurality of piece data, reconstructing and recovering the secret sharing method by collecting the piece data, a secret sharing system, and a secret sharing program. BACKGROUND

[0002] A secret sharing scheme (SSS) is widely used as an encoding for protecting important information (secret information) from threats such as destruction, leakage, and the like.

[0003] However, according to Non-Patent Literature 1, in the case where dispersing processing is performed by the secret sharing method, the total data amount of the piece data increases, although the degree differs. Thus, in the case where a large amount of data is processed, the storage efficiency of the dispersed storage can decrease.

[0004] Therefore, Non-Patent Literature 12 proposes an "n multiple (n, n) threshold method of secret dispersing method" (multiple secret sharing: hereinafter, referred to as "MSSS"), and it is shown that it can be practically implemented in Non-Patent Literature 13.

[0005] PRIOR ART DOCUMENTS

[0006] NON-PATENT LITERATURE

[0007] Non-Patent Literature 1: Shamir, Adi, "How to Share a Secret," Commun. ACM, vol. 22, no. 11, pp. 612-613, Nov. 1979 Internet URL <http: / / doi.acm.org / 10.1145 / 359168.359176>

[0008] Non-Patent Literature 2: ISO / IEC, "Information technology-Security techniques-Secret sharing-Part 2: Fundamental mechanisms," ISO / IEC 19592-2, Oct. 2017

[0009] Non-Patent Literature 3: Yamasawa Masao, Tsunoda Atsushi, Kondo Ken, Saito Toshiaki, Goto Seiji, Sato Naoki, Tsujii Shigeo, Noda Keiichi, "MELT-UP Activity toward High Reliability of Cryptocurrency (Bitcoin) and Block Chain - Centered on Secret Key Management," Proc. SCIS2018, 4F2-2. Jan. 2018.

[0010] Non-patent literature 4: Masao YAMAZAWA, Masahito GOTO, Hiroshi YAMAMOTO, Shigeo TSUJII, Shigeo NAMBA, Hisanori YOSHIMITSU, Keiichi NODA, "On the Secret Key Protection of Secure Distributed PDS System," Proc. SCIS2020, 2C2-2. Jan. 2020.

[0011] Non-patent literature 5: Masao YAMAZAWA, Masahito GOTO, Hiroshi YAMAMOTO, Yoshikazu MATSUTOMO, Kimitoshi SHIRATSU, Dairoku TOYOSHIMA, Koki SESERI, Ken KONDO, Shigeo TSUJII, "On the Requirements of Secret Sharing Method for Authentication of Distributed Population," Proc. SCIS2020, 4G-1, pp. 688-692, June 2020.

[0012] Non-patent literature 6: Masahito GOTO, Masao YAMAZAWA, Hiroshi YAMAMOTO, Ryo FUJITA, Yoshikazu MATSUTOMO, Kimitoshi SHIRATSU, Dairoku TOYOSHIMA, Koki SESERI, Ken KONDO, Shigeo TSUJII, "On the Secret Sharing Method for Authentication of Distributed Population," Proc. SCIS2020, 4G-2. pp. 693-697, June 2020.

[0013] Non-patent literature 7: Masao YAMAZAWA, Takeshi YONETSU, Masahito GOTO, Hiroshi YAMAMOTO, Shigeo TSUJII, "On the Application and Implementation Requirements of Secret Sharing Library," "Multimedia, Distributed Coordination and Mobile Symposium 2021 Proceedings," Proc. DICOMO2021, 1G-3, pp. 163-164, Jun 2021.

[0014] Non-patent literature 8: Masao YAMAZAWA, Takeshi YONETSU, Masahito GOTO, Hiroshi YAMAMOTO, Ryo FUJITA, Shigeo TSUJII, "On the Requirements in the Implementation of Secret Sharing Library," "Multimedia, Distributed Coordination and Mobile (DICOMO2022) Symposium Proceedings," Proc. DICOMO2022, 1D-1, pp. 51-54, Jul. 2022.

[0015] Non-patent literature 9: Tatsuya OKAMOTO, Hiroshi YAMAMOTO, "Modern Cryptography," Sangyo Shuppan Co., Ltd., June 30, 1997 (first edition), June 30, 2000 (first printing).

[0016] Non-Patent Literature 10: Hiroshi YAMAMOTO, "(k, L, n) threshold secret sharing system", Transactions of the Institute of Electronics, Information and Communication Engineers, vol. J68-A, no. 9, pp. 945-952, Sep. 1985, [English translation: Electronics and Communications in Japan, Part I, vol. 69, no. 9, pp. 46-54, (Scripta Technica, Inc.), Sep. 1986]

[0017] Non-Patent Literature 11: Naoaki Nishi, Katsunori TAKIZAWA, "Strong Ramp-type Threshold Secret Sharing Method by Polynomial Interpolation", Transactions of the Institute of Electronics, Information and Communication Engineers, vol. J92-A, no. 12, pp. 1009-1013, Dec. 2009.

[0018] Non-Patent Literature 12: Hiroshi YAMAMOTO, "Multi-secret sharing method", Proc. The 44th Symposium on Information Theory and its Applications (SITA2021), Nishinomiya, Hyogo, Japan, Dec. 8-10, 2021.

[0019] Non-Patent Literature 13: Masao YAMAZAWA, Takeshi YONETANI, Masashi GOTA, Hiroshi YAMAMOTO, Ryo FUJITA, Shigeo TSUJII, "Demonstration of Requirements in Implementation of Secret Sharing Library", Proc. SCIS2023, 2B2-4, Jan. 24~27, 2023

[0020] Non-Patent Literature 14: Da ICHIRAS, Kaito Tsuchiya, Yuuto Kawahara, "SHH: Super High Speed Secret Sharing Library for Object Storage", Information Processing Society of Japan Research Report Vol. 2015-CSEC-70 No. 26, Vol. 2015-SPT-14 No. 26, 2015 / 7 / 3

[0021] Non-Patent Literature 15: "[Windows 11] MD / SHA-1 / SHA-256 Hash Values to Calculate File Identity", [online], [Retrieved on November 2, 2023], Internet <URL: 1714031509391_0.html> SUMMARY

[0022] PROBLEMS TO BE SOLVED BY THE INVENTION

[0023] The secret sharing method is characterized by balancing the security of protecting important information from risks and the availability of system sustainable operation, which are two mutually contradictory requirements. However, the availability response in the "MSSS" of Non-Patent Literature 12 can be insufficient.

[0024] The present application has been achieved in order to solve such conventional problems, and the problem to be solved is to improve availability while maintaining the high-speed operability in the "MSSS" of Non-Patent Literature 12, thereby balancing the security and availability of the secret sharing method.

[0025] Solution to the problem

[0026] One embodiment of the present application is a method executed by a secret sharing system, characterized by

[0027] The secret sharing system includes:

[0028] a dispersed data generation section that generates a group of pieces obtained by encoding and dispersing a dispersed object data;

[0029] a plurality of holding sections that hold the pieces respectively; and

[0030] a data recovery section that collects the pieces held in the holding sections and reconstructs the dispersed object data based on a reconstruction dispersion number specified in advance by decoding,

[0031] The method includes:

[0032] a dispersed data generation step in which the dispersed data generation section generates a group of a plurality of pieces corresponding to the reconstruction dispersion number from the dispersed object data and a spare piece obtained by redundantly encoding the group of pieces, and disperses and holds the generated pieces and the spare piece in the holding sections respectively; and

[0033] a reconstruction step in which the data recovery section collects the pieces and the spare piece from the holding sections, and if the group of the collected pieces satisfies the reconstruction dispersion number, reconstructs the dispersed object data using the group of pieces, and if the collected pieces do not satisfy the reconstruction dispersion number, reconstructs the dispersed object data using the existing pieces and the spare piece to find the non-existing pieces.

[0034] In addition, another embodiment is a method executed by a secret sharing system, characterized by

[0035] The secret sharing system includes:

[0036] a dispersed data generating section that generates a group of pieces obtained by encoding and dispersing a dispersed object data;

[0037] a plurality of holding sections that hold each of the pieces;

[0038] a data restoring section that collects the pieces held in each of the holding sections and reconstructs the dispersed object data based on a predetermined reconstruction dispersion number by decoding,

[0039] the method includes the steps of:

[0040] a dispersed data generating step in which the dispersed data generating section generates a group of a plurality of pieces corresponding to the reconstruction dispersion number from the dispersed object data and a spare piece obtained by redundantly encoding the group of pieces, and holds each of the generated pieces and the spare piece in the holding section, wherein the spare piece is one or two; and

[0041] a reconstructing step in which the data restoring section collects the pieces and the spare piece from the holding section, and if the group of the collected pieces satisfies the reconstruction dispersion number, reconstructs the dispersed object data using the group of pieces, and if the collected pieces do not satisfy the reconstruction dispersion number, reconstructs the dispersed object data using the spare piece,

[0042] wherein the dispersed data generating step includes the steps of:

[0043] a first step of reading each of the information using a dispersed data function that holds the reconstruction dispersion number (k), the number of spare pieces (r), the data length (L) of the dispersed object data, and the information of the dispersed object data;

[0044] a second step of generating a random number and making the generated random number, the data length (L), and a data string S of the dispersed object data, the number of bytes of the data string S being a multiple of "k x 8"; and

[0045] a third step of dividing the data string S into blocks of "k x 8", and calculating pieces "W" by performing an operation on each of the divided blocks,

[0046] wherein in the third step, in the case of "the number of spare pieces (r) = 1", using formula (28), formula (29),

[0047] [Formula 28]

[0048] ,

[0049] S xyThe x-th value of block number y, where x = 1, 2, 3…k, is 8 bytes.

[0050] W zy The value of the fragment with block number y and scatter number z, where z = 1, 2, 3…k, is 8 bytes.

[0051] W (k+1)y The spare fragment,

[0052] [Number 29]

[0053] ,

[0054] On the other hand, when "the number of spare fragments (r) = 2", equations (30) and (31) are used.

[0055] [Number 30]

[0056] ,

[0057] [Number 31]

[0058] ,

[0059] S xy The x-th value of block number y, where x = 1, 2, 3…k, is 8 bytes.

[0060] W zy The value of the fragment with block number y and scatter number z, where z = 1, 2, 3…k, is 8 bytes.

[0061] W (k+1)y W (k+2)y The spare fragment.

[0062] Another aspect of the present invention is a secret sharing system, characterized by comprising:

[0063] The distributed data generation unit generates groups of fragments obtained by encoding and distributing the distributed object data;

[0064] Multiple storage units, each of which stores a specific fragment; and

[0065] The data recovery unit collects the fragments stored in each of the storage units and reconstructs the fragmented object data by decoding them based on a pre-defined number of reconstruction fragments.

[0066] The distributed data generation unit generates multiple groups of fragments corresponding to the reconstructed distributed number based on the distributed object data, as well as spare fragments obtained by redundancy of the groups of fragments. The generated fragments and the spare fragments are then distributed to the storage unit and stored thereon.

[0067] The data recovery unit collects the fragments and the spare fragments from the storage unit.

[0068] If the collected groups of the fragments satisfy the reconstructed dispersion number, then the groups of the fragments are used to reconstruct the dispersed object data.

[0069] On the other hand, if the collected fragments do not meet the reconstructed scatter number, the existing fragments and the spare fragments are used to find the non-existent fragments to reconstruct the scatter object data.

[0070] Furthermore, the present invention can also be configured as a secret sharing program that causes a computer to execute the secret sharing method.

[0071] The effects of the invention

[0072] According to the present invention, both confidentiality and usability of the secret sharing method can be achieved. Attached Figure Description

[0073] Figure 1 This is a concept diagram of the secret sharing method.

[0074] Figure 2 This is an explanatory diagram of the "(k,n) threshold method".

[0075] Figure 3 This is a diagram illustrating the principle of recovering key information S.

[0076] Figure 4 This is a chart showing a situation where recovery is impossible.

[0077] Figure 5 This is a system structure diagram of the secret sharing method involved in implementing the embodiments of the present invention.

[0078] Figure 6 This is a computer architecture diagram for implementing this secret sharing method.

[0079] Figure 7 This is a conceptual diagram of the system processing for implementing this secret sharing method.

[0080] Figure 8 This is a processing diagram of the distributed steps to implement this secret sharing method.

[0081] Figure 9 This is a processing diagram for achieving independent and uniform dispersion of the secret sharing method.

[0082] Figure 10 This is a processing diagram illustrating the recovery steps of this secret sharing method. Detailed Implementation

[0083] The secret sharing method according to the embodiments of the present invention will now be described. This secret sharing method involves the "k-fold (k, k+1) and (k, k+2) threshold method" of "MSSS" in Non-Patent Document 12. Here, it is abbreviated as "Extended Multiple Secret Sharing Scheme", or "AMSSS".

[0084] AMSSS Instructions

[0085] First, let's explain the basic idea behind the secret sharing method. When installing the secret sharing method library, the following considerations should be taken into account.

[0086] A: The scale of the data to be distributed (distributed object data)

[0087] B: Dispersion number (number of fragments)

[0088] C: Forms of sharding used in data reconstruction

[0089] D: Processing time (speed) required for secret sharing and reconstruction.

[0090] E: Security strength required for random numbers, etc.

[0091] F: Possible errors during use

[0092] G: Potential attacks on library functionality and applications that use that functionality.

[0093] H: When provided to users as a product, the following should be guaranteed: security, functionality, supported features, and features that should be intentionally limited or removed to prevent user confusion or decreased responsiveness.

[0094] In particular, the availability requirements of consideration H must be examined. That is, when implementing the secret sharing law, it is necessary to balance the two conflicting requirements of security and availability mentioned above.

[0095] When based on Figure 1 When considering the risks associated with "important information," one can typically foresee the risks of information loss due to "malfunctions or damage," as well as the risks of information being taken by unintended parties due to "leakage."

[0096] Redundancy is effective in preventing "failures and sabotage." That is, creating copies achieves risk diversification. On the other hand, regarding the risk of "leakage," creating copies actually increases the risk, so it's desirable to avoid duplication. Alternatively, one might want to ensure that even if some information is obtained, it won't be leaked.

[0097] The former is a measure to enhance usability, while the latter is a measure to enhance confidentiality; these are contradictory requirements. As shown in Non-Patent Literature 1 and 2, this dilemma can be alleviated when using the secret sharing method.

[0098] 1. Overview of the Secret Sharing Method

[0099] As shown in Non-Patent Document 9, information security measures require the achievement of "confidentiality, availability, and integrity". Figure 1 The countermeasures against "leakage" are equivalent to achieving confidentiality, while the countermeasures against "failure and damage" are equivalent to achieving availability.

[0100] Confidentiality and availability are contradictory requirements, necessitating appropriate technologies in practice, but encryption is often prioritized over confidentiality. However, when considering how to address the risks of decryption and availability inherent in encryption, secret sharing can be considered more advantageous.

[0101] (1) Shamir's Secret Sharing Method

[0102] As one of the secret sharing methods, there is the "(k, n) threshold method, n>k≥2" in non-patent literature 1. This method intuitively utilizes the theorem that there exists only one "function of degree no more than k" that passes through (k+1) distinct points in the "x, y" plane.

[0103] For example, Figure 3 As shown, when "k=1", if secret information (dispersed morphological data) S is set on the y-axis, a straight line with slope "a" is drawn through the point (0, S), and different points "x" on the x-axis are given. i ,(i=1,n)”, then we can use equation (1) to calculate the corresponding y coordinate value “y i (i=1, n)”.

[0104] [Number 1]

[0105]

[0106] Where, "i=1, n (n≡ "segment number"), "a" is any number. Furthermore, as shown in number (2), it will be associated with each "x i The corresponding "y" i The combination of "" is used for sharding.

[0107] [Number 2]

[0108]

[0109] During the recovery process, as shown in equations (3-1) and (3-2), two arbitrarily selected fragments "W" with different scatter numbers are used. p “W” r ", by combining the equations, eliminate "a". Thus, as shown in equation (4), the secret information S is obtained.

[0110] [Number 3]

[0111]

[0112] [Number 4]

[0113]

[0114] On the other hand, such as Figure 4 As shown, a single slice cannot fix the straight line and cannot recover the secret information S. When the threshold "k" is "3 or above", the function used is set to a (k-1) degree function, that is, "f(t): a (k-1) degree function of t".

[0115] Furthermore, while the above explanation used the real number plane for ease of intuitive understanding, the computation in the program implementation takes place over a finite field, namely the "Galois field: GF(q)". First, as shown below, let the secret information to be distributed be S, and let the set of random numbers used for secret distribution encoding be U.

[0116] Secret information: S∈GF(q)

[0117] • Random number: U = (U1, U2, U3, ..., U k-1 ∈GF(q)

[0118] Define f(t) using these methods as in equation (5).

[0119] [Number 5]

[0120]

[0121] Use "α" i ∈GF(q), i=1,2,…,n (where, for i≠j, αi≠αj)”, and perform the partitioning {W1, W2, W3,…,W n Encoding and decoding of}.

[0122] Encoding: W i =f(α i )∈GF(q)

[0123] Decoding: Based on Lagrange interpolation. Ik = {i1, i2, i3, ..., i...} k}

[0124] [Number 6]

[0125]

[0126] According to equation (5), "S = f(0)". Therefore, it is possible to decode using equation (7) based on equation (6-2).

[0127] [Number 7]

[0128]

[0129] (2) Features of the threshold method

[0130] According to this "(k, n) threshold method", f(t) can be determined by any "k" slices, and S can be recovered. On the other hand, S cannot be known at all by only any "k-1" slices.

[0131] Furthermore, even if "nk" units are lost, S can still be recovered, and even if "k-1" units are stolen, the information in S will not be leaked at all. That is, availability and confidentiality are both ensured in terms of protecting S. This feature can be used to explore applications such as distributing important PKI information and controlling access rights (Non-Patent Documents 4-6), and protecting cryptocurrencies (Non-Patent Document 3).

[0132] 2. Secret sharing method and total data volume

[0133] In the "(k,n) threshold secret sharing method", the secret information S is encoded and distributed into n fragments "W". j "j=1,2,…,n". On the other hand, each W j The bit length (data volume) must be at least the same as S.

[0134] Therefore, the total data size of all fragments is n times the data size of S. If the data size of the scatter objects is small, the cost of storing the fragments is relatively low, so this characteristic will not be a problem.

[0135] On the other hand, problems arise when trying to distribute and store massive amounts of data. Specifically, the cost of distributing the originally large amount of data several times becomes extremely high. Therefore, methods to reduce the total amount of data become important for practical application.

[0136] (1) Sharing method and data volume

[0137] As a secret sharing method, Non-Patent Document 2 proposes ways to reduce the amount of data and simplify calculations.

[0138] The former is proposed as the "(k, L, n) step threshold secret sharing method" (method 3 in Table 1) based on the "Shamir type" structure. On the other hand, the latter is proposed as the "(n, n) threshold secret sharing method" (method 2 in Table 1).

[0139] [Table 1]

[0140]

[0141] According to the former, the total amount of data in the distributed results is “n / L” times the amount of data in the secret information S, which can be reduced to “1 / L” of the “(k,n) threshold secret sharing method” of the “Shamir type”.

[0142] However, in the case of the "Shamir type", setting "L=n" may not guarantee security against leakage. Therefore, the total number of fragments will always exceed the amount of data in S. On the other hand, if the method of Non-Patent Documents 10 and 11 is used (method 5 in Table 1), it is possible to set "k=L=n", which allows encoding in a way that the total amount of data in the fragments is the same as the amount of data in S.

[0143] However, this method has the drawback of high computational cost. Additionally, if the "(n,n) threshold secret sharing method" is used, the computational cost is very small, but similar to the "Shamir type" method, the total data volume increases to n times the data volume of S, resulting in the inability to reduce the data volume, similar to the ladder type.

[0144] Therefore, in methods 1 to 5 of Table 1, it is difficult to simultaneously reduce both data volume and computational cost. In contrast, "MSSS" (method 6 of Table 1) in Non-Patent Document 12, as the practical results reported in Non-Patent Document 13 demonstrate, has the advantage of superior performance in both computational cost and data reduction.

[0145] (2) Examination of methods for creating spare fragments

[0146] In Method 6 of Table 1, if not all fragments are collected (even if one is missing), the original secret information S cannot be recovered. This is because the total data volume of the fragments is the same as the data volume of S, which is a necessary property from the perspective of information theory.

[0147] However, when considering applications such as backup systems for massive amounts of data, the inability to cope with the risks of failure and destruction presents a problem that could become a design flaw in the system. A way to create backup shards is desired.

[0148] Indeed, the problem can be solved by using the general strong-security "(k, L, n) thresholding method" or "k-fold (k, n) thresholding method", but it may increase the computational load of encoding / decoding.

[0149] On the other hand, in practical cases, in most cases, a spare number of shards (nk) of around 1 or 2 is sufficient. Therefore, the following preferred "k-threshold (k, n)" for the case of (nk=1, 2), i.e., spare shards derived based on "AMSSS".

[0150] 3. Encoding / Decoding Based on AMSSS

[0151] In the encoding / decoding computation based on "AMSSS", as the data body, "q=2" is used. m Let m > 3k, and let α be the primitive element of GF(q). Define the k×(k+2) matrix G using equation (8).

[0152] [Number 8]

[0153]

[0154] Furthermore, in "GF(2 m In the operation of "-1=1", since "-1=1", the subtraction sign "-" in the expression can be completely replaced by the sign "+".

[0155] (1) Encoding calculation (encoding calculation for spare 2)

[0156] First, set the secret information as "(S1, S2, ..., S..." k ), S i ∈GF(q)”

[0157] Set the shards (shares) to "(W1, W2, ..., W k+1 W k+2 ), W j ∈GF(q)",

[0158] Let G be a "k×(k+2)" matrix with values ​​on "GF(q)" as its elements. Then, it is encoded using equation (9).

[0159] [Number 9]

[0160]

[0161] [Number 10]

[0162] For the matrix G and in equation (8) (T denotes transpose), when a k×(k+1) matrix is ​​transposed... Defined as In order to implement the k-th (k,k+2) threshold method through equation (9), for any i (1≤i≤k), Any k columns of the matrix must be linearly independent. This condition is satisfied when k ≥ 4, but not when k = 2 or 3. Therefore, the following matrix is ​​used as matrix G.

[0163]

[0164] At this point, “W” can be obtained through equations (11) and (12). j ,j=1,2,…,k.

[0165] [Number 11]

[0166]

[0167] [Number 12]

[0168]

[0169] In addition, W can be obtained from equations (9) and (10) through equations (13) and (14). k+1 W k+2 .

[0170] [Number 13]

[0171]

[0172] [Number 14]

[0173]

[0174] (2) Decoding calculation (decoding calculation for backup 2)

[0175] Next, consider decoding. Given "W1, W2, ..., W..." k When “S1, S2, ..., S” can be calculated as in equations (15) and (16). k ".

[0176] [Number 15]

[0177]

[0178] [Number 16]

[0179]

[0180] According to equations (11), (12), and (15), the relationship in equation (17) holds, thus yielding the relationship in equation (16).

[0181] [Number 17]

[0182]

[0183] Additionally, consider using "W" k+1 "The decoding process is performed. Based on equations (9) to (16), we can obtain "W" using equation (18). k+1 (Here, we will describe the case where “k” is even. In the case where “k” is odd, we will use equation (16) to represent “(α-1)”.) -1 "Changed to "α" -1 That's all.

[0184] [Number 18]

[0185]

[0186] Furthermore, “β” is defined by equation (19). k ".

[0187] [Number 19]

[0188]

[0189] At this time, "W" k+1 "Given by equation (20)."

[0190] [Number 20]

[0191]

[0192] Similarly, “γ” is defined by equation (21). k ".

[0193] [Number 21]

[0194]

[0195] Therefore, "W" k+2 "Given by equation (22)."

[0196] [Number 22]

[0197]

[0198] (3) Recovery when intermediate fragments are damaged (lost)

[0199] One or two "W"s will be lost. j The decoding method for "1≤j≤k" will be explained in cases 1 and 2.

[0200] <Scenario 1>

[0201] In the absence of satisfying "W" aThe decoding method for the case of fragmentation "1≤a≤k" is to calculate "W" according to equation (23) or equation (24). a Then, decoding is performed using equation (16).

[0202] [Number 23]

[0203]

[0204] [Number 24]

[0205]

[0206] <Scenario 2>

[0207] In the absence of "W" a W b The decoding method for the case of "1≤a≤b≤k" is given by equations (25) and (26).

[0208] [Number 25]

[0209]

[0210] [Number 26]

[0211]

[0212] [Number 27]

[0213]

[0214] At this point, W is calculated according to equation (27). a W b Then, decoding is performed using equation (16).

[0215] The above text explained the case of the k-fold (k, k+2) threshold method, but in the case of the k-fold (k, k+1) threshold method, "W" is omitted. k+2 "Just consider it. That is, omit the processing of equations (21) and (22) during encoding, and perform equations (15) and (16) or equation (23) of case 1 during decoding.

[0216] Based on the asymmetry of the matrix "G" in equation (10) used for encoding, given "W1, W2, ..., W...", k Under the condition of "W1, W2, ..., W", it can decode at high speed; otherwise, the decoding time is slightly longer. Therefore, in the case of "W1, W2, ..., W", the decoding time is longer. k W k+1 W k+2 "If all are available, use "W1, W2, ..., W k "High-speed decoding can be performed."

[0217] (3) The recovery process depends on the missing part

[0218] Generate "k+2" fragments. Even if two of them are lost, the original information S can still be recovered. The recovery process differs depending on which fragment is lost, with cases 1 and 2. However, in either case, the original information S can be recovered (reconstructed) through matrix operations.

[0219] Example 1

[0220] As an example, an application of "AMSSS" in a data center backup system is illustrated. Specifically, when a secret sharing method is used in a large-capacity data backup system, data recovery may be difficult if there are no spare fragments in the system design. Therefore, in this embodiment, when constructing a backup system using spare fragments, "AMSSS," which is derived from "MSSS" in Method 6 of Table 1, is used to achieve redundancy of fragmented data.

[0221] If one wants to construct a "threshold secret sharing method" with the same function as this embodiment using existing technologies (such as method 7 in Table 1), complex calculations are required, such as the van der Mann determinant shown in Equation 27 of Non-Patent Document 10. Therefore, this method involves a large amount of computation and is difficult to implement practically. This embodiment mainly aims to solve this problem by adopting "AMSSS," which is developed based on method 6 "MSSS" in Table 1. Thus, only the calculation of Equation 10 needs to be performed, achieving high-speed processing and ease of implementation.

[0222] Example of System Structure

[0223] based on Figure 5 and Figure 6 The following describes a structural example of a secret sharing system that implements the secret sharing method described in this embodiment. System 1 is built in a data center as a cloud storage service. It leverages the high speed of the multiple secret sharing method "MSSS" by employing the aforementioned "AMSSS" and creates one or two backup shards. In this embodiment, if the necessary number of shards is gathered (if all shards exist), the distributed object data is reconstructed.

[0224] On the other hand, if the necessary number is not collected (if some fragments do not exist), the insufficient fragments are found using the fragments of the collected amount (the amount that exists) and the spare fragments, and then the scattered object data is reconstructed.

[0225] Specifically, such as Figure 5 As shown, the system 1 consists of a storage server device 2 and is connected to network 4 (see reference). Figure 6The storage server device 2 can be a single unit or a group of multiple storage server devices 2 forming a system. Furthermore, in Figure 5 As an example, the system architecture based on a single storage server device 2 is shown.

[0226] The storage server device 2 is equipped with a communication unit 21, storage devices 3-1 to 3-n, a distributed data generation unit 14, a data recovery unit 15, and a random number generation unit 16. The distributed data generation unit 14 generates groups of fragmented data (multiple fragments / spare fragments) obtained by encoding and distributing the distributed object data (secret information S), and distributes each fragment to the storage unit 22 of the storage devices 3-1 to 3-n for storage. Furthermore, when using a group of multiple storage servers 2, each fragment can be distributed to the storage device 3 of each storage server 2 for storage.

[0227] The data recovery unit 15 collects fragmented data from storage devices 3-1 to 3-n to recover scattered object data when recovering secret information S. However, as shown in Non-Patent Document 13, since each fragment is transmitted via communication and read / written in memory, it cannot be asserted that there are no errors, and sometimes missing or damaged data may occur. It is a consensus among those skilled in the art that certain error detection methods should be added at the application level. Therefore, in this embodiment, an error detection code (EDC), i.e., a checksum, is added to each fragment to achieve error verification.

[0228] If the collected fragment groups are complete and meet the required number, the secret information S is recovered from the collected fragment groups. On the other hand, if a missing or damaged fragment occurs in the collected fragment groups, and the required number is not met, the non-existent (missing / damaged) fragment is determined using the existing fragment groups and the spare fragments, and then the secret information S is recovered.

[0229] The random number generation unit 16 generates random numbers (8 bytes) according to the instructions of the scattered data generation unit 14. Furthermore, as... Figure 6 As shown, the storage server device 2 is configured with a computer as the main body, and includes a communication device 51, an input device 52, a memory (main storage device) 53, an auxiliary storage device (SSD, HDD, etc.) 54, a CPU 55, and an output device 56. The hardware 51 to 56 are connected via a bus 57, and the various parts 3, 14, 15, 16, 21, and 22 are realized through the cooperation of the hardware and software.

[0230] Processing content of Distributed Data Generation Department 14 / Data Recovery Department 15

[0231] The distributed data generation unit 14 has a distributed data generation function (function: amsss_generate) that generates fragments and spare fragments based on the distributed object data (secret information S).

[0232] On the other hand, the data recovery unit 15 has a reconstruction function (function: amsss_reconstruct) that uses sharded groups / alternate shards to reconstruct and recover scattered object data. Here, "AMSSS" is used in both the scattered data generation and reconstruction functions. Below, based on... Figure 7 Let me explain the details.

[0233] (1) Distributed data generation function

[0234] like Figure 7 As shown by arrows A and B, the distributed data generation unit 14 uses the distributed data generation function to generate fragments (1) to (k) and spare fragments (k+1) and (k+2) based on the distributed object data.

[0235] The function generates scattered data at the starting address of the one-dimensional array of the input variable, "Si(unsignedchar)". The fragments storing distributed object data (secret information S) in "Si)" can be configured with the following information by specifying them.

[0236] A: The number of fragments required to reconstruct the scattered object data (the so-called threshold "range of 2~6": hereinafter referred to as the number of fragments to be reconstructed).

[0237] B: Number of spare fragments (range 1~2: hereinafter referred to as the number of spare fragments).

[0238] C: Data length of the scattered object data (1~2) 50 (range of bytes)

[0239] Additionally, the scatter data generation function uses the starting address of the two-dimensional array of the argument (output) "Sni(unsigned char Each line of “Sni” is set to “Distributed number (1 byte) + Fragment data (total number of blocks × 8) bytes”.

[0240] The total number of blocks is "(distributed object data length + 16) ÷ (reconstructed distribution number × 8)". Here, "(distributed object data length + 16) ÷ (reconstructed distribution number × 8)" represents the number after rounding down to the nearest whole number. In this embodiment, the verification of each fragment's data and the use of hash values ​​are discussed. While calculating hash values ​​to check file identity is a well-known technique according to non-patent document 15 and other publicly available documents, the fragments described below are used as a means of confirming normality / abnormality. For example, if the reconstructed distribution number is set to "3", the spare distribution number is set to "2", and the two-dimensional array is set to "Sni", then...

[0241] Sni[0]: Distributed data of scatter number 1 + shard data of scatter number 1 (shard) + hash value

[0242] Sni[1]: Dispersion number 2 + Dispersion number 2's fragmented data (fragment) + hash value

[0243] Sni[2]: Dispersion number 3 + fragment data of dispersion number 3 (fragment) + hash value

[0244] Sni[3]: Distributed number 4 + Distributed number 4 fragment data (spare fragment) + hash value

[0245] Sni[4]: Dispersion number 5 + data of the shard with dispersion number 5 (spare shard) + hash value.

[0246] Therefore, the number of rows / columns of the two-dimensional array is:

[0247] The number of rows in a 2D array = number of reconstructed scatters + number of spare scatters

[0248] The number of columns in a two-dimensional array = the total number of blocks × 8 + 1 (bytes).

[0249] Below, based on Figure 8 To explain the processing content based on the distributed data generation function (S01~S06).

[0250] S01: When processing begins, the distributed data generation unit 14 reads the reconstructed distributed number (k: 1 byte), the spare distributed number (r: 1 byte), the data length of the distributed object data (L: 8 bytes), and the distributed object data (L bytes) set in the one-dimensional array Si of the distributed data generation function.

[0251] S02: Calculate the hash value of the scatter object data (OpenSSL SHA-256). Set the hash value obtained here as the starting address of the one-dimensional array of the scatter data generation function "CS(unsigned char CS).

[0252] S03: Output instructions to the random number generation unit 16 to generate random numbers (8 bytes) based on the time. At this time, the function "clock_gettime()" is used to obtain the time to generate random numbers.

[0253] Obtain the random number generated here and create a data string consisting of "random number + data length L of the scattered object + data of the scattered object". If the number of bytes in the created data string is not a multiple of "k×8", pad the end of the data string with "x'00'" to make it a multiple of "k×8".

[0254] S04: Make the data strings created in S03 distributed independently and uniformly as described later.

[0255] S05: Divide the data string S, which has been independently and uniformly distributed by S04, into "k×8" blocks. For each block, use "GF(2" in non-patent document 14. 64 The following operations are performed on the data to obtain the fragmented data "W(W)". zy At this point, set "α" to "GF(2)". 64 The original element of ").

[0256] Specifically, confirm the number of spare distributions r read in S01. If "the number of spare distributions r = 1", then use equations (28) and (29).

[0257] [Number 28]

[0258]

[0259] [Number 29]

[0260]

[0261] On the other hand, if “the number of spare distributions r = 2”, then equations (30) and (31) are used.

[0262] [Number 30]

[0263]

[0264] [Number 31]

[0265]

[0266] In equations (28) and (29),

[0267] S xy The x-th value (8 bytes) of block number y.

[0268] W zy : The value of the fragment with block number y and scatter number z (8 bytes)

[0269] At this point, "W" is obtained through the operations of equations (32) and (33). zy , z = 1, 2, 3, ..., k.

[0270] [Number 32]

[0271]

[0272] [Number 33]

[0273]

[0274] In addition, according to equations (28) to (31), “W” can be obtained through equations (34) and (35). (k+1)y “W” (k+2)y ".

[0275] [Number 34]

[0276]

[0277] [Number 35]

[0278]

[0279] S06: Calculate the data "W" for each fragment obtained in S05. zy The hash value of "W". Then, the data from each shard "W" is... zy The hash value is set in the two-dimensional array "Sni" of the scatter data generation function. At this time, the shard data with scatter number i and the hash value are assigned to the i-th row of the two-dimensional array "Sni", for example, if the scatter number is 2, it is assigned to the second row.

[0280] Moreover, the data in each shard "W" zy "They are separately distributed and stored in the storage sections 22 of the storage devices 3-1 to 3-n. Furthermore, in..." Figure 8 The case of "the number of spare distributions r=2" is shown. In the case of "the number of spare distributions r=1", the calculation formula of S05 is replaced by formula (28) instead of formula (30), and the last (k+2)th row is removed from the two-dimensional array "Sni" in S06, up to the (k+1)th row.

[0281] (2) Details of S04

[0282] According to "MSSS", as an algorithmic constraint used to satisfy security, for "S(secret information: distributed object data) = (S 1y S 2y S xy ), each "S" xy"It needs to be independent and uniformly distributed. Therefore, in S04, the following independent uniform distribution process is performed (S11~S14).

[0283] S11: When processing begins, perform multiplication over a finite field using appropriate matrices. For example, as shown in equation (36), suppose that the multiplication is performed modulo 2. 64 (1 word 64 bit)” uses random numbers to define the elements of matrix A, and its inverse matrix B is calculated in advance. If the value of matrix |A| is different from “2 64 "When the numbers are coprime, i.e. odd, the inverse matrix B can be found."

[0284] [Number 36]

[0285]

[0286] mod2 64 = modulo 2 64

[0287] S12: For the data string before independent uniform distribution, sequentially from front to back in "modulo 2" 64 Perform operations on matrix A to generate a new data string.

[0288] [Number 37]

[0289]

[0290] The processing example will be illustrated based on equation (37). Here, the data string before independent uniform distribution is represented as data string X(X). (1) X (2) X (3) X (4) X (5) X (6) ...X (M-2) X (M-1) X (M) Additionally, the new data string is represented as data string Y (Y (1) Y (2) Y (3) Y (4) Y (5) Y (6) ...Y (M-2) Y (M-1) Y (M Each X and Y is 8 bytes. Regarding the generation of each part of this processing example, modulo 2 is used respectively. 64 Operations on matrix A on the matrix A.

[0291] First, regarding the data string X shown in equation (37-1), (X... (1) X (2) X (3)), and perform the operation as shown in equation (37-2) to obtain (Y (1) Y (2) Y (3) ).

[0292] Next, for the data string (Y) of equation (37-3) (1) Y (2) Y (3) X (4) X (5) X (6) ...X (M-2) X (M-1) X (M) (Y) (2) Y (3) X (4) ), as shown in equation (37-4), perform the operation to obtain (Y) (2) Y (3) Y (4) Therefore, the data string (Y) of generative formula (37-5) (1) Y (2) Y (3) Y (4) X (5) X (6) ...X (M-2) X (M-1) X (M) ).

[0293] Like this, by sequentially processing (X) in the data string X from front to back (1) X (2) X (3) X (4) X (5) X (6) ...X (M-2) X (M-1) X (M) The operation is performed on every three data points to generate the data string (Y) shown on the left side of equation (37-6). (1) Y (2) Y (3) Y (4) Y (5) Y (6) ...Y (M-2 )Y (M-1) X (M) ).

[0294] Then, as shown in equation (37-7), for (Y) (M-2) Y (M-1) X (M) The operation is performed to obtain (Y) (M-2) Y (M-1) Y (M) ), and generate the data string Y(Y) of equation (37-8).(1) Y (2) Y (3) Y (4) Y (5) Y (6) ...Y (M-2) Y (M-1) Y (M) ).

[0295] S13: For the data string Y generated in S12, perform a permutation of the high 32 bits and the low 32 bits between adjacent data to generate a new data string Z.

[0296] For example, Figure 9 As shown, the data Y (1) The lower 32 bits and data Y (2) The high 32 bits are permuted to obtain the data Y. (2) The lower 32 bits and data Y (3) The high 32 bits are permuted to obtain the data Y. (3) The lower 32 bits and data Y (4) The most significant 32 bits are permuted, and the same permutation is performed on the subsequent data Y. Finally, the data Y is... (M) The lower 32 bits and data Y (1) The high 32 bits are permuted to generate the data string Z.

[0297] S14: For the data string Z generated in S13, sequentially from back to front in the modulo 2... 64 Multiplication of matrix A is performed on the matrix to generate a new data string S.

[0298] [Number 38]

[0299]

[0300] The processing example is illustrated based on equation (38). Each generation of this processing example also utilizes "modulo 2". 64 Operations on matrix A on the matrix A.

[0301] Equation (38-1) represents the permuted data string Z(Z) (1) Z (2) Z (3) Z (4) Z (5) Z (6) ...Z (M-4) Z (M-3) Z (M-2) Z (M-1) Z (M) First, as shown in equations (38-2) and (38-3), for the data string Z, (Z (M-2) Z (M-1)Z (M) ) perform the aforementioned operation to generate (S (M-2) S (M-1) S (M) Then, as shown in equations (38-4) and (38-5), for (Z) (M-3) S (M-2) S (M-1) ) perform the aforementioned operation to generate (S (M-3) S (M-2) S (M-1) ).

[0302] In this way, by performing the aforementioned operation on every three data points of the data string Z sequentially from back to front, the (Z) of equation (38-6) is generated. (1) S (2) S (3) S (4) S (5) S (6) ...S (M-4) S (M-3) S (M-2) S (M-1) S (M) ).

[0303] Then, as shown in equation (38-7), for (Z) (1) S (2) S (3) ) perform the aforementioned operation to generate (S (1) S (2) S (3) Therefore, as shown in equation (38-8), a new data string (S) is generated. (1) S (2) S (3) S (4) S (5) S (6) ...S (M-4) S (M-3) S (M-2) S (M-1) S (M) ).

[0304] (2) Refactoring Functionality

[0305] based on Figure 10 To illustrate the reconstruction function, we will use groups of sharded data (shards / spare shards) to reconstruct and restore the original sharded object data by performing a process that is essentially the opposite of the sharded data generation function.

[0306] Figure 7 Arrows C and D in the diagram indicate the state of using spare fragments (k+2) to recover the original scattered object data due to corruption or loss of the scattered data (2). Furthermore, the following information is set in the arguments of the reconstruction function.

[0307] <Input Argument>

[0308] • "N(unsigned char N)”

[0309] The starting address of a one-dimensional array containing the reconstructed scatter and the backup scatter is set to the same value as "Si" in the scatter data generation function.

[0310] • "P(unsigned char P)”

[0311] The starting address of a one-dimensional array containing the length (8 bytes) of the scattered object data (set to the same value as "Si" in the scattered data generation function).

[0312] • "CS(unsigned char CS)”

[0313] The starting address of a one-dimensional array (32 bytes) containing the hash values ​​of the scattered object data (set to the same value as the hash value "CS" returned by the scattered data generation function).

[0314] •Sni(unsigned char Sni)”

[0315] Set the starting address of the two-dimensional array that stores the fragments used for reconstruction. This two-dimensional array is set in the same way as the "Sni" of the scatter data generation function. For example, in the case of reconstructing scatters (3) and having 2 spare scatters, if the fragments with scatter numbers "1" and "3" cannot be collected due to missing or damaged fragments, then the two spare fragments are used for reconstruction. In this case, the two-dimensional array is as follows (the scatter numbers do not need to be arranged in order).

[0316] Sni[0]: Dispersion ID 2 + Dispersion ID 2's fragment data (fragment) + hash value

[0317] Sni[1]: Distributed number 4 + Distributed number 4 fragment data (spare fragment) + hash value

[0318] Sni[2]: Distributed number 5 + Distributed number 5 fragment data (spare fragment) + hash value

[0319] Furthermore, the number of rows in the two-dimensional array is the number of reconstructed scatters, and the number of columns is "total number of blocks × 8 + 1 (bytes)". This total number of blocks is the same as "Sni" of the scatter data generation function.

[0320] <Independent variable (output)>

[0321] • "Si(unsigned char Si)”

[0322] Let the starting address be a one-dimensional array that stores the decoded scattered object data. The size of the array is the same as the data length set for the independent variable (input) "P". The following describes the processing content based on the reconstruction function (S21~S23).

[0323] S21: When processing begins, collect fragmented data from storage devices 3-1 to 3-n, i.e., fragments "W1, W2, ..., W...". k "and spare fragments (W when there are 2 spare fragments)" k+1 “W” k+2 "W" is the case where there is 1 spare fragment. k+1 Then, determine the collected fragments "W1, W2, ..., W". k "Whether the number of reconstructed fragments is met. In this embodiment, the determination is performed through two stages."

[0324] Sometimes, due to malfunctions in the storage device 3, network errors, or other reasons, fragments cannot be collected within the pre-defined time, resulting in missing fragments due to timeout. Therefore, as the first step (first stage), the number of fragments corresponding to the reconstructed scatter number is checked based on the scatter number of each collected fragment. For example, if the collection of fragments with scatter numbers "1" and "3" times out, fragments with scatter numbers "1" and "3" are missing. In this case, since fragments with scatter numbers "1" and "3" do not exist, it is determined that the reconstructed scatter number is not met. On the other hand, if the number of groups of collected fragments meets the number corresponding to the reconstructed scatter number, the second step (second stage) of error checking is initiated.

[0325] In the second step (second stage), the collected fragments "W1, W2, ..., W" are calculated. k The hash value of “” is compared with the hash value set for “Sni” of the refactoring function.

[0326] If the comparison result shows that the two hash values ​​are the same, then the shard is considered normal. If all shards "W1, W2, ..., W..." are the same, then the shard is considered normal. k "If a value is deemed normal, then the number of normal values ​​corresponding to the reconstructed dispersion number is complete, thus satisfying the reconstructed dispersion number requirement. In this case, only the fragments "W1, W2, ..., W..." are considered normal." k "To decode the original scattered object data."

[0327] On the other hand, if the two hash values ​​are different, data corruption occurs, and the shard is therefore considered abnormal. For example, for the shard with scatter number "3", if the two hash values ​​are different, the shard is considered abnormal. In this case, the shard with scatter number "3" is considered non-existent due to its missing data, and the requirement to reconstruct the scatter number is not met. Furthermore, it is preferable to perform the same error check on the spare shards to confirm the normal / abnormal setting.

[0328] Furthermore, if the reconstruction scattering requirement is not met, the spare fragment "W" is used. k+1 “W” k+2 (The case where there are 2 spare distributions). For example Figure 10 S21 shows the relationship with Figure 7 Similarly, if there is no fragment with fragment number "2", the "k+2"th spare fragment is used instead to decode along with the existing fragments. However, this is preferably limited to spare fragments that have been identified as normal through error checking.

[0329] In this way, this embodiment makes the fragmented data redundant through "AMSSS", so even in fragments "W1, W2, ..., W k Even when one or two missing / damaged data points are generated, it is still possible to recover scattered object data. During recovery, based on the fragmented data set in the two-dimensional array "Sni", each block is processed in "GF(2 64 The specified operations are performed on the block to decode the scattered object data. The following represents the content of the operation (block number y is omitted and marked).

[0330] Using only "W1, W2, ..., W k "Decoding situation"

[0331] If the reconstruction dispersion number is satisfied, equations (39) and (40) are used to utilize only the fragments "W1, W2, ..., W k "Decode."

[0332] [Number 39]

[0333]

[0334] [Number 40]

[0335]

[0336] W z : The value of the fragment number z of the block to be decoded (8 bytes)

[0337] S z The z-th value (8 bytes) of the decoded block.

[0338] Here, we first calculate "(α-1)". -1 “α” -1 In the program, it is defined as an "unsigned long long" type constant to speed up the decoding process.

[0339] <Does not exist (W) a The case where 1 ≤ a ≤ k) >

[0340] In the absence of the aforementioned fragment (W) a The case of 1≤a≤k (W) a In the case of missing or incomplete data, (k-1) of the aforementioned fragments W are used first. i (1≤i≤k、i≠a) and the spare fragment (one) are used to calculate the fragment (W) a At this point, we use equations (41) to (43). Here, the upper part of equation (41) is equivalent to equation (23), and the lower part is equivalent to equation (24).

[0341] [Number 41]

[0342]

[0343] [Number 42]

[0344]

[0345] [Number 43]

[0346]

[0347] In this case, to shorten the calculation time, it is preferable to calculate "W" in advance, which is consistent with equation (41). i W k+1 W k+2 The values ​​of the relevant coefficients are set into an array. For example, with "W" k+1 W k+2 "The correlation coefficients are determined by k and a, therefore, they are initially set to a two-dimensional array with k and a as array elements. Additionally, regarding "W..." i "The correlation coefficients are determined by k, a, and i, and are therefore initially set as a three-dimensional array with these as elements. Then, the slices (W) obtained through equations (41) to (43) are used..." a ) and fragment W i (1≤i≤k,i≠a), the scattered object data is decoded by equation (40).

[0348] <Does not exist (W) a W b The case where 1 ≤ a ≤ b ≤ k) >

[0349] In the absence of fragmentation (W) a W b In the case of 1≤a≤b≤k (missing / incomplete cases), first use (k-2) fragments W. i (1≤i≤k, i≠a, i≠b) and two spare partitions are used to calculate the partition (W). a W b Here, we use equations (44) to (46). Here, equation (44) is equivalent to equation (27-1), and equation (45) is equivalent to equation (27-2).

[0350] [Number 44]

[0351]

[0352] [Number 45]

[0353]

[0354] [Number 46]

[0355]

[0356] In this case, in order to shorten the calculation time, it is preferable to calculate "W" in advance, which is consistent with equations (44) and (45). i W k+1 W k+2 The values ​​of the relevant coefficients are set into an array. For example, with "W" k+1 W k+2 The correlation coefficients are determined by k, a, and b, and are therefore initially set as a three-dimensional array with these as elements. Additionally, regarding "W..." i "The correlation coefficients are determined by k, a, b, and i, and are therefore initially set as a four-dimensional array with these as elements. Then, the slices (W) obtained through equations (44)~(46) are used..." a W b ) and fragment W i (1≤i≤k,i≠a,i≠b), the scattered object data is decoded by equation (40).

[0357] S22: Perform the inverse operation of S04, which involves independent uniform distribution, on the data string decoded in S21. Specifically,

[0358] (a) Using the inverse matrix B of matrix A, starting from the beginning of the data string, sequentially in "modulo 2"... 64 Perform matrix B operations (the inverse operation of S14) on the matrix.

[0359] (b) In the data string after the operation of the inverse matrix B, the high 32 bits and low 32 bits are reversed between adjacent data (the inverse operation of S13).

[0360] (c) For the data string after the inverse permutation, use the inverse matrix B, sequentially from back to front in the "modulo 2" matrix. 64 Perform the multiplication operation of the inverse matrix B on top (the inverse operation of S12).

[0361] S23: Extract the scattered object data from the data string generated after processing in S22 and calculate the hash value (SHA-256 of OpenSSL). Compare the hash value calculated here with the hash value set to "CS" for the reconstruction function to confirm whether the reconstructed scattered object data is correct.

[0362] That is, if the hash values ​​of the two objects are the same, the decoded scatter object data is set into a one-dimensional array "Si", and "0" is returned as the return value. On the other hand, if the hash values ​​of the two objects are different, the reconstruction failed correctly, so a value other than "0" is returned as the return value. Figure 6 The example shows a return value of "6", but it can also be any other value. In this case, it is preferable to configure the system to attempt recovery again.

[0363] In this embodiment, by employing "AMSSS," one or two spare shards can be created while leveraging the high speed of "MSSS," allowing the spare shards to be used to reconstruct distributed object data. This enables a balance between availability and confidentiality in high-capacity backup systems such as data centers.

[0364] Example 2

[0365] In Example 1, the determination of whether the group of fragments satisfies the reconstruction dispersion number was performed through two-stage steps. However, in this example, the second step is omitted, and the determination is performed only through the first step.

[0366] Therefore, it is not necessary to calculate the hash value of each data segment in S06 of this embodiment, nor is it necessary to set the hash value of each data segment in the distributed data generation and reconstruction functions.

[0367] Specifically,

[0368] (1) If all fragments “W1, W2, ..., W” can be collected from storage devices 3-1 to 3-n without timeout, k If the value is "", then it is determined that the reconstruction dispersion number is satisfied. In this case, the same equations (39) and (40) are used as in Example 1 to utilize only the fragmented data "W1, W2, ..., W". k "Decode."

[0369] However, if the fragments "W1, W2, ..., W" are used in decoding k If any fragment of the data has defects due to data corruption, etc., then in S23, a value other than "0" is returned as the return value to confirm whether the recovery was successful, and an attempt is made to recover it again. At this time, if "distribution variable = 1", then the fragmented data is replaced with the spare fragment "W" according to the order of the distribution number. k+1 Decoding is performed using equations (39) and (40). Additionally, if "distribution variable = 2", then every two fragments of data are replaced with spare fragments "W" according to the scatter numbering order. k+1 “W” k+2 ", and decoded by equations (39) and (40).

[0370] (2) On the other hand, if the fragments are “W1, W2, ... W k If the collection of any fragment in the array times out, it is determined that the reconstruction scattering number is not satisfied.

[0371] In this case, the same process as in Example 1 is performed to recover the scattered object data:

[0372] <Does not exist (W) a The case where 1 ≤ a ≤ k) >

[0373] <Does not exist (W) a The case of 1≤a≤k)>.

[0374] Furthermore, the present invention is not limited to the above-described embodiments, and can be implemented in variations within the scope of the claims. For example, it can be configured as a secret sharing program for causing a computer to execute the processing steps (S01-S06, S11-S14, S21-S23) of the secret sharing method of the embodiments.

[0375] According to this program, a secret sharing system 1 can be built on a computer. In addition, the program can be distributed in the form of being stored on a recording medium or downloaded from the Internet.

[0376] Explanation of reference numerals in the attached figures

[0377] 1: Secret Information System; 2: Storage Server Device; 3-1~3-n: Storage Device; 4: Network; 14: Distributed Data Generation Unit; 15: Data Recovery Unit; 16: Random Number Generation Unit; 21: Communication Unit; 22: Storage Unit.

Claims

1. A secret sharing method, which is a method executed by a secret sharing system, characterized in that, The secret sharing system has the following features: The distributed data generation unit generates groups of fragments obtained by encoding and distributing the distributed object data; Multiple storage units, each of which stores one of the fragments; as well as The data recovery unit collects the fragments stored in each of the storage units and reconstructs the fragmented object data by decoding them based on a pre-defined number of reconstruction fragments. The secret sharing method includes the following steps: In the distributed data generation step, the distributed data generation unit generates multiple groups of fragments corresponding to the reconstructed distributed number based on the distributed object data, as well as a spare fragment obtained by redundancy of the group of fragments, and distributes the generated fragments and the spare fragments to the storage unit for storage. as well as In the reconstruction step, the data recovery unit collects the fragments and the spare fragments from the storage unit. If the group of the collected fragments satisfies the reconstruction dispersion number, the group of fragments is used to reconstruct the scattered object data. On the other hand, if the collected fragments do not satisfy the reconstruction dispersion number, the existing fragments and the spare fragments are used to find the non-existent fragments to reconstruct the scattered object data.

2. The secret sharing method according to claim 1, characterized in that, The distributed data generation step includes the following steps: calculating the error detection code value for each generated fragment and setting it in the data recovery unit. The reconstruction step includes the following steps: determining whether each of the fragments exists by calculating the error detection code value for each of the collected fragments and comparing it with the set value.

3. The secret sharing method according to claim 1, characterized in that, The reconstruction step includes the following steps: If the number of collected fragments satisfies the reconstructed dispersion number, then the group of collected fragments is used to reconstruct the dispersed object data; and If the reconstruction using the collected groups of fragments fails, the reconstruction is performed by replacing the fragments with the spare fragments in the order of their assigned numbers.

4. A secret sharing method, which is a method executed by a secret sharing system, characterized in that: The secret sharing system has the following features: The distributed data generation unit generates groups of fragments obtained by encoding and distributing the distributed object data; Multiple storage units, each of which stores one of the fragments; as well as The data recovery unit collects the fragments stored in each of the storage units and reconstructs the fragmented object data by decoding them based on a pre-defined number of reconstruction fragments. The secret sharing method includes the following steps: The distributed data generation step involves the distributed data generation unit generating multiple groups of fragments corresponding to the reconstructed distributed data based on the distributed object data, as well as redundant spare fragments obtained by redundancy of these groups. The generated fragments and the spare fragments are then distributed and stored in the storage unit, wherein the spare fragments are one or two. In the reconstruction step, the data recovery unit collects the fragments and the spare fragments from the storage unit. If the group of the collected fragments meets the reconstruction dispersion number, the group of fragments is used to reconstruct the scattered object data. On the other hand, if the collected fragments do not meet the reconstruction dispersion number, the spare fragments are used to reconstruct the scattered object data. The distributed data generation step includes the following steps: The first step is to use the distributed data function, which stores the number of reconstructed shards k, the number of spare shards r, the data length L of the sharded object data, and the information of the sharded object data, to read in the information. The second step involves generating random numbers and creating a data string S containing the generated random numbers, the data length L, and the scattered object data, wherein the number of bytes in the data string S is a multiple of "k×8"; and The third step involves dividing the data string S into "k×8" blocks, and performing operations on each block to determine the fragment W. In the third step, when the number of spare fragments r = 1, equations (28) and (29) are used. , S xy The x-th value of block number y, where x = 1, 2, 3…k, is 8 bytes. W zy The value of the fragment with block number y and scatter number z, where z = 1, 2, 3…k, is 8 bytes. W (k+1)y The spare fragment, , On the other hand, when the number of spare fragments r=2, equations (30) and (31) are used. , , S xy The x-th value of block number y, where x = 1, 2, 3…k, is 8 bytes. W zy The value of the fragment with block number y and scatter number z, where z = 1, 2, 3…k, is 8 bytes. W (k+1)y W (k+2)y The spare fragment.

5. A secret sharing system, characterized in that, have: The distributed data generation unit generates groups of fragments obtained by encoding and distributing the distributed object data; Multiple storage units, each of which stores one of the fragments; as well as The data recovery unit collects the fragments stored in each of the storage units and reconstructs the fragmented object data by decoding them based on a pre-defined number of reconstruction fragments. The distributed data generation unit generates multiple groups of fragments corresponding to the reconstructed distributed number based on the distributed object data, as well as spare fragments obtained by redundancy of the groups of fragments. The generated fragments and the spare fragments are then distributed to the storage unit and stored. The data recovery unit collects the fragments and the spare fragments from the storage unit. If the collected groups of the fragments satisfy the reconstructed dispersion number, then the groups of the fragments are used to reconstruct the dispersed object data. On the other hand, if the collected fragments do not meet the reconstructed scatter number, the existing fragments and the spare fragments are used to find the non-existent fragments to reconstruct the scatter object data.

6. A secret sharing program that causes a computer to execute the secret sharing method according to any one of claims 1 to 4.