Archive storage method and system based on block chain
By dynamically evaluating archive access records and life cycle information, calculating dynamic access weight values, and optimizing data distribution strategies, the problem of insufficient dynamic evaluation of archive access status and life cycle in the existing technology is solved, and more efficient data distribution and storage resource utilization is achieved.
Patent Information
- Application Number
- CN202510251931.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing technology lacks dynamic assessment of the archive access status and life cycle, resulting in the failure of data distribution to accurately match access requirements, uneven resource utilization, and reducing storage efficiency and availability.
By analyzing archive access records and life cycle information, dynamic access weight values are calculated, and sorting and grouping based on the weight values, combining the capacity allocation rules of the main chain and the side chain, the data distribution strategy is optimized. Dynamically evaluate the archive cooling status, group and migrate accordingly, adjust the index path node structure, and optimize the storage path integrity and access efficiency.
It realizes dynamic evaluation of data access status, optimizes data distribution strategy, alleviates the main chain storage pressure, improves side chain resource utilization, reduces access delay caused by link changes, and improves storage efficiency and reliability.
Smart Images

Figure CN120123299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed storage, and particularly to an archival storage method and system based on blockchain. Background Art
[0002] The technical field of distributed storage includes storage technologies centered around data distribution and management. Its main objective is to achieve efficient storage and retrieval of data through multi-node collaboration, ensuring data reliability, consistency, and availability. Distributed storage technology is based on the diversity of nodes and the connectivity of the network, involving data sharding storage, replica synchronization, distributed consistency protocols, and data integrity verification. This technical field encompasses everything from the underlying architecture design of data storage to the optimization of specific management mechanisms, including distributed file systems, distributed databases, and related data storage protocols. By distributing data across different physical nodes, distributed storage technology can effectively address single-point failures and demonstrate its advantages in large-scale data management.
[0003] Among them, the archival storage method based on blockchain refers to an archival storage mechanism designed using the characteristics of blockchain technology. The main focus of this patent theme is on the trustworthy storage, traceability management, and anti-tampering requirements of archival data. By recording archival storage operations in the blockchain network and combining the immutability of the distributed ledger, secure storage of archival data is achieved. Specifically, after the archival data is block-processed, its storage location, operation records, and related index information are recorded in the blockchain. At the same time, a consensus mechanism is used to ensure data consistency, and cryptographic means are used to encrypt the archival data to protect data privacy. Throughout the process, it is supported by the underlying architecture of distributed storage, and the automated management of archival storage operations is achieved through smart contracts.
[0004] The prior art lacks dynamic assessment of archival access status and lifecycle, and data distribution fails to accurately match access requirements, often resulting in resource competition between high-frequency and low-frequency archival resources and reducing storage efficiency. The distribution rules of the main chain and side chains are fixed and difficult to adjust according to access requirements, easily leading to uneven resource utilization. The judgment criteria for cold data are imperfect, resulting in unoptimized resource occupancy and reducing the availability of high-frequency archival data on the main chain. During the migration process, the adjustment of the index path is insufficient, and link changes may cause path confusion and access delays, affecting the overall system response performance. Summary of the Invention
[0005] The objective of the present invention is to address the drawbacks existing in the prior art and propose an archival storage method and system based on blockchain.
[0006] To achieve the above objective, the present invention adopts the following technical solutions: An archival storage method based on blockchain, comprising the following steps: S1: Based on the file access records and lifecycle information, extract the access time interval and access times of the file, perform cumulative summation on the time interval data and calculate the weight value, add the weight value to the creation time and active stage of the file, perform normalization processing on the summation result, and generate the dynamic access weight value of the file; S2: Based on the dynamic access weight value of the file, sort according to the weight value, combine the capacity allocation rules of the main chain and the side chain, perform sorting and grouping based on the weight value, match the sorting and grouping result with the storage capacity, allocate the matched file to the main chain or the side chain, and establish a file distributed storage allocation table; S3: Based on the file distributed storage allocation table, extract the access time data of the files in the main chain, calculate the difference between the access interval time and the current time, compare the difference with the cumulative value of the initial time weight of the file, perform judgment marking on the cooling state of the file, and generate the cooling mark value of the main chain file; S4: Based on the cooling mark value of the main chain file, extract the cooled marked files, group the files according to the access interval and storage requirements, calculate the allocation ratio according to the available capacity of the side chain for the grouping result, select the migration order according to the allocation ratio, and update the state of the migrated files to the side chain storage to generate a file migration grouping table; S5: Based on the file migration grouping table, extract the main chain and side chain index path data of the migrated files, calculate the change value of the index path node relationship, adjust the index node structure of the migrated files, synchronously write the adjusted index structure, and establish a main - side chain storage path index table.
[0007] The dynamic access weight value of the file includes the weight value, creation time, and active stage. The file distributed storage allocation table includes the sorting and grouping result, storage capacity allocation result, and allocation situation between the main chain and the side chain. The cooling mark value of the main chain file includes the access interval time, the difference between the current time and the cumulative value of the initial time weight, and the cooling state mark. The file migration grouping table includes the access interval, storage requirements, allocation ratio, migration order, and the state of the migrated files. The main - side chain storage path index table includes the main chain index path, side chain index path, change value of the index path node relationship, and the adjusted index structure of the migrated files.
[0008] As a further solution of the present invention, the specific steps for obtaining the dynamic access weight value of the file are as follows: S101: Extract the access time interval and access times of the file, call the access time data in the file record, calculate the time interval of each access and record it, accumulate the total time interval of each file, calculate the ratio of the accumulated result to the access times, and generate a preliminary time weight value set; S102: Call the preliminary time weight value set, combine with the creation time data of each file, calculate the total time weight from the file creation time to the current time, calculate the weight in segments according to the time range of the active stage, calculate the sum of weights for each segment, and generate the sum of weights for the active stage; S103: Call the sum of weights for the active stage and the preliminary time weight value set for weighted accumulation, using the formula: ; Calculate and generate the dynamic access weight value of the file; Among them, represents the dynamic access weight value of the file, represents the cumulative access time weight of the file, represents the adjusted weight of the file creation time, represents the sum of weights for the active stage, represents the current cumulative number of active time periods of the file, represents the total number of files.
[0009] As a further solution of the present invention, the obtaining step of the file distribution storage allocation table is specifically as follows: S201: Call the dynamic access weight value of the file, sort all files in descending order according to the weight value, determine the sorting order by comparing the weight values one by one, record the index and weight value information of the sorted file list, and generate the sorted file list; S202: Call the sorted file list, group the files based on the capacity allocation rules of the main chain and the side chain, accumulate the file weight values one by one and compare with the capacity threshold, and classify the files with the accumulated weight value not exceeding the capacity into the same group to generate the preliminary grouping result of the files; S203: Match the preliminary grouping result of the files with the actual available storage capacity data, using the formula: ; Calculate the optimal storage matching for each group of files and generate the file distribution storage allocation table; Among them, represents the storage matching value of the th group of files, represents the dynamic access weight of the th file, represents the available storage capacity of the th group, represents the number of files in the th group, represents the sum of the file weight values within the group.
[0010] As a further solution of the present invention, the obtaining step of the main chain file cooling mark value is specifically as follows: S301: Extract the file records in the main chain part of the file distribution storage allocation table, read the access time data corresponding to the main chain files, calculate the access interval time of each file, save the access interval time results of all files, and generate a set of access interval times for the main chain files; S302: Call the set of access interval times for the main chain files, calculate the time difference for each file based on the difference between the current time and the last access time of the file, and compare the time difference with the cumulative value of the initial time weights of the corresponding files to generate a comparison result of file time differences; S303: According to the comparison values of each file in the comparison result of file time differences, determine whether the time difference is greater than the cooling threshold, using the formula: ; Calculate and generate the cooling flag value for the main chain files; Among them, represents the cooling flag value, represents the th time difference of the file, represents the th time difference of the file, represents the sum of all file time differences, represents the average value of all file time differences, represents the cooling threshold.
[0011] As a further solution of the present invention, the steps for obtaining the file migration grouping table are specifically as follows: S401: Extract the files marked as in the cooling state in the cooling flag value of the main chain files, read the access interval and storage requirement data of the files, establish a grouping rule according to the access interval value and the size of the storage requirement, and perform preliminary grouping on the files that meet the grouping rule to generate a preliminary file migration candidate list; S402: Call the preliminary file migration candidate list, based on the dynamic monitoring data of the available capacity of the side chain, extract the total storage requirements of each file group, calculate the ratio of the storage requirements of each group to the side chain capacity, record the ratio value of each file group and generate the corresponding migration ratio to obtain the file migration ratio allocation table; S403: According to the ratio data in the file migration ratio allocation table, extract the migration ratio and side chain capacity matching parameters of the file group, using the formula: ; Calculate the migration priority values of multiple file groups, and update the migration results to the side chain storage status according to the priority to generate the file migration grouping table; Among them, represents the The migration priority value of the group file is used to represent the urgency and order of group migration. Represents the migration ratio of the file group, reflecting the proportion of migration requirements in the cooling state. Represents the available storage capacity corresponding to the side chain, providing a quantitative measure of the feasibility of migration matching. Is the weight factor of the file group, used to adjust the flexibility and dynamic response ability of the migration priority. Is an adjustment factor, used to ensure that the denominator is non-zero and enhance the stability of the calculation.
[0012] As a further solution of the present invention, the steps for obtaining the main side-chain storage path index table are specifically as follows: S501: Extract the main-chain index path and side-chain index path data of the migrated files in the file migration grouping table, call the node change records of the index path, calculate the change values of all nodes in the index path of the migrated files, determine the adjustment requirements of each migrated file in the index structure, and generate a change table of the index node relationship of the migrated files; S502: Call the change table of the index node relationship of the migrated files, adjust the main-chain and side-chain index structures of the migrated files according to the node change values, and generate an adjusted index structure of the migrated files by rearranging and adjusting the positional relationship of the index nodes; S503: Call the adjusted index structure of the migrated files, synchronously write the adjustment results to the main-chain and side-chain storage, calculate the balanced distribution value of the index nodes of the main-chain and side-chain, and use the formula: ; Generate the main side-chain storage path index table by adjusting the index node relationship of the main-chain and side-chain according to the calculated balanced distribution value; Among them, Represents the balanced distribution value of the main-chain and side-chain index nodes, Represents the Index capacity value of the Index capacity value of the Index capacity value of the Represents the total number of index nodes participating in the calculation in the main-chain and side-chain, Represents the absolute difference between the index capacity values of the corresponding nodes of the main-chain and side-chain.
[0013] A blockchain-based file storage system, which is used to execute the above-mentioned blockchain-based file storage method. The system includes: The file dynamic weight calculation module accumulatively calculates the time interval data based on the file access time interval and the number of access times, combines the file creation time and the active stage weight value for cumulative operation, and normalizes the cumulative operation result to generate a file dynamic access weight value; Based on the dynamic access weight values of the archives, the storage allocation optimization module sorts the weight values, groups the sorting results according to the capacity parameters of the main chain and the side chain, matches the grouping results with the available storage capacity, records the matching results, and generates an archive distributed storage allocation table; Based on the archive distributed storage allocation table, the main chain status monitoring module extracts the access time information of the main chain archives, calculates the difference between the access interval and the current time, compares it with the cumulative calculation result of the initial time weight value of the archives, marks the cooling status of the archives, and generates a main chain archive cooling mark value; Based on the main chain archive cooling mark value, the side chain migration management module filters the cooled marked archives, groups the archives according to the access interval and storage requirements, calculates the allocation ratio of the grouped archives to the available capacity of the side chain, adjusts the migration order according to the allocation ratio and updates the archive storage status, and generates an archive migration grouping table; Based on the archive migration grouping table, the index path adjustment module extracts the main chain and side chain index paths of the migrated archives, calculates the change value of the index path node relationship, adjusts the index node structure of the migrated archives, updates the adjusted index data, and generates a main - side chain storage path index table.
[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, by analyzing the archive access records and life - cycle information, calculating the dynamic access weight values and normalizing them, the dynamic evaluation of the data access status is realized. Combining the capacity allocation rules of the main chain and the side chain, the data distribution strategy is optimized based on weight - value grouping and storage matching. Dynamically evaluating the cooling status of the archives and grouping and migrating accordingly effectively alleviates the storage pressure of the main chain and improves the utilization rate of side - chain resources. During the migration process, the index path node structure is adjusted to optimize the integrity of the storage path and the access efficiency, and reduce the access delay caused by link changes. This dynamic and refined management strategy improves the storage efficiency and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic diagram of the working process of the present invention; Figure 2 is a flowchart of the steps for obtaining the dynamic access weight values of the archives of the present invention; Figure 3 is a flowchart of the steps for obtaining the archive distributed storage allocation table of the present invention; Figure 4 is a flowchart of the steps for obtaining the main chain archive cooling mark value of the present invention; Figure 5 is a flowchart of the steps for obtaining the archive migration grouping table of the present invention; Figure 6 is a flowchart of the steps for obtaining the main - side chain storage path index table of the present invention. Detailed Implementation Manner
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality" is two or more unless otherwise specifically defined.
[0018] Embodiment 1 Please refer to Figure 1 , the present invention provides a technical solution: a blockchain-based file storage method, including the following steps: S1: Based on the file access records and lifecycle information, extract the access time interval and access times of the file, perform cumulative summation on the time interval data and calculate the weight value, add the weight value to the creation time and active stage of the file, perform normalization processing on the accumulated result, and generate a file dynamic access weight value; S2: Based on the file dynamic access weight value, sort according to the weight value, combine the capacity allocation rules of the main chain and the side chain, perform sorting and grouping based on the weight value, match the sorting and grouping result with the storage capacity, allocate the matched file to the main chain or the side chain, and establish a file distribution storage allocation table; S3: Based on the file distribution storage allocation table, extract the access time data of the files in the main chain, calculate the difference between the access interval time and the current time, compare the difference with the cumulative value of the initial time weight of the file, and perform judgment and marking on the cooling state of the file to generate a main chain file cooling marking value; S4: Based on the main chain file cooling marking value, extract the cooled marked files, group the files according to the access interval and storage requirements, calculate the allocation ratio according to the available capacity of the side chain for the grouping result, select the migration order according to the allocation ratio, and update the state of the migrated files to the side chain storage to generate a file migration grouping table; S5: Based on the file migration grouping table, extract the main chain and side chain index path data of the migrated files, calculate the change value of the index path node relationship, adjust the index node structure of the migrated files, synchronously write the adjusted index structure, and establish a main side chain storage path index table.
[0019] The dynamic access weight value of the file includes the weight value, creation time, and active stage. The file distribution storage allocation table includes the sorting and grouping results, storage capacity allocation results, and the allocation situation of the main chain and side chains. The main chain file cooling mark value includes the access interval time, the difference between the current time and the cumulative value of the initial time weight, and the cooling state mark. The file migration grouping table includes the access interval, storage requirements, allocation ratio, migration order, and the status of the file after migration. The main and side chain storage path index table includes the main chain index path, side chain index path, the change value of the index path node relationship, and the adjusted index structure of the migrated file.
[0020] Please refer to Figure 2 , and the specific steps for obtaining the dynamic access weight value of the file are as follows: S101: Extract the access time interval and access times of the file, call the access time data in the file record, calculate the time interval of each access and record it, accumulate the total time interval of each file, calculate the ratio of the accumulated result to the access times, and generate a preliminary time weight value set; Calculate the difference between every two consecutive access time points in the access time sequence, record the difference as the time interval, sum up all the time intervals to obtain the cumulative access time interval, count the access times of each file and generate the total access times, calculate the ratio of the cumulative time interval to the access times, and calculate the preliminary time weight value through the specific formula where is the cumulative time interval of the th file, is its total access times, and the ratios of all files are summarized to form a preliminary time weight value set.
[0021] S102: Call the preliminary time weight value set, combine it with the creation time data of each file, calculate the total time weight from the file creation time to the current time, calculate the weight in segments according to the time range of the active stage, calculate the total weight of each segment, and generate the total weight of the active stage; Divide the time range from the file creation to the current time into multiple time periods by year, count the weight of each time period, filter and accumulate the weight values in each time period according to the phased characteristics of file access activity, and use the formula to calculate the total weight of the active stage of each file, where is the time weight of the th file in the st stage, is the total number of time periods, and generate the total weight of the active stage.
[0022] S103: Call the total weight of the active stage and the preliminary time weight value set for weighted accumulation, using the formula: ; Calculate and generate the dynamic access weight value of the file; Among them, represents the dynamic access weight value of the file, represents the cumulative access time weight of the file, represents the creation time adjustment weight of the file, represents the total weight of the active phase, represents the current cumulative number of active time periods of the file, represents the total number of files.
[0023] Formula: ; The advantage of the formula is that by combining the cumulative access time weight, the creation time adjustment weight, and the total weight of the active phase, and introducing the balance factor of the number of active time periods, the calculation accuracy of the dynamic access weight of the file is improved. Details of the formula and the derivation process of the formula calculation: 1. Call the set of preliminary time weight values, set , introduce the adjustment factor , and adjust the weight of each file; 2. Call the total weight of the active phase as , and substitute it into the formula together with and when calculating the dynamic access weight; 3. Use to represent the number of active time periods, and substitute the example values into the formula for calculating , , , in turn, ; Sum up all the files, and finally get the overall value of ; This result shows that the dynamic access weight value of the file generates the dynamic access weight value of the file by comprehensively considering the importance of time accumulation and the active phase to balance the dynamic access characteristics of the current file.
[0024] Please refer to Figure 3 , the specific steps for obtaining the file distribution storage allocation table are as follows: S201: Call the dynamic access weight value of the file, sort all the files in descending order according to the weight value, determine the sorting order by comparing the weight values one by one, record the index and weight value information of the sorted file list, and generate the sorted file list; Sort the file weight values in descending order. By comparing the weight values of two files one by one, if the weight value of the previous file is greater than that of the next file, swap their positions in the sorted list. This process is repeated until the weight values are sorted. At the same time, record the sorted file indexes and their corresponding weight values. To ensure the rationality and accuracy of the sorting, a practical scenario can be combined: for example, the weight value of file A is 35 and the weight value of file B is 30, and the file A with a higher weight value is ranked before file B. Through such logical sorting, a sorted file list is generated.
[0025] S202: Call the sorted file list, group the files based on the capacity allocation rules of the main chain and the side chain, accumulate the file weight values one by one and compare them with the capacity threshold, and classify the files whose accumulated weight values do not exceed the capacity into the same group to generate a preliminary grouping result of the files; Group the files according to the capacity allocation rules of the main chain and the side chain. By reading the sorted file weight values one by one, add them item by item to the total weight value of the current group. After each addition, compare it with the current storage capacity threshold. If the accumulated weight value is less than the threshold, continue to add the current file weight value to the group. If the accumulated weight value is equal to or exceeds the threshold, end the filing of the current group and switch to the next group for the same allocation operation. During this process, record the group numbers assigned to each file. For example: the capacity threshold of the main chain is 100, and the file weight values in the sorted file list are 35, 30, 25, and 20 in sequence. Assign files 1, 2, and 3 to the main chain, stop the allocation when the total weight reaches 90, and assign the remaining files to the side chain to generate a preliminary grouping result of the files.
[0026] S203: Match the preliminary grouping result of the files with the actual available storage capacity data, and use the formula: ; Calculate the optimal storage matching of each group of files to generate a file distribution storage allocation table; Among them, represents the storage matching value of the th group of files, represents the dynamic access weight of the rd file, represents the available storage capacity of the th group, represents the number of files in the th group, represents the sum of the file weight values within the group.
[0027] Formula: ; The advantage of the formula is that by comprehensively considering the weight value of the file group and the limitation of the storage capacity, the resource matching after file grouping is more accurate, optimizing the storage allocation efficiency.
[0028] Detailed explanation of the formula and the derivation process of formula calculation: Represents the storage matching value of the th group of files, Represents the dynamic access weight of the th file, Represents the available storage capacity of the th group, Represents the number of files in the th group. The summation of the weight values of all files in the file group is expressed as follows: The specific calculation is as follows: Assume that the first group of files contains files A, B, and C with weight values of 35, 30, and 25 respectively, and the storage capacity is 100. The calculation process is as follows: 1. Calculate the total weight value of the file group: ; 2. Compare with the storage capacity: ; This result indicates that the total weight value of file group 1 matches the storage capacity requirement and does not exceed the capacity. Files A, B, and C are successfully allocated to the first group. At the same time, this matching process generates a file distribution storage allocation table by ensuring the precise correspondence between weight and capacity.
[0029] Please refer to Figure 4 , and the specific steps for obtaining the cooling mark value of the main chain file are as follows: S301: Extract the file records of the main chain part in the file distribution storage allocation table, read the access time data corresponding to the main chain files, calculate the access interval time of each file, save the results of the access interval time of all files, and generate a set of access interval times for the main chain files; First, call the access timestamp of each file record, calculate the time interval between two adjacent accesses through the timestamp. The formula for the interval is: , where represents the time interval, is the current timestamp, is the previous access timestamp. Through batch operations on multiple access time records of the main chain files, obtain the access time interval sequence corresponding to each file. Further organize the results of the access time interval for subsequent calls, arrange them in chronological order, and at the same time pair the file identifier with the interval time for storage, record it as the access time interval information table, and generate a set of access interval times for the main chain files.
[0030] S302: Invoke the main-chain file access interval time set, calculate the time difference for each file based on the difference between the current time and the last access time of the file, and compare the time difference with the initial time weight cumulative value of the corresponding file to generate a file time difference comparison result; First, extract the last access timestamp and the current timestamp of each file record, obtain the time difference using the subtraction operation, and at the same time, invoke the initial time weight cumulative value of the file, and perform the comparison through the following calculation formula: , where represents the time difference of the file, is the current timestamp, is the last access timestamp, pair and sort the calculated time difference with the initial weight cumulative value of the file. If the time difference is greater than the weight cumulative value, mark it as the cooling state. At the same time, generate a cooling state marking table, associate the result of the marking table with the main-chain file identifier, and store and sort it as the file time difference comparison result.
[0031] S303: According to the comparison value of each file in the file time difference comparison result, determine whether the time difference is greater than the cooling threshold, using the formula: ; Calculate and generate the main-chain file cooling marking value; where represents the cooling marking value, represents the th file's time difference, represents the th file's time difference, represents the sum of all file time differences, represents the average value of all file time differences, represents the cooling threshold.
[0032] Formula: ; Detailed explanation of the formula and the derivation process of the formula calculation: is the cooling marking value, used to characterize whether the file time difference meets the requirements of the cooling threshold; is the th file's time difference, obtained through data collection of the file's timestamp. For example, the file timestamps monitored from a certain data platform are 10:00 on January 15, 2025, and 10:05 on January 15, 2025, and the time difference is 300 seconds; is the th file's time difference, and the calculation method is the same as , for example, the next time difference is monitored to be 400 seconds; is the sum of all file time differences. For example, there are 5 files with time differences of 300 seconds, 400 seconds, 350 seconds, 500 seconds, and 450 seconds respectively, and the sum is 2000 seconds; is the average value of all file time differences, calculated by dividing the total time difference by the number of files. For example seconds; is the cooling threshold, determined by system settings or technical parameters in relevant documents. For example, the cooling threshold is 450 seconds; Substitute the above parameters into the formula: ; The first step is to calculate the absolute value of the numerator: ; The second step is to calculate the sum of the denominators: ; The third step is to calculate the time difference ratio and take the square root: ; The fourth step is to calculate the absolute value of the cooling threshold difference: ; The fifth step is to calculate the cooling flag value: ; Result analysis: The result shows that the cooling flag value is 0.004472. The lower the value, the smaller the deviation between the time difference and the cooling threshold. The time differences of files that meet the cooling conditions have higher adaptability to generating the cooling flag value of the main-chain file. Combining all cooling flag values can further generate a file cooling flag value sequence or other statistical results.
[0033] Please refer to Figure 5 , and the specific steps for obtaining the file migration grouping table are as follows: S401: Extract the files marked as in the cooling state from the cooling flag values of the main-chain files, read the access intervals and storage requirement data of the files, establish a grouping rule according to the access interval values and storage requirement sizes, and perform preliminary grouping on the files that meet the grouping rule to generate a preliminary file migration candidate list; According to the file access intervals and storage requirements, call each file one by one, extract the access interval records of each file, screen and record the files with an access interval period higher than the set threshold based on the access interval numerical size, and at the same time analyze the storage requirements of the files, calculate their storage priorities according to the file capacity and file importance, and use the storage requirement ratio formula Standardize the storage requirements for all files, where is the storage requirement of the file, is the total number of files participating in the grouping. Combining the standardization results and the access interval screening results, perform the grouping operation, assign serial numbers to each group and record the archiving, and complete the grouping of the files that meet the conditions to generate a preliminary file migration candidate list.
[0034] S402: Invoke the preliminary file migration candidate list, based on the dynamic monitoring data of the available capacity of the side chain, extract the total storage requirements of each file group, calculate the ratio of the storage requirements of each group to the side chain capacity, record the ratio value of each file group and generate the corresponding migration ratio to obtain the file migration ratio allocation table; Extract the total storage requirements of the grouped files, calculate the sum of the storage requirements of each group one by one, and use the formula For the group, calculate the cumulative sum of the total storage requirements of the files, where is the total storage requirement of the group, is the storage requirement of a single file, is the group number of the group. Compare the cumulative total requirements with the real-time available capacity of the side chain one by one, extract the ratio of the storage requirements of each group to the total side chain capacity, and the calculation formula is where is the ratio of the storage requirements of the group, is the real-time available capacity of the side chain. Combine the calculated ratio
[0035] S403: According to the ratio data in the file migration ratio allocation table, extract the migration ratio and side chain capacity matching parameters of the file group, and use the formula: ; Calculate the migration priority values of multiple file groups, and update the migration results to the side chain storage status according to the priority to generate the file migration grouping table; Among them, represents the migration priority value of the group of files, which is used to represent the urgency and order of the grouping migration, represents the migration ratio of the file group, which reflects the ratio of the migration requirements in the cooling state, represents the available storage capacity corresponding to the side chain, which provides a quantitative measure of the feasibility of the migration match, is the weight factor of the file group, which is used to adjust the flexibility and dynamic response ability of the migration priority, is an adjustment factor used to ensure that the denominator is non-zero and enhance the stability of calculations.
[0036] Formula: ; The advantage of the formula is that by combining the migration ratio and the available capacity of the side chain for differential calculation, and adding the weight factor and the adjustment factor , it realizes flexible adjustment of the migration priority of files, improving the accuracy and stability of the allocation process.
[0037] Detailed explanation of the formula and the derivation process of formula calculation: Let the migration ratio of the first group of files be , and the corresponding available capacity of the side chain be , the weight factor , the adjustment factor , substitute the parameter values into the formula: ; Calculate the absolute value part: ; Calculate the square root part of the denominator: ; Calculate the final result: ; This result indicates that the migration priority of the first group of files is 0.307. After comparing with the priorities of other groups, the migration order can be determined, and the group with a higher priority will perform the migration operation first to generate a file migration grouping table.
[0038] Please refer to Figure 6 , the specific steps for obtaining the main and side chain storage path index table are as follows: S501: Extract the main chain index path and side chain index path data of the migrated files in the file migration grouping table, call the node change records of the index path, calculate the change values of all nodes in the index path of the migrated files, determine the adjustment requirements of each migrated file in the index structure, and generate a migration file index node relationship change table; By calling the migration records of each file in the file migration grouping table, the main chain index path data of the file is read one by one. The main chain index path represents the node position and association relationship of the file in the main chain storage. The uniqueness of each node position needs to be verified through the integrity of its association relationship. The side chain index path records the target storage node after file migration. Each file corresponds to a target side chain index path value. According to the uniqueness of the migration records, the main chain and side chain path records are compared, and the change value of the path node is calculated. The change value of the path node uses the absolute value of the difference between the number of corresponding nodes in the index path before and after migration as the calculation standard. The absolute value can avoid the interference of positive and negative values on the accuracy of data statistics and at the same time reflect the actual degree of path change. Based on this, an index path node change table for each file is established, and the change values of the paths of each file are counted one by one. The change values are associated with the storage capacity data of the index path to obtain the distribution of the change values of the nodes. According to the proportional relationship between the change value of the node and the storage capacity, the adjustment requirements of each migrated file in the index structure are determined, and a change table of the index node relationship of the migrated file is generated.
[0039] S502: Call the change table of the index node relationship of the migrated file, adjust the main chain and side chain index structures of the migrated file according to the change value of the node. By rearranging and adjusting the positional relationship of the index nodes, an adjusted index structure of the migrated file is generated; For files with relatively large change values of each node, the index node adjustment is preferentially performed. A relatively large change value of the node indicates a relatively large change amplitude of the path node brought by file migration. During the adjustment process, the positional relationship and association relationship of the nodes in the main chain and side chain need to be reconstructed in sequence. The adjustment of the main chain index ensures the continuity of data access by rearranging the order of the path nodes. The adjustment of the side chain index ensures the integrity of data storage by reconstructing the association relationship in the storage path. The rearrangement of the node order needs to use normalization calculation to standardize the original order. The index table is re-established for the mapping relationship between the normalized node order and the storage path. For the reconstruction of the association relationship, the migrated path data needs to be compared item by item with the original storage path data, identify discontinuous data segments and insert correction points to fill the breaks in the data segments, and obtain the adjusted index structure of the migrated file.
[0040] S503: Call the adjusted index structure of the migrated file, synchronously write the adjustment result to the main chain and side chain storage, calculate the balanced distribution value of the index nodes of the main chain and side chain, using the formula: ; Adjust the index node relationship of the main chain and side chain through the calculated balanced distribution value, and generate a main and side chain storage path index table; Among them, represents the balanced distribution value of the index nodes of the main chain and side chain, represents the index capacity value of the th node in the main chain, Represents the index capacity value of the th node in the side chain, represents the total number of index nodes participating in the calculation in the main chain and the side chain, represents the absolute difference between the index capacity values of the corresponding nodes in the main chain and the side chain.
[0041] Formula: ; Detailed explanation of the formula and the derivation process of the formula calculation: The balanced distribution value of the index nodes in the main chain and the side chain is calculated by the formula. The specific explanations and calculation methods of each parameter in the formula are as follows: Represents the index capacity value of the th node in the main chain. This value is obtained by real-time monitoring of the data storage volume of the main chain nodes, and the unit is byte. Example: The index capacity value of a certain node in the main chain is determined to be 500 bytes by collecting the data storage volume.
[0042] Represents the index capacity value of the th node in the side chain. This value is obtained by real-time monitoring of the data storage volume of the side chain nodes, and the unit is byte. Example: The index capacity value of the corresponding node in the side chain is determined to be 300 bytes by collecting the data storage volume.
[0043] Represents the total number of index nodes participating in the calculation in the main chain and the side chain, which is obtained by querying the actual number of nodes participating in the calculation in the system network. For example, the system currently records that the total number of nodes participating in the calculation in the main chain and the side chain is 10 nodes.
[0044] Substitute specific numbers for calculation. Assume that the index capacity data of the main chain and the side chain are as follows: Index capacity values of main chain nodes: 500, 600, 450, 550, 500, 620, 480, 500, 510, 495 (unit: byte); Index capacity values of side chain nodes: 300, 450, 400, 420, 350, 470, 400, 310, 390, 380 (unit: byte).
[0045] First, calculate the absolute difference between the index capacity values of each pair of nodes: |500 - 300| = 200, |600 - 450| = 150, |450 - 400| = 50, |550 - 420| = 130, |500 - 350| = 150, |620 - 470| = 150, |480 - 400| = 80, |500 - 310| = 190, |510 - 390| = 120, |495 - 380| = 115; Calculate the square of the difference in the index capacity values for each pair of nodes: ; Sum all the squared results: ; Substitute the result into the formula to calculate the balanced distribution value: ; Result description: The balanced distribution value is 140.35, indicating that the distribution difference between the main chain and the side chain node index capacities fluctuates within a certain range. The larger the value, the higher the degree of uneven distribution.
[0046] A blockchain-based archival storage system, which is used to execute the above-mentioned blockchain-based archival storage method. The system includes: The archival dynamic weight calculation module accumulatively calculates the time interval data based on the archival access time interval and the number of accesses, performs an accumulative operation in combination with the archival creation time and the active phase weight value, and normalizes the accumulative operation result to generate an archival dynamic access weight value; The storage allocation optimization module sorts the weight values based on the archival dynamic access weight value, groups the sorting results according to the capacity parameters of the main chain and the side chain, matches the grouping results with the available storage capacity, records the matching results, and generates an archival distribution storage allocation table; The main chain status monitoring module extracts the access time information of the main chain archives based on the archival distribution storage allocation table, calculates the difference between the access interval and the current time, compares it with the accumulative calculation result of the archival initial time weight value, marks the archival cooling status, and generates a main chain archival cooling mark value; The side chain migration management module filters the cooled marked archives based on the main chain archival cooling mark value, groups the archives according to the access interval and the storage requirements, calculates the allocation ratio of the grouped archives to the available capacity of the side chain, adjusts the migration order according to the allocation ratio and updates the archival storage status, and generates an archival migration grouping table; The index path adjustment module extracts the main chain and side chain index paths of the migrated archives based on the archival migration grouping table, calculates the change value of the index path node relationship, adjusts the index node structure of the migrated archives, updates the adjusted index data, and generates a main-side chain storage path index table.
[0047] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. The archive storage method based on blockchain is characterized by: The following steps are involved: S1: Based on the archive access records and life cycle information, extract the access time interval and access times of the archive, accumulate and sum the time interval data and calculate the weight value, accumulate the weight value with the archive creation time and active stage, perform normalization on the accumulated results, and generate the archive dynamic access weight value; S2: Based on the dynamic access weight value of the archive, sort according to the weight value, combine the capacity allocation rules of the main chain and the side chain, sort and group based on the weight value, match the sorting and grouping results with the storage capacity, allocate the matched archives to the main chain or side chain, and establish an archive distribution storage allocation table; S3: Based on the archive distribution storage allocation table, extract the access time data of the archive in the main chain, calculate the difference between the access interval time and the current time, compare the difference with the archive initial time weight accumulation value, perform a judgment mark on the archive cooling state, and generate a main chain archive cooling mark value; S4: based on the cooling mark value of the main chain archive, extract the cooling mark archive, group the archives according to the access interval and storage demand, calculate the allocation ratio of the grouping results according to the available capacity of the side chain, select the migration order according to the allocation ratio, update the archive status after migration to the side chain storage, and generate an archive migration grouping table; S5: Based on the archive migration grouping table, extract the main chain and side chain index path data of the migration archive, calculate the index path node relationship change value, adjust the migration archive index node structure, synchronously write the adjusted index structure, and establish the main and side chain storage path index table.
2. The blockchain-based archive storage method according to claim 1, characterized in that: The dynamic access weight value of the archive includes the weight value, creation time, and active stage. The archive distribution storage allocation table includes the sorting grouping results, storage capacity allocation results, and the allocation of the main chain and the side chain. The main chain archive cooling mark value includes the access interval time, the difference between the current time and the initial time weight cumulative value, and the cooling status mark. The archive migration grouping table includes the access interval, storage demand, allocation ratio, migration order, and archive status after migration. The main and side chain storage path index table includes the main chain index path of the migrated archive, the side chain index path, the index path node relationship change value, and the adjusted index structure.
3. The blockchain-based archive storage method according to claim 2, characterized in that: The steps for obtaining the dynamic access weight value of the archive are specifically as follows: S101: extracting the access time interval and access times of the archive, calling the access time data in the archive record, calculating and recording the time interval of each access, accumulating the total time interval of each archive, calculating the ratio of the accumulated result to the access times, and generating a preliminary time weight value set; S102: calling the preliminary time weight value set, combining the creation time data of each archive, calculating the total time weight from the archive creation time to the current time, calculating the weight in segments according to the time range of the active stage, calculating the sum of the weights of each segment, and generating the sum of the active stage weights; S103: Call the active stage weight sum and the preliminary time weight value set for weighted accumulation, using the formula: ; Calculate and generate dynamic access weight value of archives; in, Represents the dynamic access weight value of the archive. Represents the cumulative access time weight of the archive, Represents the creation time adjustment weight of the archive, Represents the total weight of the active phase, Represents the number of active time periods currently accumulated in the archive. Represents the total number of files.
4. The blockchain-based archive storage method according to claim 3 is characterized in that: The steps for obtaining the archive distribution storage allocation table are specifically as follows: S201: calling the dynamic file access weight value, sorting all files in descending order according to the weight value, determining the sorting order by comparing the weight values one by one, recording the sorted file list index and weight value information, and generating a sorted file list; S202: calling the sorted archive list, grouping the archives based on the capacity allocation rules of the main chain and the side chain, accumulating the archive weights one by one and comparing them with the capacity threshold, grouping the archives whose weight accumulation values do not exceed the capacity into the same group, and generating the archive pre-grouping result; S203: Match the pre-grouping result of the archive with the actual available storage capacity data using the formula: ; Calculate the optimal storage match for each group of archives and generate an archive distribution storage allocation table; in, Representative The storage match value of the group profile, Representative Dynamic access weights for individual files, Representative The available storage capacity of the group, Representative in the The number of files in the group, Indicates the sum of the weight values of the files in the group.
5. The blockchain-based archive storage method according to claim 4 is characterized in that: The steps for obtaining the cooling mark value of the main chain archive are specifically as follows: S301: extract the archive records of the main chain part in the archive distribution storage allocation table, read the access time data corresponding to the archives in the main chain, calculate the access interval time of each archive, save the access interval time results of all archives, and generate the access interval time set of the archives in the main chain; S302: calling the main chain archive access interval time set, calculating the time difference of each archive based on the difference between the current time and the last access time of the archive, and comparing the time difference with the initial time weight accumulation value of the corresponding archive to generate an archive time difference comparison result; S303: According to the comparison value of each file in the file time difference comparison result, determine whether the time difference is greater than the cooling threshold, using the formula: ; Calculate and generate the cooling mark value of the main chain archive; in, Represents the cooling mark value, Representative The time difference of the files, Representative The time difference of the files, Represents the sum of the time differences of all files. Represents the average time difference of all archives, Represents the cooldown threshold.
6. The blockchain-based archive storage method according to claim 5, characterized in that: The steps for obtaining the file migration grouping table are specifically as follows: S401: extracting files marked as cooling status in the cooling mark value of the main chain files, reading the access interval and storage requirement data of the files, establishing a grouping rule according to the access interval value and the storage requirement size, performing preliminary grouping on the files that meet the grouping rule, and generating a preliminary file migration candidate list; S402: calling the preliminary archive migration candidate list, extracting the total storage demand of each archive group based on the dynamic monitoring data of the available capacity of the side chain, calculating the ratio of each group of storage demand to the side chain capacity, recording the ratio value of each archive group and generating the corresponding migration ratio, and obtaining the archive migration ratio allocation table; S403: According to the proportion data in the file migration ratio allocation table, the migration ratio of the file group and the side chain capacity matching parameter are extracted using the formula: ; Calculate the migration priority values of multiple archive groups, and update the migration results to the side chain storage status according to the priority, and generate an archive migration group table; in, Representative The migration priority value of the group archive is used to indicate the urgency and order of group migration. Represents the migration ratio of the file group, reflecting the proportion of migration demand in the cooling state. Represents the available storage capacity corresponding to the side chain, providing feasibility quantification of migration matching, It is the weight factor of the archive grouping, which is used to adjust the flexibility and dynamic response capability of the migration priority. is an adjustment factor used to ensure that the denominator is non-zero and enhance the stability of the calculation.
7. The blockchain-based archive storage method according to claim 6, characterized in that: The steps for obtaining the main-side chain storage path index table are specifically as follows: S501: extracting the main chain index path and side chain index path data of the migration file in the file migration grouping table, calling the node change record of the index path, calculating the change value of all nodes in the migration file index path, determining the adjustment requirement of each migration file in the index structure, and generating a migration file index node relationship change table; S502: calling the migration file index node relationship change table, adjusting the main chain and side chain index structure of the migration file according to the node change value, and generating an adjusted migration file index structure by rearranging and adjusting the position relationship of the index nodes; S503: Call the adjusted migration archive index structure, synchronously write the adjustment results to the main chain and side chain storage, and calculate the index node balance distribution value of the main chain and the side chain using the formula: ; The index node relationship between the main chain and the side chain is adjusted through the calculated balanced distribution value to generate the main chain and side chain storage path index table; in, Represents the balanced distribution value of the main chain and side chain index nodes, Represents the main chain The index capacity value of the node, Represents the side chain The index capacity value of the node, Represents the total number of index nodes participating in the calculation in the main chain and side chain. Indicates the absolute difference between the main chain and the side chain’s corresponding node index capacity values.
8. The blockchain-based archive storage system is characterized by: According to the blockchain-based archive storage method according to any one of claims 1 to 7, the system comprises: The archive dynamic weight calculation module performs cumulative calculation on the time interval data based on the archive access time interval and the number of accesses, performs cumulative calculation on the archive creation time and the active stage weight value, normalizes the cumulative calculation result, and generates the archive dynamic access weight value; The storage allocation optimization module sorts the weight values based on the dynamic access weight values of the archives, groups the sorting results according to the capacity parameters of the main chain and the side chain, matches the grouping results with the available storage capacity, records the matching results, and generates an archive distribution storage allocation table; The main chain status monitoring module extracts the access time information of the main chain archive based on the archive distribution storage allocation table, calculates the difference between the access interval and the current time, compares it with the cumulative calculation result of the archive initial time weight value, marks the archive cooling state, and generates the main chain archive cooling mark value; The side chain migration management module selects cooling mark archives based on the cooling mark value of the main chain archives, groups the archives according to the access interval and storage demand, calculates the allocation ratio of the grouped archives to the available capacity of the side chain, adjusts the migration order according to the allocation ratio and updates the archive storage status, and generates an archive migration grouping table; The index path adjustment module extracts the main chain and side chain index paths of the migration archive based on the archive migration grouping table, calculates the index path node relationship change value, adjusts the index node structure of the migration archive, updates the adjusted index data, and generates the main and side chain storage path index table.
Citation Information
Patent Citations
Cluster resource management method and device, electronic equipment and storage medium
CN112825023A
Archive management system based on block chain
CN115794739A
Archive development and utilization management system based on artificial intelligence technology
CN118427158A
System and method for data classification using machine learning during archiving
US20180373722A1
Usage Correction in a File System
US20230222097A1