Distributed storage and management method, system and equipment for super-large-scale medical and invasive data

By adopting distributed storage architecture, data compression and deduplication technology and intelligent hierarchical storage algorithms in the ultra-large-scale science and technology innovation data storage system, the shortcomings in traditional systems in terms of storage capacity, management complexity and scalability are solved, and efficient data storage and management are achieved.

CN120045144AActive Publication Date: 2025-05-27WUHAN STARLINK TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510535583.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Traditional data storage and management systems face problems such as insufficient storage capacity, high complexity of data management and poor system scalability when dealing with super-large-scale scientific and technological innovation data.

Method used

Adopting a distributed storage architecture, through data segmentation and redundant storage mechanisms, combining data compression and deduplication technology and intelligent hierarchical storage algorithms, we optimize storage space utilization and data access performance, and manage data consistency through a distributed consistency algorithm.

Benefits of technology

It improves storage efficiency and data access performance, enhances the dynamic adaptability and reliability of the system, optimizes storage space utilization and reduces storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045144A_ABST
    Figure CN120045144A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage and management, in particular to a distributed storage and management method, system and equipment for super-large-scale medical and invasive data. The method comprises the following steps: acquiring super-large-scale medical wound data based on a distributed storage architecture, segmenting the data into a plurality of small blocks, and distributing the small blocks to a plurality of storage nodes; performing compression and duplicate removal processing on the data in the storage nodes to generate an efficiently stored data sequence; based on the data sequence, carrying out automatic hierarchical storage on the data by adopting an intelligent algorithm; based on a hierarchical storage result, a node comprehensive load value is calculated by using a multi-factor collaborative optimization load balancing algorithm, and data blocks in high-load nodes are screened for preferential migration; a distributed storage architecture and an intelligent data management system are combined, and data consistency management is achieved through a distributed consistency algorithm. The reliability and expansibility of the system are enhanced, and the system is suitable for scenes such as science and technology research and development and big data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage and management, and particularly to a distributed storage and management method, system and device for ultra-large-scale scientific and technological innovation data. Background Art

[0002] With the rapid development of scientific and technological innovation activities, the scale of various types of data generated in the field of scientific and technological innovation has increased exponentially, including experimental data, simulation data, multi-modal sensor data, and historical analysis data, etc. These data have the characteristics of complex structure, strong real-time performance, and huge scale. However, traditional data storage and management systems face the following key technical problems when dealing with ultra-large-scale scientific and technological innovation data: (1) Insufficient storage capacity. The growth rate of scientific and technological innovation data far exceeds the expansion ability of traditional storage systems. Existing systems are difficult to meet the requirements in terms of both capacity and performance, resulting in the gradual emergence of data storage bottlenecks, especially in high-concurrency access scenarios.

[0003] (2) High complexity in data management. The management of ultra-large-scale scientific and technological innovation data involves the efficient storage, dynamic scheduling, and real-time access of multi-dimensional and multi-modal data. Traditional centralized storage architectures are difficult to effectively meet multiple requirements such as data consistency maintenance, distributed storage optimization, and data access efficiency.

[0004] (3) Poor system scalability. With the growth of the scale of scientific and technological innovation data, traditional storage systems are restricted by the architecture design during expansion, with insufficient horizontal expansion ability, and are unable to adapt to dynamic storage requirements and complex access patterns, resulting in a significant decline in storage performance and system reliability.

[0005] Currently, the industry is gradually exploring distributed storage architectures as an important means to meet the storage requirements of ultra-large-scale data. Through distributed node management, multi-node collaboration, and redundant storage mechanisms, the scalability and fault tolerance of storage systems are improved. At the same time, data compression and deduplication technologies can reduce the waste of storage space, and intelligent data management methods provide solutions for data hierarchical storage, load balancing, and fault recovery. However, the existing technologies still have the following deficiencies: Lack of a dynamic hierarchical strategy for high-frequency access data and low-frequency access data, resulting in insufficient data access efficiency; The algorithm performance of redundant data elimination and compression is insufficient, and it is difficult to meet the large-scale processing requirements of scientific and technological innovation data; The data consistency management method between distributed storage nodes has high complexity, resulting in a reduction in system reliability and operating efficiency. Summary of the Invention

[0006] The present invention provides a method, system and device for distributed storage and management of ultra-large-scale scientific and technological innovation data, aiming to solve the problems of insufficient storage capacity, high management complexity and poor system scalability of ultra-large-scale scientific and technological innovation data through a distributed storage architecture, data compression and deduplication technologies, and intelligent data management, so as to optimize the storage space utilization rate, improve the data management efficiency and enhance the dynamic adaptability of the system.

[0007] To solve the above technical problems, the present invention provides a method for distributed storage and management of ultra-large-scale scientific and technological innovation data, including: Based on a distributed storage architecture, obtain ultra-large-scale scientific and technological innovation data, split the data into multiple small pieces and allocate them to multiple storage nodes; Perform compression and deduplication processing on the data in the storage nodes, identify redundant data and compress the valid data to generate an efficiently stored data sequence; Based on the data sequence, use an intelligent algorithm to automatically store the data in layers, store the frequently accessed data in a high-speed storage medium, and store the infrequently accessed data in a low-cost storage medium; Based on the result of the hierarchical storage, use a multi-factor collaborative optimization load balancing algorithm to calculate the comprehensive load value of the nodes, screen the data blocks in the high-load nodes for priority migration, and select a target node to complete the data migration; Based on the distributed storage architecture and the intelligent data management system, manage data consistency through a distributed consistency algorithm, and dynamically adjust the storage resource allocation according to real-time data requirements.

[0008] Further, the load balancing algorithm includes the following steps: Collect the real-time status data of the storage nodes, including storage capacity utilization rate, bandwidth utilization rate, access frequency, and historical load fluctuation prediction value; Use the comprehensive load calculation formula Generate a comprehensive load score for the nodes; According to the node load score, screen the nodes with a load higher than the set threshold as the migration source nodes.

[0009] Further, the selection of the data migration target node includes the following steps: Calculate the data block migration priority , and select the data block with a larger storage capacity and a lower access frequency as the migration object; Through the target node selection formula , screen out the target node with the minimum comprehensive migration cost; Execute the data migration and update the node status and the data storage path index.

[0010] Further, the steps of obtaining and distributively storing the ultra-large-scale scientific and technological innovation data include: Based on a distributed computing framework, splitting the ultra-large-scale scientific and technological innovation data according to data attributes or timestamps; Dynamically allocating the data blocks to multiple storage nodes and generating a data distribution strategy according to the capacities and loads of the storage nodes.

[0011] Further, the steps of redundant storage include: Generating multiple redundant copies for each data block through a replica mechanism and dispersedly storing them to different storage nodes; Generating check information for each data block based on erasure code technology and dispersedly storing it to multiple storage nodes for data recovery.

[0012] Further, the steps of data compression and deduplication processing include: Using a feature hashing algorithm to detect redundant data in the data blocks; Deleting redundant data; Compressing the data blocks based on a content-aware compression algorithm to generate a compressed data sequence.

[0013] Further, the steps of automatic hierarchical storage include: Analyzing the access frequencies and importance of the data sequences to generate a high-priority data set and a low-priority data set; Storing the high-priority data to a high-speed storage medium and storing the low-priority data to a low-cost storage medium; Dynamically adjusting the storage paths and optimizing the hierarchical structure of the storage media.

[0014] Further, the steps of hierarchical storage adjustment include: Calculating the access frequencies of the data blocks in real time based on data access logs; Updating the data priorities according to the access frequencies and migrating the data blocks to the target storage media.

[0015] Further, a distributed storage and management system for ultra-large-scale scientific and technological innovation data includes: A data acquisition and distribution module, configured to obtain ultra-large-scale scientific and technological innovation data from multiple data sources and distribute the data to multiple storage nodes; A data compression and deduplication module, configured to perform compression and deduplication processing on the data in the storage nodes; A hierarchical storage module, configured to hierarchically store data to a high-speed storage medium or a low-cost storage medium according to the access frequencies and importance of the data; A load balancing and fault recovery module, configured to dynamically adjust the data storage distribution and automatically recover data in case of node failures; A consistency management module, which is used to achieve data consistency among multiple nodes through a distributed algorithm and dynamically adjust the storage resource configuration.

[0016] Furthermore, a distributed storage and management device for ultra-large-scale scientific and technological innovation data includes: A distributed storage architecture, which is used to obtain ultra-large-scale scientific and technological innovation data and allocate the data to multiple storage nodes; A data processing unit, which is used to perform compression and deduplication operations on the data; A hierarchical storage unit, which is used to dynamically adjust the storage hierarchical structure based on the access frequency and importance of the data; A system control unit, which is used to monitor the running status of the storage nodes and manage data consistency through a distributed consistency algorithm; A redundant storage unit, which is used to generate redundant data copies or check information and disperse and store them in multiple nodes.

[0017] The key innovations of the present invention include: (1) Distributed storage architecture design. A distributed storage system suitable for ultra-large-scale scientific and technological innovation data is constructed, combined with a data slicing and redundant storage mechanism, supporting horizontal expansion and high availability.

[0018] (2) Data compression and deduplication technology. A redundant data recognition method based on a feature hashing algorithm is proposed, and combined with content-aware compression technology, efficient compression and redundancy elimination of stored data are achieved.

[0019] (3) Intelligent hierarchical storage algorithm. By analyzing the data access frequency and importance, an intelligent hierarchical storage algorithm is adopted to store high-frequency data in high-speed storage media, and at the same time dynamically adjust the hierarchical storage structure to improve data access performance.

[0020] The following are its main beneficial effects: (1) Improved storage efficiency. Through data compression and deduplication technology, the present invention identifies and deletes redundant data in the storage nodes, and efficiently compresses the valid data, greatly reducing the storage space occupancy. Compared with traditional storage methods, the utilization rate of storage space is optimized, and at the same time the storage cost is reduced.

[0021] (2) Optimized data access performance. By adopting intelligent hierarchical storage technology, the storage structure is dynamically adjusted according to the data access frequency and importance, storing high-frequency access data in high-speed storage media and low-frequency access data in low-cost storage media, effectively shortening the access latency of high-frequency data and improving the access performance of the overall storage system.

[0022] (3) System reliability enhancement. Based on the distributed storage architecture, redundant storage of data is achieved, and combined with the node monitoring and fault recovery mechanism, when a node fails, the data can be quickly located and migrated to healthy nodes to ensure the high availability of data and the fault tolerance of the system. Description of the Drawings

[0023] Figure 1 It is a schematic flowchart of a distributed storage and management method for ultra-large-scale scientific and technological innovation data provided by an embodiment of the present application; Figure 2 It is a structural block diagram of a distributed storage and management system for ultra-large-scale scientific and technological innovation data provided by an embodiment of the present application; Figure 3 It is a structural block diagram of a device provided by an embodiment of the present application. Detailed Embodiments

[0024] Embodiment 1: Refer to Figure 1 , which is a schematic flowchart of a distributed storage and management method for ultra-large-scale scientific and technological innovation data provided by an embodiment of the present invention. The process may at least include steps S100 - S500: S100. Based on the distributed storage architecture, obtain ultra-large-scale scientific and technological innovation data and perform distributed storage, allocate the data to multiple storage nodes, and achieve redundant storage of data to ensure high availability of data.

[0025] S200. Based on the data of the storage nodes, perform data compression and deduplication processing, identify redundant data and compress valid data, optimize the utilization rate of the storage space, and generate an efficiently stored data sequence.

[0026] S300. Based on the data sequence after data compression and deduplication, use intelligent algorithms to automatically hierarchically store the data, store frequently accessed data in high-speed storage media, and store low-frequency data in media with lower costs.

[0027] S400. Based on the results of intelligent data hierarchical storage, perform load balancing and fault recovery, real-time monitor the health status of the storage nodes, and dynamically adjust data storage and migration according to the load situation.

[0028] S500. Combine the distributed storage architecture with the intelligent data management system, optimize the storage performance and perform consistency management, achieve data consistency between multiple nodes through distributed algorithms, generate an optimal storage strategy, and adjust the storage resource allocation according to real-time data requirements.

[0029] Step S100 at least includes steps S110 - S130: S110. Obtain ultra-large-scale scientific and technological innovation data, and divide the data into multiple small pieces through the distributed storage architecture and allocate them to multiple storage nodes.

[0030] First, obtain ultra-large-scale data from scientific and technological innovation data sources, including but not limited to multi-modal sensor data, experimental result data, etc. This data contains indicators in multiple dimensions, such as real-time operating condition data like temperature, pressure, stress, etc., as well as historical data.

[0031] Furthermore, based on the distributed storage architecture, use a distributed computing framework to preprocess and segment the ultra-large-scale scientific and technological innovation data. Specifically, the data is divided into multiple data blocks through specific algorithms (such as hash cutting method or timestamp cutting method), and the size of each data block is dynamically adjusted according to the storage capacity of the nodes.

[0032] Furthermore, allocate the segmented data blocks to multiple storage nodes. The node allocation adopts a load balancing algorithm. Specifically, factors such as the storage capacity, load, and bandwidth of the nodes are used as input indicators for data allocation. It can be understood that the load status of node i is measured by the ratio of its current storage capacity to the bandwidth, and the calculation formula is as follows:

[0033] Among them, is the load status of node i, is the current storage capacity of node i, is the bandwidth of node i.

[0034] S120. Perform redundant storage processing on the data blocks, and generate data redundant copies or parity information through a replica mechanism or erasure code technology, and perform backup storage on multiple nodes.

[0035] First, based on the storage allocation situation of the data blocks, use a replica mechanism or erasure code technology (such as Reed-Solomon code) to perform redundant storage on each data block. The replica mechanism generates multiple replicas of the data block and stores them on different nodes; while the erasure code technology generates parity information, so that even if part of the data is lost, the lost data can be restored through other data blocks and parity information.

[0036] Specifically, assume that the data block is stored on node , and a replica is generated and stored on node . The generation ratio of the replica is determined according to the redundant storage requirement and the available storage capacity of the node. The number of replicas is related to the size of the data block, generally 2 to 3 replicas. The calculation method of the replica is: ; Among them, represents the function of replica generation, is the original data block, is a storage node, is a replicated data block.

[0037] Further, when using erasure coding technology, the generated parity information is calculated from multiple parts in the data block. Through the formula: ; where, represents the hash function of the erasure code, is a component of the data block, and the generated parity information is stored in different nodes for recovery when data is lost.

[0038] S130. Monitor the running status of the storage nodes, detect the health status of the storage nodes in real time, generate a status report of the storage nodes, and dynamically adjust the data storage distribution according to factors such as node load and bandwidth to optimize the utilization of storage resources.

[0039] First, monitor the health status of each storage node in real time, and obtain information such as the storage capacity, bandwidth usage, operating temperature, and CPU load of each node. Specifically, through the distributed monitoring system, the node status data will be uploaded to the central monitoring system in real time.

[0040] Further, based on the real-time monitoring data, evaluate the working status of each node and generate a health assessment report for each node. This report uses the following evaluation formula according to parameters such as the node load and bandwidth: ; where, is the health assessment index of node i, is the node load, is the node bandwidth usage, is the node temperature, is the weight coefficient, indicating the relative importance of each index.

[0041] Further, based on the evaluation results, dynamically adjust the resource allocation of the storage nodes. When the health assessment index of a certain node exceeds the set threshold, the system will migrate some data of this node to other healthy nodes and reduce the load of this node. Specifically, the formula for dynamic adjustment is:

[0042] where, is the data block to be migrated, is a healthy node.

[0043] Connection description: The data splitting and distribution steps in S110 provide the basis for data distribution for subsequent redundant storage and monitoring adjustment. The splitting and distribution information of each data block is used as input data in the redundant storage processing of S120.

[0044] S130 ensures the load balance of data blocks among storage nodes by monitoring the status of nodes in real time, makes dynamic adjustments based on the health assessment report, thereby optimizing the data storage distribution and further ensuring the efficient utilization of storage resources.

[0045] Step S200 includes at least steps S210 - S230: S210. Based on the data of the storage nodes, perform data compression and deduplication processing, identify and eliminate redundant data, and generate a compressed data sequence.

[0046] First, obtain the original data blocks that have undergone distributed storage allocation from the storage nodes , where i represents the storage node number and j represents the data block number. Further, based on the node status report (the node load and bandwidth parameters generated by S130), preferentially process the node data with a health status greater than the threshold.

[0047] Further, perform duplicate data detection on the data block and generate a unique identifier based on the feature hashing algorithm . Specifically, the calculation formula is: ; where is the feature hash value of the data block, is the hash function used to identify redundant data.

[0048] Even further, compress the detected non - redundant data and generate a compressed data block using the content - aware compression algorithm . The compression ratio calculation formula is as follows:

[0049] where is the compression ratio, are the sizes of the data blocks before and after compression respectively. The generated will be used as input data for subsequent processing.

[0050] S220. Perform deduplication processing on the data block, identify the redundant parts in the data, delete the duplicate data, and generate a deduplicated data storage set.

[0051] First, based on the feature hash value generated by S210 , match the data blocks in the current storage node with the existing data. If the same hash value is found (where k, l are other data block numbers), then mark as redundant data blocks.

[0052] Further, perform deletion processing on the data blocks marked as redundant , only keep one copy, and store the deleted redundant data reference records in the index table . The format of the index table is: ; Among them, represents the storage node and location where the non-redundant data is located.

[0053] Further, define the set of data blocks retained after deduplication as , where: The deduplicated data set is passed to S230 for further processing.

[0054] S230. Based on the deduplicated and compressed data sequence, organize it into an efficient data storage format to form an optimized stored data sequence.

[0055] First, according to the deduplicated data set , rearrange each data block into an efficient storage format . Specifically, based on the status of the storage node (such as storage capacity and bandwidth utilization), reorder the data blocks to optimize the read and write efficiency of data storage.

[0056] Further, perform format conversion on the data block to generate a hierarchical storage index file . The index file contains the logical address and physical storage address of the data block, defined as follows: ; Among them, represents the address of the data block in the logical layer, represents the actual physical storage location.

[0057] Further, finally save the stored data sequence in the efficient format and the index file to the corresponding storage node, complete all steps of the data compression and deduplication module, and provide an optimized input for the subsequent intelligent hierarchical storage (S300).

[0058] Step S300 includes at least steps S310 - S330: S310. Analyze and identify high - frequency data and low - frequency data from the compressed and deduplicated data sequence according to the access frequency and importance of the data, and generate data sets with different priorities.

[0059] First, obtain the storage address and index file of each data block from the optimized data storage sequence output by the S230 module and the access frequency of the data block. The access frequency is calculated based on the log data of node real - time monitoring.

[0060] Calculate the average access frequency of each data block using the following formula according to historical access log data and real - time request data:

[0061] where, respectively represent the number of read requests and write requests of the data block , and represents the total storage time of this data block.

[0062] Generate a priority score based on the access frequency and the importance index of the data (such as the source and use of the data block, etc.): :

[0063] where, and are weight factors. Divide the scores of into a high - frequency data set H and a low - frequency data set L: ; θ is the priority threshold.

[0064] S320. Based on the analysis results of the high - frequency data and low - frequency data, use an intelligent algorithm to store the frequently accessed data in a high - speed storage medium and store the infrequently accessed data in a lower - cost storage medium.

[0065] First, after obtaining the high - frequency data set H and the low - frequency data set L, allocate storage resources according to the medium type of the storage node (the high - speed storage medium is a solid - state drive SSD, and the low - cost storage medium is a hard disk drive HDD). The target storage medium of each data block Determined according to the following rules: ; Combined with the index file generated in S230 , adjust the physical storage path of each data block , migrate the high-frequency data H to the high-speed storage medium, and migrate the low-frequency data L to the low-cost storage medium.

[0066] Based on the optimized storage path , perform the data migration operation. The migration process will dynamically adjust the migration speed and order according to the current load and health status of the node to avoid system overload.

[0067] S330. Dynamically adjust the data hierarchical storage structure according to the data access mode and storage policy, and monitor the data access situation in real time.

[0068] First, use the monitoring system in S130 to continuously collect the real-time access situation of each data block, including the read / write request frequency , and the time distribution of requests .

[0069] Based on the change of the access mode, re-evaluate the high-frequency data set and the low-frequency data set . Specifically, every certain period of time recalculate the priority score and update the data hierarchical storage set: ; For the data blocks with changed priorities , trigger the automatic migration operation to adjust their storage media to ensure that the hierarchical storage structure is consistent with the real-time access mode.

[0070] Step S400 includes at least steps S410 - S430: S410. Based on the multi-factor collaborative optimization model, by combining the real-time state of the node, the historical load trend, and the data priority, calculate the comprehensive load score of the node, and use the dynamic weight adjustment and predictive migration strategy to achieve an accurate load balancing report.

[0071] First, calculate the comprehensive load value of the node:

[0072] Among them, is the comprehensive load score of node i; is the real-time storage capacity utilization rate of node i, is the maximum storage capacity of all network nodes; is the bandwidth utilization rate of node i, is the maximum bandwidth of all network nodes; is the access frequency of node i, is the average access frequency of all network nodes; is the predicted value of the historical load fluctuation of node i, predicted based on the Long Short-Term Memory Network (LSTM); , , , is the dynamic weight coefficient, reflecting the influence weight of each factor on the comprehensive load.

[0073] Furthermore, calculate the priority of data blocks:

[0074] Among them, is the migration priority of data block j; is the storage size of data block j; is the access frequency of data block j; , is the priority weight.

[0075] Even further, select the target node for data migration:

[0076] Among them, is the minimum comprehensive migration cost of target node k; δ is the weight factor of bandwidth consumption during migration.

[0077] Specifically, the load balancing algorithm realizes the dynamic scheduling and data migration of storage nodes through a multi-factor collaborative optimization model. The specific steps are as follows: ① Calculate the comprehensive load value of nodes. Collect the real-time status data of all storage nodes, including storage capacity utilization rate, bandwidth utilization rate, access frequency, and predicted value of historical load fluctuation. Generate a load score for each node according to the comprehensive load calculation formula .

[0078] ② Calculate the data migration priority. For nodes with high load, based on the data block priority calculation formula , screen out data blocks with larger storage capacity and lower access frequency as migration targets.

[0079] ③ Select the migration target node. Among the nodes with low load, use the comprehensive migration cost formula to select the optimal target node and complete the data migration operation.

[0080] ④ Dynamic adjustment and optimization. Dynamically adjust the weight coefficient according to the real-time changes in the node status , , , and , to ensure the fast response ability of the load balancing strategy to sudden load fluctuations.

[0081] S420. Dynamically adjust the allocation and storage of data according to the load conditions of the storage nodes, and automatically migrate the data through the load balancing algorithm.

[0082] First, filter out the overloaded nodes and the low-load nodes from the load status report output by S410. The overloading condition is:

[0083] The low-load condition is:

[0084] Among them, respectively represent the high threshold and the low threshold of the load.

[0085] Furthermore, select the dataset to be migrated from , and based on the size and the importance index of the data, preferentially migrate the data with low importance and high capacity:

[0086] Assign to the low-load nodes , and allocate the target nodes according to the following rules : ; And update the storage path and the index file to complete the data migration.

[0087] S430. When a node fails or the system is abnormal, based on the redundant storage and data backup mechanism, achieve automatic fault recovery, and automatically migrate the data in the failed node to the available nodes.

[0088] First, use the monitoring system to check the health status of the storage nodes in real time . If it is found that the health status of a certain node is lower than the fault threshold , mark as a failed node :

[0089] Furthermore, extract the redundantly stored copy data set from . For the recovery process of each data block , select the optimal recovery path according to the redundant storage policy (copy mechanism or erasure code): If the copy mechanism is adopted, directly recover from the copy storage node :

[0090] If the erasure code is adopted, reconstruct the data through the parity information : ) Furthermore, migrate the recovered data to the healthy node , and update the storage path and index file: )

[0091] Step S500 includes at least steps S510 - S530: S510. Based on the distributed storage architecture and data management policy, generate a consistency management policy through a distributed consensus algorithm to avoid data loss or conflict.

[0092] First, collect the real - time status of each node from the load status report output from S410 and the storage path update information of S430 , including the current distribution of data copies stored and the node health status. Furthermore, in combination with the redundant storage mechanism, extract the copy distribution information of each data block .

[0093] Furthermore, according to the stored data and the copy distribution , verify the data consistency among multiple nodes. Specifically, detect the consistency conflict by comparing the hash values of the master node data and the copy data :

[0094] If a conflict is detected, record the set of conflict data blocks .

[0095] For the set of conflict data blocks , generate a consistency repair strategy based on a distributed consistency algorithm (such as the Paxos or Raft protocol) . The consistency repair strategy is executed according to the following rules: If the data of the primary node is complete, update the replica data: If the primary node fails, select the data of a healthy node from the redundant replicas as the new primary data: ; S520. Combine the real-time data access situation and storage resources, and adopt an intelligent scheduling mechanism to dynamically allocate storage resources and optimize storage performance.

[0096] First, from the access frequency of S310 and the node load of S410 , extract the real-time access situation of each data block . Specifically, through access log analysis, calculate the current access density of each data block :

[0097] where is the total storage time of this data block.

[0098] Furthermore, according to the access density , sort the data blocks by priority to generate a high-priority data set and a low-priority data set . The high-priority data set is allocated to high-performance storage nodes, and the low-priority data set is allocated to low-cost nodes. The selection of storage nodes is allocated according to the following rules:

[0099] where are the high and low thresholds of the access density respectively.

[0100] Furthermore, based on the above priority allocation results, migrate the data blocks to the target storage nodes and update the index file to record the new storage path .

[0101] S530. Automatically adjust the configuration of storage resources according to the consistency management strategy and storage resource allocation situation, optimize the expansion and allocation of storage according to real-time requirements, and generate an optimal storage strategy.

[0102] First, combine the consistency management strategy of S510 and the storage resource allocation result of S520 , dynamically adjust the resource configuration of storage nodes, including storage capacity , bandwidth allocation And computing power . The specific adjustment rules are as follows: For high-load nodes, release some resources: ; For low-load nodes, increase storage tasks:

[0103] Furthermore, if the current storage capacity is insufficient, trigger the expansion strategy to dynamically increase storage nodes. The expansion strategy determines by detecting the utilization rate of the current storage nodes : ; When , add new nodes and allocate storage tasks.

[0104] Combine the expansion result and the consistency management strategy to generate the optimal storage strategy , and distribute this strategy to each storage node for implementation.

[0105] Example 2: Figure 2 Show a structural block diagram of a distributed storage and management system for ultra-large-scale scientific and technological innovation data according to an embodiment of the present invention. As Figure 2 shown, the system may include: The data acquisition and distribution module 10 is used to obtain ultra-large-scale scientific and technological innovation data from multiple data sources, and perform preliminary segmentation and distribution of the data through a distributed storage architecture.

[0106] Collect data related to scientific and technological innovation research from data sources such as multimodal sensors, experimental platforms, and historical databases, including real-time working condition data (such as temperature, pressure, stress) and historical data.

[0107] Segment the collected raw data through a distributed storage architecture, divide the big data into several small blocks, and dynamically allocate them to different storage nodes using a load balancing strategy.

[0108] Implement data redundant storage, and back up data between multiple nodes through a replica mechanism or erasure code technology.

[0109] The data compression and deduplication module 20 performs compression and deduplication processing on the data in the distributed storage nodes to optimize the utilization rate of the storage space.

[0110] For the data blocks in the storage nodes, use a content-aware compression algorithm to reduce the storage size of the data while maintaining the integrity and accessibility of the data.

[0111] Through duplicate data detection technology, identify and eliminate the redundant parts between data blocks, only retain one data copy, and generate corresponding index information.

[0112] Reorganize the compressed and deduplicated data blocks into an efficient storage format.

[0113] The intelligent hierarchical storage module 30 stores data in layers based on data access frequency and importance, and dynamically adjusts the allocation of storage media.

[0114] Extract data access patterns from the compressed and deduplicated data, and divide the data into high-priority data sets and low-priority data sets according to access frequency and importance.

[0115] Adopt an intelligent storage allocation algorithm to preferentially store frequently accessed data in high-speed storage media (such as SSDs), and store infrequently accessed data in lower-cost media (such as HDDs).

[0116] Monitor data access in real time and dynamically adjust the data layering strategy.

[0117] The load balancing and fault recovery module 40 monitors the load status of storage nodes, dynamically adjusts data distribution, and automatically recovers in case of node failures.

[0118] Based on the intelligent hierarchical storage results, monitor the running status of storage nodes in real time, including storage capacity, access frequency, computing power, etc., and generate a node load status report.

[0119] When the node load exceeds the threshold, migrate some data to low-load nodes based on the load balancing algorithm, and update the storage path and index information at the same time.

[0120] In case of storage node failures, automatically recover data through the redundant storage mechanism and migrate the data of the failed node to healthy nodes.

[0121] The optimization scheduling and consistency management module 50 combines the distributed storage architecture with the intelligent data management system to optimize system performance and achieve consistent management of data among multiple nodes.

[0122] Adopt a distributed consistency algorithm to verify the data in storage nodes, and avoid conflicts and inconsistencies between data copies.

[0123] Based on the real-time data access situation and the status of storage nodes, dynamically adjust the storage resource configuration to optimize the storage performance of the system.

[0124] Generate an optimal storage strategy, dynamically expand the storage capacity according to the system operation requirements, and reasonably allocate node storage tasks.

[0125] Embodiment 3: Figure 3 Show a structural block diagram of a device according to an embodiment of the present application. As Figure 3 shown, the device includes: The distributed storage architecture 41 acquires ultra-large-scale scientific and technological innovation data and distributes the data to multiple storage nodes.

[0126] The scientific and technological innovation data is segmented through a distributed computing framework, and data blocks are generated according to data attributes or timestamps.

[0127] The data blocks are dynamically allocated to multiple storage nodes, and a distribution strategy is generated according to the storage capacity and load status of the nodes.

[0128] The data processing unit 42 performs compression and deduplication operations on the data.

[0129] The redundant data is identified using a feature hashing algorithm, and duplicate content is deleted to reduce storage space occupancy.

[0130] The data blocks are compressed based on content-aware compression technology to generate an optimized data sequence.

[0131] The hierarchical storage unit 43 dynamically adjusts the storage hierarchical structure based on the access frequency and importance of the data.

[0132] The data access logs are analyzed, priorities are assigned according to the access frequency and importance of the data, and high-priority and low-priority data sets are generated.

[0133] The frequently accessed data is stored in a high-speed storage medium (such as an SSD), and the infrequently accessed data is stored in a low-cost storage medium (such as an HDD).

[0134] The storage hierarchical structure is dynamically adjusted to ensure storage efficiency and access performance.

[0135] The system control unit 44 monitors the operating status of the storage nodes and manages data consistency through a distributed consistency algorithm.

[0136] The operating status data of the nodes is collected in real time, including storage capacity, bandwidth utilization, and load information.

[0137] Based on the distributed consistency algorithm, the data consistency between storage nodes is verified to avoid data conflicts or losses.

[0138] When a node fails, the fault is quickly located and the data is restored to a healthy node.

[0139] The redundant storage unit 45 generates redundant data copies or check information and disperses them for storage on multiple nodes.

[0140] Check information is generated based on erasure coding technology to reduce the space occupancy of redundant storage.

Claims

1. A distributed storage and management method for ultra-large-scale scientific and technological innovation data, characterized in that: The following steps are involved: Based on the distributed storage architecture, we can obtain ultra-large-scale scientific and technological innovation data, divide the data into multiple small blocks and distribute them to multiple storage nodes; Compressing and deduplicating the data in the storage node, identifying redundant data and compressing valid data, and generating a data sequence for efficient storage; Based on the data sequence, an intelligent algorithm is used to automatically store the data in layers, storing high-frequency access data in high-speed storage media and low-frequency access data in low-cost storage media; Based on the result of the tiered storage, a multi-factor collaborative optimization load balancing algorithm is used to calculate the comprehensive load value of the node, screen the data blocks in the high-load nodes for priority migration, and select the target node to complete the data migration; Based on distributed storage architecture and intelligent data management system, data consistency is managed through distributed consistency algorithm, and storage resource allocation is dynamically adjusted according to real-time data requirements.

2. The method according to claim 1, characterized in that The load balancing algorithm includes the following steps: Collect real-time status data of storage nodes, including storage capacity utilization, bandwidth utilization, access frequency, and historical load fluctuation prediction values; Using the comprehensive load calculation formula Generate node comprehensive load score; Based on the node load score, nodes with loads higher than the set threshold are selected as migration source nodes.

3. The method according to claim 1, characterized in that: The data migration target node selection comprises the following steps: Calculate the data block migration priority ,Select data blocks with larger storage capacity and lower access frequency as the migration objects; Select the formula by target node , select the target node with the minimum comprehensive migration cost; Perform data migration and update node status and data storage path index.

4. The method according to claim 1, characterized in that The steps of obtaining ultra-large-scale scientific and technological innovation data and performing distributed storage include: Based on a distributed computing framework, the ultra-large-scale scientific and technological innovation data is segmented according to data attributes or timestamps; The data blocks are dynamically allocated to multiple storage nodes, and a data distribution strategy is generated according to the capacity and load of the storage nodes.

5. The method according to claim 1, characterized in that: The redundant storage step comprises: The replication mechanism generates multiple redundant copies for each data block and stores them in different storage nodes. Based on erasure coding technology, verification information is generated for each data block and stored in multiple storage nodes for data recovery.

6. The method according to claim 1, characterized in that The steps of data compression and deduplication processing include: Using a feature hash algorithm to detect redundant data in the data block; Remove redundant data; The data block is compressed based on the content-aware compression algorithm to generate a compressed data sequence.

7. The method according to claim 1, characterized in that The steps of automatically tiering storage include: Analyze the access frequency and importance of the data sequence to generate a high-priority data set and a low-priority data set; storing the high priority data into a high-speed storage medium, and storing the low priority data into a low-cost storage medium; Dynamically adjust storage paths and optimize the hierarchical structure of storage media.

8. The method according to claim 5, characterized in that The step of adjusting the tiered storage includes: Calculating the access frequency of the data block in real time based on the data access log; The data priority is updated according to the access frequency, and the data block is migrated to the target storage medium.

9. A distributed storage and management system for ultra-large-scale scientific and technological data, characterized in that: include: A data collection and distribution module, which is used to obtain ultra-large-scale scientific and technological innovation data from multiple data sources and distribute the data to multiple storage nodes; A data compression and deduplication module, used for compressing and deduplicating the data in the storage node; A hierarchical storage module, used for storing data hierarchically to a high-speed storage medium or a low-cost storage medium according to the access frequency and importance of the data; Load balancing and fault recovery module, used to dynamically adjust data storage distribution and automatically recover data when a node fails; The consistency management module is used to achieve data consistency among multiple nodes through a distributed algorithm and dynamically adjust storage resource configuration.

10. A distributed storage and management device for ultra-large-scale scientific and technological data, characterized in that: include: Distributed storage architecture for acquiring ultra-large-scale scientific and technological data and distributing the data to multiple storage nodes; A data processing unit, used for performing compression and deduplication operations on the data; A hierarchical storage unit, configured to dynamically adjust a storage hierarchical structure based on access frequency and importance of the data; System control unit, which is used to monitor the operating status of storage nodes and manage data consistency through distributed consistency algorithms; The redundant storage unit is used to generate redundant copies of data or check information and store them in multiple nodes.

Citation Information

Patent Citations

  • Performance evaluation system and method for multi-stage RAID system

    CN117909198A

  • Data storage method based on large model

    CN119292524A

  • Method and system for safely and rapidly storing image data of imaging department

    CN119517325A

  • Data index construction method for distributed data storage

    CN119829551A

  • Datacenter storage system

    US20140115579A1

Cited By

  • Distributed block storage system performance optimization method and system

    CN120803377A

  • Distributed block storage system performance optimization methods and systems

    CN120803377B

  • Financial big data information processing system based on block chain

    CN120910886A

  • Distributed data storage dynamic optimization method based on edge computing

    CN121193759A