Data storage method based on large model
By performing intelligent sharded storage and redundant optimization on large-scale machine learning models, the problem of insufficient performance in large-scale machine learning model training is solved, efficient storage and fast access are achieved, and model training efficiency and real-time inference are improved.
Patent Information
- Application Number
- CN202411335849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-09-24
AI Technical Summary
In the process of training of large-scale machine learning models, traditional data storage systems are difficult to meet the needs of high-capacity storage, and the I/O performance is insufficient, resulting in inefficient training efficiency, and waste of storage resources and real-time inference are affected.
By extracting the internal parameters of the target pre-trained large model, using the multi-dimensional model parameter matrix and the input type parameter matrix for intelligent sharding storage strategy processing, combining the distributed storage management system to efficiently store and quickly access the training data, and optimizing storage efficiency through redundant data detection and differentiated compression.
It significantly improves the flexibility and expansion capabilities of the storage system, reduces storage costs, reduces the storage space of redundant data, improves model training speed, shortens training cycles, and enhances the real-time nature of model inference.
Smart Images

Figure CN119292524B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage management, and in particular to a data storage method based on a large model. Background Art
[0002] With the rapid development of technologies such as artificial intelligence, big data, and the Internet of Things, large-scale machine learning models are increasingly used in many fields, such as image recognition, natural language processing, and autonomous driving. However, in the face of massive data requirements and complex computing tasks, the performance of the data storage system is particularly important. In the training process of large-scale machine learning models, the data storage system not only needs to process a large amount of training data, but also needs to ensure efficient data transmission and real-time call. Therefore, the design of the data storage system directly affects the efficiency of model training and the effect of reasoning. The amount of data generated during the training of large models is huge, which requires the storage system to have high capacity characteristics. Traditional storage devices often find it difficult to meet storage requirements when facing PB-level data, and the cost is high. Secondly, data access speed is also a major problem. Large model training requires frequent reading and writing of data. If the I / O performance of the storage system is insufficient, the training efficiency will be seriously affected. However, traditional large model data storage usually uses multiple copies. Data redundancy is very common in large model training, especially in distributed environments. Data will be stored repeatedly, resulting in a waste of storage resources and low storage efficiency. The storage system cannot quickly respond to the high-frequency reading and writing needs of large models for massive data, which in turn affects the speed of model training, prolongs the training cycle, and ultimately weakens the real-time performance of model reasoning. Summary of the invention
[0003] Based on this, the present invention provides a data storage method based on a large model to solve at least one of the above technical problems.
[0004] To achieve the above purpose, a data storage method based on a large model includes the following steps:
[0005] Step S1: extract model internal parameters of the target pre-trained large model to obtain a multi-dimensional model parameter matrix and an input type parameter matrix respectively; perform input cluster sharding optimization according to the input type parameter matrix to generate modulus input sharding data; perform distributed storage node mapping on the modulus input sharding data to obtain storage modulus sharding data;
[0006] Step S2: Perform intelligent sharding storage strategy processing on the storage modulus sharding data through the multidimensional model parameter matrix, thereby obtaining an intelligent sharding storage strategy; obtain target model training data; perform intelligent sharding storage on the target model training data based on the distributed storage management system using the intelligent sharding storage strategy, and obtain model training storage data;
[0007] Step S3: performing redundant data block detection on the model training storage data to obtain training redundant data and training non-redundant data respectively; performing redundant compression optimization on the training redundant data and the training non-redundant data, and constructing redundant storage indexes to obtain dynamic redundant storage index data; performing data storage retrieval optimization based on the dynamic redundant storage index data to generate optimized storage training data;
[0008] Step S4: Perform training data call sequence reasoning on the optimized storage training data to generate data call sequence data; perform key data identification on the data call sequence data, and perform fast access processing to generate key access training data; perform data preloading processing on the key access training data to generate model training preloading cache data;
[0009] Step S5: Perform multi-channel parallel transmission of the data stream for model training pre-loaded cache data to obtain a parallel transmission data stream; transmit the parallel transmission data stream to the target pre-trained large model for model training, and perform data access hit rate feedback to generate cache hit rate data; optimize the training data storage space according to the cache hit rate data, thereby obtaining an optimized data storage strategy.
[0010] The present invention can accurately understand the model structure and parameter distribution by extracting internal parameters of the target pre-trained large model. The multidimensional model parameter matrix provides detailed parameter information, and the input type parameter matrix helps determine different types of input data and their corresponding features. Through input cluster sharding optimization, the input data can be grouped according to similarity, further reducing redundancy and improving storage efficiency. By using the multidimensional model parameter matrix for intelligent sharding storage strategy processing, the optimal storage solution can be formulated according to the characteristics of different data shards. This not only improves storage efficiency, but also reduces the delay when reading data. By obtaining the target model training data and applying the intelligent sharding storage strategy, the efficient storage and fast access of the training data in the distributed storage system are ensured. The application of this strategy significantly improves the flexibility and expansion capability of the storage system, while reducing the storage cost. Through redundant data block detection, repeated or redundant data blocks can be effectively identified, thereby avoiding unnecessary waste of storage space. Redundant compression optimization is performed on training redundant data and training non-redundant data, further reducing the storage space occupied. By constructing a redundant storage index, the specific location of the data can be quickly located, improving the efficiency of data retrieval. By performing training data call sequence reasoning on the optimized storage training data, the data access mode and frequency can be predicted, thereby optimizing the data loading sequence and helping to reduce the delay in data access. By performing key data identification on the data call sequence data, it can be determined which data is the key data frequently accessed during the training process. By performing data preloading processing on these key access training data, the data reading speed during the model training process can be significantly improved, thereby accelerating the training process. By performing multi-channel parallel transmission of data streams on the model training preload cache data, the advantages of multiple channels can be fully utilized to significantly improve the data transmission rate. Transmitting the parallel transmission data stream to the target pre-trained large model for model training can more efficiently utilize computing resources and accelerate the training process. By optimizing the training data storage space according to the cache hit rate data, the storage strategy can be dynamically adjusted to ensure the efficiency of data storage and access, thereby improving the overall training efficiency and system performance. Therefore, a data storage method based on a large model of the present invention optimizes the storage mode of input data and reduces the redundancy of data by extracting parameters of the target model and preprocessing data. The intelligent sharding storage strategy is adopted so that data can be reasonably distributed between storage nodes according to access frequency and importance, avoiding the waste of resources caused by repeated storage. At the same time, through redundant data detection and differentiated compression processing, the storage space occupied by redundant data is reduced, the storage efficiency is improved, the speed of model training is significantly increased, the training cycle is shortened, and ultimately the real-time performance of model reasoning is enhanced.
[0011] Preferably, step S1 comprises the following steps:
[0012] Step S11: extracting model internal parameters of the target pre-trained large model to obtain a multi-dimensional model parameter matrix and an input type parameter matrix, wherein the multi-dimensional model parameter matrix includes a model attention weight matrix, an embedding vector matrix, and an inter-layer activation parameter matrix;
[0013] Step S12: performing principal component dimensionality reduction processing according to the input type parameter matrix to generate low-dimensional model input type parameters;
[0014] Step S13: performing parameter feature discretization processing on the low-dimensional model input type parameters to generate discrete model input feature data;
[0015] Step S14: Perform cluster analysis based on the discrete model input feature data, and build a global metadata index to generate model input index structure data; perform cluster data sharding balance optimization based on the model input index structure data to generate modulus input shard data;
[0016] Step S15: Use the hash sharding algorithm to perform distributed storage node mapping on the modulus input shard data, thereby obtaining storage modulus shard data.
[0017] The present invention extracts model internal parameters of the target pre-trained large model, and can obtain a multidimensional model parameter matrix and an input type parameter matrix. The multidimensional model parameter matrix includes a model attention weight matrix, an embedding vector matrix, and an inter-layer activation parameter matrix, which are key components of model performance. According to the input type parameter matrix, principal component dimensionality reduction processing is performed, and high-dimensional input type parameters can be converted into low-dimensional representations. By discretizing the parameter features of the low-dimensional model input type parameters, continuous feature values can be converted into discrete category values. Cluster analysis is performed based on discrete model input feature data, and similar data can be classified into the same cluster. Cluster analysis helps to identify the intrinsic structure and pattern of data, and global metadata index construction improves the efficiency of data retrieval. Finally, cluster data sharding is balanced and optimized based on the model input index structure data, ensuring that each data shard is more balanced in size and content, and reducing the imbalance problem in the storage and access process. The modulus input shard data is mapped to distributed storage nodes using a hash sharding algorithm, and the uniform distribution of data on multiple storage nodes can be achieved. The hash sharding algorithm ensures that data shards can be efficiently mapped to different storage nodes, improving the load balancing capability and data access speed of the storage system.
[0018] Preferably, step S14 comprises the following steps:
[0019] Step S141: performing feature point similarity calculation based on the discrete model input feature data to generate input feature similarity data;
[0020] Step S142: performing cluster analysis on the discrete model input feature data by inputting feature similarity data to generate model input cluster data;
[0021] Step S143: performing hash coding processing on the model input cluster data to generate input cluster hash coding data;
[0022] Step S144: Perform hash feature mapping on the input cluster hash code data and the model input cluster data to obtain cluster hash mapping table data;
[0023] Step S145: construct a global metadata index based on the cluster hash mapping table data to generate model input index structure data;
[0024] Step S146: Utilize the model input index structure data to perform cluster data sharding balance optimization on the model input cluster data to generate modulus input shard data.
[0025] The present invention can quantify the similarity between input features by calculating the similarity of feature points on discrete model input feature data. By calculating the distance or similarity index (such as Euclidean distance, cosine similarity, etc.) between different feature points, it is ensured that similar data points can be classified into the same cluster. Data points with similar features are classified into the same cluster. Cluster analysis divides data points into several clusters through algorithms (such as K-means, DBSCAN, etc.), and the data points in each cluster are similar to each other, while the differences between clusters are large. Each cluster is mapped to a unique hash code. Hash coding is an efficient coding method that can convert complex cluster information into a short hash code to facilitate subsequent indexing and retrieval. A corresponding relationship between hash coding and actual cluster data is established, so that each hash code can quickly locate the corresponding cluster data. Cluster hash mapping table data not only improves the speed of data retrieval, but also simplifies the process of data management. Global metadata index construction is performed based on cluster hash mapping table data. The establishment of the index can speed up the speed and accuracy of data access and optimize the process of data management. The global metadata index provides a comprehensive view of the input data, making it easy to quickly locate and access specific data sets, improving the overall performance of the system. The optimization algorithm ensures that each data shard is more balanced in size and content, reducing the imbalance problem in the storage and access process. The generated modulus input shard data not only improves the load balancing ability of the storage system, but also optimizes the efficiency of data access.
[0026] Preferably, step S2 comprises the following steps:
[0027] Step S21: Perform distributed storage cluster topology analysis based on storage modulus sharding data to generate storage node topology data;
[0028] Step S22: Real-time monitoring of node resources of storage node topology data is performed through a distributed storage management system to obtain real-time load data of storage nodes;
[0029] Step S23: Based on the real-time load data of the storage node, the module input shard data is processed by the multi-dimensional model parameter matrix to obtain the intelligent shard storage strategy;
[0030] Step S24: Obtain target model training data;
[0031] Step S25: Based on the distributed storage management system, the target model training data is intelligently partitioned and stored using an intelligent partitioning storage strategy to obtain model training storage data.
[0032] The present invention can understand the connection relationship and network structure between each storage node in the distributed storage system in detail by performing distributed storage cluster topology analysis on the storage module shard data. Through topological analysis, it can ensure that the data shards are reasonably distributed between each node, avoid network bottlenecks and hot spots, and thus improve the reliability and efficiency of the system. By using the distributed storage management system to monitor the storage node topology data in real time, the current resource usage and load status of each storage node can be obtained, ensuring that the system can respond quickly under changing load conditions and optimize resource allocation. By using the storage node real-time load data and the multidimensional model parameter matrix to process the module input shard data with intelligent shard storage strategy, the optimal storage strategy can be generated, the balance and reliability of data storage are improved, the delay of data access is reduced, and the performance of the overall system is improved. Based on the distributed storage management system, the intelligent shard storage strategy is used to perform intelligent shard storage on the target model training data, and the efficient distribution of data on multiple storage nodes can be realized. Through the intelligent shard storage strategy, the training data can be reasonably distributed to each node to ensure the balance and high availability of data distribution.
[0033] Preferably, step S23 includes the following steps:
[0034] Step S231: performing storage resource requirement evaluation on the module input shard data to generate shard storage resource requirement data;
[0035] Step S232: Perform storage node performance mapping analysis on storage node real-time load data using storage node topology data to generate storage node performance mapping data; perform shard storage pressure processing on storage node performance mapping data using shard storage resource demand data to generate shard load pressure data;
[0036] Step S233: performing storage node fitness calculation on the modulus input shard data based on the shard load pressure data, thereby obtaining storage node fitness data;
[0037] Step S234: performing intelligent shard node matching on the module input shard data through the multi-dimensional model parameter matrix and the storage node fitness data to obtain intelligent shard matching data;
[0038] Step S235: Perform shard storage migration according to the smart shard matching data, thereby obtaining a smart shard storage strategy.
[0039] The present invention can quantify the storage resources required for each data shard by evaluating the storage resource requirements of the modulus input shard data. The storage resources required for each data shard can be quantified by evaluating the storage resource requirements of the modulus input shard data. The performance of each node is evaluated by analyzing the real-time load and topological structure of each node. The generated storage node performance mapping data provides the current resource usage and performance level of each node. The resource requirements of each shard and the current performance of the node are comprehensively considered, and the potential pressure of allocating each shard to different nodes is evaluated. It is ensured that the allocation of data shards can balance the load of each node and avoid the occurrence of overload. Intelligent shard matching data ensures that each shard can be allocated to the most suitable node and maximizes the use of node resources. Through this strategy, data distribution can be dynamically adjusted to allocate each shard to the best storage node. The intelligent shard storage strategy ensures efficient storage and fast access of data shards, avoiding the problems of uneven load and waste of resources. Through this strategy, the system can dynamically adjust data distribution, improve the overall performance and reliability of the storage system, and ensure efficient management and fast access of data in a distributed storage system.
[0040] Preferably, step S232 includes the following steps:
[0041] Step S2321: Perform node historical shard load analysis based on the real-time load data of the storage node to obtain node historical shard load data;
[0042] Step S2322: Perform node load fluctuation analysis on the node historical shard load data to generate storage node load fluctuation data;
[0043] Step S2323: Perform storage node performance mapping analysis on the storage node topology data using the node historical shard load data to generate storage node performance mapping data;
[0044] Step S2324: performing node load carrying capacity evaluation according to the storage node performance mapping data and the storage node load fluctuation data to generate node load capacity evaluation data;
[0045] Step S2325: Use the shard storage resource demand data to perform shard demand-resource mapping on the node load capacity assessment data, and perform shard load pressure calculation to generate shard load pressure data.
[0046] The present invention can understand the load situation of each node in the past period of time by performing historical shard load analysis on the real-time load data of the storage node. By performing node load fluctuation analysis on the node historical shard load data, the changing trend and fluctuation of the load of each node can be identified. The load fluctuation amplitude and periodic characteristics of each node in different time periods are provided. These data are helpful to identify the peak and trough periods of load. The storage node topology data is analyzed by using the node historical shard load data, and the performance of each node is evaluated by combining the historical load data and the topological structure information. The storage node performance mapping data provides the current resource usage, network connection status and overall performance level of each node, ensuring that the performance evaluation of each node is more accurate and comprehensive. The historical load situation, current performance and load fluctuation of the node are comprehensively considered, and the carrying capacity of each node under different load conditions is evaluated to ensure that each node can operate efficiently within its maximum carrying capacity to avoid overload and waste of resources. By comprehensively considering the resource requirements of each shard and the load capacity of the node, the potential pressure of allocating each shard to different nodes is evaluated to ensure that each shard can be allocated to the most suitable node, maximize the balance of the load of each node, and improve the overall performance and stability of the system. In this way, resource waste and uneven load problems can be avoided, ensuring efficient management and fast access to data in the distributed storage system.
[0047] Preferably, step S234 includes the following steps:
[0048] Step S2341: performing a model parameter training impact assessment on the multi-dimensional model parameter matrix to generate model training impact weight data;
[0049] Step S2342: performing training data stream access frequency statistics according to the model training influence weight data to obtain model training access frequency data;
[0050] Step S2343: performing input slice data priority classification on the modulus input slice data through model training access frequency data to generate model input slice classification data;
[0051] Step S2344: constructing a node storage hierarchy according to the storage node performance mapping data to generate node storage hierarchy structure data;
[0052] Step S2345: performing multiple rounds of priority matching processing on the node storage hierarchical structure data and the model input shard classification data to generate preliminary shard node allocation data;
[0053] Step S2346: Use the storage node fitness data to perform node pressure balancing correction on the preliminary shard node allocation data to generate intelligent shard node matching data.
[0054] The present invention can quantify the importance and influence of each parameter in the training process by evaluating the impact of model parameter training on a multidimensional model parameter matrix. These data provide the contribution of each parameter in the training process, which is helpful for the subsequent training access data flow frequency statistics and optimization. Through this evaluation, it can ensure that resource allocation is more reasonable, and the storage and access of key parameters are optimized. By analyzing the access frequency of each training data cluster in the training process, data slices with different importance and access frequency are distinguished, ensuring that data with higher importance can be processed preferentially during storage and access. Each slice is divided into different priority categories according to its access frequency. High-priority slices will be given priority during storage and access, while low-priority slices can be appropriately delayed. By comprehensively analyzing the performance and resource conditions of each node, a node storage hierarchy structure is constructed, and the node storage hierarchy structure data clarifies the roles and levels of different nodes in the storage system. By comprehensively considering the adaptability of the node and the current load condition, the preliminary allocation plan is corrected to ensure load balancing of each node. Ensure that each slice is stored on the best node, while avoiding node overload and resource waste.
[0055] Preferably, step S3 comprises the following steps:
[0056] Step S31: scanning the model training storage data step by step, and calculating the data repetition rate to obtain data block repetition rate data;
[0057] Step S32: using a preset data redundancy threshold to perform redundant data block detection on the data block repetition rate data, marking the data blocks whose data block repetition rate data is higher than or equal to the preset data redundancy threshold as training redundant data; marking the data blocks whose data block repetition rate data is lower than the preset data redundancy threshold as training non-redundant data;
[0058] Step S33: dividing the training redundant data into redundancy levels to obtain high redundant data, medium redundant data and low redundant data;
[0059] Step S34: perform pointer redirection deduplication processing on high-redundancy data to generate high-redundancy compressed optimized data; perform data differential compression processing on medium-redundancy data to generate medium-redundancy differential compressed data; perform light compression processing on low-redundancy data to generate low-redundancy compressed data;
[0060] Step S35: performing lossless compression processing on the training non-redundant data to generate lossless non-redundant compressed data;
[0061] Step S36: constructing a redundant storage index according to the high-redundancy compression optimization data, the medium-redundancy difference compression data, the low-redundancy compression data, and the lossless non-redundant compression data, thereby obtaining dynamic redundant storage index data;
[0062] Step S37: Optimize data storage and retrieval of the distributed storage management system through dynamic redundant storage index data to generate optimized storage training data.
[0063] The present invention can accurately identify the repetition of each data block by scanning the model training storage data step by step and calculating the data repetition rate. In this way, duplicate data can be effectively identified and managed, and the waste of storage space can be reduced. By setting a threshold, the data blocks are divided into two categories: redundant and non-redundant. Redundant data blocks refer to those data blocks with a high repetition rate, while non-redundant data blocks refer to those data blocks with a low repetition rate. By further refining the classification of redundant data, the redundant data is divided into three levels: high redundant data, medium redundant data and low redundant data, which can effectively reduce the storage space occupied by redundant data. Pointer redirection deduplication processing can eliminate duplicate data and improve storage efficiency; differential compression can reduce storage requirements while retaining data integrity; lightweight compression ensures that low redundant data can also be effectively stored. It can effectively reduce the storage space occupied by redundant data. Pointer redirection deduplication processing can eliminate duplicate data and improve storage efficiency; differential compression can reduce storage requirements while retaining data integrity; lightweight compression ensures that low redundant data can also be effectively stored. By building an index, all compressed data are uniformly managed, providing the ability to quickly locate and access, and ensuring that the compressed data can be efficiently stored and retrieved. By building an index, all compressed data can be managed uniformly, providing the ability to quickly locate and access. Dynamic redundant storage index data ensures that compressed data can be stored and retrieved efficiently, improving the overall performance of the storage system.
[0064] Preferably, step S4 comprises the following steps:
[0065] Step S41: monitoring the training process of the target pre-trained large model in real time, identifying the model training stage, and obtaining model training stage identification data;
[0066] Step S42: Tracking the training data input according to the model training phase identification data to generate model training input feature data;
[0067] Step S43: Performing training data call sequence reasoning on the optimized stored training data through model training input feature data to generate data call sequence data;
[0068] Step S44: Perform training phase resource demand analysis based on the model training phase identification data to generate training resource demand feature data;
[0069] Step S45: performing key data identification on the data call sequence data, and performing fast access processing according to the training resource demand characteristic data to generate key access training data;
[0070] Step S46: Allocate the key access training data to a fast access area, perform data preloading processing, and generate model training preloading cache data.
[0071] The present invention generates model training stage identification data by real-time monitoring of the state and training progress of the model. These data provide specific information of each training stage, including the initialization stage, the pre-training stage, the fine-tuning stage, etc. By tracking the input data of each training stage, recording the characteristics and usage of the input data, the specific characteristics and usage frequency of the input data of each stage are provided. By analyzing the input characteristic data, the calling sequence of the data in the training process is inferred. The generated data calling sequence data provides the calling time and frequency of each data block in the training process, ensuring that the data can be efficiently loaded and accessed on demand. In this way, unnecessary data loading can be reduced and the efficiency of the training process can be improved. By analyzing the resource requirements of each training stage, including computing resources, storage resources and network resources, etc., these data provide the amount and type of resources required for each stage. Identify the key data frequently accessed during the training process. Combined with the training resource requirement characteristic data, these key data are processed for fast access to ensure that the data with high frequency access can be quickly loaded and accessed. By allocating the key access training data to the fast access area, it is ensured that these data can be quickly loaded into the memory. The data pre-loading process loads the key data into the cache in advance to reduce data access delay.
[0072] Preferably, step S5 comprises the following steps:
[0073] Step S51: Allocate data transmission tasks for the model training preloaded cache data to obtain transmission task allocation data;
[0074] Step S52: Allocate data according to the transmission task to perform multi-channel parallel transmission of data streams to obtain parallel transmission data streams;
[0075] Step S53: transmitting the parallel transmission data stream to the target pre-trained large model for model training, and performing model training delay evaluation to obtain model training delay data;
[0076] Step S54: performing data consistency check on the parallel transmission data stream to obtain consistency verification data;
[0077] Step S55: performing data access hit rate feedback according to the consistency verification data to generate cache hit rate data;
[0078] Step S56: Dynamically adjust the cache using cache hit rate data and model training delay data, and optimize the training data storage space to obtain an optimized data storage strategy.
[0079] The present invention can ensure that each data block can be efficiently transmitted to the target node by allocating data transmission tasks for model training preloaded cache data, thereby ensuring the efficiency and reliability of data transmission. In this way, the delay and congestion of data transmission can be reduced, and the overall transmission efficiency can be improved. By using multiple transmission channels to transmit data simultaneously, the speed of data transmission is significantly improved. The parallel transmission data stream ensures that the data can quickly reach the target node, reducing the transmission time. By directly using the transmitted data stream for model training, it is ensured that the data can participate in the training process in time. At the same time, by evaluating the delay in the model training process, the time consumption of data transmission and processing in the training process can be accurately understood. By verifying the integrity and consistency of each data block during the transmission process, the accuracy of data transmission is ensured. The consistency verification data provides information on whether each data block is correctly transmitted, preventing data damage or loss caused by transmission errors. By analyzing the consistency verification data, the hit rate of data in the cache, that is, the access frequency and hit number of data in the cache, is evaluated. The cache hit rate data provides the efficiency and effectiveness of cache management, ensuring that the data in the cache can quickly respond to training requests. By comprehensively analyzing the cache hit rate and model training delay, the cache strategy is dynamically adjusted to ensure efficient storage and fast access of data in the cache. At the same time, optimizing the storage space for training data, reducing unnecessary storage overhead, and improving the overall performance of the storage system can significantly improve the efficiency of data storage and access, and enhance the overall performance and reliability of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 It is a schematic diagram of the steps of the data storage method based on the large model of the present invention;
[0081] Figure 2 for Figure 1 Detailed implementation steps of step S4 in FIG.
[0082] Figure 3 for Figure 1 Detailed implementation steps of step S5;
[0083] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0084] The technical method of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.
[0085] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.
[0086] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0087] To achieve this, please refer to Figures 1 to 3 The present invention provides a data storage method based on a large model, comprising the following steps:
[0088] Step S1: extract model internal parameters of the target pre-trained large model to obtain a multi-dimensional model parameter matrix and an input type parameter matrix respectively; perform input cluster sharding optimization according to the input type parameter matrix to generate modulus input sharding data; perform distributed storage node mapping on the modulus input sharding data to obtain storage modulus sharding data;
[0089] Step S2: Perform intelligent sharding storage strategy processing on the storage modulus sharding data through the multidimensional model parameter matrix, thereby obtaining an intelligent sharding storage strategy; obtain target model training data; perform intelligent sharding storage on the target model training data based on the distributed storage management system using the intelligent sharding storage strategy, and obtain model training storage data;
[0090] Step S3: performing redundant data block detection on the model training storage data to obtain training redundant data and training non-redundant data respectively; performing redundant compression optimization on the training redundant data and the training non-redundant data, and constructing redundant storage indexes to obtain dynamic redundant storage index data; performing data storage retrieval optimization based on the dynamic redundant storage index data to generate optimized storage training data;
[0091] Step S4: Perform training data call sequence reasoning on the optimized storage training data to generate data call sequence data; perform key data identification on the data call sequence data, and perform fast access processing to generate key access training data; perform data preloading processing on the key access training data to generate model training preloading cache data;
[0092] Step S5: Perform multi-channel parallel transmission of the data stream for model training pre-loaded cache data to obtain a parallel transmission data stream; transmit the parallel transmission data stream to the target pre-trained large model for model training, and perform data access hit rate feedback to generate cache hit rate data; optimize the training data storage space according to the cache hit rate data, thereby obtaining an optimized data storage strategy.
[0093] In the embodiment of the present invention, reference Figure 1 The above is a schematic diagram of the steps of the data storage method based on the large model of the present invention. In this embodiment, the data storage method based on the large model includes the following steps:
[0094] Step S1: extract model internal parameters of the target pre-trained large model to obtain a multi-dimensional model parameter matrix and an input type parameter matrix respectively; perform input cluster sharding optimization according to the input type parameter matrix to generate modulus input sharding data; perform distributed storage node mapping on the modulus input sharding data to obtain storage modulus sharding data;
[0095] In the embodiment of the present invention, taking the GPT-3 model as an example, the model is first loaded using the Transformers library of Python, and the internal parameters of the model are extracted. The model parameters mainly include weight matrix, bias vector, etc., which can be obtained using the model.parameters() method. The input type parameter matrix can be obtained by analyzing the input layer structure of the model. For example, the input of GPT-3 is a token sequence, and its type parameter matrix can represent the type of each token (such as part of speech, position encoding, etc.). After obtaining the input type parameter matrix, the K-Means clustering algorithm can be used to cluster the input types and divide inputs of similar types into the same cluster. For example, tokens with similar parts of speech such as nouns and verbs can be divided into the same cluster. Then, the input data is sharded according to the cluster to generate modulus input shard data. For example, the token sequence of each cluster can be used as a shard. Use the distributed file system HDFS to map the modulus input shard data to different storage nodes, for example, store each shard on a different DataNode to obtain storage modulus shard data.
[0096] Step S2: Perform intelligent sharding storage strategy processing on the storage modulus sharding data through the multidimensional model parameter matrix, thereby obtaining an intelligent sharding storage strategy; obtain target model training data; perform intelligent sharding storage on the target model training data based on the distributed storage management system using the intelligent sharding storage strategy, and obtain model training storage data;
[0097] In an embodiment of the present invention, for example, the weight matrices of each layer of the GPT-3 model are used to analyze the sparsity, access frequency and other characteristics of the parameter matrix. For example, the NumPy library of Python can be used to calculate the sparsity of the matrix, and the Spark framework can be used to count the access frequency of the parameter matrix during the model training process. According to the characteristics of the model parameter matrix, an intelligent sharding storage strategy is designed. For example, training with high access frequency can be stored on an SSD hard disk with faster access speed, and training with high sparsity can be compressed and stored to save storage space. Obtain target model training data, such as a large amount of text corpus. Based on the distributed storage management system Ceph, the target model training data is intelligently sharded and stored using an intelligent sharding storage strategy. For example, text corpora of different topics can be stored in different storage pools, and text data with high access frequency can be stored on storage nodes with higher performance to obtain model training storage data.
[0098] Step S3: performing redundant data block detection on the model training storage data to obtain training redundant data and training non-redundant data respectively; performing redundant compression optimization on the training redundant data and the training non-redundant data, and constructing redundant storage indexes to obtain dynamic redundant storage index data; performing data storage retrieval optimization based on the dynamic redundant storage index data to generate optimized storage training data;
[0099] In an embodiment of the present invention, a deduplication tool, such as the MapReduce framework of Apache Hadoop, is used to detect redundant data blocks for model training storage data. For example, text data can be divided into blocks according to a fixed size, and the hash value of each data block is calculated, and then the hash values of different data blocks are compared to identify duplicate data blocks, and obtain training redundant data and training non-redundant data respectively. Deduplication operation is performed on the training redundant data, for example, only one duplicate data block is retained. Compression optimization is performed on the training redundant data and the training non-redundant data, for example, the text data is compressed using the gzip algorithm. At the same time, a redundant storage index is constructed to record the storage location and compression method of the redundant data, so as to obtain dynamic redundant storage index data. For example, a LevelDB database can be used to store redundant storage indexes. According to the dynamic redundant storage index data, for example, the compressed data block is searched according to the index, and a decompression operation is performed to optimize the data storage retrieval and generate optimized storage training data.
[0100] Step S4: Perform training data call sequence reasoning on the optimized storage training data to generate data call sequence data; perform key data identification on the data call sequence data, and perform fast access processing to generate key access training data; perform data preloading processing on the key access training data to generate model training preloading cache data;
[0101] In an embodiment of the present invention, the model training code, such as the training code of the GPT-3 model, and the data access pattern during the model training process are analyzed, such as analyzing the access timestamp of each data block during the model training process, and performing training data call order reasoning on the optimized storage training data. For example, a machine learning algorithm can be used to predict the next data block to be accessed. Generate data call sequence data, such as sorting the data blocks according to the predicted access order. Perform key data identification on the data call sequence data, such as identifying frequently accessed data blocks during model training, and perform fast access processing, such as storing key data on a storage medium with faster access speed. Generate key access training data, such as storing key data in memory. Perform data preloading processing on key access training data, such as loading key data into memory before model training starts. Generate model training preload cache data, such as storing key data in CPU cache or GPU video memory.
[0102] Step S5: Perform multi-channel parallel transmission of the data stream for model training pre-loaded cache data to obtain a parallel transmission data stream; transmit the parallel transmission data stream to the target pre-trained large model for model training, and perform data access hit rate feedback to generate cache hit rate data; optimize the training data storage space according to the cache hit rate data, thereby obtaining an optimized data storage strategy.
[0103] In an embodiment of the present invention, a data transfer library, such as RDMA or NVLink, is used to perform multi-channel parallel transmission of data streams for model training preload cache data, for example, the data is divided into multiple data blocks and transmitted to the GPU simultaneously through multiple channels. A parallel transmission data stream is obtained, such as a data stream in which multiple data blocks are transmitted to the GPU simultaneously. The parallel transmission data stream is transmitted to the target pre-trained large model, such as the GPT-3 model, for model training. And a performance monitoring tool, such as nvidia-smi, is used to monitor the data access hit rate, and the data access hit rate feedback is performed to generate cache hit rate data. For example, the number of accesses and cache hits of each data block is recorded, and the cache hit rate is calculated. According to the cache hit rate data, for example, the data preloading strategy is adjusted according to the cache hit rate to optimize the training data storage space. For example, data blocks with low cache hit rates are removed from the cache, and data blocks with high cache hit rates are loaded into the cache. Thereby, an optimized data storage strategy is obtained, such as dynamically adjusting the data storage location and cache strategy according to the data access pattern.
[0104] Preferably, step S1 comprises the following steps:
[0105] Step S11: extracting model internal parameters of the target pre-trained large model to obtain a multi-dimensional model parameter matrix and an input type parameter matrix, wherein the multi-dimensional model parameter matrix includes a model attention weight matrix, an embedding vector matrix, and an inter-layer activation parameter matrix;
[0106] Step S12: performing principal component dimensionality reduction processing according to the input type parameter matrix to generate low-dimensional model input type parameters;
[0107] Step S13: performing parameter feature discretization processing on the low-dimensional model input type parameters to generate discrete model input feature data;
[0108] Step S14: Perform cluster analysis based on the discrete model input feature data, and build a global metadata index to generate model input index structure data; perform cluster data sharding balance optimization based on the model input index structure data to generate modulus input shard data;
[0109] Step S15: Use the hash sharding algorithm to perform distributed storage node mapping on the modulus input shard data, thereby obtaining storage modulus shard data.
[0110] In the embodiment of the present invention, taking the BERT model as an example, the pre-trained model is loaded using the Transformers library of Python. The embedding vector matrix is obtained using the API provided by the library, such as model.embeddings.word_embeddings.weight, and the obtained matrix data is stored in the NumPy array format to form a multidimensional model parameter matrix. The method for obtaining the input type parameter matrix is similar to the previous one. For example, the input of the BERT model is a token sequence, and its type parameter matrix can represent the type of each token (such as part of speech, position encoding, etc.), which can be obtained by analyzing the input layer structure of the model. The input type parameter matrix is subjected to principal component analysis (PCA) dimensionality reduction processing using the Scikit-learn library of Python. For example, the PCA (n_components = k) function can be used to reduce the input type parameter matrix to k dimensions, where k is a preset dimension value, and a suitable k value needs to be selected according to the actual situation. For example, the k value can be selected so that the data after dimensionality reduction can retain 95% of the variance information of the original data. For each dimension, the value range of the dimension is divided into several intervals of equal width, and each interval is assigned a discrete value. The cut function of Python's Pandas library can be used to achieve equal-width discretization. The discretized data is stored as a NumPy array of integer type to generate discrete model input feature data. The KMeans(n_clusters=m) function of the Scikit-learn library is used to divide the data into m clusters, where m is a preset number of clusters. It is necessary to select a suitable m value according to the actual situation. For example, the elbow rule can be used to determine the optimal number of clusters. After clustering is completed, a global metadata index is constructed. The index records the coordinates of the center point of each cluster, the number of data points in the cluster, and other information. The Python dictionary or list data structure can be used to store index information and generate model input index structure data. According to the model input index structure data, for example, the data size of each cluster is calculated, and the cluster data sharding is balanced and optimized. For example, a cluster with a large amount of data can be divided into multiple shards, and a cluster with a small amount of data can be merged into one shard, so that the data volume of each shard is as balanced as possible. A hash sharding algorithm, such as a consistent hashing algorithm, is used to map the modulus input shard data to distributed storage nodes. For example, you can use Python's hashlib library to calculate the hash value of each shard data and map the hash value to different storage nodes. You can use distributed coordination services such as Redis or Zookeeper to manage the information of storage nodes to obtain storage modulus shard data.
[0111] Preferably, step S14 comprises the following steps:
[0112] Step S141: performing feature point similarity calculation based on the discrete model input feature data to generate input feature similarity data;
[0113] Step S142: performing cluster analysis on the discrete model input feature data by inputting feature similarity data to generate model input cluster data;
[0114] Step S143: performing hash coding processing on the model input cluster data to generate input cluster hash coding data;
[0115] Step S144: Perform hash feature mapping on the input cluster hash code data and the model input cluster data to obtain cluster hash mapping table data;
[0116] Step S145: construct a global metadata index based on the cluster hash mapping table data to generate model input index structure data;
[0117] Step S146: Utilize the model input index structure data to perform cluster data sharding balance optimization on the model input cluster data to generate modulus input shard data.
[0118] In an embodiment of the present invention, cosine similarity is used to calculate the similarity between feature points in the discrete model input feature data. For example, the cosine_similarity function in the Python Scikit-learn library can be used to calculate the cosine similarity between feature vectors. Based on the input feature similarity data, the discrete model input feature data is clustered and analyzed using the DBSCAN clustering algorithm. The DBSCAN algorithm can cluster according to the density between feature points without pre-specifying the number of clusters. For example, the DBSCAN (eps = e, min_samples = m) function in the Scikit-learn library can be used for clustering, where the eps parameter represents the neighborhood radius and the min_samples parameter represents the minimum number of points in the neighborhood. The clustering results are stored in a list form, each element represents a cluster, and the cluster contains the feature point index belonging to the cluster, generating the model input clustering cluster data. The model input clustering cluster data is hashed using the local sensitive hashing (LSH) algorithm. The LSH algorithm can map high-dimensional data to a low-dimensional space and maintain the similarity relationship between data points. Map the feature points of each cluster to a hash bucket, and use the hash bucket number as the hash code of the cluster to generate input cluster hash code data. Perform hash feature mapping on the input cluster hash code data and the model input cluster cluster data to build a cluster hash map. A hash map is a dictionary data structure, with the key as the hash code and the value as the corresponding cluster data. For example, you can use Python's dictionary type to store hash map data to obtain cluster hash map data. Build a global metadata index based on the cluster hash map data. The index records the cluster data information corresponding to each hash code, such as the number of data points in the cluster, the coordinates of the center point of the cluster, etc. Use the model input index structure data, for example, according to the size of the cluster data corresponding to each hash code, to optimize the cluster data sharding balance of the model input cluster cluster data. For example, clusters with large data volume can be divided into multiple shards, and clusters with small data volume can be merged into one shard, so that the data volume of each shard is as balanced as possible. Generate modulus input shard data, for example, store the data of each shard as a separate file.
[0119] Preferably, step S2 comprises the following steps:
[0120] Step S21: Perform distributed storage cluster topology analysis based on storage modulus sharding data to generate storage node topology data;
[0121] Step S22: Real-time monitoring of node resources of storage node topology data is performed through a distributed storage management system to obtain real-time load data of storage nodes;
[0122] Step S23: Based on the real-time load data of the storage node, the module input shard data is processed by the multi-dimensional model parameter matrix to obtain the intelligent shard storage strategy;
[0123] Step S24: Obtain target model training data;
[0124] Step S25: Based on the distributed storage management system, the target model training data is intelligently partitioned and stored using an intelligent partitioning storage strategy to obtain model training storage data.
[0125] In the embodiment of the present invention, taking the HDFS distributed storage system as an example, the topology information of the cluster is obtained using the HDFS API (e.g., hdfsdfsadmin-report), including the number of DataNode nodes, the storage capacity of each node, the network bandwidth and other information. The XML format data returned by the API can be parsed using the Python xml.etree.ElementTree library, the node topology information can be extracted, and the node topology information can be stored in JSON format to generate storage node topology data. The monitoring tool (e.g., yarn top) provided by the distributed storage management system (e.g., Hadoop YARN) is used to perform real-time resource monitoring on the nodes in the storage node topology data. For example, the CPU usage, memory usage, disk IO speed, network traffic and other indicators of each DataNode node can be collected regularly. Based on the real-time load data of the storage node and the multidimensional model parameter matrix, the module input shard data is processed by intelligent shard storage strategy. For example, according to the access frequency characteristics of the model parameter matrix, the shard data with high access frequency can be stored on the node with lower load to balance the node load and improve data access efficiency. According to the type of the target pre-trained large model, the corresponding training data is obtained. For example, if the target model is a BERT model for natural language processing, a large amount of text corpus needs to be obtained as training data; if the target model is a ResNet model for image recognition, a large amount of image data sets need to be obtained as training data. Training data can be obtained by using web crawlers, public data sets, etc. Different types of training data can be stored in different nodes or directories according to the characteristics of the training data, such as the subject of text data, the category of image data, etc., and the data storage location can be dynamically adjusted according to the load of the node, for example, text data with high access frequency can be stored on a storage node with higher performance, and image data with large data volume can be stored on a node with larger storage capacity. Use the HDFS API (such as hdfs dfs-put) to upload the training data to the HDFS cluster, and specify the data storage location according to the intelligent sharding storage strategy to obtain the model training storage data.
[0126] Preferably, step S23 includes the following steps:
[0127] Step S231: performing storage resource requirement evaluation on the module input shard data to generate shard storage resource requirement data;
[0128] Step S232: Perform storage node performance mapping analysis on storage node real-time load data using storage node topology data to generate storage node performance mapping data; perform shard storage pressure processing on storage node performance mapping data using shard storage resource demand data to generate shard load pressure data;
[0129] Step S233: performing storage node fitness calculation on the modulus input shard data based on the shard load pressure data, thereby obtaining storage node fitness data;
[0130] Step S234: performing intelligent shard node matching on the module input shard data through the multi-dimensional model parameter matrix and the storage node fitness data to obtain intelligent shard matching data;
[0131] Step S235: Perform shard storage migration according to the smart shard matching data, thereby obtaining a smart shard storage strategy.
[0132] In an embodiment of the present invention, the analysis modulus input shard data, for example, calculates the size of each shard data, access frequency and other information. The Python os library can be used to obtain the file size, and the access frequency can be counted in combination with the historical access log. According to the characteristics of the shard data, such as size, access frequency, etc., the storage resources required for each shard data, such as storage space, network bandwidth, etc., are evaluated. Using the storage node topology data, such as the storage capacity of the node, network bandwidth and other information, the storage node real-time load data, such as the CPU usage rate, memory usage rate, disk IO speed, network traffic and other indicators of the node, the storage node performance mapping analysis is performed. For example, a linear regression model can be used to map the node load index to the node performance index, such as data reading speed, writing speed, etc. Generate storage node performance mapping data, for example, store the performance index of each node as a vector. Through the shard storage resource demand data, such as the storage space required for each shard data, network bandwidth and other information, the storage node performance mapping data is subjected to shard storage pressure processing. For example, it can be simulated to store each shard data on each node, and the changes in the node load, such as the changes in CPU usage rate and memory usage rate, can be calculated to generate shard load pressure data. Based on the shard load pressure data, such as the load index of each node after storing each shard data, the storage node fitness calculation is performed on the modulus input shard data. For example, the fitness score of each shard data and each node can be calculated according to the changes in the node load, such as the changes in CPU usage and memory usage. The higher the score, the more suitable the shard data is for storage on the node. The fitness score can be calculated using the weighted average method, such as assigning different weights to indicators such as CPU usage and memory usage, and calculating the weighted average to obtain the storage node fitness data. Through the multi-dimensional model parameter matrix, such as the access frequency of the model parameters, data dependencies and other information, and the storage node fitness data, such as the fitness score of each shard data and each node, the modulus input shard data is intelligently matched with the shard nodes. For example, according to the access frequency of the model parameters, the shard data corresponding to the parameters with high access frequency can be stored on the node with high fitness score. A greedy algorithm or a genetic algorithm can be used for intelligent matching, such as allocating each shard data to the node with the highest fitness score in turn until all shard data are allocated. According to the intelligent shard matching data, such as the storage node ID corresponding to each shard data, shard storage migration is performed. For example, the API of the distributed storage management system, such as the hdfs dfs-mv command of HDFS, can be used to migrate the shard data from the current storage node to the target storage node. Thus, the intelligent shard storage strategy is obtained, such as storing the storage node ID and storage path corresponding to each shard data as a JSON file.
[0133] Preferably, step S232 includes the following steps:
[0134] Step S2321: Perform node historical shard load analysis based on the real-time load data of the storage node to obtain node historical shard load data;
[0135] Step S2322: Perform node load fluctuation analysis on the node historical shard load data to generate storage node load fluctuation data;
[0136] Step S2323: Perform storage node performance mapping analysis on the storage node topology data using the node historical shard load data to generate storage node performance mapping data;
[0137] Step S2324: performing node load carrying capacity evaluation according to the storage node performance mapping data and the storage node load fluctuation data to generate node load capacity evaluation data;
[0138] Step S2325: Use the shard storage resource demand data to perform shard demand-resource mapping on the node load capacity assessment data, and perform shard load pressure calculation to generate shard load pressure data.
[0139] In an embodiment of the present invention, historical load data of storage nodes is obtained from a monitoring system of a distributed storage management system (e.g., ResourceManager of Hadoop YARN), such as CPU usage, memory usage, disk IO speed, network traffic, and other indicators of each node in the past period of time. A time series database (e.g., InfluxDB) can be used to store historical load data. Node load fluctuation analysis is performed on the node historical shard load data. For example, the variance or standard deviation of the load index of each node during the storage of different shard data can be calculated to measure the degree of fluctuation of the node load. The variance or standard deviation can be calculated using Python's NumPy library. Storage node load fluctuation data is generated, such as storing the degree of load fluctuation of each node as a vector. Using the node historical shard load data, such as the load index of each node during the storage of each shard data, the storage node topology data, such as the storage capacity, network bandwidth, number of CPU cores, and other information of the node, is subjected to storage node performance mapping analysis. For example, a machine learning model, such as a regression model or a neural network model, can be used to map the node load index to the node performance index, such as data reading speed, writing speed, and the like. According to the storage node performance mapping data, such as the performance index of each node, and the storage node load fluctuation data, such as the load fluctuation degree of each node, the node load carrying capacity is evaluated. For example, the maximum load pressure that the node can bear can be calculated based on the performance index and load fluctuation degree of the node, such as the maximum value of indicators such as CPU usage and memory usage. Using the shard storage resource demand data, such as the storage space and network bandwidth required for each shard data, the node load capacity evaluation data, such as the load carrying capacity index of each node, is subjected to shard demand-resource mapping. For example, the storage space required for each shard data can be mapped to the available storage space of the node, and the network bandwidth required for each shard data can be mapped to the available network bandwidth of the node. And the shard load pressure calculation is performed. For example, according to the result of the shard demand-resource mapping, the load pressure caused by each shard data on each node can be calculated, such as the increase in indicators such as CPU usage and memory usage, and the load pressure calculation can be performed using queuing theory or simulation. The shard load pressure data is generated.
[0140] Preferably, step S234 includes the following steps:
[0141] Step S2341: performing a model parameter training impact assessment on the multi-dimensional model parameter matrix to generate model training impact weight data;
[0142] Step S2342: performing training data stream access frequency statistics according to the model training influence weight data to obtain model training access frequency data;
[0143] Step S2343: performing input slice data priority classification on the modulus input slice data through model training access frequency data to generate model input slice classification data;
[0144] Step S2344: constructing a node storage hierarchy according to the storage node performance mapping data to generate node storage hierarchy structure data;
[0145] Step S2345: performing multiple rounds of priority matching processing on the node storage hierarchical structure data and the model input shard classification data to generate preliminary shard node allocation data;
[0146] Step S2346: Use the storage node fitness data to perform node pressure balancing correction on the preliminary shard node allocation data to generate intelligent shard node matching data.
[0147] In an embodiment of the present invention, a multidimensional model parameter matrix, such as information such as the gradient value and update frequency of the model parameters, is analyzed to evaluate the degree of influence of each parameter on the model training. For example, the size of the gradient value or the parameter update frequency can be used as an evaluation index. The larger the gradient value or the higher the update frequency, the greater the influence of the parameter on the model training. The gradient value or the statistical parameter update frequency can be calculated using the Python NumPy library. According to the model training influence weight data, such as the training influence weight of each parameter, the access frequency of each shard data in the training data stream is statistically analyzed. For example, according to the corresponding relationship between the model parameters and the shard data, the training influence weight of the parameter can be accumulated to the corresponding shard data to obtain the access frequency weight of each shard data. The access frequency weight of the shard data can be stored using the Python dictionary data structure. Through the model training access frequency data, such as the access frequency weight of each shard data, the modulus input shard data is input shard data priority classification. For example, according to the access frequency weight of the shard data, the shard data can be divided into three categories: high priority, medium priority, and low priority. The Python Pandas library can be used to sort and classify the shard data to generate model input shard classification data. According to the storage node performance mapping data, such as the performance indicators of each node, such as data reading speed, writing speed, etc., the node storage hierarchy is constructed. For example, according to the performance indicators of the nodes, the nodes can be divided into three levels: high-speed storage layer, medium-speed storage layer, and low-speed storage layer. Multiple rounds of priority matching processing are performed on the node storage hierarchy structure data, such as the storage layer to which each node belongs, and the model input shard classification data, such as the priority category of each shard data. For example, the high-priority shard data can be first assigned to the nodes of the high-speed storage layer, and then the medium-priority shard data can be assigned to the nodes of the medium-speed storage layer, and finally the low-priority shard data can be assigned to the nodes of the low-speed storage layer. A greedy algorithm or a genetic algorithm can be used for multiple rounds of matching to generate preliminary shard node allocation data. Using the storage node fitness data, such as the fitness score of each shard data and each node, the preliminary shard node allocation data, such as the storage node ID corresponding to each shard data, is corrected for node pressure balance. For example, the shard data allocation scheme can be adjusted according to the node fitness score and load condition to make the node load more balanced. Linear programming or simulated annealing algorithm can be used for correction to generate intelligent sharding node matching data.
[0148] Preferably, step S3 comprises the following steps:
[0149] Step S31: scanning the model training storage data step by step, and calculating the data repetition rate to obtain data block repetition rate data;
[0150] Step S32: using a preset data redundancy threshold to perform redundant data block detection on the data block repetition rate data, marking the data blocks whose data block repetition rate data is higher than or equal to the preset data redundancy threshold as training redundant data; marking the data blocks whose data block repetition rate data is lower than the preset data redundancy threshold as training non-redundant data;
[0151] Step S33: dividing the training redundant data into redundancy levels to obtain high redundant data, medium redundant data and low redundant data;
[0152] Step S34: perform pointer redirection deduplication processing on high-redundancy data to generate high-redundancy compressed optimized data; perform data differential compression processing on medium-redundancy data to generate medium-redundancy differential compressed data; perform light compression processing on low-redundancy data to generate low-redundancy compressed data;
[0153] Step S35: performing lossless compression processing on the training non-redundant data to generate lossless non-redundant compressed data;
[0154] Step S36: constructing a redundant storage index according to the high-redundancy compression optimization data, the medium-redundancy difference compression data, the low-redundancy compression data, and the lossless non-redundant compression data, thereby obtaining dynamic redundant storage index data;
[0155] Step S37: Optimize data storage and retrieval of the distributed storage management system through dynamic redundant storage index data to generate optimized storage training data.
[0156] In an embodiment of the present invention, a sliding window is used to perform step-by-step data block scanning on the model training storage data. For example, the training data can be divided into data blocks of equal size, and each data block can be scanned in turn using a fixed-size window. The data block can be read using Python's io library, and the data block scanning can be implemented using a sliding window algorithm. During the scanning process, the repetition rate of each data block is calculated. For example, the Simhash algorithm can be used to calculate the fingerprint of each data block, and the similarity between the fingerprints of different data blocks can be compared to calculate the repetition rate of the data block, and the similarity between the fingerprints can be calculated using the Hamming distance to obtain the data block repetition rate data. Using a preset data redundancy threshold, such as 0.8, redundant data block detection is performed on the data block repetition rate data. For example, data blocks with a repetition rate greater than or equal to 0.8 are marked as training redundant data, and data blocks with a repetition rate less than 0.8 are marked as training non-redundant data. The training redundant data is divided into redundancy levels. For example, based on the repetition rate of data blocks, data blocks with a repetition rate higher than 0.95 can be marked as high-redundancy data, data blocks with a repetition rate between 0.85 and 0.95 can be marked as medium-redundancy data, and data blocks with a repetition rate between 0.8 and 0.85 can be marked as low-redundancy data. The redundancy level division can be implemented using Python's if-elif-else statement. Pointer redirection deduplication processing is performed on high-redundancy data. For example, only one high-redundancy data block is retained, and other repeated data blocks are replaced with pointers pointing to the data block. The Python dictionary data structure can be used to store data block pointers to generate high-redundancy compression optimization data. Data differential compression processing is performed on medium-redundancy data. For example, only the difference between data blocks is stored, and the difference data is compressed using a compression algorithm. The Python xdelta3 library can be used for differential compression. Lightweight compression processing is performed on low-redundancy data. For example, the data block is compressed using the gzip algorithm. The Python gzip library can be used for light compression. Lossless compression processing is performed on training non-redundant data. For example, the data block is compressed using the Snappy algorithm. You can use Python's snappy library for lossless compression. Build a redundant storage index based on high-redundancy compression optimization data, medium-redundancy differential compression data, low-redundancy compression data, and lossless non-redundant compression data. The index records the storage location, compression method, redundancy level, and other information of each data block. You can use the LevelDB database or the RocksDB database to store index information. By dynamically redundantly storing index data, you can optimize data storage and retrieval for distributed storage management systems, such as HDFS. For example, when reading a data block, first query the index, find the storage location and compression method of the data block based on the index information, and then read the data block and decompress it.You can use the HDFS API to read data blocks and use the corresponding decompression algorithm to decompress the data blocks to generate optimized storage training data.
[0157] As an example of the present invention, refer to Figure 2 As shown, Figure 1 Detailed implementation steps of step S4 are shown in the flowchart. In this example, step S1 includes:
[0158] Step S41: monitoring the training process of the target pre-trained large model in real time, identifying the model training stage, and obtaining model training stage identification data;
[0159] In an embodiment of the present invention, a tool such as TensorBoard is used to monitor the training process of the target pre-trained large model in real time. For example, the changes in indicators such as the loss function value and accuracy of the model can be monitored. The Python TensorBoard library or other monitoring tools can be used to monitor the training process. According to the changing trend of the monitoring indicators, the different stages of model training are identified, for example: Initial stage: When the model training just starts, the loss function value drops rapidly and the accuracy rate increases rapidly. Stable stage: When the model training enters the stable period, the loss function value and the accuracy rate change slowly. Convergence stage: When the model training is close to convergence, the loss function value and the accuracy rate tend to be stable. Different indicator thresholds or change trends can be set according to actual conditions to identify the model training stage. For example, it can be set that when the loss function value drops to a certain level or the change amount of multiple consecutive epochs is less than a certain threshold, the model is considered to have entered a stable stage.
[0160] Step S42: Tracking the training data input according to the model training phase identification data to generate model training input feature data;
[0161] In the embodiment of the present invention, the training data used by the model in each training stage is tracked according to the model training stage identification data obtained in step S41. For example, the data sample ID used by each training batch can be recorded and associated with the model training stage identification data to generate model training input feature data.
[0162] Step S43: Performing training data call sequence reasoning on the optimized stored training data through model training input feature data to generate data call sequence data;
[0163] In an embodiment of the present invention, by inputting feature data for model training, such as the training data ID read in each training stage, training data call order reasoning is performed on the optimized stored training data, such as the training data stored in a distributed file system. For example, the calling order of the training data can be inferred according to the order in which the training data is read in different training stages. A Markov chain model or a recurrent neural network model can be used to perform call order reasoning to generate data call order data. For example, reasoning can be performed according to the following rules: Time sequence: During model training, data is usually read in sequence in the order in which the data is stored. Access frequency: Data shards with high access frequency are read continuously. Training stage: Different training stages will have different data access patterns. For example, in the early stage of model training, all data shards need to be frequently accessed, while in the later stage of model training, only some data shards need to be accessed.
[0164] Step S44: Perform training phase resource demand analysis based on the model training phase identification data to generate training resource demand feature data;
[0165] In the embodiment of the present invention, based on the model training stage identification data, such as the training stage at each time point, the training stage resource requirement analysis is performed. For example, the computing resources, storage resources, network resources, etc. required by the model at each training stage can be analyzed. A performance analysis tool, such as Profiler, can be used to analyze the resource usage of the model and generate training resource requirement feature data.
[0166] Step S45: performing key data identification on the data call sequence data, and performing fast access processing according to the training resource demand characteristic data to generate key access training data;
[0167] In an embodiment of the present invention, key data identification is performed on data call sequence data, such as a list of training data IDs stored in a predicted call order. For example, training data with a high call frequency can be marked as key data. The Python Counter class can be used to count the call frequency of training data. And according to the training resource demand characteristic data, such as the resource demand for each training stage, fast access volume processing is performed. For example, the fast access volume of key data can be determined based on the computing resource demand for the training stage, such as how much key data needs to be loaded into the memory or cache. A linear regression model or other machine learning model can be used to predict the fast access volume and generate key access training data.
[0168] Step S46: Allocate the key access training data to a fast access area, perform data preloading processing, and generate model training preloading cache data.
[0169] In an embodiment of the present invention, key access training data, such as the ID of key data and the corresponding fast access amount, are allocated to a fast access area. For example, key data can be stored in a fast access area such as a memory, an SSD hard disk, or a GPU video memory. Storage management software, such as Ceph, can be used to allocate fast access areas. And data preloading processing is performed. For example, before the model training starts, the key data is preloaded into the fast access area. Data preloading can be implemented using Python's shutil library or other file operation libraries to generate model training preload cache data.
[0170] As an example of the present invention, refer to Figure 2 As shown, Figure 1 Detailed implementation steps of step S5 in the flowchart, in this example, step S5 includes:
[0171] Step S51: Allocate data transmission tasks for the model training preloaded cache data to obtain transmission task allocation data;
[0172] In the embodiment of the present invention, the data transmission task is decomposed into multiple subtasks and assigned to different data transmission channels according to factors such as the size of the pre-loaded cache data for model training, the parallelism of the target pre-trained large model, and the available network bandwidth. For example, the cache data can be divided into multiple data blocks, and each data block can be assigned to a data transmission channel. A load balancing algorithm, such as a polling algorithm or a weighted polling algorithm, can be used to perform task allocation and obtain transmission task allocation data.
[0173] Step S52: Allocate data according to the transmission task to perform multi-channel parallel transmission of data streams to obtain parallel transmission data streams;
[0174] In an embodiment of the present invention, an RDMA library (such as librdmacm) is used to initialize four RDMA channels and establish a connection with a target computing node. Parallel transmission of data blocks: Data is allocated according to the transmission task, and each data block is transmitted to the corresponding GPU memory of the target computing node through the corresponding RDMA channel. For example, the first data block is transmitted to the video memory of GPU0 through channel 0, the second data block is transmitted to the video memory of GPU1 through channel 1, and so on. The API provided by the RDMA library (such as ibv_post_send) can be used to implement multi-channel parallel transmission of data streams to obtain parallel transmission data streams.
[0175] Step S53: transmitting the parallel transmission data stream to the target pre-trained large model for model training, and performing model training delay evaluation to obtain model training delay data;
[0176] In an embodiment of the present invention, the parallel transmission data stream is transmitted to the target pre-trained large model, such as GPT-3 or BERT, for model training. A deep learning framework, such as TensorFlow or PyTorch, can be used to load the model and train it. During the model training process, the time interval between each data block being read from the cache and the model starting to use it, i.e., the data transmission delay, is recorded.
[0177] Step S54: performing data consistency check on the parallel transmission data stream to obtain consistency verification data;
[0178] In an embodiment of the present invention, data consistency check is performed on the parallel transmission data stream. For example, the checksum of each data block can be calculated, and the checksum can be calculated again after the data transmission is completed, and the two calculation results can be compared to see whether they are consistent. The checksum can be calculated using the hashlib library of Python to obtain consistency verification data.
[0179] Step S55: performing data access hit rate feedback according to the consistency verification data to generate cache hit rate data;
[0180] In an embodiment of the present invention, data access hit rate feedback is performed based on consistency verification data, such as the result of whether the checksum of each data block is consistent. For example, if the checksum of the data block is consistent, it means that the data transmission is successful and the cache hits; otherwise, it means that the data transmission fails and the cache misses. Python's if-else statement can be used to determine whether the checksum is consistent, and to count the cache hit rate to generate cache hit rate data.
[0181] Step S56: Dynamically adjust the cache using cache hit rate data and model training delay data, and optimize the training data storage space to obtain an optimized data storage strategy.
[0182] In an embodiment of the present invention, cache is dynamically adjusted through cache hit rate data, such as cache hit rate, and model training delay data, such as the training time of each training batch. For example, if the cache hit rate is low, the cache capacity can be increased or the cache replacement strategy can be adjusted; if the model training delay is high, more data can be loaded into the cache. Cache replacement strategies such as LRU (Least Recently Used) or LFU (Least Frequently Used) can be used. The global usage of the storage system is analyzed to identify historical data with low frequency of access, and the data that has been trained and used by the model or has not been accessed is archived or deleted to ensure that all data has been used by the model, while intelligently releasing storage space to ensure the long-term stability and storage efficiency of the system under high load.
[0183] The beneficial effect of the present application is that, through the calculation of data block repetition rate and redundancy level division, a differentiated compression strategy is adopted for data blocks with different redundancy levels. For high redundant data, pointer redirection is used to deduplicate, and only one copy of the data is retained, which effectively reduces the storage space occupied; for medium redundant data, data differential compression is adopted, such as differential encoding, which only stores the difference information between data blocks to further compress the data volume; for low redundant data, a lightweight compression algorithm, such as LZ4, is adopted to strike a balance between compression efficiency and compression speed. By refining the compression of training data, the amount of data storage is effectively reduced, and the problem of storage resource waste is alleviated. In the data storage stage, the solution analyzes the access mode of model parameters and the performance differences of storage nodes, formulates an intelligent sharding storage strategy, and evenly distributes data to different nodes and storage media to avoid data hot spots and access congestion. In the data access stage, combined with the model training stage and the data call order, key data is identified and preloaded into the cache, so that the required data can be quickly obtained during the model training process, significantly reducing data access latency. Through the multi-channel parallel transmission technology of data streams, the data access speed is further improved. The solution divides the training data preloaded into the cache into multiple transmission tasks and assigns them to different data transmission channels to achieve parallel data transmission, fully utilize bandwidth resources, maximize data transmission efficiency, and solve the data storage and access bottleneck problems in the large model training process.
[0184] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is therefore intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.
[0185] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A data storage method based on a large model, characterized in that: The following steps are involved: Step S1: extracting the internal parameters of the target pre-trained large model to obtain a multi-dimensional model parameter matrix and an input type parameter matrix; Perform input cluster sharding optimization according to the input type parameter matrix to generate modulus input sharding data; Perform distributed storage node mapping on the modulus input shard data to obtain storage modulus shard data; Step S2: performing intelligent sharding storage strategy processing on the storage modulus sharding data through the multidimensional model parameter matrix, thereby obtaining an intelligent sharding storage strategy; Obtain target model training data; perform intelligent sharding storage on the target model training data using an intelligent sharding storage strategy based on a distributed storage management system to obtain model training storage data; Step S3: performing redundant data block detection on the model training storage data to obtain training redundant data and training non-redundant data respectively; Redundant compression optimization is performed on training redundant data and training non-redundant data, and redundant storage index is constructed to obtain dynamic redundant storage index data; Optimize data storage and retrieval based on dynamic redundant storage index data to generate optimized storage training data; Step S4: performing training data call sequence reasoning on the optimized stored training data to generate data call sequence data; Identify key data of data call sequence data, perform fast access processing, and generate key access training data; perform data preloading processing on key access training data to generate model training preloading cache data; Step S5: Perform multi-channel parallel transmission of the data stream for model training pre-loaded cache data to obtain a parallel transmission data stream; transmit the parallel transmission data stream to the target pre-trained large model for model training, and perform data access hit rate feedback to generate cache hit rate data; optimize the training data storage space according to the cache hit rate data, thereby obtaining an optimized data storage strategy.
2. The data storage method based on a large model according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: extracting model internal parameters of the target pre-trained large model to obtain a multi-dimensional model parameter matrix and an input type parameter matrix, wherein the multi-dimensional model parameter matrix includes a model attention weight matrix, an embedding vector matrix, and an inter-layer activation parameter matrix; Step S12: performing principal component dimensionality reduction processing according to the input type parameter matrix to generate low-dimensional model input type parameters; Step S13: performing parameter feature discretization processing on the low-dimensional model input type parameters to generate discrete model input feature data; Step S14: Perform cluster analysis based on the discrete model input feature data, and build a global metadata index to generate model input index structure data; perform cluster data sharding balance optimization based on the model input index structure data to generate modulus input shard data; Step S15: Use the hash sharding algorithm to perform distributed storage node mapping on the modulus input shard data, thereby obtaining storage modulus shard data.
3. The data storage method based on a large model according to claim 2 is characterized in that: Step S14 includes the following steps: Step S141: performing feature point similarity calculation based on the discrete model input feature data to generate input feature similarity data; Step S142: performing cluster analysis on the discrete model input feature data by inputting feature similarity data to generate model input cluster data; Step S143: performing hash coding processing on the model input cluster data to generate input cluster hash coding data; Step S144: Perform hash feature mapping on the input cluster hash code data and the model input cluster data to obtain cluster hash mapping table data; Step S145: construct a global metadata index based on the cluster hash mapping table data to generate model input index structure data; Step S146: Utilize the model input index structure data to perform cluster data sharding balance optimization on the model input cluster data to generate modulus input shard data.
4. The data storage method based on a large model according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: Perform distributed storage cluster topology analysis based on storage modulus sharding data to generate storage node topology data; Step S22: Real-time monitoring of node resources of storage node topology data is performed through a distributed storage management system to obtain real-time load data of storage nodes; Step S23: Based on the real-time load data of the storage node, the module input shard data is processed by the multi-dimensional model parameter matrix to obtain the intelligent shard storage strategy; Step S24: Obtain target model training data; Step S25: Based on the distributed storage management system, the target model training data is intelligently partitioned and stored using an intelligent partitioning storage strategy to obtain model training storage data.
5. The data storage method based on a large model according to claim 4 is characterized in that: Step S23 includes the following steps: Step S231: performing storage resource requirement evaluation on the module input shard data to generate shard storage resource requirement data; Step S232: Perform storage node performance mapping analysis on storage node real-time load data using storage node topology data to generate storage node performance mapping data; perform shard storage pressure processing on storage node performance mapping data using shard storage resource demand data to generate shard load pressure data; Step S233: performing storage node fitness calculation on the module input shard data based on the shard load pressure data, thereby obtaining storage node fitness data; Step S234: performing intelligent shard node matching on the modulus input shard data through the multi-dimensional model parameter matrix and the storage node fitness data to obtain intelligent shard matching data; Step S235: Perform shard storage migration according to the smart shard matching data, thereby obtaining a smart shard storage strategy.
6. The data storage method based on a large model according to claim 5 is characterized in that: Step S232 includes the following steps: Step S2321: Perform node historical shard load analysis based on the real-time load data of the storage node to obtain node historical shard load data; Step S2322: Perform node load fluctuation analysis on the node historical shard load data to generate storage node load fluctuation data; Step S2323: Perform storage node performance mapping analysis on the storage node topology data using the node historical shard load data to generate storage node performance mapping data; Step S2324: performing node load carrying capacity evaluation according to the storage node performance mapping data and the storage node load fluctuation data to generate node load capacity evaluation data; Step S2325: Use the shard storage resource demand data to perform shard demand-resource mapping on the node load capacity assessment data, and perform shard load pressure calculation to generate shard load pressure data.
7. The data storage method based on a large model according to claim 5, characterized in that: Step S234 includes the following steps: Step S2341: performing a model parameter training impact assessment on the multi-dimensional model parameter matrix to generate model training impact weight data; Step S2342: performing training data stream access frequency statistics according to the model training influence weight data to obtain model training access frequency data; Step S2343: performing input slice data priority classification on the modulus input slice data through model training access frequency data to generate model input slice classification data; Step S2344: constructing a node storage hierarchy according to the storage node performance mapping data to generate node storage hierarchy structure data; Step S2345: performing multiple rounds of priority matching processing on the node storage hierarchical structure data and the model input shard classification data to generate preliminary shard node allocation data; Step S2346: Use the storage node fitness data to perform node pressure balancing correction on the preliminary shard node allocation data to generate intelligent shard node matching data.
8. The data storage method based on a large model according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: scanning the model training storage data step by step, and calculating the data repetition rate to obtain data block repetition rate data; Step S32: using a preset data redundancy threshold to perform redundant data block detection on the data block repetition rate data, marking the data blocks whose data block repetition rate data is higher than or equal to the preset data redundancy threshold as training redundant data; marking the data blocks whose data block repetition rate data is lower than the preset data redundancy threshold as training non-redundant data; Step S33: dividing the training redundant data into redundancy levels to obtain high redundant data, medium redundant data and low redundant data; Step S34: perform pointer redirection deduplication processing on high-redundancy data to generate high-redundancy compressed optimized data; perform data differential compression processing on medium-redundancy data to generate medium-redundancy differential compressed data; perform light compression processing on low-redundancy data to generate low-redundancy compressed data; Step S35: performing lossless compression processing on the training non-redundant data to generate lossless non-redundant compressed data; Step S36: constructing a redundant storage index according to the high-redundancy compression optimization data, the medium-redundancy difference compression data, the low-redundancy compression data, and the lossless non-redundant compression data, thereby obtaining dynamic redundant storage index data; Step S37: Optimize data storage and retrieval of the distributed storage management system through dynamic redundant storage index data to generate optimized storage training data.
9. The data storage method based on a large model according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: monitoring the training process of the target pre-trained large model in real time, identifying the model training stage, and obtaining model training stage identification data; Step S42: Tracking the training data input according to the model training phase identification data to generate model training input feature data; Step S43: Performing training data call sequence reasoning on the optimized stored training data through model training input feature data to generate data call sequence data; Step S44: Perform training phase resource demand analysis based on the model training phase identification data to generate training resource demand feature data; Step S45: performing key data identification on the data call sequence data, and performing fast access processing according to the training resource demand characteristic data to generate key access training data; Step S46: Allocate the key access training data to a fast access area, perform data preloading processing, and generate model training preloading cache data.
10. The data storage method based on a large model according to claim 1, characterized in that: Step S5 includes the following steps: Step S51: Allocate data transmission tasks for the model training preloaded cache data to obtain transmission task allocation data; Step S52: Allocate data according to the transmission task to perform multi-channel parallel transmission of data streams to obtain parallel transmission data streams; Step S53: transmitting the parallel transmission data stream to the target pre-trained large model for model training, and performing model training delay evaluation to obtain model training delay data; Step S54: performing data consistency check on the parallel transmission data stream to obtain consistency verification data; Step S55: performing data access hit rate feedback according to the consistency verification data to generate cache hit rate data; Step S56: Dynamically adjust the cache using cache hit rate data and model training delay data, and optimize the training data storage space to obtain an optimized data storage strategy.
Citation Information
Patent Citations
Dynamic data storage method and device
CN104516912A
Deep hash pedestrian re-identification method based on data enhancement
CN110852152A