AI data loading method using dynamic metadata strategy
By converting the data set into Lafa format and dividing it into sharded files, combining dynamic metadata strategies and real-time evaluation of decision-making modules, the problem of memory overflow during large-scale data sets is solved, and efficient data loading and resource utilization are achieved.
Patent Information
- Application Number
- CN202510259371.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-13
AI Technical Summary
During the large-scale data set loading process, although the on-demand loading method avoids memory overflow problems caused by one-time data loading, the storage volume of sample indexes increases linearly with the data scale, resulting in an exponential increase in memory usage, which may eventually lead to memory exhaustion.
Using dynamic metadata strategy, data sets of different formats are uniformly converted into Lafa format through the FileWriter interface, and the data sets of Lafa format are divided into several shard files for parallel storage and retrieval. Use the decision module to evaluate the system's memory resources and dataset characteristics in real time, dynamically obtain the adapted data loading mode, calculate the mode switching threshold, and select the appropriate metadata loading logic for processing and cache.
By dynamically adjusting the data loading mode and metadata loading logic, the balance between loading performance and memory overhead is achieved, avoiding the problem of memory overflow when training large-scale data sets, and optimizing the data shuffling mechanism, significantly improving the efficiency of large-scale data loading and resource utilization.
Smart Images

Figure CN120144653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data governance, and particularly to an AI data loading method using a dynamic metadata strategy. Background Art
[0002] With the rapid development of artificial intelligence technology, the scale of deep learning models has been continuously expanding, and the demand for training large-scale datasets has been increasing day by day. Training large models not only requires constructing a complex parameter space but also relies on a vast amount of training data. However, in the current model training process, there are significant bottlenecks in loading large-scale datasets.
[0003] The traditional data loading process needs to load a vast amount of metadata into memory at one time and generate tensors that can be input into the network by combining operations such as data augmentation and format conversion. When the model training starts, data loading, network model compilation, and computing tasks often need to be carried out simultaneously. Due to the excessive memory resources occupied by dataset loading, network computing or compilation tasks often cannot obtain sufficient memory resources during the training process, resulting in OOM errors, especially when dealing with datasets of TB or even PB level, the problem is more serious.
[0004] Mainstream deep learning frameworks (such as PyTorch, TensorFlow, and MindSpore) process large-scale datasets as individual samples, associate sample indexes with file cache paths, and read them into memory in a demand-loading manner. This method avoids one-time data loading and reduces memory occupancy. However, this metadata loading logic usually means that the memory occupancy of metadata grows linearly with the increase in the number of data samples. For datasets with tens of millions or hundreds of millions of scales, the number of sample indexes will also reach a very large scale, and ultimately may lead to memory exhaustion. Summary of the Invention
[0005] The purpose of the present invention is to provide an AI data loading method using a dynamic metadata strategy to solve the following technical problems:
[0006] During the loading process of large-scale datasets, the demand-loading method, although avoiding the memory overflow problem caused by one-time data loading, the storage volume of sample indexes grows linearly with the data scale, resulting in an exponential increase in memory occupancy and ultimately leading to exhaustion of memory consumption.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] An AI data loading method using a dynamic metadata strategy includes the following steps:
[0009] Unify datasets in different formats into the Lafa format through the FileWriter interface, and divide the Lafa-format datasets into several shard files for parallel storage and retrieval. Each of the shard files includes metadata and data blocks.
[0010] Use the decision-making module to evaluate the memory resources of the system and the characteristics of the Lafa-format datasets in real time, dynamically obtain the appropriate data loading mode, and calculate the mode switching threshold. The data loading modes include fast loading, lazy loading, and slow loading.
[0011] According to the selected data loading mode, adopt the corresponding metadata loading logic to process and cache the metadata.
[0012] As a further solution of the present invention: the process of unifying the datasets in different formats into the Lafa format and performing parallel storage and retrieval is as follows:
[0013] According to the format of the input dataset, use the corresponding data reading method to unify the datasets in different formats into the Lafa format, preprocess the data, and construct a data structure that meets the requirements of the Lafa format. After converting the constructed Lafa-format data into a string form, write it into a file. According to the data volume, storage, and retrieval requirements of the generated Lafa-format file, divide the Lafa-format file into several shard files, generate corresponding index files for each shard, and store these shard files in different locations respectively.
[0014] As a further solution of the present invention: the process of preprocessing the data is as follows:
[0015] Clean the data through the OpenRefine tool, fill in the missing values of the data and filter the outliers of the data, and perform data transformation. The data transformation includes standardization, normalization, discretization, and feature encoding. Integrate the data based on the distributed data warehouse, and at the same time perform feature engineering to generate new features and dimensionality reduction, and perform reduction processing on the data.
[0016] As a further solution of the present invention: the process of dynamically obtaining the appropriate data loading mode is as follows:
[0017] The decision-making module obtains the current usage of the hardware memory in real time through the monitoring interface. The memory usage includes the used memory, available memory, and memory occupancy trend. It analyzes the metadata of the Lafa files to obtain information about the dataset, which includes the total sample size, the number of shards, and the sample distribution. By loading a number of Lafa files of the same scale, it respectively tests the memory growth rate under fast loading, lazy loading, and slow loading. It compares the memory growth rates under the three loading modes with the information of the dataset to obtain the loading mode that is most suitable for the current hardware conditions and dataset scale. The decision-making module continuously monitors the memory usage and the effect of the current loading mode, and dynamically adjusts the loading mode according to the real-time monitoring situation.
[0018] As a further solution of the present invention: The process of respectively testing the memory growth rate under fast loading, lazy loading, and slow loading is as follows:
[0019] Step 1: Select a number of Lafa files with exactly the same size, use a memory analysis tool to monitor the memory occupancy in real time, and record the initial memory value;
[0020] Step 2: Use the fast loading mode to preload the metadata of all Lafa files into the memory, traverse all Lafa file paths and store them in the memory, and obtain the memory change after loading is completed;
[0021] Step 3: Use the lazy loading mode to preload the key metadata of the Lafa files, load the detailed metadata as needed, randomly select a number of file indexes, read the file content and cache the key metadata into the memory, traverse all Lafa files, and record the memory change after loading is completed;
[0022] Step 4: Use the slow loading mode to cache the high-level shard information, dynamically generate the detailed metadata, select a number of high-level shard files, use the sample ID generator to dynamically generate the required metadata, and cache the metadata into the memory, and record the memory value after loading is completed;
[0023] Step 5: Remove the memory occupancy outliers, and according to the formula: memory growth rate = (final memory - initial memory) / total time, calculate the memory growth rates under the three loading modes respectively.
[0024] As a further solution of the present invention: The process of the sample ID generator generating local IDs is as follows:
[0025] Traverse the sample range within each shard file, bounded by start and end, where start represents the starting point of the sample and end represents the ending point of the sample. Dynamically preset shufflesize according to the memory size and the number of samples, where shufflesize represents the shuffle size. Based on the selected start and end, the sample ID generator generates local ID samples as needed, marks the range of the generated local ID samples as [start, start + shufflesize), and temporarily caches them in memory. When the number of local ID samples in memory reaches the upper limit, the sample ID generator automatically updates start to the starting point of the next interval and continues to generate new local ID intervals until all sample IDs within the shard are generated and cached.
[0026] As a further solution of the present invention: The process of dynamically generating metadata in the slow loading mode is as follows:
[0027] Define the global index range of samples within the shard file through start and end, use the sample ID generator to generate the global ID of the sample, and compare the global ID of the sample with the [start, end) ranges of all shards to obtain the partition to which the target sample ID belongs. Generate a data mapping based on the shard to which the sample ID belongs, and obtain the task information that the sample needs to execute according to the type, usage, and system configuration of the sample. Based on the shard information, data mapping information, and task information, obtain the metadata of the sample.
[0028] As a further solution of the present invention: The shuffle process in the slow loading mode is as follows:
[0029] Use the selected shuffle algorithm to shuffle each local ID temporarily cached in memory. Based on the shuffled sample ID sequence, read sample data from the storage medium batch by batch and in order according to the characteristics of the slow loading mode. When all samples within the current partition are processed, the sample ID generator will automatically update the next local ID and temporarily store it in memory, and the shuffle algorithm continues to shuffle the local IDs of the next partition.
[0030] The beneficial effects of the present invention:
[0031] The present invention uniformly converts data sets into the Lafa format through the FileWriter interface, ensuring the unity of different data set types, facilitating the interoperability of multi-source data and subsequent processing efficiency. Then, the Lafa-format files are divided into several shard files for parallel storage and retrieval, which not only enhances the scalability of the data but also facilitates the efficient positioning and access to specific data subsets. Further, the decision-making module dynamically decides the best loading mode according to the sample size and memory resource conditions, achieving a balance between loading performance and memory overhead. Finally, the metadata is processed and cached according to the selected loading mode, effectively avoiding the problem of out-of-memory when training large-scale data sets, and optimizing the data shuffling mechanism, significantly improving the large-scale data loading efficiency and resource utilization rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The present invention will be further described below with reference to the accompanying drawings.
[0033] Figure 1 is a schematic flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] Please refer to Figure 1 as shown, the present invention is an AI data loading method using a dynamic metadata strategy, including the following steps:
[0036] Select the corresponding reader according to the format of the input data. For example, for a CSV file, use a CSV reader to read the data line by line and split each line of data into fields; for a JSON file, use a JSON parser to parse it into an object tree; for an XML file, use an XML parser to parse it into a DOM (Document Object Model) tree.
[0037] After reading the data, the system maps it to the structure of the Lafa format. The Lafa format defines a set of unified data structures, including fields, records, and data sets. During the mapping process, the system maps the fields of the input data to the corresponding fields in the Lafa format according to the semantics and types of the data. For example, map the column names in the CSV file to the field names in the Lafa format, and map the key-value pairs of the JSON object to the records in the Lafa format.
[0038] After the mapping is completed, the system encodes the data to ensure that it meets the storage requirements of the Lafa format. The encoding process includes operations such as data type conversion, character encoding conversion, and compression. For example, converting string-type data to a byte array, encoding characters using UTF-8 encoding, and compressing the data using a compression algorithm to reduce storage space.
[0039] Finally, the system uses the FileWriter interface to write the encoded data into a file in the Lafa format. The FileWriter interface organizes the data according to the specifications of the Lafa format, including parts such as the file header, data blocks, and index. The file header contains metadata of the file, such as the file format version, dataset size, and field definitions. The data blocks store the actual data records, and the index is used to quickly locate the data records.
[0040] As the data volume continues to grow, a single Lafa-format file may become extremely large, which will bring performance bottlenecks to storage and retrieval. To solve this problem, we divide the Lafa-format file into several shard files for storage. Shard storage can split a large file into multiple small files, improving the flexibility and scalability of storage, and also facilitating parallel processing.
[0041] The system divides the Lafa-format file into several shard files according to factors such as the size, distribution, and access pattern of the data. The partitioning algorithm can use a simple equal-division method or can be intelligently partitioned according to the characteristics of the data. For example, for time-series data, it can be divided according to the time range; for spatial data, it can be divided according to the geographical location. Each shard file consists of two parts: metadata and data blocks. The metadata contains descriptive information about the shard file, such as the shard file number, the starting position and size of the data block, the type and range of the data, etc. The role of the metadata is to help the system quickly locate and understand the content of the data block, improving the efficiency of data retrieval. The data blocks store the actual data records, which are organized and stored according to the specifications of the Lafa format.
[0042] After partitioning the sharded files, the system stores these sharded files on different storage devices or storage nodes in parallel. Parallel storage can be achieved through multi-threading or a distributed storage system, improving the speed and efficiency of storage. For example, in a distributed file system, each sharded file can be stored on a different node, and data is transferred in parallel through the network. When retrieving data, the system determines the sharded files that need to be accessed based on the query conditions. Then, the system retrieves data from these sharded files in parallel. Parallel retrieval can be achieved through multi-threading or a distributed query engine, improving the speed and efficiency of retrieval. For example, in a distributed database, the query request can be sent to multiple nodes simultaneously, and each node processes the query request in parallel and returns the results to the main node for merging.
[0043] During the data processing process, memory resources are a key limiting factor. If all data is loaded into memory at once, it may cause a memory overflow, affecting the performance and stability of the system. Therefore, we need to dynamically select an appropriate data loading mode based on the system's memory resources and the characteristics of the dataset.
[0044] The decision-making module is responsible for evaluating the system's memory resources and the characteristics of the dataset in real time. It monitors the memory usage of the system to obtain the size and change trend of the available memory. At the same time, it analyzes the characteristics of the dataset, such as the size of the dataset, the access frequency of the data, the distribution and correlation of the data, and tests the memory growth rate under fast loading, lazy loading, and slow loading. It compares the memory growth rates under the three loading modes with the information of the dataset to obtain the loading mode that is most suitable for the current hardware conditions and dataset size.
[0045] The fast loading mode is suitable for cases where the dataset is small and the data access frequency is high. In this mode, the system loads all data into memory at startup. In this way, during subsequent data access, the system can directly obtain data from memory, avoiding frequent disk I / O operations and improving the speed of data access.
[0046] The lazy loading mode is suitable for cases where the dataset is large and the data access frequency is low. In this mode, the system does not load all data into memory at once, but loads a certain part of the data into memory only when it is needed. This can save memory resources and avoid memory overflow. For example, when a user queries a specific data record, the system reads the record from disk and loads it into memory.
[0047] The slow loading mode is applicable to the situation where the dataset is large and the data access frequency is not high, but the system has enough time for data loading. In this mode, the system will gradually load the data into the memory in the background. For example, when the system is idle, the system will read the data from the disk into the memory according to a certain strategy (such as the order of data blocks). In this way, when the user needs to access the data, most of the data is already in the memory, improving the response speed of data access.
[0048] To achieve a smooth switch between different data loading modes, the decision-making module needs to calculate the mode switch threshold. The mode switch threshold is dynamically calculated based on the memory resources of the system and the characteristics of the dataset. For example, when the available memory of the system is lower than a certain threshold, the decision-making module will switch the loading mode from fast loading to lazy loading; when the data access frequency exceeds a certain threshold, the decision-making module will switch the loading mode from lazy loading to fast loading. The calculation of the mode switch threshold needs to comprehensively consider factors such as the performance of the system, memory usage, and data access patterns to ensure that the system can maintain optimal performance in different situations.
[0049] According to the selected data loading mode, the system adopts the corresponding metadata loading logic. In the fast loading mode, the system will load all metadata into the memory at startup. In this way, during subsequent data access, the system can quickly obtain metadata information, improving the efficiency of data retrieval. In the lazy loading mode, the system will load the metadata of a certain data block only when it is needed to access that data block. This can save memory resources and avoid loading too much metadata at once. In the slow loading mode, the system will preload the metadata in batches to ensure that the metadata is ready when the user needs to access the data.
[0050] The metadata loaded into the memory needs to be processed to ensure its accuracy and consistency. Metadata processing includes operations such as data cleaning, data conversion, and data verification. Data cleaning refers to removing noise and error information in the metadata, such as deleting duplicate records, correcting incorrect field values, etc. Data conversion refers to converting the metadata from one format to another to meet the processing requirements of the system. Data verification refers to checking the integrity and validity of the metadata to ensure that the metadata conforms to the specifications and requirements of the system.
[0051] The processed metadata will be cached in the memory to improve the access speed of the metadata. The metadata cache adopts an efficient caching strategy, such as the least recently used (LRU) strategy. According to the LRU strategy, when the cache space is insufficient, the system will preferentially eliminate the least recently used metadata. This can ensure that the metadata in the cache is the most frequently accessed, improving the hit rate of the metadata. Through the metadata cache, the system can reduce the number of disk accesses and improve the overall performance of data processing.
[0052] The above has described in detail an embodiment of the present invention, but the above content is only a preferred embodiment of the present invention and cannot be considered as defining the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.
Claims
1. An AI data loading method using a dynamic metadata strategy, characterized in that: The following steps are involved: The data sets in different formats are uniformly converted into Lafa format through the FileWriter interface, and the data sets in Lafa format are divided into several shard files for parallel storage and retrieval, each of which includes metadata and data blocks; The decision module is used to evaluate the system's memory resources and the characteristics of the data set in Lafa format in real time, dynamically obtain the adapted data loading mode, and calculate the mode switching threshold. The data loading mode includes fast loading, lazy loading, and slow loading. According to the selected data loading mode, the corresponding metadata loading logic is adopted to process and cache the metadata.
2. The AI data loading method using a dynamic metadata strategy according to claim 1, characterized in that: The process of converting data sets of different formats into Lafa format and performing parallel storage and retrieval is as follows: According to the format of the input data set, the corresponding data reading method is used to uniformly convert data sets in different formats into Lafa format, and preprocess them to build a data structure that meets the requirements of the Lafa format. The constructed Lafa format data is converted into a string form and written into a file. According to the data volume, storage and retrieval requirements of the generated Lafa format file, the Lafa format file is divided into several shard files, and a corresponding index file is generated for each shard, and these shard files are stored in different locations.
3. The AI data loading method using a dynamic metadata strategy according to claim 2, characterized in that: The process of preprocessing the data is as follows: The OpenRefine tool is used to clean the data, fill in missing values, filter outliers, and transform the data. The data transformation includes standardization, normalization, discretization, and feature encoding. The data is integrated based on a distributed data warehouse, and feature engineering is performed to generate new features and reduce dimensionality, and the data is reduced.
4. The AI data loading method using a dynamic metadata strategy according to claim 1, characterized in that: The process of dynamically acquiring the adapted data loading mode is as follows: The decision module obtains the current hardware memory usage in real time through the monitoring interface. The memory usage includes used memory, available memory and memory occupancy trend. It analyzes the metadata of the Lafa file to obtain the data set information. The data set information includes the total sample size, number of shards and sample distribution. By loading several Lafa files of the same scale, the memory growth rate under fast loading, lazy loading and slow loading is tested respectively. The memory growth rate under the three loading modes is compared with the data set information to obtain the loading mode that best suits the current hardware conditions and data set size. The decision module continuously monitors the memory usage and the effect of the current loading mode, and dynamically adjusts the loading mode according to the real-time monitoring situation.
5. The AI data loading method using a dynamic metadata strategy according to claim 4, characterized in that: The process of testing the memory growth rate under fast loading, lazy loading and slow loading is as follows: Step 1: Select several Lafa files of exactly the same size, use the memory analysis tool to monitor the memory usage in real time, and record the initial memory value; Step 2: Use the fast loading mode to preload the metadata of all Lafa files into the memory, traverse all Lafa file paths and store them in the memory, and obtain the memory changes after the loading is completed; Step 3: Use lazy loading mode to preload the key metadata of Lafa files, load detailed metadata on demand, randomly select several file indexes, read the file content and cache the key metadata to memory, traverse all Lafa files, and record the memory changes after loading is completed; Step 4: Use slow loading mode to cache high-level sharding information, dynamically generate detailed metadata, select several high-level sharding files, use the sample ID generator to dynamically generate the required metadata, cache the metadata to memory, and record the memory value after loading is completed; Step 5: Remove abnormal values of memory usage and calculate the memory growth rates under the three loading modes according to the formula: memory growth rate = (final memory - initial memory) / total time consumption.
6. The AI data loading method using a dynamic metadata strategy according to claim 1, characterized in that: The process of generating a local ID by the sample ID generator is as follows: Traverse the sample range in each shard file and define it by start and end, where start indicates the starting point of the sample and end indicates the ending point of the sample. Dynamically preset shufflesize according to the memory size and the number of samples, where shufflesize indicates the shuffle size. Based on the selected start and end, the sample ID generator generates local ID samples on demand, marks the interval range of the generated local ID samples as [start, start+shufflesize), and temporarily caches them in the memory. When the number of local ID samples in the memory reaches the upper limit, the sample ID generator automatically updates start to the starting point of the next interval, and continues to generate new local ID intervals until all sample IDs in the shard are generated and cached.
7. The AI data loading method using a dynamic metadata strategy according to claim 1, characterized in that: The process of dynamically generating metadata in the slow loading mode is as follows: The global index range of the sample in the shard file is defined by start and end, and the global ID of the sample is generated by the sample ID generator. The global ID of the sample is compared with the [start, end) range of all shards to obtain the partition to which the target sample ID belongs. Based on the partition to which the sample ID belongs, a data map is generated, and according to the type, purpose and system configuration of the sample, the task information that needs to be performed by the sample is obtained. Based on the shard information, data mapping information and task information, the metadata of the sample is obtained.
8. The AI data loading method using a dynamic metadata strategy according to claim 1, characterized in that: The shuffling process of the slow loading mode is: Using the selected shuffling algorithm, each local ID temporarily cached in the memory is shuffled. Based on the shuffled sample ID sequence, according to the characteristics of the slow loading mode, the sample data is read from the storage medium in batches and in sequence. After processing all samples in the current partition, the sample ID generator will automatically update the next local ID and temporarily store it in the memory. The shuffling algorithm continues to shuffle the local ID of the next partition.
Citation Information
Cited By
Streaming data loading method and device, electronic equipment and storage medium
CN121009033A