Streaming data loading method and device, electronic equipment and storage medium

By using dynamic segmentation and sliding window buffer design in the streaming data loading method, the problem of low efficiency in memory and computing resource coordination is solved, and the optimization of controllable memory usage and continuous saturation of computing resources is achieved, thereby improving the efficiency of ultra-large-scale data processing.

CN121009033AActive Publication Date: 2025-11-25BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511107929.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-25
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

In existing technologies, the low efficiency of memory and computing resource coordination limits the efficiency of ultra-large-scale data processing.

Method used

A streaming data loading method is adopted, which adjusts the size of data blocks according to real-time memory resources and computing power through a dynamic block partitioning strategy. Combined with a sliding window buffer and an asynchronous preloading mechanism, the pipeline of data processing and loading is made parallel.

Benefits of technology

It achieves coordinated optimization of controllable memory usage and continuous saturation of computing resources, ensuring that the system can control memory usage while maintaining continuous saturation of computing devices when processing ultra-large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009033A_ABST
    Figure CN121009033A_ABST
Patent Text Reader

Abstract

The invention provides a streaming data loading method and device, electronic equipment and a storage medium. The method comprises the steps that a to-be-processed data file is acquired from a data lake; dynamically determining the dynamic block size of the data blocks according to the available memory resources of the current node and the scale of the computing equipment, and dividing the data file into a plurality of continuous data blocks according to the dynamic block size; a sliding window type buffer area is established, and the buffer area is divided into a front-end area, a rear-end area and a front-end area, wherein the front-end area is used for storing processed data blocks; the middle area is used for storing data blocks which are currently processed; the tail end area is used for storing data blocks to be preloaded; and when the computing equipment starts to process the data blocks of the middle area, asynchronously loading the next data block of the data blocks to the tail end area, and synchronously releasing the processed data blocks in the front end area when a new data block is stored in the tail end area.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a streaming data loading method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, data lakes have become the core infrastructure for processing multi-modal, super-large-scale data sets. When training deep learning models, how to efficiently load and manage massive amounts of data becomes a key challenge, especially the need to balance memory resource occupation, computing device utilization, and data supply requirements of different application scenarios (such as training and inference).

[0003] Currently, the mainstream technical solutions mainly include full preloading mode and batch loading strategy. Data lake components represented by Apache Iceberg use full loading mode to load all data files into memory to form a unified columnar storage table in the initialization phase, and the improved solution uses fixed-size data blocks to load in batches to relieve memory pressure.

[0004] However, these existing technologies have the problem of low coordination efficiency of memory and computing resources, which seriously restricts the efficiency of super-large-scale data processing. SUMMARY

[0005] The present application provides a streaming data loading method, device, electronic equipment and storage medium to solve the problem of low coordination efficiency of memory and computing resources in the prior art.

[0006] In a first aspect, the present application provides a streaming data loading method, comprising:

[0007] obtaining a data file to be processed from a data lake;

[0008] dynamically determining a dynamic block size of data blocks according to available memory resources and computing device size of a current node, and dividing the data file into a plurality of continuous data blocks according to the dynamic block size of data blocks;

[0009] establishing a sliding window buffer, the buffer being divided into: a front-end area for storing processed data blocks; a middle area for storing data blocks being processed; and an end area for storing data blocks to be preloaded;

[0010] when the computing device starts processing the data blocks in the middle area, asynchronously loading a next data block of the data block to the end area, and synchronously releasing the processed data blocks in the front-end area when a new data block is stored in the end area.

[0011] In a possible implementation, the dynamically determining the dynamic block size of the data block according to the available memory resource of the current node and the scale of the computing device comprises:

[0012] monitoring the remaining memory space in the available memory resource that is actually available for data loading;

[0013] obtaining a quantitative parameter of the scale of the computing device;

[0014] calculating an initial block value based on a proportional relationship between the remaining memory space and the quantitative parameter;

[0015] applying a size constraint condition to the initial block value to obtain the dynamic block size.

[0016] In a possible implementation, the establishing the sliding window buffer comprises:

[0017] obtaining an average processing speed of the computing device in processing a single data block, and monitoring an average loading speed of the data block from the storage device to the memory;

[0018] calculating an initial window capacity according to the average processing speed and the average loading speed;

[0019] increasing a preset window capacity adjustment base to the initial window capacity to determine a target window capacity;

[0020] initializing the sliding window buffer according to the target window capacity, so that the number of data blocks simultaneously saved in the sliding window buffer is consistent with the target window capacity.

[0021] In a possible implementation, the method further comprises:

[0022] monitoring the memory usage of the current node in real time;

[0023] in a case where the memory usage exceeds a preset high water level threshold, reducing the target window capacity to a first preset value;

[0024] in a case where the memory usage is lower than a preset low water level threshold, increasing the target window capacity to a second preset value.

[0025] In a possible implementation, the method further comprises:

[0026] monitoring a processing characteristic parameter of the data block in real time, the processing characteristic parameter being used to represent a data access mode characteristic;

[0027] determining a data supply mode according to the processing characteristic parameter, the data supply mode comprising: a training mode of storing the data chunks into a multi-level cache system, and / or an inference mode of establishing a direct transmission channel of the data chunks to a computing device.

[0028] In one possible implementation, the method further comprises:

[0029] In the training mode, two cache copies are generated and maintained for the data chunks processed by the sliding window buffer: a memory cache copy that stores part of data in the data chunk with a higher access frequency than a preset threshold, and a persistent cache copy that stores a complete data chunk in a distributed storage system.

[0030] In response to a data request of a training iterator, the part of data with a high access frequency in the memory cache is queried; in the case of a memory cache miss, the complete data chunk in the persistent cache is queried; in the case of a persistent cache miss, the data chunk is obtained from a remote storage system; the complete data chunk is written into the persistent cache for the data chunk obtained from the remote storage system; and the part of data with a higher access frequency than the preset threshold in the data chunk is extracted and written into the memory cache.

[0031] For each data chunk, access statistical information of the data chunk is periodically collected; a hotness value is calculated according to the access statistical information; and a migration strategy of the data chunk between the memory cache and the persistent cache is determined according to the hotness value corresponding to the data chunk.

[0032] In one possible implementation, the method further comprises:

[0033] A columnar storage memory mapping technology is used for the divided data chunk to establish a direct mapping channel of a disk file to a memory space.

[0034] Through the direct mapping channel, the computing device directly accesses the original binary content of the data chunk.

[0035] A shared memory pool mechanism is configured to make multiple data chunks reuse a same memory mapping area, and to dynamically manage a life cycle of the memory mapping area based on a reference counting manner, wherein: when a new data chunk is loaded, a reference count of a corresponding memory area is increased; when a data chunk processing is completed, the reference count of the corresponding memory area is decreased; and when the reference count is detected to be zero, the memory mapping is automatically released and the resource is released.

[0036] In a second aspect, the present application provides a streaming data loading device, comprising:

[0037] An acquisition module is configured to acquire a data file to be processed from a data lake.

[0038] determining a dynamic block size of the data blocks according to available memory resources of the current node and a scale of the computing device, and dividing the data file into a plurality of continuous data blocks according to the dynamic block size;

[0039] establishing a sliding window buffer, the buffer being divided into a front-end area for storing data blocks that have been processed, a middle area for storing data blocks that are being processed, and an end area for storing data blocks to be preloaded;

[0040] processing the data blocks in the middle area, asynchronously loading a next data block of the data blocks into the end area when the computing device starts processing the data blocks in the middle area, and synchronously releasing the data blocks that have been processed in the front-end area when a new data block is stored in the end area.

[0041] In one possible implementation, the determining module is specifically configured to:

[0042] monitoring a remaining memory space in the available memory resources that is actually available for data loading;

[0043] obtaining a quantitative parameter of the scale of the computing device;

[0044] calculating an initial block value based on a proportional relationship between the remaining memory space and the quantitative parameter;

[0045] applying a size constraint condition to the initial block value to obtain the dynamic block size.

[0046] In one possible implementation, the establishing module is specifically configured to:

[0047] obtaining an average processing speed of the computing device in processing a single data block, and monitoring an average loading speed of data blocks from a storage device to the memory;

[0048] calculating an initial window capacity according to the average processing speed and the average loading speed;

[0049] increasing the initial window capacity by a preset window capacity adjustment base to determine a target window capacity;

[0050] initializing the sliding window buffer according to the target window capacity, so that the number of data blocks saved in the sliding window buffer is consistent with the target window capacity.

[0051] In one possible implementation, the establishing module is further configured to:

[0052] monitoring a memory usage rate of the current node in real time;

[0053] decrease the target window size to a first preset value when the memory usage exceeds a preset high water level threshold;

[0054] increase the target window size to a second preset value when the memory usage is below a preset low water level threshold.

[0055] In one possible implementation, the apparatus further comprises a monitoring module configured to:

[0056] monitor, in real time, a processing characteristic parameter of the data chunk, the processing characteristic parameter being indicative of a data access pattern characteristic;

[0057] determine, according to the processing characteristic parameter, a data supply mode, the data supply mode comprising a training mode of storing the data chunk into a multi-level cache system and / or an inference mode of establishing a direct transmission channel of the data chunk to a computing device.

[0058] In one possible implementation, the apparatus further comprises a maintaining module configured to:

[0059] generate and maintain, in the training mode, two cache copies of the data chunk processed by the sliding window buffer: a memory cache copy storing a part of data in the data chunk with a higher access frequency than a preset threshold, and a persistent cache copy storing a complete data chunk in a distributed storage system;

[0060] in response to a data request of a training iterator, query the high-frequency access data part in the memory cache; in case of a cache miss in the memory cache, query the complete data chunk in the persistent cache; in case of a cache miss in the persistent cache, obtain the data chunk from a remote storage system; write the complete data chunk into the persistent cache for the data chunk obtained from the remote storage system; and extract the part of data in the data chunk with the higher access frequency than the preset threshold to write into the memory cache;

[0061] periodically collect access statistical information of the data chunk for each data chunk; calculate a hotness value according to the access statistical information; and determine a migration strategy of the data chunk between the memory cache and the persistent cache according to the hotness value corresponding to the data chunk.

[0062] In one possible implementation, the apparatus further comprises an accessing module configured to:

[0063] establish a direct mapping channel of a disk file to a memory space by using a columnar storage memory mapping technology for the divided data chunk;

[0064] enable the computing device to directly access raw binary content of the data chunk through the direct mapping channel;

[0065] The shared memory pool mechanism is configured to enable multiple data blocks to share the same memory mapping area, and to dynamically manage the life cycle of the memory mapping area based on a reference counting manner, wherein: the reference count of the corresponding memory area is increased when a new data block is loaded; the reference count of the corresponding memory area is decreased when the data block processing is completed; and the memory mapping is automatically released and the resource is released when the reference count is detected to be zero.

[0066] In a third aspect, the present application provides a device, comprising: a processor and a memory, the processor being configured to execute a streaming data loading program stored in the memory to implement the streaming data loading method of any one of the first aspect.

[0067] In a fourth aspect, the present application provides a storage medium, the storage medium storing one or more programs, the one or more programs being executable by one or more processors to implement the streaming data loading method of any one of the first aspect.

[0068] The above technical solution provided by the embodiments of the present application has the following advantages compared with the prior art: the method provided by the embodiments of the present application adjusts the data block size according to real-time memory resources and computing capacity through a dynamic block strategy, which not only avoids the memory pressure of full loading, but also ensures the continuous data supply of the computing device through size adaptation; the three-region (front / middle / terminal) design of the sliding window buffer combined with the asynchronous preloading mechanism enables the data processing and data loading to form a pipeline in parallel, completely eliminating the idle waiting period of the computing device; thus, the collaborative optimization of controllable memory occupation and continuous saturation work of computing resources is realized, so that the system can control memory occupation and maintain the continuous saturation working state of the computing device when processing large-scale data. BRIEF DESCRIPTION OF DRAWINGS

[0069] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate one embodiment consistent with the present application and, together with the description, serve to explain the principles of the application.

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, brief descriptions will be given below of the drawings needed to be used in the embodiments or prior art descriptions. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without any creative effort.

[0071] One or more embodiments are exemplarily illustrated by the pictures in the drawings corresponding thereto, and these exemplary illustrations do not constitute a limitation on the embodiments, and elements with the same reference numerals in the drawings represent similar elements, unless otherwise specified, and the drawings in the drawings do not constitute a proportional limitation.

[0072] Figure 1An embodiment flowchart of a stream data loading method provided by the embodiment of the present application is shown in FIG. 1.

[0073] Figure 2 An embodiment flowchart of another stream data loading method provided by the embodiment of the present application is shown in FIG. 2.

[0074] Figure 3 A flowchart of data query in a training mode provided by the embodiment of the present application is shown in FIG. 3.

[0075] Figure 4 A full flowchart of stream data loading provided by the embodiment of the present application is shown in FIG. 4.

[0076] Figure 5 An embodiment block diagram of a stream data loading device provided by the embodiment of the present application is shown in FIG. 5.

[0077] Figure 6 A structural diagram of an electronic device provided by the embodiment of the present application is shown in FIG. 6. DETAILED DESCRIPTION

[0078] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0079] The following disclosure provides many different embodiments, or examples, for implementing different structures of the present application. For the purpose of simplicity and clarity, the description of the specific examples in the following text will be described. Of course, they are only examples, and the purpose is not to limit the present application. In addition, the present application can repeatedly refer to the numbers and / or letters in different examples. Such repetition is for the purpose of simplification and clarity, and it does not indicate the relationship between the various embodiments and / or settings discussed.

[0080] To solve the technical problem of low coordination efficiency of memory and computing resources in the prior art, the present application provides a stream data loading method, which can realize the coordination optimization of controllable memory occupation and continuous saturated work of computing resources, so that the system can control the memory occupation and maintain the continuous saturated working state of the computing device when processing large-scale data.

[0081] Figure 1 An embodiment flowchart of a stream data loading method provided by the embodiment of the present application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the method comprises the following steps:

[0082] Step 101, obtaining a data file to be processed from a data lake.

[0083] Data lake refers to a large-scale distributed storage system that supports multi-modal data storage, including structured data, semi-structured data and unstructured data. Its core feature is to manage various data files stored in the underlying file system (such as HDFS (Hadoop Distributed File System), Amazon Simple Storage Service (S3)) through a metadata layer.

[0084] Data file to be processed refers to a data set that needs to be loaded into a computing system for machine learning training or inference tasks. These files usually use columnar storage formats (such as Apache Parquet, ORC (Optimized Row Columnar)) to optimize read performance.

[0085] In this embodiment, the system accesses the target data file through a standardized data lake interface (such as the API provided by Apache Iceberg and Delta Lake). During the access process, the metadata information of the data file (including partition structure, statistical information, etc.) is loaded first. Based on this metadata, the division strategy of subsequent data blocks is determined. This step establishes the entry of the data loading process and provides the original data input for subsequent dynamic block processing.

[0086] Step 102, dynamically determining the dynamic block size of the data block according to the available memory resources and computing device scale of the current node, and dividing the data file into several continuous data blocks according to the dynamic block size.

[0087] Current node refers to a physical or virtual computing unit that performs data processing tasks. Its available memory resources include physical memory and off-heap memory that are not occupied by system processes. The computing device scale represents the parallel processing capability of the node, such as the number of GPUs (Graphics Processing Unit) or the number of CPU (Central Processing Unit) cores. Dynamic block size is the size of the data block determined according to real-time resource conditions.

[0088] In an embodiment, the dynamic block size of the data block is dynamically determined according to the available memory resource of the current node and the scale of the computing device, which can include the following steps: monitoring the remaining memory space actually available for data loading in the available memory resource; obtaining a quantitative parameter of the scale of the computing device; calculating an initial block value based on the proportional relationship between the remaining memory space and the quantitative parameter; and applying a size constraint condition to the initial block value to obtain the dynamic block size.

[0089] The remaining memory space refers to the available memory capacity actually available for data loading in the current computing node, excluding the system reserved memory and the memory occupied by other running processes.

[0090] In this embodiment, the system first monitors the remaining available memory space in real time and obtains the quantitative parameter of the computing device. Then, the initial value is obtained according to the calculation formula "initial block value = remaining memory space / (quantitative parameter x adjustment coefficient)", wherein the adjustment coefficient is an empirical value (for example, the value is 4). Finally, the final dynamic block size is obtained by adjusting the size constraint condition, for example, if it is less than a preset lower limit value (such as 256 megabytes), it is set to the lower limit value to avoid memory fragmentation; if it is greater than a preset upper limit value (such as 1024 megabytes), it is limited to the upper limit value to ensure the parallel processing efficiency. Through this scheme, the system realizes the automatic adaptation of the data block size to the real-time resource status of the node, optimizes the computing efficiency while ensuring the memory safety.

[0091] In another embodiment, the method can further include the following steps: using columnar storage memory mapping technology for the divided data blocks to establish a direct mapping channel of disk files to memory space; enabling the computing device to directly access the original binary content of the data blocks through the direct mapping channel; configuring a shared memory pool mechanism to multiplex the same memory mapping area for multiple data blocks, and dynamically managing the life cycle of the memory mapping area based on a reference counting method, wherein: the reference count of the corresponding memory area is increased when a new data block is loaded; the reference count of the corresponding memory area is decreased when the data block processing is completed; and the memory mapping is automatically released and the resource is released when the reference count is detected to be zero.

[0092] The embodiment realizes efficient data loading through the inter-process communication protocol of Apache Arrow (Arrow IPC protocol), specifically including: memory mapping establishment: mapping columnar storage files (such as Apache Parquet / optimized row-column storage) on a disk to user space memory through the Arrow IPC protocol; the starting position of each data block is strictly aligned to the metadata boundary of an Arrow record batch (Record Batch); the integrity of the original columnar storage format is maintained, and structural damage caused by data cutting is avoided. Zero-copy mechanism: the computing device directly accesses the original binary data through virtual address mapping; the traditional data deserialization and memory copy operations are completely avoided; shared memory access across processes / devices is supported. Resource management optimization: multiple data blocks share the same physical memory mapping area; based on reference counting, the life cycle is automatically managed: when a new block is loaded, the reference count of the corresponding memory area is increased; when the block processing is completed, the reference count is reduced; when the reference count is detected to be zero, the mapping is automatically released and the resource is released. The technical solution significantly improves the data transmission efficiency (eliminates the copy overhead), optimizes the memory utilization (shared mapping area), enhances the system reliability (automatic resource recycling), and provides a high-performance, low-overhead memory access solution for super large-scale data processing.

[0093] In addition, in yet another embodiment, the method can further include the following steps: establishing a metadata index for each data block, the metadata index including: the starting offset of the data block in the source file; the byte length of the data block; the corresponding columnar storage format metadata boundary information; wherein the metadata index is stored and managed in a lightweight structure independent of the data block body.

[0094] The metadata index is structured information for describing the physical storage characteristics of the data block, including three key fields: the starting offset represents the starting byte position of the block in the source file, the byte length records the storage space size occupied by the block, and the columnar storage format metadata boundary information ensures that the block cutting is aligned with the inherent structure of the original data format (such as Parquet's Row Group or ORC's Stripe). The lightweight structure refers to an index storage scheme using efficient encoding formats such as JSON, binary, etc.

[0095] In this embodiment, the system independently constructs a metadata index for each dynamically generated data chunk, which is managed in an independent file or a dedicated storage area and physically separated from the data chunk body. When accessing data, the system preferentially loads the lightweight metadata index, quickly locates the physical location and structural characteristics of the target chunk by analyzing the index information, and then loads the actual data content as needed. The construction process of the index strictly follows the columnar storage format specification to ensure that the chunk boundary is always aligned with the metadata boundary of the original data. In this way, the chunk retrieval efficiency can be significantly improved, and quick positioning can be achieved through indexing to avoid full file scanning.

[0096] Step 103, a sliding window buffer is established, which is divided into: a front-end area for storing data chunks that have been processed and completed; a middle area for storing data chunks that are currently being processed; and an end area for storing data chunks to be preloaded.

[0097] The sliding window buffer is a ring queue structure with a fixed capacity, used to implement pipeline processing of data chunks. Its core feature is to divide the storage space into three functional areas: the front-end area is used to cache data chunks that have been processed and are ready to be released, the middle area stores data chunks that are currently being processed by the computing device, and the end area stores data chunks that have been preloaded and are ready to be processed. Each area is identified by a pointer, where the read pointer identifies the position of the data chunk currently being processed, and the write pointer identifies the position of the new data chunk to be written. The automatic connection of the front-end area release and the end area loading is achieved by moving the pointers, forming a first-in-first-out data flow pipeline.

[0098] In an embodiment, establishing a sliding window buffer can include the following steps: obtaining the average processing speed of a computing device processing a single data chunk, and monitoring the average loading speed of data chunks from a storage device to the memory; calculating an initial window capacity based on the average processing speed and the average loading speed; increasing the initial window capacity by a preset window capacity adjustment base to determine a target window capacity; initializing the sliding window buffer according to the target window capacity, so that the number of data chunks saved in the sliding window buffer is consistent with the target window capacity.

[0099] The average processing speed refers to the throughput (unit: MB / s) of a computing device (such as GPU / CPU) processing a single data chunk. The average loading speed reflects the transmission rate of a storage device (such as SSD (Solid State Drive) / HDD (Hard Disk Drive)) loading data chunks to the memory. The window capacity adjustment base is a buffer margin (empirical value) set to cope with speed fluctuations.

[0100] In this embodiment, the system dynamically monitors the real-time processing capacity and data loading performance of the computing device, first calculates the initial window capacity (initial window capacity = average processing speed / average loading speed), then increases the window capacity adjustment base (such as +1) to obtain the target window capacity, and finally initializes the sliding window buffer according to the capacity. The process ensures that the buffer capacity always matches the current processing capacity of the system: when the calculation speed is fast, the window is increased to maintain data supply, and when the loading speed is slow, the window is reduced to prevent memory accumulation.

[0101] In addition, in another embodiment, the method can further include the steps of: monitoring the memory usage of the current node in real time; if the memory usage exceeds a preset high water level threshold, reducing the target window capacity to a first preset value; if the memory usage is lower than a preset low water level threshold, increasing the target window capacity to a second preset value.

[0102] The memory usage refers to the percentage of the used memory in the total memory capacity in the current node; the preset high water level threshold and the preset low water level threshold are preset memory warning lines (for example, 80% and 50%) for triggering window capacity adjustment; the first preset value and the second preset value represent the minimum and maximum safe values of the window capacity (for example, 3 and 5 data blocks) respectively.

[0103] This embodiment adds a real-time feedback mechanism based on memory pressure on the basis of dynamic window adjustment: the system continuously monitors the memory usage of the node, and immediately reduces the window capacity to a first preset value (such as 3 blocks) when the high water level threshold is exceeded to prevent memory overflow; when the low water level threshold is exceeded, the capacity is expanded to a second preset value (such as 5 blocks) to fully improve the data supply capacity. This dual threshold control ensures that the system always operates within a safe memory range.

[0104] Step 104, when the computing device starts processing the data blocks of the middle area, the next data block of the data block is loaded asynchronously to the end area, and when the new data block is stored in the end area, the processed data block in the front area is released synchronously.

[0105] Asynchronous loading refers to the data prefetching operation performed by the background thread independently of the main computing process; synchronous release means that the memory resources of the data block in the front area are immediately recycled after the processing is completed. The two operations are completed through the three functional areas of the sliding window buffer: the middle area (in calculation), the end area (preloading), and the front area (to be released).

[0106] The embodiment realizes the automatic operation of the data processing pipeline: when the computing device starts processing the current data block in the middle area, the system immediately triggers the asynchronous loading of the next data block to the end area; when the new block is loaded, the old block that has been processed in the front-end area is released synchronously. This mechanism maintains the continuous movement of the sliding window through the double buffering strategy (processing and loading in parallel), ensuring that the computing device always has available data while keeping the memory usage constant.

[0107] The technical scheme provided by the embodiment of the application adjusts the data block size according to real-time memory resources and computing capacity through a dynamic block strategy, which avoids the memory pressure of full loading and ensures the continuous data supply of the computing device through size adaptation; the three-area (front-end / middle / end) design of the sliding window buffer combined with the asynchronous preloading mechanism makes the data processing and data loading form a pipeline in parallel, completely eliminating the idle waiting period of the computing device; thereby realizing the collaborative optimization of controllable memory usage and continuous saturation of computing resources, so that the system can control memory usage and maintain the continuous saturation of the computing device when processing large-scale data.

[0108] Figure 2 The embodiment flowchart of another stream data loading method provided by the embodiment of the application. Figure 2 The flowchart shown in Figure 1 Based on the flowchart shown, the following steps are included:

[0109] Step 201, real-time monitoring of the processing characteristic parameters of the data block, the processing characteristic parameters being used to represent the data access mode characteristics.

[0110] The processing characteristic parameters are dynamic indicators for quantifying the data block access mode, such as data reuse frequency (the ratio of repeated access to the same block), number of computing iterations (the number of rounds of block participation in training iterations), etc. These parameters are collected through real-time monitoring of the data access pipeline.

[0111] In this embodiment, the system continuously tracks the processing of each data block, records key characteristic parameters through a lightweight statistical module, and analyzes the parameters based on a time window (such as every 5 minutes). These parameters dynamically reflect the characteristics of the workload, for example: high-frequency reused blocks may belong to hot features, and blocks with multiple iterations are usually located in the key training sample set.

[0112] This step can provide quantitative basis for data supply mode selection (training / inference).

[0113] Step 202, determining a data supply mode according to the processing characteristic parameters, the data supply mode including: a training mode of storing the data blocks into a multi-level cache system, and / or an inference mode of establishing a direct transmission channel of the data blocks to a computing device.

[0114] The data supply mode refers to a differentiated data distribution strategy selected by the system according to data processing requirements, including two types of training mode and inference mode: the training mode optimizes data reuse through a multi-level cache system (including memory cache and persistent storage levels), and the inference mode uses a direct transmission channel (bypassing the cache and directly transmitting to the computing device) to ensure real-time performance.

[0115] In this embodiment, the system dynamically selects the optimal supply mode based on real-time monitoring of processing characteristic parameters: when high-frequency reuse or multi-round iteration features are detected (typical training scenarios), the training mode is enabled and the data blocks are cached to the multi-level storage system; when the data access presents one-time or low-latency requirements (typical inference scenarios), the direct transmission mode is switched to. Mode switching is achieved through a lightweight routing module.

[0116] Figure 2 According to the flow shown, first, based on real-time collected data reuse frequency and computing iteration times and other characteristic parameters, the system can accurately identify the workload type (training / inference) to provide quantitative basis for subsequent optimization decisions; second, through the multi-level cache system (memory + persistent storage) under the training mode, efficient reuse of hot data is achieved, improving the data supply efficiency of training tasks; finally, the direct transmission channel established under the inference mode completely avoids cache access overhead, compressing end-to-end delay to milliseconds. In this way, a single system architecture can intelligently adapt to the differentiated needs of training and inference, improving resource utilization while reducing deployment complexity, providing a unified data supply solution with high throughput and low latency for heterogeneous computing scenarios.

[0117] In another embodiment of the present application, the method can further comprise the following steps: in the training mode, generating and maintaining two cache copies for the processed data blocks in the sliding window buffer: a memory cache copy, which stores the data segments with access frequency higher than a preset threshold; and a persistent cache copy, which stores the complete data blocks in the distributed storage system; in response to a data request of a training iterator, querying the high-frequency access data segments in the memory cache; in the case of a memory cache miss, querying the complete data blocks in the persistent cache; in the case of a persistent cache miss, obtaining the data blocks from a remote storage system; writing the complete data blocks obtained from the remote storage system into the persistent cache; extracting the data segments with access frequency higher than the preset threshold from the complete data blocks and writing them into the memory cache; periodically collecting access statistical information of each data block; calculating a hotness value according to the access statistical information; and determining a migration strategy of the data block between the memory cache and the persistent cache according to the hotness value corresponding to the data block.

[0118] The memory cache copy refers to the high-frequency access data segments residing in the volatile memory (such as DRAM (Dynamic Random Access Memory)), and the persistent cache copy refers to the complete data blocks stored in the distributed storage system (such as an SSD cluster). The hotness value is a data priority score obtained by quantitative analysis of access statistical information (such as access frequency and timeliness).

[0119] In this scheme, first, two optimized copies are generated for the processed data blocks: the memory cache copy stores the high-frequency access data segments to achieve fast response, and the persistent cache copy stores the complete data blocks to ensure data integrity. Second, a hierarchical query mechanism is established to check the memory cache, the persistent cache and the remote storage in sequence, and automatically trigger data backfilling and hot spot extraction when a miss occurs. Finally, the dynamic hotness evaluation system continuously analyzes the access characteristics (such as access frequency and access timeliness) of each data block, calculates the comprehensive hotness value, and intelligently adjusts the distribution of data among the cache levels according to the hotness value.

[0120] The calculation of the hotness value includes: collecting the access frequency and the latest access timestamp of each data block; calculating the comprehensive hotness value according to a preset weighting coefficient; and implementing cache migration according to the hotness value sorting result.

[0121] In the scheme, the system periodically collects access records of each cache data block, normalizes the access frequency and the latest access timestamp, linearly weights them according to a preset weight, and generates a comprehensive hotness value in the range of 0-1. For example: hotness value = 0.7 x standardized access frequency + 0.3 x standardized time decay value. Based on the calculation result, all blocks are sorted, the top 20% high hotness blocks are upgraded to the memory cache, the last 10% low hotness blocks are downgraded to the remote storage, and the remaining blocks remain in the current cache level. In this way, based on the comprehensive hotness evaluation model of access frequency and timeliness, the cache decision can be more scientific and reasonable, and the limitations of traditional single index evaluation can be avoided.

[0122] Figure 3 A flowchart of a data query process in a training mode provided by an embodiment of the present application is shown in FIG. 1. In response to a data request of a training iterator, high-frequency data segments in the memory cache are preferentially queried. When the memory is not hit, complete blocks in the persistent cache are queried. When both the two-level cache are not hit, the original data is obtained from the remote storage system. During the query process, a "query-backfill" linkage mechanism is implemented: for the data obtained from the remote, a new cache copy is generated synchronously. In addition, the system continuously monitors and analyzes the access characteristics through the hotness analyzer, calculates the dynamic hotness score based on these characteristics, and performs intelligent migration: high-frequency hot data is preferentially retained in the memory cache; low-frequency cold data is removed from the memory cache; invalid data is eliminated from the persistent cache. Figure 3

[0123] The above mechanism realizes three optimizations: performance optimization: a three-level acceleration system of "memory-persistent-remote" is constructed to realize the gradient descent of data access delay; resource optimization: the memory usage efficiency is improved through hot data screening, and the storage load is reduced through automatic migration; management optimization: the fully automated workflow reduces the need for manual intervention and improves system stability. Figure 2 In another embodiment of the present application, the training mode further includes the following steps: triggering cache evaluation at the beginning and end of the training period of the training framework; calculating a priority score according to the access mode of the data block in the continuous training period; when the priority score meets the following conditions at the same time, the data block is promoted to a high-level cache: the score exceeds a preset promotion threshold; the access frequency ranking in the current training period enters a preset top percentage range; the access frequency in the continuous multiple training periods shows a growth trend.

[0124]

[0125] ​The training period refers to an iteration process of traversing the training dataset once in the machine learning training process; the cache evaluation refers to a process of analyzing the storage state and use efficiency of the data blocks in the cache system; the priority score is a quantitative index calculated by comprehensively analyzing the access characteristics of the data blocks in multiple training periods, and is used to evaluate the importance level of the data; the promotion threshold is the minimum score requirement set for the cache promotion operation; and the front percentage range represents the relative ranking of the access frequency of the data blocks in the current training period.

[0126] In the training mode, the embodiment realizes intelligent cache optimization management: the system continuously tracks the access behavior (including access frequency, period growth trend, etc.) of each data block in the continuous training period through the evaluation mechanism triggered at the beginning and end of the training period, and calculates the priority score based on multi-dimensional features. When a data block meets three promotion conditions at the same time, i.e., the score exceeds the preset threshold, the current period access frequency ranking is high, and a continuous growth trend is shown, the system automatically migrates it to a higher level cache (such as from SSD cache to memory cache). This decision-making mechanism based on multi-period behavior analysis can accurately identify hot data with long-term value.

[0127] Through the technical scheme of the embodiment, first, the evaluation mechanism triggered based on the training period boundary ensures that the cache decision is synchronized with the training phase, avoiding evaluation interference with the normal training process; second, multi-period access feature analysis effectively distinguishes temporary hot spots from persistent hot spots, improving the accuracy of cache decision-making; finally, the joint determination of multi-dimensional conditions prevents misjudgment caused by a single indicator, significantly improves the cache hit rate, and reduces unnecessary cache migration overhead. These effects collectively optimize the data supply efficiency in the training process.

[0128] In another embodiment of the present application, in the training mode, data migration can also be realized through the following steps: when loading data, a metadata tag containing a unique identity, loading time information and initial access state is established for each data block; a period event listener is registered in the training framework to track the training period state in real time; based on the historical period access data recorded in the metadata tag, a dynamic priority score is calculated by comprehensive analysis; according to the dynamic priority score result, combined with the current state characteristics recorded in the metadata tag, intelligent migration of data blocks between different storage levels is performed.

[0129] The metadata tag is a data structure established for each data block containing a unique identifier, a loading timestamp and an initial state; the period event listener is a trigger mechanism realized through the training period (epoch) start / stop hook of the training framework (such as PyTorch); the dynamic priority score is a numerical value calculated by a specific formula, which integrates cross-epoch access features; and the storage level includes storage media with different performance such as memory and SSD.

[0130] In this scheme, first, metadata tags containing unique identification, loading time and initial state are established for each data block at data loading; second, training state is tracked in real time through the periodic event hook of the training framework; then dynamic priority scores are calculated based on historical access data, and the scoring standard strictly follows: the current cycle access times enter the top 20%, the access increases for 3 consecutive cycles or the data is marked as preheating data, the priority is raised, when the access is not met and the score is too low, the memory is over limit and the score is lower than the median value or the survival cycle is over limit, the priority is degraded; finally, intelligent migration of data blocks between storage levels is performed according to the score results. In this way, through the periodic dynamic cache scheduling, the data access characteristics of training are accurately matched, and the optimal allocation of storage resources is realized, which significantly improves the cache hit rate and training efficiency.

[0131] In another embodiment of the present application, the training mode further includes the following steps: globally reordering the data blocks by a distributed hash algorithm; and maintaining a virtual shard mapping table at the iterator interface layer to realize cross-node data rearrangement.

[0132] The distributed hash algorithm, such as Rendezvous Hash, is used to consistently logically shard the data blocks within the cluster, ensuring that the same block is stably allocated to a fixed node in different training cycles; the virtual shard mapping table is a lightweight data structure maintained at the iterator interface layer, recording the dynamic mapping relationship between logical shards and physical nodes, supporting cross-node data rearrangement requirements.

[0133] In the training mode of the present embodiment, first, a globally unique logical shard number is assigned to all data blocks by a distributed hash algorithm, realizing initial random distribution; then a virtual shard mapping table is dynamically maintained at the iterator interface layer, and the actual storage location of the block is adjusted in real time according to the node load or network condition, while the logical shard number remains unchanged. This double mapping mechanism maintains the randomness of data distribution (beneficial to model convergence), and provides flexible rearrangement capability (adapt to cluster changes).

[0134] Figure 4 A full flow diagram of the streaming data loading provided by the embodiment of the present application is shown in FIG. 1, and the complete workflow can be divided into two parallel paths: Figure 4

[0135] Main data flow path:

[0136] Multi-modal data lake: the system reads raw data in Parquet / ORC / JSON format from data lakes such as Iceberg.

[0137] ​Streamed Segment Loading Engine: Zero-copy loading with Arrow memory format; Dynamic chunking strategy (256MB-1GB); Maintaining a sliding window buffer of 3-5 data segments.

[0138] Memory Buffer: Establishing a double buffering mechanism (Buffer A / B) to achieve asynchronous preloading.

[0139] Mode Selector: Routing data streams according to control center policy decisions.

[0140] Control Decision Path:

[0141] Control Center: Real-time monitoring of memory water level (e.g., high water level 80% / low water level 50%); analyzing workload characteristics (e.g., number of iterations / data reuse frequency); outputting cache policy decisions.

[0142] Training Mode Branch: Alluxio Cache Cluster: Implementing epoch-aware cache policy; Establishing a two-level cache of memory→SSD; Dynamically migrating hot data (e.g., Top 20% up to memory, bottom 10% downgraded); Dynamic Shuffle: Realizing global disorder based on Rendezvous hashing algorithm; GPU Computing Unit: Receiving processed data.

[0143] Inference Mode Branch: Straight-through Pipeline: Bypassing the cache system; Sequential Processor: Ensuring data sequentiality; GPU Computing Unit: Receiving processed data.

[0144] This solution realizes zero-copy data transmission through a dynamic segment loading engine and Arrow memory format, significantly improving data loading efficiency; adopts an epoch-aware intelligent cache strategy and Alluxio distributed cache system, significantly reducing epoch switching time in training scenarios through a multi-level cache migration mechanism; the innovative dual-mode data supply interface can intelligently switch between training and inference modes according to different task requirements, realizing global data disorder through a dynamic Shuffle in training mode and maintaining sequential processing in inference mode, ensuring that various deep learning tasks can obtain optimal data supply performance; based on the dynamic adjustment mechanism of memory water level and intelligent hotness analysis algorithm, efficient utilization of storage resources is realized, while the risk of memory overflow is avoided.

[0145] Figure 5 An embodiment block diagram of a streamed data loading device is provided for the embodiments of the present application. As shown in Figure 5 the device comprises:

[0146] The acquisition module 51 is configured to acquire a data file to be processed from a data lake.

[0147] The determining module 52 is configured to dynamically determine a dynamic block size of the data blocks according to available memory resources of the current node and a scale of the computing device, and divide the data file into a plurality of continuous data blocks according to the dynamic block size.

[0148] The establishing module 53 is configured to establish a sliding window buffer, and the buffer is divided into a front-end area for storing data blocks that have been processed, a middle area for storing data blocks that are being processed, and an end area for storing data blocks to be preloaded.

[0149] The processing module 54 is configured to load a next data block of the data block in the middle area to the end area when the computing device starts processing the data block in the middle area, and release the data block that has been processed in the front-end area when a new data block is stored in the end area.

[0150] In a possible implementation, the determining module is specifically configured to:

[0151] monitor a remaining memory space in the available memory resources that is actually available for data loading;

[0152] obtain a quantitative parameter of the scale of the computing device;

[0153] calculate an initial block value based on a proportional relationship between the remaining memory space and the quantitative parameter;

[0154] apply a size constraint condition to the initial block value to obtain the dynamic block size.

[0155] In a possible implementation, the establishing module is specifically configured to:

[0156] obtain an average processing speed of the computing device for processing a single data block, and monitor an average loading speed of the data block from the storage device to the memory;

[0157] calculate an initial window capacity according to the average processing speed and the average loading speed;

[0158] increase a preset window capacity adjustment base to the initial window capacity to determine a target window capacity;

[0159] initialize the sliding window buffer according to the target window capacity, so that the number of data blocks saved in the sliding window buffer is consistent with the target window capacity.

[0160] In a possible implementation, the establishing module is further configured to:

[0161] monitor a memory usage rate of the current node in real time;

[0162] decrease the target window size to a first preset value when the memory usage exceeds a preset high water level threshold;

[0163] increase the target window size to a second preset value when the memory usage is below a preset low water level threshold.

[0164] In one possible implementation, the apparatus further comprises a monitoring module configured to:

[0165] monitor, in real time, a processing characteristic parameter of the data chunk, the processing characteristic parameter being configured to represent a data access pattern characteristic;

[0166] determine, according to the processing characteristic parameter, a data supply mode, the data supply mode comprising: a training mode, in which the data chunk is stored in a multi-level cache system, and / or an inference mode, in which a direct transmission channel of the data chunk to the computing device is established.

[0167] In one possible implementation, the apparatus further comprises a maintaining module configured to:

[0168] generate and maintain, in the training mode, two cache copies of the data chunk processed by the sliding window buffer: a memory cache copy configured to save a part of data in the data chunk with an access frequency higher than a preset threshold, and a persistent cache copy configured to save a complete data chunk in a distributed storage system;

[0169] in response to a data request of a training iterator, query the high-frequency access data part in the memory cache; in the case of a memory cache miss, query the complete data chunk in the persistent cache; in the case of a persistent cache miss, obtain the data chunk from a remote storage system; write the complete data chunk obtained from the remote storage system into the persistent cache; extract the part of data in the data chunk with the access frequency higher than the preset threshold and write into the memory cache;

[0170] periodically collect, for each data chunk, access statistical information of the data chunk; calculate a hotness value according to the access statistical information; and determine a migration strategy of the data chunk between the memory cache and the persistent cache according to the hotness value corresponding to the data chunk.

[0171] In one possible implementation, the apparatus further comprises an accessing module configured to:

[0172] adopt a columnar storage memory mapping technology for the divided data chunk to establish a direct mapping channel of a disk file to a memory space;

[0173] enable the computing device to directly access raw binary content of the data chunk through the direct mapping channel;

[0174] A shared memory pool mechanism is configured to multiplex multiple data blocks in a same memory mapping region, and to dynamically manage a life cycle of the memory mapping region based on a reference counting manner, wherein: a reference count of a corresponding memory region is increased when a new data block is loaded; the reference count of the corresponding memory region is decreased when a data block is processed; and the memory mapping is automatically released and the resource is released when the reference count is detected to be zero.

[0175] As shown in Figure 6 An embodiment of the present application provides a device, comprising a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112 and the memory 113 complete mutual communication through the communication bus 114,

[0176] The memory 113 is used for storing a computer program.

[0177] In an embodiment of the present application, the processor 111 is used for executing the program stored in the memory 113, and a stream data loading method provided by any one of the preceding method embodiments is realized, comprising the following steps.

[0178] Obtaining a data file to be processed from a data lake;

[0179] Dynamically determining a dynamic block size of a data block according to available memory resources of a current node and a scale of a computing device, and dividing the data file into a plurality of continuous data blocks according to the dynamic block size;

[0180] Establishing a sliding window type buffer, and the buffer is divided into: a front-end region used for storing processed data blocks; a middle region used for storing data blocks being processed; and an end region used for storing data blocks to be preloaded;

[0181] When the computing device starts processing the data blocks in the middle region, a next data block of the data blocks is loaded into the end region asynchronously, and when a new data block is stored in the end region, the processed data blocks in the front-end region are released synchronously.

[0182] An embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize steps of a stream data loading method provided by any one of the preceding method embodiments.

[0183] The apparatus embodiments described above are only illustrative, and the units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.

[0184] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software plus a general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0185] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes", "including" and "has" are inclusive and therefore specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order in which they are described, unless specifically indicated as such. It is also to be understood that additional or alternative steps can be employed.

[0186] The above description is merely illustrative of the application and should not be taken as limiting. Numerous modifications and variations underlying the general principles of the applications can be made by those of ordinary skill in the art without departing from the spirit or scope of the application. Therefore, the application is not to be limited to the embodiments described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for loading streaming data, the method comprising: The method comprises: obtaining a data file to be processed from a data lake; dynamically determining a dynamic block size of data blocks according to available memory resources and computing device scale of a current node, and dividing the data file into a plurality of continuous data blocks according to the dynamic block size; establishing a sliding window buffer, which is divided into a front-end area for storing data blocks that have been processed, a middle area for storing data blocks that are currently being processed, and an end area for storing data blocks to be preloaded; when the computing device starts processing data blocks in the middle area, asynchronously loading a next data block of the data blocks to the end area, and synchronously releasing data blocks that have been processed in the front-end area when new data blocks are stored in the end area.

2. The method of claim 1, wherein, The dynamically determining a dynamic block size of data blocks according to available memory resources and computing device scale of a current node comprises: monitoring a remaining memory space actually available for data loading in the available memory resources; obtaining a quantitative parameter of the computing device scale; calculating an initial block value based on a proportional relationship between the remaining memory space and the quantitative parameter; applying a size constraint condition to the initial block value to obtain the dynamic block size.

3. The method of claim 1, wherein, The establishing a sliding window buffer comprises: obtaining an average processing speed of the computing device in processing a single data block, and monitoring an average loading speed of data blocks from a storage device to a memory; calculating an initial window capacity according to the average processing speed and the average loading speed; increasing a preset window capacity adjustment base to the initial window capacity to determine a target window capacity; initializing the sliding window buffer according to the target window capacity, so that the number of data blocks simultaneously saved in the sliding window buffer is consistent with the target window capacity.

4. The method of claim 3, wherein, The method further comprises: real-time monitoring of a memory usage rate of a current node; in a case where the memory usage rate exceeds a preset high water level threshold, reducing the target window capacity to a first preset value; in a case where the memory usage rate is lower than a preset low water level threshold, increasing the target window capacity to a second preset value.

5. The method of claim 1, wherein, The method further comprises: real-time monitoring of a processing characteristic parameter of the data blocks, the processing characteristic parameter being used to represent a data access mode characteristic; determining a data supply mode according to the processing characteristic parameter, the data supply mode comprising a training mode of storing the data blocks into a multi-level cache system, and / or an inference mode of establishing a direct transmission channel of the data blocks to a computing device.

6. The method of claim 5, wherein, The method further comprises: in the training mode, generating and maintaining two cache copies of data blocks processed by the sliding window buffer, an in-memory cache copy for saving part of data whose access frequency is higher than a preset threshold, and a persistent cache copy for saving complete data blocks in a distributed storage system; In response to a data request of the training iterator, query a high-frequency access data part in the memory cache; in the case of a memory cache miss, query a complete data block in the persistent cache; in the case of a persistent cache miss, obtain a data block from a remote storage system; write the complete data block into the persistent cache for the data block obtained from the remote storage system; extract a part of data with a frequency of access higher than a preset threshold in the data block and write into the memory cache; For each data block, periodically collect access statistical information of the data block; calculate a heat value according to the access statistical information; and determine a migration strategy of the data block between the memory cache and the persistent cache according to the heat value corresponding to the data block.

7. The method of claim 1, wherein, The method further comprises: Using a columnar storage memory mapping technology on the divided data blocks to establish a direct mapping channel of a disk file to a memory space; Through the direct mapping channel, the computing device directly accesses the original binary content of the data block; A shared memory pool mechanism is configured to make multiple data blocks reuse the same memory mapping area and dynamically manage the life cycle of the memory mapping area based on a reference counting manner, wherein: when a new data block is loaded, the reference count of the corresponding memory area is increased; when the data block processing is completed, the reference count of the corresponding memory area is reduced; and when the reference count is detected to be zero, the memory mapping is automatically released and the resource is released.

8. A streaming data loading apparatus, characterized by comprising: The apparatus comprises: An obtaining module configured to obtain a data file to be processed from a data lake; A determining module configured to dynamically determine a dynamic block size of a data block according to available memory resources of a current node and a computing device scale, and divide the data file into a plurality of continuous data blocks according to the dynamic block size; An establishing module configured to establish a sliding window type buffer area, the buffer area being divided into: a front-end area for storing data blocks that have been processed; a middle area for storing data blocks that are currently being processed; and an end area for storing data blocks to be preloaded; A processing module configured to, when the computing device starts processing the data blocks in the middle area, asynchronously load a next data block of the data block to the end area, and when a new data block is stored in the end area, synchronously release the data blocks that have been processed in the front-end area.

9. An apparatus, comprising: Comprise: A processor and a memory, the processor being configured to execute a streaming data loading program stored in the memory to implement the streaming data loading method in any one of claims 1-7.

10. A storage medium, characterized by The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the streaming data loading method in any one of claims 1-7.

Citation Information

Patent Citations

  • Storage device, system including the same, and operation method thereof

    CN111143234A

  • Video analysis acceleration method and system based on many-core processor

    CN113012023A

  • Memory allocation for processing sequence data

    CN116755870A

  • Data processing method and equipment

    CN118312766A

  • Target detection network hardware acceleration system based on FPGA

    CN119741589A