Method and system for quickly uploading and downloading massive small files in batches and medium

Through intelligent packaging and multi-threaded concurrency technology, massive small files are batch processed and quickly transferred, solving the problem of inefficiency in traditional methods and achieving rapid uploading and downloading of massive small files.

CN120050279AActive Publication Date: 2025-05-27SICHUAN LEWEI TECH CO LTD

Patent Information

Application Number
CN202510502932.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-27
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

When facing massive small files, traditional file upload and download methods are inefficient and frequently requested networks, resulting in waste of system resources and low transmission rates, which cannot meet the needs of fast upload and download.

Method used

Small files are batch processed through intelligent packaging modules, multiple small files are encapsulated into packaging bodies, and multi-threaded concurrent upload and download technology is used to combine metadata servers and object storage systems to achieve accurate positioning and efficient extraction.

Benefits of technology

It significantly reduces the number of network requests, reduces network overhead and server-side connection processing pressure, improves data transmission efficiency, shortens upload and download time, and meets the rapid processing needs of massive small files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050279A_ABST
    Figure CN120050279A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a system and a medium for quickly uploading and downloading massive small files in batches, and belongs to the technical field of computer storage. The method comprises an initialization stage; an uploading preprocessing stage: preprocessing a to-be-uploaded file, and judging whether the to-be-uploaded file meets a batch uploading condition or not; in the intelligent packaging stage, intelligent packaging is carried out to obtain a packaged body; in the uploading execution stage, the packaged body is divided into a plurality of file blocks, and the file blocks are matched with hash values and then uploaded to an S3 object storage system; in the downloading initialization stage, after a client receives a file downloading request of a user, information of a to-be-downloaded file specified by the user is obtained, and file matching is carried out; in the downloading execution stage, the single file downloading process is executed for the single file downloading request; and executing a batch file downloading process for the batch file downloading request. The overall efficiency of data transmission is improved, the time required for uploading is shortened, and the method is particularly suitable for a large-scale data uploading scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer storage technology, and particularly to a method, system, and medium for batch and rapid uploading and downloading of a large number of small files. Background Art

[0002] With the rapid development of information technology, the amount of data has increased explosively, and the processing of a large number of small files has become a thorny problem faced by many fields. In the big data analysis scenario, a large number of small files such as log files and sensor data files need to be frequently uploaded to the data processing platform for analysis and mining; in the enterprise-level file management system, many small files generated in daily office work, such as email attachments and small project documents, also need to be efficiently transmitted and shared within the enterprise internal network or cloud environment.

[0003] However, the traditional file uploading and downloading methods have exposed many drawbacks when dealing with a large number of small files. During the uploading process, due to the large number of small files, if each file is uploaded one by one, it will cause a large amount of network request overhead, and the establishment and disconnection of network connections frequently consume system resources and time, seriously reducing the uploading efficiency. In terms of backend storage and downloading, the traditional storage architecture is difficult to handle the high-concurrency read and write requests of a large number of small files. For the downloading operation, there is a lack of an efficient extraction and transmission strategy for small files. Especially when quickly locating and downloading a single file from a large collection of small files, it often requires traversing the entire storage structure, consuming a large amount of time and unable to meet the user's demand for quick downloading. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method, system, and medium for batch and rapid uploading and downloading of a large number of small files.

[0005] The purpose of the present invention is achieved through the following technical solutions: In the first aspect of the present invention, there is provided a method for batch and rapid uploading and downloading of a large number of small files, including the following steps: In the initialization stage, the client starts the initialization program to check the system running environment; establishes a secure connection channel with the metadata server and the S3 object storage system, and uses an encrypted communication protocol; loads the client user interface resources and displays the interactive interface for file uploading and downloading operations; In the upload preprocessing stage, when the client receives the small file upload instruction from the user, it preprocesses the file to be uploaded and determines whether the batch upload condition is met; In the intelligent packaging stage, the small files that meet the batch upload condition are intelligently packaged to obtain a package body; In the upload execution stage, the package body is split into multiple file blocks, and after matching the hash values, it is uploaded to the S3 object storage system; In the download initialization stage, when the client receives the user's file download request, it obtains the information of the file to be downloaded specified by the user and performs file matching; In the download execution stage, for a single-file download request, the single-file download process is executed; for a batch-file download request, the batch-file download process is executed.

[0006] Preferably, the system operating environment includes the operating system version, network connection status, available memory, and disk space; the encryption communication protocol is SSL / TLS; the interaction interface includes a file selection area, an upload progress bar, a download prompt box, and an operation log display area.

[0007] Preferably, the upload preprocessing stage further includes the following steps: When the client receives the user's small-file upload instruction, it scans the file path specified by the user or obtains the list of files to be uploaded through a file selection dialog box, and obtains the detailed information of each file, including the file name, file size, file type, creation time, modification time, and the directory path where the file is located; Count the total number of files to be uploaded, calculate the total size of all files, analyze the file distribution, including the number of files and the size ratio at different directory depths, and construct a mapping relationship data structure between the file directory structure and file attributes; Judge whether the number and size of the files to be uploaded meet the batch upload conditions. If the batch upload conditions are not met, skip the file packaging step and directly enter the single-file upload process; if the batch upload conditions are met, enter the file packaging module for intelligent encapsulation processing.

[0008] Preferably, the intelligent packaging stage further includes the following steps: For small files that meet the batch upload conditions, the file packaging module starts an intelligent calculation algorithm, screens files starting from the deepest-level directory according to the directory depth-first principle, and sets the initial target file number threshold and target total size threshold for the packaging body; In the order of file size from small to large, gradually add files to the packaging body. For each added file, immediately calculate its start byte position and end byte position in the packaging body, determine the range value, and associate and store it with the file unique identifier to construct a local temporary packaging body index data structure; During the process of adding files, monitor the current file number and total size of the packaging body in real time. When the preset threshold is reached, pause the file selection at the current directory level and continue to screen the upper-level directory. Repeat this process until the packaging body is constructed or all files to be uploaded have been considered; If the packaging body still does not meet the threshold requirements after traversing all files, complete the construction of the packaging body according to the actual selected files and record the packaging information; After compressing the text file, add it to the package. Merge or preprocess the image sequence files with similar characteristics and then add them to the package, and record the packaging algorithm.

[0009] Preferably, the upload execution stage further includes the following steps: After the file packaging is completed, the file upload module divides the package into multiple file blocks according to a predetermined block size, generates multiple file blocks, and assigns a unique serial number and a corresponding hash value calculation task to each file block; For each file block, create an independent upload task thread, construct an upload task thread pool containing multiple threads, and distribute the file block upload tasks to the thread pool for concurrent execution; Before uploading each file block, calculate its hash value, encapsulate the hash value and the file block data together in the upload request and send it to the S3 object storage system; after receiving the upload request, the S3 object storage system first extracts the hash value and performs a hash calculation verification on the file block data. If the two are consistent, it confirms that the file block is received successfully and stores it in the specified location; if they are inconsistent, it sends an error message to the client, requiring the client to re-upload the file block; As the file blocks are uploaded, the file upload module updates the upload progress information in real time, including the number of uploaded file blocks, the total number of file blocks, the amount of uploaded data, the total amount of data, and synchronously updates the upload progress information to the progress bar of the user interface on the client side; When all file blocks are uploaded and the S3 object storage system returns an upload success message, the file upload module sorts out the metadata information of all files, including file name, original path, file type, range information in the package, compression method; organizes the data according to the predefined data format, and then uploads this metadata information to the metadata server for storage; After receiving the metadata information, the metadata server performs data integrity verification and index construction operations, stores the information in the corresponding database table or data structure, and establishes an index relationship between the file metadata and the range information.

[0010] Preferably, the download initialization stage further includes the following steps: When the client receives the user's file download request, obtain the information of the file to be downloaded specified by the user, including file name and file hash value; if the user only provides the file name, the client first searches in the locally cached metadata information. If not found, it sends a query request to the metadata server to obtain the complete metadata information of the file, including the hash value of the package it is in and the range information in the package.

[0011] Preferably, the single-file download process includes the following steps: For a single-file download request, the client calculates the range of file blocks to be downloaded based on the obtained file range information and package body information, and directly prepares to download the file blocks according to the recorded range; The client creates a download task thread, constructs a single-file download task thread pool, and distributes the file block download tasks to the thread pool for concurrent execution; During the download of each file block, the client sends a download request for the package body hash value, file range position, and signature information to the S3 object storage system. After receiving the request, the S3 object storage system searches for the corresponding file block data based on this information and returns the file block data along with the calculated hash value to the client; After receiving the file block, the client first verifies the correctness of the hash value. If it is correct, the client assembles the file block data according to the range information of the file in the package body and gradually restores the complete file data; if the hash value is incorrect, the client requests the S3 object storage system to re-download the file block; As the file blocks are downloaded and assembled, the client updates the download progress information in real time, including the number of downloaded file blocks, the total number of file blocks, the amount of downloaded data, and the total amount of data, and synchronously updates these download progress information to the progress bar of the client's user interface; When all file blocks are downloaded and assembled into a complete file, the client performs a final integrity check on the file, such as calculating the hash value of the file again and comparing it with the original hash value. If they are the same, it means the file is downloaded successfully; if they are different, error handling is performed according to the situation, including attempting to re-download or prompting the user that the download fails and the reason for the failure.

[0012] Preferably, the batch file download process includes the following steps: For a batch file download request, the client calculates the distribution of all files in the package body based on the obtained file metadata and range information, and determines the ranges of all file blocks to be downloaded; Allocate corresponding thread numbers according to the number of files, construct a batch file download task thread pool, and distribute the file block download tasks to the thread pool for concurrent execution to perform multi-threaded breakpoint download; During the download of each file block, send a request to the S3 object storage system, receive the file block, verify the hash value, and assemble the file to ensure the correct download of each file block and the complete restoration of the file; As the file blocks are downloaded and assembled, the client updates the overall progress information of the batch download in real time, including the number of downloaded files, the total number of files, the amount of downloaded data, and the total amount of data; and displays it on the progress bar of the user interface, and at the same time records the download status and detailed information of each file in the operation log display area; After all files are downloaded, the client re-creates the corresponding directory structure on the local disk according to the original directory information in the file metadata, and stores the downloaded files on the disk according to the original directory path, ensuring that the organizational structure of the files is the same as that at the time of original upload; During the batch download process, if a network interruption or user pause operation occurs, the client records the current download status information, including the downloaded file block information and the download progress of each file, and stores this information in the local cache; when the network resumes or the user continues the operation, the client reads the download status information from the local cache and continues the download task from the breakpoint. If, during the download process, the S3 object storage system or the metadata server fails or has a response delay, the client automatically activates a fault response mechanism, including increasing the number of retries and the waiting time strategy, and attempts to connect to the server to obtain data multiple times within the first preset time; if the fault duration is greater than the second preset time, the download task is paused and a fault message is prompted to the user. After the server returns to normal, the download is continued from the breakpoint or the download task is restarted according to the user's instruction.

[0013] The second aspect of the present invention provides: A system for batch rapid uploading and downloading of a large number of small files, used to implement any one of the above methods for batch rapid uploading and downloading of a large number of small files, including: An initialization module, used to start an initialization program using the client, check the system running environment; establish a secure connection channel with the metadata server and the S3 object storage system, and adopt an encrypted communication protocol; load the client user interface resources and display the interactive interface for file uploading and downloading operations; An upload preprocessing module, used to preprocess the files to be uploaded after the client receives the small file upload instruction from the user, and determine whether the batch upload conditions are met; An intelligent packaging module, used to intelligently package the small files that meet the batch upload conditions to obtain a package; An upload execution module, used to split the package into multiple file blocks, match the hash values, and then upload them to the S3 object storage system; A download initialization module, used to obtain the information of the files to be downloaded specified by the user and perform file matching after the client receives the file download request from the user; A download execution module, used to execute the single file download process for a single file download request; for a batch file download request, execute the batch file download process.

[0014] The third aspect of the present invention provides: A computer-readable storage medium storing computer-executable instructions, which when loaded and executed by a processor, implement any of the above methods for batch fast uploading and downloading of a large number of small files.

[0015] The beneficial effects of the present invention are as follows: 1) Through the intelligent packaging module, batch processing of small files is carried out, and numerous small files are encapsulated into a package body according to multi-dimensional factors such as directory depth, file size, and quantity, effectively reducing the number of network requests. For example, originally thousands of small files may require thousands of separate network connection requests, but after packaging, it may only require several requests, greatly reducing the network load and the pressure on the server-side connection processing, improving the overall efficiency of data transmission, shortening the upload time required, and being particularly suitable for large-scale data upload scenarios such as enterprise data backup and cloud storage data synchronization and other business scenarios.

[0016] 2) In terms of downloading, with the help of the file range information stored in the metadata server, accurate positioning and efficient extraction can be achieved whether it is single-file downloading or batch downloading. For single-file downloading, according to the file hash, range, and signature information, through the object storage segmented downloading method, only the required part of the file is downloaded, avoiding unnecessary data transmission and greatly accelerating the acquisition speed of a single file. For example, when a user quickly downloads a specific document or picture, it can be obtained immediately. When batch-downloading all small files, multi-threaded breakpoint downloading combines file metadata and range information to regenerate the file and generate it on the disk according to the specified directory, not only ensuring the integrity and accuracy of the download, but also making full use of the multi-threaded concurrency advantage to achieve fast batch downloading and improving the overall download efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the method for batch fast uploading and downloading of a large number of small files; Figure 2 It is a processing flowchart of the method for batch fast uploading and downloading of a large number of small files. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.

[0019] First, an explanation of the terms of the present invention. Range information: It refers to specifying the range of the file content to be obtained through specific request headers in an HTTP request. Those skilled in the art can clearly understand its meaning.

[0020] The present invention aims to efficiently implement the batch rapid upload and download operations of a large number of small files. In terms of upload, through in-depth optimization of the front end, numerous small files are intelligently packaged, and metadata is aggregated and uploaded, effectively reducing the number of network requests. The multi-thread technology is used to achieve concurrent uploads. At the same time, combined with accurate progress display and interruption recovery mechanisms, it ensures the fluency and reliability of the user experience. The back end relies on a powerful distributed object storage system to process and store the uploaded files in parallel. Through the method of segmented downloading Range of object storage, it can quickly respond to single-file download requests, accurately extract files from the small-file package body of distributed storage, and transmit them to the user side at high speed. The achievements of this project will be widely applied in many fields such as big data analysis, cloud storage services, and enterprise-level file management, providing strong technical support for the processing of a large number of small files, significantly improving the processing efficiency and user satisfaction of related services, and strongly promoting the digital process and technological innovation development of the industry.

[0021] Refer to Figure 1 - Figure 2 , the first aspect of the present invention provides: A method for batch rapid upload and download of a large number of small files, including the following steps: Initialization stage: The client starts the initialization program, checks the system running environment; establishes a secure connection channel with the metadata server and the S3 object storage system, and uses an encrypted communication protocol; loads the client user interface resources and displays the interactive interface for file upload and download operations. Upload preprocessing stage: When the client receives the small file upload instruction from the user, it preprocesses the file to be uploaded and determines whether it meets the batch upload conditions. Intelligent packaging stage: Intelligently package the small files that meet the batch upload conditions to obtain a package body. Upload execution stage: Split the package body into multiple file blocks, match the hash values, and then upload them to the S3 object storage system. Download initialization stage: When the client receives the file download request from the user, it obtains the information of the file to be downloaded specified by the user and performs file matching. Download execution stage: For single-file download requests, execute the single-file download process; for batch file download requests, execute the batch file download process.

[0022] In this embodiment, the client is used to receive the user's small file upload action and trigger file upload; the file packaging module is used to perform intelligent calculations on a batch of small files according to directory depth, file size, file quantity, etc., and intelligently package the files into a package according to the set total number and total size, and record the position range of each file in the package; the file upload module is used to concurrently upload the package to the object storage system in a chunked upload manner, and at the same time upload the metadata information of all files and the range corresponding to each file to the metadata server for recording; the object storage system is used for S3 object storage and receives the file storage of the uploaded package; the metadata server is used to store the range information corresponding to the file, and at the same time provide the file range information during single file download; the file download module is used to download the specified content file in a segmented download manner of the object storage according to the obtained file hash, range and signature information.

[0023] For the upload link, in view of the fact that the traditional method of uploading a large number of small files one by one causes an explosive growth in the number of network requests, leading to serious network resource waste and low transmission rate problems, the present invention significantly reduces the network request frequency from the root by means of intelligent packaging of numerous small files in the front end and metadata aggregation upload strategy, and at the same time combines multi-threaded concurrent upload technology to fully exploit the potential processing capabilities of network bandwidth and front-end devices, thereby significantly accelerating the upload process.

[0024] In terms of the back-end design, the present invention constructs a powerful distributed object storage system. By means of this system, the uploaded files are processed and stored in parallel, which not only enhances the stability and reliability of storage, but also lays a solid foundation for subsequent rapid file extraction.

[0025] For the single file download requirement, adopting the object storage segmented download Range method can quickly and accurately locate and extract the target file in the small file package of distributed storage, and then transmit it to the user side at high speed, greatly shortening the download waiting time and meeting the user's expectation of quickly obtaining a single file.

[0026] Finally, it can be deeply integrated into key fields such as cloud storage services and enterprise-level file management, provide a solid and efficient technical foundation for the processing of a large number of small files, significantly promote the efficiency leap of related services in processing a large number of small files, and effectively improve the user's satisfaction with related services.

[0027] In some embodiments, the system operating environment includes the operating system version, network connection status, available memory and disk space; the encryption communication protocol is SSL / TLS; the interaction interface includes a file selection area, an upload progress bar, a download prompt box, and an operation log display area.

[0028] In this embodiment, checking the system running environment is to ensure that the basic requirements for small file upload and download operations are met. The SSL / TLS encryption communication protocol is used to ensure the confidentiality, integrity, and authenticity of data transmission, preventing information leakage and malicious tampering. The interactive interface provides users with friendly and intuitive operation guides and feedback displays. In some embodiments, the upload preprocessing stage further includes the following steps: When the client receives the user's small file upload instruction, it scans the file path specified by the user or obtains the list of files to be uploaded through a file selection dialog box, and obtains the detailed information of each file, including file name, file size, file type, creation time, modification time, and the directory path where the file is located. Count the total number of files to be uploaded, calculate the total size of all files, analyze the file distribution, including the number and size ratio of files at different directory depths, and construct a mapping relationship data structure of the file directory structure and file attributes. Judge whether the number and size of the files to be uploaded meet the batch upload conditions. If the batch upload conditions are not met, skip the file packaging step and directly enter the single file upload process; if the batch upload conditions are met, enter the file packaging module for intelligent encapsulation processing.

[0029] In this embodiment, a mapping relationship data structure of the file directory structure and file attributes is constructed for subsequent intelligent processing decisions. Judge whether the number and size of the files to be uploaded meet the preset conditions for the small file batch upload scenario. If there is a single or a small number of files (for example, the number of files is less than or equal to 5 and the total size is less than or equal to 10MB), skip the file packaging step and directly enter the single file upload process; if the batch upload conditions are met, enter the file packaging module for intelligent encapsulation processing.

[0030] In some embodiments, the intelligent packaging stage further includes the following steps: For small files that meet the batch upload conditions, the file packaging module starts an intelligent calculation algorithm. According to the directory depth-first principle, it starts screening files from the deepest-level directory, and sets the initial target file number threshold and target total size threshold for the packaging body. In the order of file size from small to large, the files are gradually added to the packaging body. Each time a file is added, its start byte position and end byte position in the packaging body are immediately calculated, the range value is determined, and it is associated with the file unique identifier for storage, constructing a local temporary packaging body index data structure. During the process of adding files, the current file number and total size of the packaging body are monitored in real time. When the preset threshold is reached, the file selection at the current directory level is paused, and the screening continues in the upper-level directory. Repeat this process until the packaging body is constructed or all files to be uploaded have been considered. If, after traversing all files, the package still does not meet the threshold requirements, the package is constructed according to the actual selected files, and the packaging information is recorded; The text files are compressed and then added to the package. The image sequence files with similar characteristics are merged or pre - processed and then added to the package, and the packaging algorithm is recorded.

[0031] In this embodiment, the initial threshold for the number of target files in the package is set to 300, and the threshold for the total target size is 256MB (these thresholds can be dynamically adjusted according to network bandwidth, server performance, and historical data statistics). For specific types of files (such as highly compressible text files, image sequence files with similar characteristics, etc.), targeted optimization strategies are adopted during packaging. For example, the text files are efficiently compressed and then added to the package, and the image sequence files are merged or pre - processed to reduce storage space occupancy and the amount of data transmitted, improving the overall packaging and uploading efficiency, and the packaging algorithm is recorded.

[0032] In some embodiments, the upload execution stage further includes the following steps: After the file packaging is completed, the file upload module divides the package into multiple file blocks according to a predetermined block size, generates multiple file blocks, and assigns a unique serial number and a corresponding hash value calculation task to each file block; For each file block, an independent upload task thread is created, an upload task thread pool containing multiple threads is constructed, and the file block upload tasks are assigned to the thread pool for concurrent execution; Before each file block is uploaded, its hash value is calculated, and the hash value and the file block data are encapsulated in the upload request and sent to the S3 object storage system; after receiving the upload request, the S3 object storage system first extracts the hash value and performs a hash calculation verification on the file block data. If the two are consistent, it confirms that the file block is successfully received and stored in the specified location; if not, it sends an error message to the client, requiring the client to re - upload the file block; As the file blocks are uploaded, the file upload module updates the upload progress information in real - time, including the number of uploaded file blocks, the total number of file blocks, the amount of uploaded data, the total amount of data, and synchronously updates the upload progress information to the progress bar of the user interface on the client; When all file blocks are uploaded and the S3 object storage system returns an upload success message, the file upload module sorts out the metadata information of all files, including file name, original path, file type, range information in the package, compression method; organizes the data according to a predefined data format, and then uploads this metadata information to the metadata server for storage; After receiving the metadata information, the metadata server performs data integrity verification and index construction operations, stores the information in the corresponding database table or data structure, and establishes an index relationship between the file metadata and the range information.

[0033] In this embodiment, the predetermined block size can be set to 64MB per block or other required sizes. Establishing an efficient index relationship between the file metadata and the range information can facilitate quick querying and positioning during subsequent file downloading and management.

[0034] In some embodiments, the download initialization stage further includes the following steps: When the client receives a file download request from the user, it obtains the information of the file to be downloaded specified by the user, including the file name and the file hash value. If the user only provides the file name, the client first searches in the locally cached metadata information. If not found, it sends a query request to the metadata server to obtain the complete metadata information of the file, including the hash value of the package it belongs to and the range information in the package.

[0035] In some embodiments, the single-file download process includes the following steps: For a single-file download request, the client calculates the range of file blocks to be downloaded based on the obtained file range information and the package information, and directly prepares to download the file block according to the recorded range. The client creates a download task thread, constructs a single-file download task thread pool, and distributes the file block download tasks to the thread pool for concurrent execution. During the download of each file block, the client sends a download request for the package hash value, the file range position, and the signature information to the S3 object storage system. After receiving the request, the S3 object storage system searches for the corresponding file block data based on this information and returns the file block data together with the calculated hash value to the client. After receiving the file block, the client first verifies the correctness of the hash value. If it is correct, it assembles the file block data according to the range information of the file in the package, gradually restoring the complete file data. If the hash value is incorrect, it requests the S3 object storage system to re-download the file block. As the file blocks are downloaded and assembled, the client real-time updates the download progress information, including the number of downloaded file blocks, the total number of file blocks, the amount of downloaded data, and the total amount of data, and synchronously updates these download progress information to the progress bar of the client's user interface. After all file chunks are downloaded and assembled into a complete file, the client performs a final integrity check on the file. For example, it recalculates the hash value of the file and compares it with the original hash value. If they are the same, it indicates that the file download is successful; if not, error handling is performed according to the situation, including attempting to redownload or prompting the user of the download failure and the reason for the failure.

[0036] In some embodiments, the batch file download process includes the following steps: For a batch file download request, the client calculates the distribution of all files in the package according to the obtained file metadata and range information, and determines the ranges of all file chunks to be downloaded. Allocate corresponding thread numbers according to the number of files, construct a thread pool for batch file download tasks, and distribute the file chunk download tasks to the thread pool for concurrent execution to perform multi-threaded breakpoint download. During the download process of each file chunk, send a request to the S3 object storage system, receive the file chunk, verify the hash value, and assemble the file to ensure the correct download of each file chunk and the complete restoration of the file. As the file chunks are downloaded and assembled, the client updates the overall progress information of the batch download in real time, including the number of downloaded files, the total number of files, the amount of downloaded data, and the total amount of data; and displays it on the progress bar of the user interface, and at the same time records the download status and detailed information of each file in the operation log display area. When all files are downloaded, the client re-creates the corresponding directory structure on the local disk according to the original directory information in the file metadata, and stores the downloaded files on the disk according to the original directory path to ensure that the organizational structure of the files is the same as when they were originally uploaded. During the batch download process, if a network interruption or user pause operation occurs, the client records the current download status information, including the downloaded file chunk information and the download progress of each file, and stores this information in the local cache; when the network resumes or the user continues the operation, the client reads the download status information from the local cache and continues to execute the download task from the breakpoint. If, during the download process, the S3 object storage system or the metadata server fails or has a response delay, the client automatically activates a fault response mechanism, including increasing the number of retries and the waiting time strategy, and attempts to connect to the server to obtain data multiple times within the first preset time; if the fault duration is greater than the second preset time, the download task is paused and the user is prompted with the fault information. After the server returns to normal, continue the download from the breakpoint or restart the download task according to the user's instruction.

[0037] Intelligent File Packaging and Indexing Technology: Through in-depth analysis of characteristics such as the directory depth, size, and quantity of a batch of small files, the present invention uses intelligent computing algorithms to encapsulate them into a packaged body. During the packaging process, the position range of each file in the packaged body is accurately recorded, and an efficient local temporary index is constructed. This technology effectively reduces the number of network requests and network overhead, and at the same time provides a basis for rapid and accurate data positioning for subsequent file upload, download, and management operations, greatly improving the overall efficiency and data manageability when processing a large number of small files.

[0038] Multi-threaded Concurrent Upload and Download Mechanism: In the upload process, the present invention divides the packaged body into blocks and uses multi-threaded concurrent upload to the object storage system, making full use of network bandwidth resources to achieve high-speed data transmission. At the same time, during download, according to the file hash, range, and signature information, threads are intelligently allocated according to the number of files to be downloaded. For single-file download, segmented download is adopted, and for batch download, multi-threaded breakpoint download strategy is used. This mechanism significantly improves the speed and stability of upload and download, effectively meets the high-concurrency requirements in the scenario of transmitting a large number of small files, and ensures the smoothness and timeliness of data transmission.

[0039] Metadata Server and Object Storage Collaboration Technology: The metadata server is responsible for storing the metadata information and range information of files, and can quickly provide key data guidance during file download. The S3 object storage focuses on receiving and storing the uploaded packaged body files. The two work together to achieve separate management of data storage and index information, ensuring both the efficiency and security of file storage, and being able to quickly locate the target file among a large number of small files through accurate metadata indexing. Whether it is a single-file or batch-file operation, it can respond efficiently, improving the reliability and performance of the entire system in processing a large number of small files.

[0040] The second aspect of the present invention provides: A system for batch and rapid upload and download of a large number of small files, used to implement any of the above methods for batch and rapid upload and download of a large number of small files, including: Initialization module, used to start the initialization program using the client, check the system running environment; establish a secure connection channel with the metadata server and the S3 object storage system, and adopt an encrypted communication protocol; load the client user interface resources and display the interactive interface for file upload and download operations; Upload preprocessing module, used to preprocess the files to be uploaded after the client receives the user's small file upload instruction, and determine whether the batch upload condition is met; Intelligent packaging module, used to intelligently package the small files that meet the batch upload conditions to obtain a packaged body; An upload execution module, configured to split a package into multiple file blocks, match hash values, and then upload them to an S3 object storage system; A download initialization module, configured to obtain information about a file to be downloaded specified by a user and perform file matching after the client receives a file download request from the user; A download execution module, configured to execute a single-file download process for a single-file download request and execute a batch-file download process for a batch-file download request.

[0041] A third aspect of the present invention provides: A computer-readable storage medium storing computer-executable instructions, which, when loaded and executed by a processor, implement any of the above methods for batch fast upload and download of a large number of small files.

[0042] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in related fields. Any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for quickly uploading and downloading a large number of small files in batches, characterized by: The following steps are involved: In the initialization phase, the client starts the initialization program and checks the system operating environment; establishes a secure connection channel with the metadata server and the S3 object storage system, using an encrypted communication protocol; loads the client user interface resources and displays the interactive interface for file upload and download operations; In the upload preprocessing stage, when the client receives the small file upload instruction from the user, it preprocesses the uploaded file and determines whether it meets the batch upload conditions; In the intelligent packaging stage, small files that meet the batch upload conditions are intelligently packaged to obtain a package body; In the upload execution phase, the package body is split into multiple file blocks and uploaded to the S3 object storage system after matching the hash values; In the download initialization phase, when the client receives the user's file download request, it obtains the information of the file to be downloaded specified by the user and performs file matching; During the download execution phase, for a single file download request, the single file download process is executed; for a batch file download request, the batch file download process is executed.

2. The method for quickly uploading and downloading a large number of small files in batches according to claim 1 is characterized in that: The system operating environment includes the operating system version, network connection status, available memory and disk space; the encrypted communication protocol is SSL / TLS; the interactive interface includes a file selection area, an upload progress bar, a download prompt box, and an operation log display area.

3. The method for quickly uploading and downloading a large number of small files in batches according to claim 1 is characterized in that: The upload preprocessing stage also includes the following steps: When the client receives the user's small file upload instruction, it scans the file path specified by the user or obtains the list of files to be uploaded through the file selection dialog box, and obtains detailed information of each file, including file name, file size, file type, creation time, modification time and directory path where the file is located; Count the total number of files to be uploaded, calculate the total size of all files, analyze the distribution of files, including the number and size ratio of files at different directory depths, and build a data structure of the mapping relationship between file directory structure and file attributes; Determine whether the number and size of files to be uploaded meet the batch upload conditions. If not, skip the file packaging step and directly enter the single file upload process; if the batch upload conditions are met, enter the file packaging module for intelligent packaging processing.

4. The method for quickly uploading and downloading a large number of small files in batches according to claim 1 is characterized in that: The intelligent packaging stage also includes the following steps: For small files that meet the batch upload conditions, the file packaging module starts the intelligent calculation algorithm, screens files from the deepest directory based on the directory depth priority principle, and sets the initial package target file quantity threshold and target total size threshold; Add files to the package body gradually in the order of file size from small to large. Each time a file is added, its starting byte position and ending byte position in the package body are calculated immediately, the range value is determined, and stored in association with the file unique identifier to build a local temporary package body index data structure; During the process of adding files, the current number of files and the total size of the package are monitored in real time. When the preset threshold is reached, the file selection of the current directory level is suspended, and the selection continues in the directory of the next level. This process is repeated until the package is built or all the files to be uploaded have been considered. If the packaged body still does not meet the threshold requirement after traversing all files, the packaged body is built according to the actual selected files and the package information is recorded; Text files are compressed and then added to the package body. Image sequence files with similar features are merged or pre-processed and then added to the package body. The packaging algorithm is recorded.

5. The method for quickly uploading and downloading a large number of small files in batches according to claim 1 is characterized in that: The upload execution phase further includes the following steps: After the file is packaged, the file upload module divides the package body into a predetermined block size to generate multiple file blocks, and assigns a unique serial number and a corresponding hash value calculation task to each file block; For each file block, create an independent upload task thread, build an upload task thread pool containing multiple threads, and assign the file block upload tasks to the thread pool for concurrent execution; Before uploading each file block, calculate its hash value, encapsulate the hash value and file block data in the upload request and send it to the S3 object storage system; after receiving the upload request, the S3 object storage system first extracts the hash value and performs hash calculation verification on the file block data. If the two are consistent, it confirms that the file block has been successfully received and stored in the specified location; if they are inconsistent, an error message is sent to the client, requiring the client to re-upload the file block; As the file blocks are uploaded, the file upload module updates the upload progress information in real time, including the number of uploaded file blocks, the total number of file blocks, the uploaded data volume, and the total data volume, and synchronously updates the upload progress information to the progress bar of the client's user interface; When all file blocks are uploaded and the S3 object storage system returns the upload success information, the file upload module organizes the metadata information of all files, including the file name, original path, file type, range information in the package body, and compression method; organizes the data according to the predefined data format, and then uploads the metadata information to the metadata server for storage; After receiving the metadata information, the metadata server performs data integrity verification and index building operations, stores the information in the corresponding database table or data structure, and establishes an index relationship between the file metadata and the range information.

6. The method for quickly uploading and downloading a large number of small files in batches according to claim 1 is characterized in that: The download initialization phase also includes the following steps: When the client receives the user's file download request, it obtains the information of the file to be downloaded specified by the user, including the file name and file hash value; if the user only provides the file name, the client first searches in the metadata information cached locally. If it is not found, it sends a query request to the metadata server to obtain the complete metadata information of the file, including the hash value of the package body and the range information in the package body.

7. The method for quickly uploading and downloading a large number of small files in batches according to claim 1 is characterized in that: The single file download process includes the following steps: For a single file download request, the client calculates the range of the file block to be downloaded based on the obtained file range information and package body information, and directly prepares to download the file block based on the recorded range; The client creates a download task thread, builds a single file download task thread pool, and assigns the file block download tasks to the thread pool for concurrent execution; During the download process of each file block, the client sends a download request to the S3 object storage system, including the package body hash value, file range location, and signature information. After receiving the request, the S3 object storage system searches for the corresponding file block data based on this information and returns the file block data and the calculated hash value to the client. After receiving the file block, the client first verifies the correctness of the hash value. If it is correct, the file block data is assembled according to the range information of the file in the package body, and the complete file data is gradually restored. If the hash value is wrong, the client requests the S3 object storage system to re-download the file block. As the file blocks are downloaded and assembled, the client updates the download progress information in real time, including the number of downloaded file blocks, the total number of file blocks, the amount of data downloaded, and the total amount of data, and synchronously updates the download progress information to the progress bar on the client's user interface; When all file blocks are downloaded and assembled into a complete file, the client performs a final integrity check on the file, such as recalculating the file's hash value and comparing it with the original hash value. If they are consistent, it means that the file has been downloaded successfully; if they are inconsistent, error handling is performed according to the situation, including attempting to re-download or prompting the user that the download failed and the reason for the failure.

8. The method for quickly uploading and downloading a large number of small files in batches according to claim 1 is characterized in that: The batch file download process includes the following steps: For batch file download requests, the client calculates the distribution of all files in the package based on the obtained file metadata and range information, and determines the range of all file blocks that need to be downloaded; Allocate the corresponding number of threads according to the number of files, build a batch file download task thread pool, assign the file block download tasks to the thread pool for concurrent execution, and perform multi-threaded breakpoint downloading; During the download process of each file block, send a request to the S3 object storage system, receive the file block, verify the hash value, and assemble the file to ensure the correct download of each file block and the complete restoration of the file; As the file blocks are downloaded and assembled, the client updates the overall progress information of the batch download in real time, including the number of files downloaded, the total number of files, the amount of data downloaded, and the total amount of data; and displays it on the progress bar of the user interface, while recording the download status and detailed information of each file in the operation log display area; When all files are downloaded, the client recreates the corresponding directory structure on the local disk according to the original directory information in the file metadata, and stores the downloaded files on the disk according to the original directory path to ensure that the file organization structure is consistent with the original upload; During the batch download process, if the network is interrupted or the user pauses the operation, the client records the current download status information, including the downloaded file block information and the download progress of each file, and stores this information in the local cache; when the network is restored or the user continues the operation, the client reads the download status information from the local cache and continues the download task from the breakpoint; If the S3 object storage system or metadata server fails or responds late during the download process, the client automatically starts the fault response mechanism, including increasing the number of retries and waiting time strategies, and trying to connect to the server to obtain data multiple times within the first preset time; if the fault duration is greater than the second preset time, the download task is paused and the user is prompted with fault information. After the server returns to normal, the download is continued from the breakpoint or the download task is restarted according to the user's instructions.

9. A system for quickly uploading and downloading a large number of small files in batches, characterized by: The method for realizing rapid batch uploading and downloading of a large number of small files as claimed in any one of claims 1 to 8 comprises: The initialization module is used to start the initialization program using the client and check the system operating environment; establish a secure connection channel with the metadata server and the S3 object storage system using an encrypted communication protocol; load the client user interface resources and display the interactive interface for file upload and download operations; The upload preprocessing module is used to preprocess the uploaded files and determine whether the batch upload conditions are met after the client receives the small file upload instruction from the user; Intelligent packaging module, used to intelligently package small files that meet the batch upload conditions to obtain a package body; The upload execution module is used to split the package into multiple file blocks and upload them to the S3 object storage system after matching the hash values; The download initialization module is used to obtain the information of the file to be downloaded specified by the user and perform file matching after the client receives the file download request from the user; The download execution module is used to execute the single file download process for a single file download request; and execute the batch file download process for a batch file download request.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by the processor, the method for quickly uploading and downloading a large number of small files in batches as described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Third-party code snippet replacement method and device, terminal and storage medium

    CN111638907A

  • Distributed batch file downloading method and device, computer equipment and storage medium

    CN112199442A

  • File batch downloading method and device

    CN115412546A

  • Method and apparatus for a distributed file storage and indexing service

    IN201918011824A

  • Forwarding element with flow learning circuit in its data plane

    US10616101B1

Cited By

  • File uploading method and device, storage medium, electronic equipment and program product

    CN120358232A

  • Downloading progress monitoring method bypassing pipeline buffer and application thereof

    CN122261946A