Cross-domain large file transmission method and device based on dynamic fragmentation

By optimizing large file transfers through dynamic segmentation and multi-threaded transmission, it solves the transmission difficulties in complex network environments and realizes efficient and secure cross-domain file transfer.

CN120639756APending Publication Date: 2025-09-12HANGZHOU SHULAN TECH CO LTD

Patent Information

Application Number
CN202510531681.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In complex network environments, the traditional TCP protocol is difficult to adapt to dynamic bandwidth changes, resulting in difficulties in transmitting large files, high packet loss rates, low cross-domain transmission efficiency, and inability to meet security and compliance requirements.

Method used

A dynamic fragmentation method is adopted to adjust the fragment size according to the file type and network conditions, combined with multi-threaded transmission and intermediate system storage, using the QUIC protocol for synchronous transmission, and optimizing the transmission process by dynamically adjusting the fragment size parameters and congestion control algorithm.

Benefits of technology

It achieves high-speed cross-domain transmission of large files, supports timeout retransmission and breakpoint resumption, improves transmission efficiency and reliability, and meets security and compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639756A_ABST
    Figure CN120639756A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-domain large file transmission method and device based on dynamic fragmentation. The invention aims to solve the problems of maximum capacity limitation, network instability, security and the like in the cross-domain transmission process of large files. According to the method, efficient and safe transmission of large files is realized by dynamically adjusting fragment size parameters and combining encryption transmission, breakpoint resume and multi-task concurrency control technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data transmission technology, and in particular to a cross-domain large file transmission method and device based on dynamic sharding. Background Art

[0002] With the rapid development of cloud computing, edge computing and Internet of Things technologies, the amount of data is also growing explosively, and the demand for data transmission is increasing.

[0003] However, existing technologies face the following challenges: 1) Packet loss may occur in complex network environments (for example, when the network connection is unstable); 2) Signal propagation between nodes may be prolonged during cross-domain transmission (for example, across physical regions and / or network domains), and the traditional TCP protocol congestion control mechanism is difficult to adapt to dynamic bandwidth changes; 3) Large files (especially those with a size of 1G or above) are difficult to transmit, and existing technologies face the dilemma of I / O bottlenecks and a sharp drop in transmission efficiency when transmitting large files, resulting in the inability to effectively receive files; 4) Data transmission cannot meet relevant security and compliance requirements.

[0004] In this regard, the present application proposes a cross-domain large file transmission method and device based on dynamic sharding to address the above challenges. Summary of the Invention

[0005] The present invention provides a file transmission method, which includes dynamically adjusting a slice size parameter during the transmission of a target file.

[0006] According to one embodiment of the present invention, dynamically adjusting the slice size parameter includes dynamically adjusting the slice size parameter according to one or more of the type of the target file, the network condition, and the transmission error rate.

[0007] According to one embodiment of the present invention, the method further comprises transmitting multiple slice files of the target file in multiple threads, wherein the transmission order of different slice files in corresponding threads is dynamically adjusted based on priority.

[0008] According to one embodiment of the present invention, the transmission is performed across domains, and the size of the target file is greater than or equal to 1G.

[0009] According to one embodiment of the present invention, the method further includes: uploading multiple segment files of the target file from the sender to the intermediate system, and then synchronizing them from the intermediate system to the receiver, wherein dynamic adjustment occurs during the uploading.

[0010] According to one embodiment of the present invention, the method further comprises dynamically adjusting the fragment size parameter according to the following formula:

[0011]

[0012] Where M is the shard size parameter before adjustment, t is the time required to transfer the preset number of shard files, e is the number of shard files with transmission errors, c is the total number of shard files that have been transferred, P is the adjusted shard size parameter, R is the baseline throughput calibration factor with a value range of 1 to 10, and a is the error rate sensitivity index with a value range of 0.5 to 2.

[0013] According to one embodiment of the present invention, the method further comprises updating R according to the following formula:

[0014]

[0015] The actual average throughput refers to the number of shard files successfully transmitted per unit time, and the theoretical maximum throughput refers to the theoretical peak value of the current network bandwidth.

[0016] The present invention also provides a file transmission device for executing the above method.

[0017] The present invention also provides a computer-readable medium having a computer program stored thereon, which implements the above-mentioned method when executed by a processor.

[0018] The above technical solution provided by the present invention supports high-speed cross-domain transmission of large files, can perform timeout retransmission and breakpoint resumption according to the task status, and is efficient, safe and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a schematic flowchart of a method for cross-domain large file transmission based on dynamic segmentation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The present invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are only provided to enable those skilled in the art to better understand and implement the present invention, rather than to limit the scope of the present invention.

[0021] As used herein, the term "including" and variations thereof should be interpreted as open-ended terms meaning "including, but not limited to." The term "based on" should be interpreted as "based, at least in part, on." The terms "one embodiment" and "an embodiment" should be interpreted as meaning "at least one embodiment." The term "another embodiment" should be interpreted as meaning "at least one other embodiment."

[0022] Refer to the following Figure 1A file transfer method based on dynamic segmentation according to an embodiment of the present invention is described. In an embodiment of the present invention, one or more steps of the method of the present invention may be omitted, performed separately, or performed in combination, and are not necessarily performed in the order described below. In an embodiment of the present invention, the method of the present invention may include additional steps. In an embodiment of the present invention, the method of the present invention is initiated automatically, by the sender, or by the receiver. In an embodiment of the present invention, the method of the present invention can be used for unidirectional transmission of a single file or for simultaneous multidirectional transmission of multiple files.

[0023] In embodiments of the present invention, "sharding" refers to cutting a continuous data stream (e.g., a single file) into multiple segments. In embodiments of the present invention, the segments generated after sharding may be referred to as "shard files." In embodiments of the present invention, a "shard size parameter" defines the size of the shard files generated by the sharding operation.

[0024] Split the target file to be transferred

[0025] In the embodiments of the present invention, various existing technologies can be used to perform slicing operations as needed. For example: (1) P2P transmission standards, including the itTorrent protocol (which divides files into 256KB slicing files) and eMule's 9.28MB slicing standard; (2) HTTP large file upload standards, including the Chunked TransferEncoding defined in HTTP / 1.1's RFC 7230 and the Blob.slice() method in the HTML5 File API; and (3) development frameworks, including the slicing functions of front-end upload components such as WebUploader and resumable upload libraries such as Resumable.js.

[0026] In an embodiment of the present invention, the method includes performing initial fragmentation based on the file type, total size, etc. of the target file. In an embodiment of the present invention, video files are initially fragmented larger (for example, 10M / piece) because a small amount of packet loss can be allowed, while medical imaging files are initially fragmented smaller (for example, 5M / piece) because reliability needs to be guaranteed. In an embodiment of the present invention, when the target file is larger than 1G, the initial fragmentation can be 10M / piece; and when the target file is smaller than 10M, no fragmentation is performed or it is divided into small fragment files (for example, 1M / piece). In an embodiment of the present invention, those skilled in the art can preset the size of the initial fragmentation as needed.

[0027] In an embodiment of the present invention, the sharding operation may generate all shard files at once before the transmission operation. In other embodiments of the present invention, the sharding operation may generate only one or more shard files to be transmitted before the transmission operation, and then gradually generate the remaining shard files as needed during the transmission operation.

[0028] Transferring fragmented files

[0029] In an embodiment of the present invention, the method can adopt multi-threading for transmission. In an embodiment of the present invention, multi-threading tasks are automatically coordinated, for example, the transmission order of the slice files is dynamically adjusted based on the priority.

[0030] In an embodiment of the present invention, the method may create a temporary file during transmission. In an embodiment of the present invention, a "temporary file" refers to metadata created during the transmission phase, which is used to track the status of the transmission task and does not store the actual content of the sliced ​​files. In an embodiment of the present invention, a temporary file may record one or more of the following: the size of the target file to be transmitted, the number of sliced ​​files, the number of sliced ​​files that have been successfully transmitted, the record of dynamic adjustment of sliced ​​files, the mapping relationship between the transmission task and the sliced ​​files, etc. In an embodiment of the present invention, the temporary file may be deleted after the transmission is completed to release the storage space.

[0031] In the embodiment of the present invention, if an exception occurs in the transmission of a certain segment file (eg, power failure), the transmission is retransmitted or resumed. In the embodiment of the present invention, the above mapping relationship is used to achieve breakpoint resume.

[0032] In an embodiment of the present invention, the method may first upload the segmented files from the sender to an intermediate system (eg, a cloud platform), and then synchronize the segmented files from the intermediate system to the receiver.

[0033] Upload

[0034] In an embodiment of the present invention, the method may create a temporary file during upload. In an embodiment of the present invention, the temporary file is associated with the upload task. In an embodiment of the present invention, after the upload of the segmented files is completed, the upload is marked as successful. The upload progress of the target file can be determined based on the number of segmented files in the temporary file and the number of successfully transferred segmented files.

[0035] Stored in intermediate systems

[0036] In an embodiment of the present invention, the inventors specifically selected the existing memory-mapped file (MMAP) storage mechanism in the art to store sharded files in an intermediate system. In an embodiment of the present invention, MMAP technology uses a page table mechanism to establish a linear mapping relationship between the disk physical address of the target file or storage object and the process virtual address space, forming a continuous virtual memory area. In an embodiment of the present invention, the mapping process includes address space allocation, offset calibration, and length matching verification steps to ensure the accurate correspondence between file data and virtual addresses. In an embodiment of the present invention, when a process accesses mapped memory through a pointer, the system automatically tracks the read and write status of the memory page, marks the modified memory page as a "dirty page", and triggers an asynchronous write-back mechanism. Within a preset time window or when the memory pressure threshold is triggered, the changed data is flushed back to the corresponding disk sector, thereby achieving system-level data persistence without calling traditional I / O interfaces. In an embodiment of the present invention, the mapping mechanism supports bidirectional synchronization. Any data changes in the kernel space to the mapped area are synchronized to the user space view in real time. At the same time, write operations in the user space trigger kernel updates through the page protection mechanism, thereby achieving zero-copy data sharing across processes. It is particularly suitable for scenarios where multiple nodes in a distributed system access shared storage media in parallel.

[0037] In an embodiment of the present invention, a user process initiates an mmap() system call to trigger a context switch from user mode to kernel mode, establishes a direct mapping relationship between the user space cache and the kernel buffer, and the kernel layer directly moves disk data to the kernel buffer through the DMA controller, relying on the Page Cache mechanism to implement zero-copy caching, and restores the user mode context after the mapping is completed.

[0038] In an embodiment of the present invention, when performing synchronous data transmission, the user process calls the write() function to enter the kernel state again, the CPU copies the data in the kernel buffer to the socket send buffer, and the DMA controller asynchronously executes the data movement from the socket buffer to the network card device to complete the network transmission. Finally, after the write() function returns, it switches to the user state context to optimize storage and transmission efficiency.

[0039] In summary, MMAP technology avoids traditional I / O data copying by directly mapping kernel buffers to user-space virtual addresses. It also utilizes dirty page flushing and bidirectional synchronization mechanisms to ensure consistency when multiple nodes access shared storage in a distributed environment. In embodiments of the present invention, those skilled in the art may also select other storage mechanisms as needed.

[0040] synchronous

[0041] In an embodiment of the present invention, after a shard file is uploaded and stored, it is synchronized to the recipient. In an embodiment of the present invention, if a task is successfully created, the file and task mappings between the intermediate system and the recipient are recorded. Subsequent synchronization of the shard file will invoke the data center's interface based on this mapping to transfer the shard file. In an embodiment of the present invention, if synchronization is successful, the synchronization record is updated to success.

[0042] In an embodiment of the present invention, the inventors specifically select the reliable data transmission method of QUIC based on UDP known in the art for synchronization. In an embodiment of the present invention, the reliable data transmission method of QUIC based on UDP, for example, steps are as follows:

[0043] During the first communication, a zero round-trip time (0RTT) connection is established. The sending client preloads the session ticket previously negotiated by the server, which contains encryption parameters and key information. The client generates a 64-bit random connection identifier and initializes the encrypted channel. It sends a Client Hello message containing the initial key derivation function parameters via a UDP packet.

[0044] After the receiving client verifies the validity of the connection identifier, it generates a session response based on the pre-shared key and completes the key negotiation and handshake protocol through a single UDP round trip; the client directly embeds the authenticated encrypted application data in the first data packet to achieve 0RTT data transmission.

[0045] After establishing a zero round-trip time (0RTT) connection, each transmitted data packet is assigned a globally unique and strictly increasing Packet Number (PN), replacing the traditional TCP sequence number mechanism. When the receiver detects packet loss, it generates an ACK frame containing the missing PN range and sends it to the sender.

[0046] After receiving the ACK, the sender calculates the exact round-trip time (RTT) based on the PN timestamp:

[0047] RTT=(Current Time-PN Timestamp)

[0048] +Ack Delay Compensation

[0049] If it is detected that PN arrives out of order, the fast retransmission mechanism is triggered to give priority to retransmitting the key data packets of the high priority stream.

[0050] In an embodiment of the present invention, a pluggable congestion control algorithm (such as Cubic or BBR) is selected during dynamic congestion control initialization. The algorithm parameters are dynamically configured through the control plane. The receiving end sends a WINDOW_UPDATE frame based on the buffer occupancy rate and dynamically adjusts the flow-level window according to the following formula:

[0051] Allowed Window

[0052] =Max(Received Offset,Last Ack Offset)

[0053] +Flow Control Credit

[0054] Connection-level flow control implements global rate limiting by aggregating the sum of each flow window. When memory pressure is detected, the congestion window exponential backoff algorithm is triggered.

[0055] In an embodiment of the present invention, data streams are divided into independently encrypted packet sequences, with each stream identifier bound to a unique service logic channel. The receiving end maintains an independent stream state machine and implements selective acknowledgment (SACK) for packets arriving out of order, allowing non-critical streams to be continuously reassembled in the background. When packet loss is detected for a specific stream, only data transmission for that stream is suspended, while other streams continue to process normally.

[0056] In an embodiment of the present invention, a 64-bit random Connection ID is assigned to each QUIC connection instead of the traditional five-tuple identifier. When the network is switched, the terminal announces the new address through a new Connection ID frame, the receiving end synchronously updates the routing table, and uses forward error correction (FEC) redundant packets to ensure data integrity during the migration process, ensuring seamless migration of the connection context.

[0057] In an embodiment of the present invention, a reliable data transmission method based on UDP QUIC synchronizes the fragmented files, uses an incremental packet sequence number and a fast retransmission mechanism to ensure data reliability, combines a pluggable congestion control algorithm with a hierarchical flow control strategy to improve transmission efficiency in high-load scenarios, and achieves seamless migration of network switching through a globally unique connection identifier and redundant error correction technology. In a cross-domain network environment, it significantly reduces handshake delays, improves throughput, supports massive concurrent stream processing, and minimizes service interruption time, thereby improving the reliability of cross-domain transmission. In an embodiment of the present invention, those skilled in the art may also select other synchronization mechanisms as needed.

[0058] Dynamically adjust shard size parameters

[0059] In an embodiment of the present invention, the method includes dynamically adjusting the slice size parameter during transmission based on factors such as the type of target file, network conditions, and transmission error rate. For example, when the network is poor, the slice size parameter is automatically adjusted to a smaller size, such as from 10M / slice to 2M / slice, or vice versa. In an embodiment of the present invention, the method includes periodically calculating the transmission time and error rate of the slice file and then adjusting the slice size parameter. In an embodiment of the present invention, the slice size parameter can be adjusted according to the following formula:

[0060]

[0061] Where M is the shard size parameter before adjustment, t is the time required to transfer the preset number of shard files, e is the number of shard files with transmission errors, c is the total number of shard files that have been transferred, P is the adjusted shard size parameter, R is the baseline throughput calibration factor with a value range of 1 to 10, and a is the error rate sensitivity index with a value range of 0.5 to 2.

[0062] In the embodiment of the present invention, the baseline throughput calibration factor (R) is used to compensate for differences between the network environment and the ideal state. For example, the value for a local area network / high stability network is 2-3, the value for a public network / general network is 1-1.5, and the value for a weak network environment is 0.5-1.

[0063] In an embodiment of the present invention, after each transmission of a predetermined number (for example, 100) of segment files is completed, the update may be performed according to the following formula:

[0064]

[0065] The actual average throughput refers to the amount of data successfully transmitted per unit time, and the theoretical maximum throughput refers to the theoretical peak value of the current network bandwidth.

[0066] In an embodiment of the present invention, the error rate sensitivity index (a) indicates the degree of sensitivity to the error rate. High reliability takes a value of 1.5 to 2, and a slight increase in the error rate significantly reduces the size of the fragmented file; the network balancing mode takes a value of 1 to 1.5, in which case the adjustment of the fragmented file size responds linearly to the change in the error rate; high throughput priority takes a value of 0.5 to 1, which can tolerate a certain error rate to maintain a large fragmented file size. In an embodiment of the present invention, medical imaging type files require high reliability, and a can be set to 1.8. Ordinary files require moderate reliability, and a can be set to 1.2. Video files can tolerate a small amount of packet loss, and a can be set to 0.6.

[0067] In an embodiment of the present invention, the method can calculate and adjust the slice size parameters according to different situations of each thread.

[0068] In an embodiment of the present invention, if all fragmented files have been generated at once prior to a transfer operation, the size of one or more fragmented files that have not been transferred can be modified during the transfer operation based on the adjusted fragment size parameter. In an embodiment of the present invention, the modification includes merging, splitting, increasing or decreasing the size of the fragmented files, or a combination thereof. In an embodiment of the present invention, if fragmented files are generated incrementally during the transfer operation, subsequent fragmentation operations are performed based on the adjusted fragment size parameter.

[0069] Verify shard files

[0070] In an embodiment of the present invention, the shard files can be verified after the transmission is completed, for example, after each shard file is transmitted or after all shard files are transmitted. In an embodiment of the present invention, common verification mechanisms in the art can be used, such as real-time MD5, SHA series, shard metadata verification, encrypted signature, transport layer verification, error correction code. In an embodiment of the present invention, the method includes correspondingly calculating a verification code (such as an MD5 value) after the sharding is completed for verification. In an embodiment of the present invention, each shard file is verified after arriving at the cloud platform and before being stored.

[0071] In an embodiment of the present invention, performing an MD5-based verification includes generating a digest of the sliced ​​data based on the MD5 algorithm after transmitting the sliced ​​files and before merging the sliced ​​files. In an embodiment of the present invention, the method includes pre-storing the digest value after slicing to facilitate integrity verification after transmission. In an embodiment of the present invention, the MD5 algorithm pads the input data to a fixed length and then appends a 64-bit binary representation of the original length; slices the padded data according to a set slice size and splits it into multiple sub-blocks; after initializing the buffer, loops through each block, combining nonlinear functions, constant tables, and bit shift operations to update register states block by block; and finally, merges the values ​​of different registers to form a verification result.

[0072] Encrypt / decrypt shard files

[0073] In an embodiment of the present invention, the segmented files may be encrypted before transmission and decrypted after transmission. Those skilled in the art may change the timing of the encryption / decryption operation as needed. In an embodiment of the present invention, encryption / decryption methods known in the art, such as asymmetric encryption mechanisms, may be used. Encrypted transmission includes the following steps:

[0074] The client initiates an RSA public key request to the server and receives a response message containing the public key parameters. Based on the logged-in context information, the client constructs key generation metadata, which consists of a user identity field (the first 10 bytes), a timestamp field (the lower 10 bytes, with the upper bits padded with 0xFF), and a session token field (the first 12 bytes). This concatenation is encrypted using the SHA-256 hash algorithm to generate a 256-bit AES symmetric key.

[0075] The server's public key is used to encrypt and encapsulate the AES key, generate a key exchange certificate, and transmit it to the server through a secure channel. After the file is divided into fixed-size data units, it is encrypted using the AES-GCM algorithm combined with the session key. Each encrypted fragment carries a unique serial number and message authentication code (MAC). The server obtains the AES key through private key decryption, decrypts the received fragments in turn, verifies the MAC value, and reconstructs the original file based on the fragment sequence number.

[0076] Merge the fragment files into the target file

[0077] In an embodiment of the present invention, various merging methods known in the art may be used to merge the fragment files back into a target file.

[0078] The methods and devices of the various embodiments of the present invention can be implemented as pure software modules (such as software programs written in Java), or as pure hardware modules (such as dedicated ASIC chips or FPGA chips) as needed, or as modules that combine software and hardware (such as a firmware system that stores fixed code).

[0079] Another aspect of the present invention is a computer-readable medium having computer-readable instructions stored thereon, which, when executed, can implement the methods of various embodiments of the present invention.

[0080] Those skilled in the art will appreciate that the foregoing is merely exemplary embodiments of the present invention and is not intended to limit the present invention. The present invention may also include various modifications and variations. Any modifications and variations made within the spirit and scope of the present invention should be included within the scope of protection of the present invention.

Claims

1. A file transmission method, characterized in that: include: During the transfer of the target file, the fragment size parameter is dynamically adjusted.

2. The file transfer method according to claim 1, wherein: The dynamically adjusting the slice size parameter includes dynamically adjusting the slice size parameter according to one or more of the type of the target file, the network condition, and the transmission error rate.

3. The file transmission method according to claim 1, wherein: Also includes: The plurality of slice files of the target file are transmitted by multithreading, wherein the transmission order of the different slice files in the corresponding threads is dynamically adjusted based on the priority.

4. The file transmission method according to claim 1, wherein: The transmission is performed across domains, and the size of the target file is greater than or equal to 1G.

5. The file transmission method according to claim 1, wherein: Also includes: The plurality of segment files of the target file are uploaded from the sender to an intermediate system, and then synchronized from the intermediate system to the receiver, wherein the dynamic adjustment occurs during the uploading.

6. The file transmission method according to claim 2, wherein: It also includes dynamically adjusting the slice size parameter according to the following formula: Where M is the shard size parameter before adjustment, t is the time required to transfer the preset number of shard files, e is the number of shard files with transmission errors, c is the total number of shard files that have been transferred, P is the adjusted shard size parameter, R is the baseline throughput calibration factor with a value range of 1 to 10, and a is the error rate sensitivity index with a value range of 0.5 to 2.

7. The file transmission method according to claim 6, wherein: It also includes updating R according to the following formula: The actual average throughput refers to the number of shard files successfully transmitted per unit time, and the theoretical maximum throughput refers to the theoretical peak value of the current network bandwidth.

8. A file transmission device for executing the method according to any one of claims 1 to 7.

9. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Priority dynamic adjustment-based remote multi-file transmission method

    CN106790633A

  • File transmission method and storage medium

    CN117579573A

  • Data security transmission system based on quantum encryption technology

    CN118842638A

  • Data transmission optimization method and device in Internet of Vehicles environment, terminal equipment and storage medium

    CN119450397A

Cited By

  • Efficient file transmission method for multi-level network

    CN122513389A