Big data file transmission method and terminal

By dividing big data files into multiple classification files and independent sub-file cluster sets, determining the access order based on the offset of the file segments, establishing multi-level associations for integrity verification, the shortcomings of traditional data transmission methods in file structure integrity and sequence guarantee in big data transmission are solved, and more efficient, compatible and secure data transmission is achieved.

CN120223689APending Publication Date: 2025-06-27FUJIAN TIANQUAN EDUCATION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510217540.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional data transmission methods are difficult to ensure the integrity of the file structure and the order of file segments during the big data transmission process, resulting in data loss or disorder, and the calculation amount is large and the retransmission efficiency is low, making it difficult to meet the high reliability requirements of large data multi-source and multi-node concurrent transmission.

Method used

By dividing big data files into multiple classification files according to file sensitivity levels and extracting content based on content sensitivity levels, an independent sub-file cluster collection is formed. Then, each basic file is divided into file segments, the file segment access order is determined based on the position offset of the file segment, and a multi-level association is established for integrity verification.

Benefits of technology

Through multi-level offset association and integrity verification of file segment access order, the compatibility and security of big data transmission are improved, and a more efficient and flexible data transmission mechanism is provided, and computing overhead and retransmission risks are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223689A_ABST
    Figure CN120223689A_ABST
Patent Text Reader

Abstract

The invention discloses a big data file transmission method and a terminal, and the method comprises the steps: dividing each file in a transmitted big data file based on a file sensitivity level, so as to obtain a plurality of classified files; performing content extraction on each classification file based on the content sensitive level to obtain an independent sub-file cluster set corresponding to a second classification label; performing file segment division on each basic file in each independent sub-file cluster set, and determining a file segment access sequence of each basic file according to the position offset of the file segment in each basic file; and establishing a first association between the first file segment of each basic file and the second classification label, establishing a second association between each file segment in each basic file and the second classification label, and performing integrity verification on transmission of the big data file according to the first association and the second association. In this way, integrity verification is performed through multi-level offset association and file segment access order.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data transmission, and particularly to a method and a terminal for transmitting big data files. Background Art

[0002] With the development of big data technology, the security and compatibility of data during transmission have become important technical research fields. In complex data processing and application environments, data usually contains a large number of files and file segments, and needs to be efficiently and reliably transmitted between multiple nodes. This kind of transmission needs to ensure the integrity of the file structure and the order of file segments to prevent data loss or disorder caused by delay, misalignment or packet loss during multi-level transmission. Therefore, ensuring the security and compatibility of big data transmission has become one of the key topics in the fields of data communication, network security and distributed computing.

[0003] Traditional data transmission methods usually use holistic verification mechanisms, such as checksum, hash verification, file size verification, etc., to ensure data integrity. These methods mainly calculate the overall verification value of the file and compare this value before and after data transmission to determine whether the data has been tampered with or lost. In addition, some systems adopt a segmented verification method, dividing the file into multiple data segments and independently verifying and transmitting them, so that fast retransmission can be achieved when some data segments are in error during the transmission process. The advantage of this method is relatively simple and low in implementation cost. However, with the expansion of the scale of big data and the complexity of data files, traditional verification methods gradually expose disadvantages such as large computational amount, low retransmission efficiency and insufficient adaptability to complex scenarios, and it is difficult to meet the high reliability requirements of multi-source and multi-node concurrent transmission of big data.

[0004] The defects of current technologies in data integrity and transmission reliability are mainly reflected in the following aspects: First, the traditional holistic verification and segmented verification methods are difficult to meet the security requirements of the multi-level file structure during big data transmission, and lack precise control over the offset and access order of each file segment in the file, which is easy to cause packet misalignment or paragraph loss. Second, when detecting and repairing packet errors, the traditional method lacks a fine offset marking mechanism, resulting in low efficiency of data error location and recovery. In addition, due to the poor compatibility of traditional methods in transmission, packets are prone to structural errors in various application environments. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: to provide a method and a terminal for transmitting big data files, which can make up for the deficiencies of traditional data transmission methods in big data compatibility and security guarantee through multi-level offset association and integrity verification of file segment access order.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A method for transmitting big data files, comprising the steps of: Obtain the big data files to be transmitted, and divide each file in the big data files based on the file sensitivity level to obtain a plurality of classified files; Extract the content of each of the classified files based on the content sensitivity level to obtain a set of independent sub-file clusters, and each set of independent sub-file clusters corresponds to a second classification label; Divide each basic file in each set of independent sub-file clusters into file segments, and determine the file segment access order of each basic file according to the position offset of the file segments in each basic file; Establish a first association between the first file segment of each basic file and the second classification label, establish a second association between each file segment in each basic file and the second classification label, and perform integrity verification on the transmission of the big data files according to the first association and the second association.

[0007] Another technical solution adopted by the present invention to solve the above technical problems is as follows: A big data file transmission terminal, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, each step of the above-mentioned big data file transmission method is implemented.

[0008] The beneficial effects of the present invention are as follows: Divide each file in the transmitted big data files based on the file sensitivity level to obtain a plurality of classified files; extract the content of each classified file based on the content sensitivity level to obtain a set of independent sub-file clusters corresponding to the second classification label; divide each basic file in each set of independent sub-file clusters into file segments, and determine the file segment access order of each basic file according to the position offset of the file segments in each basic file; establish a first association between the first file segment of each basic file and the second classification label, establish a second association between each file segment in each basic file and the second classification label, and perform integrity verification on the transmission of the big data files according to the first association and the second association. In this way, through multi-level offset association and integrity verification of the file segment access order, the deficiencies of the traditional method in terms of big data compatibility and security guarantee are made up, and a more efficient and flexible data transmission mechanism is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is a flowchart of a method for transmitting big data files according to an embodiment of the present invention; Figure 2 It is a schematic diagram of a big data file transmission terminal according to an embodiment of the present invention; Figure 3 Flow chart of the classification, segmentation and verification method for big data files according to an embodiment of the present invention; Figure 4 Flow chart of the classification processing flow for big data files according to an embodiment of the present invention; Figure 5 Flow chart of dividing the classified files into independent sub - file clusters according to an embodiment of the present invention; Figure 6 Flow chart of establishing the first association and the second association according to an embodiment of the present invention.

[0010] Label description: 1. A big data file transmission terminal; 2. A memory; 3. A processor. Detailed implementation manners

[0011] To describe in detail the technical content, achieved purpose and effects of the present invention, the following is described in combination with the implementation manners and with reference to the drawings.

[0012] Please refer to Figure 1 , an embodiment of the present invention provides a big data file transmission method, including the steps: Obtain the transmitted big data file, and divide each file in the big data file based on the file sensitivity level to obtain multiple classified files; Extract the content of each of the classified files based on the content sensitivity level to obtain a set of independent sub - file clusters, and each set of independent sub - file clusters corresponds to a second classification label; Divide each basic file in each set of independent sub - file clusters into file segments, and determine the file segment access order of each basic file according to the position offset of the file segments in each basic file; Establish a first association between the first file segment of each basic file and the second classification label, establish a second association between each file segment in each basic file and the second classification label, and perform integrity verification on the transmission of the big data file according to the first association and the second association.

[0013] As can be seen from the above description, the beneficial effects of the present invention are as follows: each file in the transmitted large data file is divided based on the file sensitivity level to obtain multiple classified files; each classified file is subjected to content extraction based on the content sensitivity level to obtain an independent sub-file cluster set corresponding to the second classification label; each basic file in each independent sub-file cluster set is divided into file segments, and the file segment access order of each basic file is determined according to the position offset of the file segments in each basic file; the first file segment of each basic file is associated with the second classification label, and each file segment in each basic file is associated with the second classification label, and the integrity verification of the transmission of the large data file is performed according to the first association and the second association. In this way, through the integrity verification of multi-level offset association and file segment access order, the deficiencies of the traditional method in terms of large data compatibility and security guarantee are made up, and a more efficient and flexible data transmission mechanism is provided.

[0014] Further, dividing each file in the large data file based on the file sensitivity level to obtain multiple classified files includes: Setting a corresponding first classification label according to the file sensitivity level; Identifying the file sensitivity level of the files in the large data file, and dividing the files with the same file sensitivity level in the large data file into the classified files corresponding to the first classification label.

[0015] As can be seen from the above description, by dividing the large data file content into different categories according to the sensitivity level and classifying the files into the corresponding classified files, the file sensitivity in each classified file is ensured to be consistent.

[0016] Further, performing content extraction on each of the classified files based on the content sensitivity level to obtain an independent sub-file cluster set includes: Setting a corresponding second classification label according to the content sensitivity level; Respectively extracting the content with the same content sensitivity level from each of the classified files as a file cluster, and obtaining an independent sub-file cluster set corresponding to the second classification label according to the file clusters extracted from each of the classified files.

[0017] As can be seen from the above description, in the process of dividing the classified files into an independent sub-file cluster set, a method for generating an independent sub-file cluster set based on the content sensitivity level is adopted to facilitate further improving the compatibility and security of large data transmission.

[0018] Further, dividing each basic file in each of the independent sub-file cluster sets into file segments includes: Divide the set of independent sub - file clusters evenly into a first number of consecutive storage sectors, where the first number is the number of classification files included in the set of independent sub - file clusters; Use the file clusters of the set of independent sub - file clusters as basic files, store the basic files in consecutive storage sectors in descending order, and perform file segment division on the basic files stored in each sector.

[0019] As can be seen from the above description, by evenly dividing the storage space of the set of independent sub - file clusters into multiple sectors, file segments can be further segmented within each sector, so as to facilitate more precise access control.

[0020] Furthermore, after performing file segment division on each basic file in each set of independent sub - file clusters, the following steps are also included: Form file groups by combining one or more consecutive file segments belonging to the same basic file in each sector.

[0021] As can be seen from the above description, by forming file groups for grouped transmission, not only the transmission risk of a single file segment is reduced, but also it is convenient to accurately locate the specific file segment and offset position in case of an exception. In this way, errors can be quickly responded to and the fault source can be located, improving the overall stability and maintainability of data transmission.

[0022] Furthermore, determining the file segment access order for each basic file according to the position offset of the file segments in each basic file includes: Establish the access order of file segments in each basic file in descending order according to the position offset of the file segments; Let the file segment access order of the i th basic file be W i , then the access order is defined as: W i ={ i 1, i 2, i 3,…, i q}, where i 1 to i q is the file segment sequence in the access order.

[0023] As can be seen from the above description, by dividing file segments and establishing the access order of file segments, reading and processing can be carried out in a specific order during transmission, thus ensuring the stability of transmission.

[0024] Furthermore, establishing a first association between the first file segment of each basic file and the second classification label includes: For the iA basic file, extract the first file segment in the access order of the file segments of the basic file, obtain the position offset of the first file segment in the storage location, use the position offset as the transfer file offset, and establish a first association between the transfer file offset and the second classification label.

[0025] As can be seen from the above description, establishing a first association according to the transfer file offset and the second classification label can be used for global transfer integrity verification at the file level to ensure that there are no errors or tampering between the classification label of the large data file and the first segment of the transferred file.

[0026] Further, establish a second association between each file segment in each basic file and the second classification label, including: Establish transfer file offset information of the basic file according to the access order of the file group, and the transfer file offset information includes the transfer file offset corresponding to each file group; Based on each of the transfer file offset information, establish a second association between the transfer file offset information and the second classification label, and generate a transfer data packet, where the transfer data packet includes a plurality of transfer file offset information and corresponding second classification labels.

[0027] As can be seen from the above description, establishing a second association according to the transfer file offset information and the second classification label can ensure that the order and position of each segment after the file is split are not tampered with, and are associated with the classification label, so that the file segments can be accessed and recombined in the correct order and position.

[0028] Further, establishing transfer file offset information of the basic file according to the access order of the file group includes: For the i th basic file, in the file segment access order, calculate and record the transfer file offset of each file group in sequence. Each file group corresponds to a file segment set, and based on the storage location offsets of the file segments in the file segment set, construct the transfer file offset information of the access order.

[0029] As can be seen from the above description, accessing each file segment in sequence according to the order of the position offsets can ensure data integrity during the transfer process.

[0030] Please refer to Figure 2 , another embodiment of the present invention provides a large data file transfer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements each step of the above-mentioned large data file transfer method.

[0031] The above-mentioned big data file transmission method and terminal of the present invention are applicable to making up for the deficiencies of traditional data transmission methods in terms of big data compatibility and security guarantee, which will be illustrated by examples below: Please refer to Figure 3 , Embodiment 1 of the present invention is as follows: Figure 3 is a flowchart showing the classification, segmentation, and verification methods of big data files according to an embodiment of the present invention. The method includes the following steps: S31. Obtain information on different file sensitivity levels of the big data file as the first classification label.

[0032] S32. Based on the first classification label, classify the big data file to obtain several classified files, and each classified file corresponds to a first classification label of one level.

[0033] S33. For each classified file, obtain several independent sub-file clusters of the classified file, and each independent sub-file cluster corresponds to a second classification label.

[0034] S34. For each set of independent sub-file clusters, divide each basic file in the set of independent sub-file clusters into several file segments according to the file size of each basic file stored in the set of independent sub-file clusters.

[0035] S35. Obtain several position offsets of each file segment according to the storage position occupied by each file segment during storage.

[0036] S36. Based on the position offsets of each file segment in each basic file, establish the file segment access order of the basic file; extract the first file segment of the basic file from the file segment access order as the first file segment, use the file segment offset of the first file segment as the transmission file offset information, and establish a first association between the transmission file offset information and the second classification label.

[0037] Specifically, the first association is to associate the offset of the file segment with the second classification label. The main purpose is to ensure the correct correspondence between the transmitted file and its classification label through the offset information of the first file segment of the file at the initial stage of transmission. Specifically, it ensures that the second classification label of the file is consistent with the position offset of its corresponding first file segment, and uses this information to establish an association between the offset and the second classification label. This association is usually used for global transmission integrity verification at the file level to ensure that there are no errors or tampering between the classification label of the big data file and the initial stage of transmission (the first segment of the file).

[0038] S37. Establish a second association between each file segment in each basic file and the second classification label.

[0039] Specifically, the second association is based on the offset information of each file segment. By extracting the offset of the file segment, it ensures that the order and position of each segment after the file is split are not tampered with and are associated with the second classification label. Its focus is on verifying the offset and managing the order at the file segment level to ensure that the file segments can be accessed and recombined in the correct order and position. Different from the global nature of the first association, the second association is at the file level to ensure that each segment can be transmitted and verified in the correct order after the file is split.

[0040] It can be seen that the first association focuses on file-level verification to ensure that the second classification label of the file matches the offset information of its first file segment. This process is used to verify the accuracy of the entire file's transmission path and classification mark and is usually carried out at the beginning stage of file transmission. The second association focuses on file segment-level verification to ensure that the order and position of each file segment after the file is split are not tampered with and the relationship between the offset and the second classification label can also be accurately matched. This process is usually executed in the middle or end stage of file transmission to ensure that each segment of the file is accessed and recombined in the correct order. Generally speaking, the first association ensures the consistency between the starting part of the file and the second classification label, and the second association ensures that each file segment after splitting is transmitted in the correct order and matches the second classification label. The two can ensure the integrity and security of data at different stages respectively.

[0041] S38. Based on the first association and the second association, perform integrity verification on the offset information of the transmitted file, where the offset information of the transmitted file corresponds to a second classification label.

[0042] In this embodiment, the purpose of integrity verification is to ensure that each file segment during the file transmission process is not tampered with, lost, or out of order, and each file segment can be correctly recombined according to the expected offset. The specific steps are as follows: (1) Before the transmission starts, according to the design of the first association and the second association, associate the offset information of the first file segment of each file and the position offset of each file segment in the entire file with the second classification label of the file.

[0043] (2) During the file transmission process, it is necessary to track the transmission status and offset information of each file segment. This means that whenever a file segment is transmitted, its actual transmission offset will be recorded and compared with the pre-set offset information. Each file segment during the transmission process should have a unique offset, and these offsets should be transmitted in the pre-determined order.

[0044] (3) At the beginning of file transfer, first verify whether the offset of the first file segment is consistent with the offset information of the transferred file. If they are consistent, it indicates that the file transfer starts from the correct starting point. Otherwise, it may indicate data tampering or disorder in the order, and the transfer process needs to be re-verified or paused. If the offset of the first file segment does not match, immediately stop the transfer, issue a warning, and re-verify and classify.

[0045] (4) During the transfer of each file segment, verify whether the offset of the current file segment is consistent with the overall order of the transferred file. Specifically, for each file segment, it is necessary to ensure the sequential transfer of file segments through its offset information. The offset of each file segment should conform to the preset order of each segment in the file. If the offset order of the file segment is incorrect, immediately mark an error and start the retransmission mechanism or perform error correction.

[0046] (5) After all file segments are transferred, according to the pre-established access order of file segments, perform a final verification on the offsets of all file segments to ensure that there are no deviations in the order and offset of all transferred file segments and they conform to the structure of the original file. This step is crucial for ensuring file integrity, and any offset mismatch needs to trigger an exception handling mechanism.

[0047] (6) After the file segments are transferred and pass the integrity verification, reassemble the file segments into a complete file based on the recorded offset information. The order, position, and content of the file segments will be restored according to the offset information to ensure that the reassembled file is exactly the same as the original file.

[0048] (7) The last step of integrity verification is to perform content verification on the reassembled file, usually by verifying the integrity of the file content through a hash value or digital signature. This step is to ensure that the file content during the transfer process has not been tampered with or damaged. If the file hash value is consistent with the hash value of the original file, it proves that the transferred file content is complete and correct.

[0049] (8) If an offset error or transfer exception is found in any step, immediately record an error log and trigger the corresponding error handling mechanism. The error handling mechanism includes measures such as requesting retransmission, reporting errors, and notifying the administrator.

[0050] Please refer to Figure 4 , Embodiment 2 of the present invention is as follows: Figure 4 is a flowchart showing the classification processing flow of big data files according to an embodiment of the present invention. To ensure the compatibility and security of big data files during the transfer process and further optimize the classification processing of big data files, in this embodiment, when classifying big data files based on classification labels, the classification strategy and processing flow are designed in detail. The method includes the following steps: S41. According to the different sensitive level information of the big data file, three first-level classification labels, namely public information, non-public information, and privacy information, are set, corresponding to different security requirements.

[0051] S42. Identify the sensitive levels of each part of the big data file.

[0052] S43. Divide the content with the same sensitive level into the corresponding classification label files, thus forming multiple classification files with the same sensitive level.

[0053] In this embodiment, the process of identifying the sensitive levels of each part of the big data file is achieved by layer-by-layer analysis, classification, and labeling of the file content. First, according to the preset sensitive level standard, use a data classification model or a rule engine to scan each part of the file to identify different types of information. Specifically, analyze the data type, field content, metadata, and file structure in the file, and judge the sensitivity of the file based on the nature of the data (such as personal information, business secrets, public data, etc.). This process can combine technical means such as natural language processing (NLP), pattern recognition, and keyword matching to extract sensitive elements from the text, numbers, or image data in the file, and determine the security level to which they belong according to these elements. Finally, all files will be marked and classified into different label files, forming multiple classification files with consistent sensitive levels, ensuring that appropriate security protection measures are taken for information with different sensitive levels during subsequent transmission. This process helps to improve the compatibility and security guarantee of big data files during transmission, ensuring that different levels of data can be processed and protected accordingly.

[0054] The division basis of the first-level classification label is the sensitivity analysis of the big data file content. Set a classification label function L ( x ) where x represents the file item in the big data file. L ( x ) outputs the sensitivity level of each file item. The formula is as follows: L ( x ) = {0, public information; 1, non-public information; 2, privacy information}.

[0055] Through this classification label function, the files included in the big data file are divided into different categories according to the sensitivity level. According to the L ( x ) output value, the file item xClassify them into corresponding classification files to ensure the consistency of the sensitivity of the content items in each classification file. It can be seen that the classification label corresponding to classification file F is determined based on the overall sensitivity level of the file, and is used to macroscopically distinguish the file sets at different security requirement levels. It can be regarded as the sensitivity level classification label at the file level, which determines the security policy framework for the overall transmission and processing of the file.

[0056] For example, assume that a big data file contains user data files, where some files are public user activity records, some files are restricted access internal transaction records, and some files are highly sensitive user personal identity information. Based on L ( x ), the public information is classified into a classification file F1, the non-public information is classified into F2, and the privacy information is classified into F3. The obtained F1, F2, and F3 each have their own security requirements. In the subsequent transmission process, different encryption and transmission protocols are adopted according to the sensitivity level corresponding to each classification file to ensure transmission compatibility and security.

[0057] Please refer to Figure 5 , Embodiment 3 of the present invention is: Figure 5 It is a flowchart showing the classification file divided into independent sub-file clusters according to an embodiment of the present invention.

[0058] To further improve the compatibility and security of big data transmission, in the process of dividing the classification file into independent sub-file clusters, an independent sub-file cluster generation method based on classification labels is adopted.

[0059] Specifically, by further subdividing the content of each classification file, the file content corresponding to each second classification label can form several independent sub-file clusters, thereby realizing finer-grained management of the content.

[0060] Let the content included in the j th classification file be F j , where F j contains the content of multiple second classification labels. For each classification file F j and the second classification labels therein, extract the content belonging to a certain second classification label from F j to form a sub-file cluster. For the p th second classification label, extract the content matching this label from F j as the p th sub-file cluster of the classification label, denoted as j th, denoted as Pp,j This process can be described by the following formula: P p,j = { y ∈ F j | L ( y ) = p} where, L ( y ) is the classification label function of the content item y. After such processing, the set of independent sub-file clusters under each second classification label p contains the content that conforms to this second classification label extracted from different classification files. In this way, a set of independent sub-file clusters { P p,1 , P p,2 , …, P p,L} is obtained, where L is the total number of classification files.

[0061] To ensure the clear division of files, the independent sub-file clusters of each second classification label are transmitted and managed separately. For example, assume there are 3 classification files F1, F2, and F3, which include the public information content label 0, the non-public information content label 2, and the privacy information content label 3. For the sub-file cluster of the public information label, the public information parts of the content can be extracted from F1, F2, and F3 respectively, and the obtained content sets are P 0,1 、 P 0,2 and P 0,3 , thus obtaining a complete set of independent sub-file clusters of public information for subsequent transmission and verification.

[0062] It can be seen that the classification label p of the set of independent sub-file clusters, although also related to the sensitive level, is further subdivided according to the content characteristics from the classification files at a finer content dimension, focusing on the internal content differences of the files, and accurately defining the security attributes and processing requirements of each sub-file cluster. This is the sensitive level classification label at the content level. The design of this structure ensures the security and independence during data transmission under different classification labels.

[0063] In this embodiment, to ensure the compatibility and security of big data transmission, by dividing each basic file in the set of independent sub-file clusters into file segments, the data transmission can be controlled and verified in a fine-grained manner. The core of this division method is based on each basic file P p,iThe file size is divided into multiple file segments, and the access order of the file segments is established so that they can be read and processed in a specific order during transmission, thereby ensuring the stability of the transmission.

[0064] The specific steps are as follows: Divide the set of independent sub-file clusters { P p,1 , P p,2 , …} into M consecutive storage sectors, where M is equal to the number of classification files contained in the independent sub-file clusters. By evenly dividing the storage space of the independent sub-file clusters into multiple sectors, file segments can be further segmented within each sector to achieve more precise access control.

[0065] In each sector, sort the files according to the file size of the basic files. Assume that each classification file F j contains several basic files, and the basic files are file clusters. Group and number them according to the file size of these basic files so that the allocation of the file group numbers conforms to the following sorting rules: F p,j,k ={ x ∈ S p,j | sin e ( x )∈ desc ( S p,j )} Among them, F p,j,k represents the p th file segment of the j th classification file of the k th classification label. sin e ( x ) represents the calculation function of the file size feature, and S p,j is the p th classification file corresponding to the j th classification label. The above formula indicates that the file segment F p,j,k is sorted in descending order according to the size of the classification file it belongs to.

[0066] Therefore, in this embodiment, file storage management is optimized through sector division and file sorting. In the storage space, first divide the independent sub-file clusters into M consecutive storage sectors, where M represents the number of classification files. The files in each sector are sorted according to the file size and numbered by the file group number to ensure that the arrangement of the files in each sector conforms to specific sorting rules.

[0067] Establishment of the access order of file segments: In each basic file, according to the storage location offset of the file segments, the access order of the file segments is established in the order of decreasing position offset. Let the access order of the file segments of the i th basic file be W i , then the access order is defined as: W i ={ i 1, i 2, i 3,…, i q}, where i from 1 to i q is the file segment sequence in the access order. According to the order of the position offsets, each file segment is accessed in turn, so as to ensure the data integrity during the transmission process.

[0068] Among them, first sort by the size of the sub-file clusters, divide them into storage sectors to form file segments, and then determine the file groups according to the number of file segments contained in each sub-file cluster and the association relationship. For example, if a sub-file cluster contains one or more consecutive file segments, these file segments constitute a file group.

[0069] Suppose there is an independent sub-file cluster, which contains files extracted from three classified files, namely public information (size 200MB), non-public information (150MB), and private information (100MB). First, the files extracted from these three sub-file clusters are divided into three storage sectors. Then, each sub-file cluster is sorted in descending order according to the size of the files, and each sub-file cluster is divided into several file segments. For example, the first sub-file cluster (200MB) contains two file groups, and each file group has two file segments. A file segment is the smallest unit when a file is stored, and each file segment contains part of the data, which is accessed and reorganized according to the position offset. In the first sub-file cluster, the two file groups are respectively marked as file group 1 and file group 2, and the access order of the file segments is sorted in descending order according to the offset to ensure that the data is read and verified in order during data transmission. Through this method, not only can the compatibility during the data transmission process be ensured, but also the consistency and security of the files among different classification labels and classified files can be guaranteed by using the access order of the file segments.

[0070] Please refer to Figure 6 , the fourth embodiment of the present invention is: Figure 6It is a flowchart showing the establishment of the first association and the second association according to an embodiment of the present invention. By extracting the first file segment of each basic file, transmission file offset information is established, and through the association between this information and the classification label, precise control of the offset information during the big data transmission process is achieved. This process includes the construction of the "first association" and the "second association" to ensure the generation of data packets and the identification of offset information, thereby improving the transmission compatibility and security. The specific steps are as follows: Extraction of the first file segment and generation of the offset: For the i ith basic file, according to the file segment access order, extract its first file segment as the first file segment of this basic file. Let the storage location offset of this file segment be Δ i , then this offset is defined as the transmission file offset, which is used to identify the starting position of this file, facilitating subsequent data positioning and access order control.

[0071] The formula is expressed as: Δ i = Offse t( F i,1 ), where F i,1 represents the first file segment of the ith basic file, and Offse t( F i,1 ) represents its offset in the storage space. Through this offset, the initial transmission position of the basic file can be identified.

[0072] Establish the first association between the transmission file offset Δ i and its corresponding second classification label p to identify the content category to which this offset belongs. This association relationship bundles the transmission file offset and the classification label together, which helps to accurately locate the transmission starting positions of different types of files during the subsequent data packet transmission process.

[0073] Transmission offset information for the file group access order: For the i ith basic file, in its file segment access order, calculate and record the transmission file offset of each file group in sequence. Each file group corresponds to a file segment set, and based on the storage location offsets of the file segments in this set, construct the transmission file offset information for the access order.

[0074] Define the file group offset Δ i,j as: Δ i,j = Offse t( F i,j ), where F i,j is the i ith basic file's jThe file segment offsets of each file group are sorted in descending order of offset to form a transmission access sequence of file segments.

[0075] After obtaining the offset information of the file segments, the offset information Δ of each file segment i,j is associated with its second classification label p to establish a second association. Based on this association, the classification label is used as the identifier of the transmission data packet, and a data packet containing the offset information of multiple transmission files is generated to correctly identify and schedule the starting position of the transmission in different classified files.

[0076] Generate transmission data packets: Finally, the offset information Δ of each file i,j is grouped according to the classification label to generate transmission data packets. The offset information of the transmission files contained in the data packets is used to quickly locate and retrieve file segments during the transmission process. Suppose the p transmission data packet generated under the i,1 th label contains the offset information set {Δ i,1 , Δ i,M}, and this data packet can ensure correct identification and scheduling among different classified files.

[0077] Suppose the access order of the file segments of the first basic file corresponding to a classification label "non-public information" is {10MB, 8MB, 5MB}, and its offset is Δ1 = 10MB. This offset is associated with the classification label "non-public information" to establish a first association. For the file group access order, assume that this file contains two file groups, and the offsets of each file group are 8MB and 5MB respectively. Then the data packet generated after establishing the second association will contain the offset information set {10MB, 8MB, 5MB}, which is convenient for identifying and accessing the order and offsets of file segments of the "non-public information" type during subsequent transmission.

[0078] Embodiment 5 of the present invention provides a specific application scenario: In practical applications, with the rapid development of big data, the sharp increase in data volume, and the data transmission requirements among multiple different systems, traditional data transmission methods often cannot effectively guarantee data integrity and transmission efficiency. To better address this challenge, this embodiment proposes an innovative data segmentation, sorting, and verification mechanism to optimize and ensure the security of data transmission from multiple dimensions such as file segments, file groups, and classification labels.

[0079] For example, in a typical big data processing application scenario, such as on a cloud computing platform, users need to upload, store, and transmit a large amount of file data. The file data may be a collection of data from different sources, involving different classification labels, such as financial data, customer information, product data, etc. The sizes and formats of these data files vary. How to ensure that these files can be processed in an efficient, compatible, and secure manner during data transmission is the key problem faced by this solution.

[0080] First, for the transmission of each file, it is processed by dividing the data file into multiple independent sub-file clusters. The specific operations include dividing each classified file into several sub-file clusters according to its classification label. These file clusters will be further divided according to file size, classification label, and other metadata, etc. In this process, the data in each file is sequentially divided into multiple "file segments", and subsequent processing is carried out based on the sizes and orders of these file segments.

[0081] For each file segment, calculate the offset of its storage location, and convert these offsets into offset information for the transmitted file. During data transmission, each file segment's position is tracked through this offset information to ensure that there is no position disorder during transmission. Especially during large-scale data transmission, this mechanism can ensure that each file segment reaches the target storage location accurately and without error.

[0082] Before data transmission, by establishing the access order of file segments, first extract the first file segment from the access order of file segments as the first file segment, and use its file segment offset as the offset information for the transmitted file for the first association. This process ensures that during file transmission, the first file segment can be preferentially processed and transmitted. Then, re-verify the order of file segments to ensure that the order of each file segment during transmission is consistent with the expected access order. If any mismatch occurs, a warning is issued and corresponding data repair is carried out.

[0083] During transmission, verify the integrity of each file. Specifically, when a file segment is transmitted over the network, check the position offset of each file segment. By comparing with the transmitted file offset, it can be determined whether the data has been damaged or lost during transmission. For example, if the transmission offset of a certain file segment does not match the expected offset, immediately mark that file segment as damaged or lost and initiate the retransmission mechanism.

[0084] In addition, to improve the transmission efficiency and security, the concept of file groups is introduced. In each file group, multiple file segments are grouped for transmission, which can reduce redundant data during transmission and improve the transmission efficiency. Moreover, the offset of each file group is associated with the classification label, making the data transmission more orderly and reliable.

[0085] Through this process, efficient file segmentation, transmission, and verification can be achieved, ensuring the correct transmission of each file segment in the big data environment and effectively preventing problems such as data loss, damage, and disorder. This method not only greatly improves the data transmission efficiency but also provides reliable guarantee in ensuring data security. Through real-time monitoring and offset verification, problems during the transmission process can be quickly responded to and corresponding repair measures can be taken, thus ensuring the integrity and compatibility of data transmission.

[0086] The method of this embodiment has important application value in the big data environment, especially in scenarios such as cloud computing, distributed storage, and other scenarios that require efficient and reliable data transmission.

[0087] Please refer to Figure 2 , Embodiment 6 of the present invention is: A big data file transmission terminal 1 includes a memory 2, a processor 3, and a computer program stored on the memory 2 and executable on the processor 3. When the processor 3 executes the computer program, each step of the above-mentioned big data file transmission method is implemented.

[0088] In summary, for the big data file transmission method and terminal provided by the present invention, each file in the transmitted big data file is divided based on the file sensitivity level to obtain multiple classified files corresponding to the first classification label; each classified file is subjected to content extraction based on the content sensitivity level to obtain an independent sub-file cluster set corresponding to the second classification label; each basic file in each independent sub-file cluster set is segmented, and the access order of the file segments of each basic file is determined according to the position offset of the file segments in each basic file; the first file segment of each basic file is associated with the second classification label to establish a first association, and each file segment in each basic file is associated with the second classification label to establish a second association, and the integrity verification of the transmission of the big data file is performed according to the first association and the second association.

[0089] In this way, through integrity verification and offset comparison, the present invention achieves the following beneficial effects: (1) During data transmission, by comparing the offsets in the storage location with those of the transmitted file, the consistency of data packets in the transmission link is ensured, preventing data segments from shifting or being damaged during transmission, and guaranteeing the integrity of the data. This enables the identification and prevention of data packet misalignment or loss even in complex transmission environments.

[0090] (2) The designed multi-level association and file segment offset comparison mechanism can effectively detect abnormal situations in transmitted data packets without having to compare the content of the entire file one by one, reducing the computational overhead required for verification. This offset-based verification method improves the verification speed while ensuring accuracy, and is particularly suitable for data transmission scenarios that require high throughput.

[0091] (3) By dividing the basic file into different file groups and transmitting them in groups, not only is the transmission risk of individual data packets reduced, but it also allows for precise positioning of specific file segments and offset positions in case of an anomaly. This enables a rapid response to errors and the location of the fault source, enhancing the overall stability and maintainability of data transmission.

[0092] (4) Through the flexible combination of transmitted file offset information, it supports different types and scales of data files and is easily compatible with existing data storage and transmission protocols. While ensuring data security, it can be implemented in a variety of application scenarios, such as file synchronization, cloud storage, or big data transmission.

[0093] (5) By using "first marker" and "second marker" to identify the status of data packets, this scheme marks when abnormal offsets in the data packets are detected. This early warning mechanism not only helps with real-time monitoring at the transport layer but also provides guidance for subsequent data recovery, reducing the response time for data verification and repair.

[0094] Generally speaking, the present invention not only enhances the security and reliability of data transmission but also optimizes the transmission efficiency, providing an efficient data verification and management solution for data-intensive applications. These effects can significantly reduce the risk of data loss or transmission errors in a variety of transmission scenarios, while also enhancing the fault tolerance and operation and maintenance efficiency.

[0095] The above are only embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent transformation made using the content of the specification and drawings of the present invention, or directly or indirectly applied in related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for transmitting large data files, characterized in that: Includes steps: Acquire the transmitted big data file, and classify each file in the big data file based on the file sensitivity level to obtain a plurality of classified files; Extracting content from each of the classified files based on the content sensitivity level to obtain an independent sub-file cluster set, each of the independent sub-file cluster sets corresponding to a second classification label; Divide each basic file in each of the independent sub-file cluster sets into file segments, and determine the file segment access order of each basic file according to the position offset of the file segment in each basic file; A first association is established between the first file segment of each basic file and the second classification tag, a second association is established between each file segment in each basic file and the second classification tag, and integrity verification is performed on the transmission of the large data file based on the first association and the second association.

2. A large data file transmission method according to claim 1, characterized in that: Each file in the big data file is divided based on the file sensitivity level to obtain a plurality of classified files, including: Setting a corresponding first classification label according to the file sensitivity level; The file sensitivity levels of the files in the big data file are identified, and the files in the big data file having the same file sensitivity level are classified into classification files corresponding to the first classification label.

3. A large data file transmission method according to claim 1, characterized in that: Extracting the contents of each of the classified files based on the content sensitivity level to obtain a set of independent sub-file clusters, including: Setting a corresponding second classification label according to the content sensitivity level; Contents with the same content sensitivity level are extracted from each of the classification files as file clusters, and a set of independent sub-file clusters corresponding to the second classification label is obtained according to the file clusters extracted from each of the classification files.

4. A large data file transmission method according to claim 3, characterized in that: Each basic file in each of the independent sub-file cluster sets is divided into file segments, including: Evenly dividing the independent sub-file cluster set into a first number of continuous storage sectors, where the first number is the number of classified files included in the independent sub-file cluster set; The file clusters of the independent sub-file cluster set are taken as basic files, the basic files are stored in a descending order in continuous storage sectors, and the basic files stored in each sector are divided into file segments.

5. A large data file transmission method according to claim 4, characterized in that: Each basic file in each of the independent sub-file cluster sets is divided into file segments, and then the following steps are further included: One or more consecutive file segments belonging to the same basic file in each sector are grouped into a file group.

6. A large data file transmission method according to claim 5, characterized in that: The file segment access order of each basic file is determined according to the position offset of the file segment in each basic file, including: In each basic file, the access order of the file segments is established in descending order according to the location offset of the file segments; Set up i The order of accessing the file segments of a basic file is: W i , then the access order is defined as: W i ={ i 1, i 2, i 3,…, i q },in i 1 to i q is a sequence of file segments in access order.

7. A large data file transmission method according to claim 6, characterized in that: Establishing a first association between the first file segment of each basic file and the second classification label includes: For i A basic file is extracted, a first file segment in the file segment access order of the basic file is extracted, a position offset of the first file segment in the storage position is obtained, the position offset is used as a transmission file offset, and a first association is established between the transmission file offset and the second classification label.

8. A large data file transmission method according to claim 6, characterized in that: Establishing a second association between each file segment in each basic file and the second classification label includes: Establishing transmission file offset information of the basic file according to the access order of the file group, wherein the transmission file offset information includes the transmission file offset corresponding to each file group; Based on each of the transmission file offset information, a second association between the transmission file offset information and the second classification label is established, and a transmission data packet is generated, wherein the transmission data packet includes a plurality of transmission file offset information and corresponding second classification labels.

9. A large data file transmission method according to claim 8, characterized in that: The transmission file offset information of the basic file is established according to the access sequence of the file group, including: For i A basic file is constructed, and in the file segment access order, the transmission file offset of each file group is calculated and recorded in turn, each file group corresponds to a file segment set, and based on the storage position offset of the file segment in the file segment set, the transmission file offset information of the access order is constructed.

10. A large data file transmission terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the processor implements the various steps of a large data file transmission method described in any one of claims 1 to 9.