A data interaction method and system based on metadata
The meta-data compression method addresses the inefficiencies of existing data compression by generating a meta-data repository file that further compresses data interactions, improving storage efficiency and transfer speed through the use of 7z algorithms.
Patent Information
- Application Number
- CN202211285466.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-10-20
AI Technical Summary
The prior art consumes a lot of computing resources when compressing big data files, and the compression ratio is limited, making it difficult to efficiently implement on personal computing devices, and the transmission efficiency is low.
The metadata-based compression method is adopted, and the data block is compressed into a 7z data folder block using the 7z algorithm, and a metadata warehouse file is generated. The metadata warehouse file is generated by traversing the number of times the metadata appears, thereby improving the compression rate.
Based on mainstream compression tools, the compression rate is further improved, the amount of downloaded data is reduced, the storage space utilization and transmission efficiency are improved, and the decompression speed is less affected.
Smart Images

Figure CN115454948B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing and data compression, and in particular to a metadata-based data interaction method and system. Background Art
[0002] With the improvement of the computing and storage performance of terminal hardware such as modern computers and mobile phones, audio and video stream data has gradually become one of the main sources of network data traffic. At the same time, the storage capacity of application data packets such as software installation packages, disk mappings, and Docker containers has also been gradually increasing. The solution of the present invention aims to more efficiently, quickly transmit, share, and publish large-volume files on the Internet.
[0003] At present, there are many mainstream compression algorithms and tools, and the compression ratio is relatively high. However, with the increase in entropy during calculation, the compression ratio decreases rapidly. In order to further improve the compression ratio, it is necessary to search from a larger data range. The compression process of the present invention requires a large amount of computing and disk space. It is very difficult to achieve these requirements in personal computing. It is necessary to utilize a large amount of giant computing resources on the cloud to achieve. After compression, the file data size is smaller. When downloading, combined with the BitTorrent sharing protocol, the data transmission efficiency can be greatly improved. Summary of the Invention
[0004] In order to solve the above technical problems existing in the prior art, the present invention proposes a metadata-based data interaction method and system to solve the above technical problems.
[0005] According to one aspect of the present invention, a metadata-based data interaction method is proposed, including:
[0006] S1: Generate a metadata warehouse file by using a metadata compression algorithm, where the metadata compression algorithm includes the following steps:
[0007] S11: Generate a virtual directory description file for each target data file;
[0008] S12: According to the extension of a single target file, call the corresponding file decomposition algorithm to decompose the target data file into several data blocks and add them to the virtual directory description file;
[0009] S13: Use the 7z algorithm to compress and convert the decomposed data blocks into 7z data folder blocks, and add 7z data folder description database nodes to the virtual directory description file;
[0010] S14: Generate a first-level data file, traverse and count the number of occurrences of metadata in the data blocks of each compressed data file, and sort to generate a metadata warehouse file;
[0011] S2: For the metadata warehouse files that have been downloaded, decompress the files based on the decompression algorithm, including:
[0012] S21: Restore the primary data file, call the metadata compression algorithm to decompress the last data block, and obtain the virtual directory description file;
[0013] S22: Based on the virtual directory description file and step S14 in the metadata compression algorithm, decompress the first data block to restore all file description data blocks;
[0014] S23: Based on the virtual directory description file and step S13 in the metadata compression algorithm, restore all data file directories to complete file decompression.
[0015] In some specific embodiments, the virtual directory description file is described in the XML file format, which is used to record the directory hierarchy of the target file, the basic information of the data file, and the file data block decomposition information.
[0016] In some specific embodiments, the data blocks include file description data blocks, original file data blocks, and 7z data folder blocks, and each data block is arranged in the storage order of the original file.
[0017] In some specific embodiments, S13 includes:
[0018] After merging all file description data blocks, call the 7z algorithm to compress and convert them into 7z data folder blocks, and add 7z data folder description database nodes to the virtual directory description file;
[0019] Call the 7z algorithm to compress and convert each original file data block into 7z data folder blocks respectively, and add 7z data folder description database nodes to the virtual directory description file.
[0020] In some specific embodiments, the metadata warehouse file is a 16GB metadata warehouse file, and the specific generation method includes:
[0021] Generate secondary data files based on the primary data files. The secondary data files are composed of a set of compressed record data, and the set of compressed record data is compressed after sorting the data blocks of the primary data files;
[0022] Merge each secondary data file one by one to generate a unique tertiary data file, sort the tertiary data file according to the repetition times, and extract the top 2 31 records with the most occurrences to generate a maximum 16GB metadata warehouse file.
[0023] In some specific embodiments, the metadata compression algorithm further includes step S15: using the metadata warehouse file as an index table, reconstructing the data blocks of all first-level data files according to the metadata size to generate a final data release file.
[0024] In some specific embodiments, before S14, there is further step S131: constructing a first-level compressed data file. The compressed data file includes file header description information, data blocks, and tail alignment data blocks. The file offset distance of the last data block is recorded in the file header description information. The last data block in the data blocks is a 7z data folder block obtained by compressing and converting the virtual directory description file by calling the 7z algorithm.
[0025] In some specific embodiments, the first-level data file is generated in the following manner: compressing all the template files in S12, S13, and S131 into a first-level data file.
[0026] In some specific embodiments, in S21, based on step S15, an inverse operation is performed to restore it to a first-level data file.
[0027] According to a second aspect of the present invention, there is provided a computer-readable storage medium having one or more computer programs stored thereon. When the one or more computer programs are executed by a computer processor, the method according to any one of the above is implemented.
[0028] According to a third aspect of the present invention, there is provided a metadata-based data interaction system, the system including:
[0029] A compression unit configured to generate a metadata warehouse file by using a metadata compression algorithm. The metadata compression algorithm includes the following steps: S11: generating a virtual directory description file for each target data file; S12: according to the extension of a single target file, calling a corresponding file decomposition algorithm to decompose the target data file into several data blocks and adding them to the virtual directory description file; S13: using the 7z algorithm to compress and convert the decomposed data blocks into 7z data folder blocks and adding 7z data folder description database nodes to the virtual directory description file; S14: generating a first-level data file, traversing and counting the number of occurrences of metadata in the data blocks of each compressed data file, and sorting to generate a metadata warehouse file;
[0030] Decompression Unit: Configured to decompress the metadata warehouse file that has been downloaded based on a decompression algorithm, including: S21: Restore the primary data file, call the metadata compression algorithm to decompress the last data block, and obtain the virtual directory description file; S22: Decompress the first data block based on the virtual directory description file and step S14 in the metadata compression algorithm to restore all file description data blocks; S23: Based on the virtual directory description file and step S13 in the metadata compression algorithm, restore all data file directories to complete file decompression.
[0031] In some specific embodiments, the virtual directory description file is described in XML file format, and is used to record the directory hierarchy of the target file, the basic information of the data file, and the file data block decomposition information; the data blocks include file description data blocks, original file data blocks, and 7z data folder blocks, and each data block is arranged in the storage order of the original file.
[0032] In some specific embodiments, S13 includes: After merging all file description data blocks, call the 7z algorithm to compress and convert them into 7z data folder blocks, and add 7z data folder description database nodes to the virtual directory description file; Call the 7z algorithm to compress and convert each original file data block into 7z data folder blocks respectively, and add 7z data folder description database nodes to the virtual directory description file.
[0033] In some specific embodiments, the metadata warehouse file is a 16GB metadata warehouse file, and the specific generation method includes: Generate a secondary data file based on the primary data file, the secondary data file consists of a compressed record data set, and the compressed record data set is compressed after sorting the data blocks of the primary data file; Merge each secondary data file one by one to generate a unique tertiary data file, sort the tertiary data file according to the repetition times, and extract the top 2 31 records with the most occurrences to generate a maximum 16GB metadata warehouse file.
[0034] In some specific embodiments, before S14, there is also step S131: Construct a primary compressed data file, the compressed data file includes file header description information, data blocks, and tail alignment data blocks, the file offset distance of the last data block is recorded in the file header description information, and the last data block in the data blocks is the 7z data folder block compressed and converted by the virtual directory description file using the 7z algorithm; The generation method of the primary data file in S14 is: Compress all the template files in S12, S13, and S131 into a primary data file; After S14, there is also S15: Use the metadata warehouse file as an index table to reconstruct the data blocks of all primary data files according to the metadata size to generate the final data release file.
[0035] The present invention provides a metadata-based data interaction method and system, which relates to the field of PB-level big data processing and data compression. In order for users to obtain target data from a metadata publishing website, they need to first download the corresponding metadata warehouse file and then download the target data. Through the algorithm of the present invention, the metadata publishing website further compresses the bytes of the interaction data through the generated metadata warehouse file on the basis of the compression ratio of current mainstream compression tools, thereby improving the storage space utilization rate of the website. Moreover, the decompression speed of the method of the present invention is less affected. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate the embodiments and together with the description are used to explain the principles of the invention. Other embodiments and many of the intended advantages of the embodiments will be readily apparent as they become better understood by reference to the following detailed description. The other features, objects, and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments read in conjunction with the accompanying drawings:
[0037] Figure 1 is a flowchart of a metadata-based data interaction method according to an embodiment of the present application;
[0038] Figure 2 is a file decompression flowchart of a specific embodiment of the present application;
[0039] Figure 3 is a framework diagram of a metadata-based data interaction system according to an embodiment of the present application;
[0040] Figure 4 is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the related invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0042] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0043] According to a metadata-based data interaction method of an embodiment of the present application, Figure 1 shows a flowchart of a metadata-based data interaction method according to an embodiment of the present application. As Figure 1As shown, the method includes:
[0044] S1: Generate a metadata warehouse file using a metadata compression algorithm, where the metadata compression algorithm includes the following steps:
[0045] S11: Generate a virtual directory description file for each target data file. The virtual directory description file is described in XML file format and is used to record the directory hierarchy of the target file, the basic information of the data file, and the file data block decomposition information.
[0046] S12: According to the extension of a single target file, call the corresponding file decomposition algorithm to decompose the target data file into several data blocks and add them to the virtual directory description file. Among them, the data blocks include file description data blocks, original file data blocks, and 7z data folder blocks, and each data block is arranged in the storage order of the original file.
[0047] S13: Use the 7z algorithm to compress and convert the decomposed data blocks into 7z data folder blocks, and add 7z data folder description database nodes to the virtual directory description file. In a specific embodiment, this step specifically includes: after merging all file description data blocks, call the 7z algorithm to compress and convert them into 7z data folder blocks, and add 7z data folder description database nodes to the virtual directory description file; call the 7z algorithm to compress and convert each original file data block into 7z data folder blocks respectively, and add 7z data folder description database nodes to the virtual directory description file.
[0048] S14: Generate a first-level data file, traverse and count the number of occurrences of metadata in the data blocks of each compressed data file, and sort to generate a metadata warehouse file.
[0049] In a specific embodiment, before S14, there is also step S131: Construct a first-level compressed data file. The compressed data file includes file header description information, data blocks, and tail alignment data blocks. The file header description information records the file offset distance of the last data block, and the last data block in the data blocks is the 7z data folder block compressed and converted by the virtual directory description file using the 7z algorithm. Among them, the generation method of the first-level data file is: compress all the template files in S12, S13, and S131 into a first-level data file.
[0050] In a specific embodiment, the metadata warehouse file is a 16GB metadata warehouse file, and the specific generation method includes:
[0051] Generate a second-level data file based on the first-level data file. The second-level data file is composed of a compressed record data set, and the compressed record data set is compressed after sorting the data blocks of the first-level data file.
[0052] Merge each secondary data file one by one to generate a unique tertiary data file, sort the tertiary data file by the number of repetitions, and extract the top 2 31 records to generate a metadata warehouse file with a maximum of 16GB.
[0053] In some specific embodiments, the metadata compression algorithm further includes step S15: using the metadata warehouse file as an index table to reconstruct the data blocks of all primary data files according to the metadata size to generate a final data release file.
[0054] S2: For the downloaded metadata warehouse file, decompress the file based on the decompression algorithm, including:
[0055] S21: Restore the primary data file, call the metadata compression algorithm to decompress the last data block to obtain a virtual directory description file. Among them, restoring the primary data file is restored based on the reverse operation of the aforementioned step S15.
[0056] S22: Based on the virtual directory description file and step S14 in the metadata compression algorithm, decompress the first data block to restore all file description data blocks;
[0057] S23: Based on the virtual directory description file and step S13 in the metadata compression algorithm, restore all data file directories to complete file decompression.
[0058] The present invention aims at a data publishing network resource site with rich data resources and compresses the PB-level data of the whole site according to metadata. The basic idea of the present invention is: call the 7z algorithm to compress each file of all data interactions one by one to generate several grouped 7z data Folder blocks, and then traverse all the compressed data, use 64 bits as the metadata size to generate 2 to the 31st power of the ordered metadata with the highest metadata utilization rate, and establish a database. Finally, re-index the target file according to the ordered metadata warehouse to obtain a cluster of release files with a compression rate between 30% and 48%.
[0059] For the data receiver in data interaction, before downloading data for the first time, it is necessary to first download the corresponding metadata warehouse and then download the compressed target file. The specific steps are as follows:
[0060] The first step is to generate a virtual directory description file. The virtual directory description file corresponding to the present invention is described in the XML file format, which is used to record the directory hierarchy of the target file, the basic information of the data file, and the file data block decomposition information.
[0061] In a specific embodiment, the nodes are divided into directory nodes, file nodes, data block nodes, and 7z data Folder description nodes. The attributes and contents of each specific node are as follows:
[0062] The directory node is a <D n="directory name"> file node.
[0063] The file node is a <F n="file name" ex="file type" sid="inverse operation algorithm ID" uuid="file unique encoding"> data block node.
[0064] Data block node <b t="数据块类型"len="块字节长">uuid of the 7z data Folder description node ,
[0065] The 7z data Folder description node <Z uuid="file unique encoding" arg="7z compression parameter" len="block byte length"> data block node uuid, data block node uuid,...
[0066] The data block types are divided into three types: file description data block, original file data block, and 7z data folder block. The content of the file description data block is mainly composed of text information. The original file data block is a block of random binary data bytes. The 7z data folder block is 7z standard folder format data that has been compressed using the 7z algorithm.
[0067] In the second step, according to its extension, the target file calls the corresponding file decomposition algorithm to decompose the target data file into several data blocks and add them to the virtual directory description file.
[0068] In a specific embodiment, each data block is arranged in the storage order of the original file. The decomposed data blocks are divided into three types: "file description" data blocks, original file data blocks, and 7z data folder blocks. For example, for pure text files such as xml, txt, docx, xlsx, http, etc., they are directly marked as "file description" data blocks. For 7z files, according to the 7z file format, they are decomposed into a header description file block, a tail folder description file block, and several 7z data folder blocks. For mp4 files, according to the mp4 file format, they are decomposed into several box data blocks, with the type marked as the original file data block. Among them, the ftyp and moov type box data blocks that play a descriptive role are registered as "file description" data blocks, and the other boxes are original file data blocks. And so on, a file decomposition algorithm for various data types is constructed.
[0069] In the third step, after all the file description data blocks are merged, the 7z algorithm is called to compress and convert them into 7z data folder blocks, and a 7z data Folder description database node is added to the virtual directory description file.
[0070] Step 4: Call the 7z algorithm for each original file data block, compress and convert it into a 7z data folder block, and add a 7z data folder description database node to the virtual directory description file.
[0071] Step 5: Construct a first-level compressed data file. The compressed data file consists of a file header description information, data blocks, and a tail alignment data block. The last data block in the data blocks is the 7z compressed folder information block of the virtual directory description file; the file header description information records the file offset distance of the last data block (preset 36 bytes) and can support PB-level files; the 7z compressed folder information block of the virtual directory description file is the virtual directory description file that calls the 7z algorithm, compresses and converts it into a 7z data folder block; the tail alignment data block is binary data according to the data block, and its length is determined by taking the remainder of the total number of bytes of the data block divided by 8. After the above steps 2 to 5, all template files are compressed into a first-level data file.
[0072] Step 6: Traverse and count the number of occurrences of metadata in the data blocks of each compressed data file, and sort to generate a 16GB metadata warehouse file. The specific algorithm is as follows:
[0073] Step 1: Generate a second-level data file based on the first-level data file. The second-level data file consists of a compressed record data set; the compressed record data set is compressed into the following format after sorting the data blocks of the first-level file:
[0074] {
[0075] {Metadata, repetition count (8 bytes, 128-bit statistics)} Each occupies 16 bytes
[0076] {Metadata, repetition count} ..
[0078] };
[0079] Step 2: Merge each second-level data file one by one to generate a unique third-level data file. The third-level data file consists of a compressed record data set; during the merging process, if the number of records in the data set of the third-level data file is approximately 2 to the 32nd power, discard the metadata record with the least repetition count;
[0080] Step 3: Sort the third-level data file according to the repetition count, extract the top 2 to the 31st power records at most, and ensure extraction to generate a maximum 16GB metadata warehouse file.
[0081] Step 7: Use the metadata warehouse file as an index table to reconstruct the data blocks of all first-level data files according to the metadata size to generate the final data release file.
[0082] In a specific embodiment, if the metadata appears in the data block, it is recorded as a metadata index; otherwise, it starts with 128 and is followed by 8 bytes of raw data:
[0083]
[0084] The decompression algorithm is adopted in the download and decompression process of the present invention. Figure 2 The file decompression flow chart of a specific embodiment of the present application is shown as Figure 2 shown, and the following steps are carried out:
[0085] In the first step, it is necessary to ensure that the metadata warehouse file has been downloaded.
[0086] In the second step, use the seventh step of the aforementioned compression algorithm to perform the reverse operation to restore it to the first-level data file.
[0087] In the third step, call the compression algorithm to decompress the last data block to obtain the virtual directory description file.
[0088] In the fourth step, according to the virtual directory description file and the fifth step in the compression algorithm, decompress the first data block to restore all file description data blocks.
[0089] In the fifth step, according to the virtual directory description file and the fourth step in the compression algorithm, restore all data files, that is, the directory, to complete the file decompression.
[0090] In a specific embodiment, the inventor of the present application conducted relevant tests on the aforementioned method: there are 2 independent target files, namely the 50GB A.7z file and the 34GB B.MP4 file. According to the aforementioned compression steps of the present invention, a total of three files are generated, namely the 16GB metadata warehouse file, the 25GB A.7z.y file, and the 15GB B.MP$.y file. If the user downloads the file A.7z.y for the first time, a total of 31GB of data needs to be downloaded, and the compression ratio is 62%. If the user has already downloaded the metadata warehouse file and then downloads the B file, a total of 15GB of data needs to be downloaded, and the compression ratio is 44%. If the user has already downloaded the metadata warehouse file and then downloads the A and B files, a total of 40GB of data needs to be downloaded, and the compression ratio is 54%. Based on the compression ratio of the current mainstream compression tools, the present invention further compresses the bytes of the interactive data through the generated metadata warehouse file, thereby improving the storage space utilization rate of the website, and the decompression speed of the method of the present invention is less affected.
[0091] Continue to refer to Figure 3 , Figure 3The framework diagram of a metadata-based data interaction system according to an embodiment of the present invention is shown. The system specifically includes a compression unit 301 and a decompression unit 302. The compression unit 301 is configured to generate a metadata warehouse file by using a metadata compression algorithm. The metadata compression algorithm includes the following steps: S11: Generate a virtual directory description file for each target data file; S12: According to the extension of a single target file, call the corresponding file decomposition algorithm to decompose the target data file into several data blocks and add them to the virtual directory description file; S13: Use the 7z algorithm to compress and convert the decomposed data blocks into 7z data folder blocks, and add a 7z data folder description database node to the virtual directory description file; S14: Generate a first-level data file, traverse and count the number of times metadata appears in the data blocks of each compressed data file, and sort to generate a metadata warehouse file; The decompression unit 302 is configured to, for the downloaded metadata warehouse file, decompress the file based on a decompression algorithm, including: S21: Restore the first-level data file, call the metadata compression algorithm to decompress the last data block, and obtain the virtual directory description file; S22: Decompress the first data block based on the virtual directory description file and step S14 in the metadata compression algorithm to restore all file description data blocks; S23: Based on the virtual directory description file and step S13 in the metadata compression algorithm, restore all data file directories to complete file decompression.
[0092] Reference is made below to Figure 4 , which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Figure 4 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0093] As Figure 4 shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0094] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including, for example, a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 410 as needed so that a computer program read therefrom is installed into the storage section 408 as needed.
[0095] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the methods of the present application are performed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, and the computer-readable storage medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0096] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0098] The modules described in the embodiments of this application can be implemented in software or in hardware.
[0099] As another aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium may be included in the electronic device described in the above embodiments; or it may exist alone without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: generate a metadata warehouse file by using a metadata compression algorithm, wherein the metadata compression algorithm includes the following steps: generating a metadata warehouse file by using a metadata compression algorithm, including: generating a virtual directory description file for each target data file; decomposing the target data file into several data blocks and adding them to the description file; compressing and converting the decomposed data blocks into 7z data folder blocks and adding 7z data folder description database nodes to the description file; generating a first-level data file, traversing and counting the occurrence times of metadata in the data blocks of the compressed data file, and sorting to generate a metadata warehouse file; for the metadata warehouse file that has been downloaded and completed, decompressing the file based on a decompression algorithm, including: restoring the first-level data file, decompressing the last data block to obtain a virtual directory description file; decompressing the first data block and restoring all file description data blocks; restoring all data file directories to complete file decompression.
[0100] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present application.
Claims
1. A data interaction method based on metadata, characterized in that, Including: S1: Generate a metadata warehouse file using a metadata compression algorithm, where the metadata compression algorithm includes the following steps: S11: Generate a virtual directory description file for each target data file; S12: According to the extension of a single target file, call the corresponding file decomposition algorithm to decompose the target data file into several data blocks and add them to the virtual directory description file; S13: Use the 7z algorithm to compress and convert the decomposed data blocks into 7z data folder blocks, and add 7z data folder description database nodes to the virtual directory description file; S14: Generate a first-level data file, traverse and count the number of occurrences of metadata in the data blocks of each compressed data file, and sort to generate a metadata warehouse file; S2: For the downloaded metadata warehouse file, decompress the file based on the decompression algorithm, including: S21: Restore the first-level data file, call the metadata compression algorithm to decompress the last data block, and obtain the virtual directory description file; S22: Based on the virtual directory description file and step S14 in the metadata compression algorithm, decompress the first data block to restore all file description data blocks; S23: Based on the virtual directory description file and step S13 in the metadata compression algorithm, restore all data file directories to complete file decompression; The metadata warehouse file is a 16GB metadata warehouse file, and the specific generation method includes: Generate a second-level data file according to the first-level data file. The second-level data file consists of a compressed record data set, and the compressed record data set is compressed after sorting the data blocks of the first-level data file; Merge each of the secondary data files one by one to generate a unique tertiary data file, sort the tertiary data file by the number of repetitions, and extract the top 2 31 records to generate a metadata warehouse file with a maximum of 16GB.
2. The data interaction method based on metadata according to claim 1, wherein The virtual directory description file is described in the XML file format, and is used to record the directory hierarchy of the target file, the basic information of the data file, and the file data block decomposition information.
3. The data interaction method based on metadata according to claim 1, wherein The data blocks include file description data blocks, original file data blocks, and 7z data folder blocks, and each data block is arranged in the storage order of the original file.
4. The data interaction method based on metadata according to claim 3, wherein The S13 includes: After merging all the file description data blocks, call the 7z algorithm to compress and convert them into 7z data folder blocks, and add 7z data folder description database nodes to the virtual directory description file; Call the 7z algorithm to compress and convert each original file data block into 7z data folder blocks respectively, and add 7z data folder description database nodes to the virtual directory description file.
5. The data interaction method based on metadata according to claim 1, wherein The metadata compression algorithm further includes step S15: Use the metadata warehouse file as an index table to reconstruct the data blocks of all the first-level data files according to the metadata size to generate a final data release file.
6. The data interaction method based on metadata according to claim 1, wherein Before S14, there is also step S131: constructing a first-level compressed data file, where the compressed data file includes file header description information, data blocks, and a tail alignment data block. The file offset distance of the last data block is recorded in the file header description information, and the last data block in the data blocks is the 7z data folder block obtained by compressing and converting the virtual directory description file using the 7z algorithm.
7. The data interaction method based on metadata according to claim 6, characterized in that, The first-level data file is generated in the following way: Compressing all the template files in S12, S13, and S131 into a first-level data file.
8. The data interaction method based on metadata according to claim 5, characterized in that In S21, based on step S15, perform the inverse operation to restore to the first-level data file.
9. A computer-readable storage medium having one or more computer programs stored thereon, characterized in that, When the one or more computer programs are executed by a computer processor, they implement the method according to any one of claims 1 to 8.
10. A metadata-based data interaction system, characterized in that, The system includes: A compression unit configured to generate a metadata warehouse file using a metadata compression algorithm. The metadata compression algorithm includes the following steps: S11: Generating a virtual directory description file for each target data file; S12: According to the extension of a single target file, calling the corresponding file decomposition algorithm to decompose the target data file into several data blocks and adding them to the virtual directory description file; S13: Using the 7z algorithm to compress and convert the decomposed data blocks into 7z data folder blocks and adding 7z data folder description database nodes to the virtual directory description file; S14: Generating a first-level data file, traversing and counting the occurrence times of metadata in the data blocks of each compressed data file, and sorting to generate a metadata warehouse file. A decompression unit configured to, for the downloaded metadata warehouse file, decompress the file based on a decompression algorithm, including: S21: Restoring the first-level data file, calling the metadata compression algorithm to decompress the last data block to obtain the virtual directory description file; S22: Based on the virtual directory description file and step S14 in the metadata compression algorithm, decompressing the first data block to restore all file description data blocks; S23: Based on the virtual directory description file and step S13 in the metadata compression algorithm, restoring all data file directories to complete file decompression. The metadata warehouse file is a 16GB metadata warehouse file. The specific generation method includes: Generating a second-level data file according to the first-level data file. The second-level data file is composed of a compressed record data set, and the compressed record data set is compressed after sorting the data blocks of the first-level data file; Merging each of the second-level data files one by one to generate a unique third-level data file, sorting the third-level data file according to the repetition times, and extracting the first 231 records with the most occurrences to generate a maximum 16GB metadata warehouse file.
11. The data interaction system based on metadata according to claim 10, wherein The virtual directory description file is described in the XML file format and is used to record the directory hierarchy of the target file, the basic information of the data file, and the file data block decomposition information; The data blocks include file description data blocks, original file data blocks, and 7z data folder blocks, and each data block is arranged in the storage order of the original file.
12. The data interaction system based on metadata according to claim 11, wherein The S13 includes: after merging all the file description data blocks, calling the 7z algorithm to compress and convert them into 7z data folder blocks, and adding 7z data folder description database nodes to the virtual directory description file; calling the 7z algorithm to compress and convert each of the original file data blocks into 7z data folder blocks, and adding 7z data folder description database nodes to the virtual directory description file.
13. The data interaction system based on metadata according to claim 10, wherein Before the S14, there is also a step S131: constructing a first-level compressed data file, which includes file header description information, data blocks, and tail alignment data blocks. The file offset distance of the last data block is recorded in the file header description information. The last data block in the data blocks is the 7z data folder block obtained by compressing and converting the virtual directory description file using the 7z algorithm. The generation method of the first-level data file in the S14 is: compressing all the template files in S12, S13, and S131 into a first-level data file. After the S14, there is also an S15: using the metadata warehouse file as an index table to reconstruct the data blocks of all the first-level data files according to the metadata size to generate a final data release file.
14. The data interaction system based on metadata according to claim 13, wherein In the S21, based on the step S15, an inverse operation is performed to restore it to the first-level data file.
Citation Information
Patent Citations
Hybrid-cloud data storage method, hybrid-cloud data storage device, corresponding equipment and cloud system
CN107330337A
Backup file recovery method for heterogeneous database
CN107391306A