A method for optimizing the storage of massive file data
By analyzing the differences in data blocks, obtaining redundant data blocks, calculating data redundancy and similarity, building a storage priority model, solving the problem of the impact of redundant data in massive file data storage, and achieving efficient storage optimization.
Patent Information
- Application Number
- CN202411953875.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-27
AI Technical Summary
The prior art in the storage of massive file data, due to the existence of redundant data, the storage priority classification errors are caused, storage costs are increased, and storage efficiency is reduced.
By analyzing the differences between data blocks, the redundant data blocks in each file are obtained, and the data redundancy and redundancy similarity are determined based on the proportion of redundant data blocks, a computing model of deduplication and storage priority is built, and the storage strategy is optimized.
Effectively identify and remove redundant data, improve data storage efficiency, optimize resource allocation, eliminate the impact of redundant data on storage priority, and improve the storage efficiency of massive file data.
Smart Images

Figure CN119883119B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of optimized storage of massive file data, and particularly to an optimized storage method for massive file data. Background Art
[0002] When facing massive file data, how to optimize storage is a very important task, especially in scenarios such as big data, cloud computing, and high-concurrency access. An effective optimized storage method can help enterprises reduce storage costs, improve data access speed, and optimize the overall performance of the system. Determining the storage priority of data is a key step in optimizing data storage, improving access efficiency, and reducing costs. Therefore, it is necessary to classify and grade the data, and set different storage strategies according to factors such as data redundancy characteristics.
[0003] Currently, for the storage strategy of massive data, different priorities are generally divided based on the functional characteristics of the data and stored in the order of priority. However, due to the relatively large scale of the data, the redundancy degree of the data is relatively greater. The existence of redundant data will affect the division of storage priority, thereby leading to errors in the priority division strategy based on the functional characteristics of the data. At the same time, the existence of redundant data will increase storage costs and reduce the storage efficiency of massive file data. Summary of the Invention
[0004] In order to solve the above technical problems, this application provides an optimized storage method for massive file data to solve the existing problems.
[0005] An optimized storage method for massive file data of this application adopts the following technical solutions:
[0006] An embodiment of this application provides an optimized storage method for massive file data, and this method includes the following steps:
[0007] Obtain all classes of files to be stored and their respective access frequencies, and allocate multiple data blocks for each file in each class of data to be stored for storage;
[0008] Within each class of files to be stored, based on the differences between any two data blocks in each file, and the differences between any two data blocks between each file and the remaining files, obtain redundant data blocks of various sizes; based on the distribution of all redundant data blocks in each file, determine the data redundancy degree of each file;
[0009] Based on the proportion of redundant data blocks of the same size between each file and the rest of the files, determine the redundancy index between each file and the rest of the files within each type of file to be stored, and in combination with the data redundancy degree, determine the redundancy similarity between each file and the rest of the files within each type of file to be stored, so as to determine the redundancy coefficient of each type of file to be stored; based on the proportion of the remaining data blocks except the redundant data blocks in each file, and the redundancy coefficient, determine the deduplication degree of each file within each type of file to be stored;
[0010] Based on the access frequency of each type of file to be stored and the deduplication degree, determine the priority storage coefficient of each file within each type of file to be stored, and in combination with the number of all data blocks in each type of file to be stored, determine the storage priority of each type of file to be stored, and store all types of files to be stored.
[0011] Preferably, the method for obtaining redundant data blocks of various sizes in each file is as follows:
[0012] The data form in each data block is a binary number. Among all the files of each type of file to be stored, perform a bitwise exclusive OR operation on the binary numbers between any two data blocks of the same size. Mark any one of the two data blocks whose binary result of the operation has all bits as 0 as a redundant data block. Traverse all data blocks to obtain all the redundant data blocks in each type of file to be stored, where the already marked redundant data blocks are not marked repeatedly, and count the redundant data blocks of various sizes in each file.
[0013] Preferably, the data redundancy degree of each file is the proportion of all redundant data blocks in each file among all data blocks.
[0014] Preferably, the method for determining the redundancy index between each file and the rest of the files within each type of file to be stored is as follows:
[0015] Within each type of file to be stored, use the ratio of the number of types of redundant data blocks of the same size between each file and the rest of the files to the total number of types of all different-sized redundant data blocks as the redundancy index between each file and the rest of the files within each type of file to be stored.
[0016] Preferably, the expression for the redundancy similarity between each file and the rest of the files within each type of file to be stored is: In the formula, represents the redundancy similarity between file i and file j within the kth type of file to be stored; represents the redundancy index between file i and file j within the kth type of file to be stored; β i 、β jrespectively represent the data redundancy of file i and file j; norm() represents the normalization function; ε represents a preset constant greater than 0.
[0017] Preferably, the method for determining the redundancy coefficient of each type of file to be stored is as follows:
[0018] Calculate the mean of the redundancy similarity between each file in each type of file to be stored and all the other files, and use the average level of the means of the redundancy similarities of all files as the redundancy coefficient of each type of file to be stored.
[0019] Preferably, the method for determining the deduplication degree of each file in each type of file to be stored is as follows:
[0020] Calculate the proportion of the remaining data blocks except the redundant data blocks in each file in each type of file to be stored in all the data blocks, and denote it as the non-repetitive proportion of each file in each type of file to be stored;
[0021] The deduplication degree of each file in each type of file to be stored is the ratio of the redundancy coefficient of each type of file to be stored to the non-repetitive proportion of the corresponding file in each type of file to be stored.
[0022] Preferably, the expression of the priority storage coefficient of each file in each type of file to be stored is: In the formula, γ k,i represents the priority storage coefficient of file i in the kth type of file to be stored; β k represents the access frequency of the kth type of file to be stored; ω k,i represents the deduplication degree of file i in the kth type of file to be stored; norm() represents the normalization function; σ represents a preset constant greater than 0.
[0023] Preferably, the expression of the storage priority of each type of file to be stored is: In the formula, U k represents the storage priority of the kth type of file to be stored; represents the mean of the priority storage coefficients of all files in the kth type of file to be stored; C k represents the number of all data blocks in the kth type of file to be stored.
[0024] Preferably, the storage of all types of files to be stored includes:
[0025] Sort each type of file to be stored in descending order of storage priority and store it in the storage device.
[0026] This application has at least the following beneficial effects:
[0027] This application analyzes the differences between data blocks to obtain redundant data blocks in each file, and constructs a data redundancy degree based on the proportion of all redundant data blocks in each file. The beneficial effect is that by calculating the data redundancy degree, the repeated parts in the data can be identified, so as to remove redundant data and improve the storage efficiency of the data. This application constructs a redundancy similarity by analyzing the redundancy degree between different files. The beneficial effect is that it can more accurately evaluate the redundancy degree between data, so as to formulate a more reasonable storage strategy to ensure that key data is stored preferentially and improve the data storage efficiency. This application constructs a deduplication degree based on the redundancy similarity and the proportion of the remaining data blocks outside the redundant data blocks. The beneficial effect is that it helps to optimize resource allocation and improve the data storage efficiency. This application constructs a storage priority by comprehensively considering the access frequency, the deduplication degree, and the number of all data blocks in various files to be stored, and stores all types of files to be stored. The beneficial effect is that it can eliminate the influence of redundant data on the storage priority and improve the data storage efficiency. This application improves the storage efficiency of massive file data by analyzing the influence of redundant data on the storage priority. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] Figure 1 It is a flowchart of the steps of a method for optimizing the storage of massive file data provided by an embodiment of the present application;
[0030] Figure 2 It is a schematic diagram of the process of extracting the deduplication degree provided by an embodiment of the present application;
[0031] Figure 3 It is a schematic diagram of the process of obtaining the storage priority provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] In order to further elaborate on the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific implementation manner, structure, features and effects of a method for optimizing the storage of massive file data proposed according to the present application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0034] The following specifically describes the specific solution of a method for optimizing the storage of massive file data provided by this application in conjunction with the accompanying drawings.
[0035] A method for optimizing the storage of massive file data provided by an embodiment of this application. Specifically, the following is a method for optimizing the storage of massive file data. Please refer to Figure 1 , the method includes the following steps:
[0036] Step S1: Obtain all types of files to be stored and their respective access frequencies, and allocate multiple data blocks for each file in each type of data to be stored for storage.
[0037] With the popularization and application of technologies such as the Internet, Internet of Things, and social media, the generation of massive file data has grown exponentially. Massive file data includes various types such as text files, images, videos, and audio. Currently, different types of data are generally stored in the form of binary strings in storage devices, but large storage systems generally distinguish storage spaces, and the smallest unit of distinction is a data block. Therefore, in order to facilitate the storage of large-scale data, data blocks need to be allocated for large-scale data.
[0038] Therefore, obtain all types of files to be stored and the access frequencies of each type of file to be stored, and allocate multiple data blocks for each file in each type of data to be stored for storage. Among them, the data in the data block is in the form of binary numbers.
[0039] It should be noted that all types of files to be stored refer to files in different formats. In this embodiment, all types of files to be stored include text files, image files, video files, and audio files. Among them, the access frequency of each type of file to be stored refers to the overall access frequency of that type of file to be stored. The process of obtaining the access frequency is a well-known technology, and its specific acquisition process will not be elaborated here.
[0040] Step S2: Within each type of file to be stored, based on the differences between any two data blocks in each file, and the differences between any two data blocks between each file and the rest of the files, obtain redundant data blocks of various sizes in each file; based on the distribution of all redundant data blocks in each file, determine the data redundancy of each file.
[0041] When storing each type of file to be stored, generally, the storage priority of the corresponding type of file is determined by analyzing the functional characteristics of the data contained in the file, such as access frequency, data importance, etc. For large-scale data with redundant characteristics, if the storage priority order is determined only based on the functional characteristics of the data, the optimal storage efficiency cannot be achieved.
[0042] Therefore, by analyzing the redundant features of large-scale data, eliminating the redundant data in the data, and excluding the influence of redundant data on the data storage priority, the data storage efficiency can be improved. Specifically:
[0043] (1) In various files to be stored, based on the differences between any two data blocks in each file and the differences between any two data blocks in each file and the rest of the files, redundant data blocks of various sizes in each file are obtained.
[0044] Internal redundancy is a common phenomenon in data storage. It is not only related to the characteristics of file formats, but also involves multiple aspects such as data backup and data integration. For example, in a text file, some text content may be copied and pasted to different places; in a video file, there may be multiple identical frames; in an audio file, to support playback at different bitrates, multiple audio streams with the same content may be stored; these phenomena are all cases of internal data redundancy. The redundant data occupies additional storage space. Since the storage resources are occupied by unnecessary data, the priority strategy based on functional characteristics may not work as expected. In addition, due to the fact that users may save multiple copies of the same data in different locations, or duplicate data copies may be generated when sharing data between different business departments or applications, resulting in external data redundancy. The existence of external redundancy also has a certain impact on the storage priority.
[0045] Therefore, in order to screen out the redundant data in the file to exclude the interference of redundant data on the storage priority, by analyzing the differences between any two data blocks in each file in various files to be stored and the differences between any two data blocks in each file and the rest of the files, redundant data blocks of various sizes in each file are obtained. Specifically:
[0046] Since the data in each data block is in the form of binary numbers, in all files of various files to be stored, perform a bitwise exclusive OR operation on the binary numbers between any two data blocks of the same size. Mark any one of the two data blocks with all bits being 0 in the binary result of the operation as a redundant data block. Traverse all data blocks to obtain all redundant data blocks in various files to be stored. For the already marked redundant data blocks, do not make repeated markings, and count the redundant data blocks of various sizes in each file.
[0047] It should be noted that since the sizes of data blocks are different, the lengths of the binary numbers they store are also different. Data blocks of different sizes must not be redundant data blocks. The redundant data blocks here refer to data blocks with the same length of the binary numbers stored in the data blocks, otherwise the bitwise exclusive OR operation cannot be performed.
[0048] Among them, the calculation steps of the bitwise exclusive OR operation are well-known technologies, and the specific calculation process will not be elaborated here.
[0049] (2) Based on the distribution of all redundant data blocks in each file, determine the data redundancy of each file.
[0050] If there are more redundant data blocks in a file, it means that the data redundancy of the file is greater, and it is more necessary to eliminate the redundant data in the file. Therefore, by analyzing the proportion of redundant data blocks in the file, the data redundancy is determined, so as to judge the redundancy situation of the file. Specifically:
[0051] The data redundancy of each file is the proportion of all redundant data blocks in each file among all data blocks.
[0052] According to the data redundancy of each file, it can be understood that the more redundant data blocks in the file, the greater the proportion of all redundant data blocks in the file among all data blocks, that is, the greater the data redundancy of the file; on the contrary, the fewer redundant data blocks in the file, the smaller the proportion of all redundant data blocks in the file among all data blocks, that is, the smaller the data redundancy of the file.
[0053] Step S3: Based on the proportion of redundant data blocks of the same size between each file and the rest of the files, determine the redundancy index between each file and the rest of the files within each type of file to be stored, and combine the data redundancy to determine the redundancy similarity between each file and the rest of the files within each type of file to be stored, so as to determine the redundancy coefficient of each type of file to be stored; based on the proportion of the remaining data blocks except the redundant data blocks in each file, and the redundancy coefficient, determine the deduplication degree of each file within each type of file to be stored.
[0054] During the data migration process, data copies of the old system may be temporarily or permanently retained to ensure a smooth transition of the migration, resulting in external data redundancy. When sharing data between different applications, duplicate data copies may be generated, which will also lead to external data redundancy. For large-scale data with redundancy characteristics, if the storage order is only determined based on functional characteristics, redundant data may be wrongly given a higher storage priority, resulting in unreasonable allocation of storage resources. Therefore, only determining the priority storage order based on the functional characteristics of the data will not achieve the optimal storage efficiency.
[0055] Therefore, by analyzing the redundancy similarity between files, the interference of redundant data between files is excluded, so as to improve the storage efficiency of massive file data. Specifically:
[0056] (1) Based on the proportion of redundant data blocks of the same size between each file and the rest of the files, determine the redundancy index between each file and the rest of the files within each type of file to be stored, which is used to characterize the redundancy degree between files. Specifically:
[0057] Within each type of file to be stored, take the ratio of the number of types of redundant data blocks of the same size between each file and the rest of the files to the total number of types of all different-sized redundant data blocks as the redundancy index between each file and the rest of the files within each type of file to be stored.
[0058] From the redundancy index between each file and the rest of the files within each type of file to be stored, it can be understood that if the number of types of redundant data blocks of the same size between files is larger, the redundancy index is larger, indicating a greater redundancy degree between files; conversely, if the number of types of redundant data blocks of the same size between files is smaller, the redundancy index is smaller, indicating a smaller redundancy degree between files.
[0059] (2) Further, based on the proportion of redundant data blocks of the same size between each file and the rest of the files, and in combination with the data redundancy degree, determine the redundancy similarity between each file and the rest of the files within each type of file to be stored. Specifically:
[0060] The redundancy similarity between file i and file j within the kth type of file to be stored The expression of is: In the formula, represents the redundancy similarity between file i and file j within the kth type of file to be stored; represents the redundancy index between file i and file j within the kth type of file to be stored; β i 、β j respectively represent the data redundancy degrees of file i and file j; norm() represents the normalization function; ε represents a preset constant greater than 0, which is used to prevent the denominator from being 0. The value of ε is set artificially. In this embodiment, the value of ε is 0.01. On the premise of ensuring that the denominator is not 0 and does not overly affect the calculation result, the implementer can set it according to the specific situation by himself, and this embodiment does not make special restrictions.
[0061] According to the redundancy similarity between file i and file j within the kth type of file to be stored It can be understood that the larger the ratio of the number of types of redundant data blocks of the same size between file i and file j within the kth type of file to be stored to the total number of types of all different sizes in file i and file j, that is The larger it is, the more redundant data blocks of the same size files i and j have, the greater the redundancy similarity between files i and j, and the smaller the difference in data redundancy between files i and j, the greater the redundancy similarity; conversely, the smaller the ratio of the number of types of redundant data blocks of the same size between files i and j in the k-th type of files to be stored to the total number of all different sizes in files i and j, that is The smaller it is, the fewer redundant data blocks of the same size files i and j have, the smaller the redundancy similarity between files i and j, and the greater the difference in data redundancy between files i and j, the smaller the redundancy similarity.
[0062] (3) Further, based on the redundancy similarity between each file and the rest of the files in each type of files to be stored, to determine the redundancy coefficient of each type of files to be stored, specifically:
[0063] Calculate the mean of the redundancy similarities between each file and all the other files in each type of files to be stored, and take the average level of the means of the redundancy similarities of all files as the redundancy coefficient of each type of files to be stored.
[0064] It can be understood from the redundancy coefficients of each type of files to be stored that if the redundancy similarity between the files in the files to be stored is greater, that is, the mean of the redundancy similarities between all files is greater, then the redundancy coefficient of the files to be stored is greater, and the storage priority level of the data is smaller, and it is more necessary to perform deduplication operations on the files; conversely, if the redundancy similarity between the files in the files to be stored is smaller, that is, the mean of the redundancy similarities between all files is smaller, then the redundancy coefficient of the files to be stored is smaller, and the storage priority level of the data is greater.
[0065] (4) Based on the redundancy coefficients of each type of files to be stored, and combined with the proportion of the remaining data blocks except the redundant data blocks in each file, to determine the deduplication degree of each file in each type of files to be stored, specifically:
[0066] Calculate the proportion of the remaining data blocks except the redundant data blocks in each file in all data blocks in each type of files to be stored, denoted as the non-repetitive proportion of each file in each type of files to be stored;
[0067] The deduplication degree of each file in each type of files to be stored is the ratio of the redundancy coefficient of each type of files to be stored to the non-repetitive proportion of the corresponding file in each type of files to be stored.
[0068] It can be understood from the deduplication degree of each file in various files to be stored that the more redundant data blocks in the file, the lower its storage priority level may be, and the more necessary it is to deduplicate the file. That is, the smaller the non-repetitive proportion of the file and the larger the redundancy coefficient, the greater the deduplication degree, and the smaller the storage priority level of the file data; conversely, the fewer redundant data blocks in the file, the higher its storage priority level may be, that is, the larger the non-repetitive proportion of the file and the smaller the redundancy coefficient, the smaller the deduplication degree, and the larger the storage priority level of the file data.
[0069] Preferably, the schematic diagram of the deduplication degree extraction process provided in this embodiment is as Figure 2 shown.
[0070] Step S4: Based on the access frequencies of various files to be stored and the deduplication degree, determine the priority storage coefficient of each file in various files to be stored, and combine the number of all data blocks in various files to be stored to determine the storage priorities of various files to be stored, and store all types of files to be stored.
[0071] If the deduplication degree of a file is greater, it means that there are more redundant data in it originally, and a lower storage priority can be set for it. If the repetition rate of the data file is low and the deduplication degree is small, it means that most of the data is unique, and the data storage priority needs to be improved.
[0072] Therefore, based on the access frequencies of various files to be stored and the deduplication degree, determine the priority storage coefficient of each file in various files to be stored, and combine the number of all remaining data blocks except redundant data blocks in various files to be stored to determine the storage priorities of various files to be stored, and store all types of files to be stored. Specifically:
[0073] (1) Based on the access frequencies of various files to be stored and the deduplication degree, determine the priority storage coefficient of each file in various files to be stored, specifically:
[0074] The priority storage coefficient γ k,i of file i in the kth type of file to be stored is expressed as: In the formula, β k represents the access frequency of the kth type of file to be stored; ω k,i represents the deduplication degree of file i in the kth type of file to be stored; norm() represents the normalization function; σ represents a preset constant greater than 0, which is used to prevent the denominator from being 0. The value of σ is set artificially. In this embodiment, the value of σ is 0.01. On the premise of ensuring that the denominator is not 0 and not overly affecting the calculation result, the implementer can set it according to the specific situation by himself, and this embodiment does not make special restrictions.
[0075] It can be understood from the priority storage coefficients of each file in various types of files to be stored that the more redundant data blocks there are in a file, the greater the degree of deduplication of the file. If the access frequency is smaller at this time, its priority storage coefficient is lower; conversely, the fewer redundant data blocks there are in a file, the smaller the degree of deduplication of the file. If the access frequency is larger at this time, its priority storage coefficient is higher.
[0076] (2) Further, based on the priority storage coefficients of each file in various types of files to be stored and in combination with the number of all data blocks except redundant data blocks in various types of files to be stored, determine the storage priorities of various types of files to be stored. Specifically:
[0077] The storage priority U k of the k-th type of file to be stored has the following expression: In the formula, represents the average value of the priority storage coefficients of all files in the k-th type of file to be stored; C k represents the number of all data blocks in the k-th type of file to be stored.
[0078] Among them, since the storage device will allocate at least one data block to each file, the denominator cannot be 0.
[0079] It can be understood from the storage priorities of various types of files to be stored that the larger the priority storage coefficients of all files in the same type of file to be stored, the fewer data blocks are allocated to the files in this type of file to be stored, indicating that the average storage priority and the smaller data scale correspond to a larger storage priority. Then the storage priority of this type of file to be stored is higher, and this type of file to be stored will be stored first; conversely, the smaller the priority storage coefficients of all files in the same type of file to be stored, the more data blocks are allocated to the files in this type of file to be stored, then the storage priority of this type of file to be stored is lower.
[0080] Preferably, the schematic diagram of the process for obtaining the storage priority provided in this embodiment is as Figure 3 shown.
[0081] (3) Further, store them based on the storage priorities of various types of files to be stored. Specifically: Sort various types of files to be stored in descending order of storage priority and store them in the storage device.
[0082] It should be noted that: The above sequence of embodiments of the present application is only for description and does not represent the advantages and disadvantages of the embodiments. And the above describes specific embodiments of this specification. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be beneficial.
[0083] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
[0084] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; modifying the technical solutions recorded in the foregoing embodiments, or equivalently replacing some of the technical features, does not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A method for optimizing storage of massive file data, characterized in that: The method comprises the following steps: Obtain all types of files to be stored and their respective access frequencies, and allocate multiple data blocks for each file in each type of data to be stored for storage; In various types of files to be stored, based on the difference between any two data blocks in each file and the difference between any two data blocks between each file and other files, redundant data blocks of various sizes in each file are obtained; based on the distribution of all redundant data blocks in each file, the data redundancy of each file is determined; Based on the proportion of redundant data blocks of the same size between each file and the remaining files, determine the redundancy index between each file in each type of files to be stored and the remaining files, and determine the redundancy similarity between each file in each type of files to be stored and the remaining files in combination with the data redundancy, so as to determine the redundancy coefficient of each type of files to be stored; based on the proportion of remaining data blocks in each file except the redundant data blocks, and the redundancy coefficient, determine the deduplication degree of each file in each type of files to be stored; Based on the access frequency of each type of files to be stored and the degree of deduplication, the priority storage coefficient of each file in each type of files to be stored is determined, and combined with the number of all data blocks in each type of files to be stored, the storage priority of each type of files to be stored is determined, and all types of files to be stored are stored.
2. The method for optimizing storage of massive file data according to claim 1, characterized in that: The method for obtaining redundant data blocks of various sizes in each file is as follows: The data in each data block is in the form of a binary number. In all files of various types of files to be stored, a bitwise XOR operation is performed on the binary numbers between any two data blocks of the same size, and any data block of the two data blocks in which all bits of the binary result of the operation are 0 is marked as a redundant data block. All data blocks are traversed to obtain all redundant data blocks in various types of files to be stored, wherein the marked redundant data blocks are not marked repeatedly, and redundant data blocks of various sizes in each file are obtained by statistics.
3. The method for optimizing storage of massive file data according to claim 1, characterized in that: The data redundancy of each file is the proportion of all redundant data blocks in all data blocks in each file.
4. The method for optimizing storage of massive file data according to claim 1, characterized in that: The method for determining the redundancy index between each file in each type of files to be stored and the other files is as follows: In each type of files to be stored, the ratio of the number of types of redundant data blocks of the same size between each file and the other files to the total number of types of redundant data blocks of all different sizes is used as the redundancy index between each file and the other files in each type of files to be stored.
5. The method for optimizing storage of massive file data according to claim 1, characterized in that: The expression of the redundant similarity between each file in the various types of files to be stored and the other files is: In the formula, represents the redundant similarity between file i and file j in the kth category of files to be stored; represents the redundancy index between file i and file j in the kth type of files to be stored; β i , β j They represent the data redundancy of file i and file j respectively; norm() represents the normalization function; ε represents a constant preset to be greater than 0.
6. The method for optimizing storage of mass file data according to claim 1, characterized in that: The method for determining the redundancy coefficient of each type of files to be stored is: The mean value of the redundant similarity between each file in each type of files to be stored and all other files is calculated, and the average value of the means of the redundant similarity of all files is used as the redundant coefficient of each type of files to be stored.
7. The method for optimizing storage of mass file data according to claim 1, characterized in that: The method for determining the deduplication degree of each file in the various types of files to be stored is: Calculate the proportion of the remaining data blocks in all data blocks in each file of each type of files to be stored, excluding the redundant data blocks, and record it as the non-duplicate proportion of each file in each type of files to be stored; The degree of deduplication of each file in each type of files to be stored is the ratio of the redundancy coefficient of each type of files to be stored to the non-duplicate proportion of each file in the corresponding type of files to be stored.
8. The method for optimizing storage of massive file data according to claim 1, characterized in that: The expression of the priority storage coefficient of each file in the various types of files to be stored is: In the formula, γ k,i represents the priority storage coefficient of file i in the kth category of files to be stored; β k represents the access frequency of the kth type of files to be stored; ω k,i represents the degree of deduplication of file i in the kth category of files to be stored; norm() represents the normalization function; σ represents a constant preset to be greater than 0.
9. The method for optimizing storage of mass file data according to claim 1, characterized in that: The expression of the storage priority of each type of files to be stored is: Where U k Indicates the storage priority of the kth type of files to be stored; represents the mean value of the priority storage coefficients of all files in the kth category of files to be stored; C k Indicates the number of all data blocks in the k-th type of files to be stored.
10. The method for optimizing storage of mass file data according to claim 1, characterized in that: The storing of all types of files to be stored includes: Sort the various files to be stored in the storage device in descending order of storage priority.
Citation Information
Patent Citations
High-performance hierarchical storage system supporting heterogeneous storage
CN107943867A
Distributed storage system data storage method, apparatus, system, and storage medium
CN109241023A