File Processing Method, Apparatus, Computer Device, Storage Medium, and Program Product
By configuring multiple sample granularities for file block sampling and summary calculation, the problem of slow file summary calculation speed in the prior art is solved, efficient file matching processing is achieved, and the accuracy of matching results is improved.
Patent Information
- Application Number
- CN202210527085.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-05-16
AI Technical Summary
The prior art requires traversing all the contents of the file when calculating file summary information, resulting in slow calculation speed, which in turn affects file matching efficiency, especially for large files.
By configuring multiple sampling granularities, the target file is sampled by blocks, and the sampling file is summarized to obtain the summary information of the target file. Then, the matching process is performed based on the summary information of the target file and the summary information in the reference file set.
It improves the calculation speed of file summary information, improves file matching efficiency, and performs matching processing under multiple sample granularity to improve the accuracy of matching results.
Smart Images

Figure CN115114242B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a file processing method, a file processing device, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In a file processing scenario, it is often necessary to perform matching processing between files, and then corresponding business processing can be performed on the files according to the matching results. The matching processing between files can be performed through the digest information of the files. When the digest information of two files matches, it can be determined that the two files match. Currently, when calculating the digest information of a file, it is necessary to traverse all the content of the file, which makes the calculation speed of the digest information of the file slow, resulting in low file matching efficiency. Especially for some files with large volumes, the impact of traversing all the content of the file to calculate the digest information on the file matching efficiency is more obvious. Therefore, how to improve the calculation speed of file digest information to improve file matching efficiency has become a current research hotspot. Summary of the Invention
[0003] Embodiments of the present application provide a file processing method, device, computer device, storage medium, and program product, which can improve the calculation speed of file digest information to improve file matching efficiency.
[0004] On the one hand, embodiments of the present application provide a file processing method, which includes:
[0005] Obtain a target file to be processed, where the target file includes multiple file blocks, and the target file is configured with M sampling granularities, and M is a positive integer;
[0006] At least one of the M sampling granularities, perform file block sampling processing on the target file, and calculate the digest of the sampled file obtained by the file block sampling processing to obtain the digest information of the target file;
[0007] Obtain the digest information of each reference file in the reference file set for matching the target file, and based on the digest information of each reference file and the digest information of the target file, perform matching processing on the target file in the reference file set;
[0008] After obtaining the matching result of the target file, perform business processing on the target file according to the matching result.
[0009] Correspondingly, embodiments of the present application provide a file processing device, which includes:
[0010] An acquisition unit, configured to acquire a target file to be processed, where the target file includes multiple file blocks, and the target file is configured with M sampling granularities, and M is a positive integer;
[0011] A processing unit, configured to perform file block sampling processing on the target file at at least one of the M sampling granularities, and calculate a digest of the sampled file obtained by the file block sampling processing to obtain digest information of the target file;
[0012] An acquisition unit, configured to acquire digest information of each reference file in a reference file set for matching processing of the target file;
[0013] The processing unit is further configured to perform matching processing on the target file in the reference file set based on the digest information of each reference file and the digest information of the target file;
[0014] The processing unit is further configured to, after obtaining a matching result of the target file, perform service processing on the target file according to the matching result.
[0015] In one implementation, the M sampling granularities are arranged in order. After acquiring the target file to be processed, the processing unit is further configured to perform the following steps:
[0016] Traverse the M sampling granularities in sequence;
[0017] Wherein, each time a sampling granularity is traversed, a file matching process is triggered to be executed. A file matching process includes the following operations: file block sampling processing, digest calculation, acquiring digest information of each reference file, and performing matching processing on the target file;
[0018] After executing the file matching process at the currently traversed sampling granularity, the processing unit is further configured to perform the following steps:
[0019] If it is determined through the file matching process at the currently traversed sampling granularity that there is no reference file in the reference file set that matches the target file, pruning processing is performed on the sampling granularities in the M sampling granularities that have not been traversed, so as to stop traversing the M sampling granularities, and a matching result indicating that the matching processing fails is generated.
[0020] In one implementation, after executing the file matching process at the currently traversed sampling granularity, the processing unit is further configured to perform the following steps:
[0021] If it is determined through the file matching process at the currently traversed sampling granularity that there is a reference file in the reference file set that matches the target file, the traversal progress of the M sampling granularities is detected;
[0022] If the detected traversal progress indicates that all M sampling granularities have been traversed, a matching result for indicating successful matching processing is generated;
[0023] If the detected traversal progress indicates that the M sampling granularities have not been traversed, the file matching process for the next sampling granularity of the currently traversed sampling granularity is executed.
[0024] In one implementation, in the file matching process for the currently traversed sampling granularity, the digest information of the target file is obtained by calculating the digest of the sampled file of the target file at the currently traversed sampling granularity; the digest information of any reference file is obtained by calculating the digest of the sampled file of the corresponding reference file at the currently traversed sampling granularity;
[0025] For the file matching process at the currently traversed sampling granularity, the processing unit, when performing matching processing on the target file in the reference file set based on the digest information of each reference file and the digest information of the target file, is specifically configured to perform the following steps:
[0026] If there is a reference file in the reference file set whose digest information matches the digest information of the target file, it is determined that there is a reference file in the reference file set that matches the target file;
[0027] If there is no reference file in the reference file set whose digest information matches the digest information of the target file, it is determined that there is no reference file in the reference file set that matches the target file.
[0028] In one implementation, in the file matching process for the currently traversed sampling granularity, if the currently traversed sampling granularity is ranked first among the M sampling granularities, the reference file set includes: files in the file library that match the file name or file volume of the target file;
[0029] If the currently traversed sampling granularity is ranked non-first among the M sampling granularities, the reference file set includes: reference files that match the target file determined through the file matching process for the previous sampling granularity of the currently traversed sampling granularity.
[0030] In one implementation, the processing unit, when performing file block sampling processing on the target file at at least one sampling granularity among the M sampling granularities, is specifically configured to perform the following steps:
[0031] At each of the M sampling granularities, file block sampling processing is respectively performed on the target file to obtain M sampled files of the target file, with one sampled file corresponding to one sampling granularity;
[0032] Among them, the abstract information of the target file includes: the abstract calculation results of each of the M sampled files of the target file; the abstract information of any reference file includes: the abstract calculation results of the sampled files of the corresponding reference file at each sampling granularity.
[0033] In one implementation, the processing unit, when performing matching processing on the target file in the reference file set based on the abstract information of each reference file and the abstract information of the target file, is specifically used to perform the following steps:
[0034] For any reference file in the reference file set, if the abstract calculation result in the abstract information of any reference file matches the abstract calculation result in the abstract information of the target file according to the sampling granularity, it is determined that any reference file matches the target file;
[0035] If there is an abstract calculation result in the abstract information of any reference file and the abstract information of the target file that does not match according to the sampling granularity, it is determined that the target file does not match any reference file;
[0036] If there is a reference file in the reference file set that matches the target file, a matching result for indicating successful matching processing is generated;
[0037] If there is no reference file in the reference file set that matches the target file, a matching result for indicating failed matching processing is generated.
[0038] In one implementation, for the i-th sampling granularity among the M sampling granularities, when the processing unit performs file block sampling processing on the target file at at least one of the M sampling granularities, it is specifically used to perform the following steps:
[0039] Determine the sampling interval corresponding to the i-th sampling granularity;
[0040] According to the sampling interval corresponding to the i-th sampling granularity, perform file block sampling processing on the target file to obtain the sampled file block of the target file at the i-th sampling granularity;
[0041] Perform splicing processing on the sampled file blocks of the target file at the i-th sampling granularity to obtain the sampled file of the target file at the i-th sampling granularity, where i is a positive integer less than or equal to M.
[0042] In one implementation, the reference file set has been backed up to the cloud, and the target file is a file to be identified for the backup status; the service processing includes file status synchronization processing; when the processing unit performs service processing on the target file according to the matching result, it is specifically used to perform the following steps:
[0043] If the matching result indicates that the matching process fails, mark the target file with an unbacked-up status identifier;
[0044] If the matching result indicates that the matching process succeeds, mark the target file with a backed-up status identifier.
[0045] In one implementation, the target file is the file currently being moved, and the move operation is used to move the target file to the target folder; the reference file set is stored in the target folder; the service processing includes file move processing; when the processing unit is used to perform service processing on the target file according to the matching result, it is specifically used to perform the following steps:
[0046] If the matching result indicates that the matching process succeeds, reject the move operation to prevent the target file from being moved to the target folder;
[0047] If the matching result indicates that the matching process fails, accept the move operation and move the target file to the target folder.
[0048] Correspondingly, an embodiment of the present application provides a computer device, which includes a processor and a computer-readable storage medium, where:
[0049] The processor is adapted to implement a computer program;
[0050] The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the above file processing method.
[0051] Correspondingly, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is read and executed by the processor of the computer device, the computer device is caused to perform the above file processing method.
[0052] Correspondingly, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the above file processing method.
[0053] In the embodiments of the present application, for a target file to be processed, the target file may include multiple file blocks, and the target file may be configured with one or more sampling granularities. The target file can be subjected to file block sampling processing at at least one of the one or more sampling granularities, and a digest calculation can be performed on the sampled file obtained by the sampling processing to obtain the digest information of the target file. Then, based on the digest information of the target file and the digest information of each reference file in the reference file set, a matching process can be performed on the target file in the reference file set. After obtaining the matching result of the target file, the target file can be processed according to the matching result; in the embodiments of the present application, by configuring one or more sampling granularities to perform file block sampling processing on the target file, the file volume of the sampled file obtained by the file block sampling processing is smaller than the file volume of the target file, thereby improving the calculation speed of the file digest information and enhancing the file matching efficiency; moreover, when the number of configured sampling granularities is greater than or equal to two, the file can be matched at multiple sampling granularities, and the accuracy of the file matching result can also be improved. Description of the Drawings
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0055] Figure 1 It is a schematic structural diagram of a logical structure of a file provided by an embodiment of the present application;
[0056] Figure 2 It is a schematic flowchart of a file processing method provided by an embodiment of the present application;
[0057] Figure 3 It is a schematic diagram of the calculation process of digest information under a single sampling granularity provided by an embodiment of the present application;
[0058] Figure 4a It is a schematic diagram of file block sampling processing provided by an embodiment of the present application;
[0059] Figure 4b It is a schematic diagram of another file block sampling processing provided by an embodiment of the present application;
[0060] Figure 5 It is a schematic flowchart of a file matching method with multiple sampling granularities and pruning processing provided by an embodiment of the present application;
[0061] Figure 6aIt is a schematic diagram showing successful matching of a file matching method with multiple sampling granularities and pruning processing provided by an embodiment of the present application;
[0062] Figure 6b It is a schematic diagram showing failed matching of a file matching method with multiple sampling granularities and pruning processing provided by an embodiment of the present application;
[0063] Figure 7 It is a schematic flowchart of another file processing method provided by an embodiment of the present application;
[0064] Figure 8 It is a schematic diagram showing the calculation process of abstract information under multiple sampling granularities provided by an embodiment of the present application;
[0065] Figure 9a It is a schematic diagram showing successful matching of a file matching method integrating multiple sampling granularities provided by an embodiment of the present application;
[0066] Figure 9b It is a schematic diagram showing failed matching of a file matching method integrating multiple sampling granularities provided by an embodiment of the present application;
[0067] Figure 10 It is a schematic diagram of the interface of a file status synchronization scenario provided by an embodiment of the present application;
[0068] Figure 11 It is a schematic diagram of the architecture of a file processing system provided by an embodiment of the present application;
[0069] Figure 12 It is a schematic diagram of the structure of a file processing device provided by an embodiment of the present application;
[0070] Figure 13 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0071] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0072] Embodiments of this application relate to files. The files related to the embodiments of this application refer to computer files, which are collections of information stored on a computer with a computer hard disk as the carrier. The files related to the embodiments of this application may include, but are not limited to, any of the following: text documents, pictures, audio, video, and application programs. The embodiments of this application do not limit the types of files. Files usually have file extensions in the format of ".XXX" or ".XXXX", and file extensions can be used to indicate file types. For example, files with file extensions ".jpg" or ".jpeg" are pictures, files with file extensions ".doc" or ".docx" are text documents, files with file extensions ".mp3" are audio, and files with file extensions ".mp4" are video. Files can be divided in units of file blocks (which can also be called logical blocks). File blocks can be divided into a sequence of file blocks in units of file blocks. And during the file division process, the part less than one file block can be divided as one file block. It can be understood that a file can include multiple file blocks, and a file is a sequence of file blocks composed of multiple file blocks. As Figure 1 shown, a file is a sequence of file blocks composed of N file blocks, and the N file blocks are respectively file block 1, file block 2, file block 3, …, file block N, where N is a positive integer greater than 1. Among them, a file block is the basic unit of file system I / O (Input / Output) operations. A file system refers to the method by which an operating system organizes files on a storage device (such as a computer hard disk), and the software mechanism in the operating system responsible for managing and storing file information is called a file system.
[0073] Embodiments of this application relate to file matching. File matching refers to the operation of determining whether files match by comparing whether the digest information of the files is the same (i.e., identical). Taking a target file and a reference file as an example, if the digest information of the target file is the same as the digest information of the reference file, it can be determined that the target file and the reference file match, that is, the target file and the reference file are the same file. If the digest information of the target file is different from the digest information of the reference file, it can be determined that the target file and the reference file do not match, that is, the target file and the reference file are not the same file. The digest information of a file can be the digest calculation result obtained by performing a digest calculation on the content of the file using a digest algorithm. Among them, the digest algorithm can include, but is not limited to, any of the following: MD5 (Message-Digest Algorithm), SHA-1 (SecureHash Algorithm-1), SHA-256 (Secure Hash Algorithm-256), and SHA-512 (Secure Hash Algorithm-512). The embodiments of this application do not limit the types of digest algorithms.
[0074] Based on the relevant descriptions of files and file matching, an embodiment of the present application provides a file processing method. In this file processing method, one or more sampling granularities are configured for a file, and file block sampling processing (specifically, downsampling processing of file blocks) can be performed on the file at one or more sampling granularities. After obtaining a sampled file through file block sampling processing, a digest calculation can be performed on the sampled file to obtain the digest information of the file, so that file matching processing can be performed based on the digest information of the file. By configuring the sampling granularity to perform downsampling processing on file blocks, the downsampling processing of file blocks can be understood as an operation of extracting some file blocks from the file at a certain sampling interval and splicing them to form a new file. The sampling interval is related to the sampling granularity, and the new file formed by splicing is the sampled file. In this way, the file volume of the sampled file obtained by the downsampling processing of file blocks is smaller than the file volume of the original file. Compared with the method of traversing all the content of the file to calculate the digest information of the file, the file processing method provided by the embodiment of the present application can improve the calculation speed of the digest information of the file, thereby improving the file matching efficiency.
[0075] Among them, the sampling granularity can be used to characterize the degree of downsampling processing of file blocks. The larger the sampling granularity (it can also be understood as the coarser the sampling granularity), the larger the sampling interval determined according to the sampling granularity, the higher the degree of downsampling processing of file blocks, the greater the difference in file volume between the sampled file and the original file, and the greater the difference in file content between the sampled file and the original file. On the contrary, the smaller the sampling granularity (it can also be understood as the finer the sampling granularity), the smaller the sampling interval determined according to the sampling granularity, the lower the degree of downsampling processing of file blocks, the smaller the difference in file volume between the sampled file and the original file, and the smaller the difference in file content between the sampled file and the original file. By configuring different sampling granularities in the embodiment of the present application, file matching processing can be performed based on the digest information of the file at different sampling granularities, which can improve the accuracy of the file matching result.
[0076] The following combines Figures 2 - 10 to introduce the file processing method provided by the embodiment of the present application in detail.
[0077] An embodiment of the present application provides a file processing method, which mainly introduces the process of file block sampling processing, as well as the file matching methods of multi-sampling granularity and pruning processing. This file processing method can be executed by a computer device, and the computer device can be a terminal such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch, a vehicle-mounted terminal, a smart voice interaction device, and a smart home appliance. Please refer to Figure 2 , this file processing method may include the following steps S201 - step S206:
[0078] S201, obtain a target file to be processed.
[0079] The target file may include multiple file blocks, and the target file is configured with M sampling granularities. The sampling granularity may be a specific value, and the M sampling granularities may be arranged in order, specifically, they may be arranged in descending order of value (i.e., in order of decreasing granularity), and M is a positive integer.
[0080] S202. At at least one of the M sampling granularities, perform file block sampling processing on the target file.
[0081] S203. Calculate the digest of the sampled file obtained by the file block sampling processing to obtain the digest information of the target file.
[0082] In steps S202 - S203, for the i-th sampling granularity among the M sampling granularities, the process of performing file block sampling processing on the target file at the i-th sampling granularity and calculating the digest of the sampled file obtained by the file block sampling processing can be referred to Figure 3 , specifically, it can be referred to the following description: First, the sampling interval corresponding to the i-th sampling granularity can be determined; in one implementation, there may be a power operation relationship between the sampling interval and the sampling granularity. The sampling interval can be the result of a power operation with a target value (the target value can be an integer greater than or equal to 0) as the base and the sampling granularity as the exponent. Taking the target value as 2 as an example, if the i-th sampling granularity is represented as K i , then the sampling interval corresponding to the i-th sampling granularity can be represented as 2 Ki ; in another implementation, the sampling granularity can be directly determined as the sampling interval. For example, if the i-th sampling granularity is represented as K i , then the sampling interval corresponding to the i-th sampling granularity can be represented as K i . Second, according to the sampling interval corresponding to the i-th sampling granularity, perform file block sampling processing on the target file to obtain the sampled file blocks of the target file at the i-th sampling granularity. The sampled file blocks of the target file at the i-th sampling granularity can be concatenated, specifically, they can be concatenated in the arrangement order of the sampled file blocks in the target file to obtain the sampled file of the target file at the i-th sampling granularity; then, a digest algorithm can be used to calculate the digest of the sampled file at the i-th sampling granularity to obtain the digest information of the target file at the i-th sampling granularity, and i is a positive integer less than or equal to M.
[0083] Among them, the method of performing file block sampling processing on the target file according to the sampling interval corresponding to the i-th sampling granularity to obtain the sampled file blocks of the target file at the i-th sampling granularity may include any one of the following:
[0084] The first method for file block sampling processing is to group the file blocks included in the target file according to the sampling interval corresponding to the i-th sampling granularity, obtaining at least one file block group of the target file. The number of file blocks included in each file block group is less than or equal to the sampling interval corresponding to the i-th sampling granularity. Then, a sampled file block can be sampled from the same corresponding position of each file block group. That is to say, the arrangement position of each sampled file block in the corresponding file block group is the same, obtaining the sampled file blocks of the target file at the i-th sampling granularity. By sequentially splicing the sampled file blocks of the target file at the i-th sampling granularity, the sampled file of the target file at the i-th sampling granularity can be obtained. Take Figure 4a the schematic diagram of the file block sampling processing shown as an example. The target file includes 12 file blocks, namely file block 1, file block 2,..., file block 12. The i-th sampling granularity is 2, and the sampling interval corresponding to the i-th sampling granularity is 4. File blocks 1 - 4 can be used as the first file block group, file blocks 5 - 8 can be used as the second file block group, and file blocks 9 - 12 can be used as the third file block group. The file block 2 arranged in the second position in the first file block group is used as the first sampled file block, the file block 6 arranged in the second position in the second file block group is used as the second sampled file block, and the file block 10 arranged in the second position in the third file block group is used as the third sampled file block. Then, file blocks 2, 6, and 10 can be sequentially spliced to obtain the sampled file of the target file at the i-th sampling granularity. The sampled file of the target file at the i-th sampling granularity includes file blocks 2, 6, and 10.
[0085] The second method for file block sampling processing is to select the first sampled file block in the target file (generally, the first file block of the target file can be selected as the first sampled file block). After that, taking the first sampled file block as a reference, the file block at an interval of the number of the first sampled file block equal to the sampling interval corresponding to the i-th sampling granularity is used as the second sampled file block. Then, taking the second sampled file block as a reference, the file block at an interval of the number of the second sampled file block equal to the sampling interval corresponding to the i-th sampling granularity is used as the third sampled file block. Sampling continues in this way until the last file block of the target file is reached, obtaining the sampled file blocks of the target file at the i-th sampling granularity. Then, by sequentially splicing the sampled file blocks of the target file at the i-th sampling granularity, the sampled file of the target file at the i-th sampling granularity can be obtained. Take Figure 4bTaking the schematic diagram of the file block sampling process shown as an example, the target file includes 12 file blocks, namely file block 1, file block 2, …, file block 12. The i-th sampling granularity is 2, and the sampling interval corresponding to the i-th sampling granularity is 4. After selecting file block 1 as the first sampling file block, file block 6, which is 4 file blocks away from the first sampling file block, is used as the second sampling file block, and file block 11, which is 4 file blocks away from the second sampling file block, is used as the third sampling file block. Then, file block 1, file block 6, and file block 11 can be sequentially spliced to obtain the sampling file of the target file at the i-th sampling granularity. The sampling file of the target file at the i-th sampling granularity includes file block 2, file block 6, and file block 10.
[0086] S204, obtain the summary information of each reference file in the reference file set used for matching the target file.
[0087] S205, based on the summary information of each reference file and the summary information of the target file, perform matching processing on the target file in the reference file set.
[0088] The file matching method provided by the embodiments of the present application may include a file matching method with multiple sampling granularities and pruning processing. The file matching method with multiple sampling granularities and pruning processing means: taking a single sampling granularity as a unit, according to the order of different sampling granularities in M sampling granularities from large to small (i.e., from coarse to fine), iteratively perform the matching process between the reference file set and the target file. When it is determined in the current iteration that the matching process between the reference file set and the target file fails, stop the subsequent iterations of the current iteration. When it is determined in the current iteration that the matching process between the reference file set and the target file is successful, continue the subsequent iterations of the current iteration until the last sampling granularity is iterated. More specifically, the file matching method with multiple sampling granularities and pruning processing may include: after obtaining the target file to be processed, sequentially traverse M sampling granularities. Each time a sampling granularity is traversed, a file matching process can be triggered to execute; after the file matching process under the currently traversed sampling granularity is completed, if it is determined through the file matching process under the currently traversed sampling granularity that there is no reference file in the reference file set that matches the target file, pruning processing can be performed on the sampling granularities in M sampling granularities that have not been traversed, so as to stop traversing M sampling granularities, not execute the file matching processes under the sampling granularities in M sampling granularities that have not been traversed, and generate a matching result indicating that the matching process fails; if it is determined through the file matching process under the currently traversed sampling granularity that there is a reference file in the reference file set that matches the target file, the traversal progress of M sampling granularities can be detected. If the detected traversal progress indicates that all M sampling granularities have been traversed, a matching result indicating that the matching process is successful can be generated. If the detected traversal progress indicates that M sampling granularities have not been traversed, the file matching process under the next sampling granularity of the currently traversed sampling granularity can be executed.
[0089] Among them, a file matching process may include the following operations: file block sampling processing, digest calculation, obtaining the digest information of each reference file in the reference file set, and performing a matching process on the target file, that is, a file matching process may include step S202-step S205; that is to say, after obtaining the target file to be processed, sequentially traverse M sampling granularities. Each time a sampling granularity is traversed, it can trigger the execution of file block sampling processing on the target file under at least one sampling granularity in M sampling granularities (here, at least one sampling granularity specifically may refer to the currently traversed sampling granularity), performing digest calculation on the sampled file obtained by the file block sampling processing to obtain the digest information of the target file, obtaining the digest information of each reference file in the reference file set for performing a matching process on the target file, and based on the digest information of each reference file and the digest information of the target file, performing a matching process on the target file in the reference file set.
[0090] Moreover, in the file matching process at the currently traversed sampling granularity, if the currently traversed sampling granularity ranks first among the M sampling granularities, that is, the currently traversed sampling granularity is the first sampling granularity among the M sampling granularities, the reference file set may include: files in the file library that match the file name or file volume of the target file; if the currently traversed sampling granularity ranks non-first among the M sampling granularities, that is, the currently traversed sampling granularity is other than the first sampling granularity among the M sampling granularities, the reference file set may include: reference files that match the target file determined through the file matching process at the previous sampling granularity of the currently traversed sampling granularity.
[0091] In addition, in the file matching process at the currently traversed sampling granularity, the digest information of the target file is obtained by calculating the digest of the sampled file of the target file at the currently traversed sampling granularity; the digest information of any reference file is obtained by calculating the digest of the sampled file of the corresponding reference file at the currently traversed sampling granularity. For the file matching process at the currently traversed sampling granularity, the process of matching the target file in the reference file set based on the digest information of each reference file and the digest information of the target file may include: if there is a reference file in the reference file set whose digest information matches (i.e., is the same as) the digest information of the target file, it may be determined that there is a reference file in the reference file set that matches the target file; if there is no reference file in the reference file set whose digest information matches the digest information of the target file, it may be determined that there is no reference file in the reference file set that matches the target file.
[0092] The following will combine Figure 5 to sort out the file matching method of multi-sampling granularity and pruning processing. Before sorting out the file matching method of multi-sampling granularity and pruning processing, some concepts are clarified here: ① The target file can be configured with a sampling granularity combination, and the sampling granularity combination can be expressed as [K1, K2, …, K i , …, K M , and the sampling granularity combination may include M sampling granularities, and the values of the M sampling granularities gradually decrease. Among them, the i-th sampling granularity can be expressed as K i . ② The digest information of the target file q at the M sampling granularities can be expressed as [S1(q), S2(q),.., S i (q), …, S M (q)], where the digest information of the target file q at the i-th sampling granularity can be expressed as S i(q). ③In the first iteration (i.e., the file matching process at the first sampling granularity), the reference files in the reference file set Q1 are the files in the file library that match (i.e., are the same as) the file volume or file name of the target file; in non-first iterations (taking the i-th iteration, i.e., the file matching process at the i-th sampling granularity, i≠1 as an example), the reference files in the reference file set Q i are determined in the (i - 1)-th iteration (i.e., the file matching process at the (i - 1)-th sampling granularity), and are the reference files in the reference file set Q i-1 in the (i - 1)-th iteration that match the target file. ④The abstract information of each reference file in the reference file set can be included in the abstract information set R of the reference file set. The abstract information set R of the reference file set can include the abstract information of each reference file in the reference file set at M sampling granularities. In the i-th iteration, the abstract information set of the reference file set Q i at the i-th sampling granularity can be expressed as R i , and the abstract information set R i can include the abstract information of each reference file in the reference file set Q i at the i-th sampling granularity.
[0093] Based on the above concepts, taking the i-th iteration (i.e., the file matching process at the i-th sampling granularity) as an example, the file matching process at the i-th sampling granularity can include:
[0094] (1) Obtain the abstract information set R i of the reference file set Q i at the i-th sampling granularity. The abstract information set R i of the reference file set Q i at the i-th sampling granularity can include the abstract information of each reference file in the reference file set Q i at the i-th sampling granularity K i .
[0095] (2) Calculate the abstract information S i (q) of the target file q at the i-th sampling granularity. Specifically, at the i-th sampling granularity K i , perform file block sampling processing on the target file q, and perform abstract calculation on the sampled file obtained by the file block sampling processing to obtain the abstract information S i of the target file q at the i-th sampling granularity K i (q).
[0096] (3) Compare the abstract information S i (q) of the target file q at the i-th sampling granularity K i with the abstract information set R i of the reference file set Qi Specifically, by matching the target file q at the i-th sampling granularity K i Summary information under S i (q), with reference document set Q i The summary information set R i By matching the various summary information in the reference document set Q i The summary information set R at the i-th sampling granularity i Does the target file q exist in the i-th sampling granularity K? i Summary information under S i (q) to determine the reference document set Q i Check whether there is a reference file that matches the target file (that is, the reference file corresponding to the matching item).
[0097] (4) If the reference document set Q i The summary information set R at the i-th sampling granularity i There is a target file q in the i-th sampling granularity K i Summary information under S i (q), i.e., the reference document set Q i If there is a reference file matching the target file in the sampling granularity combination, it is determined whether all sampling granularities in the sampling granularity combination have been traversed. If all sampling granularities in the sampling granularity combination have been traversed, a matching result indicating that the matching process is successful is generated; if all sampling granularities in the sampling granularity combination have not been traversed, then i=i+1 is set, and the file matching process under the i+1th sampling granularity is executed. The reference file set Q in the file matching process under the i+1th sampling granularity is i+1 Based on the reference document set Q i The reference file that matches the target file (ie, the reference file corresponding to the matching item) is determined.
[0098] (5) If the reference document set Q i The summary information set R at the i-th sampling granularity i The target file q does not exist in the i-th sampling granularity K i Summary information under S i (q), i.e., the reference document set Q i If there is no reference file matching the target file in the M sampling granularities, pruning is performed. Pruning means stopping traversing the M sampling granularities, that is, stopping the file matching process under the sampling granularities that have not been traversed in the sampling granularity combination in advance, and generating a matching result indicating that the matching process has failed.
[0099] by Figure 6aTaking the schematic diagram of the file matching method with multi-sampling granularity and pruning processing shown as an example, the sampling granularity combination includes three sampling granularities [2, 1, 0]. In the file matching process under the first sampling granularity, the digest information of the target file under the first sampling granularity is 123456a, and the reference file set includes three reference files. The digest information set of the reference file set under the first sampling granularity is [123456a, 123456a, 123456c], that is, the digest information of reference file 1 under the first sampling granularity is 123456a, the digest information of reference file 2 under the first sampling granularity is 123456a, and the digest information of reference file 3 under the first sampling granularity is 123456c. Through the file matching process under the first sampling granularity, it can be determined that reference file 1 and reference file 2 in the reference file set are the reference files that match the target file, and the reference file set in the file matching process under the second sampling granularity can be determined according to reference file 1 and reference file 2; similarly, in the file matching process under the second sampling granularity, the digest information of the target file under the second sampling granularity is 1234567a, and the digest information set of the reference file set under the second sampling granularity is [1234567a, 1234567b], that is, the digest information of reference file 1 under the second sampling granularity is 1234567a, and the digest information of reference file 2 under the second sampling granularity is 1234567b. Through the file matching process under the second sampling granularity, it can be determined that reference file 1 in the reference file set is the reference file that matches the target file, and the reference file set in the file matching process under the third sampling granularity can be determined according to reference file 1; similarly, in the file matching process under the third sampling granularity, the digest information of the target file under the third sampling granularity is 12345678a, and the digest information set of the reference file set under the third sampling granularity is [12345678a], that is, the digest information of reference file 1 under the third sampling granularity is 12345678a. Through the file matching process under the third sampling granularity, it can be determined that reference file 1 in the reference file set is the reference file that matches the target file. At this time, all the sampling granularities in the sampling granularity combination have been traversed, and a matching result indicating successful matching processing can be generated.
[0100] Taking Figure 6bTaking the schematic diagram of the file matching method with multi-sampling granularity and pruning processing shown as an example, the sampling granularity combination includes three sampling granularities [2, 1, 0]. In the file matching process under the first sampling granularity, the digest information of the target file under the first sampling granularity is 123456a, and the reference file set includes three reference files. The digest information set of the reference file set under the first sampling granularity is [123456b, 123456c, 123456d], that is, the digest information of reference file 1 under the first sampling granularity is 123456b, the digest information of reference file 2 under the first sampling granularity is 123456c, and the digest information of reference file 3 under the first sampling granularity is 123456d. Through the file matching process under the first sampling granularity, it can be determined that there is no reference file in the reference file set that matches the target file. At this time, pruning processing can be performed, and the file matching processes under the second and third sampling granularities can be stopped in advance, and a matching result indicating the failure of the matching process can be generated.
[0101] In summary, regarding the file matching method with multi-sampling granularity and pruning processing, among the M sampling granularities included in the sampling granularity combination, for the sampling granularity with a large value (i.e., a coarse granularity), the larger the file volume difference between the sampled file obtained by performing file block sampling under the sampling granularity with a large value and the original file, the greater the content difference between the sampled file and the original file, and the faster the calculation speed of the digest information. However, the accuracy during matching is lower, and the probability of file collision (i.e., the digest information of two files is the same, but the two files are not the same file) is higher; for the sampling granularity with a small value (i.e., a fine granularity), the smaller the file volume difference between the sampled file obtained by performing file block sampling under the sampling granularity with a small value and the original file, the smaller the content difference between the sampled file and the original file, the higher the accuracy during matching, and the lower the probability of file collision. However, the calculation speed of the digest information is slower. The embodiments of the present application configure multiple sampling granularities with different numerical sizes for the file, comprehensively consider the file matching efficiency and the file matching accuracy rate, and have the ability to balance the file matching efficiency and the file matching accuracy rate. Moreover, in the file matching process under any sampling granularity, if it is determined through the file matching process under this sampling granularity that there is no reference file in the reference file set that matches the target file, then pruning processing can save the time consumption of the file matching processes under the sampling granularities that have not been traversed in the sampling granularity combination, and accelerate the file matching efficiency in a statistical sense. In addition, when there is a power operation relationship between the sampling interval and the sampling granularity, and in the case of using the file block sampling processing method of file block grouping, the last sampling granularity among the M sampling granularities can be configured as 0, so that the sampled file obtained by sampling under the last sampling granularity is the same file as the original file, further improving the file matching accuracy.
[0102] S206. After obtaining the matching result of the target file, perform business processing on the target file according to the matching result.
[0103] After obtaining the matching result of the target file, business processing can be performed on the target file according to the matching result of the target file. The matching result can include a matching result indicating successful matching processing or a matching result indicating failed matching processing.
[0104] In the embodiments of the present application, by configuring multiple sampling granularities with different values for the file, the file matching efficiency and the file matching accuracy rate are comprehensively considered. It has the ability to balance the file matching efficiency and the file matching accuracy rate, and can improve the file matching efficiency on the premise of ensuring the file matching accuracy. Moreover, in the file matching process at any sampling granularity, if it is determined through the file matching process at this sampling granularity that there is no reference file in the reference file set that matches the target file, then through pruning processing, the time consumption of the file matching process at the sampling granularities that have not been traversed in the sampling granularity combination can be saved, and the file matching efficiency is accelerated statistically. The file matching method with multiple sampling granularities and pruning processing provided by the embodiments of the present application is more superior in terms of the accuracy rate and the file matching efficiency during the file matching process for large-volume files.
[0105] The embodiments of the present application also provide a file processing method. This file processing method mainly introduces the file matching method integrating multiple sampling granularities, the file matching method of a single sampling granularity, and the business processing of files based on the matching result, etc. This file processing method can be executed by a computer device. The computer device can be a terminal such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch, a vehicle-mounted terminal, a smart voice interaction device, and a smart home appliance. Please refer to Figure 7 , this file processing method can include steps S701 - step S709:
[0106] S701. Obtain the target file to be processed.
[0107] The target file can include multiple file blocks, and the target file is configured with M sampling granularities. The sampling granularity can be a specific value, and the M sampling granularities can be arranged in order, specifically in the order of decreasing numerical value (i.e., from coarser to finer granularity), and M is a positive integer.
[0108] S702. Perform file block sampling processing on the target file at at least one of the M sampling granularities.
[0109] S703. Calculate the digest of the sampling file obtained by the file block sampling processing to obtain the digest information of the target file.
[0110] In steps S702 - step S703, the aboveFigure 2 In the illustrated embodiment, the file matching method with multi-sampling granularity and pruning processing is mainly introduced. In addition to the file matching method with multi-sampling granularity and pruning processing, the file matching method provided by the embodiments of the present application may further include a file matching method that synthesizes multi-sampling granularity. The file matching method that synthesizes multi-sampling granularity means that the summary information of a file includes the summary calculation results of the file under each of multiple sampling granularities. When performing file matching, the summary calculation results of the file under each of the multiple sampling granularities are comprehensively compared. In the file matching method that synthesizes multi-sampling granularity, in at least one of the M sampling granularities, the process of performing file block sampling on the target file may include: performing file block sampling on the target file respectively under each of the M sampling granularities to obtain M sampled files of the target file, and one sampled file corresponds to one sampling granularity. In the file matching method that synthesizes multi-sampling granularity, the process of calculating the summary information of the target file by performing summary calculation on the sampled files obtained by file block sampling may include: performing summary calculation on each of the M sampled files of the target file respectively to obtain the summary calculation results of each sampled file of the target file. The summary information of the target file may include the summary calculation results of the sampled files of the target file under each sampling granularity. For example, the target file is configured with three sampling granularities, namely the first sampling granularity, the second sampling granularity, and the third sampling granularity. The summary information of the target file may include the summary calculation results of the sampled files of the target file under the first sampling granularity, the summary calculation results of the sampled files of the target file under the second sampling granularity, and the summary calculation results of the sampled files of the target file under the third sampling granularity.
[0111] It should be noted that in the process of calculating the summary information of the target file, the summary calculation processes of the target file under each sampling granularity may be parallel, that is, the summary calculation processes of the target file under each sampling granularity can be executed synchronously. Or, the summary calculation processes of the target file under each sampling granularity may be serial, that is, the M sampling granularities can be traversed in sequence. For each traversed sampling granularity, a summary calculation process is executed. After the summary calculation process of the target file under the currently traversed sampling granularity is completed, the summary calculation process of the target file under the next sampling granularity of the currently traversed sampling granularity is executed until all M sampling granularities are traversed. For example Figure 8Schematic diagram of the calculation process of the summary information under the multi-sampling granularity shown. The target file can be sampled in file blocks at the i-th sampling granularity to obtain the sampled file of the target file at the i-th sampling granularity. Then, the sampled file of the target file at the i-th sampling granularity can be subjected to summary calculation to obtain the summary calculation result of the target file at the i-th sampling granularity. After obtaining the summary calculation result of the target file at the i-th sampling granularity, it can be determined whether M sampling granularities have been traversed. If M sampling granularities have not been traversed, then i = i + 1 can be set, and the sampled file of the target file at the (i + 1)-th sampling granularity can be subjected to summary calculation. If M sampling granularities have been traversed, then the summary information of the target file can be determined according to the summary calculation results of the target file at each of the M sampling granularities.
[0112] S704, obtain the summary information of each reference file in the reference file set used for matching the target file.
[0113] In the file matching method integrating multi-sampling granularity, the reference files included in the reference file set are the files in the file library that match the file name or file volume of the target file. The summary information of any reference file in the reference file set can include the summary calculation results of the sampled files of the corresponding reference file at each of the M sampling granularities.
[0114] S705, for any reference file in the reference file set, if the summary calculation result in the summary information of any reference file matches the summary calculation result in the summary information of the target file according to the sampling granularity, it is determined that any reference file matches the target file.
[0115] In the file matching method that combines multiple sampling granularities, if the digest calculation results in the digest information of any reference file match the digest calculation results in the digest information of the target file according to the sampling granularity, it can be determined that any reference file matches the target file; matching according to the sampling granularity can be understood as: the digest calculation results of the target file at each sampling granularity match (i.e., are the same as) the digest calculation results of the reference file at the corresponding sampling granularity, then it is determined that the target file matches the reference file. For example, the target file and the reference file are configured with three sampling granularities, namely the first sampling granularity, the second sampling granularity, and the third sampling granularity. The digest calculation result of the target file at the first sampling granularity is 123456a, the digest calculation result of the target file at the second sampling granularity is 1234567a, and the digest calculation result of the target file at the third sampling granularity is 12345678a. The digest calculation result of the reference file at the first sampling granularity is 123456a, the digest calculation result of the reference file at the second sampling granularity is 1234567a, and the digest calculation result of the reference file at the third sampling granularity is 12345678a. It can be seen that the digest calculation result of the target file at the first sampling granularity matches the digest calculation result of the reference file at the first sampling granularity, the digest calculation result of the target file at the second sampling granularity matches the digest calculation result of the reference file at the second sampling granularity, and the digest calculation result of the target file at the third sampling granularity matches the digest calculation result of the reference file at the third sampling granularity, then it can be determined that the target file matches the reference file.
[0116] S706, if there are digest calculation results in the digest information of any reference file and the digest information of the target file that do not match according to the sampling granularity, it is determined that the target file does not match any reference file.
[0117] In the file matching method that combines multiple sampling granularities, if there are digest calculation results that do not match according to the sampling granularity in the digest information of any reference file and the target file, it is determined that the target file does not match any reference file; it can be understood that among the M sampling granularities, there is a sampling granularity such that the digest calculation result of the target file under this sampling granularity does not match the digest calculation result of the reference file under this sampling granularity, then it can be determined that the target file does not match the reference file. For example, the target file and the reference file are configured with three sampling granularities, namely the first sampling granularity, the second sampling granularity, and the third sampling granularity. The digest calculation result of the target file under the first sampling granularity is 123456a, the digest calculation result of the target file under the second sampling granularity is 1234567a, and the digest calculation result of the target file under the third sampling granularity is 12345678a. The digest calculation result of the reference file under the first sampling granularity is 123456b, the digest calculation result of the reference file under the second sampling granularity is 1234567b, and the digest calculation result of the reference file under the third sampling granularity is 12345678a. It can be seen that the digest calculation result of the target file under the first sampling granularity does not match the digest calculation result of the reference file under the first sampling granularity, and the digest calculation result of the target file under the second sampling granularity does not match the digest calculation result of the reference file under the second sampling granularity, then it can be determined that the target file does not match the reference file.
[0118] S707. If there is a reference file in the reference file set that matches the target file, a matching result indicating successful matching processing is generated.
[0119] S708. If there is no reference file in the reference file set that matches the target file, a matching result indicating failed matching processing is generated.
[0120] In steps S707 - S708, in the file matching method that combines multiple sampling granularities, if there is a reference file in the reference file set that matches the target file, a matching result indicating successful matching processing can be generated. If there is no reference file in the reference file set that matches the target file, a matching result indicating failed matching processing can be generated. Take Figure 9aTaking the schematic diagram of the file matching method with comprehensive multi-sampling granularity shown as an example, the reference file set includes three reference files, namely reference file 1, reference file 2, and reference file 3. If the digest calculation results of reference file 1 and the target file at three sampling granularities are matched according to the sampling granularity, then reference file 1 matches the target file. If the digest calculation results of reference file 2 and the target file at three sampling granularities do not match, then reference file 2 does not match the target file. If the digest calculation results of reference file 3 and the target file at three sampling granularities do not match, then reference file 3 does not match the target file. That is to say, if there is a reference file 1 in the reference file set that matches the target file, a matching result indicating successful matching processing can be generated. Taking Figure 9b Taking the schematic diagram of the file matching method with comprehensive multi-sampling granularity shown as an example, the reference file set includes three reference files, namely reference file 1, reference file 2, and reference file 3. If the digest calculation results of reference file 1 and the target file at three sampling granularities do not match, then reference file 1 does not match the target file. If the digest calculation results of reference file 2 and the target file at three sampling granularities do not match, then reference file 2 does not match the target file. If the digest calculation results of reference file 3 and the target file at three sampling granularities do not match, then reference file 3 does not match the target file. That is to say, if there is no reference file in the reference file set that matches the target file, a matching result indicating failed matching processing can be generated.
[0121] In summary, regarding the file matching method with comprehensive multi-sampling granularity, among the M sampling granularities, at the sampling granularity with a large value (i.e., coarse granularity), the calculation speed of the digest information is fast, but the accuracy during matching is low, and the probability of file collision is higher; at the sampling granularity with a small value (i.e., fine granularity), the accuracy during matching is high, the probability of file collision is low, but the calculation speed of the digest information is slow. In the embodiments of the present application, multiple sampling granularities with different numerical sizes are configured for the file, comprehensively considering the file matching efficiency and the file matching accuracy rate, and can improve the file matching efficiency on the premise of ensuring the file matching accuracy rate.
[0122] The above content focuses on introducing the file matching method with comprehensive multi-sampling granularity. In addition to the file matching method with comprehensive multi-sampling granularity, the file matching method provided by the embodiments of the present application may also include a file matching method with single-sampling granularity. The file matching method with single-sampling granularity means that the digest information of a file is obtained by calculating the digest of the sampled file of the file at a single sampling granularity. When performing file matching, the file matching method of comparing the digest information of the file at a single sampling granularity can be understood as that the file matching method with single-sampling granularity is a special case of the file matching method with comprehensive multi-sampling granularity when M = 1. In the file matching method with single-sampling granularity, the single sampling granularity can be called the target sampling granularity. The digest information of the target file at the target sampling granularity can be calculated. The reference files included in the reference file set are the files in the file library that match the file name or file size of the target file. The digest information of any reference file in the reference file set is obtained by calculating the digest of the sampled file of the corresponding reference file at the target sampling granularity. If there is a reference file in the reference file set whose digest information matches (i.e., is the same as) the target file, it can be determined that there is a reference file in the reference file set that matches the target file, and further a matching result indicating successful matching processing can be generated. If there is no reference file in the reference file set whose digest information matches (i.e., is the same as) the target file, it can be determined that there is no reference file in the reference file set that matches the target file, and further a matching result indicating failed matching processing can be generated.
[0123] In summary, regarding the file matching method with single-sampling granularity, the file matching method with single-sampling granularity compares the digest information of the file at a single sampling granularity. Therefore, the advantage of the file matching method with single-sampling granularity in terms of file matching efficiency is very obvious. However, the improvement of the file matching efficiency by the file matching method with single-sampling granularity is at the cost of the loss of the accuracy of the digest information representation, and the probability of file collision in the file matching method with single-sampling granularity is very high. By comparing the file matching methods with multi-sampling granularity and pruning processing, the file matching method with comprehensive multi-sampling granularity, and the file matching method with single-sampling granularity, it is not difficult to find that the file matching methods with multi-sampling granularity and pruning processing, and the file matching method with comprehensive multi-sampling granularity make up for the lack of the file matching method with single-sampling granularity in terms of file matching accuracy.
[0124] Generally speaking, the file matching method provided in the embodiments of the present application may include any one of a file matching method with multiple sampling granularities and pruning processing, a file matching method integrating multiple sampling granularities, and a file matching method with a single sampling granularity; the embodiments of the present application may select any one of the file matching methods according to actual requirements for file matching processing: for example, the file matching method may be selected according to the file size. When the file size of the target file is greater than the size threshold, a file matching method with multiple sampling granularities and pruning processing may be selected. When the file size of the target file is less than or equal to the size threshold, a file matching method integrating multiple sampling granularities may be selected; another example is that the file matching method may be selected according to the file processing requirements. For file processing requirements with high file matching efficiency, a file matching method with a single sampling granularity may be selected. For file processing requirements with high file matching accuracy, a file matching method with multiple sampling granularities and pruning processing may be selected, or a file matching method integrating multiple sampling granularities may be selected.
[0125] S709. After obtaining the matching result of the target file, perform service processing on the target file according to the matching result.
[0126] After obtaining the matching result of the target file, service processing may be performed on the target file according to the matching result of the target file. The matching result may include a matching result indicating successful matching processing or a matching result indicating failed matching processing.
[0127] The service processing may include file status synchronization processing, which refers to the processing of keeping the file backup status in a computer device (i.e., a terminal) synchronized with the file backup status in the cloud. Specifically, in the service processing scenario of file status synchronization processing, the target file is the file in the terminal whose backup status needs to be identified. Each reference file in the reference file set has been backed up to the cloud. The reference file set is the file in the cloud database (i.e., the file library mentioned above) that matches the file name or file size of the target file. Performing service processing on the target file according to the matching result may include: if the matching result indicates successful matching processing, it means that the target file has been backed up to the cloud, and then the target file may be marked with a backed-up identifier; if the matching result indicates failed matching processing, it means that the target file has not been backed up to the cloud, and then the target file may be marked with a non-backed-up identifier; as Figure 10 shown in the schematic diagram of the relevant interface in the service processing scenario of file status synchronization processing, the file named "XX-1 file" is marked with a non-backed-up identifier 1001 (for example Figure 10 "not backed up" in), and the file named "XX-2 file" is marked with a backed-up identifier 1002 (for example Figure 10The file named "XX-3 File" is marked with the "Backed up" mark in []. In addition, after marking the target file with the "Not backed up" mark, if a backup operation on the target file is detected, the target file can be backed up to the cloud for storage in response to the backup operation on the target file. By performing file status synchronization processing on the matching results obtained through the file matching method provided in the embodiments of this application, not only can the accuracy of file status synchronization be ensured, but also the efficiency of file status synchronization can be improved.
[0128] Alternatively, the service processing may include file backup processing, which refers to the process of immediately backing up a file to the cloud when it is detected that the file has not been backed up to the cloud. Specifically, in the service processing scenario of file backup processing, the target file is the file in the terminal whose backup status needs to be identified. Each reference file in the reference file set has been backed up to the cloud. The reference file set is the file in the cloud database (i.e., the file library mentioned above) that matches the file name or file volume of the target file. Performing service processing on the target file according to the matching result may include: if the matching result indicates that the matching process is successful, it means that the target file has been backed up to the cloud, and no backup processing is performed on the target file; if the matching result indicates that the matching process fails, it means that the target file has not been backed up to the cloud, and the target file can be immediately backed up to the cloud for storage. Through the matching result obtained by the file matching method provided in the embodiments of this application, file backup processing can be accurately and efficiently performed.
[0129] Alternatively, the service processing may include file moving processing, which refers to the process of moving a file to the corresponding folder. Specifically, in the service processing scenario of file moving processing, the target file is the file currently being moved. The moving operation can be used to move the target file to the target folder. The reference file set is stored in the target folder (i.e., the file library mentioned above). The reference file set is the file in the target folder that matches the file name or file volume of the target file. Performing service processing on the target file according to the matching result may include: if the matching result indicates that the matching process is successful, the moving operation can be rejected to prevent the target file from being moved to the target folder; if the matching result indicates that the matching process fails, the moving operation can be accepted to move the target file to the target folder. Through the matching result obtained by the file matching method provided in the embodiments of this application, file moving processing can be accurately and efficiently performed.
[0130] It should be noted that the abstract information of each reference file in the reference file set in the embodiments of the present application can be directly obtained. For example, in the business processing scenarios of file status synchronization processing and file backup processing, every time a file is backed up to the cloud, the cloud will associate and store the abstract information of the file as the unique identifier of the file. That is to say, the cloud stores the abstract information of each reference file in the reference file set. When file matching processing is required, the terminal can directly obtain the abstract information of each reference file in the reference file set from the cloud. Similarly, in the file movement scenario, every time the terminal stores a file, it will associate and store the abstract information of the file as the unique identifier of the file. That is to say, the terminal stores the abstract information of each reference file in the reference file set. When file matching processing is required, the abstract information of each reference file in the reference file set can be directly obtained from the terminal. In this way, the file matching efficiency can be further improved. Or, the abstract information of each reference file in the reference file set in the embodiments of the present application is obtained by performing file block sampling processing and abstract calculation processing after obtaining the reference file set. For example, in the business processing scenarios of file status synchronization processing and file backup processing, when file matching processing is required, the reference file set can be obtained from the cloud, and then the terminal can perform file block sampling processing and abstract calculation on each reference file in the reference file set to obtain the abstract information of each reference file in the reference file set. Similarly, in the file movement scenario, when file matching processing is required, the reference file set can be obtained from the terminal, and then file block sampling processing and abstract calculation can be performed on each reference file in the reference file set to obtain the abstract information of each reference file in the reference file set. In this way, the file matching accuracy can be further improved.
[0131] In the embodiments of the present application, for the file matching method with comprehensive multi-sampling granularity, during the file matching process, the abstract calculation results at each sampling granularity can be compared, and the advantage in file matching accuracy is relatively obvious. For the file matching method with single sampling granularity, during the file matching process, the abstract calculation result at a single sampling granularity is compared, and the advantage in file matching efficiency is relatively obvious. The embodiments of the present application can select different file matching methods for file processing according to different file processing requirements, and the file processing methods are flexible and diverse. In addition, the file processing methods provided in the embodiments of the present application can be applied to file business processing scenarios such as file status synchronization processing, file backup processing, and file synchronization processing, and can accurately and efficiently perform file business processing.
[0132] The above Figures 2 - 10The illustrated embodiments elaborate on the method of the embodiments of the present application. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, a file processing system of the embodiments of the present application is provided below.
[0133] As Figure 11 shown in the file processing system, the file processing system may include a terminal 1101 and a server 1102 (or the file processing system may include a terminal 1101 and not include a server 1102, Figure 11 taking the file processing system including a terminal 1101 and a server 1102 as an example); the server mentioned in the embodiments of the present application may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services; the terminal mentioned in the embodiments of the present application may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, a vehicle-mounted terminal, a smart voice interaction device, and a smart home appliance, etc., but is not limited thereto; the terminal and the server may be directly connected by a wired communication method, or may be indirectly connected by a wireless communication method, and the present application does not make any restrictions here.
[0134] In the file service processing scenario of the above file status synchronization processing or file backup processing, the file processing system may include a terminal and a server. The file is backed up to the server (i.e., the cloud), that is, the server stores the backed-up file. When there is a need for file status synchronization processing or file backup processing of the target file in the terminal, the terminal can send a file processing request to the server. The file processing request may carry the file name or file volume of the target file. When the server receives the file processing request sent by the terminal, in response to the file processing request, it can search for the file that matches the file name or file volume of the target file from the database of the server (i.e., the file library mentioned above), and determine the reference file set according to the found file. If the server stores the digest information of each reference file in the reference file set, the server can send the digest information of each reference file in the reference file set to the terminal. If the server does not store the digest information of each reference file in the reference file set, the server can obtain the digest information of each reference file in the reference file set through file block sampling processing and digest calculation, and send the digest information of each reference file in the reference file set to the terminal. Or, if the server does not store the digest information of each reference file in the reference file set, the server can send each reference file in the reference file set to the terminal, and the terminal can obtain the digest information of each reference file in the reference file set through file block sampling processing and digest calculation. Then, the terminal can perform a matching process on the target file in the reference file set by using any one of the file matching methods of multi-sampling granularity and pruning processing, comprehensive multi-sampling granularity file matching method, and single-sampling granularity file matching method based on the digest information of the target file and the digest information of each reference file in the reference file set. After obtaining the matching result of the target file, the terminal can perform file status synchronization processing or file backup processing on the target file according to the matching result of the target file.
[0135] In the file service processing scenario of the above file movement processing, the file processing system may include a terminal, and the file is stored in the terminal. When there is an operation requirement to move a target file to a target folder in the terminal, the terminal may search for a file in the target folder (i.e., the file library mentioned above) that matches the file name or file size of the target file, and determine a reference file set based on the found file. If the terminal stores the digest information of each reference file in the reference file set, the terminal may, based on the digest information of the target file and the digest information of each reference file in the reference file set, adopt any one of the file matching methods of multi-sampling granularity and pruning processing, comprehensive multi-sampling granularity, and single-sampling granularity to perform matching processing on the target file in the reference file set. After obtaining the matching result of the target file, the terminal may perform file movement processing on the target file according to the matching result. If the terminal does not store the digest information of each reference file in the reference file set, the terminal may obtain the digest information of each reference file in the reference file set through file block sampling processing and digest calculation. Then, the terminal may, based on the digest information of the target file and the digest information of each reference file in the reference file set, adopt any one of the file matching methods of multi-sampling granularity and pruning processing, comprehensive multi-sampling granularity, and single-sampling granularity to perform matching processing on the target file in the reference file set. After obtaining the matching result of the target file, the terminal may perform file movement processing on the target file according to the matching result.
[0136] Under the system architecture of the file processing system provided in the embodiments of the present application, by adopting the file matching method provided in the embodiments of the present application, business processing can be efficiently and accurately performed in file service processing scenarios such as file status synchronization processing, file backup processing, and file movement processing.
[0137] It can be understood that the file processing system described in the embodiments of the present application is for more clearly explaining the technical solutions of the embodiments of the present application, and does not constitute a limitation to the technical solutions provided in the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the blockchain network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.
[0138] It should be noted that the file processing method provided in the embodiments of the present application may be combined with cloud computing technology and cloud storage technology in cloud basic technologies.
[0139] Among them:
[0140] Cloud computing technology is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". Cloud computing technology refers to the delivery and usage model of IT (Internet Technology) infrastructure, which means obtaining the required resources in a on-demand and easily scalable manner through the network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services in a on-demand and easily scalable manner through the network. Such services can be related to IT and software, the Internet, or other services. Cloud computing technology is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing. In the embodiments of this application, cloud computing technology can be adopted to accelerate the file block sampling process, digest calculation, and matching process for the target file when performing file block sampling processing, digest calculation, and matching process for the target file, thereby further improving the file matching efficiency.
[0141] Cloud storage technology is a new concept extended and developed from the concept of cloud computing technology. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through functions such as cluster applications, grid technology, and distributed storage file systems, and collaborates through application software or application interfaces to jointly provide data storage and business access functions to the outside world. Currently, the storage method of the storage system is as follows: create a logical volume. When creating a logical volume, physical storage space is allocated for each logical volume, and this physical storage space may be composed of the disks of a certain storage device or several storage devices. The client stores data on a certain logical volume, that is, stores the data on the file system. The file system divides the data into many parts, and each part is an object. The object not only contains data but also contains additional information such as data identity (Identity, ID). The file system writes each object into the physical storage space of the logical volume respectively, and the file system will record the storage location information of each object. Thus, when the client requests to access the data, the file system can enable the client to access the data according to the storage location information of each object. The server (i.e., the cloud) of the embodiment of the present application or the terminal of the embodiment of the present application can adopt cloud storage technology to store files and the summary information of the files, so as to improve the storage efficiency of the files and the file summary information.
[0142] The above has elaborated in detail the method and system of the embodiment of the present application. In order to facilitate better implementation of the above solutions of the embodiment of the present application, correspondingly, the device of the embodiment of the present application is provided below.
[0143] Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of a file processing device provided by an embodiment of the present application. The file processing device can be set in the computer device provided by the embodiment of the present application, and the computer device can be the terminal mentioned above. Figure 12 The file processing device shown can be a computer program (including program code) running in a computer device. The file processing device can be used to execute Figure 2 or Figure 7 part or all of the steps in the method embodiment shown. Please refer to Figure 12 , and the file processing device can include the following units:
[0144] An obtaining unit 1201, configured to obtain a target file to be processed. The target file includes a plurality of file blocks, and the target file is configured with M sampling granularities, where M is a positive integer;
[0145] A processing unit 1202 is configured to perform file block sampling on a target file at at least one of M sampling granularities, and calculate a digest of the sampled file obtained by the file block sampling to obtain the digest information of the target file.
[0146] An acquisition unit 1201 is configured to acquire the digest information of each reference file in a reference file set used for matching the target file.
[0147] The processing unit 1202 is further configured to perform a matching process on the target file in the reference file set based on the digest information of each reference file and the digest information of the target file.
[0148] The processing unit 1202 is further configured to perform a service process on the target file according to the matching result after obtaining the matching result of the target file.
[0149] In one implementation, the M sampling granularities are arranged in order. After acquiring the target file to be processed, the processing unit 1202 is further configured to perform the following steps:
[0150] Traverse the M sampling granularities in sequence.
[0151] Wherein, each time a sampling granularity is traversed, a file matching process is triggered to be executed. One file matching process includes the following operations: file block sampling, digest calculation, acquiring the digest information of each reference file, and performing a matching process on the target file.
[0152] After executing the file matching process at the currently traversed sampling granularity, the processing unit 1202 is further configured to perform the following steps:
[0153] If it is determined through the file matching process at the currently traversed sampling granularity that there is no reference file in the reference file set that matches the target file, pruning processing is performed on the sampling granularities in the M sampling granularities that have not been traversed to stop traversing the M sampling granularities, and a matching result indicating that the matching process fails is generated.
[0154] In one implementation, after executing the file matching process at the currently traversed sampling granularity, the processing unit 1202 is further configured to perform the following steps:
[0155] If it is determined through the file matching process at the currently traversed sampling granularity that there is a reference file in the reference file set that matches the target file, the traversal progress of the M sampling granularities is detected.
[0156] If the detected traversal progress indicates that all the M sampling granularities have been traversed, a matching result indicating that the matching process is successful is generated.
[0157] If the detected traversal progress indicates that M sampling granularities have not been traversed, the file matching process for the next sampling granularity of the currently traversed sampling granularity is executed.
[0158] In one implementation, in the file matching process for the currently traversed sampling granularity, the digest information of the target file is obtained by calculating the digest of the sampled files of the target file at the currently traversed sampling granularity; the digest information of any reference file is obtained by calculating the digest of the sampled files of the corresponding reference file at the currently traversed sampling granularity.
[0159] For the file matching process at the currently traversed sampling granularity, the processing unit 1202, when used to perform matching processing on the target file in the reference file set based on the digest information of each reference file and the digest information of the target file, is specifically used to perform the following steps:
[0160] If there is a reference file in the reference file set whose digest information matches the digest information of the target file, it is determined that there is a reference file in the reference file set that matches the target file;
[0161] If there is no reference file in the reference file set whose digest information matches the digest information of the target file, it is determined that there is no reference file in the reference file set that matches the target file.
[0162] In one implementation, in the file matching process for the currently traversed sampling granularity, if the currently traversed sampling granularity ranks first among the M sampling granularities, the reference file set includes: files in the file library that match the file name or file volume of the target file;
[0163] If the currently traversed sampling granularity ranks non-first among the M sampling granularities, the reference file set includes: reference files determined to match the target file through the file matching process for the previous sampling granularity of the currently traversed sampling granularity.
[0164] In one implementation, the processing unit 1202, when used to perform file block sampling processing on the target file at at least one of the M sampling granularities, is specifically used to perform the following steps:
[0165] At each of the M sampling granularities, file block sampling processing is respectively performed on the target file to obtain M sampled files of the target file, with one sampled file corresponding to one sampling granularity;
[0166] Among them, the digest information of the target file includes: the digest calculation results of each of the M sampled files of the target file; the digest information of any reference file includes: the digest calculation results of the sampled files of the corresponding reference file at each sampling granularity.
[0167] In one implementation, when the processing unit 1202 is used to perform matching processing on the target file in the reference file set based on the summary information of each reference file and the summary information of the target file, it is specifically used to execute the following steps:
[0168] For any reference file in the reference file set, if the summary calculation result in the summary information of any reference file matches the summary calculation result in the summary information of the target file according to the sampling granularity, it is determined that any reference file matches the target file;
[0169] If there is a summary calculation result in the summary information of any reference file and the summary information of the target file that does not match according to the sampling granularity, it is determined that the target file does not match any reference file;
[0170] If there is a reference file in the reference file set that matches the target file, a matching result indicating successful matching processing is generated;
[0171] If there is no reference file in the reference file set that matches the target file, a matching result indicating failed matching processing is generated.
[0172] In one implementation, for the i-th sampling granularity among the M sampling granularities, when the processing unit 1202 is used to perform file block sampling processing on the target file at at least one of the M sampling granularities, it is specifically used to execute the following steps:
[0173] Determine the sampling interval corresponding to the i-th sampling granularity;
[0174] According to the sampling interval corresponding to the i-th sampling granularity, perform file block sampling processing on the target file to obtain the sampled file blocks of the target file at the i-th sampling granularity;
[0175] Perform splicing processing on the sampled file blocks of the target file at the i-th sampling granularity to obtain the sampled file of the target file at the i-th sampling granularity, where i is a positive integer less than or equal to M.
[0176] In one implementation, the reference file set has been backed up to the cloud, and the target file is a file to be identified for the backup status; the service processing includes file status synchronization processing; when the processing unit 1202 is used to perform service processing on the target file according to the matching result, it is specifically used to execute the following steps:
[0177] If the matching result indicates failed matching processing, mark the target file with an unbacked-up status identifier;
[0178] If the matching result indicates successful matching processing, mark the target file with a backed-up status identifier.
[0179] In one implementation, the target file is the file currently being moved, and the move operation is used to move the target file to the target folder; the reference file set is stored in the target folder; the service processing includes file move processing; when the processing unit 1202 is used to perform service processing on the target file according to the matching result, it is specifically used to perform the following steps:
[0180] If the matching result indicates that the matching process is successful, the move operation is rejected to prevent the target file from being moved to the target folder;
[0181] If the matching result indicates that the matching process fails, the move operation is accepted and the target file is moved to the target folder.
[0182] According to an embodiment of the present application, Figure 2 or Figure 7 The method steps involved in the method embodiments shown may be executed by each unit in the file processing device shown. For example, Figure 12 The steps S201 and S204 shown in may be executed by the acquisition unit 1201 shown in, Figure 2 The steps S202 - S203, and the steps S205 - S206 shown in may be executed by the processing unit 1202 shown in. Another example, Figure 12 The steps S201 and S204 shown in may be executed by the acquisition unit 1201 shown in, Figure 2 The steps S202 - S203, and the steps S205 - S206 shown in may be executed by the processing unit 1202 shown in. Another example, Figure 12 The steps S201 and S204 shown in may be executed by the acquisition unit 1201 shown in, Figure 7 The steps S701 and S704 shown in may be executed by the acquisition unit 1201 shown in, Figure 12 The steps S701 and S704 shown in may be executed by the acquisition unit 1201 shown in, Figure 7 The steps S702 - S703, and the steps S705 - S709 shown in may be executed by the processing unit 1202 shown in. Figure 12 The steps S702 - S703, and the steps S705 - S709 shown in may be executed by the processing unit 1202 shown in.
[0183] According to another embodiment of the present application, Figure 12 Each unit in the file processing device shown can be separately or all combined into one or several other units to form, or some of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the file processing device may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by multiple units collaborating.
[0184] According to another embodiment of the present application, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method embodiments shown in Figure 2 or Figure 7 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), a read-only storage medium (ROM), etc., to construct a file processing device as shown in Figure 12 and to implement the file processing method of the embodiments of the present application. The computer program can be recorded on a computer-readable storage medium, for example, and loaded into the above computing device through the computer-readable storage medium and run therein.
[0185] In the embodiments of the present application, for a target file to be processed, the target file may include multiple file blocks, the target file may be configured with one or more sampling granularities, and the target file can be sampled at least one of the one or more sampling granularities, and a digest calculation is performed on the sampled file obtained by the sampling process to obtain the digest information of the target file. Then, based on the digest information of the target file and the digest information of each reference file in the reference file set, a matching process is performed on the target file in the reference file set. After obtaining the matching result of the target file, the target file can be processed according to the matching result; in the embodiments of the present application, by configuring one or more sampling granularities to perform file block sampling on the target file, the file volume of the sampled file obtained by the file block sampling process is smaller than the file volume of the target file, so that the calculation speed of the file digest information can be improved and the file matching efficiency can be enhanced; moreover, when the number of configured sampling granularities is greater than or equal to two, the file can be matched at multiple sampling granularities, and the accuracy of the file matching result can also be enhanced.
[0186] Based on the above method, system, and device embodiments, the embodiments of the present application provide a computer device, and the computer device can be the terminal mentioned above. Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a computer device provided by the embodiments of the present application. Figure 13 The computer device shown at least includes a processor 1301, an input interface 1302, an output interface 1303, and a computer-readable storage medium 1304. Among them, the processor 1301, the input interface 1302, the output interface 1303, and the computer-readable storage medium 1304 can be connected through a bus or other means.
[0187] The input interface 1302 can be used to obtain a target file, the M sampling granularities configured for the target file, each reference file in the reference file set, or the summary information of each reference file in the reference file set. The output interface 1303 can be used to output the service processing result of the target file (for example, output a backup flag, or output a non-backup flag, etc.).
[0188] The computer-readable storage medium 1304 can be stored in the memory of the computer device. The computer-readable storage medium 1304 is used to store a computer program. The computer program includes computer instructions. The processor 1301 is used to execute the program instructions stored in the computer-readable storage medium 1304. The processor 1301 (or CPU (Central Processing Unit, central processing unit)) is the computing core and control core of the computer device. It is suitable for implementing one or more computer instructions, and is specifically suitable for loading and executing one or more computer instructions to implement the corresponding method flow or corresponding function.
[0189] The embodiment of the present application also provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device, and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device, and of course can also include the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the computer device is stored in this storage space. And, one or more computer instructions suitable for being loaded and executed by the processor are also stored in this storage space. These computer instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0190] One or more computer instructions stored in the computer-readable storage medium 1304 can be loaded and executed by the processor 1301 to implement the above related Figure 2 or Figure 7 The corresponding steps of the vehicle navigation method shown. In a specific implementation, the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 as follows:
[0191] Obtain a target file to be processed. The target file includes multiple file blocks. The target file is configured with M sampling granularities, and M is a positive integer;
[0192] At at least one of the M sampling granularities, perform file block sampling on the target file, and calculate the digest of the sampled file obtained by the file block sampling to obtain the digest information of the target file;
[0193] Obtain the digest information of each reference file in the reference file set used for matching the target file;
[0194] Based on the digest information of each reference file and the digest information of the target file, perform matching processing on the target file in the reference file set;
[0195] After obtaining the matching result of the target file, perform service processing on the target file according to the matching result.
[0196] In one implementation, the M sampling granularities are arranged in order. After obtaining the target file to be processed, the computer instructions in the computer-readable storage medium 1304 are loaded by the processor 1301 and are also used to execute the following steps:
[0197] Traverse the M sampling granularities in sequence;
[0198] Wherein, every time a sampling granularity is traversed, a file matching process is triggered to execute. A file matching process includes the following operations: file block sampling, digest calculation, obtaining the digest information of each reference file, and performing matching processing on the target file;
[0199] After executing the file matching process under the currently traversed sampling granularity, the computer instructions in the computer-readable storage medium 1304 are loaded by the processor 1301 and are also used to execute the following steps:
[0200] If it is determined through the file matching process under the currently traversed sampling granularity that there is no reference file in the reference file set that matches the target file, pruning processing is performed on the sampling granularities in the M sampling granularities that have not been traversed to stop traversing the M sampling granularities, and a matching result indicating that the matching process fails is generated.
[0201] In one implementation, after executing the file matching process under the currently traversed sampling granularity, the computer instructions in the computer-readable storage medium 1304 are loaded by the processor 1301 and are also used to execute the following steps:
[0202] If it is determined through the file matching process under the currently traversed sampling granularity that there is a reference file in the reference file set that matches the target file, detect the traversal progress of the M sampling granularities;
[0203] If the detected traversal progress indicates that all the M sampling granularities have been traversed, generate a matching result indicating that the matching process is successful;
[0204] If the detected traversal progress indicates that M sampling granularities have not been traversed, then execute the file matching process for the next sampling granularity of the currently traversed sampling granularity.
[0205] In one implementation, in the file matching process for the currently traversed sampling granularity, the digest information of the target file is obtained by calculating the digest of the sampled files of the target file at the currently traversed sampling granularity; the digest information of any reference file is obtained by calculating the digest of the sampled files of the corresponding reference file at the currently traversed sampling granularity.
[0206] For the file matching process at the currently traversed sampling granularity, the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301. When performing the matching process for the target file in the reference file set based on the digest information of each reference file and the digest information of the target file, it is specifically used to execute the following steps:
[0207] If there is a reference file in the reference file set whose digest information matches the digest information of the target file, then it is determined that there is a reference file in the reference file set that matches the target file.
[0208] If there is no reference file in the reference file set whose digest information matches the digest information of the target file, then it is determined that there is no reference file in the reference file set that matches the target file.
[0209] In one implementation, in the file matching process for the currently traversed sampling granularity, if the currently traversed sampling granularity is ranked first among the M sampling granularities, then the reference file set includes: files in the file library that match the file name or file volume of the target file.
[0210] If the currently traversed sampling granularity is ranked non-first among the M sampling granularities, then the reference file set includes: reference files that match the target file determined by the file matching process at the previous sampling granularity of the currently traversed sampling granularity.
[0211] In one implementation, the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301. When performing file block sampling processing on the target file at at least one sampling granularity among the M sampling granularities, it is specifically used to execute the following steps:
[0212] At each of the M sampling granularities, perform file block sampling processing on the target file respectively to obtain M sampled files of the target file, with one sampled file corresponding to one sampling granularity.
[0213] Among them, the summary information of the target file includes: the summary calculation results of each of the M sampled files of the target file; the summary information of any reference file includes: the summary calculation results of the sampled files of the corresponding reference file at each sampling granularity.
[0214] In one implementation, the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301. When performing matching processing on the target file in the reference file set based on the summary information of each reference file and the summary information of the target file, it is specifically used to execute the following steps:
[0215] For any reference file in the reference file set, if the summary calculation result in the summary information of any reference file matches the summary calculation result in the summary information of the target file according to the sampling granularity, it is determined that any reference file matches the target file;
[0216] If there is a summary calculation result in the summary information of any reference file and the summary information of the target file that does not match according to the sampling granularity, it is determined that the target file does not match any reference file;
[0217] If there is a reference file in the reference file set that matches the target file, a matching result indicating successful matching processing is generated;
[0218] If there is no reference file in the reference file set that matches the target file, a matching result indicating failed matching processing is generated.
[0219] In one implementation, for the i-th sampling granularity among the M sampling granularities, the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301. When performing file block sampling processing on the target file at at least one of the M sampling granularities, it is specifically used to execute the following steps:
[0220] Determine the sampling interval corresponding to the i-th sampling granularity;
[0221] According to the sampling interval corresponding to the i-th sampling granularity, perform file block sampling processing on the target file to obtain the sampled file blocks of the target file at the i-th sampling granularity;
[0222] Perform splicing processing on the sampled file blocks of the target file at the i-th sampling granularity to obtain the sampled file of the target file at the i-th sampling granularity, where i is a positive integer less than or equal to M.
[0223] In one implementation, the reference file set has been backed up to the cloud, and the target file is a file whose backup status is to be identified; the service processing includes file status synchronization processing; when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform service processing on the target file according to the matching result, it is specifically used to perform the following steps:
[0224] If the matching result indicates that the matching process fails, the target file is marked with an unbacked-up status flag;
[0225] If the matching result indicates that the matching process succeeds, the target file is marked with a backed-up status flag.
[0226] In one implementation, the target file is a file currently being moved, and the move operation is used to move the target file to a target folder; the reference file set is stored in the target folder; the service processing includes file move processing; when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform service processing on the target file according to the matching result, it is specifically used to perform the following steps:
[0227] If the matching result indicates that the matching process succeeds, the move operation is rejected to prevent the target file from being moved to the target folder;
[0228] If the matching result indicates that the matching process fails, the move operation is accepted and the target file is moved to the target folder.
[0229] In the embodiments of the present application, for the target file to be processed, the target file may include multiple file blocks, the target file may be configured with one or more sampling granularities, and the file block sampling process may be performed on the target file at at least one of the one or more sampling granularities, and the sampling file obtained by the sampling process is subjected to digest calculation to obtain the digest information of the target file. Then, based on the digest information of the target file and the digest information of each reference file in the reference file set, the target file is matched in the reference file set. After obtaining the matching result of the target file, the service processing may be performed on the target file according to the matching result; in the embodiments of the present application, by configuring one or more sampling granularities to perform file block sampling on the target file, the file volume of the sampling file obtained by the file block sampling process is smaller than the file volume of the target file, so that the calculation speed of the file digest information can be improved and the file matching efficiency can be enhanced; moreover, when the number of configured sampling granularities is greater than or equal to two, the file can be matched at multiple sampling granularities, and the accuracy of the file matching result can also be enhanced.
[0230] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the file processing method provided in the above various alternative manners.
[0231] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A file processing method, characterized in that, The method includes: Obtaining a target file to be processed, where the target file includes multiple file blocks, and the target file is configured with M sampling granularities, and M is a positive integer; Performing file block sampling processing on the target file at at least one of the M sampling granularities, and calculating a digest of the sampled file obtained by the file block sampling processing to obtain the digest information of the target file; the process of performing file block sampling processing on the target file at the i-th sampling granularity among the M sampling granularities includes: determining the sampling interval corresponding to the i-th sampling granularity; performing file block sampling processing on the target file according to the sampling interval corresponding to the i-th sampling granularity to obtain the sampled file blocks of the target file at the i-th sampling granularity; splicing the sampled file blocks of the target file at the i-th sampling granularity to obtain the sampled file of the target file at the i-th sampling granularity, where i is a positive integer less than or equal to M; Obtaining the digest information of each reference file in the reference file set for matching the target file, and performing matching processing on the target file in the reference file set based on the digest information of each reference file and the digest information of the target file; After obtaining the matching result of the target file, performing service processing on the target file according to the matching result.
2. The method according to claim 1, characterized in that, The M sampling granularities are arranged in order. After obtaining the target file to be processed, the method further includes: sequentially traversing the M sampling granularities; Wherein, each time a sampling granularity is traversed, a file matching process is triggered to execute once, and a file matching process includes the following operations: file block sampling processing, digest calculation, obtaining the digest information of each reference file, and performing matching processing on the target file; After executing the file matching process at the currently traversed sampling granularity, the method further includes: if it is determined through the file matching process at the currently traversed sampling granularity that there is no reference file in the reference file set that matches the target file, pruning the sampling granularities among the M sampling granularities that have not been traversed to stop traversing the M sampling granularities, and generating a matching result indicating that the matching process fails.
3. The method according to claim 2, wherein After executing the file matching process at the currently traversed sampling granularity, the method further includes: If it is determined through the file matching process at the currently traversed sampling granularity that there is a reference file in the reference file set that matches the target file, detecting the traversal progress of the M sampling granularities; If the detected traversal progress indicates that all of the M sampling granularities have been traversed, generating a matching result indicating that the matching process is successful; If the detected traversal progress indicates that the M sampling granularities have not been traversed completely, executing the file matching process at the next sampling granularity of the currently traversed sampling granularity.
4. The method according to claim 2, wherein In the file matching process at the currently traversed sampling granularity, the digest information of the target file is obtained by calculating the digest of the sampled file of the target file at the currently traversed sampling granularity; the digest information of any reference file is obtained by calculating the digest of the sampled file of the corresponding reference file at the currently traversed sampling granularity; For the file matching process at the currently traversed sampling granularity, based on the digest information of each reference file and the digest information of the target file, performing matching processing on the target file in the reference file set includes: If there is a reference file in the reference file set whose digest information matches the digest information of the target file, it is determined that there is a reference file in the reference file set that matches the target file; If there is no reference file in the reference file set whose digest information matches the digest information of the target file, it is determined that there is no reference file in the reference file set that matches the target file.
5. The method according to claim 2, wherein In the file matching process at the currently traversed sampling granularity, if the currently traversed sampling granularity ranks first among the M sampling granularities, the reference file set includes: files in the file library that match the file name or file volume of the target file; If the currently traversed sampling granularity ranks non-first among the M sampling granularities, the reference file set includes: reference files determined to match the target file through the file matching process at the previous sampling granularity of the currently traversed sampling granularity.
6. The method according to claim 1, wherein Performing file block sampling processing on the target file at at least one sampling granularity among the M sampling granularities includes: At each of the M sampling granularities, performing file block sampling processing on the target file respectively to obtain M sampled files of the target file, with one sampled file corresponding to one sampling granularity; Among them, the digest information of the target file includes: the digest calculation results of each of the M sampled files of the target file; the digest information of any reference file includes: the digest calculation results of the sampled files of the corresponding reference file at each sampling granularity.
7. The method according to claim 6, characterized in that, Based on the digest information of each reference file and the digest information of the target file, performing matching processing on the target file in the reference file set includes: For any reference file in the reference file set, if the digest calculation result in the digest information of the any reference file matches the digest calculation result in the digest information of the target file according to the sampling granularity, it is determined that the any reference file matches the target file; If there is a digest calculation result in the digest information of the any reference file and the digest information of the target file that does not match according to the sampling granularity, it is determined that the target file does not match the any reference file; If there is a reference file in the reference file set that matches the target file, a matching result indicating successful matching processing is generated; If there is no reference file matching the target file in the reference file set, a matching result indicating the failure of the matching process is generated.
8. The method according to claim 1, wherein All of the reference file sets have been backed up to the cloud, and the target file is a file for which the backup status is to be identified; the service processing includes file status synchronization processing. The processing of the target file according to the matching result includes: If the matching result indicates the failure of the matching process, the target file is marked with an unbacked-up status identifier. If the matching result indicates the success of the matching process, the target file is marked with a backed-up status identifier.
9. The method according to claim 1, wherein, The target file is a file on which a moving operation is currently being performed, and the moving operation is used to move the target file to a target folder. The reference file set is stored in the target folder. The service processing includes file moving processing; the processing of the target file according to the matching result includes: If the matching result indicates the success of the matching process, the moving operation is rejected to prevent the target file from being moved to the target folder. If the matching result indicates the failure of the matching process, the moving operation is accepted, and the target file is moved to the target folder.
10. A file processing device, characterized in that, The file processing device includes: An acquisition unit configured to acquire a target file to be processed, the target file including a plurality of file blocks, and the target file being configured with M sampling granularities, where M is a positive integer. A processing unit configured to perform file block sampling processing on the target file at at least one of the M sampling granularities, and calculate a digest of the sampled file obtained by the file block sampling processing to obtain digest information of the target file; the process of performing file block sampling processing on the target file at the i-th sampling granularity among the M sampling granularities includes: determining a sampling interval corresponding to the i-th sampling granularity; performing file block sampling processing on the target file according to the sampling interval corresponding to the i-th sampling granularity to obtain a sampled file block of the target file at the i-th sampling granularity; and performing splicing processing on the sampled file blocks of the target file at the i-th sampling granularity to obtain a sampled file of the target file at the i-th sampling granularity, where i is a positive integer less than or equal to M. The acquisition unit is further configured to acquire digest information of each reference file in the reference file set used for matching the target file. The processing unit is further configured to perform matching processing on the target file in the reference file set based on the digest information of each reference file and the digest information of the target file. The processing unit is further configured to, after obtaining the matching result of the target file, perform service processing on the target file according to the matching result.
11. A computer device, characterized in that, The computer device includes: A processor adapted to implement a computer program. A computer-readable storage medium storing a computer program, the computer program being adapted to be loaded and executed by the processor to perform the file processing method according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by a processor to perform the file processing method according to any one of claims 1-9.
13. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by a processor, the file processing method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Data management method and system of distributed type file system
CN108052649A
File verification method and device, electronic equipment and computer readable storage medium
CN113468863A