File scanning method, electronic equipment, medium and product
By chunking the files and matching them with the cache library based on the data block attribute information, unscanned data blocks are determined and scanning tasks are generated, the high overhead problem caused by repeated scans in the storage system is solved, incremental virus scanning is achieved, and scanning efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510607168.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-12
AI Technical Summary
The scanning overhead caused by repeated scanning in existing storage systems is large, especially when there are a large number of duplicate similar file contents in massive files. The traditional scanning method takes a long time and takes up high computing resources. Small-scale updates still require full-text scanning, which lacks a fine incremental recognition mechanism.
The target file is processed into data blocks, and matches the data block cache library based on the attribute information of the data block (such as scan status, scan period and hash data), and determines that the data block is not scanned, and generates scanning tasks for incremental scanning to avoid repeated scanning.
Through the block processing and matching mechanism, the scanning overhead caused by repeated scanning is reduced, the storage and calculation overhead is reduced, incremental virus scanning is realized, and scanning efficiency and accuracy are improved.
Smart Images

Figure CN120470586A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a file scanning method, electronic equipment, medium, and product. Background Art
[0002] Conventional antivirus methods for storage systems scan at a file-level. Each scan requires a comprehensive scan of numerous files. This results in a large number of duplicate and similar files in the storage system, leading to repeated scans. Even small changes require rescanning the entire file, resulting in a long scan time. Repeated scanning of partial file content incurs significant scanning overhead.
[0003] Therefore, how to avoid repeated scanning to reduce scanning overhead is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0004] The present application provides a file scanning method, electronic device, medium and product to at least solve the problem of high scanning overhead caused by repeated scanning in the related art.
[0005] This application provides a file scanning method, including:
[0006] Obtain the target file to be scanned and the data block cache library;
[0007] Dividing the target file into blocks to obtain at least one data block and corresponding attribute information; wherein the attribute information at least includes a data block scanning status, a scanning period, and hash data of the data block;
[0008] Performing matching processing on at least one data block and a data block cache based on attribute information of at least one data block to determine an unscanned data block;
[0009] A scanning task is generated based on the unscanned data blocks for scanning processing.
[0010] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned file scanning methods when executing the computer program.
[0011] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned file scanning methods are implemented.
[0012] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned file scanning methods when executed by a processor.
[0013] The present application addresses the issue of increased scanning overhead due to the high repetitiveness of scanning, which occurs when conventional storage systems scan the entire file at file granularity. The present application divides the target file to be scanned into blocks to obtain at least one data block. Before the actual scan, the file granularity is reduced to the data block granularity, and the corresponding blocks are subsequently scanned. The data blocks are then matched against the data block cache based on their attribute information (including at least the data block scan status, scan duration, and hash data), identifying unscanned data blocks that do not have a matching match in the data cache. These unscanned data blocks are then subjected to the actual subsequent scanning process. When a file is modified or updated, the present application matches at least one block after the block processing with the data block cache to identify unscanned data blocks that have not been scanned repeatedly, thereby avoiding repeated scanning. Therefore, the high scanning overhead caused by repeated scanning can be reduced, thereby reducing storage overhead and the computational overhead of repeated scanning, achieving the technical effect of incremental virus scanning. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0015] Figure 1 A schematic diagram of antivirus for a storage system provided in an embodiment of the present application;
[0016] Figure 2 A flowchart of a file scanning method provided in an embodiment of the present application;
[0017] Figure 3 A schematic diagram of a file block processing method provided in an embodiment of the present application;
[0018] Figure 4 A schematic diagram of a scanning task generation provided in an embodiment of the present application;
[0019] Figure 5 A schematic diagram of a data block cache management library provided in an embodiment of the present application;
[0020] Figure 6 A schematic diagram of an actual scanning range corresponding to an overlapping area provided in an embodiment of the present application;
[0021] Figure 7 A schematic diagram of a document scanning method provided in an embodiment of the present application;
[0022] Figure 8A structural diagram of a document scanning device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0024] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0025] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0026] In modern network and storage environments, malicious programs such as viruses, Trojans, and ransomware are constantly emerging. To protect massive amounts of files, antivirus engines are often deployed to perform file scanning and threat detection. The core technologies of such engines include:
[0027] a) Virus signature matching;
[0028] 1) Most traditional antivirus engines rely on a "signature library" to identify malicious programs. The signature library contains a large number of "signatures" of viruses or malicious files, which are often expressed as a byte sequence or pattern.
[0029] 2) When the engine scans a file, it compares the file content byte by byte with the pattern in the signature library. If a match is successful, it is determined to be infected or suspicious.
[0030] b) Feature library maintenance;
[0031] Antivirus vendors regularly update or expand their signature databases (virus databases) to cover the latest discovered virus samples or variants.
[0032] In the deployed storage system, it is usually necessary to perform anti-virus protection on the huge amount of stored file contents. To achieve this function, the system is usually designed to deploy an anti-virus engine cluster (scanning software) to connect to the storage system. Figure 1 A schematic diagram of antivirus for a storage system provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the two use a specific communication protocol to transmit scan requests and related files to be scanned. The scan requests issued by the storage system are based on the file-level scanning unit, which has some disadvantages:
[0033] 1) Full scan overhead is high: When performing virus detection on massive files in large enterprises or cloud storage, each file must be repeatedly read and matched with the signature library, which is time-consuming and requires a large amount of data I / O and CPU computing.
[0034] 2) Failure to fully utilize duplicate content: There may be a large amount of duplicate or similar file content (such as different versions of the same file) in the storage system, but traditional scanning methods will still read and compare these identical parts multiple times.
[0035] 3) Small-scale updates lead to full-text rescanning: When only a very small number of bytes in a file are modified, traditional methods still require rescanning the entire file, lacking a more sophisticated incremental recognition mechanism.
[0036] The document scanning method provided in this application can solve the above technical problems.
[0037] Figure 2 A flowchart of a file scanning method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the method includes:
[0038] S11: Obtain the target file to be scanned and the data block cache library;
[0039] S12: Divide the target file into blocks to obtain at least one data block and corresponding attribute information; wherein the attribute information at least includes a data block scanning status, a scanning period, and hash data of the data block;
[0040] S13: performing matching processing on the at least one data block and the data block cache based on the attribute information of the at least one data block to determine the unscanned data blocks;
[0041] S14: Generate a scanning task according to the unscanned data blocks and perform scanning processing.
[0042] The target file to be scanned provided in this embodiment can be the target file to be scanned corresponding to the update event triggered when the storage system detects a file write operation or version update. The data block cache stores multiple scanned data blocks, each of which corresponds to the scan time, hash data, and the remaining scan period for subsequent scans.
[0043] The target file is segmented to obtain at least one data block. The segmentation operation here can split the file into blocks based on a fixed data block size and identify which data blocks have actually changed when the file content changes (write operation or version update), providing a basis for subsequent incremental scanning.
[0044] During the block processing, attribute information corresponding to the data block is also assigned. The attribute information here includes at least the data block scanning status, scanning period, and hash data of the data block. The scanning status refers to the possible scanning conditions of the current target file during the historical incremental scan, such as "UNSCANNED (unscanned status)", "CLEAN (clean status, i.e., non-infected status)", "INFECTED (infected status)", "ERROR (error status)", etc. The scanning period is based on whether the scanning period has been reached after the completion time of the last scan or whether the scanning cycle has expired. The hash data of the data block is determined based on a hash function. The specific hash function used is not limited in this application, but it needs to be the same as the hash function corresponding to the hash data obtained in the data block cache library to facilitate comparison in the subsequent matching process.
[0045] In step S13, at least one data block and the data block cache are matched based on the attribute information of at least one data block to determine the unscanned data block. In this embodiment, the at least one data block after the corresponding block processing is matched with the hash data corresponding to the stored data block in the data block cache. The matching strategy is to make a judgment based on the attribute information of the data block. The hash data of at least one data block is first judged and compared with the hash data in the data block cache. It should be noted that the data block cache in this embodiment takes into account that the data block occupies a certain amount of memory. Only three parameters, namely hash data, scan status and scan period, are stored in this library. The three parameters are stored in this embodiment using a mapping relationship. The judgment is first to compare the hash data. If they are the same, it is determined that the data block has been scanned. If they are not the same, it is determined that the data block corresponding to the target hash data that is different from the hash data in the data block cache is the unscanned data block. In addition, under the same circumstances, it is necessary to check the scanning status and scanning period corresponding to the same hash data. This is because when the file is first processed in this embodiment, the scanning status of the divided data blocks is defaulted to the non-scanning state, or when the file changes, the scanning status of the corresponding data block is changed to the non-scanning state. Matching is performed here to update the scanning status and scanning period corresponding to the data block, so as to determine the subsequent parameter management in the data block cache library, and whether the data block will be skipped or postponed for scanning. The conventional scanning process is to perform version and cache management at the file granularity, and directly perform hash cache management on the file during the scanning process. This embodiment first performs a management process of dividing the file into data blocks, and then performs subsequent hash cache management based on the data block hash data, the scanning status and the scanning period. This is significantly different from the conventional technical solution.
[0046] In step S14, the unscanned data blocks are generated into subsequent scanning tasks for scanning processing.
[0047] Through the embodiments of the present application, conventional storage systems scan the entire file at file granularity, and the scanning overhead is increased due to high scanning repetition. The present application divides the target file to be scanned into blocks to obtain at least one data block. Before the actual scan, the file granularity is reduced to the data block granularity, and the corresponding blocks are subsequently scanned. The data blocks are matched with the data block cache based on the data block's attribute information (including at least the data block scan status, scan duration, and data block hash data) to determine unscanned data blocks that do not have a corresponding match in the data cache. The unscanned data blocks are then subjected to the actual scanning process. When a file is changed or updated, the present application matches at least one block-processed data block with the data block cache to determine unscanned data blocks that have not been scanned repeatedly, thereby avoiding repeated scanning. Therefore, the high scanning overhead caused by repeated scanning can be reduced, thereby reducing storage overhead and the computational overhead of repeated scanning, and achieving the technical effect of incremental virus scanning.
[0048] In some embodiments, the target file is divided into blocks to obtain at least one data block and corresponding attribute information, including:
[0049] Get the preset data block size;
[0050] Divide the target file into blocks according to a preset data block size to obtain at least one data block;
[0051] Determine the starting position, actual length, and corresponding hash data of the data block in the target file according to the preset data block size;
[0052] Determine the scanning state corresponding to the data block and the scanning period corresponding to the last scanning completion time; wherein, the scanning state corresponding to the first block is an unscanned state and has no last scanning completion time.
[0053] Specifically, the preset data block size can be based on the data block size corresponding to daily scanning, or a set fixed data block size, etc., and is not limited here. The target file is divided into blocks according to the preset data block size to obtain at least one data block. The data block size here is the same as the preset data block size. Since there will be residual data such as the data block size of the last data block being smaller than the preset data block size during the block division process, the last data block can be discarded, or when the next target file arrives, it can be merged with the last data block of the next target file to obtain a new data block to participate in the subsequent matching process. This is not limited here and can be set according to actual conditions.
[0054] The data chunk index (Chunk_index) is set based on the position of at least one data chunk in the target file. This index corresponds to the sequential index of the data chunk in the target file, or is calculated based on the starting position of the data chunk in the file. The starting position (Offset), actual length (Length), and corresponding hash data (Hash_value) of the data chunk in the target file are determined based on the preset data chunk size. The starting position is the position of the first data chunk in the sequential position corresponding to the preset data chunk size. The actual length is the actual length of the data chunk, and the hash data is a hash value derived from the data in the data chunk.
[0055] Determine the scan status (Scan_status) of the data block and the scan deadline corresponding to the last scan completion time (Last_scan_time).
[0056] In addition, the latest version number and chunk list of each target file can also be recorded. The fields are: File_id, Current_version, Chunk_list (including several ChunkInfo objects).
[0057] This embodiment provides a method for performing block processing on a target file to obtain at least one data block and corresponding attribute information, which lays a foundation for subsequent matching processing. By matching the corresponding attribute information during the block processing, the block management capability of the data block is improved.
[0058] In some embodiments, when the target file is an updated file, the process of determining the hash data of the data block includes:
[0059] Get the change event flag of the target file;
[0060] Determine the changed data block area according to the starting position of the changed data block of the change event flag in the target file and the actual length;
[0061] In the target file, the target data block of the target file is updated with corresponding hash data according to the changed data block area.
[0062] Specifically, when the storage system detects a file write operation or version update, it triggers a change event for the target file to indicate which interval has been modified. Example fields include File_id, Offset, Length, and Timestamp.
[0063] When considering incremental file updates, the storage system captures a change event, indicating that a write operation occurred within a certain interval on the target file. Therefore, it is necessary to obtain the target file's change event flag (ChangeEvent). The changed data block area is determined based on the starting position and actual length of the changed data block indicated by the change event flag. The corresponding hash data is then updated based on all data blocks in the changed data block area to locate the update area based on the change event flag.
[0064] In some embodiments, determining the changed data block area according to the starting position and actual length of the changed data block of the change event flag in the target file includes:
[0065] Determine the starting data block according to the starting position and the preset data block size;
[0066] Determine the end data block according to the starting position, actual length and preset data block size;
[0067] The data blocks between the start data block and the end data block are used as the changed data block area.
[0068] Specifically, the calculation process is as follows:
[0069] The starting data block is: start_chunk=offset / / chunk_size;
[0070] The ending data block is: end_chunk=(offset+length-1) / / chunk_size.
[0071] The data blocks between the start data block and the end data block are used as the changed data block area.
[0072] Figure 3 A schematic diagram of a file block processing provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the target file obtained by reading or writing a file is stored in the storage system or file system. Once a change event or new file is detected, subsequent file segmentation is performed. For new files, initial segmentation is performed, followed by block-level splitting. When a modified block is detected, incremental monitoring is performed, updating the block's hash data and scan status, and subsequent segmentation and hash calculation are performed. A record represents a block of the target file and is organized as a list or data store to facilitate subsequent updates, queries, and maintenance. Finally, external module access processing, i.e., subsequent scanning, is performed.
[0073] This embodiment provides real-time updating of hash data of the changed data block area to facilitate subsequent matching processing. The real-time update here is based on the change mark of the file to locate the write operation interval, reducing the capture time caused by traversal and improving data processing efficiency.
[0074] In some embodiments, matching the at least one data block with the data block cache based on the attribute information of the at least one data block to determine the unscanned data block includes:
[0075] Obtaining a first target data block in a target file whose data block scanning status is an unscanned state;
[0076] Matching the first target data block in the data block cache to determine whether there is a data block marked as non-infected and not expired according to the hash data, the scan status, and the scan deadline of the first target data block;
[0077] If not, the first target data block is determined to be safe and is regarded as an unscanned data block.
[0078] The matching process in this embodiment primarily involves matching the attribute information of the first target data block with parameters stored in the data block cache to determine whether an identical data block exists. This matching process initially utilizes hash data. If a match is found, it is determined that the data block cache already stores the same data block corresponding to the hash data. An initial screening process can then be performed based on the attribute information (non-infected state) of the data blocks stored in the data block cache, followed by a secondary screening process corresponding to the scan period. Alternatively, an initial screening process based on the scan period can be performed, followed by a secondary screening process corresponding to the non-infected state. If no such data block exists, it is determined that the first target data block has not been scanned and is considered an unscanned data block.
[0079] The matching processing between the data block cache library corresponding to this embodiment and at least one data block uses the hash data, scanning status and scanning period of the data block for comparison, which improves the matching efficiency and the overall accuracy of the data block matching processing compared to only comparing the specific data corresponding to the data block.
[0080] In some embodiments, determining whether there is a data block marked as non-infected and not expired by matching in a data block cache according to the hash data, the scan status, and the scan deadline of the first target data block includes:
[0081] Obtaining first hash data of a first target data block;
[0082] Determining whether there is second hash data identical to the first hash data in the data block cache according to the first hash data;
[0083] If so, determining the second target data block corresponding to the second hash data;
[0084] Checking whether the corresponding unscanned state is a non-infected state according to the second target data block;
[0085] If so, determining whether the scan period of the second target data block has not expired;
[0086] If it is not expired, the second target data block is determined to be safe, the first target data block will be updated according to the non-infected state and the unexpired scanning period, and the scanning needs to be skipped or postponed;
[0087] If not, it is determined that the first target data is safe.
[0088] Specifically, the first hash data of the first target data block is obtained and compared with the hash data of data blocks in the data block cache to determine whether there is a second hash data that is identical to the first hash data. If so, the data block is conventionally considered to be the same as the first target data block. However, in this embodiment, the subsequent scan status and scan deadline are further compared to check whether the corresponding scan status is clean based on the second target data block. If so, the second target data block is further determined to be clean and unscanned, confirming that the first target data block can be scanned for subsequent tasks and the unscanned status of the first target data block needs to be updated. The scan deadline of the second target data block is then determined to be not expired. If so, the scan of the second target data block can be skipped or postponed. In this case, the scan status and scan deadline of the first target data block are required to participate in subsequent scan matching, but it cannot be scanned as the first target data block. If not, the first target data block is determined to be safe and is considered an unscanned data block.
[0089] This embodiment provides a comparison based on hash data, followed by a detailed comparison of the scan status and scan period to update the scan status and scan period of the first target data block, thereby facilitating subsequent scan matching and improving the accuracy of the matching process.
[0090] In some embodiments, obtaining a first target data block whose data block scanning status is an unscanned state includes:
[0091] Acquire a first initial target data block whose data block scanning status is an unscanned state;
[0092] Obtaining the non-expired time and scanning cycle time of the first initial target data block;
[0093] If the scanning cycle time is greater than the non-expired time, removing the second initial target data block whose scanning cycle time is greater than the non-expired time from the first initial target data block; and determining the remaining first initial target data block as the first target data block;
[0094] If the scanning cycle time is less than or equal to the non-expired time, all first initial target data blocks are used as first target data blocks.
[0095] Considering that a data block is in an unscanned state, but the corresponding non-expired time conflicts with the scanning cycle time, if the scanning cycle time is greater than the non-expired time, it means that the current data block is in an unscanned state before the scanning cycle time arrives, but will soon be in an expired state. In this case, the corresponding scanning state is not an unscanned state. In this way, even if it is subsequently screened out, there may be a certain risk that it will become a scanned state during periodic scanning, which in a certain sense increases the probability of repeated scanning. In this embodiment, the second initial target data block whose scanning cycle time is greater than the non-expired time needs to be removed from the first initial target data block to determine the remaining first initial target data block as the first target data block.
[0096] When the scanning cycle time is less than or equal to the non-expired time, all first initial target data blocks are used as first target data blocks.
[0097] The comparison between the scanning cycle time and the non-expired time provided in this embodiment further improves the accuracy of the scanning task and reduces the occurrence of unnecessary repeated scanning.
[0098] Figure 4 A schematic diagram of a scanning task generation provided in an embodiment of the present application is shown as follows: Figure 4 As shown in the figure, after the file is segmented, the unscanned state or other scan status of the data blocks that have changed is checked. A corresponding scan matching process is required, which compares the hash data, scan status, and scan expiration in the data block cache. For each changed data block, the following steps are performed: 1. Check the corresponding hash data; 2. If it is safe and not expired, skip the scan; 3. Otherwise, generate a scan task. Generating or skipping a scan task requires generating a scan object to be delivered to the task queue or scheduling module for distribution to the scanning nodes.
[0099] In some embodiments, after generating a scan task based on the unscanned data block and before performing the scan process, the method further includes:
[0100] Constructing scan objects based on unscanned data blocks;
[0101] Add scanning software priority and timestamp information to the scan object;
[0102] Determine scanning priority based on scanning software priority and timestamp information;
[0103] Scan objects according to their scan priority.
[0104] Specifically, a scanning object is constructed based on the unscanned data block, and scanning software priority and timestamp information are attached to the scanning object. The scanning software priority corresponds to which software is preferentially used for virus engine processing, and the timestamp information corresponds to when the scan is performed and the corresponding scanning time.
[0105] The final scanning priority is determined based on the scanning software priority and timestamp information, taking into account the priority scanning order used by different scanning software. In this embodiment, in order to improve the scanning processing efficiency, it is necessary to comprehensively consider the timestamp information to comprehensively utilize the sequence of scanning time to improve the scanning utilization rate of the scanning software.
[0106] In some embodiments, determining the scanning priority based on the scanning software priority and the timestamp information includes:
[0107] Set the scanning software priority to the first weight parameter;
[0108] Determining a second weight parameter according to the timestamp information;
[0109] The scanning priority is determined according to the scanning software priority, the first weight parameter, the timestamp information and the second weight parameter.
[0110] Specifically, the first weight parameter is set according to the priority of the scanning software to determine the order of importance of the software. The second weight parameter is determined according to the timestamp information, which can be based on the order of arrival of time, or based on the order of completion of the scanning time, etc., which can be used as a reference for setting the second weight parameter. It should be noted that the sum of the first weight parameters corresponding to the priority of each scanning software is 1, and the sum of the second weight parameters corresponding to its timestamp information is 1. The first weight parameters and the second weight parameters corresponding to the scanning software priority and the timestamp information are added together to obtain the final data of each software, and the final data are compared by data size to determine the final scanning priority. The data block can be scanned based on the final scanning priority.
[0111] This embodiment provides weight parameters corresponding to the scanning software priority and timestamp information, which are disrupted to comprehensively evaluate the scanning priority and improve the authority of the judgment of software priority use.
[0112] In some embodiments, the process of establishing a data block cache library includes:
[0113] Pre-write the data blocks in the target scanning task that are in the non-infected state or infected state into the data block cache;
[0114] When the scan status is non-infected, the sum of the current time and the validity period is used as the expiration time;
[0115] When the scanning status is infected, the corresponding data block will be recorded as permanent.
[0116] Specifically, during the creation and management of the data block cache, data blocks with a scan status of either clean or infected in the target scan task are pre-written to the data block cache. When the scan status is clean, the expiration time is the sum of the current time and the validity period, i.e., the expire_time is set to "current time + policy period." When the scan status is infected (INFECTED), the corresponding data block is recorded as permanent.
[0117] I. Write cache;
[0118] When a scan task is completed and the hash_value of the block is determined to be CLEAN or INFECTED, the result summary and arbitration module will call this module to write or update the HashCache.
[0119] If marked as CLEAN, the expire_time is set to "current time + policy period"; if marked as INFECTED, a longer or permanent status can be recorded.
[0120] II. Read and judge skip;
[0121] When the "Incremental Scan Task Generation Module" determines whether to scan, it reads HashCache[hash_value] and compares scan_result with expire_time.
[0122] If it is still valid, there is no need to scan again; if it has expired or the virus database has been significantly upgraded, rescanning is required.
[0123] III. Expired or invalid;
[0124] The system can periodically inspect the HashCache and clear or set entries to "UNKNOWN" if they have expired.
[0125] After the virus database is updated or the engine version is significantly upgraded, you can also set the expiration_time of all CLEAN status entries to 0 in batches to prompt subsequent rescanning.
[0126] IV. Remove duplicate meaning;
[0127] In scenarios where multiple files or multiple versions share the same block, as long as the hash value of a block is determined to be safe once, other blocks with the same hash value can be skipped directly.
[0128] Significantly reduces repeated calls to the feature scanning engine and improves overall processing efficiency.
[0129] Figure 5 A schematic diagram of a data block cache management library provided in an embodiment of the present application is shown as follows: Figure 5 As shown, a task is generated or the scan is skipped; the actual scan is sent to the antivirus engine or scanning node to obtain the scan results, record the scan results, and update the write. Notify the block or overall scan status to update the process.
[0130] The data block cache management process provided in this embodiment improves the management and processing efficiency of data blocks while also improving the subsequent matching efficiency of data blocks.
[0131] In some malicious files, virus signatures (characteristic strings) may be distributed at the intersection of two adjacent blocks. If only the independent content of each block is scanned, there is a theoretical risk of "missing cross-block signatures." In some embodiments, a scanning task is generated based on the unscanned data blocks for scanning, including:
[0132] Determine an overlapping area between a preceding unscanned data block and a succeeding unscanned data block between adjacent unscanned data blocks;
[0133] The starting position, data block length and overlapping area of the current unscanned data block are used as the actual scanning range of the unscanned data block;
[0134] Starting from the first actual scanning range, each unscanned data block and the data block corresponding to the actual scanning range are scanned.
[0135] Specifically, when defining blocks, a small overlap area with the previous (or next) block is added to each block, for example, 512B to 2KB. This way, if a virus signature happens to cross the original block boundary, it will be fully detected in this overlapped segment. A fixed block size, chunk_size (e.g., 1MB), is set, along with an additional overlap_size (e.g., 1KB). The actual scan range for block n is [offset_n, offset_n + chunk_size + overlap_size], where the overlap_size portion overlaps with the start of block (n+1). The overlap area between the previous and next unscanned blocks is determined between adjacent unscanned blocks. The starting position, block length, and overlap area of the current unscanned block are used as the actual scan range for the unscanned blocks. Starting from the first actual scan range, each unscanned block and the data blocks corresponding to the actual scan range are scanned. Feature continuity across block boundaries is guaranteed at the block level, eliminating the need for special adaptation for existing feature scanning engines.
[0136] In some embodiments, the actual scanning range is determined by:
[0137] The starting position of the currently unscanned data block is used as the first critical point of the actual scanning range;
[0138] The sum of the starting position of the currently unscanned data block, the length of the data block, and the length of the virus signature in the overlapping area is processed as a second critical point, wherein the second critical point is greater than the first critical point;
[0139] The range between the first critical point and the second critical point is taken as the actual scanning range.
[0140] Specifically, Figure 6 A schematic diagram of an actual scanning range corresponding to an overlapping area provided in an embodiment of the present application is shown as follows: Figure 6 As shown, the actual block scan range is determined by summing the starting position of the currently unscanned data block, the data block length, and the virus signature length of the overlapped area. An overlap_size (e.g., 1KB, 2KB) can be set in the system configuration, based on typical virus signature lengths. Because the "keyword string" length in most traditional antivirus engine signature libraries typically ranges from a few bytes to hundreds of bytes, setting an appropriately sized overlap effectively covers the vast majority of cross-block signatures.
[0141] This embodiment provides a method for determining the actual scanning range by taking into account that a virus signature may exist at the junction of two adjacent blocks, thereby avoiding omission of cross-block features and improving the security of scanning.
[0142] In many antivirus storage engines, file type identification is crucial for subsequent signature matching, decompression processing, and script analysis. If the storage system physically or logically divides a file into multiple data blocks and only passes one block to the engine, the engine may not have sufficient file context to determine the file type, thus affecting scanning accuracy.
[0143] In some embodiments, generating a scanning task based on the unscanned data block for scanning processing includes:
[0144] Obtain the file encoding information corresponding to the file to which the unscanned data block belongs;
[0145] Classify and scan based on file encoding information and corresponding unscanned data blocks.
[0146] Specifically, in most file type identifications, the key "magic number" or header signature is located in the tens to hundreds of bytes at the beginning of the file; even if block-level splitting is performed, the complete file header information can be contained in block 0, and the header content of the block can be passed to the engine during scanning or scheduling to allow it to identify the type; in this way, even if the subsequent file data is divided into blocks, it will not prevent the engine from knowing the type of the scanned file. Before performing a feature scan, the file header can be read in its entirety and passed to the engine for type judgment; then, an incremental scan can be performed on each block of the file. If type identification or deep analysis requires the entire file, the relevant agent in the storage system can merge the blocks in the background and present them to the engine, but this does not negate the advantages of incremental scanning in most scenarios.
[0147] The file header information identification provided by this embodiment through the key "magic number" of the file encoding information or the header signature facilitates the identification and processing of the file type, avoids the problem of long identification time caused by subsequent identification and processing of multiple data blocks, and improves the identification efficiency.
[0148] Figure 7 A schematic diagram of a document scanning provided in an embodiment of the present application is shown as follows: Figure 7 As shown, each file in the storage system is segmented and versioned. When a change occurs, the scan status and hash data in the data block cache are compared to determine the unscanned data blocks. Scan tasks are generated and then scheduled accordingly. Tasks are distributed so that the engines in each node can perform local scanning and obtain scan results.
[0149] During the file segmentation and version management process, create data structures:
[0150] I.ChunkInfo;
[0151] File_id: unique identifier of the file;
[0152] Chunk_index: the sequence number of the current block in the file, or calculated based on offset;
[0153] Offset: The starting position of the block in the file;
[0154] Length: the actual length of the block;
[0155] Hash_value: hash of the block content;
[0156] Scan_status: the scan status of the current block, such as "UNSCANNED", "CLEAN", "INFECTED", "ERROR", etc.
[0157] Last_scan_time: The time when the last scan was completed.
[0158] II.FileVersionTable;
[0159] Record the latest version number of "each file" and the corresponding block list. Example fields:
[0160] File_id;
[0161] Current_version;
[0162] Chunk_list: contains several ChunkInfo objects.
[0163] III.ChangeEvent;
[0164] When the storage system detects a file write operation or version update, a change event is triggered to indicate which interval has been modified. Example fields:
[0165] File_id;
[0166] Offset;
[0167] Length;
[0168] Timestamp.
[0169] The specific process is:
[0170] 1. Initialize the block;
[0171] When a new file is detected on the system, or when a file is included in the virus scan management for the first time:
[0172] The file is split into fixed chunk_sizes (e.g., 1MB, 4MB, etc.) based on the file_size. A hash value is calculated for each chunk; a ChunkInfo is created and recorded in the FileVersionTable. All chunks in the file are initially marked as UNSCANNED.
[0173] 2. Incremental update;
[0174] The storage system captures a ChangeEvent(file_id, offset, length), indicating that a write operation occurred on a file in the range [offset, offset+length]. The module calculates which blocks are involved in this range (from start_chunk=offset / / chunk_size to end_chunk=(offset+length-1) / / chunk_size), rereads the contents of these blocks, and calculates a new hash new_hash. If new_hash!=old_hash, the corresponding ChunkInfo.hash_value is updated and its scan_status is set to UNSCANNED. Otherwise, it remains unchanged. Finally, the updated block information is written back to the FileVersionTable.
[0175] During the generation of the incremental scan task, create the data structure:
[0176] ScanTask:
[0177] File_id;
[0178] Chunk_index;
[0179] Hash_value;
[0180] Priority;
[0181] Timestamp.
[0182] TaskQueue:
[0183] A list of scan tasks or message queues to be distributed;
[0184] For subsequent scheduling modules or nodes to extract.
[0185] HashCache (only used for query in this module, actual management is in the "Cache and Deduplication Module");
[0186] Its structure: hash_value->(scan_result, last_scan_time, expire_time, ...).
[0187] When the data blocks under the file change, the cache is compared to determine whether a new scan task needs to be issued.
[0188] Here are the steps:
[0189] 1. Listen for or get the "UNSCANNED block" event:
[0190] From the previous module's "File Blocking and Version Management", you know that a block's scan_status == UNSCANNED; or periodically scan the UNSANNED blocks in the FileVersionTable;
[0191] 2. Compare to the global hash cache (HashCache);
[0192] Check whether hash_value is marked as "CLEAN" in the cache and has not expired; if the cache exists and shows that the block content is safe, the scan can be skipped or postponed; otherwise, a ScanTask is generated.
[0193] 3. Generate scan report;
[0194] For the blocks confirmed to need to be scanned (file_id, chunk_index, hash_value), a ScanTask object is constructed and priority or timestamp information is attached; it is delivered to the system's TaskQueue (or scheduling module) to wait for subsequent antivirus engine processing.
[0195] In the data block cache library, a global hash cache needs to be maintained to record the mapping of "hash of each block content to scan result" to skip or accelerate repeated scans of known safe content; it is also responsible for managing the expiration and update strategy of cache results.
[0196] The data structure created:
[0197] I.HashCache;
[0198] Hash_value(key);
[0199] Scan_result:CLEAN / INFECTED / UNKOWN / ERROR;
[0200] Last_scan_time: records the time when the most recent scan was completed;
[0201] Expire_time: The expiration time of the cache entry.
[0202] II.CachePolicy;
[0203] Configure cache expiration policies, such as "valid within n days" and "expired after virus database update";
[0204] Can be stored in global configuration.
[0205] The steps are as follows:
[0206] I. Write cache:
[0207] When a scan task completes and determines the block's hash value to be CLEAN or INFECTED, the result aggregation and arbitration module calls this module to write or update the HashCache. If marked CLEAN, the expiration_time is set to "current time + policy period"; if marked INFECTED, a longer or permanent state can be recorded.
[0208] II. Read and judge skip:
[0209] When the "Incremental Scan Task Generation Module" determines whether to scan, it reads HashCache[hash_value] and compares scan_result with expire_time. If it is still valid, there is no need to rescan. If it has expired or the virus database has been significantly updated, a rescan is required.
[0210] III. Expired or invalid:
[0211] The system can regularly inspect the HashCache and clear or set any entries that have expired to "UNKNOWN." After a virus database update or a major engine version upgrade, the system can also uniformly set the expiration_time of all CLEAN status entries to 0 in batches, prompting subsequent rescanning.
[0212] IV. Deduplication significance:
[0213] In scenarios where multiple files or multiple versions share the same block, as long as the hash value of a block is determined to be safe once, other blocks with the same hash value can be skipped directly; this greatly reduces repeated calls to the feature scanning engine and improves overall processing efficiency.
[0214] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0215] The embodiment of the present application also provides a document scanning device, Figure 8 A structural diagram of a document scanning device provided in an embodiment of the present application is shown as follows: Figure 8 As shown, the device includes:
[0216] An acquisition module 11 is used to acquire a target file to be scanned and a data block cache library;
[0217] The block processing module 12 is used to process the target file into blocks to obtain at least one data block and corresponding attribute information; wherein the attribute information at least includes a data block scanning status and a scanning period;
[0218] A matching processing module 13 is configured to perform matching processing on at least one data block and a data block cache based on attribute information of at least one data block to determine unscanned data blocks;
[0219] The generating module 14 is configured to generate a scanning task according to the unscanned data blocks and perform scanning processing.
[0220] For the description of the features in the embodiment corresponding to the file scanning device, reference can be made to the relevant description of the embodiment corresponding to the file scanning method, and no further details will be given here.
[0221] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned file scanning method embodiments.
[0222] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned file scanning method embodiments when running.
[0223] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0224] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned file scanning method embodiments are implemented.
[0225] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned file scanning method embodiments are implemented.
[0226] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0227] The above is a detailed introduction to the document scanning method, electronic device, medium, and product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core concept of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.
Claims
1. A file scanning method, characterized in that: include: Obtain the target file to be scanned and the data block cache library; Dividing the target file into blocks to obtain at least one data block and corresponding attribute information; wherein the attribute information at least includes a data block scanning status, a scanning period, and hash data of the data block; Performing matching processing on at least one data block and a data block cache based on attribute information of at least one data block to determine an unscanned data block; A scanning task is generated based on the unscanned data blocks for scanning processing.
2. The document scanning method according to claim 1, wherein: The target file is divided into blocks to obtain at least one data block and corresponding attribute information, including: Get the preset data block size; Divide the target file into blocks according to a preset data block size to obtain at least one data block; Determine the starting position, actual length, and corresponding hash data of the data block in the target file according to the preset data block size; Determine the scanning state corresponding to the data block and the scanning period corresponding to the last scanning completion time; wherein, the scanning state corresponding to the first block is an unscanned state and has no last scanning completion time.
3. The document scanning method according to claim 2, wherein: When the target file is an updated file, the process of determining the hash data of the data block includes: Obtaining a change event flag of the target file; Determine a changed data block area according to the starting position and actual length of the changed data block of the change event flag in the target file; In the target file, corresponding hash data is updated for the target data block of the target file according to the changed data block area.
4. The document scanning method according to claim 3, wherein: Determining a changed data block area according to the starting position and actual length of the changed data block of the change event flag at the target file includes: Determine a starting data block according to the starting position and the preset data block size; Determine the end data block according to the starting position, the actual length and the preset data block size; The data blocks between the start data block and the end data block are used as the changed data block area.
5. The document scanning method according to claim 1, wherein: Matching the at least one data block with the data block cache based on the attribute information of the at least one data block to determine the unscanned data blocks includes: Acquire a first target data block in the target file whose data block scanning status is unscanned; Matching the data block cache according to the hash data, scan status, and scan deadline of the first target data block to determine whether there is a data block marked as non-infected and not expired; If not, the first target data block is determined to be safe and is regarded as an unscanned data block.
6. The document scanning method according to claim 5, characterized in that: Matching the data block cache according to the hash data, the scan status, and the scan deadline of the first target data block to determine whether there is a data block marked as non-infected and not expired, including: Obtaining first hash data of the first target data block; Determining, in the data block cache based on the first hash data, whether there is second hash data that is identical to the first hash data; If so, determining a second target data block corresponding to the second hash data; checking whether the corresponding unscanned state is a non-infected state according to the second target data block; If so, determining whether the scanning period of the second target data block has not expired; If it is not expired, the second target data block is determined to be safe, the first target data block is updated according to the non-infected state and the unexpired scanning period, and the scanning is skipped or postponed; If not, it is determined that the first target data is safe.
7. The document scanning method according to claim 6, wherein: Acquiring a first target data block whose data block scanning status is an unscanned state, including: Acquire a first initial target data block whose data block scanning status is an unscanned state; Obtaining the non-expired time and scanning cycle time of the first initial target data block; If the scanning cycle time is greater than the non-expired time, removing the second initial target data block whose scanning cycle time is greater than the non-expired time from the first initial target data block; and determining the remaining first initial target data block as the first target data block; If the scanning cycle time is less than or equal to the non-expired time, all first initial target data blocks are used as first target data blocks.
8. The document scanning method according to any one of claims 1 to 7, characterized in that: After generating a scan task based on the unscanned data blocks and before performing the scan process, the following steps are also included: constructing a scan object according to the unscanned data block; Adding scanning software priority and timestamp information to the scanning object; Determining a scanning priority based on the scanning software priority and timestamp information; The scanning object is scanned according to the scanning priority.
9. The document scanning method according to claim 8, wherein: Determining a scanning priority according to the scanning software priority and timestamp information includes: Setting the scanning software priority to a first weight parameter; determining a second weight parameter according to the timestamp information; The scanning priority is determined according to the scanning software priority, a first weight parameter, the timestamp information, and a second weight parameter.
10. The document scanning method according to claim 1, wherein: The process of establishing the data block cache library includes: Pre-writing data blocks in the target scanning task whose scanning status is non-infected or infected into the data block cache; When the scanning state is a non-infected state, the sum of the current time and the validity period is processed as the expiration time; When the scanning state is an infected state, the corresponding data block is recorded as a permanent state.
11. The document scanning method according to claim 1, wherein: Generate scanning tasks based on unscanned data blocks for scanning processing, including: Determine an overlapping area between a preceding unscanned data block and a succeeding unscanned data block between adjacent unscanned data blocks; Taking the starting position, data block length and overlapping area of the currently unscanned data block as the actual scanning range of the unscanned data block; Starting from the first actual scanning range, scanning is performed on each of the unscanned data blocks and the data blocks corresponding to the actual scanning range.
12. The document scanning method according to claim 1, wherein: Generate scanning tasks based on unscanned data blocks for scanning processing, including: Obtaining file encoding information corresponding to the file to which the unscanned data block belongs; Classification scanning is performed according to the file encoding information and the corresponding unscanned data blocks.
13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the file scanning method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the file scanning method according to any one of claims 1 to 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the document scanning method according to any one of claims 1 to 12 are implemented.
Citation Information
Cited By
Network virus detection method and system, storage medium and electronic equipment
CN121056212A