Disk scanning method and device, equipment, medium and program product
By reading and parsing the metafiles in the master file table within the NTFS file system, the problem of time-consuming disk scanning is solved, enabling fast and efficient disk scanning.
Patent Information
- Application Number
- CN202410559879.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-07
- Publication Date
- 2025-11-11
AI Technical Summary
In existing technologies, disk scanning takes a long time, mainly due to excessive I/O operations and time consumption caused by the recursive enumeration method.
By adopting the linear storage characteristic of metafiles in the master file table of the NTFS file system, file attributes are directly obtained by reading and parsing the first type of metafile, avoiding recursive enumeration at each level, and file attributes are directly parsed at the application layer using the fixed metafile format of the NTFS file system.
It significantly reduces disk scan time, shortening it by nearly 30 times compared to traditional methods, and improves scan speed.
Smart Images

Figure CN120928994A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of disk scanning, and in particular to a disk scanning method, apparatus, device, medium, and program product. Background Technology
[0002] Some security software requires scanning for junk files on the disk. Common disks include local disk (C), local disk (D), local disk (E), and local disk (F).
[0003] In related technologies, file and folder enumeration is performed by calling the operating system's API (Application Programming Interface). The API's function is to return the files and folders within the directory hierarchy provided by a given directory path. If a folder is encountered, the API needs to be called again to enumerate the files and folders at the next directory level, until all files under the initially input path have been enumerated. After returning the files using the API, another API provided by the operating system needs to be called to parse the file's attribute information.
[0004] The relevant technology uses a recursive enumeration method to perform disk scanning, which is time-consuming. Summary of the Invention
[0005] This application provides a disk scanning method, apparatus, device, medium, and program product that can accelerate disk scanning speed. The technical solution includes the following:
[0006] According to one aspect of this application, a method for scanning a disk is provided, the disk employing the new technology file system NTFS, the method comprising the following steps.
[0007] During the disk scanning process, the first type of metafile in the NTFS master file table is read. The first type of metafile stores the attribute information of the files on the disk. The metafiles corresponding to files at different directory levels are stored linearly in the master file table.
[0008] By parsing the first type of metafile, the target attributes of the files on the disk are obtained;
[0009] Files whose target attributes match the attribute conditions are identified as the scan results of the disk.
[0010] In an optional embodiment, during the disk scan, the NTFS master file table is traversed and read according to a fixed data block size; in each read data block, each metafile is traversed and parsed, and each metafile includes the first type of element; by parsing the first type of metafile, the target attributes of the file on the disk are obtained. Each metafile in the master file table has the same size, and the size of a data block is an integer multiple of the size of a metafile.
[0011] In an optional embodiment, the NTFS master file table is jointly stored by multiple storage areas on the disk. The region location offset and region length of each of the multiple storage areas are obtained; each storage area is located using its region location offset; within each storage area, a fixed-size data block is traversed and read until all data within the region length has been traversed.
[0012] In an optional embodiment, the starting cluster position of the master file table in the disk is obtained in the first sector of the disk; the first metafile of the master file table is read in the starting cluster position; and the region position offset and region length of each of the plurality of storage regions are obtained by parsing the first metafile.
[0013] In an optional embodiment, for each read metafile, the position number and parent position number of each metafile are obtained, wherein the position number and the parent position number are obtained by parsing each metafile;
[0014] Based on the location number, parent location number, and related name of each metafile, path concatenation is performed to obtain the complete file path of the file on the disk. For the first type of metafile, the related name is the file name; for the second type of metafile, the related name is the folder name. The second type of metafile is used to store the attribute information of the folders on the disk.
[0015] The location number indicates the position of each metafile in the main file table, the parent location number indicates the position of the parent metafile in the main file table, each metafile corresponds to a file or folder, and the parent metafile is the metafile of the parent directory of the file or folder.
[0016] In an optional embodiment, for each of the read metafiles, the position number, the parent position number, and the file name of the first type of metafile are stored in a first mapping file;
[0017] In addition, the location number, the parent location number, and the folder name of the second type of meta-file are stored in the second mapping file;
[0018] Based on the first mapping file and the second mapping file, path concatenation is performed to obtain the complete file path of the file on the disk.
[0019] In an optional embodiment, the data structure in the first mapping file is a key-value pair, each first key-value pair in the first mapping file corresponds to a metafile, the key of each first key-value pair is the position number, and the value of each first key-value pair is file information, the file information including the position number, the parent position number, and the file name;
[0020] The data structure in the second mapping file is a key-value pair. Each second key-value pair in the second mapping file corresponds to a metafile. The key of each second key-value pair is the location number, and the value of each second key-value pair is folder information, which includes the location number, the parent location number, and the folder name.
[0021] In an optional embodiment, the target attribute of the file includes the file size, and files whose file size exceeds a size threshold are identified as the scan results of the disk.
[0022] In an optional embodiment, the target attribute of the file includes the file creation time, and files whose creation time is earlier than the critical time are identified as the scan results of the disk.
[0023] In an optional embodiment, the target attribute of the file includes the file name, and the file whose file name matches the search term is identified as the scan result of the disk.
[0024] In an optional embodiment, file information of files whose target attributes meet the attribute conditions is displayed, the file information including at least one of file size, file name, file creation time, and full file path.
[0025] According to another aspect of this application, a disk scanning device is provided, the disk employing the new technology file system NTFS, the device comprising the following modules.
[0026] The reading module is used to read the first type of metafile in the NTFS master file table during the scanning process of the disk. The first type of metafile is used to store the attribute information of files on the disk. The metafiles corresponding to files at different directory levels are stored linearly in the master file table.
[0027] The parsing module is used to obtain the target attributes of the files on the disk by parsing the first type of metafile;
[0028] The determination module is used to determine files whose target attributes meet the attribute conditions as the scan results of the disk.
[0029] According to one aspect of this application, a computer device is provided, comprising: a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the disk scanning method described above.
[0030] According to another aspect of this application, a computer-readable storage medium is provided, which stores a computer program that is loaded and executed by a processor to implement the disk scanning method described above.
[0031] According to another aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the disk scanning method described above.
[0032] The beneficial effects of the technical solutions provided in this application include at least the following:
[0033] By reading the first type of metafile in the master file table, which stores attribute information of files on the disk, and then parsing the first type of metafile, the target attributes of the files on the disk are obtained; files whose target attributes meet the attribute conditions are identified as scan results. In the master file table, each file and folder corresponds to a metafile used to record attribute information. The metafiles in the master file table are arranged linearly, one after another. Related technologies require reading the files and subfolders of different directory levels first, and then reading the files of the next level in the subfolders, requiring a recursive approach to completely traverse the files. However, in the master file table, the metafiles of the previous and next levels are not stored hierarchically; they are at the same level. This application reads and parses the linearly stored first type of metafile in the master file table, without recursion, which reduces the time spent obtaining file attributes and thus speeds up disk scanning.
[0034] Furthermore, this application utilizes the characteristics of the NTFS file system. The NTFS file system has a fixed metafile format, so the read metafile can be directly parsed at the application layer to extract file attributes, without needing to use the API functions provided by the operating system, thus reducing time. In related technologies, operating system API calls go from user mode to kernel mode, the kernel mode calls the disk driver to obtain data, and then returns the result from kernel mode to user mode. This multi-layered data transmission and reception process is time-consuming.
[0035] To illustrate, taking a 400GB disk as an example, this application can complete the disk scan in just about 4 seconds, which is nearly 30 times faster than related technologies. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a schematic diagram of the hardware structure of a disk in related technologies.
[0038] Figure 2 This is a schematic diagram of a simple operating system architecture provided in one embodiment of this application.
[0039] Figure 3 This is a schematic diagram illustrating the principle of a disk scanning method provided in one embodiment of this application.
[0040] Figure 4 This is a flowchart of a disk scanning method provided in one embodiment of this application.
[0041] Figure 5 This is a schematic diagram of linear storage provided in one embodiment of this application.
[0042] Figure 6 This is a flowchart of a disk scanning method provided in another embodiment of this application.
[0043] Figure 7 This is a flowchart of a method for obtaining the full path of a file provided in one embodiment of this application.
[0044] Figure 8 This is a schematic diagram of the first and second mapping files provided in one embodiment of this application.
[0045] Figure 9 This is a schematic diagram of a method for concatenating complete file paths according to an embodiment of this application.
[0046] Figure 10 This is a flowchart of a disk scanning method provided in another embodiment of this application.
[0047] Figure 11 This is a flowchart of a disk scanning method provided in another exemplary embodiment of this application.
[0048] Figure 12 This is a schematic diagram of the interface of a security software provided in one embodiment of this application.
[0049] Figure 13 This is a schematic diagram of the scanning results of a large file provided in one embodiment of this application.
[0050] Figure 14 This is a flowchart of a disk scanning method provided in another embodiment of this application.
[0051] Figure 15 This is a structural block diagram of a disk scanning device provided in one embodiment of this application.
[0052] Figure 16 This is a structural block diagram of a computer device provided in one embodiment of this application. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0054] First, a brief introduction to the terms used in the embodiments of this application will be given.
[0055] Disk: A disk is composed of multiple stacked platters. (Refer to reference.) Figure 1 The disk is divided into many concentric circles along the radius of the disk 100. Each concentric circle represents a track 101. Each track 101 is composed of many sector areas. Each sector area is a sector 102. The sector 102 is the smallest unit of data storage on the disk, usually 512 bytes in size. Tracks of the same radius on different platters form cylinders.
[0056] For a disk, its capacity equals the number of heads * the number of tracks (cylinders) * the number of sectors per track * the number of bytes per sector. **Heads:** Each platter typically has two sides, one top and one bottom, corresponding to one head each, for a total of two heads. Therefore, the head used indicates which side the data is on. **Tracks:** Tracks are numbered from the outermost circle of the platter to the innermost circle: track 0, track 1, etc. Concentric circles near the spindle are used to dock the heads and do not store data. **Cylinders:** Equal to the number of tracks. Concentric tracks of the same radius on all platters form a "cylinder." A series of tracks stacked vertically together form a cylinder. **Sectors:** Each track is divided into many sector regions, and each track has the same number of sectors.
[0057] Sector and cluster: A sector is the smallest unit of storage on a disk, while a cluster is the unit the operating system uses to read files from the disk. In other words, a cluster is the basic unit of file storage on a disk, and this is consistent for all files. In Windows operating systems, it is often called a cluster, while in Linux operating systems, it is often called a block. Typically, the storage capacity of a cluster (block) is larger than that of a sector. For a disk, the size of a cluster (block) and sector is fixed after formatting. Normally, by default, one cluster consists of 8 sectors, the default size of one sector is 512 bytes, and the default size of one cluster is 4KB.
[0058] Disk partitioning: The main storage device for information in a computer is the disk. However, a disk cannot be used directly; it must be partitioned. These partitions are called disk partitions. Common partitions include local disks (C:), local disks (D:), and local disks (E:). The operating system is usually placed on the local disk (C:), which is also known as the system disk.
[0059] NTFS (New Technology File System): A file system is a method or data structure used by an operating system to define files on a disk or partition; that is, a method for organizing files on a disk or partition. In Windows operating systems, common file systems include FAT (File Allocation Table), FAT32 (32-bit File Allocation Table), and NTFS. The NTFS file system contains a Master File Table (MFT), which is a hidden metadata file within the NTFS file system. For the NTFS file system, every file and folder on the disk has a metadata file stored in this Master File Table, which is composed of several contiguous disk regions.
[0060] For partitioned disks, such as local disk (C) and local disk (D), each disk has a master file table. The master file table stores metadata information for all files and folders on the current disk, and each metadata file is a fixed 1KB in size. The master file table occupies a certain amount of space for a disk, the specific size depending on the number of files on the disk. The master file table for the current disk is not necessarily a contiguous area; it may be divided into different contiguous areas.
[0061] Disk driver: A disk driver is a wrapper around the physical interface of disk hardware. The operating system uses this driver to control the hardware. The operating system indirectly controls the hardware through the interface provided by the driver, while the driver directly controls the hardware through the actual physical interface.
[0062] In related technologies, disk scanning involves enumerating files and folders by calling the operating system's API (Application Programming Interface). The API returns the files and folders at that directory level after a path is passed in. If a folder is encountered, the API is called again to enumerate files and folders at the next directory level, until all files under the initially passed path have been enumerated. This recursive enumeration method for disk scanning is time-consuming.
[0063] In related technologies, security software is available, often referred to as PC managers, software managers, antivirus software, etc. When these programs perform disk scans, their application-layer API interfaces are scheduled by the operating system, in conjunction with references... Figure 2 , Figure 2 The operating system architecture is illustrated in the diagram. During a disk scan, the computer device first goes through the scheduling of multiple modules in the application layer 201, then passes a read data request to the kernel layer 202, and finally to the disk driver 203. After retrieving the data from the disk driver 203, it goes back through the kernel layer 202 and then back to the application layer 201. This process involves multiple kernel switches and generates multiple I / O operations (read and write operations). Furthermore, a single system call only returns files and directories under the input path, not files and directories in subdirectories. Therefore, for upper-layer applications, to completely enumerate all files in the input path, they need to call system API functions many times, which in turn leads to I / O operations and time consumption. In summary, the disk scanning scheme of related technologies is a very I / O-intensive and time-consuming process.
[0064] Figure 3This illustration shows a schematic diagram of the principle of a disk scanning method provided in an exemplary embodiment of this application.
[0065] In the NTFS file system, a master file table (300) is stored on the disk. This master file table includes multiple metafiles, which are stored linearly within the master file table, one after another. Each file or folder on the disk corresponds to a metafile in the master file table, which stores the file's or folder's attribute information. Furthermore, metafiles corresponding to files (or folders) at different directory levels are stored linearly in the master file table; the master file table (300) does not restrict the hierarchical relationship of the metafiles, and each metafile is independent. For example, if folder A includes file a and subfolder B, and subfolder B includes file b, in the master file table, the metafiles of file a and file b do not have a hierarchical relationship but are stored flatly in a linear fashion—this is linear storage.
[0066] On the disk, as the number of files increases, the main file table 300 will include multiple sub-tables, and these sub-tables will be stored in multiple storage areas. At this point, the main file table 300 will be stored on the disk by multiple contiguous storage areas, as illustrated below. Figure 3 The disk contains storage regions 1, 2, and N, where N is a positive integer. During the disk scan, the region offsets and lengths of multiple storage regions are first obtained, and then the multiple storage regions are traversed.
[0067] For a storage region, the system locates the starting cluster based on the region's offset and then begins traversing and reading fixed-size data blocks, one block at a time, until all data within the region's length has been traversed. Optionally, the data block size can be 512kb or 128kb. A data block size is an integer multiple of a metafile size, which is fixed at 1kb in the NTFS file system. In the NTFS file system, a cluster is typically 4kb. A cluster is the basic unit of file storage on disk, and each read operation reads a power of 2 times the size of one cluster.
[0068] After reading a data block, combine it with the reference. Figure 3 The metadata will be parsed one by one. Figure 3 M metafiles are shown. Each metafile in the NTFS file system has a fixed size of 1kb, meaning that metafile information is parsed every 1kb. In the NTFS file system, a file or folder has a metafile that records its attribute information. Files are generated by the operating system or user operations, and folders can also be called directories. A folder may contain both files and subfolders.
[0069] If a metafile stores file attribute information (i.e., it's a Type 1 metafile), the attribute information includes file name, file creation time, file size, whether it's a file, location number, and parent location number. If a metafile stores folder attribute information (i.e., it's a Type 2 metafile), the attribute information includes folder name, folder creation time, folder size, whether it's a folder, location number, and parent location number. The location number indicates the position of each metafile in the master file table, and each metafile's location number is unique. The parent location number indicates the position of the parent metafile in the master file table. Each metafile corresponds to a file or folder, and the parent metafile is the metafile of the parent directory of the file or folder.
[0070] After parsing each metafile, for the first type of metafile, the attribute information of the files on the disk will be obtained. In this application, files whose target attributes meet the attribute conditions are identified as disk scan results. For example, files larger than 30MB are identified as large files, and the purpose of this disk scan is to scan for large files. For example, files created more than three months ago are identified as early files, and the purpose of this disk scan is to scan for early files. For example, files whose file names match the search terms are identified as search results, and the purpose of this disk scan is to search for files on the disk.
[0071] After parsing each metafile, the complete file path will be concatenated. For the first type of metafile, the parsed location number (FRN), parent location number (parent FRN), and file name are obtained. For the second type of metafile, the parsed location number (FRN), parent location number (parent FRN), and folder name are obtained.
[0072] Reference Figure 3 Metafile 1 stores the attribute information of the file (abc.txt). Based on the parent position number 04 of metafile 1, metafile 4 is found. Metafile 4 stores the attribute information of the folder (File). Based on the parent position number 07 of metafile 4, metafile 7 is found. Metafile 7 stores the attribute information of the folder (Windows). Based on the names of the found metafiles, the complete file path of the file C:\Windows\File\abc.txt is finally constructed.
[0073] After constructing the complete file path, the system will also display the file information of the scanned file, including the file name, file size, file creation time, and complete file path. The file name, file size, and file creation time are obtained through metadata parsing, while the complete file path is obtained by concatenation.
[0074] In this embodiment of the application, the disk scanning method provided will be executed by a computer device. Optionally, the computer device is equipped with security software (also commonly known as antivirus software, software manager, software assistant, etc.). The security software provides disk scanning function, which can scan locally to obtain large files, old files, target search files, etc.
[0075] Optionally, the computer device includes at least one of the following: desktop computer, laptop computer, smartphone, smartwatch, in-vehicle terminal, wearable device, smart TV, tablet computer, e-book reader, MP3 player, and MP4 player.
[0076] It should be noted that all information (including but not limited to the obtained master file table), data (including but not limited to data used for analysis, stored data, and displayed data), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the meta-files involved in this application were obtained with full authorization. Furthermore, regarding related information, the relevant information processor will adhere to the principles of legality, legitimacy, and necessity, clearly define the purpose, method, and scope of the relevant information processing, obtain the consent of the relevant information subject, and take necessary technical and organizational measures to ensure the security of the relevant information.
[0077] Figure 4 A flowchart illustrating a disk scanning method provided in an exemplary embodiment of this application is shown. The disk employs the NTFS file system. The method is illustrated by way of execution by a computer device and includes:
[0078] Step 420: During the disk scan, read the first type of metafile in the NTFS master file table. The first type of metafile stores the attribute information of the files on the disk. The metafiles corresponding to files at different directory levels are stored linearly in the master file table.
[0079] The Master File Table (MFT) is the core of the NTFS file system. Each file and folder has a corresponding metafile (often called a file record) in the MFT, which records the attribute information of the file or folder. In the MFT, metafiles are arranged linearly, meaning they are arranged one after another. Even for files at different directory levels, the metafiles are still arranged linearly. For example, if folder A includes file a and subfolder B, and subfolder B includes file b, in the MFT, the metafiles of file a and file b do not have a hierarchical relationship; instead, they are stored flat, like lines, in the MFT – this is linear storage. (See reference...) Figure 5 , Figure 5 This shows a sequence of metafiles. Figure 5 The metafiles corresponding to folder A, file a, subfolder B, and file b are not stored hierarchically, but are stored at the same level.
[0080] The first type of metafile is the metafile corresponding to the file type, which stores the file's attribute information, such as file name, file size, and file creation time. This application also involves a second type of metafile, which is the metafile corresponding to the folder type. This second type of metafile stores the folder's attribute information, such as folder name, folder size, and folder creation time.
[0081] The files in this application are generated by the operating system or by user operations. Operating system-generated files directly affect the normal operation of the system, and most cannot be altered arbitrarily; their existence plays a crucial role in maintaining the stability of the computer system. User-generated files include user-created documents, game data, and application installation packages.
[0082] The disk refers to a partitioned disk, such as a local disk (C), local disk (D), local disk (E), etc. Each partitioned disk stores a master file table, and the master file tables of different partitions do not interfere with each other. Optionally, the disk in this application is a system disk, that is, a disk on which the operating system is installed, commonly a local disk (C). Optionally, the method of this application can be used for a full disk scan, that is, to perform a scan on each partition. During the full disk scan, the scan of one partition is completed before scanning the next partition, or multiple threads can be set to perform disk scans on multiple partitions simultaneously.
[0083] Step 440: Obtain the target attributes of the files on the disk by parsing the first type of metafile;
[0084] In the master file table, the first type of metafile is the metafile of the file type. That is, the first type of metafile is the metafile corresponding to the file. The first type of metafile stores the attribute information of the file on the disk.
[0085] In this application, the target attributes of the file on the disk are obtained by parsing the first type of metafile.
[0086] Step 460: Files whose target attributes match the attribute conditions are identified as disk scan results.
[0087] Optionally, the target attribute includes file size, and the attribute condition is a condition that the file size must meet. In this embodiment, the file size of the files on the disk is obtained by parsing the first type of metafile. Then, files whose file size exceeds a size threshold are identified as disk scan results. That is, this embodiment provides a method for scanning large files on the disk. The size threshold corresponding to large files can be arbitrarily set. Therefore, the method for scanning large files provided in this embodiment has strong versatility.
[0088] Optionally, the target attributes of the file include the file creation time, and the attribute conditions are the conditions that the file creation time must meet. In this embodiment, by parsing the first type of metafile, the file creation time of the files on the disk is obtained. Then, files whose creation time is earlier than the critical time are identified as the disk scan results. That is, this embodiment provides a method for scanning early files. The critical time corresponding to the early files can be arbitrarily set. Therefore, the method for scanning early files provided in this embodiment has strong versatility.
[0089] Optionally, the target attributes of the file include the file name, and the attribute conditions are the conditions that the file name must satisfy. In this embodiment, by parsing the first type of metafile, the file names of the files on the disk are obtained. Then, the files whose file names match the search terms are determined as the disk scan results. That is, this embodiment provides a file retrieval method. Optionally, the way to match file names with search terms can be arbitrary and varied. Therefore, the file retrieval method provided in this embodiment has strong versatility.
[0090] In one embodiment, the file information of the scanned files is also sent to the application layer, where the application layer displays the final scan results. Optionally, the file information of files whose target attributes match the attribute conditions is displayed. The file information includes at least one of the following: file name, file size, and file creation time.
[0091] The application layer, located at the top of the operating system architecture, typically hosts various applications developed using high-level programming languages, such as security software, map software, contact management, music players, address books, and so on.
[0092] In summary, by reading the first type of metafile in the master file table (which stores attribute information of files on the disk) and then parsing it, the target attributes of the files on the disk are obtained. Files whose target attributes match the attribute conditions are identified as scan results. In the master file table, each file and folder corresponds to a metafile that records attribute information. The metafiles in the master file table are arranged linearly, one after another. Related technologies require reading the files and subfolders at different directory levels first, and then reading the files at the next level within the subfolders—a recursive approach is needed to completely traverse the files. However, in the master file table, there is no distinction between files at different levels; their metafiles are at the same level. This application reads and parses the linearly stored first type of metafile in the master file table, without recursion, which reduces the time spent obtaining file attributes and thus speeds up disk scanning.
[0093] Furthermore, this application utilizes the characteristics of the NTFS file system. The NTFS file system has a fixed metafile format, so the read metafile can be directly parsed at the application layer to extract file attributes, without needing to use the API functions provided by the operating system, thus reducing time. In related technologies, operating system API calls go from user mode to kernel mode, the kernel mode calls the disk driver to obtain data, and then returns the result from kernel mode to user mode. This multi-layered data transmission and reception process is time-consuming.
[0094] To illustrate, taking a 400GB disk as an example, this application can complete the disk scan in just about 4 seconds, which is nearly 30 times faster than related technologies.
[0095] In one embodiment, steps 420 and 440 can be replaced by: during disk scanning, traversing and reading the NTFS master file table according to a fixed data block size; in each read data block, traversing and parsing each metafile, each metafile including a first type of element; and obtaining the target attributes of the file on the disk by parsing the first type of metafile. Here, each metafile in the master file table has the same size, and the size of a data block is an integer multiple of the size of a metafile.
[0096] In this embodiment, the metafiles in the data block are stored one after another. Optionally, the data block size is 512kb or 128kb, and each metafile is 1kb in size. That is, in this embodiment, the metafile information will be parsed once every 1kb in the data block. For the first type of metafile, the attribute information of the file is obtained by parsing.
[0097] In this embodiment, a data block includes multiple metafiles. This means that the attribute information of multiple files or folders can be obtained with a single read operation (IO operation), significantly reducing the number of IO operations. For example, if there are nearly 3 million files on the system disk and the main file table is close to 300MB, assuming each data block read is 512KB, approximately 600 read operations would be required. Related technologies would require approximately 3 million read operations to actually traverse and scan 3 million files. Clearly, the number of read operations in this embodiment is not on the same order of magnitude as related technologies; the number of read operations is negligible.
[0098] Optionally, the master file table can be stored in multiple storage areas on the disk. Figure 6 A flowchart illustrating a disk scanning method provided in an exemplary embodiment of this application is shown, illustratively described by way of the method being performed by a computer device. The method includes:
[0099] Step 610: Obtain the starting cluster position of the master file table in the first sector of the disk;
[0100] In this embodiment, after initiating a disk scan, the first sector (BootSecond) of the disk is read first. For the NTFS file system, the first sector is the boot sector, storing crucial information about the current disk driver. This includes the number of sectors in a cluster, the number of bytes in a sector, and the starting cluster position of the master file table. The master file table requires disk space and is shared by multiple storage areas, each containing multiple clusters. The master file table is stored at a specific location on the disk, and this location marks the beginning of the master file table's starting cluster.
[0101] Step 620: Read the first metafile of the master file table from the starting cluster position;
[0102] After obtaining the starting cluster position, the file pointer is moved to the starting cluster position. This position represents the $MFT metafile. The $MFT metafile is the first metafile of the main file table. The data area of the $MFT metafile stores information about the distribution area (multiple storage areas) of the main file table. Since the main file table is stored by multiple storage areas, the multiple storage areas will be represented by their respective area position offsets and area lengths.
[0103] Step 630: By parsing the first metafile, the region location offset and region length of each of the multiple storage regions are obtained;
[0104] By parsing the $MFT metafile, we can obtain the region location offset and region length of each of the multiple storage regions. The region location offset is the cluster starting position of this region, that is, the position of the first cluster.
[0105] Step 640: Locate each storage region by offsetting the region location of each storage region;
[0106] Based on the region location offset of each storage location, locate the first cluster of each storage region.
[0107] Step 650: In each storage area, traverse and read data blocks of a fixed size until all data within the area length has been traversed.
[0108] In each storage region, a fixed 512kb (or 128kb) data block is read. If the length of a storage region is less than 512kb, a 512kb data block is read only once. If the length of a storage region is greater than 512kb, 512kb data blocks are read continuously until all data within the region length has been traversed.
[0109] Optionally, a fixed-size data block can be read by traversing through the Readfile function.
[0110] Step 660: Parse the metafile one by one in each read data block;
[0111] In the NTFS file system, each metafile is fixed at 1KB in size. In this application, metafile information is parsed every 1KB. Within a data block, there may be a first-class metafile and a second-class metafile. The first-class metafile corresponds to the file type and stores file attribute information, such as file name, file size, and file creation time.
[0112] The second type of metafile is the metafile corresponding to the folder (directory) type. The second type of metafile stores the folder's attribute information, such as the folder name, folder size, folder creation time, etc.
[0113] Step 670: Obtain the target attributes of the file on the disk by parsing the first type of metafile;
[0114] By parsing the first type of metafile, the file's attribute information will be obtained, including target attributes. Optionally, the target attributes may include at least one of the following: file size, file creation time, and file name.
[0115] Step 680: Files whose target attributes match the attribute conditions are identified as disk scan results;
[0116] Optionally, the target attribute includes file size, identifying files whose size exceeds a certain threshold as disk scan results. Optionally, the target attribute includes file creation time, identifying files whose creation time is earlier than the threshold as disk scan results. Optionally, the target attribute includes file name, identifying files whose names match the search terms as disk scan results.
[0117] Step 690: Display the file information of files whose target attributes meet the attribute conditions. The file information includes at least one of the following: file name, file size, and file creation time.
[0118] In summary, the above embodiments provide a scheme for reading and parsing the first type of metafile from the master file table. In conclusion, this application will replace the actual traversal scanning process of related technologies by scanning the master file table, which greatly reduces scanning time and I / O consumption.
[0119] exist Figure 6 In the optional embodiment shown, step 690 will display the file information of files whose target attributes meet the attribute conditions. This file information is attribute information that can be directly obtained by parsing the first type of metafile. In the following embodiments, the complete file path will also be displayed; the complete file path cannot be directly obtained by parsing the first type of metafile.
[0120] based on Figure 6 In the optional embodiment shown, step 670 is further performed... Figure 7 Steps 701 and 702 are shown. Step 680 and step 703 are the same steps. Step 690 is replaced by step 704. Figure 7 This is a flowchart illustrating a method for generating a complete file path according to an exemplary embodiment of this application, exemplified by the method being executed by a computer device. The method includes:
[0121] Step 701: For each metafile read, obtain the position number and parent position number of each metafile. The position number and parent position number are obtained by parsing each metafile.
[0122] In step 660, the metafiles in the data blocks are parsed one by one. For each metafile, attribute information is parsed, which includes the file name, file size, file creation time, folder name, folder size, and folder creation time mentioned above, as well as a location number and a parent location number. Each file or folder has its own location number and parent location number in its corresponding metafile.
[0123] A location number, also known as a file record number (FRN), indicates the position of each metafile in the master file table. Each metafile has a unique location number. Each metafile corresponds to a file or folder; therefore, the location number can be used to distinguish between different files or folders.
[0124] The parent position number indicates the position of the parent metafile in the main file table. Each metafile corresponds to a file or folder, and the parent metafile is the metafile of the parent directory of the file or folder.
[0125] For example, if metafile A has a position number of 01 and a parent position number of 05, and metafile A is used to store the attribute information of file a, and the parent directory (upper directory level) of file a is folder b, then we can know that metafile A has a position number of 01 in the main file table, metafile B has a position number of 05 and a position number of 05 in the main file table, and metafile B is used to store the attribute information of folder b.
[0126] In one embodiment, for each read metafile, the location number, parent location number, and file name of the first type of metafile are stored in a first mapping file. And the location number, parent location number, and folder name of the second type of metafile are stored in a second mapping file.
[0127] The first type of metafile is the metafile corresponding to the file type, and it is used to store the attribute information of files on the disk. In this embodiment, the file information is stored in the first mapping file (first map file), which includes the file name, file size, file creation time, location number, and parent location number. The data structure in the first mapping file is key-value pairs. Each first key-value pair in the first mapping file corresponds to a metafile, the key of each first key-value pair is the location number, and the value of each first key-value pair is the file information.
[0128] Reference Figure 8 , Figure 8 Part (A) shows the first mapping file. The data in the mapping file (map file) is structured as key-value pairs (key, value). In the first mapping file, the key of the key-value pair is FR N, i.e., the position index, and the value is file information. One key-value pair corresponds to one metafile.
[0129] The second type of metafile corresponds to the folder type and is used to store the attribute information of folders on the disk. Folder information is stored in the second mapping file (the first map file), including folder name, folder size, folder creation time, location number, and parent location number. The data structure in the second mapping file is key-value pairs. Each second key-value pair in the second mapping file corresponds to a metafile, with the key being the location number and the value being the folder information.
[0130] Reference Figure 8 , Figure 8 Part (B) shows the second mapping file. The data in the mapping file (map file) is structured as key-value pairs (key, value). In the second mapping file, the key of the key-value pair is FR N, i.e., the location index, and the value is folder information. One key-value pair corresponds to one metafile.
[0131] It is understandable that using FRN as the key for key-value pairs is advantageous because FRNs are unique; each file or folder has only one FRN, and one FRN corresponds to only one file or folder. Therefore, using FRN as the key facilitates query operations in the mapping file.
[0132] Step 702: Based on the location number, parent location number, and related name of each metafile, perform path concatenation to obtain the complete file path of the file on the disk;
[0133] For the first type of metafile, the related name is the file name; for the second type of metafile, the related name is the folder name. The first type of metafile is used to store the attribute information of files, and the second type of metafile is used to store the attribute information of folders.
[0134] In this embodiment, the position number and parent position number link the metafile and its parent metafile. The system searches upwards layer by layer based on the position number and parent position number, then concatenates the file names or folder names stored within the found metafiles according to directory relationships to obtain the complete file path on the disk. In the master file table, the physical addresses of the metafile and its parent metafile do not need to be consecutive; they are independent metafiles. The metafile and its parent metafile are linked by their position numbers and parent position numbers. That is, the position number of the metafile may precede or follow the position number of the parent metafile. This way, when a file is moved from one folder to another, no operation needs to be performed on the file's data itself; only its parent position number needs to be changed.
[0135] illustrative, for reference only Figure 9Metafile A stores the attribute information of the file (abc.txt). Based on the parent location number 04 of metafile A, metafile D is found. Metafile D stores the attribute information of the folder (File). Based on the parent location number 07 of metafile D, metafile G is found. Metafile G stores the attribute information of the folder (Windows). Using the names of the found metafiles, the complete file path of the file C:\Windows\File\abc.txt is finally constructed.
[0136] In one embodiment, based on the first and second mapping files, path concatenation is performed to obtain the complete file path of the file on the disk. From the first mapping file, the path number, parent path number, and file name of each first-type metafile are extracted. From the second mapping file, the path number, parent path number, and folder name of each second-type metafile are extracted. For the initial metafile in the first mapping file, the parent file number is found. Then, a key-value search is performed in the second mapping file according to the parent file number. If the attribute information of the found metafile contains the parent file number, the key-value search continues in the second mapping file according to its parent file number, and so on, until the finally found metafile does not have a parent file number. Finally, the initial metafile and all found metafiles are concatenated in reverse order to obtain the complete file path of the file corresponding to the initial metafile.
[0137] It is understandable that by storing the attribute information of the first type of metafile and the second type of metafile in two mapping files respectively, the attribute information of the initial file is placed in one mapping file and the attribute information of the folder is placed in the other mapping file. When concatenating the complete file paths of all files, it is only necessary to traverse the first mapping file and then query the second mapping file one by one. The path generation process of all files is orderly and it is not easy to have omissions or duplicate generation.
[0138] Step 703: Files whose target attributes match the attribute conditions are identified as disk scan results;
[0139] Optionally, the target attribute includes file size, identifying files whose size exceeds a certain threshold as disk scan results. Optionally, the target attribute includes file creation time, identifying files whose creation time is earlier than the threshold as disk scan results. Optionally, the target attribute includes file name, identifying files whose names match the search terms as disk scan results.
[0140] Step 704: Display the file information of files whose target attributes match the attribute conditions. The file information includes at least one of the following: file name, file size, file creation time, and full file path.
[0141] In the file information, the file name, file size, and file creation time can be obtained by parsing the first type of metafile, while the complete file path is obtained by concatenating steps 701 and 702 above.
[0142] In this application, the file information will be sent to the application layer, which will then display the file information of the file whose target attributes match the attribute conditions.
[0143] In summary, the above embodiments obtain the complete file path by concatenating the parsed location number (FRN), parent location number (parent FRN), and name (file name or folder name). This fully utilizes the information obtainable from parsing the metafile, creating a coupling stage between the process of generating the complete file path and the process of obtaining the target attribute. The process of generating the complete file path is quite ingenious and simple.
[0144] Furthermore, it will also display file information, including at least one of the following: file name, file size, file creation time, and full file path, consistent with the display content commonly found in related technologies. In other words, by scanning the main file table, this application maintains consistency with related technologies in both the accuracy of the scan results and the information displayed, while surpassing related technologies in terms of scan time and the number of I / O operations consumed.
[0145] Figure 10 A flowchart illustrating a disk scanning method provided in an exemplary embodiment of this application is shown, illustratively described by way of the method being performed by a computer device. The method includes:
[0146] Step 1001, Begin;
[0147] Step 1002: Obtain the system disk path;
[0148] When performing a large file cleanup scan on the system drive (C drive), the system drive path, such as C drive, will be obtained.
[0149] Step 1003: Read the data from the first sector (BootSector) of the system disk;
[0150] During the scan, the first sector of the system disk will be read first, because for the NTFS file system, the first sector is the Boot sector, which stores key information about the current drive, such as how many sectors are in a cluster of the current file system, how many bytes are in a sector, and the starting cluster position of the MFT table.
[0151] Step 1004: Parse and obtain the cluster size, sector size, and $MFT metafile;
[0152] From the first sector, we need to obtain the cluster size (how many sectors a cluster contains, and how many bytes each sector contains) and the starting cluster position of the MFT table, because the entire scanning process in this application requires scanning the metafiles in the MFT table one by one. It's worth noting that the MFT table is a very important structure in the NTFS file system. For the NTFS file system, each file has a metafile to store the file's attribute information, and these metafiles are all of a fixed size of 1KB. The MFT table is the important data table used to aggregate these metafiles.
[0153] Step 1005: Adjust the file pointer to the beginning of the $MFT metafile and read 1kb of data;
[0154] After obtaining the MFT starting cluster, move the file pointer to the starting cluster position. This position represents the $MFT metafile. Here, it should be noted that the $MFT metafile is the first metafile in the MFT metafile table, and the size of the $MFT metafile is 1kb.
[0155] Step 1006: Parse the data attribute regions of the $MFT metafile to obtain the distribution regions of the MFT table;
[0156] The data area of this $MFT metafile stores the regional data of the MFT table. Since the MFT table may be divided into blocks, these blocks are represented by cluster position offset (cluster start position) and length. Therefore, parsing the $MFT metafile is mainly to obtain the regional information of the distribution area of the entire MFT table.
[0157] Step 1007: Obtain the cluster start position and region length of each region from the distribution region of the MFT table to form a region linked list;
[0158] Based on the cluster starting position and region length of each distribution area, a region linked list is formed. This linked list is used to locate each distribution area by pointer.
[0159] Step 1008: Traverse the region linked list;
[0160] After parsing the $MFT metafile and obtaining the region information of the MFT table, the metafile traversal process is started. The file pointer is located by the cluster position offset mentioned earlier, and the end of the traversal is controlled by the region length.
[0161] Step 1009: Read 512kb of data, and then start parsing the metafile every 1kb.
[0162] When traversing each region, a fixed 512KB data block is read, and then metadata information is parsed every 1KB, as the metadata file has a fixed length of 1KB in the NTFS file system. The metadata file contains various attributes representing file name, creation time, size, and whether it is a file or folder. Crucially important is the FRN (File Record Number), which represents the metadata file's unique position in the table, allowing different files to be distinguished. Additionally, each file or folder has its parent directory's FRN. Because the NTFS file system does not include path information in the metadata file's attributes, obtaining the path requires traversing the parent directory level by level until the complete path is constructed.
[0163] Step 1010: Determine if it is a file type;
[0164] Based on the attribute information parsed from the current metafile, determine whether the current metafile is a file type. The file type corresponds to the metafile and the file, and the original file stores the file's attribute information.
[0165] Step 1011: Parse and obtain the folder name;
[0166] If the current metafile is not a file type, it means the current metafile is a folder type. Parse to obtain the folder name of the current metafile.
[0167] Step 1012: Parse and obtain the folder FRN;
[0168] Parse and retrieve the folder FRN of the current metadata file.
[0169] Step 1013, use the folder information<FRN,Info> Store the data in the folder "map" in the correct format;
[0170] During the traversal, there are two maps to store the parsed information. One is the directory map (i.e., the folder map). The key of the directory map is FRN, and the value is the folder information, which includes the folder name and the folder FRN.
[0171] Step 1014: Parse and obtain the file name;
[0172] If the current metafile is a file type, then parse and obtain the filename of the file corresponding to the current metafile.
[0173] Step 1015: Parse and obtain the file creation time;
[0174] Step 1016: Parse and obtain the file size;
[0175] Step 1017: Parse and obtain the file FRN and its parent FRN;
[0176] Step 1018, use the file information<FRN,Info> Store the data in the file map in the form of pairs;
[0177] During the traversal, another map is the file map. The key of the file map is the FRN, and the value is the file information, which includes the file name, file creation time, file size, file FRN, and parent FRN.
[0178] Step 1019: Determine whether all regions have been traversed.
[0179] Step 1020: Iterate through the file map and folder map to find the complete file path;
[0180] Step 1021: Return the results to the large file cleaning and scanning module to complete the overall scanning process;
[0181] After scanning all regions, the paths will be concatenated. Once concatenated, the file information that meets the size criteria, including file name, full file path, file creation time, and file size, will be passed to the upper-level large file adaptation layer for final result display.
[0182] Step 1022, End.
[0183] In summary, the above embodiments, based on the characteristics of the NTFS file system, complete a full disk scan by scanning the MFT table. For a normal 400GB system disk with 3 million files, the above embodiments only need to read 300MB of data through the Readfile interface to complete the traversal of the entire MFT table, saving a lot of time compared to the traditional method of enumerating 3 million files.
[0184] Furthermore, in the above embodiments, a fixed-size data block is read each time via Readfile. Since the NTF S file system has a fixed metafile format, the read data can be directly parsed by the code to obtain file attributes without using the API functions provided by the Windows system. Therefore, the parsing is completed entirely at the application layer. Compared with the traditional solution that requires traversing and enumerating the files and then using the API provided by the system to obtain the file size, time and other attributes, there is a significant time optimization.
[0185] In one embodiment, the disk scanning method provided in this application can be used to scan large files.
[0186] Figure 11A flowchart illustrating a disk scanning method provided in an exemplary embodiment of this application is shown. The disk uses the NTFS file system. The method is illustrated by example, performed by a computer device, and includes:
[0187] Step 1120: During the disk scan, read the first type of metafile in the NTFS master file table. The first type of metafile stores the attribute information of the files on the disk. The metafiles corresponding to files at different directory levels are stored linearly in the master file table.
[0188] Step 1140: Obtain the file size of the file on the disk by parsing the first type of metafile;
[0189] In one embodiment, during disk scanning, the NTFS master file table is traversed and read according to a fixed data block size; in each read data block, each metafile is traversed and parsed, and each metafile includes a first type of element; by parsing the first type of metafile, the file size of the file on the disk is obtained; wherein, each metafile in the master file table has the same size, and the size of a data block is an integer multiple of the size of a metafile.
[0190] In an optional embodiment, the NTFS master file table is stored jointly by multiple storage regions on the disk. The region location offset and region length of each of the multiple storage regions are obtained; each storage region is located using its region location offset; within each storage region, fixed-size data blocks are traversed and read until all data within the region length has been traversed.
[0191] In an optional embodiment, the starting cluster position of the master file table on the disk is obtained in the first sector of the disk; the first metafile of the master file table is read in the starting cluster position; and the region position offset and region length of each of the multiple storage regions are obtained by parsing the first metafile.
[0192] In an optional embodiment, for each read metafile, the position number and parent position number of each metafile are obtained, which are obtained by parsing each metafile;
[0193] Based on the location number, parent location number, and related name of each metafile, path concatenation is performed to obtain the complete file path of the file on the disk. For the first type of metafile, the related name is the file name; for the second type of metafile, the related name is the folder name. The second type of metafile is used to store the attribute information of the folder on the disk.
[0194] The position number indicates the position of each metafile in the main file table, the parent position number indicates the position of the parent metafile in the main file table, each metafile corresponds to a file or folder, and the parent metafile is the metafile of the parent directory of the file or folder.
[0195] In an optional embodiment, for each read metafile, the location number, parent location number, and file name of the first type of metafile are stored in a first mapping file; and the location number, parent location number, and folder name of the second type of metafile are stored in a second mapping file; based on the first mapping file and the second mapping file, path concatenation is performed to obtain the complete file path of the file on the disk.
[0196] In an optional embodiment, the data structure in the first mapping file is a key-value pair. Each first key-value pair in the first mapping file corresponds to a metafile. The key of each first key-value pair is a position number, and the value of each first key-value pair is file information, which includes a position number, a parent position number, and a file name.
[0197] The data structure in the second mapping file is key-value pairs. Each second key-value pair in the second mapping file corresponds to a metafile. The key of each second key-value pair is the location number, and the value of each second key-value pair is folder information, which includes the location number, the parent location number, and the folder name.
[0198] Step 1160: Files whose size is greater than the size threshold are identified as disk scan results.
[0199] Optionally, a size threshold of 300MB is set; files larger than 300MB are defined as large files. (Refer to the reference.) Figure 12 , Figure 12 A schematic diagram of the interface of security software is shown. Figure 12 The image shows the relevant operation area 1201 for large file cleanup provided by security software. In response to a trigger operation received by the "Select Cleanup" control in the relevant operation area 1201, the scanned file items are displayed. Figure 13 The diagram shows the file entries of the large files obtained from the scan, along with the format, creation time, size, and cleanup suggestions for each file entry. The file name (file entry), creation time, and size are obtained directly from parsing the first type of metafiles in the master file table or after processing.
[0200] In an optional embodiment, file information for files whose size is greater than a size threshold is displayed. The file information includes at least one of the following: file size, file name, file creation time, and full file path.
[0201] In summary, by reading the first type of metafile in the master file table (which stores attribute information of files on the disk) and then parsing it, the file size of the files on the disk is obtained. Files with a size greater than a threshold are identified as scan results. In the master file table, each file and folder corresponds to a metafile that records attribute information. The metafiles in the master file table are arranged linearly, one after another. Related technologies require reading the files and subfolders of different directory levels first, and then reading the files of the next level within the subfolders—a recursive approach is needed to completely traverse the files. However, in the master file table, there is no distinction between files at different levels; their metafiles are at the same level. This application reads and parses the linearly stored first type of metafile in the master file table, without recursion, which reduces the time spent obtaining file attributes and thus speeds up disk scanning.
[0202] To illustrate, taking a 400GB disk as an example, this application can complete the disk scan in just about 4 seconds, which is nearly 30 times faster than related technologies.
[0203] In one embodiment, the disk scanning method provided in this application can be used to scan earlier files.
[0204] Figure 14 A flowchart illustrating a disk scanning method provided in an exemplary embodiment of this application is shown. The disk uses the NTFS file system. The method is illustrated by example, performed by a computer device, and includes:
[0205] Step 1420: During the disk scan, read the first type of metafile in the NTFS master file table. The first type of metafile stores the attribute information of the files on the disk. The metafiles corresponding to files at different directory levels are stored linearly in the master file table.
[0206] Step 1440: Obtain the file creation time of the file on the disk by parsing the first type of metafile;
[0207] In one embodiment, during disk scanning, the NTFS master file table is traversed and read according to a fixed data block size; in each read data block, each metafile is traversed and parsed, and each metafile includes a first type of element; by parsing the first type of metafile, the file creation time of the file on the disk is obtained; wherein, each metafile on the master file table has the same size, and the size of a data block is an integer multiple of the size of a metafile.
[0208] In an optional embodiment, the NTFS master file table is stored jointly by multiple storage regions on the disk. The region location offset and region length of each of the multiple storage regions are obtained; each storage region is located using its region location offset; within each storage region, fixed-size data blocks are traversed and read until all data within the region length has been traversed.
[0209] In an optional embodiment, the starting cluster position of the master file table on the disk is obtained in the first sector of the disk; the first metafile of the master file table is read in the starting cluster position; and the region position offset and region length of each of the multiple storage regions are obtained by parsing the first metafile.
[0210] In an optional embodiment, for each read metafile, the position number and parent position number of each metafile are obtained, which are obtained by parsing each metafile;
[0211] Based on the location number, parent location number, and related name of each metafile, path concatenation is performed to obtain the complete file path of the file on the disk. For the first type of metafile, the related name is the file name; for the second type of metafile, the related name is the folder name. The second type of metafile is used to store the attribute information of the folder on the disk.
[0212] The position number indicates the position of each metafile in the main file table, the parent position number indicates the position of the parent metafile in the main file table, each metafile corresponds to a file or folder, and the parent metafile is the metafile of the parent directory of the file or folder.
[0213] In an optional embodiment, for each read metafile, the location number, parent location number, and file name of the first type of metafile are stored in a first mapping file; and the location number, parent location number, and folder name of the second type of metafile are stored in a second mapping file; based on the first mapping file and the second mapping file, path concatenation is performed to obtain the complete file path of the file on the disk.
[0214] In an optional embodiment, the data structure in the first mapping file is a key-value pair. Each first key-value pair in the first mapping file corresponds to a metafile. The key of each first key-value pair is a position number, and the value of each first key-value pair is file information, which includes a position number, a parent position number, and a file name.
[0215] The data structure in the second mapping file is key-value pairs. Each second key-value pair in the second mapping file corresponds to a metafile. The key of each second key-value pair is the location number, and the value of each second key-value pair is folder information, which includes the location number, the parent location number, and the folder name.
[0216] Step 1460: Files whose creation time is earlier than the critical time are identified as disk scan results.
[0217] For example, if a file was created three months ago, then the file creation time is considered to be earlier than the critical time.
[0218] In an optional embodiment, file information for files whose size is greater than a size threshold is displayed. The file information includes at least one of the following: file size, file name, file creation time, and full file path.
[0219] In summary, by reading the first type of metafile in the master file table (which stores attribute information of files on the disk) and then parsing it, the file creation time of the files on the disk is obtained. Files whose creation time is earlier than a critical time are identified as scan results. In the master file table, each file and folder corresponds to a metafile used to record attribute information. The metafiles in the master file table are arranged linearly, one after another. Related technologies require reading the files and subfolders of different directory levels first, and then reading the files of the next level within the subfolders, requiring a recursive approach to completely traverse the files. However, in the master file table, the metafiles of the previous and next levels are not stored hierarchically; they are at the same level. This application reads and parses the linearly stored first type of metafile in the master file table without recursion, which reduces the time spent obtaining file attributes and thus speeds up disk scanning.
[0220] To illustrate, taking a 400GB disk as an example, this application can complete the disk scan in just about 4 seconds, which is nearly 30 times faster than related technologies.
[0221] Figure 15 This application shows a structural block diagram of a disk scanning apparatus provided in an exemplary embodiment, wherein the disk employs the new technology file system NTFS, and the apparatus includes:
[0222] The reading module 1501 is used to read the first type of metafile in the master file table on NTFS during the disk scanning process. The first type of metafile stores the attribute information of the files on the disk. The metafiles corresponding to files at different directory levels are stored linearly in the master file table.
[0223] Parsing module 1502 is used to obtain the target attributes of files on the disk by parsing the first type of metafile;
[0224] The determination module 1503 is used to determine files whose target attributes meet the attribute conditions as the disk scan results.
[0225] In an optional embodiment, the reading module 1501 is further configured to traverse and read the NTFS master file table according to a fixed data block size during disk scanning; the parsing module 1502 is further configured to traverse and parse each metafile in each read data block, each metafile including a first type of element; and obtain the target attributes of the file on the disk by parsing the first type of metafile. Each metafile in the master file table has the same size, and the size of a data block is an integer multiple of the size of a metafile.
[0226] In an optional embodiment, the master file table is jointly stored by multiple storage areas on the disk. The read module 1501 is further configured to obtain the area position offset and area length of each of the multiple storage areas; locate each storage area using the area position offset of each storage area; and traverse and read fixed-size data blocks in each storage area until all data within the area length has been traversed.
[0227] In an optional embodiment, the reading module 1501 is further configured to obtain the starting cluster position of the master file table in the first sector of the disk; read the first metafile of the master file table in the starting cluster position; and obtain the region position offset and region length of each of the multiple storage regions by parsing the first metafile.
[0228] In an optional embodiment, the device further includes a processing module 1504. The processing module 1504 is further configured to, for each read metafile, obtain the position number and parent position number of each metafile, which are obtained by parsing each metafile; based on the position number, parent position number, and associated name of each metafile, perform path concatenation to obtain the complete file path of the file on the disk; for the first type of metafile, the associated name is the file name; for the second type of metafile, the associated name is the folder name; the second type of metafile is used to store the attribute information of folders on the disk; wherein, the position number is used to indicate the position of each metafile in the master file table, the parent position number is used to indicate the position of the parent metafile in the master file table, each metafile corresponds to a file or folder, and the parent metafile is the metafile of the parent directory of the file or folder.
[0229] In an optional embodiment, the processing module 1504 is further configured to, for each read metafile, store the location number, parent location number, and file name of the first type of metafile through a first mapping file; and store the location number, parent location number, and folder name of the second type of metafile through a second mapping file; and perform path concatenation based on the first mapping file and the second mapping file to obtain the complete file path of the file on the disk.
[0230] In an optional embodiment, the data structure in the first mapping file is a key-value pair. Each first key-value pair in the first mapping file corresponds to a metafile. The key of each first key-value pair is a position number, and the value of each first key-value pair is file information, which includes a position number, a parent position number, and a file name.
[0231] The data structure in the second mapping file is key-value pairs. Each second key-value pair in the second mapping file corresponds to a metafile. The key of each second key-value pair is the location number, and the value of each second key-value pair is folder information, which includes the location number, the parent location number, and the folder name.
[0232] In an optional embodiment, the target attributes of the file include the file size, and the determination module 1503 is further configured to determine files whose file size exceeds a size threshold as disk scan results.
[0233] In an optional embodiment, the target attributes of the file include the file creation time, and the determination module 1503 is further configured to determine files whose file creation time is earlier than the critical time as the disk scan result.
[0234] In an optional embodiment, the target attributes of the file include the file name, and the determination module 1503 is further configured to determine the file whose file name matches the search term as the disk scan result.
[0235] In an optional embodiment, the apparatus further includes a display module 1505. The display module 1505 is used to display file information of files whose target attributes meet the attribute conditions. The file information includes at least one of the following: file size, file name, file creation time, and full file path.
[0236] In summary, by reading the first type of metafile in the master file table (which stores attribute information of files on the disk) and then parsing it, the target attributes of the files on the disk are obtained. Files whose target attributes match the attribute conditions are identified as scan results. In the master file table, each file and folder corresponds to a metafile that records attribute information. The metafiles in the master file table are arranged linearly, one after another. Related technologies require reading the files and subfolders of different directory levels first, and then reading the files of the next level within the subfolders—a recursive approach is needed to completely traverse the files. However, in the master file table, the metafiles of the previous and next levels are stored hierarchically, and they are at the same level. This application reads and parses the linearly stored first type of metafile in the master file table, without recursion, which reduces the time spent obtaining file attributes and thus speeds up disk scanning.
[0237] To illustrate, taking a 400GB disk as an example, this application can complete the disk scan in just about 4 seconds, which is nearly 30 times faster than related technologies.
[0238] Figure 16 A structural block diagram of a computer device 1600 provided in an exemplary embodiment of this application is shown. The computer device 1600 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The computer device 1600 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.
[0239] Typically, computer device 1600 includes a processor 1601 and a memory 1602.
[0240] Processor 1601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1601 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1601 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0241] The memory 1602 may include one or more computer-readable storage media, which may be non-transitory. The memory 1602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1602 is used to store at least one instruction, which is executed by the processor 1601 to implement the disk scanning method provided in the method embodiments of this application.
[0242] In some embodiments, the computer device 1600 may also optionally include a peripheral device interface 1603 and at least one peripheral device. The processor 1601, memory 1602, and peripheral device interface 1603 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1603 via a bus, signal line, or circuit board. For example, the peripheral device may include at least one of the following: a radio frequency circuit 1604, a display screen 1605, a camera assembly 1606, an audio circuit 1607, and a power supply 1608.
[0243] Peripheral interface 1603 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1601 and memory 1602. In some embodiments, processor 1601, memory 1602 and peripheral interface 1603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1601, memory 1602 and peripheral interface 1603 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0244] The radio frequency (RF) circuit 1604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1604 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1604 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1604 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0245] Display screen 1605 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1605 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1601 for processing. In this case, display screen 1605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1605, disposed on the front panel of computer device 1600; in other embodiments, there may be at least two display screens, disposed on different surfaces of computer device 1600 or in a folded design; in still other embodiments, display screen 1605 may be a flexible display screen, disposed on a curved or folded surface of computer device 1600. Furthermore, display screen 1605 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1605 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0246] The camera assembly 1606 is used to acquire images or videos. Optionally, the camera assembly 1606 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1606 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0247] The audio circuit 1607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 1601 for processing, or to the radio frequency circuit 1604 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the computer device 1600. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1601 or the radio frequency circuit 1604 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1607 may also include a headphone jack.
[0248] Power supply 1608 is used to supply power to the various components in computer device 1600. Power supply 1608 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1608 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0249] In some embodiments, the computer device 1600 further includes one or more sensors 1609. The one or more sensors 1609 include, but are not limited to, an accelerometer 1610, a gyroscope 1611, a pressure sensor 1612, an optical sensor 1613, and a proximity sensor 1614.
[0250] Accelerometer 1610 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by computer device 1600. For example, accelerometer 1610 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1601 can control display screen 1605 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1610. Accelerometer 1610 can also be used for games or for acquiring user motion data.
[0251] The gyroscope sensor 1611 can detect the orientation and rotation angle of the computer device 1600. The gyroscope sensor 1611 can work in conjunction with the accelerometer sensor 1610 to acquire 3D motion data from the user on the computer device 1600. Based on the data acquired by the gyroscope sensor 1611, the processor 1601 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0252] The pressure sensor 1612 can be disposed on the side bezel of the computer device 1600 and / or on the lower layer of the display screen 1605. When the pressure sensor 1612 is disposed on the side bezel of the computer device 1600, it can detect the user's grip signal on the computer device 1600, and the processor 1601 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1612. When the pressure sensor 1612 is disposed on the lower layer of the display screen 1605, the processor 1601 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1605. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0253] Optical sensor 1613 is used to collect ambient light intensity. In one embodiment, processor 1601 can control the display brightness of display screen 1605 based on the ambient light intensity collected by optical sensor 1613. For example, when the ambient light intensity is high, the display brightness of display screen 1605 is increased; when the ambient light intensity is low, the display brightness of display screen 1605 is decreased. In another embodiment, processor 1601 can also dynamically adjust the shooting parameters of camera assembly 1606 based on the ambient light intensity collected by optical sensor 1613.
[0254] The proximity sensor 1614, also known as a distance sensor, is typically located on the front panel of the computer device 1600. The proximity sensor 1614 is used to detect the distance between the user and the front of the computer device 1600. In one embodiment, when the proximity sensor 1614 detects that the distance between the user and the front of the computer device 1600 is gradually decreasing, the processor 1601 controls the display screen 1605 to switch from a screen-on state to a screen-off state; when the proximity sensor 1614 detects that the distance between the user and the front of the computer device 1600 is gradually increasing, the processor 1601 controls the display screen 1605 to switch from a screen-off state to a screen-on state.
[0255] Those skilled in the art will understand that Figure 16 The structure shown does not constitute a limitation on the computer device 1600, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0256] This application also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the disk scanning method provided in the above method embodiments.
[0257] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the disk scanning method provided in the above-described method embodiments.
[0258] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0259] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0260] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A disk scanning method, characterized in that, The disk uses the new NTFS file system, and the method includes: During the disk scanning process, the first type of metafile in the NTFS master file table is read. The first type of metafile stores the attribute information of the files on the disk. The metafiles corresponding to files at different directory levels are stored linearly in the master file table. By parsing the first type of metafile, the target attributes of the files on the disk are obtained; Files whose target attributes match the attribute conditions are identified as the scan results of the disk.
2. The method according to claim 1, characterized in that, During the disk scanning process, the first type of metafile in the NTFS master file table is read, and the target attributes of the files on the disk are obtained by parsing the first type of metafile, including: During the disk scan, the NTFS master file table is traversed and read according to a fixed data block size; In each data block read, each metafile is traversed and parsed, and each metafile includes the first type of element; By parsing the first type of metafile, the target attributes of the files on the disk are obtained; In this context, each metafile in the master file table has the same size, and the size of a data block is an integer multiple of the size of a metafile.
3. The method according to claim 2, characterized in that, The NTFS master file table is stored jointly by multiple storage areas on the disk; During the disk scan, the NTFS master file table is traversed and read according to a fixed data block size, including: Obtain the region position offset and region length of each of the multiple storage regions; Positioning is achieved by offsetting the location of each of the storage regions; In each of the storage areas, a fixed-size data block is read and traversed until all data within the length of the area has been traversed.
4. The method according to claim 3, characterized in that, The step of obtaining the region location offset and region length of each of the plurality of storage regions includes: Obtain the starting cluster position of the master file table in the first sector of the disk; Read the first metafile of the main file table from the starting cluster position; By parsing the first metafile, the region position offset and region length of each of the multiple storage regions are obtained.
5. The method according to any one of claims 2 to 4, characterized in that, The method further includes: For each metafile read, obtain the position number and parent position number of each metafile, which are obtained by parsing each metafile. Based on the location number, parent location number, and related name of each metafile, path concatenation is performed to obtain the complete file path of the file on the disk. For the first type of metafile, the related name is the file name; for the second type of metafile, the related name is the folder name. The second type of metafile is used to store the attribute information of the folders on the disk. The location number indicates the position of each metafile in the main file table, the parent location number indicates the position of the parent metafile in the main file table, each metafile corresponds to a file or folder, and the parent metafile is the metafile of the parent directory of the file or folder.
6. The method according to claim 5, characterized in that, The method further includes: For each of the read metafiles, the position number, the parent position number, and the file name of the first type of metafile are stored in the first mapping file; In addition, the location number, the parent location number, and the folder name of the second type of meta-file are stored in the second mapping file; The process of concatenating paths based on the location number, parent location number, and related name of each metafile to obtain the complete file path of the file on the disk includes: Based on the first mapping file and the second mapping file, path concatenation is performed to obtain the complete file path of the file on the disk.
7. The method according to claim 6, characterized in that, The data structure in the first mapping file is a key-value pair. Each key-value pair in the first mapping file corresponds to a metafile. The key of each key-value pair is the position number, and the value of each key-value pair is file information, which includes the position number, the parent position number, and the file name. The data structure in the second mapping file is a key-value pair. Each second key-value pair in the second mapping file corresponds to a metafile. The key of each second key-value pair is the location number, and the value of each second key-value pair is folder information, which includes the location number, the parent location number, and the folder name.
8. The method according to any one of claims 1 to 4, characterized in that, The target attributes of the file include file size, and determining files whose target attributes meet the attribute conditions as the disk scan results includes: Files whose size exceeds a certain threshold are identified as the results of the disk scan.
9. The method according to any one of claims 1 to 4, characterized in that, The target attributes of the file include the file creation time. The step of identifying files whose target attributes match the attribute conditions as the disk scan results includes: Files whose creation time is earlier than the critical time are identified as the disk scan results.
10. The method according to any one of claims 1 to 4, characterized in that, The target attributes of the file include the file name, and determining the file whose target attributes meet the attribute conditions as the disk scan result includes: Files whose names match the search terms are identified as the scan results of the disk.
11. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Display file information of files whose target attributes meet the attribute conditions, including at least one of file size, file name, file creation time, and full file path.
12. A disk scanning device, characterized in that, The disk uses the new NTF-S file system, and the device includes: The reading module is used to read the first type of metafile in the NTFS master file table during the scanning process of the disk. The first type of metafile stores the attribute information of the files on the disk. The metafiles corresponding to files at different directory levels are stored linearly in the master file table. The parsing module is used to obtain the target attributes of the files on the disk by parsing the first type of metafile; The determination module is used to determine files whose target attributes meet the attribute conditions as the scan results of the disk.
13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the disk scanning method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is loaded and executed by a processor to implement the disk scanning method as described in any one of claims 1 to 11.
15. A computer program product, characterized in that, The computer program product stores a computer program that is loaded and executed by a processor to implement the disk scanning method as described in any one of claims 1 to 11.