File fragment type identification method based on context content awareness
By introducing the BERT-BiLSTM model and contextual relationship modeling, combined with the file fragment data set with strong authenticity, the problems of insufficient contextual relationships and lack of authenticity in file fragment recognition in the existing technology are solved, and high-precision file fragment type recognition is achieved, supporting file recovery and digital forensics.
Patent Information
- Application Number
- CN202510056363.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art fails to fully utilize the context dependence characteristics between file fragments in file fragment recognition, which makes it difficult for classification models to capture the global characteristics of fragments, and the existing data sets lack authenticity and cannot effectively reflect the distribution of file fragments in storage devices.
By introducing the BERT-BiLSTM model, combining contextual relationship modeling and sequence dependency feature processing capabilities, a file fragment data set simulates the distribution of real files of the hard disk is constructed, the structure, location and adjacent relationship characteristics of the file fragment are extracted, and deep learning model training is carried out to identify different types of file fragments.
It significantly improves the accuracy and applicability of file fragment type recognition, overcomes the problems of insufficient contextual relationships and lack of authenticity of data sets in the prior art, and provides efficient and accurate file recovery and digital forensic solutions.
Smart Images

Figure CN119989043A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a context-based file fragmentation type identification method, which is suitable for scenarios such as file recovery, digital forensics, and storage system optimization. It can effectively simulate the file fragmentation distribution characteristics in a storage environment and classify and identify file fragmentation types through a deep learning model. Background Art
[0002] In modern storage systems, files are often divided into multiple fragments due to read and write operations, insufficient storage space, etc., and each fragment is stored in a non-contiguous manner on the disk or other storage media. This file fragmentation phenomenon brings great challenges to tasks such as file recovery and digital forensics.
[0003] Existing file fragmentation identification technologies mainly rely on the following methods:
[0004] Feature matching-based methods: By analyzing the characteristics of file fragments, such as byte distribution, entropy value, etc., they try to classify or reorganize the fragments. However, these methods have limited performance when faced with complex fragment distributions, especially when dealing with fragment types with strong contextual relationships.
[0005] Methods based on machine learning and deep learning: Some studies use machine learning and deep learning models to extract features from file fragments, and finally input them into the model for classification. However, since the contextual dependency characteristics between fragments are not fully considered, the accuracy of these classification models in real file recovery scenarios is not high.
[0006] Using randomly shuffled file fragmentation datasets to train the model: Existing methods are trained using artificially constructed fragmentation datasets, but these datasets usually lack sufficient simulation of the actual storage environment and cannot truly reflect the file fragmentation distribution in the storage device.
[0007] The above prior art has the following deficiencies:
[0008] Insufficient contextual relationships: Failure to fully utilize the contextual dependency characteristics between file fragments makes it difficult for the classification model to capture the global characteristics of the fragments.
[0009] The dataset lacks authenticity: Existing fragmented datasets are difficult to truly reflect the distribution characteristics of fragments in storage devices, which affects the applicability of classification models.
[0010] In actual hard disk storage, adjacent file fragments often belong to the same file or similar file types, so the adjacent relationship between file fragments may provide important clues. Therefore, a new technical method is urgently needed to simulate the file fragment distribution characteristics in a real storage environment, combine the excellent sequence semantic extraction and understanding capabilities of the model in the field of natural language processing (NLP), improve the accuracy and applicability of file fragment sequence type recognition, and thus overcome the shortcomings of existing technologies. Summary of the invention
[0011] This paper introduces the BERT-BiLSTM model, combines contextual relationship modeling and sequence dependency feature processing capabilities, and significantly improves the recognition accuracy of file fragment types. This innovative file fragment type recognition method overcomes the limitations of existing technologies and provides an efficient and accurate solution for data recovery and file forensics.
[0012] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0013] A method for identifying file fragment types based on contextual relationships, characterized by comprising the following steps:
[0014] a) Construct a file fragmentation dataset that simulates the real file distribution on a hard disk;
[0015] b) extracting features from the file fragments in the constructed dataset;
[0016] c) Use the BERT-BiLSTM model to learn the contextual relationship of file fragments;
[0017] d) Classifying and identifying different types of file fragments based on the context-dependent characteristics.
[0018] Further, preferably, the step a) comprises the following sub-steps:
[0019] (1) The original file is divided into two parts according to a set division ratio, wherein the division ratio r is generated by a predefined probability distribution function, and the preselected probability density function is as follows;
[0020]
[0021] (2) Calculating the storage distance between the two fragments after segmentation, the distance d is generated by a predefined probability distribution function, and the preselected probability density function is as follows;
[0022]
[0023] (3) The obtained file fragments are placed in the simulated storage space, the positions of the fragments in the simulated storage are determined by the segmentation ratio r and the segmentation distance d, and the original file information, offset position and adjacent relationship of the fragments are recorded.
[0024] Further, preferably, in the step b), the feature extraction of the file fragments includes:
[0025] (1) Extract the structural features of file fragments, including the byte distribution, entropy value, and content pattern of the fragments;
[0026] (2) Extract physical location features, including the storage location, size, and distance to adjacent fragments of the fragment;
[0027] (3) Extract adjacent relationship features to describe the logical connection and similarity between fragments.
[0028] Furthermore, preferably, in the step c), the BERT module introduces feature extraction data of file fragments as input features, position information in the sequence as additional features, and optimizes the attention mechanism in combination with position weights, thereby improving the modeling capability of the global contextual characteristics of the file fragment sequence; the BiLSTM module further combines the context embedding and timing characteristics of the fragments, integrates the output results of the bidirectional sequence through residual connections, and enhances the ability to capture dependencies between fragments.
[0029] Furthermore, preferably, in the step d), the classifier is a multi-layer fully connected network, whose activation function is ReLU, and a regularization mechanism is introduced after each layer of the network to prevent overfitting, and the final classification result is output through the Softmax function; the classifier can adapt to the classification requirements of single type and mixed types in file fragments, and significantly improve the classification accuracy.
[0030] Further, preferably, the classification results include the following types: determined type and mixed type (MIX). Determined fragment types include common file types such as JPG, PNG, MP3, etc. Mixed type fragments indicate that the file fragment sequence contains more than one file fragment type, and the classification results are determined by combining contextual relationship characteristics and adjacent features.
[0031] Furthermore, preferably, a sliding window mechanism is set to analyze the file fragment sequence, the window size is set to W, the sliding step is set to 1, and for each fragment sequence in the sliding window {f 1 ,f 2 ,…,f w}, when the classification result is "MIX" type, the end position of the current window is used as the starting point of the new sliding window, starting from f w The position starts sliding again to avoid double counting.
[0032] Furthermore, preferably, for the square component fragments classified as mixed types, independent features of each type of fragments are further extracted through a multi-task learning model, wherein the model includes an independent feature separation network and a joint classification module to generate detailed classification results.
[0033] Furthermore, preferably, the construction of the file fragmentation data set includes a predefined segmentation ratio range of 30% to 70%, wherein a 30% probability of not performing segmentation and a 70% probability of performing segmentation according to a uniform distribution, and the ratio range of the two parts of the file after segmentation is 30% to 70%, the storage distance range of the two parts of the file fragments after segmentation is 10 to 200 storage blocks, and supports dynamic adjustment of multiple storage block sizes.
[0034] Further, preferably, the file systems supported by the method include FAT32, NTFS and EXT4, wherein a different contextual relationship modeling strategy is used for each file system to optimize classification performance.
[0035] Furthermore, preferably, the contextual relationship modeling introduces position encoding and adjacent fragment interaction modules into the BERT-BiLSTM model to improve the adaptability of fragment classification in complex storage environments.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] Improved classification accuracy: By introducing the BERT-BiLSTM model and combining the contextual relationship and sequence dependency characteristics, the accuracy and stability of file fragment classification are improved.
[0038] Real environment simulation: Construct a fragmented data set that simulates the real file distribution of a hard disk, taking into account the segmentation ratio, storage distance, and adjacent relationship, to enhance the adaptability to the actual storage environment.
[0039] Complex type support: It can identify single-type and mixed-type file fragments, and perform multi-task learning fine-grained classification on mixed types to achieve accurate processing of complex-type file fragments.
[0040] Strong versatility: Supports multiple file systems (FAT32, NTFS, EXT4), and uses different dataset construction strategies and model hyperparameter selection for each file system to optimize classification performance.
[0041] Dynamic storage features: The segmentation ratio and storage distance are controlled through probability distribution, and dynamic adjustment of multiple storage block sizes is supported, which improves the adaptability to different storage devices.
[0042] Reduced misclassification rate: The introduction of position encoding and adjacent fragment interaction modules improves the model's adaptability to fragment classification tasks in complex storage environments and significantly reduces the misclassification rate.
[0043] Efficient classification processing: The classifier is combined with a multi-layer fully connected network and finally outputs the classification result through the Softmax function, ensuring efficient and stable classification performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flow chart of the method for identifying file fragment types based on contextual relationship of the present invention;
[0045] Figure 2 It is a schematic diagram of the file segmentation strategy for constructing a data set in the present invention;
[0046] Figure 3 It is a statistical graph of average continuous length of file fragments in a data set constructed in an embodiment of the present invention;
[0047] Figure 4 It is a diagram of the BERT-BiLSTM model constructed by the present invention; DETAILED DESCRIPTION
[0048] The present invention significantly improves the recognition accuracy of file fragment types in personal hard disks by introducing the BERT-BiLSTM model, combining contextual relationship modeling and sequence dependency feature processing capabilities. At the same time, the data set construction simulates the real storage environment, and enhances the authenticity and diversity of file fragment data by controlling the file segmentation ratio and the distance between storage fragments. Since modern mainstream file systems are based on sequential storage of file contents, the files stored in them have certain contextual relationships, so the classification model can run under a variety of file systems (such as FAT32, NTFS and EXT4), and supports the recognition of file fragments of 512 and 4096 bytes in size, adapting to the complex needs of different personal hard disk storage environments. Through position coding and adjacent fragment interaction modules, the fragment classification ability in complex storage environments is further optimized, and the misclassification rate is significantly reduced. This innovative method for identifying file fragment types overcomes the limitations of the prior art, provides a feasible solution for the efficient recovery of lost or damaged data in personal hard disks, and can be applied to electronic data forensics technology to help forensics personnel quickly locate and identify key information. It provides important support for data integrity protection in complex storage environments.
[0049] The specific implementation is as follows:
[0050] See also Figure 1 , which shows a flow chart of the file fragment type identification method of the present invention, the method comprising the following steps:
[0051] a) Construct a file fragmentation dataset that simulates the real file distribution on a hard disk;
[0052] b) extracting features from the file fragments in the constructed dataset;
[0053] c) Use the BERT-BiLSTM model to learn the contextual relationship of file fragments;
[0054] d) Classifying and identifying different types of file fragments based on the context-dependent characteristics.
[0055] Further, preferably, the step a) comprises the following sub-steps:
[0056] (1) The original file is divided into two parts according to a set division ratio, wherein the division ratio r is generated by a predefined probability distribution function, and the preselected probability density function is as follows;
[0057]
[0058] (2) Calculating the storage distance between the two fragments after segmentation, the distance d is generated by a predefined probability distribution function, and the preselected probability density function is as follows;
[0059]
[0060] (3) Place the obtained file fragments in the simulated storage space, determine the location of the fragments in the simulated storage by the segmentation ratio r and the segmentation distance d, and record the original file information, offset position, and adjacent relationship of the fragments
[0061] Further, preferably, in the step b), the feature extraction of the file fragments includes:
[0062] (1) Extract the structural features of file fragments, including the byte distribution, entropy value, and content pattern of the fragments;
[0063] (2) Extract physical location features, including the storage location, size, and distance to adjacent fragments of the fragment;
[0064] (3) Extract adjacent relationship features to describe the logical connection and similarity between fragments.
[0065] Furthermore, preferably, in the step c), the classifier is a multi-layer fully connected network, whose activation function is ReLU, and a regularization mechanism is introduced after each layer of the network to prevent overfitting, and the final classification result is output through the Softmax function; the classifier can adapt to the classification requirements of single type and mixed types in file fragments, and significantly improve the classification accuracy.
[0066] Furthermore, preferably, in the step d), the classifier is a multi-layer fully connected network, whose activation function is ReLU, and a regularization mechanism is introduced after each layer of the network to prevent overfitting, and the final classification result is output through the Softmax function; the classifier can adapt to the classification requirements of single type and mixed types in file fragments, and significantly improve the classification accuracy.
[0067] Further, preferably, the classification results include the following types: document type fragments, picture type fragments, video type fragments and mixed type fragments, and the classification results are determined by combining contextual relationship characteristics and adjacent features.
[0068] Furthermore, preferably, for file fragments classified as mixed types, independent features of each type of fragments are further extracted through a multi-task learning model, wherein the model includes an independent feature separation network and a joint classification module to generate detailed classification results.
[0069] Furthermore, preferably, the construction of the file fragmentation data set includes a predefined segmentation ratio range of 30% to 70%, wherein a 30% probability of not performing segmentation and a 70% probability of performing segmentation according to a uniform distribution, and the ratio range of the two parts of the file after segmentation is 30% to 70%, the storage distance range of the two parts of the file fragments after segmentation is 10 to 200 storage blocks, and supports dynamic adjustment of multiple storage block sizes.
[0070] Further, preferably, the file systems supported by the method include FAT32, NTFS and EXT4, wherein a different contextual relationship modeling strategy is used for each file system to optimize classification performance.
[0071] Furthermore, preferably, the contextual relationship modeling introduces position encoding and adjacent fragment interaction modules into the BERT-BiLSTM model to improve the adaptability of fragment classification in complex storage environments.
[0072] The following is explained with reference to embodiments.
[0073] Embodiment 1: A method for identifying the type of personal hard disk file fragments based on contextual relationships
[0074] This embodiment aims at the problem of identifying file fragments in personal hard disks and proposes a method for identifying file fragment types based on contextual relationship modeling. Since the data types stored in personal hard disks are diverse, such as documents, pictures, videos, etc., the data will usually be divided into fragments when the file is deleted or the hard disk is damaged. This method constructs a data set that simulates the real storage environment and combines it with a deep learning model to achieve efficient classification and recovery of personal hard disk file fragments.
[0075] First, the dataset is constructed by extracting multiple file types from the GovDocs dataset as data sources, including common document files (such as PDF, Word), image files (such as JPEG, PNG) and directly readable files (such as html). The original file is segmented according to the segmentation ratio r ranging from 30% to 70%, of which 30% of the probability is not segmented, and 70% of the probability is segmented according to uniform distribution, and the ratio of the two parts of the segmented file ranges from 30% to 70%. Subsequently, the storage distance d (10 to 200 storage block lengths) between the two parts after segmentation is generated according to the uniform distribution. The segmented file is stored in the simulated storage space according to the generated ratio r and distance d, and the original file information, size, offset position of the fragment and the relationship with the adjacent fragments are recorded.
[0076] The present invention chooses to divide the file into two segments, mainly based on the following reasons: First, it is common for files to be divided into two segments in the storage system, especially when writing data, due to insufficient storage space or file system operations, some fragments are stored in another area. Two-segment segmentation can effectively simulate the fragmentation phenomenon of most storage devices, close to the real scene. Secondly, two-segment segmentation simplifies the data generation and analysis process, and avoids the combination problem of complex multi-segment segmentation. Simple two-segment segmentation is easy to control variables, such as segmentation ratio and storage distance, thereby reducing design complexity. The file is divided into two segments according to the ratio, and segment 1 and segment 2 are distributed in the storage space at random distances, respectively, showing the dispersion and discontinuity of file fragments in a real storage environment. Figure 2 The average continuous length of file fragments in the constructed data set is shown. This distribution reflects the advantages of the present invention in simulating a real storage environment and provides high-quality data support for subsequent classification and identification.
[0077] In terms of feature extraction, the method constructs the representation vector of the fragment by extracting the structural features of the file fragment (byte frequency distribution, entropy value, original byte embedding vector), position features (such as the position information of the fragment in the input sequence), and adjacent relationship features (similarity of logical connections between fragments). This further enriches the input information of the classification model. By extracting these features and combining them with contextual relationship modeling, the present invention significantly improves the classification accuracy of complex fragments.
[0078] In the classification model design, the BERT-BiLSTM model is used to learn the contextual relationship of fragments, such as Figure 4As shown in the figure. In the classification model design, the BERT-BiLSTM model is used to learn the contextual relationship of fragments. The BERT module captures the global contextual characteristics of the fragment sequence through the attention mechanism, and the BiLSTM module further extracts the temporal relationship and potential dependency characteristics of the fragment sequence. Finally, the classifier uses a multi-layer fully connected network, takes the embedded features of the model as input, and outputs the classification result through the Softmax function. The model optimizes parameters for the multi-type file fragmentation characteristics in personal hard drives. The BERT module uses fine-tuned pre-trained weights, the BiLSTM contains two layers, each with 128 hidden units, and the classifier is a three-layer fully connected network.
[0079] The advantages of this method are reflected in the following aspects: the accuracy of classification is significantly improved through contextual relationship modeling; the applicability of the model in a real storage environment is enhanced by simulating a fragmented data set of a real storage environment; it supports fine-grained analysis of mixed types of fragments and can effectively solve the problem of file fragment classification in complex scenarios. This example verifies the effectiveness of this method in identifying and recovering file fragments on personal hard disks, and provides a high-quality solution for the secure storage and efficient recovery of personal data.
[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A file fragment type identification method based on contextual content perception, characterized in that: The following steps are involved: a) Construct a file fragmentation dataset that simulates the real file distribution on a hard disk; b) extracting features from the file fragments in the constructed dataset; c) Use a model based on BERT (Bidirectional Encoder Representation Transformer) and BiLSTM (Bidirectional Long Short-Term Memory Network) to learn the contextual relationship of file fragments; d) Classifying and identifying different types of file fragments based on the context-dependent characteristics.
2. The file fragment type identification method according to claim 1, characterized in that: The step a) comprises the following sub-steps: (1) The original file is divided into two parts according to a set division ratio, wherein the division ratio r is generated by a predefined probability distribution function, and the preselected probability density function is as follows; (2) Calculating the storage distance between the two fragments after segmentation, the distance d is generated by a predefined probability distribution function, and the preselected probability density function is as follows; (3) The obtained file fragments are placed in the simulated storage space, the positions of the fragments in the simulated storage are determined by the segmentation ratio r and the segmentation distance d, and the original file information, offset position and adjacent relationship of the fragments are recorded.
3. The file fragment type identification method according to claim 2, characterized in that: In the step b), the feature extraction of the file fragments includes: (1) Extract the structural features of file fragments, including the byte frequency distribution, entropy value, and content pattern of the fragments; (2) Input the file fragments into the convolutional encoder neural network and extract the corresponding byte embedding vectors; (3) Extract physical location features, including the storage location, size, and distance to adjacent fragments of the fragment.
4. The file fragment type identification method according to claim 3, characterized in that: In the step c), the BERT module introduces the feature extraction data of the file fragments as input features, the position information in the sequence as an additional feature, and optimizes the attention mechanism in combination with the position weight, thereby improving the modeling capability of the global context characteristics of the file fragment sequence; the BiLSTM module further combines the context embedding and timing characteristics of the fragments, integrates the output results of the bidirectional sequence through residual connections, and enhances the ability to capture dependencies between fragments.
5. The file fragment type identification method according to claim 4, characterized in that: In the step d), the classifier is a multi-layer fully connected network, whose activation function is ReLU, and a regularization mechanism is introduced after each layer of the network to prevent overfitting, and the final classification result is output through the Softmax function; the classifier can output a certain type or mixed type after recognizing the input file fragment sequence.
6. The file fragment type identification method according to claim 1, characterized in that: The classification results include determining the fragment type and mixed type fragments (MIX). The determined fragment types include common file types such as JPG, PNG, MP3, etc. Mixed type fragments indicate that the file fragment sequence contains more than one file fragment type. The classification results are determined by combining contextual relationship characteristics and adjacent features.
7. The file fragment type identification method according to claim 5, characterized in that: In step d), a sliding window mechanism is set to analyze the file fragment sequence, the window size is set to W, the sliding step is set to 1, and for each fragment sequence {f1, f2, ..., f w }, when the classification result is "MIX" type, the end position of the current window is used as the starting point of the new sliding window, starting from f w The position starts sliding again to avoid double counting.
8. The file fragment type identification method according to claim 6, characterized in that: For file fragments classified as mixed types, independent features of each type of fragment are further extracted through a multi-task learning model, which includes an independent feature extraction network and a joint classification module to generate detailed classification results.
9. The file fragment type identification method according to claim 1, characterized in that: The construction of the file fragment data set includes a predefined segmentation ratio range of 30% to 70%, wherein a 30% probability of not performing segmentation and a 70% probability of performing segmentation according to a uniform distribution, and a ratio range of 30% to 70% between the two parts of the file after segmentation, a storage distance range of 10 to 200 storage blocks between the two parts of the file fragments after segmentation, and supports dynamic adjustment of multiple storage block sizes.
10. The file fragment type identification method according to claim 1, characterized in that: The file systems supported by the method include FAT32, NTFS and EXT4, where different dataset construction strategies and model hyperparameter selections are used for each file system to optimize classification performance.
11. The file fragment type identification method according to claim 1, characterized in that: The contextual relationship modeling introduces position encoding and adjacent fragment interaction modules into the BERT-BiLSTM model to improve the adaptability of fragment classification in complex storage environments.