Mobile hard disk storage space optimization method based on AI algorithm
Through the combination of XGBoost model and Gaussian hybrid model, intelligent management of mobile hard disk storage space is achieved, solving the problems of wasted storage resources and inefficient data management in the existing technology, and improving the utilization rate of storage space and the operation efficiency of file system.
Patent Information
- Application Number
- CN202510896390.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing mobile hard disk storage space optimization methods are difficult to analyze and manage intelligently in an automated manner, resulting in waste of storage resources and inefficient data management, which cannot meet the needs of intelligent storage optimization in massive data scenarios.
The XGBoost model is used for invalid data classification, combined with the Gaussian hybrid model for clustering, and duplicate files are identified using file fingerprint vectors and Jaccard similarity check methods, and the file value is evaluated through the importance index, and the storage space is dynamically managed.
It improves storage space utilization, reduces disk redundancy, improves the operating efficiency of the file system and the intelligent level of system management, ensures that high-value files are retained first, and automatically cleans up low-value files.
Smart Images

Figure CN120406854A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hard disk storage space optimization, and specifically provides a method for optimizing the storage space of a mobile hard disk based on an AI algorithm. Background Art
[0002] With the development of information technology and the growth of data storage requirements, mobile hard disks have become important data storage media for individual and enterprise users. However, over time, the storage space in mobile hard disks is gradually occupied by a large number of redundant files, invalid data, and low-value content, resulting in problems such as wasted storage resources, decreased data management efficiency, and slower reading speeds.
[0003] Existing methods for optimizing the space of mobile hard disks mainly rely on traditional manual cleaning, simple file deduplication, or screening methods based on file names and sizes, and it is difficult to perform intelligent analysis and automatic optimization for different types of files. In addition, existing methods often cannot combine artificial intelligence algorithms to dynamically evaluate and intelligently manage stored content, lacking in-depth calculations of file importance and similarity, resulting in limited optimization effects. At the same time, traditional storage optimization lacks a flexible index management structure and cannot dynamically adjust the index after clustering analysis, deduplication processing, and importance ranking, easily causing excessive system overhead or incomplete optimization, and unable to meet the intelligent and automated storage optimization requirements in the current massive data scenario.
[0004] Therefore, there is an urgent need to propose a method for optimizing the storage space of a mobile hard disk that combines artificial intelligence algorithms with an efficient index structure to achieve intelligent management and optimization of file metadata, thereby improving the utilization rate of the storage space of the mobile hard disk. Summary of the Invention
[0005] The present invention provides a method for optimizing the storage space of a mobile hard disk based on an AI algorithm. The present invention uses the XGBoost model to remove invalid data, ensuring that subsequent clustering only processes valid and representative data, thereby improving the clustering quality. The Gaussian mixture model is used for unsupervised clustering to accurately group file metadata and further optimize file management. Combining the file fingerprint vector and the Jaccard similarity-based duplicate checking method can efficiently identify and delete duplicate files, significantly reducing disk space redundancy and improving storage efficiency. At the same time, the importance index dynamically evaluates the value of files by comprehensively considering access times, file sizes, and time factors, ensuring that high-value files are preferentially retained and automatically cleaning low-value files, further optimizing the use of storage space and system management.
[0006] To achieve the above object, the present invention provides the following technical solutions: A method for optimizing the storage space of a mobile hard disk based on an AI algorithm, including: Scan the file system of the external hard drive, obtain the metadata of all files through the NTFS index table, and establish a first-level metadata index chain; Use the pre-trained XGBoost model to perform binary classification on the metadata. The classification includes valid data and invalid data. Delete the invalid data and its metadata, and update the NTFS index table and the first-level metadata index chain; Cluster the metadata through the Gaussian mixture model, establish multiple second-level metadata index chains according to the clustering results, and delete the first-level metadata index chain; Obtain the file fingerprint vectors of the metadata on each second-level metadata index chain and perform vector duplicate checking. Delete the duplicate files and their metadata according to the duplicate checking results, and update the NTFS index table and the second-level metadata index chain; Calculate the importance index of all the metadata in the second-level metadata index chain, establish a metadata B+ tree according to the importance index, and delete the second-level metadata index chain.
[0007] Furthermore, the pre-training steps of the XGBoost model include: Collect the metadata of the external hard drive files to construct a training data set; Divide the file metadata into two categories: valid data and invalid data and label them to form a supervised learning training set; Adopt the K-fold cross-validation method to optimize the hyperparameters of the XGBoost model, and use the gradient boosting decision tree for training to maximize the classification accuracy; When the improvement of the F1-score of the validation set of the XGBoost model in 10 consecutive iterations is less than the error threshold, the model training is completed.
[0008] Furthermore, the clustering steps of the Gaussian mixture model include: Select the feature vectors of the file metadata as the clustering input; Adopt the independent component analysis method to reduce the dimension of the feature vectors to obtain the reduced-dimensional feature vectors; Adopt the Gaussian mixture model to perform unsupervised clustering on the reduced-dimensional feature vectors, use the expectation maximization algorithm to estimate the parameters of each Gaussian distribution, and dynamically select the optimal number of clustering clusters through the Bayesian information criterion; Calculate the internal file similarity of each clustering cluster, verify the clustering quality, and combine the clustering results with the file type feature mapping relationship to label the type labels of each clustering cluster.
[0009] Furthermore, the vector duplicate checking steps include: Generate file fingerprint vectors according to the file type, calculate the Jaccard similarity between the file fingerprint vectors, and judge whether the files are duplicates; Set a similarity threshold. If the similarity exceeds the threshold, mark it as a duplicate file.
[0010] Furthermore, the calculation formula for the importance index of metadata is as follows: ; Wherein, represents the importance index, represents the number of accesses, represents the maximum number of accesses, represents the current time, represents the file creation time, represents the file size, represents the most recent modification time of the file, and and represent weight coefficients.
[0011] Furthermore, the optimization method further includes: When a new file is written, write verification is performed; if the permission passes, it is written to the external hard drive, and the NTFS index table and the metadata B+ tree are updated; otherwise, writing to the external hard drive is rejected; When a file is deleted, the NTFS index table and the metadata B+ tree are updated.
[0012] Furthermore, the write verification includes: Calculate the importance index corresponding to the metadata of the new file; Find the left and right metadata closest to the metadata B+ tree according to the importance index; Perform vector duplicate checking on the metadata of the new file, the left metadata, and the right metadata; if the duplicate checking passes, the write permission is obtained; otherwise, the write permission is rejected.
[0013] Furthermore, the write verification further includes: When only the left metadata can be found according to the importance index of the metadata of the new file, and the importance index is less than the upper limit of the importance index, the write permission is obtained; otherwise, the write permission is rejected; When only the right metadata can be found according to the importance index of the metadata of the new file, and the importance index is greater than the lower limit of the importance index, the write permission is obtained; otherwise, the write permission is rejected.
[0014] The beneficial effects of the present invention are: 1. Removing invalid data through the XGBoost model can effectively improve the quality of the dataset, ensuring that the subsequent clustering process only processes valid and representative data. By accurately classifying valid and invalid data, XGBoost maximizes the classification accuracy, thus avoiding the interference of invalid data on the clustering results and the decline in clustering quality. On this basis, using the Gaussian mixture model for unsupervised clustering can perform efficient clustering grouping according to the feature vectors of file metadata, ensuring that the clustering results are more refined and accurate.
[0015] 2. By generating file fingerprint vectors and using Jaccard similarity to check for vector duplicates, it is possible to efficiently detect and identify duplicate content between files. This method accurately determines whether files are duplicates by calculating the similarity of fingerprint vectors and sets a similarity threshold to ensure that duplicate files can be promptly identified and marked. Combining the updated NTFS index table and the secondary metadata index chain, this process significantly reduces disk space redundancy, optimizes storage management, and improves the operating efficiency of the file system.
[0016] 3. The importance index quantifies the actual value of a file in the storage system by comprehensively considering multiple key factors such as the access frequency of the file, file size, creation time, and most recent modification time. By assigning weights to these factors, the importance index can dynamically evaluate the importance of files, ensuring that the system preferentially retains high-value and frequently used files and automatically clears low-value or infrequently accessed files. The introduction of the importance index not only improves the utilization efficiency of storage space, avoiding the occupation of precious resources by invalid files, but also enhances the intelligent management ability of the system. Brief Description of the Drawings
[0017] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of the mobile hard disk storage space optimization method based on AI algorithms provided by the present invention; Figure 2 is another flowchart of the mobile hard disk storage space optimization method based on AI algorithms provided by the present invention. Detailed Embodiments
[0018] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0019] The mobile hard disk storage space optimization method based on AI algorithms, as Figure 1 shown, includes: S110: Scan the file system of the external hard drive, obtain the metadata of all files through the NTFS index table, and establish a first-level metadata index chain; S120: Use the pre-trained XGBoost model to perform binary classification on the metadata. The classification includes valid data and invalid data. Delete the invalid data and its metadata, and update the NTFS index table and the first-level metadata index chain; Further, the pre-training steps of the XGBoost model include: Collect the metadata of the external hard drive files to construct a training data set; Divide the file metadata into two categories: valid data and invalid data and label them to form a supervised learning training set; Adopt the K-fold cross-validation method to optimize the hyperparameters of the XGBoost model, and use the gradient boosting decision tree for training to maximize the classification accuracy; When the improvement of the F1-score of the validation set of the XGBoost model in 10 consecutive iterations is less than the error threshold, the model training is completed.
[0020] Collect all the file metadata information in the external hard drive. The metadata information includes, but is not limited to, features such as file name, file size, file type, creation time, modification time, access frequency, path level depth, extension name, and approximate duplication degree, etc., to construct an initial training data set; According to the actual business scenario requirements, divide the file metadata in the training data set into two categories, "valid data" and "invalid data", either by manual annotation or combined with historical cleaning experience, to form a training sample set with supervised learning attributes, where valid data refers to file metadata with use value or high importance, and invalid data refers to file metadata that has not been accessed for a long time, has a high repetition rate, occupies space but has no actual value; Adopt the K-fold cross-validation method (K = 5 or K = 10), perform stratified sampling and division on the training samples, cycle train the XGBoost model, evaluate the model performance during each fold of validation, and dynamically adjust the hyperparameters of the model, including learning rate, maximum depth, subsample ratio, column sampling ratio, and regularization parameter, to maximize the classification accuracy of the model on the validation set; After completing the preliminary training and hyperparameter adjustment, introduce an early stopping strategy into the training process, and set the improvement amplitude of the F1-score of the model on the validation set in 10 consecutive iterations to be lower than the preset error threshold (such as 0.001) as the termination condition; If this condition is met, it is determined that the model converges and the training process stops.
[0021] The validity binary classification of the metadata of the external hard drive files is performed through the XGBoost model, and combined with the K-fold cross-validation and hyperparameter optimization in the pre-training step, the model can fully learn the discriminant features of valid data and invalid data during the training process. At the same time, using the improvement of the F1-score of the continuous iterative validation set being less than the error threshold as the stopping condition can effectively prevent the model from overfitting or underfitting, ensuring that the model has good generalization ability. This method not only improves the accuracy and stability of invalid data recognition, but also can dynamically update the NTFS index table and the metadata index chain during the cleaning process, thereby significantly improving the storage space utilization efficiency, reducing the disk redundancy burden, and laying a foundation for the subsequent clustering and deduplication steps.
[0022] S130: Cluster the metadata through the Gaussian mixture model, establish multiple secondary metadata index chains according to the clustering results, and delete the primary metadata index chain; Further, the Gaussian mixture model clustering step includes: Select the feature vector of the file metadata as the clustering input; Use the independent component analysis method to reduce the dimension of the feature vector to obtain a reduced-dimensional feature vector; Use the Gaussian mixture model to perform unsupervised clustering on the reduced-dimensional feature vector, use the expectation-maximization algorithm to estimate the parameters of each Gaussian distribution, and dynamically select the optimal number of clustering clusters through the Bayesian information criterion; Calculate the internal file similarity of each clustering cluster, verify the clustering quality, and combine the clustering results with the file type feature mapping relationship to label the type label of each clustering cluster.
[0023] Select the preprocessed and labeled valid file metadata, and convert it into a numerical feature vector. The feature vector includes, but is not limited to, file size, access frequency, modification period, directory depth, and other quantization parameters reflecting file behavior and characteristics. This feature vector serves as the input data for Gaussian mixture model clustering. To reduce the computational complexity of the high-dimensional feature space and improve the clustering efficiency, an independent component analysis method is used to perform dimensionality reduction on the feature vector, extracting the principal component variables that are statistically independent of each other to obtain the dimensionality-reduced feature vector set. The ICA algorithm separates the mixed signals by maximizing the non-Gaussianity measure, which helps to enhance the separability and feature representativeness of the clustering input data. Use the Gaussian mixture model to perform unsupervised clustering on the dimensionality-reduced feature vector set, and use the expectation-maximization algorithm to iteratively estimate parameters such as the mean, covariance matrix, and mixing weights of each Gaussian distribution. The EM algorithm gradually optimizes the model convergence by calculating the expected value of the hidden variable distribution in the E step and re-maximizing the parameter likelihood function in the M step. By dynamically adjusting the number of clustering clusters and using the Bayesian information criterion to evaluate each candidate model, the minimum value of BIC is used as the optimal model selection criterion to automatically determine the optimal number of clustering clusters. This process effectively avoids the deviation problem caused by artificially specifying the number of clustering clusters. Calculate the average similarity within each clustering cluster and the separation degree between clusters to verify the clustering quality. Combine the clustering results with the existing file type feature mapping relationship, and assign corresponding type labels (such as "picture type", "document type", "audio type", "temporary file type", etc.) to each clustering cluster to provide a data basis for subsequent index chain construction and deduplication strategy.
[0024] By introducing the Gaussian mixture model for unsupervised clustering of file metadata, combined with dimensionality reduction using the independent component analysis method, key information is effectively retained and the computational complexity is reduced. Use the expectation-maximization algorithm to dynamically estimate the parameters of each Gaussian distribution, and adaptively select the optimal number of clustering clusters through the Bayesian information criterion to ensure that the clustering results are more scientific and accurate. At the same time, verify the clustering quality by calculating the similarity of files within the clustering cluster, and complete the labeling of the clustering cluster labels in combination with file type features, making the clustering results interpretable and targeted.
[0025] S140: Obtain the file fingerprint vectors of the metadata on each secondary metadata index chain and perform vector duplicate checking. Delete the duplicate files and their metadata according to the duplicate checking results, and update the NTFS index table and the secondary metadata index chain. Furthermore, the vector duplicate checking step includes: Generate file fingerprint vectors according to the file type, calculate the Jaccard similarity between the file fingerprint vectors, and determine whether the files are duplicates. Set a similarity threshold. If the similarity exceeds the threshold, mark it as a duplicate file.
[0026] According to the file type corresponding to each file metadata, select a matching fingerprint extraction algorithm to generate a file fingerprint vector. For example, for text files, the Shingling algorithm combined with the MinHash method is used to generate a fixed-length fingerprint vector; for image files, the perceptual hashing (pHash) or differential hashing (dHash) algorithm is used to generate an image fingerprint; for audio files, the Mel frequency cepstral coefficients (MFCC) plus hashing encoding is used to generate a fingerprint vector; for compressed packages or binary files, extract byte sequence features and perform segmented hashing processing to form a general fingerprint vector; calculate the Jaccard similarity between the file fingerprint vectors in the same clustering cluster or index chain pairwise; the calculation formula of the Jaccard similarity is: ; wherein, respectively represent the fingerprint vector sets of two files, represents the number of elements in the intersection of the two sets, represents the number of elements in the union; subsequently, a similarity threshold is preset (set to 0.85 in this embodiment), and after the calculation is completed, a judgment is made according to the similarity result. If the Jaccard similarity between any two file fingerprint vectors exceeds the threshold, it is determined that the file is a duplicate file. For duplicate files, mark them as "duplicate status" and update the metadata index chain identification field; finally, according to the duplicate file marking result, combined with the file with a higher importance index that is preferentially retained, delete the duplicate files with low importance and their metadata, and update the NTFS index table and the metadata index chain in a timely manner to ensure the integrity of the file system structure and the accuracy of the deduplication result.
[0027] By generating the corresponding file fingerprint vector according to the file type and using the Jaccard similarity to quantitatively compare the similarity between the fingerprint vectors, cross-type and high-precision file content duplicate detection can be achieved. Setting a reasonable similarity threshold enables the duplicate file marking to have the ability of automation and flexible adjustment, effectively avoiding manual misjudgment and omission. This method not only significantly improves the intelligent level and accuracy of file deduplication, reduces the waste of storage space, but also can improve the overall system performance and storage optimization effect in subsequent index update and file management.
[0028] S150: Calculate the importance index of all metadata in the secondary metadata index chain, establish a metadata B+ tree according to the importance index, and delete the secondary metadata index chain.
[0029] Furthermore, the calculation formula of the importance index of metadata is: ; wherein, represents the importance index, represents the number of access times, represents the maximum number of accesses, represents the current time, represents the file creation time, represents the file size, represents the most recent modification time of the file, 、 and represent the weight coefficients.
[0030] By comprehensively considering multiple key factors such as the access times, file size, file creation time, and most recent modification time of the file, and assigning appropriate weight coefficients to each factor, a quantitative evaluation of the file value is achieved. By combining the access frequency and time correlation, this index can accurately reflect the actual usage value and priority of the file in the current storage system, effectively distinguishing important files from infrequently used or outdated files.
[0031] Embodiment 2 As another alternative embodiment of this application, referring to Figure 2 , this embodiment is mainly an extended solution to the fault rapid location method described in the above Embodiment 1. As shown in Figure 2 , this method includes: S210: Scan the file system of the external hard drive, obtain the metadata of all files through the NTFS index table, and establish a first-level metadata index chain; S220: Use the pre-trained XGBoost model to perform binary classification on the metadata. The classification includes valid data and invalid data. Delete the invalid data and its metadata, and update the NTFS index table and the first-level metadata index chain; S230: Cluster the metadata through the Gaussian mixture model, establish multiple second-level metadata index chains according to the clustering results, and delete the first-level metadata index chain; S240: Obtain the file fingerprint vectors of the metadata on each second-level metadata index chain and perform vector duplicate checking. Delete the duplicate files and their metadata according to the duplicate checking results, and update the NTFS index table and the second-level metadata index chain; S250: Calculate the importance index of all metadata in the second-level metadata index chain, establish a metadata B+ tree according to the importance index, and delete the second-level metadata index chain.
[0032] For the detailed process of steps S210 - S250, reference can be made to the relevant introduction of steps S110 - S150 in Embodiment 1, which will not be elaborated here.
[0033] S260: When a new file is written, perform write verification; if the permission passes, write to the external hard drive and update the NTFS index table and the metadata B+ tree; otherwise, reject writing to the external hard drive; S270: When a file deletion is executed, update the NTFS index table and the metadata B+ tree.
[0034] Further, the write verification includes: Calculate the importance index corresponding to the new file metadata; Find the closest left metadata and right metadata in the metadata B+ tree according to the importance index; Perform vector duplicate checking on the new file metadata, the left metadata, and the right metadata; if the duplicate checking passes, obtain the write permission; otherwise, reject the write permission.
[0035] When it is detected that there is a new file to be written to the external hard drive, extract the metadata information of the new file, and according to the extracted metadata, use a preset importance index calculation formula to calculate the importance index corresponding to the new file metadata; secondly, based on the constructed metadata B+ tree structure, use this importance index as the keyword to perform a range search operation in the B+ tree to locate the left metadata node and the right metadata node closest to this importance index; obtain the new file metadata fingerprint vector, and perform vector duplicate checking with the file fingerprint vectors of the left metadata and the right metadata respectively, and compare the Jaccard similarity results with the left metadata and the right metadata with a preset similarity threshold respectively. If any group of similarity results exceeds the similarity threshold (set to 0.85 in this embodiment), it is determined that the file is a duplicate or highly similar file, and the write permission is rejected; if neither side triggers the threshold condition, that is, the duplicate checking passes, the system determines that the new file can be written, grants the write permission, and updates the NTFS index table and the metadata B+ tree node after the file is written, inserts a new metadata node, and rebalances the B+ tree.
[0036] By calculating the importance index of the new file metadata before file writing, quickly locating adjacent importance index nodes based on the B+ tree structure, and combining the vector duplicate checking method to efficiently compare the new file with adjacent metadata, it is possible to prevent the writing of duplicate or low-value files at the source. This process not only improves the system's automated management ability, reduces the occupation of redundant data, and ensures the efficient use of disk space, but also dynamically maintains the orderliness and accuracy of the metadata B+ tree structure, further optimizing the long-term healthy operation and resource scheduling efficiency of the storage system.
[0037] Further, the write verification further includes: When only the left metadata can be found according to the importance index of the new file metadata and the importance index is less than the importance index upper limit, obtain the write permission; otherwise, reject the write permission; When only the right metadata can be found according to the importance index of the new file metadata and the importance index is greater than the importance index lower limit, obtain the write permission; otherwise, reject the write permission.
[0038] By introducing a boundary condition determination mechanism in the write verification, when only the left or right metadata node can be found, intelligent judgment is made by combining the importance index with the upper and lower limit constraints to ensure that strict write control logic is still available in the case of index boundaries. This design effectively avoids data miswriting or duplicate writing in extreme cases and improves the robustness and intelligence level of the system.
[0039] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for optimizing the storage space of a portable hard drive based on an AI algorithm, characterized in that, Including: Scan the file system of the external hard drive, obtain the metadata of all files through the NTFS index table, and establish a first-level metadata index chain; Use the pre-trained XGBoost model to perform binary classification on the metadata. The classification includes valid data and invalid data. Delete the invalid data and its metadata, and update the NTFS index table and the first-level metadata index chain; Cluster the metadata through the Gaussian mixture model, establish multiple second-level metadata index chains according to the clustering results, and delete the first-level metadata index chain; Obtain the file fingerprint vectors of the metadata on each second-level metadata index chain and perform vector duplicate checking. Delete the duplicate files and their metadata according to the duplicate checking results, and update the NTFS index table and the second-level metadata index chain; Calculate the importance index of all metadata in the second-level metadata index chain, establish a metadata B+ tree according to the importance index, and delete the second-level metadata index chain.
2. The method for optimizing the storage space of a mobile hard disk based on an AI algorithm according to claim 1, wherein The pre-training steps of the XGBoost model include: Collect the metadata of the external hard drive files and construct a training data set; Divide the file metadata into two categories: valid data and invalid data and label them to form a supervised learning training set; Adopt the K-fold cross-validation method to optimize the hyperparameters of the XGBoost model, and use the gradient boosting decision tree for training to maximize the classification accuracy; When the improvement of the F1-score of the validation set in the XGBoost model is less than the error threshold in 10 consecutive iterations, the model training is completed.
3. The method for optimizing the storage space of a mobile hard disk based on the AI algorithm according to claim 1, wherein, The clustering steps of the Gaussian mixture model include: Select the feature vectors of the file metadata as the clustering input; Adopt the independent component analysis method to reduce the dimension of the feature vectors to obtain the reduced-dimensional feature vectors; Adopt the Gaussian mixture model to perform unsupervised clustering on the reduced-dimensional feature vectors, use the expectation maximization algorithm to estimate the parameters of each Gaussian distribution, and dynamically select the optimal number of clustering clusters through the Bayesian information criterion; Calculate the internal file similarity of each clustering cluster, verify the clustering quality, and combine the clustering results with the file type feature mapping relationship to label the type labels of each clustering cluster.
4. The method for optimizing the storage space of a mobile hard disk based on an AI algorithm according to claim 1, wherein The vector duplicate checking steps include: Generate file fingerprint vectors according to the file type, calculate the Jaccard similarity between the file fingerprint vectors, and judge whether the files are repeated; Set a similarity threshold. If the similarity exceeds the threshold, mark it as a repeated file.
5. The method for optimizing the storage space of a mobile hard disk based on the AI algorithm according to claim 1, wherein, The calculation formula for the importance index of the metadata is: ; Among them, represents the importance index, represents the number of accesses, represents the maximum number of accesses, represents the current time, represents the file creation time, represents the file size, represents the file's most recent modification time, 、 and represent the weight coefficients.
6. The method for optimizing the storage space of a mobile hard disk based on an AI algorithm according to claim 1, characterized in that, The optimization method further includes: When a new file is written, perform a write verification; if the permission passes, write it to the external hard drive and update the NTFS index table and the metadata B+ tree; otherwise, reject writing to the external hard drive; When a file is deleted, update the NTFS index table and the metadata B+ tree.
7. The method for optimizing the storage space of a mobile hard disk based on the AI algorithm according to claim 6, wherein The write verification includes: Calculate the importance index corresponding to the new file metadata; Find the left and right metadata closest to the new file metadata in the metadata B+ tree according to the importance index; Perform vector duplicate checking on the new file metadata, the left metadata, and the right metadata; if the duplicate checking passes, obtain the write permission; otherwise, reject the write permission.
8. The method for optimizing the storage space of a mobile hard disk based on the AI algorithm according to claim 7, characterized in that, The write verification further includes: When only the left metadata can be found according to the importance index of the new file metadata, and the importance index is less than the upper limit of the importance index, obtain the write permission; otherwise, reject the write permission; When only the right metadata can be found according to the importance index of the new file metadata, and the importance index is greater than the lower limit of the importance index, obtain the write permission; otherwise, reject the write permission.