File storage management method and system based on computer system resource use

By monitoring the system resource usage and file access frequency in real time, evaluating the accuracy of priority determination of AI algorithms, and dynamically adjusting weights, the problems of misjudgment risks and decision deviations in the AI ​​algorithm when determining file priority are solved, and the system's resource utilization efficiency and data migration stability are improved.

CN119938601AInactive Publication Date: 2025-05-06WEIHAI OCEAN VOCATIONAL COLLEGE
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202411977210.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, AI algorithms have a risk of misjudgment when determining file priority, which may lead to false deletion or mistransfer of key files, affecting system operation, and the self-learning characteristics of the AI ​​model may cause data migration decisions to gradually deviate from the initial setting, reducing the reliability of the system.

Method used

By monitoring the resource usage of computer systems in real time, setting migration thresholds to enter resource management mode, monitoring file access frequency in high-load and low-load scenarios in real time, calculating the comprehensive access frequency exception index and file call frequency fluctuation index, evaluating the accuracy of priority determination of AI algorithms, and dynamically adjusting the weight of file access frequency and recovery call frequency.

Benefits of technology

It significantly improves the system's effectiveness in storage resource utilization, file priority determination accuracy and data migration stability, reduces the risk of deviation in AI judgment, improves the intelligence level of data migration decisions, and enhances the long-term reliability and resource management efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938601A_ABST
    Figure CN119938601A_ABST
Patent Text Reader

Abstract

The invention discloses a file storage management method and system based on computer system resource use, and particularly relates to the technical field of file storage management. Triggering a data migration operation through a preset storage space upper limit, distinguishing file access characteristics in high and low load scenes, dynamically evaluating the accuracy of file priority judgment by an AI algorithm, judging whether priority judgment is reasonable or not by calculating an access frequency abnormal index and calling frequency fluctuation, and dividing judgment results into an accurate class and an inaccurate class; for an accurate judgment result, a management strategy of high and low priority files is further optimized; and for the inaccurate judgment result, the weight of the file access frequency and the weight of the recovery calling frequency are fed back to the AI model, the judgment capability of the AI model is continuously optimized, the risk of file priority misjudgment is effectively reduced, and the accuracy of a data migration decision and the stability of system storage resources are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of file storage management, and in particular to a file storage management method and system based on computer system resource usage. Background Art

[0002] File storage management based on computer system resource usage is a method for optimizing file storage management based on the actual usage of system resources. It dynamically adjusts the allocation and management of file storage by monitoring resources such as system memory, CPU load, and disk usage to ensure efficient use of resources. In specific implementations, file storage management based on resource usage usually involves automated storage allocation, file compression, and archiving operations. For example, when the system detects that the storage space is approaching the upper limit, it can automatically trigger data cleanup or migrate low-priority files to the cloud or external storage to free up local space. Specifically, the system will first identify the priority of the file. For example, high-priority files may be data that is currently in use or frequently accessed by the system, while low-priority files may be old files that have not been used for a long time, temporary files, or backup files. For these low-priority files, the system can automatically move them to external storage media such as cloud storage or external hard drives. By migrating these low-priority files, the system can free up valuable local storage space, thereby ensuring that the access to high-priority files and the storage of new data are not affected. This automated management process not only saves storage resources, but also makes the overall operation of the system more efficient and stable.

[0003] The prior art has the following deficiencies:

[0004] Some file storage management systems may use AI algorithms to intelligently determine and migrate file priorities. However, AI algorithms may have the risk of misjudgment under abnormal circumstances. For example, certain types of files are considered low priority but are actually system-critical files. If the AI ​​algorithm makes mistakes in identifying file features, it may lead to the accidental deletion or migration of critical files, affecting system operation. In addition, the self-learning characteristics of the AI ​​model may also cause data migration decisions to gradually deviate from the initial settings. For example, if AI mistakenly learns to define certain critical files as low priority, the system will repeatedly migrate these files by mistake, resulting in the unpredictability of long-term system operation and a significant reduction in data reliability. Once such an error occurs, it will be very complicated to repair, and may even require resetting the entire AI model and priority settings. Summary of the invention

[0005] The purpose of the present invention is to provide a file storage management method and system based on computer system resource usage to solve the shortcomings of the background technology.

[0006] In order to achieve the above object, the present invention provides the following technical solution: a file storage management method based on computer system resource usage, comprising the following steps:

[0007] S1: Based on the file storage management system, the local storage resources are monitored in real time. The upper limit of the storage space usage is preset as the migration threshold for triggering the data migration operation. When the file storage management storage usage exceeds the migration threshold, the resource management mode is entered;

[0008] S2: Monitor the actual access frequency of migrated files in the high-load scenario and the low-load scenario in the resource management mode in real time, record the actual access status of all low-priority files, count the proportion of files whose access frequency exceeds the access frequency threshold, and calculate the comprehensive access frequency anomaly index in the high-load scenario and the low-load scenario;

[0009] S3: After the file is migrated to the cloud, the file storage management system records each recovery call operation and continuously monitors the call status. Based on the fluctuation of the number of low-priority file calls within a fixed observation period, the abnormal degree of the low-priority file call frequency is determined;

[0010] S4: Based on the abnormal degree of low-priority file call frequency and the access comprehensive frequency abnormality index, the accuracy of the AI ​​algorithm in determining the file priority is evaluated. Based on the evaluation results, the file priority determination results are divided into accurate determination results and inaccurate determination results;

[0011] S5: Based on the access patterns of the files corresponding to the accurate judgment results, further refine the high-priority and low-priority management strategies. For inaccurate judgment results, feed the judgment result information back to the AI ​​model, and dynamically adjust the weights of file access frequency and recovery call frequency based on the accuracy abnormality of the inaccurate judgment results within a fixed time period.

[0012] Preferably, in S2, the access frequency anomaly rate is calculated and recorded in the high-load and low-load scenarios respectively, and after comprehensive analysis, the comprehensive access frequency anomaly index in the high-load and low-load scenarios is calculated. The calculation method of the comprehensive access frequency anomaly index is:

[0013] Collect and organize the access frequency anomaly rates in high-load and low-load scenarios within the W time period to form time series data, determine the smoothing coefficient α, and set the value range to 0<α<1. Set the initial anomaly index S0 to the anomaly rate value at the first time point as the starting point of the smoothing calculation. Calculate the comprehensive access frequency anomaly index MS at each time point t. The expression is: MS = α × access frequency anomaly rate t +(1-α)×S t-1 ;MS t-1 Indicates the comprehensive access frequency anomaly index at time t-1.

[0014] Preferably, in S3, the file call frequency fluctuation index is generated after analyzing the fluctuation of the number of low priority file calls within a fixed observation period, and the method for obtaining the file call frequency fluctuation index is:

[0015] Collect the call frequency time series data of the file, set X = {x1, x2, ..., x n}, where x n represents the number of calls at time n, selects the wavelet function, and selects the decomposition layer J according to the frequency distribution and fluctuation characteristics of the data. The number of layers determines the scale of the wavelet decomposition. The call frequency time series X is decomposed by multi-layer wavelet to obtain detail components and approximate components at different scales. For the j-th layer of decomposition, the call frequency series X is decomposed into an approximate coefficient A J and several detail coefficients D j , the expression is: A J Indicates the approximate components of the signal, reflecting the low-frequency components, D j Represents the detail component of the jth layer, reflecting the high-frequency component; selects the detail coefficient of the medium and high-frequency layers to reflect the short-term fluctuation of the call frequency, and calculates the file call frequency fluctuation index, the expression is: Where KM is the file call frequency fluctuation index.

[0016] Preferably, in S4, the accuracy of the AI ​​algorithm in determining the file priority is evaluated based on the abnormal degree of the low-priority file call frequency and the access comprehensive frequency abnormality index. According to the evaluation result, the file priority determination result is divided into an accurate determination result and an inaccurate determination result, specifically:

[0017] The comprehensive access frequency anomaly index and the file call frequency fluctuation index are converted into the first eigenvector, and the first eigenvector is used as the input of the machine learning model. The machine learning model uses each group of first eigenvectors to predict the accuracy value label of the AI ​​algorithm for file priority judgment as the prediction target, and takes minimizing the sum of the prediction errors of the accuracy value labels of all AI algorithms for file priority judgment as the training target. The machine learning model is trained until the sum of the prediction errors converges and the model training is stopped. The accuracy value of the AI ​​algorithm for file priority judgment is determined according to the model output results, wherein the machine learning model is a polynomial regression model.

[0018] Preferably, in S4, the acquired accuracy value of the file priority determination by the AI ​​algorithm is compared with the readiness standard threshold. If the accuracy value of the file priority determination by the AI ​​algorithm is greater than or equal to the readiness standard threshold, it means that the accuracy of the file priority determination by the AI ​​algorithm is high. In this case, a high-accuracy determination signal is generated, and the determination result of the file priority is classified as an accurate determination result. If the accuracy value of the file priority determination by the AI ​​algorithm is less than the readiness standard threshold, it means that the accuracy of the file priority determination by the AI ​​algorithm is low. In this case, a low-accuracy determination signal is generated, and the determination result of the file priority is classified as an inaccurate determination result.

[0019] Preferably, in S5, for inaccurate judgment results, that is, the accuracy value of the file priority judgment of the AI ​​algorithm generated within a fixed time period is less than the readiness standard threshold, the judgment result information is fed back to the AI ​​model, and the accuracy values ​​of the file priority judgment of the AI ​​algorithm generated within a subsequent fixed time period that are less than the readiness standard threshold are collected, and a data set is established, the mean and standard deviation of the data set are calculated, and after analyzing the abnormal degree of accuracy of the inaccurate judgment results within the fixed time period, the weights of the file access frequency and the recovery call frequency are dynamically adjusted according to the analysis results.

[0020] Preferably, if the accuracy value mean is greater than or equal to the reference threshold of the accuracy value mean, and the accuracy value standard deviation is less than the reference threshold of the accuracy value standard deviation, the AI ​​model has high accuracy and stable accuracy in most cases, maintains the current weight setting, reduces the sensitivity of file access frequency and recovery call frequency, and maintains system stability;

[0021] If the mean accuracy value is greater than or equal to the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is greater than or equal to the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is high, but the accuracy fluctuation is large. Increase the file access frequency weight and reduce the recovery call frequency weight;

[0022] If the mean accuracy value is less than the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is greater than or equal to the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is low and the accuracy fluctuates greatly. Increase the weight of the file access frequency and reduce the weight of the recovery call frequency.

[0023] If the mean accuracy value is less than the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is less than the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is low but the fluctuation is small, and the weights of increasing the file access frequency and the recovery call frequency are balanced.

[0024] The present invention also provides a file storage management system based on computer system resource usage, including a storage monitoring module, a load scenario management module, a call frequency monitoring module, a determination accuracy evaluation module and a priority management optimization module;

[0025] Storage monitoring module: It monitors local storage resources in real time based on the file storage management system. It sets the upper limit of storage space usage as the migration threshold for triggering data migration operations. When the file storage management storage usage exceeds the migration threshold, it enters the resource management mode.

[0026] Load scenario management module: monitors the actual access frequency of migrated files in high-load scenarios and low-load scenarios in the resource management mode in real time, records the actual access status of all low-priority files, counts the proportion of files whose access frequency exceeds the access frequency threshold, and calculates the comprehensive access frequency anomaly index in high-load scenarios and low-load scenarios;

[0027] Call frequency monitoring module: After the files are migrated to the cloud, the file storage management system records each recovery call operation and continuously monitors the call status. It determines the abnormality of the low-priority file call frequency based on the fluctuation of the number of low-priority file calls within a fixed observation period;

[0028] Judgment accuracy assessment module: evaluates the accuracy of the AI ​​algorithm's judgment of file priority based on the abnormal degree of low-priority file call frequency and the access comprehensive frequency abnormality index. Based on the evaluation results, the file priority judgment results are divided into accurate judgment results and inaccurate judgment results;

[0029] Priority management optimization module: Based on the access patterns of files corresponding to accurate judgment results, high-priority and low-priority management strategies are further refined. For inaccurate judgment results, the judgment result information is fed back to the AI ​​model. Based on the abnormal degree of accuracy of the inaccurate judgment results within a fixed time period, the weights of file access frequency and recovery call frequency are dynamically adjusted.

[0030] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0031] 1. The present invention significantly improves the system's storage resource utilization, file priority determination accuracy, and data migration stability through an intelligent file storage management method based on the use of computer system resources. By real-time monitoring of storage resources and setting migration thresholds to trigger resource management modes, the solution effectively avoids system risks caused by insufficient storage space. At the same time, for detailed monitoring of high and low load scenarios, the solution can dynamically analyze the access characteristics of low-priority files, calculate the access frequency anomaly index and the call frequency fluctuation index, and accurately judge changes in file access requirements. Based on these indicators, the system evaluates the priority determination accuracy of the AI ​​algorithm, and feeds back accurate and inaccurate determination results to the AI ​​model to ensure that the priority of files in each scenario is reasonable.

[0032] 2. The present invention uses the comprehensive access frequency anomaly index and call frequency fluctuation index as feature vectors and uses a machine learning model to iteratively optimize the accuracy of file priority determination, thereby ensuring continuous improvement of priority determination. For inaccurate determination results, by introducing a feedback mechanism and dynamic weight adjustment, the system automatically adjusts the weights of file access frequency and recovery call frequency based on the accuracy mean and standard deviation, gradually improving the stability and accuracy of the AI ​​model, effectively reducing the risk of AI judgment deviation, and improving the intelligence level of data migration decision-making, thereby enhancing the long-term reliability and resource management efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0034] Figure 1 The figure is a flow chart of the method of the present invention.

[0035] Figure 2 It is a system module diagram of the present invention. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] Example 1, please refer to Figure 1 and Figure 2As shown, the file storage management method based on computer system resource usage described in this embodiment includes the following steps:

[0038] S1: Based on the file storage management system, the local storage resources are monitored in real time. The upper limit of the storage space usage is preset as the migration threshold for triggering the data migration operation. When the file storage management storage usage exceeds the migration threshold, the resource management mode is entered;

[0039] S2: Monitor the actual access frequency of migrated files in the high-load scenario and the low-load scenario in the resource management mode in real time, record the actual access status of all low-priority files, count the proportion of files whose access frequency exceeds the access frequency threshold, and calculate the comprehensive access frequency anomaly index in the high-load scenario and the low-load scenario;

[0040] S3: After the file is migrated to the cloud, the file storage management system records each recovery call operation and continuously monitors the call status. Based on the fluctuation of the number of low-priority file calls within a fixed observation period, the abnormal degree of the low-priority file call frequency is determined;

[0041] S4: Based on the abnormal degree of low-priority file call frequency and the access comprehensive frequency abnormality index, the accuracy of the AI ​​algorithm in determining the file priority is evaluated. Based on the evaluation results, the file priority determination results are divided into accurate determination results and inaccurate determination results;

[0042] S5: Based on the access patterns of the files corresponding to the accurate judgment results, further refine the high-priority and low-priority management strategies. For inaccurate judgment results, feed the judgment result information back to the AI ​​model, and dynamically adjust the weights of file access frequency and recovery call frequency based on the accuracy abnormality of the inaccurate judgment results within a fixed time period.

[0043] In S1, the local storage resources are monitored in real time based on the file storage management system. The upper limit of the storage space usage is preset as the migration threshold for triggering the data migration operation. When the file storage management storage usage exceeds the migration threshold, the resource management mode is entered, specifically:

[0044] When the file storage management system starts, the storage monitoring module is initialized and continuous monitoring of local storage resources (including storage space, memory, CPU, etc.) is started. The monitoring frequency is set, such as checking the storage resource usage once every minute or every second, to ensure that changes in storage capacity can be identified in a timely manner. Key parameters related to storage resources are collected, including the current storage space usage, remaining capacity, file distribution, etc., so as to subsequently determine the pressure of storage resources. System resource status logs are generated and updated in real time, including file storage status, memory and CPU usage, etc. Logs can be used to trace and analyze the resource usage of the system in different time periods.

[0045] The system administrator or automated program presets an upper limit on the use of storage space (such as 90% of the total capacity) as the migration threshold based on the total capacity of the storage space and the system load. This threshold will serve as the boundary for triggering data migration operations. In scenarios where the load fluctuates greatly, a dynamic threshold adjustment mechanism can be set, such as lowering the threshold to 80% during peak periods and restoring it to 90% during off-peak periods, so as to manage storage resources more flexibly and prevent storage pressure caused by emergencies. To prevent the system from being overloaded when the migration threshold is reached, an alarm can be activated when the storage usage rate approaches the threshold (such as 85%), prompting the administrator or automated system to prepare for data migration.

[0046] The file storage management system regularly analyzes the real-time monitoring data to determine whether the current storage usage exceeds the preset migration threshold. When the storage space usage reaches or exceeds the migration threshold (for example, the usage reaches 90%), the system triggers to enter the resource management mode and starts to perform automated data migration or cleanup operations. When the storage exceeds the threshold, the system can suspend or limit low-priority file access operations to ensure that the normal reading and writing of high-priority files are not affected. This protection measure becomes invalid after entering the resource management mode.

[0047] In resource management mode, the system categorizes files according to pre-set priority rules. Usually, files that are currently accessed frequently are marked as high priority, while files that are not used or accessed less frequently are marked as low priority. The system first cleans up low-priority temporary files and cache files, deleting them from local storage to free up some space. For some low-priority uncompressed files, compression operations can be performed to reduce the space occupied. After cleaning and compression, if the storage space is still tight, the system will select the data that needs to be migrated from the low-priority files.

[0048] The system selects the migration target (such as the cloud or external storage hard drive) according to the preset strategy and starts the data transfer. The system migrates files through a secure transmission protocol (such as HTTPS or FTP) to ensure the security of the files during the transmission process.

[0049] After the migration is complete, the system updates the directory index of the file and replaces the local storage path with the external storage path to ensure that the user or system can quickly locate the file when it is subsequently accessed. The system generates a migration operation log, recording detailed information such as the name of the migrated file, the original storage location, the migration target location, and the migration time for future query and maintenance. After the migration operation is completed, the system re-checks the usage of local storage to determine whether it is below the migration threshold; if it is not below the threshold, further cleanup or migration operations can be performed according to the system settings.

[0050] When the local storage resource usage drops below the migration threshold, the system automatically exits the resource management mode and returns to the daily monitoring state. If the read and write permissions of low-priority files were previously restricted, the system will remove the restrictions and restore normal file access permissions to ensure smooth subsequent file read and write operations. The operation records under the resource management mode are organized and archived to form a detailed log report, which is convenient for system administrators to analyze and optimize resource management strategies in the future.

[0051] S2: Monitor the actual access frequency of migrated files in high-load scenarios and low-load scenarios in the resource management mode in real time, record the actual access status of all low-priority files, count the proportion of files whose access frequency exceeds the access frequency threshold, and calculate the access frequency abnormality rate.

[0052] High-load scenarios refer to time periods when system resource usage (especially storage, CPU, and memory) is high and file access is frequent. This scenario usually occurs when data volume surges, user activity is high, or system tasks are intensive.

[0053] Common characteristics of high-load scenarios: High CPU, memory, and storage resource usage: For example, CPU and memory usage exceeds 80%, and storage space exceeds the threshold. Frequent file read and write operations: For example, the system performs batch file processing tasks, large amounts of data backup, or users access files in large quantities. Typical time periods: For example, daytime on weekdays, the system's daily backup period, and when large amounts of data are uploaded or downloaded.

[0054] Set high load thresholds (e.g. 85%) for CPU, memory, and storage usage. Once the system resource usage exceeds this threshold, it is considered high load. Time division: Based on historical data analysis, you can divide the system into typical high load time periods (e.g. 10:00-12:00, 14:00-18:00 every day), and enable high load monitoring by default during these time periods.

[0055] A low-load scenario refers to a period of time when system resource usage is low and file access frequency is not high, which usually occurs during off-peak hours or when user activity is low.

[0056] Common characteristics of low-load scenarios: Low resource usage: For example, CPU and memory usage is less than 40%, and storage space usage is below the threshold. Fewer file read and write operations: For example, during non-working hours, or when the business system is idle. Typical time period: Non-working hours (such as evenings and weekends) or rest time after system tasks are completed.

[0057] Set low load thresholds for CPU, memory, and storage (such as CPU and memory usage are both below 40%). Once the system resource usage is below this threshold, it is considered low load. Based on historical data, divide the system's typical low load time periods (such as 20:00-8:00 in the evening) and enable low load monitoring by default during these time periods.

[0058] After dividing the high-load and low-load scenarios, monitor and record the actual access to low-priority files in real time, focusing on whether the access frequency exceeds the preset threshold. In high-load and low-load scenarios, record the access logs of all low-priority files separately, including the file access timestamp: the specific time of each access. Access source: record whether it is a user access, system task call, or other process call. Number of accesses: the number of multiple accesses to a single file in the same time period. Set the access frequency threshold according to the file type and scenario. For example, in a high-load scenario, the access frequency threshold can be set to more than 5 times per hour; in a low-load scenario, the access frequency threshold can be set to more than 1 time per hour. The specific threshold should be flexibly adjusted according to the system business needs and file access rules.

[0059] Calculate the access frequency exception rate, count the proportion of files that exceed the threshold, and in each scenario, count the number of low-priority files whose access frequency exceeds the preset threshold. Count the total number of low-priority files in each scenario. Use the formula to calculate the access frequency exception rate of low-priority files in high-load and low-load scenarios:

[0060] The access frequency anomaly rate is calculated and recorded in high-load and low-load scenarios respectively. After comprehensive analysis, the comprehensive access frequency anomaly index in high-load and low-load scenarios is calculated. The calculation method of the comprehensive access frequency anomaly index is:

[0061] Collect and organize the access frequency anomaly rates in high-load scenarios and low-load scenarios within the W time period to form time series data, and confirm the smoothing coefficient α. This parameter controls the degree of smoothing, and the value range is 0<α<1. Usually, a larger α value (such as 0.7) means more sensitivity to recent anomaly rate fluctuations, and a smaller α value (such as 0.3) pays more attention to long-term trends. The initial anomaly index S0 is usually set to the anomaly rate value at the first time point as the starting point for the smoothing calculation. Calculate the comprehensive access frequency anomaly index MS at each time point t, and the expression is: MS = α × access frequency anomaly rate t +(1-α)×S t-1 ;MS t-1 Indicates the comprehensive access frequency anomaly index at time t-1.

[0062] A higher comprehensive access frequency anomaly index indicates that the access frequency anomaly rate in the current time period is high, which may reflect potential problems such as incorrect priority determination and insufficient resources. A lower comprehensive access frequency anomaly index indicates that the anomaly rate in the current time period is low, the file priority determination of the AI ​​algorithm is more accurate, and the system resource usage is relatively normal.

[0063] S3: After the files are migrated to the cloud, the file storage management system records each recovery call operation and continuously monitors the call status. It determines the abnormality of the low-priority file call frequency based on the fluctuation of the number of low-priority file calls within a fixed observation period.

[0064] The file storage management system records the specific information of the call operation every time the cloud file is called, including: Call timestamp: Accurately record the time when the file is called. Record the source of the file call request (such as application, user, system process). If the file is called multiple times in a short period of time, it can be accumulated at a set time interval (such as every hour). Store the call record in the database or monitoring log for subsequent analysis.

[0065] According to system requirements and file call characteristics, set a fixed observation period. For example, daily, weekly or monthly observation periods can better monitor the fluctuation of the number of calls. Within the observation period, set the threshold of the file call frequency. For example, when the observation period is one week, set the normal range of the number of low-priority file calls, such as the call frequency should not exceed 5 times; exceeding this may indicate that the priority of the file is incorrect.

[0066] At the end of each observation period, count the total number of times low-priority files were called for recovery during that period. Compare the number of calls in different periods to detect fluctuations in the frequency of calls. For example, if a file is called significantly more frequently in one period than in previous periods, it may indicate that its priority is wrong. Set a fluctuation range threshold. For example, if the number of calls in the current period increases by 50% or more than the previous period, it is considered an abnormal fluctuation.

[0067] After analyzing the fluctuation of the number of low-priority file calls within a fixed observation period, a file call frequency fluctuation index is generated. The method for obtaining the file call frequency fluctuation index is as follows:

[0068] Collect the call frequency time series data of the file, set X = {x1, x2, ..., x n}, where x n Indicates the number of calls at time n. Standardize or denoise the time series data to remove outliers or noise to improve the accuracy of wavelet decomposition.

[0069] Select a suitable wavelet function. Commonly used ones include Daubechies wavelet (db4), Haar wavelet, etc. Different wavelet functions are suitable for different feature extraction requirements. Generally speaking, db4 wavelet is more commonly used in signal analysis. Select the decomposition layer J according to the frequency distribution and fluctuation characteristics of the data. The number of layers determines the scale of wavelet decomposition. Usually it can range from 2 to 5 layers, and the number of decomposition layers is adjusted according to the periodicity and volatility of the call data. Perform multi-layer wavelet decomposition on the call frequency time series X to obtain detail components and approximate components at different scales. For the j-th layer decomposition, the call frequency sequence X is decomposed into an approximate coefficient A J and several detail coefficients D j , the expression is: A J Indicates the approximate components of the signal, reflecting the low-frequency (long-term trend) components, D j Represents the detail component of the jth layer, reflecting the high-frequency (short-term fluctuation) component. Select the detail coefficient of the medium and high-frequency layer (such as D j To D j-1 ) to reflect the short-term fluctuation of the call frequency and calculate the file call frequency fluctuation index. The expression is: Where KM is the file call frequency fluctuation index.

[0070] When the file call frequency fluctuation index is larger, it means that the call frequency of low-priority files fluctuates more violently, that is, during the observation period, the number of calls of the file fluctuates up and down to a large extent. In this case, the access demand of the file may be far beyond expectations, indicating that it may not be reasonable to set it to low priority at present. Frequent calls may mean that the file has a certain real-time or importance to the system or user, and it is worth re-evaluating its priority. An excessively large fluctuation index usually indicates that the usage demand of the file in the short term is unstable and may be affected by specific business needs or access peaks, which will lead to frequent cloud call back operations, increase system resource consumption and data access delays. Therefore, low-priority files with large fluctuation indexes should be considered to be adjusted to high priority so that they can quickly obtain local resource support when they are called.

[0071] On the contrary, when the file call frequency fluctuation index is smaller, it means that the call frequency of low-priority files is relatively stable during the observation period, and there is no obvious access fluctuation. This shows that it is reasonable to set the file to low priority at present, and the system does not need to allocate too many local resources to it because its access demand is low and stable. A smaller fluctuation index means that the file has lower real-time requirements for the system and may be long-term backup or occasionally accessed data. For low-priority files with a smaller fluctuation index, the system can continue to keep them in the cloud or external storage to avoid occupying local resources, while also reducing unnecessary data migration costs. By identifying these files with small fluctuation indexes, the system can allocate storage resources more effectively and improve the efficiency of resource management.

[0072] S4: According to the abnormal degree of low-priority file call frequency and the access comprehensive frequency abnormality index, the accuracy of the AI ​​algorithm in determining file priority is evaluated. Based on the evaluation results, the file priority determination results are divided into accurate determination results and inaccurate determination results.

[0073] The comprehensive access frequency anomaly index and the file call frequency fluctuation index are converted into the first eigenvector, and the first eigenvector is used as the input of the machine learning model. The machine learning model uses each group of first eigenvectors to predict the accuracy value label of the AI ​​algorithm for file priority judgment as the prediction target, and takes minimizing the sum of the prediction errors of the accuracy value labels of all AI algorithms for file priority judgment as the training target. The machine learning model is trained until the sum of the prediction errors converges and the model training is stopped. The accuracy value of the AI ​​algorithm for file priority judgment is determined according to the model output results, wherein the machine learning model is a polynomial regression model.

[0074] The method for obtaining the accuracy value of the AI ​​algorithm for file priority judgment is: from the first eigenvector training data of the trained machine learning model, obtain the corresponding function expression: CQ=F(MS, KM); where F is the output function of the model, MS is the comprehensive access frequency anomaly index, KM is the file call frequency fluctuation index, and CQ is the accuracy value of the AI ​​algorithm for file priority judgment.

[0075] The acquired accuracy value of the AI ​​algorithm for file priority determination is compared with the readiness standard threshold. If the accuracy value of the AI ​​algorithm for file priority determination is greater than or equal to the readiness standard threshold, it means that the AI ​​algorithm has high accuracy in file priority determination. In this case, a high-accuracy determination signal is generated, and the determination result of the file priority is classified as an accurate determination result. If the accuracy value of the AI ​​algorithm for file priority determination is less than the readiness standard threshold, it means that the accuracy of the AI ​​algorithm for file priority determination is low. In this case, a low-accuracy determination signal is generated, and the determination result of the file priority is classified as an inaccurate determination result.

[0076] S5: Based on the access patterns of the files corresponding to the accurate judgment results, further refine the high-priority and low-priority management strategies. For inaccurate judgment results, feed the judgment result information back to the AI ​​model, and dynamically adjust the weights of file access frequency and recovery call frequency based on the accuracy abnormality of the inaccurate judgment results within a fixed time period.

[0077] High-priority files should be stored in local storage first to reduce access delays and ensure the response speed of high-frequency file calls. At the same time, if local storage space is limited, a tiered storage strategy (such as a cache or fast storage device) can be used to store key high-priority files to improve access efficiency.

[0078] Dynamic resource allocation: Frequently called high-priority files: For high-priority files with extremely high access frequency (such as multiple calls per day), you can set a higher cache frequency and prioritize more CPU and memory resources to ensure the speed and stability of file retrieval.

[0079] High-priority files that are called periodically: If certain high-priority files are called periodically (such as being called every Monday), they can be automatically loaded into the cache before the estimated calling time to optimize the response speed.

[0080] Access control: Based on the importance and access frequency of files, access rights to high-priority files are more strictly managed to avoid unnecessary calls. Different levels of access rights are assigned to different user groups to ensure the security of core files.

[0081] Hot backup and disaster recovery management: Implement hot backup and real-time mirroring for extremely high priority files to prevent data loss due to hardware failure or system crash. These files can be synchronously backed up to redundant devices or off-site storage to ensure that the files can be restored at any time.

[0082] For inaccurate judgment results, that is, the accuracy value of the file priority judgment of the AI ​​algorithm generated within a fixed time period is less than the readiness standard threshold, the judgment result information is fed back to the AI ​​model, and the accuracy values ​​of the file priority judgment of the AI ​​algorithm generated within a subsequent fixed time period that are less than the readiness standard threshold are collected, and a data set is established, and the mean and standard deviation of the data set are calculated. After analyzing the abnormal degree of accuracy of the inaccurate judgment results within the fixed time period, the weights of the file access frequency and the recovery call frequency are dynamically adjusted according to the analysis results.

[0083] If the mean accuracy value is greater than or equal to the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is less than the reference threshold of the standard deviation of the accuracy value, the AI ​​model has high accuracy and stable accuracy in most cases, and the current weight setting is maintained, but the sensitivity of the file access frequency and the recovery call frequency can be slightly reduced to avoid over-response and maintain system stability;

[0084] If the mean accuracy value is greater than or equal to the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is greater than or equal to the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is high, but the accuracy fluctuates greatly, indicating that the accuracy is unstable in some scenarios. Increase the weight of file access frequency and reduce the weight of recovery call frequency to reduce the impact of frequent changes on the system and improve the stability of the model in different scenarios;

[0085] If the mean accuracy value is less than the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is greater than or equal to the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is low, and the accuracy fluctuates greatly, indicating that the model cannot accurately determine the priority in most cases. Significantly increase the weight of file access frequency, reduce the weight of recovery call frequency, prioritize file priority based on access frequency, ensure the core stability of the determination, and retrain the model later;

[0086] If the mean accuracy value is less than the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is less than the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is low but the fluctuation is small, indicating that although the model judgment result is inaccurate, the consistency is high. Balance the weights of increasing the file access frequency and the recovery call frequency, and add new feature variables or increase the amount of data in model training to improve the overall accuracy.

[0087] In this embodiment, first, the system sets the upper limit of storage space usage as the migration threshold, and when the threshold is reached, it enters the resource management mode. In the resource management mode, the system monitors the access frequency of low-priority files in high-load and low-load scenarios in real time, and calculates the access frequency anomaly index. After the file is migrated to the cloud, the system records each recovery call operation, and judges the degree of abnormality of the call frequency based on the fluctuation of the number of calls within a fixed observation period. Subsequently, the system evaluates the accuracy of the AI ​​algorithm's judgment of file priority based on the degree of abnormality of the low-priority file call frequency and the access frequency anomaly index, and divides the judgment results into "accurate" or "inaccurate". For accurate judgment results, the high and low priority management strategies are refined based on the access mode; for inaccurate judgment results, the information is fed back to the AI ​​model, and the weights of the file access frequency and the recovery call frequency are dynamically adjusted based on the degree of accuracy anomaly, and the intelligent judgment of file priority is gradually optimized.

[0088] Embodiment 2, a file storage management system based on computer system resource usage described in this embodiment includes a storage monitoring module, a load scenario management module, a call frequency monitoring module, a determination accuracy evaluation module and a priority management optimization module;

[0089] Storage monitoring module: It monitors local storage resources in real time based on the file storage management system. It sets the upper limit of storage space usage as the migration threshold for triggering data migration operations. When the file storage management storage usage exceeds the migration threshold, it enters the resource management mode.

[0090] Load scenario management module: monitors the actual access frequency of migrated files in high-load scenarios and low-load scenarios in the resource management mode in real time, records the actual access status of all low-priority files, counts the proportion of files whose access frequency exceeds the access frequency threshold, and calculates the comprehensive access frequency anomaly index in high-load scenarios and low-load scenarios;

[0091] Call frequency monitoring module: After the files are migrated to the cloud, the file storage management system records each recovery call operation and continuously monitors the call status. It determines the abnormality of the low-priority file call frequency based on the fluctuation of the number of low-priority file calls within a fixed observation period;

[0092] Judgment accuracy assessment module: evaluates the accuracy of the AI ​​algorithm's judgment of file priority based on the abnormal degree of low-priority file call frequency and the access comprehensive frequency abnormality index. Based on the evaluation results, the file priority judgment results are divided into accurate judgment results and inaccurate judgment results;

[0093] Priority management optimization module: Based on the access patterns of files corresponding to accurate judgment results, high-priority and low-priority management strategies are further refined. For inaccurate judgment results, the judgment result information is fed back to the AI ​​model. Based on the abnormal degree of accuracy of the inaccurate judgment results within a fixed time period, the weights of file access frequency and recovery call frequency are dynamically adjusted.

[0094] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0095] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0096] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0097] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A file storage management method based on computer system resource usage, characterized in that: The steps include: S1: Based on the file storage management system, the local storage resources are monitored in real time. The upper limit of the storage space usage is preset as the migration threshold for triggering the data migration operation. When the file storage management storage usage exceeds the migration threshold, the resource management mode is entered; S2: Monitor the actual access frequency of migrated files in the high-load scenario and the low-load scenario in the resource management mode in real time, record the actual access status of all low-priority files, count the proportion of files whose access frequency exceeds the access frequency threshold, and calculate the comprehensive access frequency anomaly index in the high-load scenario and the low-load scenario; S3: After the file is migrated to the cloud, the file storage management system records each recovery call operation and continuously monitors the call status. Based on the fluctuation of the number of low-priority file calls within a fixed observation period, the abnormal degree of the low-priority file call frequency is determined; S4: Based on the abnormal degree of low-priority file call frequency and the access comprehensive frequency abnormality index, the accuracy of the AI ​​algorithm in determining the file priority is evaluated. Based on the evaluation results, the file priority determination results are divided into accurate determination results and inaccurate determination results; S5: Based on the access patterns of the files corresponding to the accurate judgment results, further refine the high-priority and low-priority management strategies. For inaccurate judgment results, feed the judgment result information back to the AI ​​model, and dynamically adjust the weights of file access frequency and recovery call frequency based on the accuracy abnormality of the inaccurate judgment results within a fixed time period.

2. A file storage management method based on computer system resource usage according to claim 1, characterized in that: In S2, the access frequency anomaly rate is calculated and recorded in the high-load and low-load scenarios respectively, and the comprehensive access frequency anomaly index in the high-load and low-load scenarios is calculated after comprehensive analysis. The calculation method of the comprehensive access frequency anomaly index is: Collect and organize the access frequency anomaly rates in high-load scenarios and low-load scenarios within the W time period to form time series data, determine the smoothing coefficient α, and set the initial anomaly index S0 to the anomaly rate value at the first time point as the starting point of the smoothing calculation. At each time point t, calculate the comprehensive access frequency anomaly index MS, and the expression is: MS = α × access frequency anomaly rate t + (1-α) × St-1; MSt-1 represents the comprehensive access frequency anomaly index at time t-1.

3. A file storage management method based on computer system resource usage according to claim 2, characterized in that: In S3, the fluctuation of the number of low-priority file calls within a fixed observation period is analyzed to generate a file call frequency fluctuation index. The method for obtaining the file call frequency fluctuation index is: Collect the call frequency time series data of the file, set X = {x1, x2, ..., x n }, where x n represents the number of calls at time n, selects the wavelet function, and selects the decomposition layer J according to the frequency distribution and fluctuation characteristics of the data. The number of layers determines the scale of the wavelet decomposition. The call frequency time series X is decomposed by multi-layer wavelet to obtain detail components and approximate components at different scales. For the j-th layer of decomposition, the call frequency series X is decomposed into an approximate coefficient A J and several detail coefficients D j , the expression is: A J Indicates the approximate components of the signal, reflecting the low-frequency components, D j Represents the detail component of the jth layer, reflecting the high-frequency component; selects the detail coefficient of the medium and high-frequency layers to reflect the short-term fluctuation of the call frequency, and calculates the file call frequency fluctuation index, the expression is: Where KM is the file call frequency fluctuation index.

4. A file storage management method based on computer system resource usage according to claim 3, characterized in that: In S4, the accuracy of the AI ​​algorithm in determining the file priority is evaluated based on the abnormal degree of the low-priority file call frequency and the access comprehensive frequency abnormality index. Based on the evaluation results, the file priority determination results are divided into accurate determination results and inaccurate determination results, specifically: The comprehensive access frequency anomaly index and the file call frequency fluctuation index are converted into the first eigenvector, and the first eigenvector is used as the input of the machine learning model. The machine learning model uses each group of first eigenvectors to predict the accuracy value label of the AI ​​algorithm for file priority judgment as the prediction target, and takes minimizing the sum of the prediction errors of the accuracy value labels of all AI algorithms for file priority judgment as the training target. The machine learning model is trained until the sum of the prediction errors converges and the model training is stopped. The accuracy value of the AI ​​algorithm for file priority judgment is determined according to the model output results, wherein the machine learning model is a polynomial regression model.

5. A file storage management method based on computer system resource usage according to claim 4, characterized in that: In S4, the acquired accuracy value of the AI ​​algorithm for file priority determination is compared with the readiness standard threshold. If the accuracy value of the AI ​​algorithm for file priority determination is greater than or equal to the readiness standard threshold, it means that the AI ​​algorithm has high accuracy in file priority determination. In this case, a high-accuracy determination signal is generated, and the determination result of the file priority is classified as an accurate determination result. If the accuracy value of the AI ​​algorithm for file priority determination is less than the readiness standard threshold, it means that the accuracy of the AI ​​algorithm for file priority determination is low. In this case, a low-accuracy determination signal is generated, and the determination result of the file priority is classified as an inaccurate determination result.

6. A file storage management method based on computer system resource usage according to claim 1, characterized in that: In S5, for inaccurate judgment results, that is, the accuracy value of the file priority judgment of the AI ​​algorithm generated within a fixed time period is less than the readiness standard threshold, the judgment result information is fed back to the AI ​​model, and the accuracy values ​​of the file priority judgment of the AI ​​algorithm generated within a subsequent fixed time period that are less than the readiness standard threshold are collected, and a data set is established, and the mean and standard deviation of the data set are calculated. After analyzing the abnormal degree of accuracy of the inaccurate judgment results within the fixed time period, the weights of the file access frequency and the recovery call frequency are dynamically adjusted according to the analysis results.

7. A file storage management method based on computer system resource usage according to claim 6, characterized in that: If the mean accuracy value is greater than or equal to the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is less than the reference threshold of the standard deviation of the accuracy value, the AI ​​model has high accuracy and stable accuracy in most cases, and the current weight setting is maintained, the sensitivity of the file access frequency and the recovery call frequency is reduced, and the system stability is maintained; If the mean accuracy value is greater than or equal to the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is greater than or equal to the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is high, but the accuracy fluctuation is large. Increase the file access frequency weight and reduce the recovery call frequency weight; If the mean accuracy value is less than the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is greater than or equal to the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is low and the accuracy fluctuates greatly. Increase the weight of the file access frequency and reduce the weight of the recovery call frequency. If the mean accuracy value is less than the reference threshold of the mean accuracy value, and the standard deviation of the accuracy value is less than the reference threshold of the standard deviation of the accuracy value, the overall accuracy of the AI ​​model is low but the fluctuation is small, and the weights of increasing the file access frequency and the recovery call frequency are balanced.

8. A file storage management system based on computer system resource usage, used to implement a file storage management method based on computer system resource usage as claimed in any one of claims 1 to 7, characterized in that: It includes storage monitoring module, load scenario management module, call frequency monitoring module, judgment accuracy assessment module and priority management optimization module; Storage monitoring module: It monitors local storage resources in real time based on the file storage management system. It sets the upper limit of storage space usage as the migration threshold for triggering data migration operations. When the file storage management storage usage exceeds the migration threshold, it enters the resource management mode. Load scenario management module: monitors the actual access frequency of migrated files in high-load scenarios and low-load scenarios in the resource management mode in real time, records the actual access status of all low-priority files, counts the proportion of files whose access frequency exceeds the access frequency threshold, and calculates the comprehensive access frequency anomaly index in high-load scenarios and low-load scenarios; Call frequency monitoring module: After the files are migrated to the cloud, the file storage management system records each recovery call operation and continuously monitors the call status. It determines the abnormality of the low-priority file call frequency based on the fluctuation of the number of low-priority file calls within a fixed observation period; Judgment accuracy assessment module: evaluates the accuracy of the AI ​​algorithm's judgment of file priority based on the abnormal degree of low-priority file call frequency and the access comprehensive frequency abnormality index. Based on the evaluation results, the file priority judgment results are divided into accurate judgment results and inaccurate judgment results; Priority management optimization module: Based on the access patterns of files corresponding to accurate judgment results, high-priority and low-priority management strategies are further refined. For inaccurate judgment results, the judgment result information is fed back to the AI ​​model. Based on the abnormal degree of accuracy of the inaccurate judgment results within a fixed time period, the weights of file access frequency and recovery call frequency are dynamically adjusted.

Citation Information

Cited By

  • Data storage control method and device, storage medium and electronic equipment

    CN120704617A

  • Data optimization storage method and system based on artificial intelligence

    CN120723735A

  • Intelligent contract-driven automatic file version updating system

    CN121050742A

  • Smart contract driven automatic file version updating system

    CN121050742B

  • Photomagnetic storage method and system for archiving based on virtual file system

    CN121349370A