File Storage Real-time Monitoring Method and System
By establishing a multi-dimensional monitoring index system and early warning prediction model, the problems of false alarms, missed reports and inefficiency of traditional storage monitoring methods in large-scale distributed environments are solved, and comprehensive monitoring and efficient operation and maintenance of the storage system are achieved.
Patent Information
- Application Number
- CN202510578049.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Traditional storage monitoring methods cannot adapt to dynamically changing business loads in large-scale distributed storage environments, resulting in frequent false alarms and missed reports. Single-dimensional monitoring cannot reflect the overall health of the system, manual processing efficiency is low, lacks predictability, and cannot handle abnormalities in a timely manner.
Establish a multi-dimensional monitoring index system, collect storage cluster information through distributed monitoring agents, generate distributed storage warning feature data sets, combine early warning prediction models to predict risks, realize automatic alarm rating and disposal, and use adaptive feature extraction and two-stage submission to ensure consistent execution of instructions.
It realizes comprehensive monitoring of the storage system, improves the accuracy and timeliness of early warnings, reduces false alarms and missed reports, improves operation and maintenance efficiency and system reliability, and has adaptive learning capabilities.
Smart Images

Figure CN120086098B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data monitoring and alarm, and in particular to a method and system for real-time monitoring of file storage. Background Art
[0002] In current file storage systems, the scale of storage clusters continues to expand with the rapid growth of data volumes and increasing business demands. Traditional storage monitoring methods rely primarily on simple resource usage threshold monitoring, such as setting fixed thresholds for CPU usage, memory utilization, and storage space. When monitoring indicators exceed preset thresholds, alarms are triggered, and administrators manually address them based on the alarm information. This approach can meet basic monitoring needs in small-scale storage systems and enables timely detection and resolution of storage system anomalies.
[0003] However, in large-scale distributed storage environments, traditional monitoring methods have significant shortcomings. First, fixed threshold monitoring cannot adapt to dynamically changing business loads, easily generating a large number of false positives or missed reports. Second, single-dimensional resource monitoring fails to reflect the overall health of the storage system and ignores the correlations between resource indicators. Third, manual alarm processing is inefficient and difficult to respond to frequent alarm events, resulting in system anomalies not being addressed promptly. Finally, the lack of predictive monitoring means that only reactive responses can be made after problems occur, failing to detect and prevent potential risks in advance. Summary of the Invention
[0004] The present application provides a real-time monitoring method and system for file storage, which is used to implement intelligent monitoring and early warning in a distributed file storage system. By establishing a multi-dimensional monitoring indicator system and combining it with an early warning prediction model, system anomalies are predicted, and automatic classification processing of alarms and automatic handling of anomalies are achieved, thereby improving the operation and maintenance efficiency and reliability of the storage system.
[0005] In a first aspect, the present application provides a method for real-time monitoring of file storage. The method for real-time monitoring of file storage includes: collecting operation status information, storage security status information, and system alarm trigger information of a file storage cluster through a distributed monitoring agent, allocating monitoring tasks to each monitoring agent node, and forming an original monitoring alarm data stream after establishing a distributed monitoring network; aligning the original monitoring alarm data stream according to a monitoring period, and adaptively extracting storage alarm features, storage security features, and system status features to generate a distributed storage early warning feature dataset; combining an early warning prediction model based on the distributed storage early warning feature dataset to predict storage security risks, system alarm risks, and data access risks, and forming a storage system early warning result; dividing storage security alarm levels, system alarm levels, and access alarm levels according to the storage system early warning result, selecting a distributed early warning signal transmission channel according to different levels, and forming a storage system hierarchical alarm signal; generating a storage security handling instruction, a system maintenance instruction, and an access control instruction according to the storage system hierarchical alarm signal, coordinating the resources of each storage node to issue distributed instructions, and establishing a storage system handling instruction; evaluating the execution effect of the storage system handling instruction, and adjusting monitoring parameters and alarm strategies according to the evaluation result to form a distributed storage monitoring optimization scheme.
[0006] In a second aspect, the present application provides a system for real-time monitoring of file storage. The system for real-time monitoring of file storage includes:
[0007] An allocation module, configured to collect operation status information, storage security status information, and system alarm trigger information of a file storage cluster through a distributed monitoring agent, allocate monitoring tasks to each monitoring agent node, and form an original monitoring alarm data stream after establishing a distributed monitoring network;
[0008] An alignment module, configured to align the original monitoring alarm data stream according to a monitoring period, and adaptively extract storage alarm features, storage security features, and system status features to generate a distributed storage early warning feature dataset;
[0009] A prediction module, configured to combine an early warning prediction model based on the distributed storage early warning feature dataset to predict storage security risks, system alarm risks, and data access risks, and form a storage system early warning result;
[0010] A division module, configured to divide storage security alarm levels, system alarm levels, and access alarm levels according to the storage system early warning result, select a distributed early warning signal transmission channel according to different levels, and form a storage system hierarchical alarm signal;
[0011] The issuing module is used to generate storage security disposal instructions, system maintenance instructions, and access control instructions according to the storage system graded alarm signals, coordinate the resources of each storage node to perform distributed instruction issuance, and establish storage system disposal instructions;
[0012] The monitoring module is used to evaluate the execution effect of the storage system processing instructions, adjust the monitoring parameters and alarm strategy according to the evaluation results, and form a distributed storage monitoring optimization plan.
[0013] The technical solution provided by this application achieves comprehensive monitoring of the storage system's operational status by dynamically collecting multi-dimensional status information from a storage cluster through distributed monitoring agents, avoiding information omissions caused by single-metric monitoring. Data alignment and adaptive feature extraction methods are used to process raw monitoring data, ensuring not only temporal consistency across multiple sources but also automatically adjusting feature extraction strategies based on data variation, improving the representativeness and accuracy of the feature data. System risk prediction and analysis are performed based on an early warning prediction model. The collaborative work of multiple prediction units—including an early warning perception unit for real-time analysis of current status, an early warning memory unit for storing historical patterns, and an early warning reasoning unit for risk inference—significantly enhances the accuracy and timeliness of early warnings. An adaptive alarm classification mechanism dynamically assigns alarm levels based on early warning results and selects differentiated signal transmission channels, ensuring timely delivery of high-risk alarms while avoiding wasteful signal transmission resources. In the action instruction generation phase, action strategies are intelligently generated by correlating and matching historical action plans with the current system status. A two-phase submission process ensures consistent execution of distributed instructions. Finally, through execution performance evaluation and parameter optimization mechanisms, monitoring and early warning strategies are continuously improved, giving the entire monitoring system adaptive learning capabilities. This solution fully utilizes the advantages of the early warning prediction model in time series data analysis and risk prediction. Through the organic combination of algorithms such as feature extraction, pattern recognition and risk reasoning, it realizes the transition from passive monitoring to active early warning. At the same time, the adaptive optimization mechanism of the early warning model ensures the stability and reliability of the algorithm in different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0015] Figure 1 This is a schematic diagram of an embodiment of a method for real-time monitoring of file storage in an embodiment of the present application;
[0016] Figure 2This is a flow chart showing how to allocate monitoring tasks to monitoring agent nodes and establish a distributed monitoring network to generate an original monitoring alarm data stream in an embodiment of the present application.
[0017] Figure 3 This is a schematic diagram of parameter types of the alarm strategy optimization list in an embodiment of the present application;
[0018] Figure 4 This is a schematic diagram of an embodiment of a real-time monitoring system for file storage in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The embodiments of the present application provide a method and system for real-time monitoring of file storage. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0020] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the present application, a method for real-time monitoring of file storage includes:
[0021] Step S101: Collect the running status information, storage security status information, and system alarm trigger information of the file storage cluster through the distributed monitoring agent, assign the monitoring task to each monitoring agent node, and form the original monitoring alarm data stream after establishing the distributed monitoring network;
[0022] Step S102: aligning the original monitoring alarm data stream according to the monitoring period, extracting storage alarm features, storage security features, and system status features in an adaptive manner, and generating a distributed storage early warning feature dataset;
[0023] Step S103: Based on the distributed storage warning feature data set and the combined warning prediction model, storage security risk, system alarm risk, and data access risk are predicted to form a storage system warning result;
[0024] Step S104: Classify the storage security alarm levels, system alarm levels, and access alarm levels according to the storage system warning results, select the distributed warning signal transmission channels according to different levels, and form the storage system hierarchical alarm signals.
[0025] Step S105: Generate storage security handling instructions, system maintenance instructions, and access control instructions according to the storage system hierarchical alarm signals, coordinate the resources of each storage node to perform distributed instruction distribution, and establish storage system handling instructions.
[0026] Step S106: Evaluate the execution effect of the storage system handling instructions, adjust the monitoring parameters and alarm strategies according to the evaluation results, and form a distributed storage monitoring optimization plan.
[0027] It can be understood that the execution subject of this application can be a file storage real-time monitoring system, or a terminal or a server. Specifically, it is not limited here. This application embodiment is described by taking the server as the execution subject as an example.
[0028] Specifically, collect operation status information, including basic performance indicators such as CPU usage rate, memory occupancy rate, and disk I / O performance; storage security status information covers security-related parameters such as data access permissions, data integrity check values, and file system permission settings; system alarm trigger information records alarm data such as the occurrence time, abnormal type, and influence range of abnormal events. After the monitoring agent nodes perform preliminary processing on this information, the distributed monitoring network aggregates the data collected by each node to form the original monitoring alarm data stream.
[0029] When processing the original monitoring alarm data stream, align the data according to the monitoring cycle to eliminate the difference in data collection time of different nodes. Align the data through the sliding time window mechanism, and the window size is dynamically adjusted according to the monitoring cycle. The aligned data extracts storage alarm features (such as I / O delay abnormality, storage space alarm, etc.), storage security features (such as unauthorized access, data integrity damage, etc.), and system status features (such as too high load, service response timeout, etc.) respectively through the adaptive feature extraction method. The data after feature extraction constitutes the distributed storage warning feature data set.
[0030] The warning prediction model performs risk prediction based on the distributed storage warning feature data set. The warning prediction model works in coordination with the warning perception unit, warning memory unit, and warning reasoning unit: the warning perception unit processes real-time data and identifies the current system status; the warning memory unit stores historical warning data and establishes a warning pattern library; the warning reasoning unit combines the real-time status and historical patterns to perform risk reasoning. The warning prediction results include prediction indicators such as storage security risk (data security threat level), system alarm risk (system abnormality possibility), and data access risk (access abnormality probability).
[0031] When dividing the alarm levels according to the warning results of the storage system, calculate the storage security risk score, system warning risk score, and access risk score respectively to form risk score data. Divide them into a high-risk range, a medium-risk range, and a low-risk range according to the degree of risk, and correspondingly generate the storage security alarm level, system warning level, and access alarm level. Adopt different signal transmission channels for alarms of different levels: ensure reliable signal delivery through redundant transmission in the high-risk range, use message queue transmission in the medium-risk range, and adopt ordinary data channels for transmission in the low-risk range. In the link of generating disposal instructions, parse the hierarchical alarm signals of the storage system into three types of alarm data: storage security, system maintenance, and access control, and establish an alarm disposal response pool. Generate corresponding disposal instructions by associating and matching historical disposal solutions. Mark the priority of the disposal instructions, construct an instruction dependency table, and determine the instruction execution order. Statistically analyze the resource status of storage nodes, calculate the load distribution, and formulate a node resource allocation plan. Adopt a two-phase commit method to ensure the consistency of instruction issuance and monitor the instruction execution status in real time.
[0032] Evaluate the execution effect of the disposal instructions, statistically analyze indicators such as instruction response time, resource utilization rate, and problem resolution rate, and generate an execution effect report. Quantitatively evaluate the disposal instructions and calculate the disposal instruction score. For instruction types with scores lower than the threshold, extract the corresponding monitoring parameters for optimization. Compare the monitoring parameter optimization table with historical monitoring data, calculate the parameter deviation value, and form a monitoring parameter adjustment plan. Modify parameters such as alarm trigger thresholds, sampling frequencies, and data alignment periods according to the adjustment plan to form a distributed storage monitoring optimization plan.
[0033] For example, in the monitoring scenario of a file storage system, collect abnormal access records in the file access log, including data such as the number of access failures, unauthorized access attempts, and file integrity verification failures. Perform real-time analysis on this data through the warning perception unit to calculate the current security risk indicators. The warning memory unit records historical security event data and establishes a security risk pattern library. The warning reasoning unit combines the current risk indicators and historical patterns to predict possible future security risks. When it is detected that the frequency of unauthorized access attempts on a certain storage node suddenly increases and the file integrity verification failure rate rises, the system will raise the security risk level of this node and trigger the corresponding disposal mechanism. Select an appropriate signal transmission channel according to the risk level to ensure that the alarm information is delivered to the management node in a timely manner. After receiving the alarm, the management node calls the security reinforcement strategy in the historical disposal solution library to generate targeted security disposal instructions. After the execution of the disposal instructions is completed, the system evaluates the disposal effect and optimizes relevant monitoring parameters and alarm strategies according to the evaluation results.
[0034] In this embodiment, distributed monitoring agents dynamically collect multi-dimensional status information from storage clusters, enabling comprehensive monitoring of the storage system's operational status and avoiding information omissions caused by single-metric monitoring. Data alignment and adaptive feature extraction methods are used to process raw monitoring data, ensuring not only temporal consistency across multiple sources but also automatically adjusting feature extraction strategies based on data variation, improving the representativeness and accuracy of feature data. System risk prediction and analysis are conducted based on an early warning prediction model. The collaborative work of multiple prediction units, including an early warning perception unit that analyzes current status in real time, an early warning memory unit that stores historical patterns, and an early warning reasoning unit that infers risk, significantly enhances the accuracy and timeliness of early warnings. An adaptive alarm classification mechanism dynamically assigns alarm levels based on early warning results and selects differentiated signal transmission channels, ensuring timely delivery of high-risk alarms while avoiding wasteful signal transmission resources. In the generation of action instructions, action strategies are intelligently generated by correlating and matching historical action plans with the current system status. A two-phase submission approach ensures consistent execution of distributed instructions. Finally, through execution performance evaluation and parameter optimization mechanisms, monitoring and early warning strategies are continuously improved, giving the entire monitoring system adaptive learning capabilities. This solution fully utilizes the advantages of the early warning prediction model in time series data analysis and risk prediction. Through the organic combination of algorithms such as feature extraction, pattern recognition and risk reasoning, it realizes the transition from passive monitoring to active early warning. At the same time, the adaptive optimization mechanism of the early warning model ensures the stability and reliability of the algorithm in different application scenarios.
[0035] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0036] (1) Obtain the read and write operation logs, storage capacity usage data, and I / O performance data of the file storage cluster through the distributed monitoring agent, and partition the file storage cluster according to the node load;
[0037] (2) Allocate read and write operation logs, storage capacity usage data, and I / O performance data to different monitoring agent nodes, and use a local cache mechanism to store the collected data;
[0038] (3) Based on the collected data in the local cache mechanism, the data transmission tasks between monitoring agent nodes are dynamically scheduled and allocated;
[0039] (4) Assign tasks to monitoring agent nodes based on dynamic scheduling allocation results, and the monitoring agent nodes collect operating status information, store security status information, and system alarm trigger information;
[0040] (5) Aggregate the operation status information, storage safety status information, and system alarm trigger information to the monitoring network center node to form the original monitoring alarm data stream.
[0041] Specifically, as Figure 2 shown, this is a flowchart of the original monitoring alert data stream formed after the monitoring tasks are assigned to each monitoring agent node and a distributed monitoring network is established in the embodiment of the present application. This flowchart shows the complete process from the collection of original data to the formation of the monitoring alert data stream: First, the distributed monitoring agents collect basic monitoring data, including read and write operation logs, storage capacity data, and I / O performance data. Subsequently, the collected data is distributed to different monitoring agent nodes and stored through a local caching mechanism. The local cache is divided into two levels: memory cache and disk cache. After dynamic task scheduling based on the cached data, the monitoring agent nodes collect operation status information, storage security information, and alert trigger information in parallel. Finally, the collected monitoring data is uniformly summarized to the central node to form the original monitoring alert data stream. Through a multi-level parallel processing and dynamic scheduling mechanism, the entire process realizes the real-time monitoring of the file storage system.
[0042] The distributed monitoring agents obtain key operation data by connecting to each node of the file storage cluster. Read the read and write operation logs of the file storage cluster, record operation information such as file creation, modification, and deletion, and at the same time collect storage capacity usage data to monitor the space occupancy of each storage node, and collect I / O performance data, including indicators such as read and write speeds, response times, and queue lengths. Based on the collected data, according to the load indicators such as CPU utilization rate, memory occupancy rate, and disk load of each node, the file storage cluster is partitioned, and nodes with similar loads are partitioned into the same partition for subsequent task allocation and resource scheduling. The collected data needs to be reasonably distributed to different monitoring agent nodes for processing. The read and write operation logs are sliced according to time sequence, and the logs in consecutive time periods are distributed to adjacent monitoring agent nodes; the storage capacity usage data is grouped according to the physical location of the storage nodes and distributed to the nearest monitoring agent nodes; the I / O performance data is divided according to the data type, and performance indicators of the same type are distributed to dedicated monitoring agent nodes. Each monitoring agent node is equipped with a local caching mechanism, which adopts a two-level structure of memory cache and disk cache to temporarily store the collected data. The memory cache is used to store recent high-frequency access data, and the disk cache is used to store historical data. The two-level cache maintains data synchronization through prefetching and eviction policies.
[0043] Data transmission between monitoring agent nodes requires dynamic scheduling and allocation. The load status of each node is calculated based on factors such as data access frequency, data volume, and node processing capacity, as recorded in the local cache mechanism. When a node's load exceeds a threshold, part of its data transmission tasks is reallocated to a node with a lower load. For cross-node data access requests, data is preferentially retrieved from the local cache. If the local cache misses, it is retrieved from the caches of other nodes. During dynamic scheduling, resource utilization at each node is continuously monitored, and task allocation strategies are adjusted promptly. Based on the results of dynamic scheduling and allocation, monitoring agent nodes collect different types of monitoring data. Operational status information includes basic performance indicators such as the node's CPU utilization, memory usage, and network throughput, which are obtained through regular sampling. Storage security status information includes security-related data such as file access permissions, data integrity check values, and abnormal access records, which are collected through the security audit module. System alarm trigger information includes alarm events such as hardware failures, service anomalies, and performance bottlenecks, which are obtained through the alarm detection module.
[0044] The data collected by each monitoring agent node is aggregated to the central node of the monitoring network. The central node receives the data streams from each monitoring agent and aligns the data in time series to eliminate time differences between different nodes. It integrates operating status information, storage security status information, and system alarm trigger information according to a predefined data format to form the original monitoring alarm data stream.
[0045] For example, when a node in a file storage cluster begins to experience anomalies, an increase in file access latency is recorded in the node's read and write operation logs, and I / O performance data shows a sudden increase in the length of the read and write queues. The monitoring agent responsible for monitoring the node stores this abnormal data in the local cache and triggers dynamic scheduling of data transmission tasks. Because the node is already highly loaded, some monitoring tasks are reallocated to other monitoring agent nodes. The newly assigned monitoring agent node begins collecting operational status information for the node and finds that CPU usage continues to rise and memory usage is approaching the upper limit. At the same time, security status information indicates abnormal file access patterns, and the system alarm module also generates a performance bottleneck alarm. This information is transmitted to the central node in real time, and after time alignment, it generates an original monitoring alarm data stream containing multi-dimensional information such as performance anomalies, security risks, and system alarms.
[0046] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0047] (1) Divide the original monitoring alarm data stream into time series, mark the data within the monitoring period with timestamps, and form time series monitoring data;
[0048] (2) Divide the time series monitoring data into operating status sequence, storage safety sequence, and alarm trigger sequence according to the data type, and perform time alignment calibration on the data series;
[0049] (3) Extract storage performance indicators, resource utilization indicators, and system load indicators based on the operating status sequence to generate storage alarm features;
[0050] (4) Extract data access patterns, security authentication records, and operation behavior characteristics based on storage security sequences to generate storage security features;
[0051] (5) Extract alarm frequency characteristics, alarm duration characteristics, and alarm correlation characteristics based on the alarm trigger sequence to generate system status characteristics;
[0052] (6) Integrate and associate storage alarm features, storage security features, and system status features, combine features according to the monitoring cycle, and generate a distributed storage warning feature data set.
[0053] Specifically, the raw monitoring alarm data stream contains continuous data collected from each monitoring node. This data stream is segmented into multiple time segments according to preset monitoring periods (such as 5-minute or 10-minute intervals). A unified timestamp is added to the data within each time segment to record the specific time point at which the data was generated. Timestamps are assigned with millisecond-level accuracy to ensure data timeliness. Through time series segmentation and timestamp tagging, the disordered raw data is converted into ordered time series monitoring data.
[0054] Time series monitoring data is further classified and processed according to data type, dividing it into three types of data series. The operating status series records the runtime status data of the file storage system, including performance indicators and resource usage; the storage security series contains monitoring records related to data security, such as access control and data integrity; and the alarm trigger series stores records of various alarm events. These three types of data series are time-aligned and calibrated to eliminate time deviations between different data sources and establish a unified time base. Three key indicators are extracted from the operating status series: storage performance indicators reflect the read and write performance of the storage system, including IOPS (input and output operations per second), throughput, and response time; resource utilization indicators record system resource usage, including CPU utilization, memory utilization, and storage space utilization; and system load indicators represent the overall system load, including process queue length and average system load. These indicators are cleaned and normalized to form standardized storage alarm features.
[0055] Processing storage security sequences involves extracting features from three dimensions: data access patterns analyze the temporal patterns, spatial distribution, and frequency of file access; security authentication records include security audit information such as user authentication, permission verification, and access control; and operational behavior features describe the type, frequency, and impact of file operations. By extracting and analyzing these features, storage security features are generated that reflect the security status of the storage system. Alarm trigger sequences require the extraction of three types of features: alarm frequency features count the number of alarms and alarm density per unit time; alarm persistence features record the duration and recovery time of alarm events; and alarm correlation features analyze the temporal and causal relationships between different alarms. Through correlation analysis and pattern recognition, these features generate system status features that reflect the overall state of the system.
[0056] Integrate and correlate storage alarm features, storage security features, and system status features. Use feature vectors to represent each feature, and establish a correlation matrix between them. Combine features across time dimensions according to the monitoring cycle to construct a multidimensional feature space. Through feature combination and correlation analysis, generate a distributed storage warning feature dataset containing complete monitoring information.
[0057] In the data backup scenario of a file storage system, when executing the data backup task, the original monitoring alarm data stream records various indicators during the backup process. Through time series division, the monitoring data is segmented into 10-minute cycles and precise timestamps are added. After data classification, it is observed in the operating status sequence that the backup task causes the I / O performance of the storage system to degrade, manifested as reduced IOPS and increased response time; the storage security sequence shows that the file access pattern during the data backup process has changed, with intensive read operations; and the alarm trigger sequence records an increase in the frequency of performance alarms. Features are extracted from these sequences: the storage alarm feature shows that the I / O queue length continues to increase; the storage security feature indicates that data access is concentrated in the backup source directory; and the system status feature reflects that the alarm events have obvious temporal correlation. These features are integrated into the early warning feature dataset, which fully records the impact of the backup task on the storage system.
[0058] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0059] (1) Establish time series warning training samples based on the distributed storage warning feature dataset, and construct storage alarm features, storage security features, and system status features as feature input sequences;
[0060] (2) The feature input sequence is subjected to data association analysis through the early warning perception unit to generate early warning perception features; the historical early warning data is memorized and stored through the early warning memory unit to form early warning memory features; the feature inference is performed through the early warning reasoning unit to generate early warning reasoning features;
[0061] (3) Input the warning perception features, warning memory features, and warning reasoning features into the warning prediction model, calculate the weights of the feature data, and generate a warning parameter matrix;
[0062] (4) Based on the warning parameter matrix, the newly input distributed storage warning feature data set is forward calculated to obtain the storage security risk value, system alarm risk value, and data access risk value;
[0063] (5) Optimize the parameters of the warning parameter matrix through back propagation and adjust the weight parameters of each unit in the warning prediction model until the prediction accuracy meets the requirements;
[0064] (6) Use the optimized early warning prediction model to predict and analyze storage security risks, system alarm risks, and data access risks to form storage system early warning results.
[0065] Specifically, the early warning prediction model is a deep neural network structure consisting of an early warning perception unit, an early warning memory unit, and an early warning reasoning unit. The early warning perception unit uses a convolutional neural network to process feature input sequences. The early warning memory unit stores historical warning patterns based on a long-short-term memory network. The early warning reasoning unit performs feature fusion and risk reasoning through a fully connected neural network. Training the early warning prediction model first requires the construction of time-series early warning training samples. Historical data is selected from the distributed storage early warning feature dataset as training data, and the stored alarm features, stored safety features, and system status features are organized into a feature input sequence in chronological order. The feature input sequence undergoes data normalization to eliminate dimensional differences between different features and convert them into a unified numerical range.
[0066] The warning perception unit receives a feature input sequence and extracts the spatiotemporal correlations between these features through multi-layer convolution operations. The size and number of convolution kernels are determined by the time span and dimensionality of the features, while the step size of the sliding window determines the granularity of feature extraction. After processing in the convolutional and pooling layers, warning perception features are generated that reflect the correlation patterns between the various features.
[0067] The warning memory unit is responsible for long-term storage of historical warning data. It employs an LSTM (Long Short-Term Memory) network structure, consisting of three control units: an input gate, a forget gate, and an output gate. The input gate controls the writing of new information, the forget gate clears irrelevant historical information, and the output gate manages the reading of memory information. This gating mechanism forms a warning memory feature that encompasses historical warning patterns.
[0068] The early warning inference unit performs feature inference analysis based on a fully connected neural network. It receives early warning perception features and early warning memory features as input and transforms and combines them using a multi-layer perceptron. Each layer of neurons is equipped with an activation function, introducing nonlinear transformation capabilities. This generates early warning inference features that reflect the system's risk status.
[0069] The early warning prediction model uses three types of features (warning perception features, warning memory features, and warning reasoning features) as input and performs feature fusion and weight calculation through a multi-layer neural network. The connection weights between neurons in each layer form the early warning parameter matrix, which records the importance and influence of each feature. The early warning parameter matrix is the core parameter of the model for risk prediction.
[0070] During the actual prediction phase, the newly collected distributed storage warning feature dataset is fed into the trained warning prediction model. Forward calculations are performed using the warning parameter matrix to generate three types of risk prediction values: a storage security risk value reflects the level of security threats to the system, a system alarm risk value indicates the likelihood of system anomalies, and a data access risk value indicates the probability of data access anomalies.
[0071] The model optimization process utilizes a backpropagation algorithm. The error between the predicted and observed values is propagated back through the network structure, calculating the gradient of each neuron layer. Based on this gradient information, the weights in the warning parameter matrix are adjusted to gradually reduce the prediction error. Model optimization is complete when the prediction accuracy reaches a preset threshold. The optimized warning prediction model provides real-time predictions for three types of storage system risks. The model first senses the current system state, then conducts inference analysis based on historical warning experience, and outputs a warning result reflecting the overall risk status of the system. This warning result includes information such as risk level, risk trend, and risk impact scope.
[0072] For example, when user access patterns change abnormally, the early warning perception unit analyzes access logs to detect features such as a sudden increase in access frequency and abnormal access permissions. The early warning memory unit extracts similar abnormal access patterns from historical data. The early warning reasoning unit comprehensively analyzes the current state and historical patterns to predict potential security risks. The early warning prediction model calculates a specific risk value based on the weights of various features. For example, an increase in the data access risk value indicates the possibility of unauthorized access.
[0073] It should be noted that the early warning prediction model consists of three functional units: early warning perception unit, early warning memory unit and early warning reasoning unit.
[0074] The early warning perception unit receives and stores alarm features, storage security features, and system status features as inputs. This unit transforms and combines the input features through a multi-layer perceptron structure. Each layer of the perceptron contains multiple processing nodes, and information is transmitted between the nodes through weighted connections. The early warning perception unit performs a sliding window analysis on the input features to extract the correlation and temporal patterns between the features. The processed data forms early warning perception features, reflecting the current operating state of the system. The early warning memory unit is constructed based on the memory mechanism of the neural network and is specifically used to store and manage historical early warning data. This unit includes three sub-modules: input control, memory storage, and output control. The input control module is responsible for screening and encoding new early warning data; the memory storage module maintains an early warning pattern library and stores the feature vectors of historical early warning events; the output control module manages the extraction and update of memory information. The early warning memory unit transforms historical early warning patterns into early warning memory features, providing empirical support for risk prediction.
[0075] The early warning reasoning unit is the decision-making center of the model and adopts a deep neural network architecture. The input layer of this unit receives the feature data from the perception unit and the memory unit and performs feature transformation through multiple hidden layers. Each hidden layer consists of multiple neurons, and the connection weights between the neurons are optimized through the training process. The output layer of the reasoning unit generates early warning reasoning features, comprehensively expressing the risk state of the system.
[0076] These three functional units are connected through feature transfer and weight matrices. The output features of the early warning perception unit are first transmitted to the early warning memory unit for pattern matching, and then, together with the historical features extracted by the memory unit, are input into the early warning reasoning unit. The early warning reasoning unit assigns weights according to the importance of the features and obtains the early warning result through weighted combination. The parameter optimization of the entire early warning prediction model uses the backpropagation method, calculates the prediction error, and adjusts the weights layer by layer. The training data of the model comes from the historical monitoring records of the file storage system, including known abnormal events and corresponding feature data. Through continuous iterative optimization, the model gradually masters the law of system risk evolution and improves the accuracy of early warning. For example, when the file storage system experiences frequent read / write delays, the early warning perception unit first captures the features of degraded I / O performance, such as an increase in the read / write queue length and an extension of the response time. After being processed by the perception unit, these feature data are matched with the historical performance alarm patterns stored in the early warning memory unit. The early warning reasoning unit comprehensively analyzes the current performance features and historical alarm patterns and predicts the development trend of the system alarm risk. The early warning prediction model calculates the specific risk value according to the weights of each feature, guiding the system to take performance optimization measures in a timely manner.
[0077] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0078] (1) Risk scoring is performed on the storage system warning results, and the storage security risk score, system alarm risk score, and access risk score are calculated to generate risk score data;
[0079] (2) Divide the risk score data into high-risk interval, medium-risk interval, and low-risk interval according to the preset risk threshold, and generate storage security warning level, system warning level, and access warning level accordingly;
[0080] (3) Establish a security warning channel for the storage security warning level, a maintenance warning channel for the system warning level, and a control warning channel for the access warning level, forming a distributed warning signal transmission channel;
[0081] (4) The alarm signals of the high-risk area are sent through multi-path transmission, the alarm signals of the medium-risk area are sent through message queue, and the alarm signals of the low-risk area are sent through ordinary data channels;
[0082] (5) Monitor the network status of the distributed warning signal transmission channel, evaluate the signal transmission quality, and select the optimal transmission path;
[0083] (6) The alarm signals of storage security alarm level, system alarm level and access alarm level are sent down through the corresponding distributed early warning signal transmission channel to form a storage system hierarchical alarm signal.
[0084] Specifically, in the real-time file storage monitoring method, a comprehensive scoring model is used to calculate the risk score of the storage system warning results:
[0085]
[0086] This formula calculates the overall risk score, where the parameters have the following meanings: Indicates the overall risk score of the storage system, with a value range of [0,1]; It represents the weight coefficient of the i-th warning factor, reflecting the importance of the factor; It represents the frequency coefficient of the i-th warning factor, describing the frequency of the alarm; It represents the time impact coefficient of the i-th warning factor, reflecting the duration of the alarm; It represents the danger level coefficient of the i-th warning factor, reflecting the severity of the alarm; n represents the total number of warning factors.
[0087] When a warning event is detected, the specific parameter values for each warning factor are determined. For example, in the storage security risk score, warning factors include unauthorized access, data integrity breaches, and security policy violations. For each factor, the frequency coefficient (e.g., the number of alarms per unit time), time impact coefficient (e.g., the duration of the alarm), and criticality coefficient (e.g., the range of affected data) are calculated. Then, based on the security risk assessment criteria, a weighting coefficient is determined to produce a security risk score.
[0088] Network status monitoring and transmission path selection use a multi-dimensional evaluation model:
[0089]
[0090] This formula is used to select the optimal transmission path. The meanings of the parameters are as follows: Indicates the score value of the optimal transmission path, used for path selection; Indicates the weight coefficient of bandwidth evaluation, reflecting the importance of bandwidth factors; represents the bandwidth utilization of the jth path, describing the transmission capacity of the path; Indicates the weight coefficient of delay evaluation, reflecting the impact of delay factors; represents the transmission delay of the jth path and records the delay of data transmission; Indicates the weight coefficient of quality assessment, reflecting the importance of transmission quality; represents the transmission quality indicator of the jth path, which can be the packet loss rate or error rate; m represents the total number of optional transmission paths.
[0091] Regularly collect performance indicators for each transmission path. Obtain bandwidth utilization, transmission delay, and quality indicators through network detection, and set corresponding weight coefficients based on the characteristics of the warning signal (such as real-time requirements, reliability requirements, etc.). The evaluation formula comprehensively considers these factors and selects the path with the highest score as the transmission channel for the warning signal. Based on the risk score data, the warning signals are graded. Set the high-risk interval threshold to 0.8, the medium-risk interval threshold to 0.5, and those below 0.5 are classified as low-risk intervals. Establish corresponding warning signal transmission channels for different risk intervals. Storage security alarm levels are transmitted through the security warning channel, using encrypted transmission to ensure data security; system alarm levels are transmitted through the maintenance warning channel, focusing on ensuring the real-time nature of the signal; access alarm levels are transmitted through the control warning channel, focusing on signal reliability.
[0092] Alarm signals for high-risk areas are transmitted using multiple paths, with signals sent simultaneously through primary and backup paths to ensure that the signals are successfully delivered through at least one path. Alarm signals for medium-risk areas are transmitted using message queues, with the message queue's persistence mechanism ensuring that the signals are not lost. Alarm signals for low-risk areas are transmitted through ordinary data channels, using the standard TCP / IP protocol stack for data transmission. Distributed early warning signal transmission channels are monitored in real time, collecting metrics such as bandwidth usage, transmission delay, and packet loss rate for each channel. The status of each transmission path is regularly tested using network probe packets, and the path availability and performance metrics are recorded. Based on the path status assessment results, the optimal transmission path is selected for each early warning signal.
[0093] For example, if a file system anomaly is detected on a storage node, the risk score for the anomaly is first calculated. Assuming an increase in the frequency of unauthorized access (F = 0.8), a prolonged anomaly duration (T = 0.9), and a large impact range (D = 0.7), and considering the security factor weight (w = 0.9), a higher risk score is calculated. Based on this score, the alarm is classified as high risk and transmitted through a security warning channel. During transmission, the network status is continuously monitored. If a sudden increase in latency on the primary transmission path is detected, a switch to a backup path is immediately initiated to ensure timely delivery of the alarm signal.
[0094] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0095] (1) Analyze the storage system's hierarchical alarm signals into three types of alarm data: storage security, system maintenance, and access control, and establish an alarm handling response pool;
[0096] (2) Correlate and match the data in the alarm response pool with historical disposal plans to generate storage security disposal instructions, system maintenance instructions, and access control instructions respectively;
[0097] (3) Prioritize storage security disposal instructions, system maintenance instructions, and access control instructions, build an instruction dependency table, and form an instruction execution sequence;
[0098] (4) Based on the instruction execution sequence, the resource occupancy status of each storage node is counted, the node load distribution is calculated, and a node resource allocation plan is generated;
[0099] (5) Combine the node resource allocation plan with the instruction execution sequence and issue task instructions to each storage node through a two-phase submission method;
[0100] (6) Monitor the execution process of task instructions in real time, record the instruction response status, and establish storage system disposal instructions.
[0101] Specifically, the alarm parser parses storage system hierarchical alarm signals, extracting basic information such as the alarm type, content, and severity. Based on the alarm type, the alarm data is categorized into three categories: storage security (e.g., permission anomalies, data leaks), system maintenance (e.g., performance degradation, resource exhaustion), and access control (e.g., unauthorized access, unauthorized operations). A separate data structure is created for each type of alarm data, including fields such as the alarm ID, alarm time, alarm description, and impact range, forming a structured alarm response pool.
[0102] The alarm response pool, as a centralized storage and processing center for alarm data, requires data linkage with a historical response plan library. This library stores verified, effective response plans, each containing trigger conditions, response steps, and execution parameters. Using a feature matching algorithm, the feature vector of the current alarm data is compared with the feature templates of the historical plans, and a similarity score is calculated. The historical plan with the highest similarity is selected as the response basis. Based on the alarm type, corresponding response instructions are generated: storage security response instructions formulate protective measures against security threats, system maintenance instructions address system-level anomalies, and access control instructions regulate data access. Priority analysis and dependency analysis are performed on the generated response instructions. Priority analysis considers factors such as alarm urgency, impact, and resource consumption, assigning a priority tag to each instruction. Dependency analysis identifies execution order constraints between instructions. For example, certain security hardening instructions must be executed after system maintenance is completed, and access control modifications must take effect after security policy updates. By constructing an instruction dependency table, the preconditions and subsequent impacts between instructions are recorded, forming an instruction execution sequence that satisfies dependency constraints.
[0103] The resource requirements of storage nodes are assessed based on the instruction execution sequence. The CPU resources, memory capacity, network bandwidth, and other parameters required for each instruction are counted, and the node load distribution is calculated based on the current resource usage of each node. Based on the load balancing principle, a node resource allocation plan is formulated to ensure that the execution of disposal instructions does not lead to node resource exhaustion. A two-phase commit protocol is used to ensure the consistent execution of distributed instructions. In the first phase (preparation phase), pre-execution instructions are sent to each node, and the node verifies resource conditions and locks the required resources. In the second phase (commit phase), after confirming that all nodes are ready, the execution instructions are officially issued. If any node reports insufficient resources during the preparation phase, the resource locks of all nodes are rolled back and the allocation plan is readjusted.
[0104] Implement real-time monitoring during the execution of instructions and record the changes in the execution status of the instructions. The monitoring content includes information such as the instruction start time, execution progress, resource consumption, and completion status. When an abnormal instruction execution is detected, take timely intervention measures, such as adjusting execution parameters, switching to alternative solutions, etc. All monitoring records and execution results are recorded in the storage system's disposal instruction database to provide a basis for subsequent scheme optimization.
[0105] For example: When it is detected that the space utilization rate of a certain storage node exceeds the threshold, a system maintenance class alarm is triggered. After the alarm disposal response pool receives the alarm signal, it matches the disposal scheme of data migration from the historical scheme library. The generated system maintenance instruction includes steps such as source node selection, target node screening, and data block migration. Determine the access control policy and security protection measures during the migration process through dependency analysis. After evaluating the storage space and network bandwidth of each node, select the node with lower load as the migration target. Use two-phase commit to ensure the atomicity of the migration operation, and monitor the migration progress and resource usage in real time to complete the secure migration of data.
[0106] In a specific embodiment, the process of executing step S106 may specifically include the following steps:
[0107] (1) Perform execution tracking on the storage system disposal instructions, respectively count the instruction response time, resource utilization rate, and problem-solving rate, and generate an instruction execution effect report;
[0108] (2) Based on the instruction execution effect report, quantitatively evaluate the disposal effects of the storage security disposal instructions, system maintenance instructions, and access control instructions to form a disposal instruction score;
[0109] (3) For the instruction types with disposal instruction scores lower than the threshold, extract their corresponding monitoring parameters and establish a monitoring parameter optimization table;
[0110] (4) Compare and analyze the monitoring parameter optimization table with the historical monitoring data, calculate the parameter deviation value, and generate a monitoring parameter adjustment scheme;
[0111] (5) According to the monitoring parameter adjustment scheme, correct the alarm trigger threshold, sampling frequency, and data alignment period to form an alarm strategy optimization list;
[0112] (6) Integrate the alarm strategy optimization list and the monitoring parameter adjustment scheme to construct a distributed storage monitoring optimization scheme.
[0113] Specifically, an execution tracking mechanism was established to monitor the entire execution process of storage system instructions. Instruction response time was calculated by recording the time the instruction was issued and the time it was completed. Resource utilization was measured by collecting metrics such as CPU usage, memory utilization, and I / O bandwidth. Problem resolution rates were calculated based on the resolution of abnormal conditions. This data was compiled into a structured instruction execution performance report, which included execution details and performance metrics for each instruction.
[0114] Based on the instruction execution effect report, a quantitative assessment is conducted, and scoring criteria are established for storage security disposal instructions, system maintenance instructions, and access control instructions. Scoring dimensions include execution efficiency (the quality of instruction response time), resource consumption (the rationality of resource utilization), and disposal effectiveness (the level of problem resolution). A weight coefficient is assigned to each dimension, and a comprehensive calculation is performed to determine the disposal instruction score. The scoring results reflect the disposal effectiveness of different types of instructions. For instructions whose disposal instruction scores fall below the preset threshold, their performance bottlenecks need to be analyzed. The monitoring parameters corresponding to these instructions are extracted, including configuration items such as sampling frequency, alarm thresholds, and data alignment settings. A monitoring parameter optimization table is established to record the current value, reference value, and optimization direction of the parameters to be optimized. These parameters directly affect the accuracy and real-time performance of monitoring and are key factors in improving disposal effectiveness.
[0115] Compare and analyze the monitoring parameter optimization table with historical monitoring data, and calculate the parameter deviation value. By looking back at historical data, find out the parameter configuration that performs better in similar scenarios as a reference benchmark for optimization. Calculate the difference between the current parameters and the reference values, while taking into account changes in the system operating environment, and generate a practical monitoring parameter adjustment plan. Modify specific monitoring parameters according to the adjustment plan. The adjustment of the alarm trigger threshold takes into account changes in system load to avoid excessive false alarms; the correction of the sampling frequency needs to balance monitoring accuracy and resource overhead; the optimization of the data alignment cycle focuses on the real-time requirements of data synchronization. These corrections form a complete alarm strategy optimization list, which records in detail the adjustment value and reason for each parameter. Figure 3 A schematic diagram of the parameter types of the alarm strategy optimization list in an embodiment of the present application.
[0116] Integrate the alert strategy optimization checklist with the monitoring parameter adjustment plan to build a distributed storage monitoring optimization solution. This optimization plan includes the specific steps for parameter adjustment, the implementation sequence, and verification methods to ensure the controllability and effectiveness of the optimization process. This closed-loop optimization process continuously improves the performance of the monitoring system through continuous evaluation and adjustment.
[0117] Taking the performance monitoring scenario of a file storage system as an example: When an I / O performance degradation warning occurs, the disposal instructions include measures such as load balancing and cache optimization. Through execution tracking, it is found that the response time of the load balancing instruction is relatively long and the resource consumption is relatively high, resulting in a relatively low score for the disposal instruction. Analyzing the corresponding monitoring parameters, it is found that the sampling frequency of the performance data is too high, causing a large amount of resource consumption. Comparing the historical monitoring data, under similar load conditions, a lower sampling frequency can also meet the monitoring requirements. Based on this, a parameter adjustment plan is generated to reduce the sampling frequency and optimize the data alignment period to make the monitoring more efficient. These optimization measures are integrated into the storage monitoring optimization plan, and the optimization effect is verified through implementation.
[0118] In a specific embodiment, the process of executing the step of calculating the load distribution of computing nodes may specifically include the following steps:
[0119] (1) Conduct a resource requirement analysis for the instruction execution sequence, count the CPU occupancy, memory usage, and storage space occupancy, and generate a node resource statistics table;
[0120] (2) Extract the real-time resource status of each storage node from the node resource statistics table, calculate the CPU utilization rate, memory utilization rate, and storage utilization rate, and form a node load status matrix;
[0121] (3) Conduct historical data analysis on the node load status matrix, draw the load change curves of each node, and generate a node load change trend table;
[0122] (4) Calculate the remaining resource capacity of each storage node based on the node load change trend table, count the number of allocable resources, and form a resource capacity distribution map;
[0123] (5) Perform matching analysis on the resource capacity distribution map and the instruction execution sequence, group the instructions according to the remaining resources of the nodes, and establish an instruction allocation table;
[0124] (6) According to the instruction allocation table, conduct resource allocation planning for each storage node to generate a node resource allocation plan.
[0125] Specifically, for the resource requirement analysis of the instruction execution sequence, collect the resource usage of each storage node, including CPU occupancy (recording the actual number of processor cores used, the number of threads used, and the process queue length), memory usage (counting the size of the allocated memory space, cache usage, and page swap rate), and storage space occupancy (calculating the used disk capacity, I / O queue length, and disk read / write rate). These data are sorted into a structured node resource statistics table according to the node ID, resource type, and timestamp, and each record in the table contains complete resource usage information.
[0126] The calculation of the node load status adopts the following formula:
[0127]
[0128] in: Represents the total score value of the node load status matrix; Indicates the real-time CPU usage of the i-th node; Indicates the CPU resource utilization weight factor; Indicates the real-time memory usage of the i-th node; Indicates the utilization weight factor of memory resources; Indicates the real-time storage usage of the i-th node; Indicates the utilization weight factor of storage resources; Indicates the total number of storage nodes.
[0129] This formula uses a weighted calculation to create a load state matrix, where the weight factors reflect the importance of different resource types. For example, CPU utilization is calculated by dividing the number of currently occupied cores by the total number of cores. The weight factor is typically set to 0.4, indicating its importance in the overall assessment.
[0130] Perform historical data analysis on the node load status matrix, collect load data from the past period (e.g., 24 hours), and plot a load change curve through time series analysis. The curve records the dynamic trend of resource usage, forming a node load change trend table.
[0131] When extracting data from node resource statistics, consider the timeliness and integrity of the data. For each storage node, first verify that the data collection timestamp is within the validity period (typically the last 5 minutes) to ensure the latest resource status data is used. The extracted data is then preprocessed, including outlier filtering (removing data that significantly deviates from the normal range), missing value interpolation (interpolating missing values using valid data from nearby time points), and data normalization (converting metrics of different dimensions to a unified calculation benchmark).
[0132] When analyzing historical data of the node load status matrix, a sliding time window approach is used to process the data. The window size is set based on the monitoring period, typically 24 hours, with a sliding step of 5 minutes. Within each time window, statistical characteristics of resource usage are calculated, including the mean, standard deviation, peak values, and valley values. These statistics are used to plot load change curves, reflecting dynamic trends in resource usage. Furthermore, time series analysis methods (such as moving average and exponential smoothing) are used to predict short-term resource usage trends, providing predictive support for resource allocation.
[0133] The following formula is used to calculate the remaining resource capacity of a node:
[0134]
[0135] Wherein: represents the total available resources of the node; represents the total capacity value of the j-th type of resource; represents the current occupancy rate of the j-th type of resource; represents the resource elasticity adjustment coefficient; represents the resource reliability coefficient; V represents the number of resource types. This formula comprehensively considers the total amount of resources, the current occupancy situation, and the reliability factor, and calculates the actually allocable amount of resources. The result forms a resource capacity distribution map, which intuitively shows the resource distribution of each node.
[0136] When matching and analyzing the resource capacity distribution map with the instruction execution sequence, a multi-dimensional evaluation method is adopted. First, according to the resource requirement characteristics of the instructions, the instructions are classified into compute-intensive (mainly consuming CPU resources), memory-intensive (mainly consuming memory resources), and I / O-intensive (mainly consuming storage resources). Then, for each type of instruction, a scoring standard for resource matching is set, including resource sufficiency (the ratio of the remaining resources of the node to the required resources of the instruction), load balance (the difference between the current load of the node and the average load of the cluster), and performance stability (the degree of fluctuation of the resource usage of the node).
[0137] An instruction allocation table is established according to the scoring results. The table records the priority, resource requirements, target node, and reserved resource amount of each instruction. The generation of the allocation table needs to meet both resource constraints (not exceeding the available resources of the node) and performance constraints (ensuring the load balance of the node). For resource conflict situations, adjustments are made according to the instruction priority, and the resource preemption mechanism is triggered if necessary.
[0138] The finally generated node resource allocation plan includes a complete execution plan, which clearly specifies the set of instructions to be executed by each node, the resource allocation quota, and the execution time window. The plan also includes resource monitoring thresholds for timely detection and handling of resource anomalies during execution. The entire resource allocation process forms a closed loop, ensuring the efficiency and reliability of resource usage through continuous monitoring and adjustment.
[0139] In a specific embodiment, the process of performing the step of grouping instructions according to the remaining resources of the node may specifically include the following steps:
[0140] (1) Perform resource threshold partitioning on the resource capacity distribution map, divide it into a high-availability resource area, a medium-availability resource area, and a low-availability resource area, and form a resource partition table;
[0141] Based on the resource partition table, extract the resource distribution of each node, count the number of nodes in each resource area, and generate a node resource distribution matrix;
[0142] (3) Mark the resource requirement levels of the instruction execution sequence, classify the instructions according to their requirements for the CPU, memory, and storage space, and establish an instruction resource requirement table;
[0143] (4) Perform a matching calculation between the node resource distribution matrix and the instruction resource requirement table, sort the instructions according to the resource matching degree to form an instruction priority list;
[0144] (5) For each instruction in the instruction priority list, select the optimal node in the corresponding resource area for resource pre-allocation to generate a resource allocation mapping table;
[0145] (6) Integrate the corresponding relationship between the instructions and the nodes according to the resource allocation mapping table to establish an instruction allocation table.
[0146] Specifically, perform a partitioning process on the resource capacity distribution map, set a partitioning threshold according to the resource availability of the nodes, classify the nodes with sufficient resources (availability rate exceeding 80%) into the high-availability resource area, the nodes with moderate resources (availability rate between 50% - 80%) into the medium-availability resource area, and the nodes with tight resources (availability rate below 50%) into the low-availability resource area. The partitioning information is recorded in the resource partitioning table, which contains fields such as node ID, resource type, availability rate, and the affiliated area. When extracting the node distribution from the resource partitioning table, it is necessary to count the number of nodes and their resource status in each resource area. The counting process is carried out separately according to the resource type (CPU, memory, storage space), and the distribution characteristics of various resources in different areas are recorded. These statistical data are organized into a node resource distribution matrix, where the rows of the matrix represent the resource areas, the columns represent the resource types, and the matrix elements record the number of nodes and the total resource volume under the corresponding conditions.
[0147] For each instruction sequence to be executed, the resource requirements of each instruction must be analyzed. Resource requirement analysis is conducted along three dimensions: CPU demand (number of processor cores and computation time), memory demand (memory size and access frequency), and storage demand (disk space and I / O bandwidth). Based on the size of the demand, instructions are categorized into three levels: high demand (a single resource requirement exceeding 60% of the node's total resource requirements), medium demand (a requirement between 30% and 60%), and low demand (a requirement below 30%). The classification results are recorded in an instruction resource requirement table, which contains the instruction ID, resource requirement vector, and requirement level tag. The node resource distribution matrix is matched against the instruction resource requirement table, using a resource matching scoring mechanism. This matching calculation considers several factors: resource satisfaction (the ratio of available node resources to instruction requirements), load balancing (load change after node allocation), and resource affinity (the instruction's preference for specific resource types). Instructions are prioritized based on the matching scores to form an instruction priority list. This list is sorted by score, and the resource region for each instruction is also recorded.
[0148] For each instruction in the instruction priority list, the optimal execution node is selected within its corresponding resource region. The node selection process first considers resource sufficiency to ensure that the node has sufficient resources to meet the instruction requirements; then, it evaluates load balancing to avoid over-concentration of resource allocation; and finally, it considers the node's historical execution history, prioritizing nodes with stable execution. The selection results are recorded in a resource allocation mapping table, which contains information such as the instruction ID, target node, and pre-allocated resource amount. The corresponding relationships in the resource allocation mapping table are integrated to generate a complete instruction allocation table. The allocation table serves as the basis for resource scheduling and details the set of instructions to be executed by each node, the allocated quota for each resource, and the priority order of instruction execution. The table also includes resource reservation tags to prevent over-allocation of resources.
[0149] For example, when multiple data migration tasks are required, storage nodes are divided into zones based on available free space, identifying a set of nodes suitable as migration targets. The resource requirements of the migration instructions are analyzed, including the source node's read bandwidth, the target node's write space, and network transmission resources. Through matching calculations, large-capacity migration tasks are preferentially assigned to nodes in high-availability resource zones, while small-capacity migration tasks are assigned to nodes in medium-availability resource zones. Finally, the instruction allocation table clearly defines the source node, target node, and resource quota for each migration task, guiding the orderly progress of the migration process.
[0150] The above describes the file storage real-time monitoring method in the embodiment of the present application. The following describes the file storage real-time monitoring system in the embodiment of the present application. Figure 4 In one embodiment of the present application, a system for real-time monitoring of file storage includes:
[0151] The allocation module is used to collect the operating status information, storage security status information, and system alarm trigger information of the file storage cluster through distributed monitoring agents, distribute monitoring tasks to each monitoring agent node, and form the original monitoring alarm data flow after establishing a distributed monitoring network;
[0152] An alignment module is used to align the original monitoring alarm data stream according to the monitoring period, extract storage alarm features, storage security features, and system status features in an adaptive manner, and generate a distributed storage early warning feature data set;
[0153] A prediction module, configured to predict storage security risks, system alarm risks, and data access risks based on the distributed storage warning feature data set combined with the warning prediction model, and form a storage system warning result;
[0154] A classification module is used to classify the storage system warning results into storage security warning levels, system warning levels, and access warning levels, and select distributed warning signal transmission channels according to different levels to form a storage system hierarchical warning signal;
[0155] The issuing module is used to generate storage security disposal instructions, system maintenance instructions, and access control instructions according to the storage system graded alarm signals, coordinate the resources of each storage node to perform distributed instruction issuance, and establish storage system disposal instructions;
[0156] The monitoring module is used to evaluate the execution effect of the storage system processing instructions, adjust the monitoring parameters and alarm strategy according to the evaluation results, and form a distributed storage monitoring optimization plan.
[0157] Through the collaborative cooperation of the above-mentioned various components, the multi-dimensional status information of the storage cluster is dynamically collected by the distributed monitoring agent, realizing the comprehensive monitoring of the running status of the storage system and avoiding the problem of information omission caused by single-index monitoring. By processing the original monitoring data through data alignment and adaptive feature extraction methods, not only the temporal consistency of multi-source data is ensured, but also the feature extraction strategy can be automatically adjusted according to the characteristics of data changes, improving the representativeness and accuracy of the feature data. Based on the early warning prediction model, the system risks are predicted and analyzed. Through the collaborative work of multiple prediction units, including the early warning perception unit for real-time analysis of the current status, the early warning memory unit for storing historical patterns, and the early warning reasoning unit for risk inference, the accuracy and timeliness of the early warning are significantly enhanced. An adaptive alarm classification mechanism is adopted to dynamically divide the alarm levels according to the early warning results and select different signal transmission channels, which not only ensures the timely delivery of high-risk alarms but also avoids the waste of signal transmission resources. In the link of generating disposal instructions, through the associated matching with historical disposal solutions and combining with the current system status, the disposal strategy is intelligently generated, and the consistent execution of distributed instructions is ensured through the two-phase commit method. Finally, through the execution effect evaluation and parameter optimization mechanism, the monitoring and early warning strategy is continuously improved, enabling the entire monitoring system to have the ability of adaptive learning. This solution makes full use of the advantages of the early warning prediction model in time series data analysis and risk prediction. Through the organic combination of algorithms such as feature extraction, pattern recognition, and risk inference, the transformation from passive monitoring to active early warning is realized. At the same time, the adaptive optimization mechanism of the early warning model ensures the stability and reliability of the algorithm in different application scenarios.
[0158] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0159] The above is the case. The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A real-time monitoring method for file storage, characterized in that, The described file storage real-time monitoring method includes: Collecting the operation status information, storage security status information, and system alarm trigger information of the file storage cluster through distributed monitoring agents, allocating monitoring tasks to each monitoring agent node, and forming an original monitoring alarm data stream after establishing a distributed monitoring network; Aligning the original monitoring alarm data stream according to the monitoring period, and adopting an adaptive method to extract storage alarm features, storage security features, and system status features to generate a distributed storage early warning feature dataset; Combining an early warning prediction model based on the distributed storage early warning feature dataset to predict storage security risks, system alarm risks, and data access risks, and forming a storage system early warning result, including: establishing a time series early warning training sample according to the distributed storage early warning feature dataset, and constructing storage alarm features, storage security features, and system status features into a feature input sequence; performing data correlation analysis on the feature input sequence through an early warning perception unit to generate early warning perception features; storing historical early warning data through an early warning memory unit to form early warning memory features; performing feature reasoning through an early warning reasoning unit to generate early warning reasoning features; inputting the early warning perception features, early warning memory features, and early warning reasoning features into the early warning prediction model to calculate the weights of the feature data and generate an early warning parameter matrix; based on the early warning parameter matrix, performing forward calculation on the newly input distributed storage early warning feature dataset to obtain storage security risk values, system alarm risk values, and data access risk values; optimizing the parameters of the early warning parameter matrix through backpropagation, adjusting the weight parameters of each unit in the early warning prediction model until the prediction accuracy meets the requirements; using the optimized early warning prediction model to predict and analyze storage security risks, system alarm risks, and data access risks to form a storage system early warning result; Dividing the storage security alarm level, system alarm level, and access alarm level according to the storage system early warning result, and selecting a distributed early warning signal transmission channel according to different levels to form a storage system hierarchical alarm signal; Generating a storage security disposal instruction, a system maintenance instruction, and an access control instruction according to the storage system hierarchical alarm signal, coordinating the resources of each storage node to issue distributed instructions, and establishing a storage system disposal instruction; Evaluating the execution effect of the storage system disposal instruction, and adjusting the monitoring parameters and alarm strategies according to the evaluation result to form a distributed storage monitoring optimization plan.
2. The real-time monitoring method for file storage according to claim 1, wherein The process of collecting the operation status information, storage security status information, and system alarm trigger information of the file storage cluster through distributed monitoring agents, allocating monitoring tasks to each monitoring agent node, and forming an original monitoring alarm data stream after establishing a distributed monitoring network includes: Obtaining the read / write operation logs, storage capacity usage data, and I / O performance data of the file storage cluster through distributed monitoring agents, and partitioning the file storage cluster according to the node load conditions; Allocating the read / write operation logs, storage capacity usage data, and I / O performance data to different monitoring agent nodes, and storing the collected data using a local cache mechanism; Based on the collected data in the local cache mechanism, dynamically schedule and allocate the data transmission tasks between monitoring agent nodes; Perform task allocation for monitoring agent nodes according to the dynamic scheduling and allocation results. The monitoring agent nodes collect operation status information, storage security status information, and system alarm trigger information; Summarize the operation status information, storage security status information, and system alarm trigger information to the monitoring network center node to form an original monitoring alarm data stream.
3. The file storage real-time monitoring method according to claim 1, characterized in that: Align the original monitoring alarm data stream according to the monitoring period, and adopt an adaptive method to extract storage alarm features, storage security features, and system status features to generate a distributed storage warning feature dataset, including: Perform time series division on the original monitoring alarm data stream, mark the data within the monitoring period with timestamps to form time series monitoring data; Divide the time series monitoring data into an operation status sequence, a storage security sequence, and an alarm trigger sequence according to the data type, and perform time alignment calibration on the data sequences; Extract storage performance indicators, resource utilization indicators, and system load indicators for the operation status sequence to generate storage alarm features; Extract data access patterns, security authentication records, and operation behavior characteristics for the storage security sequence to generate storage security features; Extract alarm frequency features, alarm duration features, and alarm correlation features for the alarm trigger sequence to generate system status features; Integrate and correlate the storage alarm features, storage security features, and system status features, and perform feature combination according to the monitoring period to generate a distributed storage warning feature dataset.
4. The real-time monitoring method for file storage according to claim 1, characterized in that, Divide the storage security alarm level, system alarm level, and access alarm level according to the storage system warning result, and select a distributed warning signal transmission channel according to different levels to form a storage system hierarchical alarm signal, including: Perform risk scoring on the storage system warning result, calculate the storage security risk score, system alarm risk score, and access risk score respectively to generate risk scoring data; Divide the risk scoring data into a high-risk interval, a medium-risk interval, and a low-risk interval according to a preset risk threshold, and correspondingly generate a storage security alarm level, a system alarm level, and an access alarm level; Establish a security warning channel for the storage security alarm level, establish a maintenance warning channel for the system alarm level, and establish a control warning channel for the access alarm level to form a distributed warning signal transmission channel; Send the alarm signals in the high-risk interval through a multi-path transmission method, send the alarm signals in the medium-risk interval through a message queue method, and send the alarm signals in the low-risk interval through a normal data channel; Monitor the network status for the distributed warning signal transmission channel, evaluate the signal transmission quality, and select the optimal transmission path; Send the alarm signals of the storage security alarm level, system alarm level, and access alarm level through the corresponding distributed warning signal transmission channels to form a storage system hierarchical alarm signal.
5. The file storage real-time monitoring method according to claim 1, characterized in that: Generate storage security handling instructions, system maintenance instructions, and access control instructions according to the hierarchical alarm signals of the storage system, coordinate the resources of each storage node for distributed instruction issuance, and establish storage system handling instructions, including: Parse the hierarchical alarm signals of the storage system to obtain three types of alarm data: storage security, system maintenance, and access control, and establish an alarm handling response pool; Associate and match the data in the alarm handling response pool with historical handling solutions to generate storage security handling instructions, system maintenance instructions, and access control instructions respectively; Mark the priorities of the storage security handling instructions, system maintenance instructions, and access control instructions, construct an instruction dependency relationship table, and form an instruction execution sequence; Based on the instruction execution sequence, count the resource occupancy of each storage node, calculate the node load distribution, and generate a node resource allocation plan; Combine the node resource allocation plan with the instruction execution sequence, and issue task instructions to each storage node through a two-phase commit method; Monitor the execution process of the task instructions in real time, record the instruction response status, and establish storage system handling instructions.
6. The real-time monitoring method for file storage according to claim 1, wherein Evaluate the execution effect of the storage system handling instructions, and adjust the monitoring parameters and alarm strategies according to the evaluation results to form a distributed storage monitoring optimization plan, including: Track the execution of the storage system handling instructions, count the instruction response time, resource utilization rate, and problem resolution rate respectively, and generate an instruction execution effect report; Based on the instruction execution effect report, quantitatively evaluate the handling effects of the storage security handling instructions, system maintenance instructions, and access control instructions to form a handling instruction score; For the instruction types with handling instruction scores lower than the threshold, extract their corresponding monitoring parameters and establish a monitoring parameter optimization table; Compare and analyze the monitoring parameter optimization table with historical monitoring data, calculate the parameter deviation value, and generate a monitoring parameter adjustment plan; According to the monitoring parameter adjustment plan, correct the alarm trigger threshold, sampling frequency, and data alignment period to form an alarm strategy optimization list; Integrate the alarm strategy optimization list and the monitoring parameter adjustment plan to construct a distributed storage monitoring optimization plan.
7. The real-time monitoring method for file storage according to claim 5, characterized in that, Based on the instruction execution sequence, count the resource occupancy of each storage node, calculate the node load distribution, and generate a node resource allocation plan, including: Conduct a resource requirement analysis for the instruction execution sequence, count the CPU occupancy, memory usage, and storage space occupancy, and generate a node resource statistics table; Extract the real-time resource status of each storage node from the node resource statistics table, calculate the CPU utilization rate, memory utilization rate, and storage utilization rate, and form a node load status matrix; Conduct historical data analysis on the node load status matrix, draw the load change curve of each node, and generate a node load change trend table; Based on the node load change trend table, calculate the remaining resource capacity of each storage node, count the number of allocable resources, and form a resource capacity distribution map; Match and analyze the resource capacity distribution map with the instruction execution sequence, group the instructions according to the remaining resource amount of the nodes, and establish an instruction allocation table; Resource allocation planning is carried out for each storage node according to the instruction allocation table to generate a node resource allocation plan.
8. The file storage real-time monitoring method according to claim 7, characterized in that: Matching and analyzing the resource capacity distribution map with the instruction execution sequence, grouping the instructions according to the remaining resource amount of the nodes, and establishing an instruction allocation table, including: Performing resource threshold partitioning on the resource capacity distribution map, dividing it into a high-availability resource area, a medium-availability resource area, and a low-availability resource area to form a resource partitioning table; Extracting the resource distribution of each node based on the resource partitioning table, counting the number of nodes in each resource area, and generating a node resource distribution matrix; Marking the resource requirement levels of the instruction execution sequence, grading the instructions according to the requirements for CPU, memory, and storage space, and establishing an instruction resource requirement table; Performing matching calculations on the node resource distribution matrix and the instruction resource requirement table, sorting the instructions according to the resource matching degree to form an instruction priority list; For each instruction in the instruction priority list, select the optimal node in the corresponding resource area for resource pre-allocation to generate a resource allocation mapping table; Integrate the corresponding relationship between instructions and nodes according to the resource allocation mapping table to establish an instruction allocation table.
9. A file storage real-time monitoring system for implementing the file storage real-time monitoring method according to any one of claims 1-8, characterized in that, The file storage real-time monitoring system includes: An allocation module, which is used to collect the operation status information, storage security status information, and system alarm trigger information of the file storage cluster through a distributed monitoring agent, allocate monitoring tasks to each monitoring agent node, and form an original monitoring alarm data stream after establishing a distributed monitoring network; An alignment module, which is used to align the original monitoring alarm data stream according to the monitoring period, and adopt an adaptive method to extract storage alarm features, storage security features, and system status features to generate a distributed storage warning feature data set; A prediction module, which is used to combine a warning prediction model based on the distributed storage warning feature data set to predict storage security risks, system alarm risks, and data access risks to form a storage system warning result, including: establishing a time-series warning training sample according to the distributed storage warning feature data set, and constructing the storage alarm feature, storage security feature, and system status feature into a feature input sequence; performing data correlation analysis on the feature input sequence through a warning perception unit to generate a warning perception feature; storing the historical warning data through a warning memory unit to form a warning memory feature; performing feature reasoning through a warning reasoning unit to generate a warning reasoning feature; inputting the warning perception feature, warning memory feature, and warning reasoning feature into the warning prediction model to calculate the weights of the feature data and generate a warning parameter matrix; based on the warning parameter matrix, performing forward calculation on the newly input distributed storage warning feature data set to obtain a storage security risk value, a system alarm risk value, and a data access risk value; optimizing the parameters of the warning parameter matrix through backpropagation, adjusting the weight parameters of each unit in the warning prediction model until the prediction accuracy meets the requirements; using the optimized warning prediction model to predict and analyze storage security risks, system alarm risks, and data access risks to form a storage system warning result; A partitioning module, which is used to partition the storage security alarm level, system alarm level, and access alarm level according to the warning result of the storage system, select a distributed warning signal transmission channel according to different levels, and form a hierarchical alarm signal of the storage system; A distribution module, which is used to generate storage security handling instructions, system maintenance instructions, and access control instructions according to the hierarchical alarm signal of the storage system, coordinate the resources of each storage node to perform distributed instruction distribution, and establish storage system handling instructions; A monitoring module, which is used to evaluate the execution effect of the storage system handling instructions, adjust the monitoring parameters and alarm strategies according to the evaluation results, and form an optimized distributed storage monitoring plan.
Citation Information
Patent Citations
Intelligent monitoring method and intelligent monitoring system
CN107294764A
Intelligent task alarm rule self-learning method and system based on support priority
CN119441832A