Distributed storage system automatic capacity expansion method and system and readable storage medium
By monitoring system storage pressure and predicting future needs, configuring preset storage nodes in hot standby state and performing incremental migration, the problems of low expansion efficiency and low availability of distributed storage systems are solved, and seamless expansion and efficient operation and maintenance are achieved.
Patent Information
- Application Number
- CN202511092877.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The capacity expansion process of existing distributed storage systems is inefficient and error-prone. Automatic capacity expansion affects online services, increases operation and maintenance costs and complexity, and reduces system availability.
By monitoring system storage pressure, predicting future storage needs, configuring preset storage nodes and setting them to hot standby status, using incremental migration strategies to migrate data, switching to normal working status, and using machine learning algorithm models to optimize the expansion process.
It achieves seamless access to new nodes, shortens failover time, improves system availability and reliability, reduces the impact on online services, reduces manual intervention, and improves expansion efficiency and user experience.
Smart Images

Figure CN120596034A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computers, and in particular to a method and system for automatic capacity expansion of a distributed storage system and a readable storage medium. Background Art
[0002] A distributed storage system is a storage system that stores data in a dispersed manner on different computer nodes and communicates and collaborates through a network to achieve high availability, scalability, and fault tolerance.
[0003] There are two ways to expand distributed storage systems: manual configuration, which involves multiple steps, such as identifying system bottlenecks, planning new nodes, and configuring network connections. Automatic expansion is also available. Current automatic expansion solutions often require service interruptions or performance degradation because they require data redistribution when adding new nodes, a process that can consume significant resources.
[0004] Furthermore, distributed file storage is decentralized, and capacity expansion requires consideration of data distribution balance, leading to the reallocation of all object data. Object data on the original storage device may be relocated to the new storage device, changing its storage location. When the system subsequently attempts to read data, it will not be able to retrieve it if the data is no longer in the original storage location. Users can modify logic code or configuration files to instruct the system to read data from the new storage location, but this requires stopping business operations, impacting system performance.
[0005] During the implementation of this application, the inventors discovered that the prior art has at least the following problems:
[0006] 1) The manual expansion process is inefficient and error-prone;
[0007] 2) Automatic capacity expansion affects online services during the expansion process, resulting in a poor user experience.
[0008] 3) Data redistribution causes changes in the storage location of business data, which affects online business, increases operation and maintenance costs and complexity, and reduces system availability. Summary of the Invention
[0009] To this end, embodiments of the present application provide a distributed storage system automatic expansion method, system, and readable storage medium, which can solve the technical problems of low efficiency and low availability in existing distributed storage system expansion. The specific technical solutions are as follows:
[0010] In a first aspect, an embodiment of the present application provides a method for automatically expanding a distributed storage system, the method comprising: Collect system operation data and use machine learning algorithm models to predict future storage needs based on the system operation data; When future storage requirements exceed the system's current storage capacity, configure preset storage nodes; Detection system storage pressure: When the system storage pressure exceeds the preset system storage threshold, the expansion mode is triggered; Set the preset storage node to hot standby state; Update the placement rule according to the identification information of the preset storage node; Start the background data migration task and migrate the data to the preset storage node using the incremental migration strategy; When the background data migration task is completed, switch the preset storage node to normal working status.
[0011] Preferably, detecting the system storage pressure comprises: monitoring storage-related parameters and / or access-related parameters of the system; The system storage pressure is obtained according to storage-related parameters and / or access-related parameters.
[0012] Preferably, system operation data is collected and a machine learning algorithm model is used to predict future storage needs based on the system operation data, including: Collect system operation data, including storage system logs, performance monitoring data, and user behavior data; Clean the system operation data to obtain valid data; Extract features from valid data to form structured data; Input structured data into a machine learning algorithm model to predict future storage needs.
[0013] Preferably, setting the preset storage node to hot standby state is: Set the system configuration and network configuration of the preset storage node, and set the preset storage node to hot standby status.
[0014] Preferably, the background data migration tasks include: Filter the data to be migrated first based on the preset data priority migration policy: Migrate the priority migration data to the preset storage node using the incremental migration strategy.
[0015] Preferably, the method further comprises: Verify the accuracy and consistency of data migrated to the preset storage node by the background data migration task.
[0016] Preferably, verifying the accuracy and consistency of the data migrated from the preset storage node by the background data migration task includes: Perform verification and comparison, metadata consistency check, random sampling comparison, business process verification and / or log audit on the source node data of background data migration and the preset storage node data; If the data on the background data migration source node is inconsistent with the preset storage node data, restore or re-migrate the data based on the storage system log.
[0017] Preferably, after the background data migration task is completed and the preset storage node is switched to a normal working state, the following steps are included: Send expansion completion information to the administrator.
[0018] In a second aspect, an embodiment of the present application provides a distributed storage system automatic expansion system, the system comprising: The prediction module is used to collect system operation data and use machine learning algorithm models to predict future storage needs based on the system operation data; A judgment module is used to configure a preset storage node when future storage demand exceeds the current storage capacity of the system; Detection module, used to detect system storage pressure: A trigger module is used to trigger the expansion mode when the system storage pressure exceeds a preset system storage threshold; A state switching module is used to set a preset storage node to a hot standby state; An updating module, configured to update a placement rule according to identification information of a preset storage node; The data migration module is used to start the background data migration task and migrate the data to the preset storage node using the incremental migration strategy; The state switching module is used to switch the preset storage node to a normal working state after the background data migration task is completed.
[0019] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the aforementioned methods for automatically expanding the capacity of a distributed storage system.
[0020] In summary, compared with the prior art, the technical solutions provided by the embodiments of the present application have at least the following beneficial effects:
[0021] 1. By monitoring system storage pressure and predicting future storage needs, the system intelligently triggers the expansion process based on the system storage pressure and future storage needs through preset storage thresholds and machine learning algorithm models. New node configuration is completed in advance, allowing new nodes to seamlessly connect to the system, shortening failover time and improving system availability and reliability. This reduces the impact on online systems, improves user experience and business continuity, and through intelligent expansion, reduces manual intervention, improves expansion efficiency, and reduces operation and maintenance costs.
[0022] 2. When the system detects that the system storage pressure exceeds the preset storage threshold and that future storage demand exceeds the system's current storage capacity, the system and network configurations of the preset storage nodes are configured. Appropriate hardware devices are selected, and necessary software and configurations are installed. This allows the preset storage nodes to seamlessly connect to the existing network, reducing the impact of the expansion process on online services and improving expansion efficiency.
[0023] 3. Setting a data priority migration strategy can make the expansion process more orderly, select data with less impact on online services for background migration, and improve system availability;
[0024] 4. After the expansion is completed, the consistency and accuracy of the source node data and the preset storage node data are verified, which effectively reduces the impact on online services and improves the consistency of system data. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flow chart of a distributed storage system automatic expansion method provided in one embodiment of the present application.
[0026] Figure 2 This is one of the flow charts of a distributed storage system automatic expansion method provided by another embodiment of the present application.
[0027] Figure 3 This is the second flow chart of a distributed storage system automatic expansion method provided by another embodiment of the present application.
[0028] Figure 4 This is the third flow chart of a distributed storage system automatic expansion method provided by another embodiment of the present application.
[0029] Figure 5 This is the fourth flow chart of a distributed storage system automatic expansion method provided by another embodiment of the present application.
[0030] Figure 6 This is the fifth flow chart of a distributed storage system automatic expansion method provided by another embodiment of the present application.
[0031] Figure 7 This is the sixth flow chart of a distributed storage system automatic expansion method provided by another embodiment of the present application.
[0032] Figure 8 This is a structural diagram of an automatic capacity expansion system for a distributed storage system provided in one embodiment of the present application. DETAILED DESCRIPTION
[0033] This specific embodiment is merely an explanation of the present application and is not a limitation of the present application. After reading this specification, those skilled in the art may make non-creative modifications to the present embodiment as needed, but as long as they are within the scope of the claims of the present application, they are protected by the patent law.
[0034] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0035] In addition, the term "and / or" in this application is simply a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this application, unless otherwise specified, generally indicates that the related objects are in an "or" relationship.
[0036] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.
[0037] In the present application, the term "at least one" means one or more, and the term "plurality" means three or more. For example, a plurality of first positions means three or more first positions.
[0038] The embodiments of the present application are described in further detail below with reference to the accompanying drawings.
[0039] Reference Figure 1 In one embodiment of the present application, a method for automatically expanding the capacity of a distributed storage system is provided. The main steps of the method are described as follows. The sequence of steps in this embodiment is only one of the steps in the present application:
[0040] S1: Collect system operation data and use machine learning algorithm models to predict future storage needs based on the system operation data;
[0041] S2: When future storage demand exceeds the system's current storage capacity, configure a preset storage node;
[0042] S3: Detection system storage pressure:
[0043] S4: When the system storage pressure exceeds the preset system storage threshold, the expansion mode is triggered;
[0044] S5: Set the preset storage node to hot standby state;
[0045] S6: Update the placement rule according to the identification information of the preset storage node;
[0046] S7: Start the background data migration task and migrate the data to the preset storage node using the incremental migration strategy;
[0047] S8: When the background data migration task is completed, the preset storage node is switched to a normal working state.
[0048] In this embodiment, by actively collecting system operational data, such as storage system logs, performance monitoring data, and user behavior data, existing machine learning algorithm models, including but not limited to time series prediction algorithms (such as ARIMA, SARIMA, and LSTM), regression algorithms (such as linear regression and ridge regression), and ensemble algorithms (such as random forest and gradient boosting), are employed. The collected data is used to train, validate, test, and evaluate the machine learning algorithm models. The final model is then used to predict future storage demand. In this embodiment, the trained model is deployed in a production environment, the system status is monitored in real time, and new data is regularly collected through a feedback loop to ensure the model's predictive accuracy. In this embodiment, the model is evaluated using performance metrics such as mean squared error, root mean square error, and mean absolute error to measure model performance, and cross-validation is used to ensure the model's generalization ability.
[0049] In this embodiment, by comparing the future storage demand predicted by the machine learning algorithm model with the current storage capacity of the system, when the future storage demand exceeds the current storage capacity of the system, it is determined that the storage space of the current distributed storage system will be insufficient, and a preset storage node is configured in advance. The preset storage node of this embodiment is a storage hardware device that does not undertake storage tasks in this system, which is adapted to the current and future storage needs of the system of this embodiment, is compatible with the existing system, has reliable performance and is cost-effective, such as a new host. The configuration of the preset storage node in this application is to configure the software, parameters, network, etc. of the preset storage node so that the preset storage node can be connected to the network and adapted to this system.
[0050] Through the settings of this embodiment, the distributed storage system detects the system storage pressure regularly or at preset intervals. The system storage pressure in this embodiment can be reflected by any one or several of the following data: average load, CPU usage, storage utilization, I / O request rate, network bandwidth usage, storage node load, data redundancy and backup strategy and other indicators. In this embodiment, the system storage threshold is preset in advance. For example, the system storage threshold is set to 80%. When the system detects in real time that the current system storage pressure exceeds 80%, it means that the system may face insufficient storage space, and the expansion mode is triggered at this time. In other embodiments of the present application, the system storage pressure can also be reflected by two or more indicators such as storage utilization and I / O read and write times. The scheme of judging the system storage pressure based on various indicators related to data storage and reading and writing of the system is within the scope of this application. In other implementation methods of the present application, when the future storage demand is greater than the current storage capacity of the system and the system storage pressure exceeds the preset system storage threshold, the preset storage node configuration is performed and the expansion mode is triggered.
[0051] When expansion mode is triggered, the preset storage node is placed in hot standby mode. In this state, the preset storage node does not participate in the current system's data storage. Simultaneously, the preset storage node's identification information is updated within the placement rules, enabling the preset storage node to receive system-stored data. Placement rules are key strategies or algorithms used in distributed storage systems to determine the storage location of data objects. Updating placement rules allows newly added preset storage nodes to participate in data storage, utilizing their storage capabilities while maintaining system performance and reliability. The update process involves identifying the preset storage node as a new node and assigning it a unique identifier such as an IP address, hostname, or UUID. The existing rules are then modified in the storage system's management interface or configuration file to incorporate the new node's identification information. Data is then redistributed according to the updated placement rules, and testing is performed to ensure that the new node can correctly receive and process data. The identification information uniquely identifies each node, enabling the system to accurately route data to a specific node according to the placement rules. In this embodiment, data placement rules can select storage pools with more remaining capacity or dynamically select based on the storage pool's load, depending on the actual situation.
[0052] After that, the background data migration task is started to migrate part of the data originally stored in the distributed storage system from the existing node to the preset storage node (i.e., the new node). In this embodiment, in order to avoid affecting the online business, an incremental migration strategy is used to gradually migrate the data to reduce the impact on system performance.
[0053] When the background data migration task is completed, the preset storage node is switched from the hot standby state to the normal working state, that is, the preset storage node (new node) is connected to the distributed storage system, starts to receive new data storage requests, and completes automatic expansion.
[0054] Therefore, in this embodiment, by continuously monitoring the system storage pressure, continuously collecting system operation data, and predicting future storage needs, the expansion process is intelligently triggered based on the system storage pressure and future storage needs through preset storage thresholds and machine learning algorithm models, and the new node configuration is completed in advance, so that the new node can be seamlessly connected to the system, shortening the fault switching time, improving the system's availability and reliability, reducing the impact on the online system, improving user experience and business continuity, and through intelligent expansion, reducing manual intervention, improving expansion efficiency, and reducing operation and maintenance costs.
[0055] Reference Figure 2 In another embodiment of the present application, step S3 includes:
[0056] S31: monitoring storage-related parameters and / or access-related parameters of the system;
[0057] S32: Obtaining system storage pressure according to storage-related parameters and / or access-related parameters.
[0058] In this embodiment, the system's storage-related parameters, i.e., parameters associated with system storage data, include storage utilization, I / O request rate, network bandwidth usage, etc. Access-related parameters include system storage pool usage and access frequency, etc. This embodiment collects relevant parameters periodically, such as every minute. In other embodiments of the present application, other collection frequencies may be used, such as once an hour.
[0059] System storage pressure is determined by storage-related parameters and / or access-related parameters. In one example of this embodiment, the system storage utilization rate is used as the system storage pressure. This directly indicates the available space in the current system and reduces the need for premature capacity expansion. In other examples of this embodiment, system storage pressure can be determined using both storage-related and access-related parameters, such as storage utilization rate and storage pool access frequency. In this case, the determined system storage pressure is more predictable and can be more accurately determined based on the current user profile (derived from storage utilization rate and access frequency).
[0060] Reference Figure 3 In another embodiment of the present application, step S1 includes:
[0061] S11: Collect system operation data, including storage system logs, performance monitoring data, and user behavior data;
[0062] S12: Clean the system operation data to obtain valid data;
[0063] S13: Extract features from valid data to form structured data;
[0064] S14: Input structured data into the machine learning algorithm model to predict future storage needs.
[0065] In this embodiment, the system operation data is information collected from multiple data sources, including storage system logs, performance monitoring data, and user behavior data. These data are cleaned to remove invalid, erroneous, and duplicate information to ensure the quality and accuracy of the data. Feature extraction is then performed on the valid data to identify key features that have a significant impact on storage requirements. In this embodiment, these features include timestamps, storage utilization, I / O read and write times, and file size. In other embodiments of the present application, other attributes such as data type and size may also be used, which are not limited here. In this embodiment, data is optimized through feature selection, feature conversion, and feature construction in the feature extraction stage. First, correlation analysis, information gain, and other methods are used to screen out the features that have the greatest impact on the prediction target to improve the predictive ability of the model. Then, the data is normalized or standardized to make it more suitable for algorithmic processing of the machine learning algorithm model. In addition, in other examples of this implementation, new features such as historical growth rates, seasonal changes, etc. can be created to further optimize the data to enrich the information dimension of the data. By performing operations such as cleaning and feature extraction on the system operation data, the system operation data is transformed into structured data. The structured data is suitable as the input value of the machine learning algorithm model, which reduces the complexity of the system operation data and improves the prediction accuracy of the machine learning algorithm model.
[0066] Reference Figure 4 In another embodiment of the present application, step S5 is:
[0067] S51: Setting the system configuration and network configuration of the preset storage node, and setting the preset storage node to a hot standby state.
[0068] The system configuration of a computer storage device mainly includes software configuration and hardware configuration. In this embodiment, when the system detects that the system storage pressure exceeds the preset storage threshold and that future storage demand exceeds the current storage capacity of the system, the system configuration and network configuration of the preset storage node are configured, appropriate hardware devices are selected, and necessary software and configurations are installed, so that the preset storage node (new node) can be seamlessly connected to the existing network, reducing the impact of the expansion process on online services and improving expansion efficiency.
[0069] In another embodiment of the present application, the background migration task includes:
[0070] A00: Filter the data to be migrated first based on the preset data priority migration policy:
[0071] A01: Migrate the priority migration data to the preset storage node using the incremental migration strategy.
[0072] Prioritized migration strategies can be set based on the actual system usage scenario. In one solution of this application, data can be prioritized based on data characteristics, such as access frequency (prioritizing cold data), data size (large files vs. small files), data type (such as backup or archived data), and node load (migrating from high-load nodes). Setting a data-prioritized migration strategy can streamline the expansion process, select data with less impact on online services for background migration, and improve system availability.
[0073] Reference Figure 5 In another embodiment of the present application, the method further includes:
[0074] S9: Verify the accuracy and consistency of the data migrated by the background data migration task to the preset storage node.
[0075] Through the configuration of this embodiment, the consistency of system data can be improved.
[0076] Reference Figure 6 In another embodiment of the present application, step S9 includes:
[0077] S91: Perform checksum comparison, metadata consistency check, random sampling comparison, business process verification, and / or log audit on the source node data of the background data migration and the preset storage node data;
[0078] S92: If the background data migration source node data is inconsistent with the preset storage node data, restore or re-migrate the data according to the storage system log.
[0079] In this embodiment, the source node data and the preset storage node data are checked and compared by writing data node verification, that is, the checksum of the data written by the source node is verified with the checksum of the preset storage node. In other embodiments of the present application, other existing data verification and comparison methods can be used, which will not be elaborated here.
[0080] In this embodiment, metadata consistency check is performed by first setting consistency rules, such as data integrity rules, and then using a comparison algorithm or comparison tool to compare the source node data with the data of the preset storage node to generate a difference report. In other embodiments of the present application, other existing methods can also be used.
[0081] In this embodiment, the source node data is randomly sampled and compared with the preset storage node data. The data range for comparison and the number of samples to be drawn can be determined first, and then systematic random sampling is performed. The sampled data are then compared to judge the consistency of the data. In other embodiments of the present application, other sampling and comparison methods can also be used.
[0082] In this embodiment, when verifying a business process, the business process is first simulated. Based on the actual business process, the data transmission and processing between the source node and the preset storage node are simulated. Secondly, the data integrity and accuracy are verified. During the simulation, the integrity and accuracy of the data are verified to ensure that they meet business requirements, and any data loss, corruption, or inconsistency is checked. Other business process verification methods may also be used in other embodiments of the present application.
[0083] In this embodiment, log auditing includes log collection: comprehensively collecting log information of source nodes and storage nodes, including operation logs, error logs, etc.; log parsing: parsing and filtering the collected log information to extract useful information; correlation analysis: performing correlation analysis on the parsed log information to identify potential security risks, abnormal operations, etc.
[0084] In different implementations of this embodiment, the verification process can be formed by combining one, two, or more of the following methods: checksum comparison, metadata consistency check, random sampling comparison, business process verification, and log auditing. These methods are not described in detail here. If the source node data is inconsistent with the preset storage node data, data recovery or re-migration can effectively improve the consistency of the system data.
[0085] Reference Figure 7 In another embodiment of the present application, step S8 includes:
[0086] S10: Send expansion completion information to the administrator.
[0087] In this embodiment, the administrator is notified of the completion of the expansion through email, text message, or system log, so as to conduct further inspection and confirmation, thereby improving the accuracy and stability of the display.
[0088] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0089] Reference Figure 8In one embodiment of the present application, a distributed storage system automatic expansion system is provided. The distributed storage system automatic expansion system corresponds one-to-one to a distributed storage system automatic expansion method in the above embodiment. The distributed storage system automatic expansion system includes:
[0090] The prediction module is used to collect system operation data and use machine learning algorithm models to predict future storage needs based on the system operation data;
[0091] A judgment module is used to configure a preset storage node when future storage demand exceeds the current storage capacity of the system;
[0092] Detection module, used to detect system storage pressure:
[0093] A trigger module is used to trigger the expansion mode when the system storage pressure exceeds a preset system storage threshold;
[0094] A state switching module is used to set a preset storage node to a hot standby state;
[0095] An updating module, configured to update a placement rule according to identification information of a preset storage node;
[0096] The data migration module is used to start the background data migration task and migrate the data to the preset storage node using the incremental migration strategy;
[0097] The state switching module is used to switch the preset storage node to a normal working state after the background data migration task is completed.
[0098] Each module of the distributed storage system automatic expansion system described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0099] In one embodiment of the present application, a computer device is provided, which may be a server. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device can be implemented by any type of volatile or non-volatile storage device, or a combination thereof. Volatile or non-volatile storage devices include, but are not limited to, magnetic disks, optical disks, EEPROM (Electrically Erasable Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), magnetic storage, flash memory, and PROM (Programmable Read-Only Memory). The memory of the computer device provides an environment for the operation of the operating system and computer programs stored therein. The network interface of the computer device is used to communicate with external terminals via a network connection. When executed by the processor, the computer program implements the steps of the distributed storage system automatic expansion method described in the above embodiment.
[0100] In one embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When executed by a processor, the computer program implements the steps of the method for automatically expanding the capacity of a distributed storage system described in the above embodiment. Computer-readable storage media include ROM (Read-Only Memory), RAM (Random-Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic disks, floppy disks, and the like.
[0101] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device of the present application can be divided into different functional units or modules to complete all or part of the functions described above.
Claims
1. A distributed storage system automatic expansion method, characterized in that: The method comprises: Collecting system operation data and using a machine learning algorithm model to predict future storage needs based on the system operation data; When the future storage demand exceeds the current storage capacity of the system, configuring a preset storage node; Detection system storage pressure: When the system storage pressure exceeds a preset system storage threshold, the expansion mode is triggered; Setting the preset storage node to a hot standby state; Updating the placement rule according to the identification information of the preset storage node; Start the background data migration task and migrate the data to the preset storage node using the incremental migration strategy; When the background data migration task is completed, the preset storage node is switched to a normal working state.
2. A distributed storage system automatic expansion method according to claim 1, characterized in that: The detection system storage pressure includes: Monitoring system storage-related parameters and / or access-related parameters; The system storage pressure is obtained according to the storage-related parameters and / or the access-related parameters.
3. The method for automatically expanding a distributed storage system according to claim 1, wherein: The collecting of system operation data and predicting future storage demand using a machine learning algorithm model based on the system operation data includes: Collecting system operation data, including storage system logs, performance monitoring data, and user behavior data; Performing data cleaning on the system operation data to obtain valid data; Extracting features from the valid data to form structured data; The structured data is input into the machine learning algorithm model to predict the future storage demand.
4. The method for automatically expanding a distributed storage system according to claim 1, wherein: The setting of the preset storage node to the hot standby state is: The system configuration and network configuration of the preset storage node are set, and the preset storage node is set to a hot standby state.
5. The method for automatically expanding the capacity of a distributed storage system according to claim 1, wherein: The background data migration tasks include: Filter the data to be migrated first based on the preset data priority migration policy: The priority migration data is migrated to the preset storage node using an incremental migration strategy.
6. A distributed storage system automatic expansion method according to claim 1, characterized in that: The method further comprises: Verify the accuracy and consistency of the data migrated from the preset storage node by the background data migration task.
7. A distributed storage system automatic expansion method according to claim 6, characterized in that: The verifying the accuracy and consistency of the data migrated from the preset storage node by the background data migration task includes: Perform verification and comparison, metadata consistency check, random sampling comparison, business process verification and / or log audit on the source node data of the background data migration and the preset storage node data; If the background data migration source node data is inconsistent with the preset storage node data, the data is restored or re-migrated according to the storage system log.
8. The method for automatically expanding the capacity of a distributed storage system according to claim 1, wherein: After the background data migration task is completed, the preset storage node is switched to a normal working state, including: Send expansion completion information to the administrator.
9. A distributed storage system automatic expansion system, characterized in that: The system comprises: A prediction module, configured to collect system operation data and predict future storage demand using a machine learning algorithm model based on the system operation data; A judgment module, configured to configure a preset storage node when the future storage demand exceeds the current storage capacity of the system; Detection module, used to detect system storage pressure: A trigger module, configured to trigger a capacity expansion mode when the system storage pressure exceeds a preset system storage threshold; A state switching module, configured to set the preset storage node to a hot standby state; An updating module, configured to update a placement rule according to identification information of the preset storage node; The data migration module is used to start the background data migration task and migrate the data to the preset storage node using an incremental migration strategy; The state switching module is used to switch the preset storage node to a normal working state after the background data migration task is completed.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for automatic capacity expansion of a distributed storage system according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Distributed automatic expansion method in high-performance computing
CN118838735A
Upgrade migration method, device and equipment and computer readable storage medium
CN119088409A
Dynamic adjustment method and system for storage resource pool
CN119847454A
Distributed storage method and system
WO2018000993A1