Method for determining cluster architecture and hardware configuration of distributed storage system and related equipment

By obtaining user needs and utilizing the hardware product information library, the cluster architecture and hardware configuration of the distributed storage system are automatically determined, which solves the problem of the inability to quickly determine the cluster architecture and hardware configuration in the existing technology and realizes the design of a high-performance and low-cost distributed storage system.

CN120687036APending Publication Date: 2025-09-23SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510809269.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies are unable to quickly determine the cluster architecture and hardware configuration of a distributed storage system based on user needs, resulting in high complexity for users when selecting a suitable distributed storage system.

Method used

By obtaining user needs and utilizing the hardware product information library, the cluster architecture and hardware configuration of the distributed storage system are automatically determined, including the number of core components, storage pool type, and hardware configuration parameters. The read and write performance and cost are calculated, and candidate solutions that meet the needs are screened.

Benefits of technology

It simplifies the complexity of users in selecting distributed storage system cluster architecture and hardware configuration, provides a fast and convenient solution to meet capacity, performance and cost requirements, and realizes high-performance and low-cost cluster architecture and hardware configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687036A_ABST
    Figure CN120687036A_ABST
Patent Text Reader

Abstract

The invention relates to a method for determining cluster architecture and hardware configuration of a distributed storage system and related equipment. The method comprises the following steps: acquiring a user demand of the distributed storage system; according to the storage capacity requirement or the storage access mode requirement, a plurality of storage schemes are determined by searching in a hardware product information base and selecting storage pool types and storage pool configuration parameters; determining the number of core components of the distributed storage system to obtain the cluster architecture under each storage scheme; the read-write performance of the cluster architecture under each storage scheme and the cost of hardware configuration in the storage scheme are calculated, and candidate storage schemes meeting the cost requirement and the read-write performance requirement and the corresponding cluster architecture are screened from the multiple storage schemes. Therefore, the cluster architecture and the hardware configuration of the distributed storage system meeting the user requirements can be automatically and quickly determined for the user based on the user requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of distributed storage, and in particular to a method and related devices for determining a cluster architecture and hardware configuration of a distributed storage system. Background Art

[0002] A distributed storage system stores data across multiple independent devices. This system utilizes a scalable architecture, sharing the storage load across multiple storage servers. This not only improves system reliability, availability, and access efficiency, but also facilitates scalability.

[0003] Storage can be divided into block storage, object storage, and file storage according to its type. Among various distributed storage systems, HDFS (Hadoop Distributed File System), GPFS (General Parallel File System), and GFS (Google File System) belong to file storage, Swift belongs to object storage, and Ceph, through its architecture and principles, can simultaneously support block storage, file storage, and object storage to meet the diverse needs of different application scenarios. These distributed storage systems each include multiple core components. For example, the core components of the Ceph system include Monitor (MON, cluster monitoring), OSD (Object Storage Device, object storage daemon), MDS (Metadata Server, metadata service), MGR (Manager, cluster management), etc. These core components work together to achieve high availability, strong consistency, and automatic management of distributed storage clusters.

[0004] Existing technologies for distributed storage system cluster architecture design and hardware configuration selection primarily offer various methods for estimating storage performance and general principles for hardware selection. These methods and principles typically focus solely on theoretical calculations of storage performance and performance evaluations of individual hardware components, failing to quickly determine the cluster architecture and hardware configuration for distributed storage systems based on user needs. Summary of the Invention

[0005] In view of this, the present disclosure proposes a method and related devices for determining the cluster architecture and hardware configuration of a distributed storage system, which can automatically and quickly determine the cluster architecture and hardware configuration of a distributed storage system that meets user needs based on user needs.

[0006] According to one aspect of the present disclosure, a method for determining a cluster architecture and hardware configuration of a distributed storage system is provided, wherein the cluster architecture includes at least the number of core components of the distributed storage system, the storage pool type, and storage pool configuration parameters, and the hardware configuration includes at least the type of hard disk, the capacity of a single hard disk, and the number of hard disks. The method includes: obtaining user requirements for a distributed storage system to be constructed, the user requirements including storage capacity requirements, cost requirements, and read / write performance requirements, and may also include storage access method requirements; determining multiple storage solutions based on the storage capacity requirements, or based on the storage access method requirements, by searching in a hardware product information library and selecting storage pool types and storage pool configuration parameters, the storage solutions including the hardware configuration, the storage pool type, and the storage pool configuration parameters; determining the number of core components of the distributed storage system for each of the multiple storage solutions to obtain the cluster architecture under each storage solution; calculating the read / write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution, and screening candidate storage solutions and corresponding cluster architectures that meet the cost requirements and the read / write performance requirements from the multiple storage solutions.

[0007] In a possible implementation, the method further includes: after obtaining the user needs, obtaining parameter information of multiple publicly available hardware products to construct or update the hardware product information library, wherein the multiple hardware products include a central processing unit, memory, hard disk, and network card.

[0008] In one possible implementation, the multiple storage solutions are determined based on the storage capacity requirement by searching in a hardware product information library and selecting storage pool types and storage pool configuration parameters, including: searching for parameter information of multiple hard disks in the hardware product information library based on the storage capacity requirement to determine the type of hard disk and the capacity of a corresponding single hard disk; selecting one or more storage pool types and corresponding storage pool configuration parameters from a plurality of preset storage pool types and corresponding storage pool configuration parameters, and calculating the total storage space required for each storage pool type and corresponding storage pool configuration parameters based on the storage capacity requirement; determining the number of hard disks for each storage pool type and corresponding storage pool configuration parameters based on the total storage space and the capacity of a single hard disk; and determining different combinations of the determined hard disk type and the capacity of a corresponding single hard disk, storage pool type and corresponding storage pool configuration parameters, and the number of hard disks as the multiple storage solutions.

[0009] In one possible implementation, the hardware configuration also includes the number of central processing unit (CPU) cores and the memory size, and the method also includes: determining the number of CPU cores based on the number of at least one core component of the distributed storage system; determining the memory size based on the total storage space and the number of at least one core component; and adding the number of CPU cores and memory size to the corresponding storage solution.

[0010] In one possible implementation, when determining multiple storage solutions by searching in a hardware product information library and selecting storage pool types and storage pool configuration parameters, the network structure of the distributed storage system is also selected to determine that the cluster architecture also includes the multiple storage solutions of the network structure.

[0011] In one possible implementation, the calculation of the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution includes: calculating the read and write performance based on the hardware configuration and the cluster architecture in each storage solution; determining parameter information of the network card based on the read and write performance; selecting the model and quantity of the network card for the cluster architecture by searching in the hardware product information library; and calculating the cost of the hardware configuration including the model and quantity of the network card.

[0012] In one possible implementation, after screening multiple candidate storage solutions that meet the cost requirements and the read-write performance requirements from the multiple storage solutions, the method further includes: sorting the multiple candidate storage solutions according to the read-write performance of the cluster architecture under each of the multiple candidate storage solutions and the cost of the hardware configuration in each storage solution; selecting one or more optimal candidate storage solutions from the sorted multiple candidate storage solutions, and providing the selected candidate storage solutions and the corresponding cluster architecture to the user.

[0013] According to another aspect of the present disclosure, a device for determining the cluster architecture and hardware configuration of a distributed storage system is provided, wherein the cluster architecture includes at least the number of core components of the distributed storage system, the storage pool type, and storage pool configuration parameters, and the hardware configuration includes at least the type of hard disk, the capacity of a single hard disk, and the number of hard disks. The device includes: a user demand acquisition module configured to acquire user demand for the distributed storage system to be built, wherein the user demand includes storage capacity demand, cost demand, and read / write performance demand, and can also include storage access method demand; a storage solution determination module configured to determine the hardware configuration according to the storage capacity demand, or can also determine the storage access method demand, by determining the hardware configuration according to the hardware configuration. A search is performed in a product information library, and multiple storage solutions are determined by selecting storage pool types and storage pool configuration parameters, where the storage solutions include the hardware configuration, the storage pool type, and the storage pool configuration parameters; a cluster architecture determination module is configured to determine the number of core components of the distributed storage system for each storage solution in the multiple storage solutions, so as to obtain the cluster architecture under each storage solution; a candidate solution screening module is configured to calculate the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution, and screen candidate storage solutions and corresponding cluster architectures that meet the cost requirements and the read and write performance requirements from the multiple storage solutions.

[0014] According to another aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0015] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0016] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program implements the steps of the above method when executed by a processor.

[0017] According to various aspects of the present disclosure, based on user requirements such as storage capacity, cost requirements, read-write performance requirements, and storage access method requirements, combined with a hardware product information library, a cluster architecture and hardware configuration of a distributed storage system that meets user requirements can be automatically generated. Moreover, the generated cluster architecture and hardware configuration are selected based on a comprehensive consideration of read-write performance and cost, which can simplify the complexity of users in selecting suitable hardware for the cluster, and provide users with a quick and convenient recommendation of the cluster architecture and hardware configuration of the distributed storage system, especially a cluster architecture and hardware configuration of a distributed storage system that can meet capacity requirements while having high performance and low cost.

[0018] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0020] Figure 1 A flowchart of a method for determining a cluster architecture and hardware configuration of a distributed storage system according to an embodiment of the present disclosure is shown.

[0021] Figure 2 A block diagram illustrating an apparatus for determining a cluster architecture and hardware configuration of a distributed storage system according to an embodiment of the present disclosure is shown.

[0022] Figure 3 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0023] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0024] As used herein, the terms "comprises," "comprising," "having," or variations thereof are open ended and include one or more stated features, integers, elements, steps, parts, or functions, but do not preclude the presence or addition of one or more other features, integers, elements, steps, parts, functions, or groups thereof.

[0025] When an element is referred to as being "connected," "coupled," "responsive" or variations thereof to another element, it can be directly connected, coupled or responsive to the other element or intervening elements may be present.

[0026] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another element / operation. Therefore, without departing from the teachings of the present invention, the first element / operation in some embodiments may be referred to as the second element / operation in other embodiments.

[0027] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0028] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0029] As mentioned above, the existing technology cannot quickly determine the cluster architecture and hardware configuration of the distributed storage system according to user needs. However, in response to the storage access method, read and write performance and capacity requirements proposed by users, it is very necessary to quickly provide a cluster architecture and hardware configuration of a distributed storage system that meets user needs while being cost-effective. Therefore, the embodiment of the present disclosure proposes a method for determining the cluster architecture and hardware configuration of a distributed storage system, which can simplify the complexity of users in selecting a suitable cluster architecture and hardware configuration in a distributed storage system (such as a Ceph system), and design the cluster architecture and hardware configuration of the distributed storage system in an automated manner according to the read and write performance, capacity and storage access method requirements of the user, evaluate and calculate the read and write performance and cost under the cluster architecture and hardware configuration in the distributed storage system, thereby providing a hardware configuration and cluster architecture that meets the read and write performance requirements, cost requirements and capacity requirements, helping users to quickly build a distributed storage cluster and put it into production while achieving cost optimization.

[0030] The method of the embodiment of the present disclosure can be deployed on various terminal devices through software or hardware modification. The terminal device involved in the embodiment of the present disclosure may refer to a device with a wireless connection function and / or a wired connection function. The wireless connection function refers to the ability to connect to other devices through wireless connection methods such as wifi and Bluetooth. The terminal device involved in the embodiment of the present disclosure may also communicate with other devices through a wired connection function. The terminal device involved in the embodiment of the present disclosure may be a touch screen, a non-touch screen, or a screenless device. The touch screen device can be controlled by clicking, sliding, etc. on the display screen with a finger or a stylus. The non-touch screen device can be connected to an input device such as a mouse, keyboard, touch panel, etc., and the terminal device can be controlled through the input device. For example, a device without a screen can be a Bluetooth speaker without a screen. For example, the terminal device of the present application may include but is not limited to a user equipment (UE), a mobile device, a user terminal, a terminal, a handheld device, a tablet computer, a laptop computer, a handheld computer, a computing device, etc.

[0031] The method of the embodiment of the present disclosure can also be deployed on a server, which can be located in the cloud or locally, and can be a physical device or a virtual device, such as a virtual machine, a container, etc., with a wireless communication function, wherein the wireless communication function can be set in the chip (system) or other parts or components of the server. It can refer to a device with a wireless connection function, and the wireless connection function means that it can be connected to other servers or terminal devices through wireless connection methods such as Wi-Fi and Bluetooth. The server involved in the embodiment of the present disclosure can also have the function of communicating through a wired connection. For example, the server of the embodiment of the present disclosure can be located in the cloud, communicate with the terminal device, receive user requirements sent by the terminal device, and use the method deployed on the server to determine the cluster architecture and hardware configuration of the distributed storage system that meets the user's requirements, and return it to the terminal device to display the determined cluster architecture and hardware configuration of the distributed storage system to the user in the terminal device.

[0032] In the embodiments of the present disclosure, the cluster architecture includes at least the number of core components of the distributed storage system (such as the core components of the Ceph system include MON, OSD, MDS, MGR, etc.), the storage pool type (such as full copy type or erasure code type) and storage pool configuration parameters (such as the number of copies under the full copy type, the number of data blocks k and check blocks m under the erasure code type); the hardware configuration includes at least the type of hard disk (such as solid state drive (SSD), mechanical hard disk (HDD)), the capacity of a single hard disk, and the number of hard disks.

[0033] Figure 1FIG. 1 is a flow chart showing a method for determining a cluster architecture and hardware configuration of a distributed storage system according to an embodiment of the present disclosure. Figure 1 As shown, the method includes: steps S11 to S14.

[0034] In step S11, user requirements of the distributed storage system to be built are obtained, wherein the user requirements include storage capacity requirements, cost requirements, and read / write performance requirements, and may also include storage access method requirements.

[0035] The storage capacity requirement represents the total storage capacity required by the user for the entire cluster. The read / write performance requirement includes the read / write bandwidth requirement and / or the input / output operations per second (IOPS) requirement. The read / write bandwidth requirement represents the total read / write bandwidth required by the user for all hard disks in the cluster, and the IOPS requirement represents the total IOPS required by all hard disks in the cluster. The cost requirement represents the acceptable cost of setting up the entire cluster.

[0036] It should be understood that storage capacity requirements, read / write performance requirements, and cost requirements can be ranges. For example, the storage capacity requirement can be a user-given storage capacity range (e.g., [1PB, 2PB]), the read / write performance requirement can be a user-given read / write performance range, which can include a read / write bandwidth range (e.g., [10GB / s, 20GB / s]) and / or an IPOS range (e.g., [500,000, 600,000]), and the cost requirement can also be a user-given cost range (e.g., [2 million yuan, 3 million yuan]). In this case, the determined cluster architecture and hardware configuration can be a cluster architecture and hardware configuration that meets the aforementioned storage capacity range, read / write performance range, and cost range. Of course, the user can also provide one or more specific storage capacity values ​​(such as 1PB) as storage capacity requirements, one or more read and write performance values ​​(i.e., one or more read and write bandwidths, such as 10GB / s; one or more IPOS, such as 500,000) as read and write performance requirements, and one or more cost values ​​(such as 2 million yuan) as cost requirements. The embodiments of the present disclosure do not limit this.

[0037] It can be known that a distributed storage system such as Ceph can support three storage access modes: block storage, file storage, and object storage. Different storage access modes have different requirements for hardware configuration. Therefore, users can also give storage access mode requirements for the distributed storage system to be built according to actual needs. Of course, users can also not give storage access mode requirements. Accordingly, a corresponding cluster architecture and hardware configuration can be generated for each storage access mode supported by the distributed storage system. It should be understood that if the distributed storage system to be built does not support multiple storage access modes, for example, GFS only supports file storage, MinIO and Swift only support object storage, etc., then the storage access mode requirements can be defaulted to the storage access mode supported by the distributed storage system to be built, and the embodiments of the present disclosure do not limit this.

[0038] It should be understood that the embodiments of the present disclosure do not limit the method of obtaining user needs. For example, a questionnaire survey or direct communication with users can be used to determine user needs, and the embodiments of the present disclosure do not limit this.

[0039] In step S12, multiple storage solutions are determined by searching in a hardware product information library and selecting storage pool types and storage pool configuration parameters based on storage capacity requirements, or possibly also based on storage access method requirements.

[0040] The hardware product information library includes product information of various hardware products, including model, price, and parameter information. Various hardware products include central processing units (CPUs), memory, hard disks, and network cards. The parameter information of hardware products includes performance parameters used to describe the hardware products. For example, the parameter information of the CPU may include the number of CPU cores, the parameter information of the memory may include the memory capacity (i.e., memory size), the parameter information of the hard disk may include the hard disk capacity, read / write bandwidth, and read / write IPOS, and the parameter information of the network card may include the network bandwidth of the network card.

[0041] In practical applications, product information of mainstream hardware products on the market can be regularly collected, including models, prices, and parameter information of central processing units, memory, hard drives, network cards, etc., to build or update the hardware product information library. Alternatively, after obtaining user requirements, parameter information of various publicly available hardware products can be obtained to build or update the hardware product information library. In this way, the hardware products and product information in the hardware product information library can be kept up-to-date and comprehensive. The disclosed embodiments do not restrict the method for obtaining product information for each hardware product in the hardware product information library.

[0042] In one possible implementation, step S12, based on storage capacity requirements, determines multiple storage solutions by searching in a hardware product information library and selecting a storage pool type and storage pool configuration parameters. The determination may include:

[0043] Step S121 , searching parameter information of multiple hard disks in a hardware product information library according to storage capacity requirements to determine the hard disk type and the corresponding capacity of a single hard disk;

[0044] Step S122: Select one or more storage pool types and corresponding storage pool configuration parameters from a plurality of preset storage pool types and corresponding storage pool configuration parameters, and calculate the total storage space required for each storage pool type and corresponding storage pool configuration parameters based on the storage capacity requirement;

[0045] Step S123, for each storage pool type and corresponding storage pool configuration parameters, determine the number of hard disks according to the total storage space and the capacity of a single hard disk;

[0046] Step S124 : Determine different combinations of the determined hard disk type and corresponding single hard disk capacity, storage pool type and corresponding storage pool configuration parameters, and the number of hard disks as multiple storage solutions.

[0047] In step S121, multiple hard drive types with capacities similar to the storage capacity requirements can be searched from the hardware product information library, and the capacities of individual hard drives of different types can be obtained. It should be understood that different types of hard drives can have multiple models, and multiple models of hard drives of each type can be searched, thereby obtaining the capacities of individual hard drives of each model of the different types. Of course, all or some hard drives in the hardware product information library can also be directly selected to generate a storage solution, and this is not limited in the present embodiment.

[0048] In step S122, the storage pool type includes an erasure code type and a full copy type. The storage pool type can also be called a redundancy mechanism. The full copy type means that each data to be stored is completely copied into multiple copies (such as three copies) and stored on different nodes or hard disks. Each copy of data is called a copy. In the case where the storage pool type is a full copy type, a storage capacity requirement of 1PB actually requires a total storage space of 1PB×the number of copies. For example, if the number of copies is 3, a storage capacity requirement of 1PB actually requires a total storage space of 1PB×3=3PB. The number of copies can be described as a storage pool configuration parameter of the full copy type. The erasure code type refers to dividing the data to be stored into multiple data blocks and check blocks by encoding, providing redundancy with lower storage overhead. Specifically, it can be set to divide the data to be stored into k data blocks and generate m check blocks, with a total number of blocks being k+m. The erasure code scheme k+m can be described as a storage pool configuration parameter of the erasure code type. Currently, commonly used storage pool configuration parameters for erasure coding include, for example, 4+2 (i.e., 4 data blocks + 2 parity blocks) and 6+2 (i.e., 4 data blocks + 2 parity blocks). Under erasure coding, a 1PB storage capacity requirement actually requires a total storage space of 1PB×(k+m) / k. For example, if k+m is 6+2, a 1PB storage capacity requirement actually requires a total storage space of 1PB×(8 / 6)=1.33PB.

[0049] It should be understood that different types of hard drives can have different models, and different models of hard drives of the same type can have different capacities, different read / write bandwidths, and different IOPS. In step S123, assume that the capacity of a single hard drive of a certain type and model is 10TB, and the storage capacity requirement is 1PB. In this case, if the storage pool type is full replica and the number of replicas is 3, the total storage space actually required can be calculated as 1PB × 3 = 3PB. Since the capacity of a single hardware device is 10TB, the number of hard drives required is 3PB / 10TB = 3 × 1024 / 10 = 307.2, meaning that at least 308 hard drives are required. In actual use, not every hard drive is fully populated, meaning that hard drives have a certain space utilization ratio. For example, a space utilization ratio of 70% or 80% means that 70% or 80% of the hard drive capacity is actually used for storage. Therefore, the number of hard drives can be further calculated based on the hard drive space utilization ratio. For example, if the hardware space utilization rate is 70%, the number of hard drives is 3PB / (10×0.7)TB=3×1024 / (10×0.7)≈438.9, which means that at least 439 hard drives are required. When the storage pool type is erasure coded, if the erasure code scheme k+m is 6+2, then the storage capacity requirement of 1PB actually requires a total storage space of 1PB×(8 / 6)=1.33PB. Since the capacity of a single hard drive is 10TB, the number of hard drives is 1.33PB / 10TB=1.33×1024 / 10=136.2, which means that at least 137 hard drives are required. Similarly, the number of hard drives for erasure coding can be further calculated based on the hard drive space utilization rate. For example, if the hard disk space utilization rate is 70%, the number of hard disks is 1.33PB / (10×0.7)TB=1.33×1024 / (10×0.7)≈194.6, that is, at least 195 hard disks are required.

[0050] It should be understood that in step S124, based on different hard disk types, different hard disk models, different storage pool types, and different storage pool configuration parameters of different storage pool types (different numbers of copies of the full copy type, different erasure code schemes k+m under the erasure code type), a variety of storage solutions that meet storage capacity requirements can be obtained by traversing and permuting the selected hard disk types, hard disk models, storage pool types, and storage pool configuration parameters. Each storage solution may include hardware configuration, storage pool type, and corresponding storage pool configuration parameters, wherein the hardware configuration of each storage solution may include hard disk type, hard disk model, capacity of a single hard disk, and number of hard disks. At least one of the hard disk type, hard disk model, capacity of a single hard disk, number of hard disks, storage pool type, and corresponding storage pool configuration parameters may be different in different storage solutions.

[0051] In practical applications, after calculating the number of hard drives in each storage solution, the number of servers required for each storage solution can be calculated based on the maximum number of hard drives that can be installed in a single server. It is understandable that the number of hard drives that can be installed in a single server is limited, and the maximum number of hard drives that can be installed varies between different types of servers. Therefore, the number of servers required for each storage solution can be calculated based on the maximum number of hard drives that can be installed in various types of servers on the market. For example, the maximum number of hard drives that can be installed in a single server can be 16, 13, 12, 10, etc. If a storage solution has 439 hard drives, and the maximum number of hard drives that can be installed in a single server is 10, at least 44 servers are required. If the maximum number of hard drives that can be installed in a single server is 16, at least 439 / 16≈27.4≈28 servers are required. It should be understood that for each storage solution, different server numbers can be calculated based on the maximum number of hard drives that can be installed in a single server, thereby obtaining a wider range of storage solutions. Therefore, the hardware configuration of each storage solution can also include the number of servers, and the number of servers can vary between different storage solutions. Of course, users can also specify the number of servers in their requirements. If this is the case, the hardware configuration for each storage solution can directly include this number of servers, along with the number of hard drives required for each server. For example, if the user requirements specify 10 servers and a storage solution specifies 137 hard drives, then each server should have 14 hard drives.

[0052] In step S13 , for each storage solution among the multiple storage solutions, the number of core components of the distributed storage system is determined to obtain a cluster architecture under each storage solution.

[0053] It should be understood that different core components are different in different distributed storage systems, and different core components are required for different storage access methods. For example, when the distributed storage system is a Ceph system, the core components include OSD, MON, MGR, MDS, etc. If the storage access method of block storage or object storage is adopted, the core components required in the Ceph system are OSD, MON, MGR, and if the storage access method of file storage is adopted, the core components required in the Ceph system are OSD, MON, MGR, and MDS. For another example, when the distributed storage system is a Lustre system, the core components include MDS and OSS (Object Storage Service, object storage service), etc. Such core components determine that the Lustre system only supports the storage access method of file storage.

[0054] In actual applications, the number of core components in the distributed storage system under each storage solution can be designed based on the number of hard disks and the number of servers in each storage solution determined in step S12. For example, for the Ceph system, one hard disk usually corresponds to 1 to 3 OSDs, and the entire system usually requires 1 to 3 MONs, MGRs, and MDSs. Assuming that one hard disk corresponds to 3 OSDs, the number of hard disks in a certain storage solution is 130, and the number of servers is 10, then the number of hard disks installed in a single server is 13, then each server requires 3×13=39 OSDs, and a total of 39×10=390 OSDs are required. Assuming that the storage access method specified in the user requirements is file storage, the number of MONs under the storage solution can be selected as 3, and the number of MGRs and MDSs can be selected as 2, thus obtaining the cluster architecture under the storage solution.

[0055] Considering that the distributed storage system has certain requirements for the central processing unit (CPU) and memory in addition to the hard disk, in one possible implementation, the hardware configuration of the distributed storage system may further include the number of CPU cores and the memory size. The method further includes:

[0056] Determine the number of CPU cores based on the number of at least one core component of the distributed storage system;

[0057] Determine the memory size based on the total storage space and the number of at least one core component;

[0058] Add the number of CPU cores and memory size to the corresponding storage plan.

[0059] It should be understood that different CPU models may have different numbers of cores, different memory models may also have different sizes, different core components in a distributed storage system may have different requirements for CPU and memory, and different storage access methods may also require different core components. Therefore, after determining the number of core components in the distributed storage system under each storage solution based on the number of hard drives and the number of servers in each storage solution, the CPU model and the number of CPU cores (i.e., the number of cores), as well as the memory model and memory size, can be designed based on the number of core components in the distributed storage system under each storage solution.

[0060] For example, in a Ceph system, a single hard drive typically supports 1 to 3 OSDs. Each OSD requires approximately 1 to 4 CPU cores and at least 2GB of memory. Other core components, such as the MON, MGR, and MDS, each require approximately 2 to 3 CPU cores and 10 to 15GB of memory. Assuming a single hard drive supports 3 OSDs, a storage solution with 130 hard drives and 10 servers, the number of hard drives installed in a single server is 13, and each server requires 3 × 13 = 39 OSDs. In a file-based storage access mode, assuming each OSD requires 4 CPU cores, the total number of CPU cores required for the OSDs in a single server is 39 × 4 = 156. If a single server is deployed with 1 MON, 1 MGR, and 1 MDS, and the minimum number of cores required for each MON, MGR, and MDS is 2, the total number of CPU cores required for the single server is 156 + 2 + 2 + 2 = 162. Assuming the capacity of a single hard drive in the storage solution is 15TB, and based on the calculation that each 1TB of hard drive capacity corresponds to 1GB of memory and each OSD requires at least 2GB of memory, we can determine that the memory required for a single OSD is 15GB. Assuming that the memory required by the MON, MGR, and MDS is also 15GB, the total memory required for a single server is 15×39+15+15+15=630GB. Considering that memory also has a certain space utilization ratio in actual applications, in order to meet the 630GB memory requirement, in practice, more memory than 630GB is required. For example, if the memory space utilization ratio is 80%, the memory required for a single server is (15×39+15+15+15) / 0.8=787.5GB. When using block or object storage, the calculations can be performed by removing the minimum number of cores and memory required by the MDS. That is, the total number of processor cores required for a single server is 156 + 2 + 2 = 160 cores, and the total memory required for a single server is 15 × 39 + 15 + 15 = 615 GB. Similarly, considering memory space utilization, the required memory size for a single server is (15 × 39 + 15 + 15) / 0.8 = 768.8 GB. Therefore, the hardware configuration for each storage solution can also include the number of CPU cores and memory required for a single server.

[0061] It is understandable that different CPU models have different numbers of cores, and different memory models also have different sizes. If the number of CPU cores and memory size required by a single server in a certain storage solution are known, the CPU model and memory model that meet the number of CPU cores and memory size required by a single server in the storage solution can be searched from the hardware product information library. Among them, for the CPU, it can be searched that the number of cores of a single CPU of a certain model is greater than or equal to the number of CPU cores indicated in the storage solution, or it can be searched that the number of cores of two or more CPUs of a certain model combined is greater than or equal to the number of CPU cores indicated in the storage solution, and the determined CPU model can also be added to the hardware configuration of the corresponding storage solution. For the memory, it can be searched that the single memory of a certain model is greater than or equal to the memory size required by a single server indicated in the storage solution, or it can be searched that the memory size of two or more memories of a certain model combined is greater than or equal to the memory size required by a single server indicated in the storage solution, and the determined memory model can be added to the hardware configuration of the corresponding storage solution. For each storage solution, you can search the hardware product information above to find multiple CPU models and memory models that meet the number of CPU cores and memory size required for a single server in the hardware configuration of each storage solution. Thus, by combining CPU models and memory models, you can get more storage solutions. The hardware configuration of each storage solution can also include CPU models and memory models.

[0062] It should be understood that the number of core components deployed by a single hard disk can be different, the requirements for the number of CPU cores and memory size of different core components in the same distributed storage system can be different, and the core components required for different storage access methods are also different. By arranging and combining these variables, the number of multiple core components can be determined based on the number of hard disks in a single server in each storage solution, and then the combination of multiple CPU cores and memory sizes can be determined. In addition, since the number of cores of different CPU models is different, the size of memory of different models is also different, so by arranging and combining the above variables, more storage solutions can be obtained. At this time, at least one of the number of CPU cores, memory size, CPU model and memory model in different storage solutions is different.

[0063] It can be seen that the network architecture of a distributed storage system generally includes a front-end and back-end network separation architecture and a front-end and back-end network integration architecture. A front-end and back-end network separation architecture separates the front-end and back-end networks. In this architecture, each server may require two network cards. A back-end network integration architecture integrates the front-end and back-end networks. In this architecture, each server may require one network card. The front-end network is used for client access to the system, while the back-end network is used for communication between OSDs on different servers within the system. A network separation architecture physically isolates the front-end and back-end networks, preventing mutual interference. Therefore, when selecting multiple storage solutions by searching the hardware product information library and selecting the storage pool type and storage pool configuration parameters, the cluster architecture also includes multiple storage solutions for network architecture by selecting the network architecture of the distributed storage system. In other words, the cluster architecture may also include whether the network architecture is a front-end and back-end network separation architecture or an integration architecture. Furthermore, the number of network cards required can be determined based on the type of network architecture used in the cluster architecture. In this case, the hardware configuration for the corresponding storage solution may also include the number of network cards. For example, if a distributed storage system uses a front-end and back-end network separation architecture, each server uses two network cards. If a storage solution has 10 servers, a total of 20 network cards are required. If a distributed storage system uses a front-end and back-end network separation architecture, each server uses one network card. If a storage solution has 10 servers, a total of 10 network cards are required. The choice of network card model depends on the read and write bandwidth generated by the hard drives in the cluster.

[0064] In step S14, the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution are calculated, and candidate storage solutions and corresponding cluster architectures that meet the cost requirements and read and write performance requirements are screened from multiple storage solutions.

[0065] Calculating the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution may include:

[0066] Calculate read and write performance based on the hardware configuration and cluster architecture of each storage solution;

[0067] Determine the network card parameter information based on read and write performance;

[0068] Select the model and quantity of network cards used for the cluster architecture by searching in the hardware product information library;

[0069] Calculate the cost of the hardware configuration including the model and number of network cards.

[0070] As mentioned above, different types of hard drives can have different models, and different models of hard drives of the same type can have different capacities, different read / write bandwidths, and different read / write IOPS. The read / write performance of the cluster architecture under each storage solution can include the total read / write bandwidth and total read / write IOPS of all hard drives in the cluster under each storage solution. Therefore, for the hardware configuration of any storage solution, given the number of servers in the storage solution's hardware configuration, the number of hard drives in a single server, and the read / write bandwidth and read / write IOPS of the hard drives indicated by the hard drive models in the storage solution, the total read / write bandwidth (cluster read / write bandwidth) and total read / write IOPS (cluster read / write IOPS) generated by all hard drives in all servers under the storage solution can be calculated. For example, suppose a storage solution has 14 servers in its hardware configuration, with 10 hard drives on each server. The hard drive model in the hardware configuration of the storage solution indicates that the read and write bandwidth of a single hard drive is 250 MB / s and the read and write IOPS is 1000. In addition, the storage pool type in the cluster architecture of the storage solution uses an erasure code type and a 4+2 erasure code scheme (i.e., storage pool configuration parameters). Then, under this erasure code scheme, (4+2) / 4=1.5 times the hard drive redundancy will be generated. The calculation process of the cluster read and write bandwidth corresponding to this storage solution can be: 14×10×250 / 1.5≈22.8 GB / s, and the calculation process of the cluster read and write IOPS corresponding to this storage solution can be: 14×10×100 / 1.5≈9333.3. If the storage pool type in the cluster architecture of this storage solution uses a full-replication model with 3 replicas, this full-replication model generates 3x disk redundancy. The corresponding cluster read / write bandwidth for this storage solution can be calculated as: 14×10×250 / 3≈11.4 GB / s, and the corresponding cluster read / write IOPS can be calculated as: 14×10×100 / 3≈4666.7. Considering that hard drives experience some read / write performance degradation in real applications, meaning that the actual read / write performance of a hard drive can be lower than the performance indicated by its parameters, the cluster read / write bandwidth and cluster read / write IOPS can be calculated based on the hard drive's read / write performance degradation factor. For example, if the hard disk's read and write performance degradation coefficient is 0.7, the calculation process for the cluster read and write bandwidth corresponding to the storage solution under the erasure code type can also be: 14×10×250×0.7 / 1.5≈16.3GB / s. The calculation process for the cluster read and write IOPS corresponding to this storage solution can also be: 14×10×100×0.7 / 1.5≈6333.3.

[0071] It should be understood that for each of the multiple storage solutions, the read and write performance corresponding to each storage solution can be obtained by referring to the calculation process of the above-mentioned total read and write bandwidth and total read and write IOPS. Furthermore, the parameter information of the network card (i.e., the network bandwidth of the network card) can be determined to be greater than or equal to the total read and write bandwidth in the read and write performance, and then the network card model that meets the determined parameter information of the network card can be searched from the hardware product information library as the model of the network card used for the cluster architecture. At the same time, as mentioned above, based on whether the network structure adopted in the cluster architecture is a front-end and back-end network separation structure or a front-end and back-end network non-separation structure, the number of network cards used for the cluster architecture can be obtained.

[0072] As described above, the hardware product information library can include the price of each hardware product. Therefore, based on the hardware product information library, the price of each hardware component in each storage solution's hardware configuration can be determined. Combined with the quantity of each hardware component in each storage solution's hardware configuration, the hardware cost of each storage solution's hardware configuration can be calculated. Specifically, the cost of the hardware configuration, including the model and number of network cards, can be calculated. For example, if a storage solution's hardware configuration includes 10 servers, the price of the CPU, memory, and network card for each server is 69,000 yuan, the price of a single hard drive is 11,700 yuan, and each server has 12 hard drives installed, then the total cost for the 10 servers is: 69 + 1.17 × 12 × 10 = 2,094,000 yuan. This 2,094,000 yuan can be used as the hardware cost of the storage solution's hardware configuration, and this hardware cost can be directly used as the cost of the storage solution's hardware configuration.

[0073] Considering that in addition to hardware costs, distributed storage systems also generate energy consumption during operation, the cost of the hardware configuration of each storage solution can also be calculated in combination with the energy consumption costs that may be generated by the cluster. Therefore, in one possible implementation, the above calculation of the cost of the hardware configuration including the model and number of network cards can include: determining the hardware cost corresponding to the hardware configuration of each storage solution based on the price of each hardware in the hardware configuration of each storage solution; determining the energy consumption cost corresponding to the hardware configuration of each storage solution based on the unit energy consumption of each hardware in the hardware configuration of each storage solution; and determining the sum of the hardware cost and energy consumption cost corresponding to the hardware configuration of each storage solution as the cost of the hardware configuration of each storage solution.

[0074] In practical applications, the unit energy consumption of each hardware component represents the energy consumption (or power consumption) of the hardware per unit time. Given the unit energy consumption parameters of each hardware component in any storage solution's hardware configuration, combined with the number of components in the storage solution's hardware configuration and the unit energy consumption cost (e.g., the electricity cost per kilowatt-hour), the energy consumption cost of the storage solution's hardware configuration over its service life can be calculated. This energy consumption cost can be used as the energy cost corresponding to the storage solution's hardware configuration. For example, assume that a single server in a storage solution's hardware configuration, including its processor, memory, network interface card, and hard drive, consumes 1 kilowatt-hour of electricity per unit time (e.g., 1 hour), and the electricity cost per kilowatt-hour is 10 yuan. Then, the 10 servers in the storage solution's hardware configuration can incur a total electricity cost of 100 yuan per unit time. Assuming that each server has a service life of 1 year, the energy consumption cost of these 10 servers in a year is 1 × 365 × 24 × 100 = 876,000 yuan. This 876,000 yuan can be used as the energy cost corresponding to the storage solution's hardware configuration. Furthermore, the sum of the energy consumption cost and the hardware cost corresponding to the hardware configuration of the storage solution may be used as the cost of the hardware configuration of the storage solution.

[0075] As described above, the read and write performance requirements include read and write bandwidth requirements and / or IOPS requirements, wherein the read and write bandwidth requirement represents the total read and write bandwidth of all hard disks in the entire cluster required by the user, and the IOPS requirement represents the total IOPS of all hard disks in the entire cluster required by the user. In addition, the cost requirement represents the cost incurred by building the entire cluster that is acceptable to the user. Therefore, after calculating the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution, candidate storage solutions and corresponding cluster architectures that meet the cost requirements and read and write performance requirements can be screened from a variety of storage solutions. Among them, if the read and write performance requirement only indicates the read and write bandwidth requirement, then meeting the read and write performance requirement is to meet the read and write bandwidth requirement; if the read and write performance requirement only indicates the read and write IOPS requirement, then meeting the read and write performance requirement is to meet the read and write IOPS requirement; if the read and write performance requirement indicates both the read and write bandwidth requirement and the read and write IOPS requirement, then meeting the read and write performance requirement is to meet both the read and write bandwidth requirement and the read and write IOPS requirement.

[0076] As described above, the read / write performance requirement can be a read / write performance range, so storage solutions with read / write performance within the read / write performance range can be screened to determine candidate storage solutions. For example, storage solutions with read / write bandwidth within the read / write bandwidth range [10GB / s, 20GB / s] and read / write IOPS within the IPOS range [500,000, 600,000] can be screened to determine candidate storage solutions. Alternatively, the read / write performance requirement can be a read / write performance value. In this case, storage solutions with a specified floating range above or below the read / write performance value can be screened based on the read / write performance value to determine candidate storage solutions. For example, if the read / write bandwidth requirement is 10GB / s and the read / write IPOS requirement is 5000, storage solutions with a read / write bandwidth within the range of 10±1GB / s (i.e., [9GB / s, 11GB / s]) and a read / write IOPS within the range of 5000±200 (i.e., [4800, 5200]) can be screened to determine candidate storage solutions. Alternatively, when the read / write performance requirement is a read / write performance value, storage solutions with read / write performance greater than or equal to the read / write performance value may be directly screened out to determine candidate storage solutions. For example, if the read / write bandwidth requirement is 10 GB / s and the read / write IPOS requirement is 5000, then candidate storage solutions may be directly determined based on the read / write bandwidth being greater than or equal to 10 GB / s and the read / write IPOS being greater than or equal to 5000. This is not a limitation of the embodiments of the present disclosure.

[0077] As described above, the cost requirement can be a cost interval, so storage solutions with costs within the cost interval can be screened out to determine candidate storage solutions. For example, storage solutions with costs within [2 million, 3 million] can be screened out to determine candidate storage solutions. Alternatively, the cost requirement can also be a cost value, in which case storage solutions within a specified floating amount above and below the cost value can be screened out based on the cost value to determine candidate storage solutions. For example, if the cost requirement is 2 million, storage solutions with costs within the range of 2 million ± 100,000 yuan (that is, [1.9 million yuan, 2.1 million yuan]) can be screened out to determine candidate storage solutions. Alternatively, when the cost requirement is a cost value, storage solutions with costs less than or equal to the cost value can also be directly screened out to determine candidate storage solutions. For example, if the cost requirement is 2 million yuan, storage solutions with costs less than or equal to 2 million yuan can be directly screened out to determine candidate storage solutions, and this embodiment of the present disclosure does not limit this.

[0078] It should be understood that based on the above storage solution screening method that meets read / write performance requirements and cost requirements, candidate storage solutions that meet both cost requirements and read / write performance requirements can be screened out, and the corresponding cluster architecture can be obtained.

[0079] According to the method of the embodiment of the present disclosure, it is possible to utilize the hardware product information library in combination with user needs to automatically generate the hardware configuration and cluster architecture of the distributed storage system that meets the user's needs. Moreover, the generated cluster architecture and hardware configuration are selected based on the comprehensive read and write performance and cost, which can simplify the complexity of users in selecting suitable hardware for the cluster, and provide users with a quick and convenient recommendation of the cluster architecture and hardware configuration of the distributed storage system, especially it can generate the cluster architecture and hardware configuration of the distributed storage system that meets capacity requirements and has high performance and low cost.

[0080] In actual applications, after screening out multiple candidate storage solutions that meet both read / write performance requirements and cost requirements, one or more optimal candidate storage solutions can be further selected from the multiple candidate storage solutions and provided to the user. In other words, cost and read / write performance can be further combined to select low-cost and high-performance candidate storage solutions and provide them to the user. Therefore, in one possible implementation, after screening out multiple candidate storage solutions that meet both cost requirements and read / write performance requirements from multiple storage solutions, the method can further include:

[0081] sorting the candidate storage solutions based on the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in each storage solution;

[0082] One or more optimal candidate storage solutions are selected from the sorted candidate storage solutions, and the selected candidate storage solutions and the corresponding cluster architecture are provided to the user.

[0083] Among them, sorting multiple candidate storage solutions according to the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in each storage solution can include: determining a comprehensive score for each candidate storage solution according to the read and write performance of the cluster architecture under each candidate storage solution and the cost of the hardware configuration in each candidate storage solution, and sorting multiple candidate storage solutions according to the comprehensive score of each candidate storage solution.

[0084] As described above, the read and write performance requirements include read and write performance intervals, and the cost requirements include cost requirements. Therefore, the above-mentioned determination of the comprehensive score of each candidate storage solution based on the read and write performance of the cluster architecture under each candidate storage solution and the cost of the hardware configuration in each candidate storage solution may include: for any candidate storage solution, subtracting the read and write performance of the cluster architecture under the candidate storage solution from the highest value of the read and write performance interval to obtain a first difference; subtracting the cost of the hardware configuration in the candidate storage solution from the lowest value of the cost interval to obtain a second difference; and using the standard deviation between the first difference and the second difference as the comprehensive score of the candidate storage solution. The comprehensive score determined in this way can comprehensively evaluate each candidate storage solution from the two dimensions of cost and read and write performance, which is conducive to screening out high-performance and low-cost candidate storage solutions based on the comprehensive scores of each candidate storage solution.

[0085] Among them, if the read and write performance requirements include read and write bandwidth requirements and read and write IOPS requirements, then subtracting the read and write performance of the cluster architecture under the candidate storage solution from the highest value of the read and write performance range to obtain the first difference includes: subtracting the total read and write bandwidth of the cluster architecture under the candidate storage solution from the highest value of the read and write bandwidth range to obtain the bandwidth difference, and subtracting the total read and write IOPS of the cluster architecture under the candidate storage solution from the highest value of the read and write IOPS range to obtain the IOPS difference. Furthermore, the above-mentioned use of the standard deviation between the first difference and the second difference as the comprehensive score of the candidate storage solution may include: using the standard deviation between the bandwidth difference, the IOPS difference, and the second difference as the comprehensive score of the candidate storage solution. If the read / write performance requirements only include read / write bandwidth requirements or read / write IOPS requirements, then subtracting the read / write performance of the cluster architecture under the candidate storage solution from the highest value of the read / write performance range to obtain the first difference includes: subtracting the total read / write bandwidth of the cluster architecture under the candidate storage solution from the highest value of the read / write bandwidth range to obtain the bandwidth difference, or subtracting the total read / write IOPS of the cluster architecture under the candidate storage solution from the highest value of the read / write IOPS range to obtain the IOPS difference. Furthermore, the above-mentioned use of the standard deviation between the first difference and the second difference as the comprehensive score of the candidate storage solution may include: using the standard deviation between the bandwidth difference and the second difference as the comprehensive score of the candidate storage solution, or using the standard deviation between the IOPS difference and the second difference as the comprehensive score of the candidate storage solution.

[0086] Among them, if the read and write performance requirement is the read and write performance value, and the cost requirement is the cost value, then the above-mentioned determination of the comprehensive score of each candidate storage solution based on the read and write performance of the cluster architecture under each candidate storage solution and the cost of the hardware configuration in each candidate storage solution may include: for any candidate storage solution, subtracting the read and write performance of the cluster architecture under the candidate storage solution from the read and write performance requirement to obtain a first difference; subtracting the cost of the hardware configuration in the candidate storage solution from the cost requirement to obtain a second difference; and using the standard deviation between the first difference and the second difference as the comprehensive score of the candidate storage solution.

[0087] It should be understood that the comprehensive score calculation method described above can be used to calculate the comprehensive score of each of the multiple candidate storage solutions, and then the comprehensive scores can be sorted and the candidate storage solution with the highest comprehensive score can be selected and provided to the user. Alternatively, the top N candidate storage solutions with the highest comprehensive scores can be selected and provided to the user, and this is not limited in the present embodiment. A higher comprehensive score indicates higher read and write performance and lower cost.

[0088] In actual applications, when providing the selected candidate storage solutions and the corresponding cluster architecture to the user, the cost of the hardware configuration of the candidate storage solution and the read-write performance under the cluster architecture can also be provided. In order to provide the above information to the user intuitively, the selected candidate storage solutions and the corresponding cluster architecture, cost, read-write performance and other information can be displayed to the user in the form of a chart. For example, a Ceph architecture diagram can be used to display the distribution and quantity of components such as OSD, MON, MGR, MDS, the network structure type and the storage pool type. A hardware configuration details list can be used to list the specific models and quantities of the CPU, memory, hard disk, network card, etc. used. In addition, the predicted performance can be displayed, that is, the predicted total read and write IOPS and total read and write bandwidth. The cost can also be displayed, including the hardware cost and the possible energy consumption cost. It should be understood that if two or more candidate storage solutions are provided to the user, the hardware configuration, cluster architecture, cost, read and write performance and other information of each candidate storage solution can be displayed separately. The embodiment of the present disclosure does not limit the display method of the above information.

[0089] It should be understood that the method proposed in the embodiment of the present disclosure can be applied to various distributed storage systems, and users can select the type of distributed storage system they need based on actual needs. For example, if the user has particularly high requirements for the performance of the file system, then the Lustre system may be a better choice; if the user needs a simple object storage solution, the MinIO system may be a better choice. Furthermore, the method of the embodiment of the present disclosure can be used to generate candidate storage solutions and corresponding cluster architectures. In other words, the method of the embodiment of the present disclosure can also first guide the user to select the type of distributed storage system they need, and then generate candidate storage solutions and corresponding cluster architectures under the corresponding distributed storage system based on the user's needs.

[0090] The method according to the embodiments of the present disclosure can utilize a hardware product information library to collect hardware information about various hardware products on the market, evaluate the performance and cost of different candidate storage solutions based on read / write performance and cost requirements, and perform a performance and cost comparison of various candidate storage solutions to select the optimal candidate storage solution. By scoring and ranking candidate storage solutions based on performance and cost, it is possible to comprehensively consider multiple factors to recommend the optimal candidate storage solution, thereby achieving a balance between read / write performance and cost.

[0091] Figure 2 A block diagram of an apparatus for determining a cluster architecture and hardware configuration of a distributed storage system according to an embodiment of the present disclosure is shown. Figure 2 As shown, the device includes:

[0092] The user demand acquisition module 201 is configured to acquire user requirements of the distributed storage system to be built, wherein the user requirements include storage capacity requirements, cost requirements, read / write performance requirements, and storage access method requirements;

[0093] The storage solution determination module 202 is configured to determine multiple storage solutions based on the storage capacity requirement, or further based on the storage access method requirement, by searching in a hardware product information library and selecting a storage pool type and storage pool configuration parameters, wherein the storage solution includes the hardware configuration, the storage pool type, and the storage pool configuration parameters.

[0094] The cluster architecture determination module 203 is configured to determine the number of core components of the distributed storage system for each of the multiple storage solutions, so as to obtain the cluster architecture under each storage solution;

[0095] The candidate solution screening module 203 is configured to calculate the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution, and screen candidate storage solutions and corresponding cluster architectures that meet the cost requirements and the read and write performance requirements from the multiple storage solutions.

[0096] In one possible implementation, the device also includes: a product information library construction and update module, which is used to obtain parameter information of multiple publicly available hardware products after obtaining the user needs to construct or update the hardware product information library, and the multiple hardware products include central processing units, memory, hard disks and network cards.

[0097] In one possible implementation, the multiple storage solutions are determined based on the storage capacity requirement by searching in a hardware product information library and selecting storage pool types and storage pool configuration parameters, including: searching for parameter information of multiple hard disks in the hardware product information library based on the storage capacity requirement to determine the type of hard disk and the capacity of a corresponding single hard disk; selecting one or more storage pool types and corresponding storage pool configuration parameters from a plurality of preset storage pool types and corresponding storage pool configuration parameters, and calculating the total storage space required for each storage pool type and corresponding storage pool configuration parameters based on the storage capacity requirement; determining the number of hard disks for each storage pool type and corresponding storage pool configuration parameters based on the total storage space and the capacity of a single hard disk; and determining different combinations of the determined hard disk type and the capacity of a corresponding single hard disk, storage pool type and corresponding storage pool configuration parameters, and the number of hard disks as the multiple storage solutions.

[0098] In one possible implementation, the hardware configuration also includes the number of central processing unit (CPU) cores and the memory size, and the device also includes: a CPU core number determination module, used to determine the number of CPU cores based on the number of at least one core component of the distributed storage system; a memory size determination module, used to determine the memory size based on the total storage space and the number of at least one core component; and an adding module, used to add the number of CPU cores and memory size to the corresponding storage solution.

[0099] In one possible implementation, when determining multiple storage solutions by searching in a hardware product information library and selecting storage pool types and storage pool configuration parameters, the network structure of the distributed storage system is also selected to determine that the cluster architecture also includes the multiple storage solutions of the network structure.

[0100] In one possible implementation, the calculation of the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution includes: calculating the read and write performance based on the hardware configuration and the cluster architecture in each storage solution; determining parameter information of the network card based on the read and write performance; selecting the model and quantity of the network card for the cluster architecture by searching in the hardware product information library; and calculating the cost of the hardware configuration including the model and quantity of the network card.

[0101] In one possible implementation, after screening multiple candidate storage solutions that meet the cost requirements and the read-write performance requirements from the multiple storage solutions, the method further includes: sorting the multiple candidate storage solutions according to the read-write performance of the cluster architecture under each of the multiple candidate storage solutions and the cost of the hardware configuration in each storage solution; selecting one or more optimal candidate storage solutions from the sorted multiple candidate storage solutions, and providing the selected candidate storage solutions and the corresponding cluster architecture to the user.

[0102] According to the device of the embodiment of the present disclosure, it is possible to utilize the hardware product information library in combination with user needs to automatically generate the hardware configuration and cluster architecture of the distributed storage system that meets the user's needs. Moreover, the generated cluster architecture and hardware configuration are selected based on the comprehensive read and write performance and cost, which can simplify the complexity of the user in selecting suitable hardware for the cluster, and provide the user with a quick and convenient recommendation of the cluster architecture and hardware configuration of the distributed storage system, especially it can generate the cluster architecture and hardware configuration of the distributed storage system that meets the capacity requirements and has high performance and low cost.

[0103] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0104] An embodiment of the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0105] An embodiment of the present disclosure further provides a non-volatile computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.

[0106] An embodiment of the present disclosure further provides a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program implements the steps of the above method when executed by a processor.

[0107] Figure 3 FIG1 shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Figure 3 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0108] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.

[0109] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0110] A computer-readable storage medium can be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0111] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0112] The computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, server instructions, server-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The computer readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, by utilizing state information of computer-readable program instructions to personalize and customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0113] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0114] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a server, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0115] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0116] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0117] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for determining the cluster architecture and hardware configuration of a distributed storage system, characterized in that: The cluster architecture includes at least the number of core components of the distributed storage system, the storage pool type, and storage pool configuration parameters; the hardware configuration includes at least the hard disk type, the capacity of a single hard disk, and the number of hard disks; and the method includes: Obtain user requirements for the distributed storage system to be built, including storage capacity requirements, cost requirements, read / write performance requirements, and possibly storage access method requirements; Determining multiple storage solutions based on the storage capacity requirement, or further based on the storage access method requirement, by searching in a hardware product information library and selecting a storage pool type and storage pool configuration parameters, wherein the storage solution includes the hardware configuration, the storage pool type, and the storage pool configuration parameters; For each of the multiple storage solutions, determine the number of core components of the distributed storage system to obtain the cluster architecture under each storage solution; The read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution are calculated, and candidate storage solutions and corresponding cluster architectures that meet the cost requirements and the read and write performance requirements are screened from the multiple storage solutions.

2. The method according to claim 1, characterized in that The method further comprises: After obtaining the user's needs, obtain the parameter information of multiple publicly available hardware products to build or update the hardware product information library. The various hardware products include central processing units, memory, hard disks and network cards.

3. The method according to claim 2, characterized in that According to the storage capacity requirement, multiple storage solutions are determined by searching in a hardware product information library and selecting a storage pool type and storage pool configuration parameters, including: According to the storage capacity requirement, searching parameter information of multiple hard disks in the hardware product information library to determine the type of hard disk and the capacity of a corresponding single hard disk; selecting one or more storage pool types and corresponding storage pool configuration parameters from a plurality of preset storage pool types and corresponding storage pool configuration parameters, and calculating the total storage space required for each storage pool type and corresponding storage pool configuration parameters according to the storage capacity requirement; For each storage pool type and corresponding storage pool configuration parameters, determine the number of hard disks according to the total storage space and the capacity of a single hard disk; Different combinations of the determined hard disk type and the corresponding capacity of a single hard disk, the storage pool type and the corresponding storage pool configuration parameters, and the number of hard disks are determined as the multiple storage solutions.

4. The method according to claim 3, characterized in that The hardware configuration also includes the number of CPU cores and memory size, and the method further includes: Determining the number of CPU cores according to the number of at least one core component of the distributed storage system; Determining the memory size based on the total storage space and the number of the at least one core component; Add the stated number of CPU cores and memory size to the corresponding storage plan.

5. The method according to claim 4, characterized in that When determining multiple storage solutions by searching in the hardware product information library and selecting storage pool types and storage pool configuration parameters, the network structure of the distributed storage system is also selected to determine that the cluster architecture also includes the multiple storage solutions of the network structure.

6. The method according to claim 5, characterized in that The calculating the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution includes: Calculating the read and write performance based on the hardware configuration and the cluster architecture in each storage solution; Determining parameter information of the network card according to the read and write performance; Selecting the model and quantity of the network card used for the cluster architecture by searching in the hardware product information library; The cost of the hardware configuration including the model and quantity of the network cards is calculated.

7. The method according to any one of claims 1 to 6, characterized in that When multiple candidate storage solutions that meet the cost requirement and the read / write performance requirement are screened from the multiple storage solutions, the method further includes: sorting the multiple candidate storage solutions according to the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in each storage solution; One or more optimal candidate storage solutions are selected from the sorted candidate storage solutions, and the selected candidate storage solutions and the corresponding cluster architecture are provided to the user.

8. A device for determining the cluster architecture and hardware configuration of a distributed storage system, characterized in that: The cluster architecture includes at least the number of core components of the distributed storage system, the storage pool type, and storage pool configuration parameters; the hardware configuration includes at least the hard disk type, the capacity of a single hard disk, and the number of hard disks; and the device includes: A user demand acquisition module is configured to acquire user requirements of the distributed storage system to be built, wherein the user requirements include storage capacity requirements, cost requirements, read / write performance requirements, and storage access method requirements; a storage solution determination module configured to determine a plurality of storage solutions based on the storage capacity requirement or the storage access method requirement by searching in a hardware product information library and selecting a storage pool type and storage pool configuration parameters, the storage solution including the hardware configuration, the storage pool type, and the storage pool configuration parameters; a cluster architecture determination module configured to determine, for each of the multiple storage solutions, the number of core components of the distributed storage system to obtain the cluster architecture under each storage solution; The candidate solution screening module is configured to calculate the read and write performance of the cluster architecture under each storage solution and the cost of the hardware configuration in the storage solution, and screen candidate storage solutions and corresponding cluster architectures that meet the cost requirements and the read and write performance requirements from the multiple storage solutions.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the operations of the method according to any one of claims 1 to 7.

10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the operation of the method according to any one of claims 1 to 7 is realized.