File storage method and device, electronic equipment and storage medium
By selecting appropriate resource pools and storage devices in the distributed storage system, data migration is avoided, bandwidth reduction and stability issues during expansion are resolved, and system performance is improved.
Patent Information
- Application Number
- CN202510819975.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
AI Technical Summary
During the expansion process of existing distributed storage systems, the system's available bandwidth decreases and stability is affected. Especially during data migration, user access is affected and system throughput decreases.
When a file to be stored is detected, the target resource pool is selected from the existing resource pool and the expansion resource pool based on the space utilization of multiple resource pools, and the target storage device is determined based on the file attributes, avoiding data migration and directly writing the file to the target storage device.
It improves system bandwidth, avoids bandwidth consumption during data migration, and improves system stability and performance. Bandwidth increases linearly with the increase of storage devices.
Smart Images

Figure CN120669920A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to fields such as data storage. More specifically, the present disclosure provides a file storage method, device, electronic device, storage medium, and computer program product. Background Art
[0002] Distributed storage systems can provide cloud servers with low-latency, persistent, highly reliable, and highly elastic storage services. Summary of the Invention
[0003] The present disclosure provides a file storage method, apparatus, electronic device, storage medium, and computer program product.
[0004] According to one aspect of the present disclosure, a file storage method is provided, comprising: in response to detecting a file to be stored, determining a target resource pool from a plurality of resource pools according to the space utilization rates of the respective plurality of resource pools; the plurality of resource pools each comprising a plurality of storage devices; determining a target storage device from a plurality of storage devices in the target resource pool according to attributes of the file; and writing the file to the target storage device; wherein the plurality of resource pools comprise at least one existing resource pool and at least one expanded resource pool, and the at least one expanded resource pool is obtained by expansion when the space utilization rate of the at least one existing resource pool satisfies a predetermined expansion condition.
[0005] According to another aspect of the present disclosure, a file storage device is provided, comprising: a first determination module, a second determination module, and a write module. The first determination module is used to determine a target resource pool from a plurality of resource pools in response to detecting a file to be stored, based on the space utilization rates of the plurality of resource pools; the plurality of resource pools each include a plurality of storage devices. The second determination module is used to determine a target storage device from a plurality of storage devices in the target resource pool based on the attributes of the file. The write module is used to write the file to the target storage device. The plurality of resource pools include at least one existing resource pool and at least one expanded resource pool, and the at least one expanded resource pool is obtained by expansion when the space utilization rate of at least one existing resource pool meets a predetermined expansion condition.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided in the present disclosure when executed by a processor.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1 is a schematic diagram of an application scenario of the file storage method and device according to an embodiment of the present disclosure;
[0012] Figure 2 is a schematic flow chart of a file storage method according to an embodiment of the present disclosure;
[0013] Figure 3 is a schematic diagram of a file storage method according to an embodiment of the present disclosure;
[0014] Figure 4 is a schematic structural block diagram of a file storage device according to an embodiment of the present disclosure; and
[0015] Figure 5 It is a structural block diagram of an electronic device used to implement the file storage method of an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0017] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0018] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0019] A distributed storage system can include multiple machines, each of which can be configured with one or more storage devices to store data. These storage devices can be disks. As the amount of stored data increases, the remaining available space in the distributed storage system gradually decreases. When storage space is insufficient, additional storage nodes can be added.
[0020] However, the performance of the expanded distributed storage system needs to be improved. For example, consider a distributed storage system consisting of 10 servers. After a period of use, the space utilization of each server is high. Capacity expansion is now possible, for example, by adding 10 more servers to store data.
[0021] In one technical solution, because the space utilization of the old 10 servers is high, even with fast data write speeds, if an exception such as a write error occurs, the old servers will not be able to write the entire data to be stored. Therefore, the 10 new servers will be used first. As can be seen, the available bandwidth of the distributed storage system in this way is the bandwidth of the 10 new servers, not the bandwidth of all 20 servers, resulting in lower system available bandwidth.
[0022] In another technical solution, to utilize the bandwidth of 20 machines during data storage, data migration can be performed. Specifically, some data from 10 old machines can be migrated to 10 new machines. After the data migration, all 20 machines have more free space, so even with faster write speeds, the bandwidth of all 20 machines can be utilized. However, the data migration process itself consumes some bandwidth. Therefore, during the data migration process, the actual bandwidth available to users is less than the total bandwidth of the 20 machines, affecting user access and reducing system throughput.
[0023] The presently disclosed embodiments provide a file storage method for storing files in multiple resource pools, the multiple resource pools including at least one existing resource pool and at least one expanded resource pool, the at least one expanded resource pool being expanded when the space utilization of at least one existing resource pool satisfies predetermined expansion conditions. After detecting a file to be stored, a target resource pool can be determined from the multiple resource pools based on their respective space utilizations. A target storage device can then be determined from multiple storage devices in the target resource pool based on the attributes of the file, and the file can then be written to the target storage device.
[0024] Using the above file storage method, expansion is performed when an existing resource pool meets the predefined expansion conditions. During data storage, a target resource pool is first selected, and the files to be stored are then written to the target resource pool. It should be noted that, in practice, expansion can be performed even when the existing resource pool's space utilization is relatively low (for example, at 70% utilization). This allows the existing resource pool to still have ample remaining storage space after expansion, thus avoiding prioritizing the use of storage devices in the expanded resource pool. Furthermore, data migration is not required after expansion.
[0025] On the one hand, since the target resource pool is selected not just from the existing resource pool but from the existing resource pool and the expanded resource pool, the system's available bandwidth is not the bandwidth of the expanded resource pool, but the combined bandwidth of the existing and expanded resource pools. Therefore, compared to a solution that prioritizes only the expanded storage devices, the system bandwidth can be increased. On the other hand, data migration is not required after expansion, which avoids bandwidth consumption during the data migration process. Therefore, compared to data migration solutions, the system bandwidth can be increased, as well as system stability. It can be seen that the bandwidth of a distributed storage system using the above method increases linearly with the increase in storage devices, thereby improving the performance of the distributed storage system.
[0026] The technical solutions provided by the present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] Figure 1 Schematic diagram of an application scenario of the file storage method and device according to an embodiment of the present disclosure.
[0028] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0029] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a terminal device 101, a scheduling device 102, storage nodes 103, 104, 105 and a network.
[0030] The network is used to provide a medium for communication links between the terminal device 101, the scheduling device 102, and the storage nodes 103, 104, 105. The network may include various connection types, such as wired and / or wireless communication links, etc.
[0031] The user can use the terminal device 101 to interact with the scheduling device 102 via the network to receive or send messages, etc. The terminal device 101 can be any electronic device with a display screen and supporting web browsing, including but not limited to a smartphone, tablet computer, laptop computer, desktop computer, etc.
[0032] The scheduling device 102 can be a server or other electronic device. The scheduling device 102 is used to receive the file to be stored sent by the terminal device 101, and then determine the storage location, such as the identifier of the storage device in the storage node 103, 104, 105, and write the file to the target storage device.
[0033] The storage nodes 103, 104, 105 may be servers or other electronic devices. One storage node 103, 104, 105 is a machine. One storage node 103, 104, 105 may be installed with multiple storage devices, which may be disks.
[0034] The scheduling device 102 and any storage node 103 , 104 , 105 may be integrated into one electronic device, or the scheduling device 102 and the storage nodes 103 , 104 , 105 may be deployed in different electronic devices respectively.
[0035] It should be noted that the file storage method provided in the embodiment of the present disclosure can generally be executed by the scheduling device 102. Accordingly, the file storage apparatus provided in the embodiment of the present disclosure can generally be set in the scheduling device 102.
[0036] It should be understood that Figure 1 The number of terminal devices, scheduling devices, and storage nodes in the embodiment is merely illustrative. Any number of terminal devices, scheduling devices, and storage nodes may be provided as required.
[0037] Figure 2 is a schematic flowchart of a file storage method according to an embodiment of the present disclosure.
[0038] like Figure 2 As shown, the file storage method 200 may include operations S210 to S230.
[0039] In operation S210 , in response to detecting a file to be stored, a target resource pool is determined from a plurality of resource pools according to space utilization rates of the plurality of resource pools, each of the plurality of resource pools including a plurality of storage devices.
[0040] For example, a distributed storage system includes multiple machines, each of which can be configured with one or more storage devices to store data. The storage devices can be disks. The multiple machines in the distributed storage system can be pre-partitioned to obtain one or more existing resource pools.
[0041] In addition, during use, it can be determined whether the space utilization of an existing resource pool meets a predetermined expansion condition, and the predetermined expansion condition can be that the space utilization reaches a set threshold. If the space utilization of an existing resource pool meets the predetermined expansion condition, expansion processing can be performed. Expansion processing includes adding one or more new machines to a distributed storage system, thereby obtaining one or more expanded resource pools. For example, if the space utilization of an existing resource pool is large, an expanded resource pool can be obtained through expansion processing, and the expanded resource pool can include 3 new machines. Two expanded resource pools can also be obtained through expansion processing, and the expanded resource pools each include 5 new machines. In the process of a machine joining a resource pool, the identifier of the resource pool to which the machine belongs can be configured, thereby declaring the resource pool to which the machine belongs.
[0042] The plurality of resource pools include the existing resource pool before expansion and the expanded resource pool after expansion. A resource pool with the lowest space utilization rate can be selected from the plurality of resource pools as the target resource pool.
[0043] In operation S220 , a target storage device is determined among a plurality of storage devices in a target resource pool according to attributes of the file.
[0044] For example, the attributes of a file may include the file's identification, size, type, permissions, etc. The attributes of the file may be converted into a hash value through a hash function, and then the hash value is mapped to the number of the target storage device, and then the target storage device is determined by the number.
[0045] In operation S230 , the file is written to the target storage device.
[0046] For example, a request is sent to a machine configured with a target storage device to write a file to the target storage device.
[0047] According to the file storage method provided by the embodiment of the present disclosure, the method expands the existing resource pool when it meets the predetermined expansion conditions. During the data storage process, the target resource pool is first selected, and then the file to be stored is written into the target resource pool. It should be noted that in the actual expansion process, the expansion can be carried out when the space utilization rate of the existing resource pool is relatively low (for example, the space utilization rate reaches 70%). In this way, after the expansion, the existing resource pool still has a large amount of remaining storage space, thereby avoiding the priority use of only the storage devices in the expanded resource pool. In addition, there is no need to relocate data after the expansion.
[0048] On the one hand, since the target resource pool is selected not just from the existing resource pool but from the existing resource pool and the expanded resource pool, the system's available bandwidth is not the bandwidth of the expanded resource pool, but the combined bandwidth of the existing and expanded resource pools. Therefore, compared to a solution that prioritizes only the expanded storage devices, the system bandwidth can be increased. On the other hand, data migration is not required after expansion, which avoids bandwidth consumption during the data migration process. Therefore, compared to data migration solutions, the system bandwidth can be increased, as well as system stability. It can be seen that the bandwidth of a distributed storage system using the above method increases linearly with the increase in storage devices, thereby improving the performance of the distributed storage system.
[0049] Next, the process of determining a target resource pool from multiple resource pools is described.
[0050] In this embodiment, the utilization of the resource pools can be compared with a first utilization threshold, which can be 70%, 80%, 85%, or the like. If the space utilization of at least one resource pool among the multiple resource pools is less than or equal to the first utilization threshold, a resource pool can be selected from the at least one resource pool as the target resource pool, and the selection method can be random. If the space utilization of multiple resource pools is greater than the first utilization threshold, the resource pool with the lowest space utilization among the multiple resource pools can be determined as the target resource pool.
[0051] For example, the resource pools can be traversed and the traversed resource pools can be added to the target queue. In addition, a comparison can be made to see whether the space utilization of the traversed resource pools is less than or equal to a first utilization threshold. If so, the resource pool is also added to the target set; otherwise, it is not added to the target set. Next, it can be determined whether the target set is empty. If the target set is not empty, a resource pool can be randomly selected from the target set as the target resource pool. If the target set is empty, the resource pool with the lowest utilization can be selected from the target queue as the target resource pool.
[0052] In this embodiment, when some resource pools have low utilization, the target resource pool can be randomly selected. This prevents a large number of files to be stored from rushing to the resource pool with the lowest utilization when write speeds are high, which could cause a transient hotspot. By randomly selecting, the load is distributed among the resource pools with low utilization. When all resource pools have high utilization, the resource pool with the lowest utilization is selected, thereby delaying the resource pools from reaching the critical point where they need to be expanded.
[0053] According to another embodiment of the present disclosure, the predetermined expansion condition includes: the space utilization of at least one existing resource pool is greater than or equal to a second utilization threshold. The second utilization threshold is, for example, 80%, 70%, etc., which is not limited in this embodiment.
[0054] According to another embodiment of the present disclosure, if the process of determining a target resource pool from multiple resource pools is based on a first utilization threshold, and the predetermined capacity expansion condition includes: the space utilization of at least one existing resource pool is greater than or equal to a second utilization threshold, in this case, the second utilization threshold may be lower than the first utilization threshold. For example, the first utilization threshold is 80% and the second utilization threshold is 70%.
[0055] It should be noted that if the first utilization threshold is lower than the second utilization threshold, for example, the first utilization threshold is 70% and the second utilization threshold is 80%, then the utilization of the existing resource pool reaches 80% for expansion. After expansion, since the utilization of the existing resource pools is greater than 70%, the target set will only include the newly expanded resources, resulting in the priority use of the newly expanded resource pool and the failure to utilize the existing resource pool. Therefore, this embodiment sets the first utilization threshold relatively high and the second utilization threshold relatively low, thereby ensuring that the target resource pool can be selected from both the existing resource pool and the expanded resource pool.
[0056] According to another embodiment of the present disclosure, the process of determining a target storage device from among multiple storage devices in a target resource pool based on file attributes may include: determining a target replication group from among the multiple replication groups based on a hash value determined based on the file attributes and the number of replication groups (RGs) in the target resource pool, the target replication group including replica meta information. Then, determining the target storage device based on the replica meta information.
[0057] For example, file attributes include a file identifier and a file offset. These identifiers and file offsets can be concatenated, added, or otherwise processed to produce a string, which is then hashed to produce a hash value. The hash value can then be mapped to a target replication group. For example, if the target resource pool includes several replication groups, the hash value can be modulo-ed by the number of replication groups in the target resource pool to produce a remainder. Each remainder corresponds to a target replication group.
[0058] It should be noted that during the processing process, large files can be processed in blocks, for example, each 64MB or other predetermined data size is a file block, thereby dividing the file into multiple file blocks. Each file block corresponds to an offset, so that a file is associated with multiple hash values. The hash value of each file block is mapped to a target replication group, achieving distributed load balancing for large files. For small files, the file can be processed without block processing. In this way, only one hash value is calculated for each file, and this hash value is mapped to only one target replication group, meeting the centralized storage requirements of small files.
[0059] A replication group includes three or another number of replicas, divided into primary and secondary replicas. Multiple replicas within the same replication group are stored in a single resource pool, but different replicas are stored on different storage devices on different machines. After determining the target replication group, the target storage device can be determined and data written. This embodiment determines a hash value based on file attributes, and then accurately determines the file's storage location based on the hash value. This allows for lightweight computation to determine the storage location, improving file storage efficiency.
[0060] In one example, the replica meta information of each replica within the same target replication group can be determined based on the target replication group. The replica meta information records the storage location of the replica, namely, the machine ID and storage device ID of the replica. Based on the machine ID and storage device ID, the target storage device can be determined. A request can then be sent to the machine where the target storage device is located, thereby writing data to the target storage device.
[0061] In another example, the replica meta information of the primary replica can be determined based on the target replication group. This replica meta information can then be used to determine the machine and storage device where the primary replica resides. During the actual data writing process, a request can be sent to the machine where the target storage device storing the primary replica resides, thereby writing the data to the target storage device. Furthermore, upon receiving the request, the machine storing the primary replica can send the data to be stored to the machine storing the secondary replica, thereby writing the data to the corresponding storage device on that machine.
[0062] In addition, during the actual data writing process, if it is detected that at least two replicas in the same replication group have been written successfully, the writing process for the replication group can be terminated. If other replication groups have not completed data writing for the time being, the replication group can synchronize data with other replication groups that have completed writing.
[0063] Next, the process of splitting a resource pool is described.
[0064] In this embodiment, multiple machines can be divided to obtain an existing resource pool. A machine includes one or more storage devices, and the storage devices in the same machine are referred to as a device group. For example, a reference value for the number of device groups can be determined based on the maximum number of replication groups in a single resource pool, the number of replicas included in a single replication group, the maximum amount of data in a single replica, the number of storage devices included in a single device group, the capacity of a single storage device, and predetermined expansion conditions. Then, based on the reference value for the number of device groups, multiple device groups are divided to obtain at least one existing resource pool. After the division, each existing resource pool includes at least one device group, each device group includes at least one storage device, and the number of device groups in a single existing resource pool is less than or equal to the reference value for the number of device groups.
[0065] For example, the predetermined capacity expansion condition is that the space utilization of at least one existing resource pool is greater than or equal to the second utilization threshold. For example, a resource pool includes multiple replication groups, and a single replication group includes multiple replicas. A resource pool includes one or more machines, each of which is equipped with one or more storage devices. The storage devices in the same machine are called a device group, and the storage devices can be disks. The reference value for the number of device groups can be calculated using the following formula:
[0066]
[0067] in, Indicates the reference value of the number of equipment groups. Indicates the maximum amount of data for a single replica, for example It's 30GB. Indicates the maximum number of replication groups in a resource pool, for example It’s 30,000. Indicates the number of replicas included in a single replication group, for example is 3. Indicates the capacity of a single storage device, such as It's 7TB. Indicates the number of storage devices installed on a single machine, for example is 8. represents the second utilization threshold, e.g. It is 80%.
[0068] After obtaining the reference value of the number of device groups, the resource pool can be divided according to the reference value of the number of device groups. For example, when the reference value of the number of device groups is 10, each predetermined number of machines can be divided into a resource pool, and the predetermined number is less than or equal to 10.
[0069] In practice, the number of replication groups in a single resource pool can be around 30,000. Each replication group can include one master replica and two slave replicas. Multiple replicas in the same replication group are stored in a single resource pool, but different replicas in the same replication group are stored on different storage devices on different machines. Furthermore, the data volume of a single replica should not be too large to avoid prolonged replica recovery. For example, if the data volume of a single replica does not exceed 30GB, and a machine is configured with eight storage devices, each with a capacity of 7TB, then the number of machines in a resource pool can be no more than 60.
[0070] In this embodiment, the number of machines in the resource pool can be accurately determined based on the reference value of the number of device groups, which not only ensures performance constraints such as the replication group size, but also reserves a certain amount of storage space through predetermined expansion conditions to prevent storage exhaustion.
[0071] The above describes the process of splitting a resource pool.
[0072] Figure 3 It is a schematic diagram of a file storage method according to an embodiment of the present disclosure.
[0073] In this embodiment, the distributed storage system includes multiple machines, each of which can be configured with one or more storage devices to store data. The storage devices can be disks. The multiple machines in the distributed storage system can be pre-partitioned to create one or more existing resource pools 311. During use, if the space utilization of an existing resource pool 311 meets predetermined expansion conditions, expansion processing can be performed to create an expanded resource pool 312. These existing resource pools 311 and expanded resource pools 312 constitute multiple resource pools 313.
[0074] When a file storage instruction is received, a target resource pool 320 may be determined from the multiple resource pools 313 based on the space utilization of each of the multiple resource pools 313. For example, if the space utilization of at least one resource pool in the multiple resource pools 313 is less than or equal to a first utilization threshold, a resource pool is selected from the at least one resource pool as the target resource pool 320. If the space utilization of all the multiple resource pools 313 is greater than the first utilization threshold, the resource pool with the lowest space utilization among the multiple resource pools 313 is determined as the target resource pool 320.
[0075] After determining the target resource pool 320, the target storage device 350 can be determined from the multiple storage devices in the target resource pool 320 based on the file attribute 340 of the file 330. For example, a hash value can be determined based on the file attribute 340 of the file 330, and then a modulo operation can be performed on the hash value and the number of multiple replication groups in the target resource pool 320 to obtain a remainder. The identifier of the target replication group is determined based on the remainder, thereby determining the target data group for storing the file 330. Since each replication group includes replica meta information, such as replica meta information of the primary replica, it can also include replica meta information of the secondary replica. Based on the replica meta information, the identifier of the machine used to store the replica and the identifier of the storage device can be determined, thereby determining the target storage device 350. Next, the file 330 can be written to the target storage device 350.
[0076] Figure 4 4 is a schematic structural block diagram of a file storage device according to an embodiment of the present disclosure.
[0077] like Figure 4 As shown, the file storage device 400 may include a first determination module 410 , a second determination module 420 and a writing module 430 .
[0078] The first determination module 410 is configured to, in response to detecting a file to be stored, determine a target resource pool from a plurality of resource pools based on space utilization rates of the plurality of resource pools, each of which includes a plurality of storage devices. The plurality of resource pools may include at least one existing resource pool and at least one expanded resource pool, wherein the at least one expanded resource pool is obtained by expanding the capacity of the at least one existing resource pool when the space utilization rate of the at least one existing resource pool satisfies a predetermined expansion condition.
[0079] The second determining module 420 is configured to determine a target storage device from a plurality of storage devices in the target resource pool according to attributes of the file.
[0080] The writing module 430 is used to write the file to the target storage device.
[0081] According to an embodiment of the present disclosure, the first determination module includes: a first determination submodule and a second determination submodule. The first determination submodule is configured to select a resource pool from the at least one resource pool as a target resource pool in response to detecting that the space utilization of at least one resource pool among the multiple resource pools is less than or equal to a first utilization threshold. The second determination submodule is configured to determine the resource pool with the lowest space utilization among the multiple resource pools as the target resource pool in response to detecting that the space utilization of the multiple resource pools is greater than the first utilization threshold.
[0082] According to an embodiment of the present disclosure, the predetermined capacity expansion condition includes: the space utilization of at least one existing resource pool is greater than or equal to a second utilization threshold, wherein the second utilization threshold is less than the first utilization threshold.
[0083] According to an embodiment of the present disclosure, the second determination module includes: a third determination submodule and a fourth determination submodule. The third determination submodule is configured to determine a target replication group from multiple replication groups based on a hash value determined based on file attributes of a file and the number of replication groups in a target resource pool, the target replication group including replica meta information. The fourth determination submodule is configured to determine a target storage device based on the replica meta information.
[0084] According to an embodiment of the present disclosure, at least one existing resource pool is obtained through a reference value determination module and a division module: the reference value determination module is used to determine a reference value for the number of device groups based on the maximum number of replication groups in a single resource pool, the number of replicas included in a single replication group, the maximum amount of data in a single replica, the number of storage devices included in a single device group, the capacity of a single storage device, and predetermined expansion conditions. The division module is used to divide multiple device groups based on the reference value for the number of device groups to obtain at least one existing resource pool. Each existing resource pool includes at least one device group, each device group includes at least one storage device, and the number of device groups in the existing resource pool is less than or equal to the reference value for the number of device groups.
[0085] According to an embodiment of the present disclosure, the file attributes include an identifier of the file and an offset of the file.
[0086] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned file storage method.
[0087] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above-mentioned file storage method.
[0088] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, including a computer program, which implements the above-mentioned file storage method when executed by a processor.
[0089] Figure 5 1 is a block diagram of an electronic device used to implement the file storage method of an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0090] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. Computing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to bus 504.
[0091] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0092] The computing unit 501 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the file storage method. For example, in some embodiments, the file storage method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the file storage method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the file storage method by any other suitable means (e.g., via firmware).
[0093] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0094] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0095] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0096] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0097] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0098] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0099] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0100] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A file storage method, comprising: In response to detecting a file to be stored, determining a target resource pool from the plurality of resource pools according to respective space utilization rates of the plurality of resource pools; Each of the plurality of resource pools includes a plurality of storage devices; determining a target storage device from the plurality of storage devices in the target resource pool according to the attribute of the file; as well as Writing the file to the target storage device; The multiple resource pools include at least one existing resource pool and at least one expanded resource pool. The at least one expanded resource pool is obtained by expansion when the space utilization rate of the at least one existing resource pool meets a predetermined expansion condition.
2. The method according to claim 1, wherein In response to detecting the file to be stored, determining a target resource pool from the multiple resource pools according to respective space utilization rates of the multiple resource pools includes: In response to detecting that the space utilization rate of at least one resource pool among the plurality of resource pools is less than or equal to a first utilization rate threshold, selecting a resource pool among the at least one resource pool as the target resource pool; and In response to detecting that the space utilization rates of the multiple resource pools are all greater than the first utilization rate threshold, a resource pool with the smallest space utilization rate among the multiple resource pools is determined as the target resource pool.
3. The method according to claim 2, wherein: The predetermined capacity expansion conditions include: The space utilization rate of the at least one existing resource pool is greater than or equal to the second utilization rate threshold; The second utilization threshold is smaller than the first utilization threshold.
4. The method according to claim 1, wherein Determining a target storage device from the plurality of storage devices in the target resource pool according to the attribute of the file includes: Determining a target replication group from the plurality of replication groups according to a hash value determined based on a file attribute of the file and the number of the plurality of replication groups in the target resource pool, the target replication group including replica meta information; and The target storage device is determined according to the replica meta information.
5. The method according to claim 4, wherein The file attributes include a file identifier and an offset of the file.
6. The method according to claim 1, wherein The at least one existing resource pool is obtained in the following manner: Determine a reference value for the number of device groups based on the maximum number of replication groups in a single resource pool, the number of replicas included in a single replication group, the maximum amount of data in a single replica, the number of storage devices included in a single device group, the capacity of a single storage device, and the predetermined expansion condition; as well as Dividing the plurality of device groups according to the reference value of the number of device groups to obtain at least one existing resource pool; Each existing resource pool includes at least one device group, each device group includes at least one storage device, and the number of device groups in the existing resource pool is less than or equal to the reference value of the number of device groups.
7. A file storage device comprising: a first determining module configured to, in response to detecting a file to be stored, determine a target resource pool from the plurality of resource pools according to respective space utilization rates of the plurality of resource pools; Each of the plurality of resource pools includes a plurality of storage devices; A second determining module is configured to determine a target storage device from among the plurality of storage devices in the target resource pool according to the attribute of the file; as well as A writing module, configured to write the file into the target storage device; The multiple resource pools include at least one existing resource pool and at least one expanded resource pool. The at least one expanded resource pool is obtained by expansion when the space utilization rate of the at least one existing resource pool meets a predetermined expansion condition.
8. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.