Data copy storage method and system, medium, electronic equipment and program product

By dividing the storage space into storage areas equal to the number of nodes in the super-converged architecture and using a software-defined storage subsystem for management, the problem that cluster storage pools cannot achieve IO localization in the HCI scenario is solved, and IO access efficiency and storage performance are improved.

CN120255827AActive Publication Date: 2025-07-04JINAN INSPUR DATA TECH CO LTD

Patent Information

Application Number
CN202510744585.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

In the HCI scenario, cluster storage pools cannot achieve IO localization, resulting in low IO access efficiency.

Method used

By determining the distributed nodes corresponding to the virtual machine in the super-converged architecture, and dividing the storage space into storage areas equal to the number of nodes during the disk formatting process, ensuring that the data replicas are stored on the corresponding nodes, and using the software-defined storage subsystem for refined management, realizing efficient localized storage of the data replicas.

Benefits of technology

Improves IO access efficiency, optimizes storage performance, enhances the applicability and competitiveness of hyperconverged architectures in large-scale virtualization environments without adding additional hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255827A_ABST
    Figure CN120255827A_ABST
Patent Text Reader

Abstract

The invention discloses a data copy storage method and system, a medium, electronic equipment and a program product, and relates to the technical field of server hyper-convergence, and the method comprises the steps: determining N distributed nodes corresponding to M virtual machines supported by a hyper-convergence architecture application; in a disk formatting process, dividing a to-be-allocated storage space associated with the super fusion architecture application according to the N distributed nodes so as to equally divide the to-be-allocated storage space into N storage areas; determining a corresponding relationship between the N storage areas and the N distributed nodes; wherein each storage area at least corresponds to one distributed node; and performing storage indication on the data copies generated by the M virtual machines according to the corresponding relationship, thereby solving the problem that a cluster storage pool cannot realize IO localization in an HCI scene, and realizing the technical effect of efficient localization storage of the data copies in the HCI scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of server hyper-convergence, and in particular, to a method, system, medium, electronic device, and program product for storing data replicas. Background Art

[0002] The hyper-converged infrastructure (HCI) integrates virtualization and software-defined storage (SDS) to optimize the IT infrastructure. SDS enhances data reliability, availability, and scalability by dispersing and replicating data across multiple hosts, surpassing traditional centralized storage. SDS provides block, object, and file services to meet diverse application requirements. Virtualization technology improves the utilization efficiency of hardware resources, enhances system security and service availability, and is an important user of SDS. The two operate in the same node in HCI, and data is directly stored on local disks. In addition, data replicas are usually stored on disks in the form of files or data blocks. For the scenario where the backend disk is a hardware of HDD (Hard Disk Drive) type, there will inevitably be performance bottleneck problems.

[0003] In related technologies, SSD (Solid State Drive) cache disks are often introduced for performance optimization. However, the backend storage of traditional server virtualization systems is usually provided by centralized storage. In the HCI system, the virtualization service and software-defined storage are deployed on the same physical node, and data replicas are ultimately stored on the local disks of the physical node. When accessing IO (Input / Output) through a virtual machine, there will be a problem of whether the IO accesses across nodes. Therefore, the related storage pool virtualization solutions cannot achieve IO localization in the cluster storage pool in the HCI scenario.

[0004] In view of the problem that the cluster storage pool in the HCI scenario cannot achieve IO localization in related technologies, no effective solution has been proposed yet. Summary of the Invention

[0005] The embodiments of the present application provide a method, system, medium, electronic device, and program product for storing data replicas to at least solve the problem that the cluster storage pool in the HCI scenario cannot achieve IO localization in related technologies.

[0006] According to an embodiment of the present application, a method for storing data replicas is provided, including: determining N distributed nodes corresponding to M virtual machines supported by a hyper-converged architecture application; during the process of disk formatting, dividing the storage space to be allocated associated with the hyper-converged architecture application according to the N distributed nodes, so as to equally divide the storage space to be allocated into N storage areas; determining the correspondence between the N storage areas and the N distributed nodes; wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; storing and indicating the data replicas generated by the M virtual machines according to the correspondence.

[0007] According to another embodiment of the present application, a data replica storage system is provided, including: a cluster file subsystem, configured to determine N distributed nodes corresponding to M virtual machines supported by a hyper-converged architecture application; during the process of disk formatting, dividing the storage space to be allocated associated with the hyper-converged architecture application according to the N distributed nodes, so as to equally divide the storage space to be allocated into N storage areas; a software-defined storage subsystem, connected to the cluster file subsystem, configured to determine the correspondence between the N storage areas and the N distributed nodes; wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; storing and indicating the data replicas generated by the M virtual machines according to the correspondence.

[0008] According to still another embodiment of the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0009] According to still another embodiment of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0010] According to still another embodiment of the present application, a computer program product is further provided, including a computer program, and the computer program implements the steps in any one of the above method embodiments when executed by a processor.

[0011] Through this application, first, N distributed nodes corresponding to M virtual machines supported by the hyper-converged architecture are determined. Then, during the process of disk formatting, the storage space to be allocated associated with the hyper-converged architecture application is divided according to the N distributed nodes, so as to equally divide the storage space to be allocated into N storage areas, ensuring that the number of divisions of the storage space to be allocated is consistent with the number of distributed cluster nodes. At the same time, the corresponding relationship between the N storage areas and the N distributed nodes is obtained. On this basis, the software-defined storage subsystem is instructed to place one of the virtual machine replicas to the corresponding node. Subsequently, the technical effect of efficient local storage of data replicas in the HCI scenario is achieved. Through the above method, the storage space is finely managed, ensuring the matching of the distribution of data replicas with the nodes where the virtual machines are located, reducing cross-node data transmission, improving the IO access efficiency, and optimizing the storage performance without increasing additional hardware costs, enhancing the applicability and competitiveness of the hyper-converged architecture in large-scale virtualization environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 is the hardware structure block diagram of the server device for the data replica storage method according to an embodiment of the present application;

[0014] Figure 2 is the flowchart of the data replica storage method according to an embodiment of the present application;

[0015] Figure 3 is the architecture schematic diagram for supporting the realization of cluster file system IO localization according to an embodiment of the present application;

[0016] Figure 4 is the schematic diagram of the link resource organizational structure according to an embodiment of the present application;

[0017] Figure 5 is the schematic diagram of the change in the relationship between virtual disk data and SDS replicas according to an embodiment of the present application;

[0018] Figure 6 is the structure block diagram of the data replica storage system according to an embodiment of the present application;

[0019] Figure 7 is the computer system structure block diagram of the electronic device according to an embodiment of the present application;

[0020] Among them, 102 in the above figure is a processor, 104 is a memory, 106 is a transmission device, 108 is an input / output device, 62 is a cluster file subsystem, 64 is a software-defined storage subsystem, 800 is a computer system, 801 is a CPU, 802 is a ROM, 803 is a RAM, 804 is a bus, 805 is an I / O interface, 806 is an input part, 807 is an output part, 808 is a storage part, 809 is a communication part, 810 is a driver, and 811 is a removable medium. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0022] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0023] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0024] As an optional implementation manner, the method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 is a hardware structure block diagram of a server device for a method of storing data copies in an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in Figure 1The structure shown is only illustrative and does not limit the structure of the above server device. For example, the server device may further include more or fewer components than those shown in Figure 1 or have a different configuration from that shown in Figure 1 .

[0025] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the data copy storage method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0026] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0027] In this embodiment, a method for storing data copies is provided. Figure 2 is a flowchart of the method for storing data copies according to the embodiments of the present application, as shown in Figure 2 . The process includes the following steps:

[0028] Step S202, determining N distributed nodes corresponding to M virtual machines supported by the hyper-converged architecture application;

[0029] It can be understood that in a hyper-converged infrastructure (HCI), multiple virtual machines (M) run on a distributed cluster composed of several physical nodes. Each physical node is both the host of the virtualized environment and a node of software-defined storage (SDS), jointly providing computing and storage resources. Through the above steps, the number of active virtual machines (M) in the current HCI system and the total number (N) of physical nodes (i.e., distributed nodes) supporting the operation of these virtual machines can be identified and determined. The M virtual machines can be any number, while the N nodes are the physical basis of the hyper-converged infrastructure. There is a relationship where M is less than or equal to N, meaning that each virtual machine corresponds to at least one distributed node, but the number of nodes may exceed the number of virtual machines, providing guarantees for system expansion and redundancy. This determination process is a prerequisite for the subsequent implementation of the IO localization strategy, ensuring that data can be optimized and distributed for specific nodes.

[0030] Step S204, during the process of disk formatting, divide the to-be-allocated storage space associated with the super-converged infrastructure application according to the N distributed nodes, so as to equally divide the to-be-allocated storage space into N storage areas;

[0031] It can be understood that in a traditional HCI architecture, the file system data of virtual machines may be disorderly distributed on each node of the cluster, resulting in frequent cross-node access during IO operations and limited performance. To overcome this challenge, by introducing a region division function in the disk formatting stage, the to-be-allocated storage space is evenly sliced into N storage areas (allocation zones), and each area corresponds to a specific distributed node. For example, when there are two nodes in the system, the available space in the storage pool will be sliced into two areas. The first storage area, zone1, corresponds to the first node, node1, and the second storage area, zone2, corresponds to the second node, node2. This division method ensures that when virtual machines on each node perform data read and write operations, they can preferentially access the data copies located on the local disk, significantly reducing network transmission overhead and improving IO performance. At the same time, the region division information will be passed to the SDS, further guiding the SDS to ensure that at least one copy is located on the corresponding node when creating data copies, matching the distribution strategy of the virtual machines and optimizing the overall system performance.

[0032] Step S206, determine the corresponding relationship between the N storage areas and the N distributed nodes; where each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers;

[0033] Optionally, by defining the binding rules between each allocation zone and a specific node, clear guidance is provided for subsequent data replica allocation. For example, assume N = 2, that is, there are two nodes, then the storage space will also be divided into two zones (zone1 and zone2), zone1 will correspond to node1, and zone2 will correspond to node2. This correspondence ensures the orderly distribution of data among different nodes, rather than randomly or blindly scattered throughout the cluster, laying the foundation for data localization.

[0034] Step S208, perform storage indication on the data replicas generated by the M virtual machines according to the correspondence.

[0035] It can be understood that during the virtual machine IO operation, the storage location of the data replica is indicated based on this correspondence. This means that when the virtual machines (M) perform read and write operations, the generated data replicas will be guided to be stored in the storage area corresponding to the node where the virtual machine is located to achieve the localization of IO access. Specifically, when creating a virtual disk or when a virtual machine IO write operation is triggered, the cluster file system will call the optimized space allocation interface and preferentially allocate disk space to the data replica located on the target node. For example, if virtual machine A is located on node1, then its data replica will be stored as much as possible on the physical disk corresponding to zone1, rather than distributed on any other nodes in the network. This method minimizes cross-node data transmission, improves the efficiency and performance of IO operations. Especially when reading data within a node, it is hardly affected by network latency.

[0036] Through the above method, first determine the N distributed nodes corresponding to the M virtual machines supported by the hyper-converged architecture application. Then, during the process of disk formatting, divide the storage space to be allocated associated with the hyper-converged architecture application according to the N distributed nodes, so as to equally divide the storage space to be allocated into N storage areas, ensuring that the number of divided storage spaces is consistent with the number of distributed cluster nodes. At the same time, obtain the correspondence between the N storage areas and the N distributed nodes. On this basis, instruct the software-defined storage subsystem to place one of the virtual machine replicas to the corresponding node according to this correspondence. Subsequently, the technical effect of efficient local storage of data replicas in software-defined storage is achieved. Through the above method, the storage space is finely managed, ensuring the matching of the distribution of data replicas with the nodes where the virtual machines are located, reducing cross-node data transmission, improving the efficiency of IO access, and optimizing the storage performance without increasing additional hardware costs, enhancing the applicability and competitiveness of the hyper-converged architecture in a large-scale virtualization environment.

[0037] In an exemplary embodiment, storing instructions for data replicas generated by the M virtual machines according to the corresponding relationship includes: determining physical node number information corresponding to a target virtual machine that generates the data replicas; determining a target distributed node associated with the target virtual machine based on the physical node number information, and determining a target storage area corresponding to the target distributed node through the corresponding relationship; setting the target storage area as the storage location of the data replicas of the target virtual machine to store instructions for the data replicas.

[0038] Optionally, the mechanism for storing instructions for the data replicas mainly focuses on determining the specific physical node number information where the "target virtual machine" that generates the data replicas is located. Briefly, in a hyper-converged architecture, since virtual machines can span multiple nodes, it is crucial to identify the physical nodes where they reside, which is the basis for localizing data replicas. Specifically, when virtual machine A generates a data write operation, it will first identify the physical node where A is located, assumed to be node1. Subsequently, based on the physical node number information of node1, the distributed node associated with virtual machine A will be determined, which is the same node node1 in the current scenario. This ensures that data replicas can be stored on the same or the nearest node to the virtual machine, reducing cross-node data transmission. Next, the previously established correspondence between the storage area and the distributed node will be invoked to find the target storage area corresponding to node1, referred to as zone1 here. Storage will use zone1 as the preferred storage location for the data replicas of target virtual machine A, thus implementing storage instructions for the data replicas. When virtual machine A needs to write data, its data replicas will be placed within zone1, ensuring that at least one replica is located on the node where the target virtual machine is located, to achieve localization of IO access. After setting the target storage area, locking the storage location of the data replicas in the area matching the node corresponding to the target virtual machine not only simplifies data management but also significantly reduces network latency and performance loss caused by cross-node access to data replicas. Especially in large data centers where the number of virtual machines (M) is much smaller than the number of nodes (N), such a mechanism can allocate storage resources more efficiently and improve overall performance.

[0039] Through the above process, the data replicas of the virtual machines are accurately matched with the storage areas of the physical nodes. This not only improves the management of data replicas in the HCI architecture but also greatly enhances the speed and stability of IO operations by reducing unnecessary network communication, which is a key step in achieving high-performance and low-latency storage in a hyper-converged environment.

[0040] In an exemplary embodiment, after setting the target storage area as the storage location of the target virtual machine data copy, the above method further includes: detecting the storage location setting results of the N storage areas; in the case where the setting results indicate that there is at least one storage area whose storage location is not set, marking the at least one storage area as an alternate area for the data copy; in the case where the setting results indicate that there is no storage area whose storage location is not set, sending a target message indicating that the storage configuration is valid to the management object corresponding to the storage system.

[0041] Optionally, if it is detected that there are storage areas that are not correctly configured, that is, there are storage areas not associated with specific nodes, these unconfigured areas will be marked as alternate areas. The role of the alternate areas is to quickly provide alternative storage space when the main storage area resources are fully loaded or a failure occurs, ensuring the storage security and continuity of the virtual machine data copy. This mechanism enhances the elasticity and flexibility of data management. When all N storage areas have been correctly set with storage locations, that is, a complete mapping with the nodes has been achieved, a confirmation message will be sent to the storage management component, indicating that the storage configuration is valid. This step confirms that the storage is ready to store and manage the data copy according to the IO localization strategy, marking the completion of the configuration phase and allowing entry into the normal operating state.

[0042] In summary, through the above implementation manner, it actively detects whether all N storage areas have been correctly set with storage locations according to the node information. Dynamically adjusts the storage layout according to the detection results, provides alternate storage areas to cope with discontinuous resource allocation or failures, and timely notifies the management component when the configuration is valid, providing guarantee for efficient operation and management.

[0043] In an exemplary embodiment, before determining the correspondence between the N storage areas and the N distributed nodes, the above method further includes: setting data transfer interfaces for the N storage areas and configuring the function logic of the data transfer interfaces; determining the first quantity of the first type of storage areas among the N storage areas for which data transfer interfaces have been set; determining the second quantity of the second type of storage areas among the N storage areas for which data transfer interfaces have not been set based on the first quantity; synchronizing the second quantity to the cluster file subsystem to instruct the cluster file subsystem to increase the data transfer interfaces according to the second quantity.

[0044] It can be understood that by setting data transfer interfaces in each storage area (N in number), these interfaces are responsible for the transfer of data replicas between different nodes. Configuring the functional logic of the data transfer interfaces means defining how the interfaces work, including rules for data transfer, priorities, encryption methods, etc., to ensure the efficiency and security of data during transfer. Then, based on the first quantity, the number of storage areas that have not been set with data transfer interfaces will be calculated, that is, the second quantity of the second type of storage area. The aim is to identify the configuration gap and provide a clear target for subsequent interface addition. Finally, the information of the second quantity will be synchronized to the cluster file subsystem, instructing the subsystem to add the corresponding number of data transfer interfaces according to this quantity. The cluster file subsystem dynamically adds interfaces according to the new requirements, which can effectively support the cross-node transfer requirements of data replicas. Especially when new storage areas require data replicas, the new data transfer interfaces will ensure that the data can be transferred to the target nodes quickly and accurately.

[0045] In summary, before configuring the correspondence between storage areas and nodes, the data transfer mechanism between the cluster file system and the SDS is established and optimized. By pre-setting data transfer interfaces and dynamically increasing the number of interfaces according to the actual configuration, the system can more effectively manage and optimize the distribution of data replicas between nodes and achieve the goal of IO localization.

[0046] In an exemplary embodiment, after synchronizing the second quantity to the cluster file subsystem to instruct the cluster file subsystem to increase the data transfer interfaces according to the second quantity, the above method further includes: associating the added data transfer interfaces with the second type of storage areas that have not been set with data transfer interfaces; in the case of completion of the association, according to the area parameter information of the second type of storage area and the physical node number information of the virtual machine corresponding to the area parameter information; supporting the data storage function of the virtual machine according to the area parameter information and the physical node number information.

[0047] In simple terms, after determining that the number of data transmission interfaces needs to be increased (i.e., the second number), the newly added interfaces are associated with the second-class storage areas that have not yet been configured with interfaces. This means that each newly added interface will be assigned to a specific storage area to support the data transmission needs of the area. This association process ensures that each storage area has a corresponding data transmission interface, so that data copies can be transmitted across nodes. Once the data transmission interface is associated with the second-class storage area, a more detailed configuration is performed based on the parameter information of the storage area and the number information of the virtual machine on the physical node. The regional parameter information generally includes the size and location of the storage area, and the SDS node information associated with it, while the physical node number information of the virtual machine is used to determine the specific physical node where the virtual machine is located. Finally, these configuration information will be used to support the data storage function of the virtual machine, ensuring that the virtual machine can read and write data based on the characteristics of the node and storage area where it is located. Specifically, when the virtual machine needs to store data, according to the physical node number of the virtual machine and the parameter information of the storage area, the data copy is preferentially stored in the storage area associated with the node where the virtual machine is located, so as to realize IO localization, reduce the delay of data transmission and the network burden, and thus improve data access performance. If the available space in the current storage area is insufficient, other associated or spare storage areas will be considered to ensure the storage requirements of the data.

[0048] In summary, through the above embodiments, the close cooperation between the data transmission interface and the storage area, as well as the reasonable matching between the storage area and the virtual machine physical node are ensured. Through these steps, not only the security and reliability of data storage are enhanced, but also the access efficiency of the virtual machine to the storage resources is maximized, and the data storage performance under the hyper-converged architecture is improved.

[0049] In an exemplary embodiment, the data storage function of the virtual machine is supported according to the area parameter information and the physical node number information, including: determining an unallocated storage area that meets the virtual machine disk space requirements through the area parameter information; binding the unallocated storage area to the physical node number information where the virtual machine is located, so as to store a copy of the data generated by the virtual machine in the unallocated storage area, and generate a corresponding data storage record.

[0050] It is understandable that when a virtual machine needs to create a new disk or expand the capacity of an existing disk, one or more currently unallocated storage areas are determined based on the requirements of the virtual machine (such as the required disk space size) and the parameter information of the storage area (such as the remaining space in the storage area and the partition layout). These areas have enough space to meet the disk requirements of the virtual machine. When selecting a storage area, priority is given to areas that match the physical node where the virtual machine is located to promote IO localization. Once a suitable unallocated storage area is found, a storage binding operation is performed, that is, the storage area is bound to the specific physical node where the virtual machine is located. This means that data copies generated by the virtual machine will be preferentially stored in this storage area related to its node. The binding process records the physical node number of the virtual machine and the corresponding storage area information to form a data storage record for subsequent tracking and optimization of data management operations. After the storage binding is completed, data copies generated by the virtual machine are directly stored in the already bound unallocated storage area, reducing network latency caused by cross-node data transmission. At the same time, a data storage record is automatically generated, which contains detailed information about the storage operation, such as the virtual machine ID storing the data, the location of the data copy (i.e., the bound storage area), and the timestamp of the storage operation, etc., which helps to maintain data consistency and quickly find or migrate data copies when needed.

[0051] In summary, through the above operations, the optimization of virtual machine data storage is effectively achieved, ensuring that data copies can be stored on the physical node closest to the virtual machine while meeting performance requirements, thereby reducing cross-node data communication and improving the efficiency of IO operations. At the same time, the generation of storage binding and data storage records provides convenience for the subsequent management and operation and maintenance of the system, ensuring the transparency and controllability of data management. This strategy is particularly critical in a hyper-converged architecture because it can fully utilize the advantages of software-defined storage (SDS), while taking into account the flexibility of the virtualized environment and the high performance of data access.

[0052] In an exemplary embodiment, after storing and indicating the data copies generated by the M virtual machines according to the corresponding relationship, the above method further includes: in the case of determining that a new virtual machine needs to be created or the storage size corresponding to the M virtual machines needs to be expanded, obtaining the remaining local storage size of the physical node corresponding to each of the M virtual machines; updating the configuration information of the disk space corresponding to the M virtual machines according to the storage size and the remaining local storage size.

[0053] It can be understood that when a new virtual machine needs to be deployed in the hyper-converged architecture or the storage requirements of an existing virtual machine change (such as an expansion of disk capacity), it is necessary to re-evaluate the allocation of storage resources to ensure that data replicas can continue to be stored in accordance with the principle of I / O localization. The local storage resource status of each physical node where the virtual machine is located is collected and analyzed. Specifically, the size of the remaining available storage space on each node is obtained. This information is crucial for the reasonable allocation of storage resources because it directly determines which nodes can be used to create new virtual machines or expand disk space. With the information on the remaining storage space of each node, the disk space configuration corresponding to each virtual machine can be dynamically updated based on this data and the storage requirements of the virtual machine. For example, when creating a new virtual machine, the system will select the best physical node based on the available storage space, then allocate disk space for it on that node, and at the same time update the configuration of the cluster file system to ensure that the data replicas of the new virtual machine are preferentially stored on the local node. If it is to expand the storage capacity of an existing virtual machine, by checking the remaining local storage size of the node where the virtual machine is located, if it is sufficient, it can be directly expanded; if it is insufficient, other nodes with sufficient remaining storage space need to be found, and the storage is expanded by migrating some data or adjusting the data replica distribution strategy, and at the same time the configuration information is updated to reflect the new storage layout.

[0054] In summary, through the above embodiments, it is ensured that even when the number of virtual machines or storage requirements change, the storage indication of data replicas can still follow the I / O localization strategy, thereby reducing cross-node data transmission and improving the efficiency of I / O operations. At the same time, the system updates the configuration based on the real-time storage resource status, can more flexibly respond to fluctuations in storage requirements, and ensures the high performance and high availability of the virtualized environment under the hyper-converged architecture.

[0055] In an exemplary embodiment, after storing and indicating the data replicas generated by the M virtual machines according to the corresponding relationship, the method further includes: when it is determined that at least one of the M virtual machines performs I / O access, determining whether the target data replica corresponding to the at least one virtual machine is located in the local storage area; when it is determined that the target data replica is located in the local storage area, prohibiting the migration process of the target data replica; when it is determined that the target data replica is not located in the local storage area, triggering the migration of the target data replica, where the migration is used to add at least one replica corresponding to the target data replica to the local storage area of the distributed node corresponding to the at least one virtual machine.

[0056] Optionally, when the system detects that any one or more of the M virtual machines are performing IO access operations, such as reading or writing data on the disk, it will start a series of checking and adjustment mechanisms to ensure that the data copy currently being accessed is located in the local storage area, that is, the storage area matching the physical node where the virtual machine is currently running. The system will check whether the target data copy (i.e., the data copy currently being accessed) has been stored in the local storage area of the node where the virtual machine is located. This is a key step to ensure high-efficiency IO operations, because the localization of data copies can reduce network transmission latency and improve IO access speed. If the target data copy is already in the local storage area, the system will not migrate it, that is, keep its current storage location unchanged. This strategy avoids unnecessary data migration operations, saves computing and network resources, and maintains the high performance state of IO access. On the contrary, if the target data copy is not in the local storage area, the system will trigger the data copy migration process. Specifically, the system will migrate one or more copies of the target data copy to the local storage area of the node where the virtual machine is located to achieve the localization of IO access. This operation is automatically completed through the mechanism of software-defined storage (SDS) without manual intervention, so as to ensure that the IO operations of the virtual machine can be performed on the local storage, reduce cross-node data transmission, and significantly improve the efficiency and performance of data access.

[0057] In summary, through the above embodiments, it is possible to continuously monitor and adjust the storage location of the virtual machine data copy to ensure that the IO operation is completed on the local storage as much as possible, thereby optimizing the storage performance under the hyper-converged architecture and improving the overall response speed and user experience of the system.

[0058] Among them, the execution subject of the above steps can be a server, a terminal, etc., but not limited thereto.

[0059] To facilitate the understanding of the embodiments of the present application, relevant scenarios are now explained, but they do not limit the present application.

[0060] In the related art, in the HCI scenario, there will be a problem of whether the IO crosses nodes when the virtual machine performs IO access. Specifically, when the virtual machine IO data copy is located on the current node: for the IO read operation, there is no need to forward the performance through the network, and its IO path is essentially the process of reading the local disk, with excellent performance; when performing the IO write operation, there must be a copy located on the current node, and only need to forward the data of other copies to other nodes, which is better than the scenario where the IO data copy is located on a non-current node in terms of both performance and network resource occupancy. When all the virtual machine IO data copies are located on other nodes: for the IO read operation, it needs to be read from other nodes through the network, which will be affected by network performance and the data forwarding process; when performing the IO write operation, all copies need to be forwarded to other nodes, with a large amount of data forwarding and high network resource occupancy.

[0061] To avoid the occurrence of the above problems, as an alternative implementation, an optional embodiment of the present application proposes an implementation method for supporting the localization of cluster file system I / O. By adding a region division function to the data allocation manager of the cluster file system, it supports dividing the data to-be-allocated space into the same number as the number of distributed cluster nodes during disk formatting. Secondly, the divided region information is transmitted to the SDS, and the SDS places one of the replicas on the corresponding node according to the region division. Thirdly, an interface is added to the data space allocator of the cluster file system to support passing in the node number information during space allocation. Finally, when creating a virtual disk or when virtual machine I / O writing triggers space allocation, the cluster file system calls the data allocation interface for space allocation, and preferentially allocates the disk space to the corresponding replica, thereby achieving I / O localization. The patent solution solves the problem that the cluster storage pool cannot achieve I / O localization in the HCI scenario by modifying the data space allocation strategy of the cluster file system and supporting the linkage between the data block allocation of the cluster file system and the replica data block allocation of the SDS.

[0062] As an alternative implementation, Figure 3 FIG. is a schematic diagram of an architecture for supporting the implementation of cluster file system I / O localization according to an embodiment of the present application, which includes the mapping relationship between the virtual disk data blocks in the cluster storage pool and the SDS underlying storage in a typical hyper-converged scenario, as follows: The disk of a virtual machine is essentially a file in the cluster storage pool. The Hypervisor usually provides the virtual disk file to the virtual machine in the form of a block device via the virtio-blk or virtio-scsi protocol. The backend of the cluster storage pool is usually a large-sized block device (hundreds of GB or several TB), and is usually provided by the SDS via the iscsi (Internet Small Computer Systems Interface) protocol in the HCI scenario. The SDS node runs on the physical node together with the virtualization system as a service component, and the SDS and the virtualization system perform data plane interaction via the iscsi protocol. The SDS takes over the physical disks on the node for saving data replicas, and usually uses high-performance disks such as NVMe (Non-Volatile Memory Express) and SSD as the cache layer, and uses HDD disks to store a large number of replica files. Note that since the actual distribution of the virtual disk data is managed by the SDS, the actual location of its data replicas may be on the local disks of any node in the cluster.

[0063] Specifically, Figure 3There are two virtual machines (the first virtual machine VM1 and the second virtual machine VM2) connected to the cluster storage pool in []. Data is transferred between the cluster storage pool and the SDS through a NIC (Network Interface Card), and the SDS mainly includes an SDS Core (software-defined storage core) and an iscsi target (target storage resource using the iscsi protocol). The SDC stores and processes the received data through a cache disk, an HDD disk, etc.

[0064] Optionally, Figure 4 is a schematic diagram of the relationship between virtual disk data and SDS replicas according to an embodiment of the present application. In Figure 4 virtual disk 1 (vdisk1) and virtual disk 2 (vdisk2) are virtual disks corresponding to two virtual machines located on different nodes respectively. The distribution of data blocks in the virtual disk is controlled by the file system data allocator, and then the data blocks are distributed in the CFS (Ceph File Syste, file system data allocator, abbreviated as CFS) storage pool space. In fact, the data of the two virtual disks may be mixed and distributed on the LUN (Logical Unit Number, abbreviated as LUN). The SDS will place a replica of the LUN on the local disk of any node. The figure shows a two-node scenario. For a multi-node scenario, its data distribution may not be on the first node node1 or the second node node2.

[0065] Optionally, Figure 5 is a schematic diagram of the change in the relationship between virtual disk data and SDS replicas according to an embodiment of the present application. First, an area division function is added to the data allocation manager of the cluster file system to support dividing the data to-be-allocated space into the same number as the number of distributed cluster nodes during disk formatting. Taking the figure as an example, when formatting, the available area corresponding to the CFS is evenly sliced into multiple allocation zones, and the number of zones is the same as the number of cluster nodes. The first allocation zone zone1 is the allocation zone for the first node node1, and the second allocation zone zone2 is the allocation zone for the second node node2.

[0066] Secondly, the divided area information is transmitted to the SDS, and the SDS places one of the replicas on the corresponding node according to the area division. The SDS provides an interface. After the CFS storage pool is divided, the zone information will be transmitted to the SDS system through the control plane of the HCI system. The SDS system will allocate a replica of an SDS LUN to the corresponding SDS node according to the zone division. As shown in the above figure, the actual replica distribution of CFS zone1 is on the local disk of the current node, i.e., SDS node1.

[0067] Secondly, add an interface to the data space allocator of the cluster file system to support passing in node number information during space allocation.

[0068] Optionally, to ensure the efficiency of the native interface for space allocation in the cluster file system, the interface generation control can be performed through the following code:

[0069] 。

[0070] It should be noted that in the above code, the input parameter only contains the information about the size of the space to be allocated. Add node information to it and during actual space allocation, allocate from the allocation zone where the specified node is located first. Note that when the space to be allocated in the current zone is insufficient, it will be allocated from other zones to ensure the availability of the function first. Since the space requested by virtual machines on each node cannot guarantee the same size, this strategy allows some nodes to request space from other zones when the space application quota is large.

[0071] Finally, when creating a virtual disk or when virtual machine IO writing triggers space allocation, the cluster file system will call the optimized allocation interface to allocate data blocks, and preferentially allocate disk space to the corresponding replicas of the current node. In this way, when the virtual machine is running, it can ensure that most virtual machine IOs on the current node are localized.

[0072] In summary, by modifying the data space allocation strategy of the cluster file system to support the linkage between data block allocation in the cluster file system and data block allocation in SDS replicas, the problem that the cluster storage pool cannot achieve IO localization in the HCI scenario is solved. Adopting this solution enables full utilization of the advantages of the cluster storage pool in the HCI scenario, such as features like HA, load balancing, and snapshot support, and also supports the SDS IO localization feature to achieve higher IO performance. This solution can significantly improve the overall performance of HCI products without increasing hardware costs, effectively enhancing the product competitiveness.

[0073] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0074] In this embodiment, a storage system for data replicas is further provided. This system is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0075] Figure 6 is a structural block diagram of a storage system for data replicas according to an embodiment of the present application. As Figure 6 shown, the system includes:

[0076] A cluster file subsystem 62, configured to determine N distributed nodes corresponding to M virtual machines supported by a hyper-converged architecture application; during the process of disk formatting, divide the storage space to be allocated associated with the hyper-converged architecture application according to the N distributed nodes, so as to equally divide the storage space to be allocated into N storage areas;

[0077] A software-defined storage subsystem 64, connected to the cluster file subsystem, configured to determine the correspondence between the N storage areas and the N distributed nodes; wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; and perform storage indication on the data replicas generated by the M virtual machines according to the correspondence.

[0078] Through the above system, first determine N distributed nodes corresponding to M virtual machines supported by the hyper-converged architecture application, and then during the process of disk formatting, divide the storage space to be allocated associated with the hyper-converged architecture application according to the N distributed nodes, so as to equally divide the storage space to be allocated into N storage areas, ensuring that the number of divided storage spaces is consistent with the number of distributed cluster nodes, and at the same time obtaining the correspondence between the N storage areas and the N distributed nodes. On this basis, instruct the software-defined storage subsystem to place one of the virtual machine replicas at the corresponding node according to this correspondence. Subsequently, the technical effect of efficient local storage of data replicas in software-defined storage is achieved. Through the above method, the storage space is finely managed, ensuring the matching of the distribution of data replicas with the nodes where the virtual machines are located, reducing cross-node data transmission, improving the IO access efficiency, and optimizing the storage performance without increasing additional hardware costs, enhancing the applicability and competitiveness of the hyper-converged architecture in a large-scale virtualization environment.

[0079] In an exemplary embodiment, the above software-defined storage subsystem is further configured to determine the physical node number information corresponding to the target virtual machine that generates the data copy; determine the target distributed node associated with the target virtual machine based on the physical node number information, and determine the target storage area corresponding to the target distributed node through the corresponding relationship; set the target storage area as the storage location of the data copy of the target virtual machine to indicate the storage of the data copy.

[0080] In an exemplary embodiment, after setting the target storage area as the storage location of the data copy of the target virtual machine, the above software-defined storage subsystem is further configured to detect the storage location setting results of the N storage areas; in the case where the setting results indicate that at least one storage area has no storage location set, mark the at least one storage area as an alternate area for the data copy; in the case where the setting results indicate that no storage area has no storage location set, send a target message indicating that the storage configuration is valid to the management object corresponding to the storage system.

[0081] In an exemplary embodiment, before determining the corresponding relationship between the N storage areas and the N distributed nodes, the above cluster file subsystem is further configured to set data transfer interfaces for the N storage areas and configure the function logic of the data transfer interfaces; determine the first quantity of the first type of storage areas among the N storage areas for which data transfer interfaces have been set; determine the second quantity of the second type of storage areas among the N storage areas for which data transfer interfaces have not been set based on the first quantity; synchronize the second quantity to the cluster file subsystem to instruct the cluster file subsystem to increase the data transfer interfaces according to the second quantity.

[0082] In an exemplary embodiment, after synchronizing the second quantity to the cluster file subsystem to instruct the cluster file subsystem to increase the data transfer interfaces according to the second quantity, the above cluster file subsystem is further configured to associate the increased data transfer interfaces with the second type of storage areas for which data transfer interfaces have not been set; in the case where the association is completed, according to the area parameter information of the second type of storage areas and the physical node number information of the virtual machines corresponding to the area parameter information; support the data storage function of the virtual machines according to the area parameter information and the physical node number information.

[0083] In an exemplary embodiment, the above cluster file subsystem is further configured to determine the unallocated storage areas that meet the virtual machine disk space requirements through the area parameter information; perform storage binding between the unallocated storage areas and the physical node number information where the virtual machines are located, so as to store the data copies generated by the virtual machines in the unallocated storage areas and generate corresponding data storage records.

[0084] In an exemplary embodiment, after the above-mentioned cluster file subsystem stores and indicates data replicas generated by the M virtual machines according to the corresponding relationship, when it is determined that a new virtual machine needs to be created or the storage size corresponding to the M virtual machines needs to be expanded, the remaining local storage size of the physical node corresponding to each of the M virtual machines is obtained; the configuration information of the disk space corresponding to the M virtual machines is updated according to the storage size and the remaining local storage size.

[0085] In an exemplary embodiment, after the above-mentioned cluster file subsystem stores and indicates data replicas generated by the M virtual machines according to the corresponding relationship, when it is determined that at least one of the M virtual machines performs an IO access, it is determined whether the target data replica corresponding to the at least one virtual machine is located in the local storage area; when it is determined that the target data replica is located in the local storage area, migration processing of the target data replica is prohibited; when it is determined that the target data replica is not located in the local storage area, migration of the target data replica is triggered, where the migration is used to add at least one replica corresponding to the target data replica to the local storage area of the distributed node corresponding to the at least one virtual machine.

[0086] In an exemplary embodiment, the above-mentioned cluster file subsystem includes: a data block allocation unit, configured to determine the storage space to be allocated for the M virtual machines supported by the hyper-converged architecture application, and evenly divide the storage space to be allocated into N allocable regions.

[0087] It should be noted that the data block allocation unit is a key component in the cluster file subsystem, and its main responsibility is to manage the storage space allocation strategy to meet the storage resource requirements of virtual machines. In the hyper-converged architecture application, this unit will:

[0088] Determine the storage space to be allocated: First, the data block allocation unit needs to determine how much storage space can be allocated to the M virtual machines. This usually involves an assessment of the current storage resources, including the available storage capacity on all nodes and the existing data distribution.

[0089] Evenly divide into N allocable regions: After determining the total storage space to be allocated, the data block allocation unit evenly divides these spaces into N allocable regions, where N is equal to the number of nodes in the distributed cluster. Each region will be assigned to a specific node to support subsequent data block allocation and storage operations, thereby promoting local storage of data.

[0090] By evenly dividing the storage space into regions equal in number to the nodes, the data block allocation unit can ensure that each node has a fair allocation of storage resources, and at the same time provides a basis for implementing local storage of data replicas. When a virtual machine generates a data replica, the system can preferentially store the replica in the allocable region corresponding to the node where the virtual machine is located, reducing cross-node data transmission and improving the performance of IO operations. The implementation of this strategy requires close cooperation between the cluster file subsystem and software-defined storage (SDS). By passing region division information and node information, SDS can adjust the distribution of its data replicas based on this information to support IO localization.

[0091] In summary, the role of the data block allocation unit in the hyper-converged architecture is crucial. It not only manages the allocation of storage space but also provides technical support for achieving high efficiency and low latency in virtual machine data access by dividing the storage space into allocable regions that match the number of nodes. This mechanism helps maintain the high-performance state of the system while ensuring the flexible use of storage resources by virtual machines.

[0092] In an exemplary embodiment, the above-mentioned cluster file subsystem includes: an interface addition unit for adjusting the number of interfaces of the cluster file subsystem according to the physical node number information of M virtual machines passed in during space allocation.

[0093] The Cluster File System (CFS) plays a core role in the hyper-converged architecture. It is responsible for managing the data block allocation and file operations in the storage pool. In traditional implementation methods, the number of interfaces of CFS is usually fixed and related to the configuration during system initialization. However, with the increase or change in the number of virtual machines and the rising demand for IO localization, the original interface configuration may no longer meet the requirements of high-performance storage access. To solve this problem, the "interface increment unit" in this embodiment plays a crucial role: by adjusting the number of interfaces based on the physical node number information. That is, when the system is about to allocate space for M virtual machines, the interface increment unit will receive the number information representing the physical nodes where the virtual machines are located. Based on this information, it can dynamically adjust the number of interfaces of the cluster file system to ensure that the virtual machines on each node can effectively utilize the local storage resources. Secondly, it can also optimize the data access path: by adding interfaces that match the number of nodes, the system can route data requests more directly to the correct physical nodes, reducing the unnecessary transmission of data requests between nodes, thus significantly improving the efficiency and performance of data access. That is to say, the interface increment unit can ensure that when a virtual machine requests storage space, the corresponding data block allocation operation will give priority to the local storage resources of the node, which is closely related to the IO localization strategy and can effectively reduce the cross-node data transmission delay and improve the response speed of IO operations. Further, the dynamic adjustment ability of this interface increment unit also allows the system to automatically adjust the number of interfaces when nodes are added or removed to adapt to the dynamic changes of nodes in the hyper-converged architecture, ensuring the high availability and scalability of the system.

[0094] In summary, through the implementation of the interface increment unit function, the cluster file system can more flexibly handle the changes of virtual machines and physical nodes. By increasing the number of interfaces that match the nodes, optimizing the data access path, and supporting the IO localization strategy, it can ensure the fast access of virtual machines to storage resources while maintaining the high efficiency and performance of the storage pool. This dynamic adjustment and optimization strategy is the key to realizing intelligent storage management in the hyper-converged architecture and helps to improve the overall system performance and response ability.

[0095] In an exemplary embodiment, the above storage system further includes: a policy component for obtaining in real time the storage resources and load data corresponding to each of the N distributed nodes, and adjusting the splitting size of the storage space to be allocated after formatting according to the storage resources and the load data.

[0096] It can be understood that the policy component plays the role of a decision maker in the storage system. Its core task is to optimize the space splitting strategy by considering the real-time state of each distributed node during the storage space allocation process. The specific functions are as follows:

[0097] (1) Real-time resource and load monitoring: The policy component periodically or proactively obtains the storage resource usage (such as remaining disk space, cache utilization rate, etc.) and load data (such as CPU utilization rate, network bandwidth usage, etc.) of N distributed nodes. This information can reflect the current processing capacity and storage capacity of the nodes.

[0098] (2) Dynamically adjust the segmentation size: Based on the collected resource and load data, the policy component can intelligently adjust the segmentation size of the storage space to be allocated. For example, when the storage resources of some nodes are in short supply while there is still a large amount of free space on other nodes, the policy component can increase the segmentation ratio of the nodes with rich free resources and reduce the segmentation burden of the nodes with resource shortages.

[0099] (3) Optimize space allocation: The dynamic adjustment of the segmentation size helps to optimize the allocation of storage space. The policy component can ensure that data copies are preferentially stored on nodes with sufficient resources and low load, thereby improving the efficiency of IO operations and reducing performance degradation caused by node overload.

[0100] (4) Promote load balancing: Through the intervention of the policy component, the segmentation of the storage space is no longer a simple average distribution, but is adjusted according to the actual resources and load conditions of the nodes. This dynamic segmentation strategy helps to achieve load balancing of storage resources and avoid a single node becoming a performance bottleneck.

[0101] (5) Support dynamic adjustment and fault recovery: The intelligent feature of the policy component is also reflected in its ability to quickly respond to dynamic changes in the storage system. For example, in the case of node failure or sudden increase in load, it can automatically adjust the segmentation strategy to ensure that the reliability and access performance of the data are not affected.

[0102] In summary, through the real-time monitoring and intelligent decision-making of the policy component, the storage system can better adapt to the continuous changes of resources and loads in the distributed environment, ensuring that the segmentation of the storage space takes into account both data security redundancy and optimal performance. This strategic adjustment is particularly important for the hyper-converged architecture because it can maximize the utilization of existing resources, improve the overall storage efficiency and data access speed of the system, while reducing the complexity of management and maintenance.

[0103] In an exemplary embodiment, the above cluster file subsystem further includes: a data space allocation unit for preferentially allocating the disk space in the storage system to the target distributed node corresponding to the target data copy on the basis of determining that the target data copy has been generated.

[0104] That is to say, the data space allocation unit is a module in the cluster file subsystem responsible for data storage space allocation. Its main task is to ensure that the disk space allocation policy in the storage system gives priority to the local storage resources of the distributed nodes corresponding to the target data replicas after the generation of the target data replicas. Specifically: The data space allocation unit first needs to identify which data replicas are target replicas, that is, the data sets that need to prioritize local storage. This is usually obtained based on the access patterns of virtual machines, data access frequencies, and the hot and cold states of the data. After identifying the target data replicas, the data space allocation unit will further analyze the storage locations of the target replicas to determine which replicas should be preferentially stored on the local distributed nodes to reduce cross-node IO transfers.

[0105] When allocating storage space, the data space allocation unit will give priority to the local storage resources of the target distributed nodes. This means that when a virtual machine requests additional storage space or writes data, the new data blocks will first try to be saved on the local disk of the node, rather than being forwarded to other nodes through the network. The linkage between the data space allocation unit and the data replica layout enables the allocation of storage space to no longer be independent of the replica distribution, but to be dynamically adjusted according to the replica layout strategy, ensuring the local storage and access of data. The priority allocation strategy of the data space allocation unit directly improves the IO performance because the local storage of data blocks reduces the distance and time of data transmission, improving the speed and efficiency of data access.

[0106] Through the above mechanism, the data space allocation unit can ensure the efficient storage and access of virtual machine data, while supporting the local storage strategy of data replicas under the hyper-converged architecture. The introduction of such a unit is of great significance for optimizing the resource utilization of the storage system, improving the IO performance, and enhancing the overall reliability and availability of the system.

[0107] In an exemplary embodiment, the above software-defined storage subsystem further includes: a configuration unit for placing at least one replica of the same logical block on the local disk of a specific node according to the area parameters corresponding to the N storage areas, where the specific node matches the node distribution of the N distributed nodes of the M virtual machines in the storage pool.

[0108] The configuration unit receives N storage area partitioning information from the cluster file system, and each area corresponds to a specific node in the virtual machine distribution. These area parameters contain important details about how to distribute data among nodes, including storage capacity, data distribution strategies, etc. Based on the received area parameters, the configuration unit determines how to place data replicas. For any given logical block (i.e., a data block in the virtual machine file system), the configuration unit ensures that at least one replica is placed on the node corresponding to the virtual machine distribution, which is called the "specific node". When selecting the specific node for replica placement, the configuration unit will give priority to using the local disk of that node for storage rather than the disks of remote nodes. This reduces the network latency of data transmission and enhances the performance and efficiency of data access. The replica placement strategy of the configuration unit matches the node distribution of the N distributed nodes of the M virtual machines in the storage pool, ensuring that regardless of how the virtual machines migrate or redistribute, data access can be kept local, further optimizing the IO performance. This replica placement method based on node distribution also supports dynamic adjustment. When the virtual machine distribution changes, the configuration unit can rearrange the replica placement to adapt to the new node distribution and maintain the efficiency and locality of data access.

[0109] Through the functions of the configuration unit, SDS can effectively support the optimization of data replica layout in the hyper-converged architecture, ensure the efficient utilization of storage resources, and at the same time improve the IO access speed of virtual machines and the storage performance of the entire system by reducing the data transmission distance. This intelligent data management strategy is an important application of software-defined storage technology in the hyper-converged environment, which helps to achieve higher resource utilization efficiency and service levels.

[0110] In an exemplary embodiment, the above software-defined storage subsystem further includes: a replication management unit, configured to monitor the working state of the data replicas; in the case where the working state indicates that a data replica is lost or faulty, perform replica recovery through other distributed nodes except the current distributed node where the data replica is located; in the case where the working state indicates that the data replica is not lost or faulty, determine that the storage state of the data replica is normal and no other processing operations need to be performed.

[0111] In an exemplary embodiment, the above software-defined storage subsystem further includes: a cache management unit, configured to store the first type of data blocks with an access frequency greater than or equal to a preset frequency of the M virtual machines in the cache devices corresponding to the N distributed nodes, and store the second type of data blocks with an access frequency less than the preset frequency of the M virtual machines in the non-cache devices corresponding to the N distributed nodes.

[0112] In an exemplary embodiment, the above software-defined storage subsystem further includes: an interface generation unit, configured to receive the number of interfaces sent by the cluster file subsystem when it is determined that the division of the cluster file subsystem is completed, and generate a new data synchronization interface between the software-defined storage subsystem and the cluster file subsystem according to the number of interfaces.

[0113] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above-mentioned modules are all located in the same target processor; or, the above-mentioned modules are respectively located in different target processors in any combination form.

[0114] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0115] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks, etc., all kinds of media that can store computer programs.

[0116] An embodiment of the present application also provides an electronic device, including a target memory and a target processor. A computer program is stored in the target memory, and the target processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0117] Optionally, Figure 7 is a block diagram of the computer system structure of the electronic device according to the embodiment of the present application. As Figure 7 shown, the computer system 800 includes a CPU 801 (Central Processing Unit, central processing unit, abbreviated as CPU), which can execute various appropriate actions and processes according to the program stored in the ROM 802 (Read-Only Memory, read-only memory, abbreviated as ROM) or the program loaded from the storage part 808 into the RAM 803 (Random Access Memory, random access memory, abbreviated as RAM). In the RAM 803, various programs and data required for system operation are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other through a bus 804. The I / O interface 805 (Input / Output interface, input / output interface, i.e., I / O interface) is also connected to the bus 804.

[0118] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as required. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 810 as required so that a computer program read therefrom is installed into the storage section 808 as required.

[0119] In an exemplary embodiment, the above-mentioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above-mentioned central processing unit (CPU 801), and the input / output device is connected to the above-mentioned central processing unit (CPU 801).

[0120] An embodiment of the present application further provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.

[0121] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, where the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.

[0122] An embodiment of the present application further provides a computer program, the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.

[0123] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be elaborated herein.

[0124] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0125] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of this application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, this application is not limited to any specific combination of hardware and software.

[0126] The above has introduced in detail a storage system, method, medium, electronic device, and program product for a data copy provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for storing data replicas, characterized in that, Including: Determine N distributed nodes corresponding to M virtual machines supported by the hyper-converged architecture application; During the process of disk formatting, divide the storage space to be allocated associated with the hyper-converged architecture application according to the N distributed nodes, so as to equally divide the storage space to be allocated into N storage areas; Determine the corresponding relationship between the N storage areas and the N distributed nodes; wherein, each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; Perform storage indication on the data replicas generated by the M virtual machines according to the corresponding relationship.

2. The storage method of the data copy according to claim 1, characterized in that Performing storage indication on the data replicas generated by the M virtual machines according to the corresponding relationship includes: Determine the physical node number information corresponding to the target virtual machine that generates the data replica; Based on the physical node number information, determine the target distributed node associated with the target virtual machine, and determine the target storage area corresponding to the target distributed node through the corresponding relationship; Set the target storage area as the storage location of the data replica of the target virtual machine, so as to perform storage indication on the data replica.

3. The storage method of the data copy according to claim 2, wherein After setting the target storage area as the storage location of the data replica of the target virtual machine, the method further includes: Detect the storage location setting results of the N storage areas; In the case that the setting result indicates that there is at least one storage area without a set storage location, mark the at least one storage area as an alternate area for the data replica; In the case that the setting result indicates that there is no storage area without a set storage location, send a target message indicating that the storage configuration is valid to the management object corresponding to the storage system.

4. The storage method of the data copy according to claim 1, wherein Before determining the corresponding relationship between the N storage areas and the N distributed nodes, the method further includes: Set data transfer interfaces for the N storage areas and configure the function logic of the data transfer interfaces; Determine the first quantity of the first type of storage areas in the N storage areas that have set data transfer interfaces; Based on the first quantity, determine the second quantity of the second type of storage areas in the N storage areas that have not set data transfer interfaces; Synchronize the second quantity to the cluster file subsystem to instruct the cluster file subsystem to increase the data transfer interfaces according to the second quantity.

5. The storage method of the data copy according to claim 4, characterized in that, After synchronizing the second quantity to the cluster file subsystem to instruct the cluster file subsystem to increase the data transfer interfaces according to the second quantity, the method further includes: Associate the increased data transfer interfaces with the second type of storage areas that have not set data transfer interfaces; In the case of completion of the association, according to the area parameter information of the second type of storage areas and the physical node number information of the virtual machines corresponding to the area parameter information; Support the data storage function of the virtual machine according to the area parameter information and the physical node number information.

6. The storage method of the data copy according to claim 5, wherein Supporting the data storage function of the virtual machine according to the area parameter information and the physical node number information includes: Determine the unallocated storage areas that meet the disk space requirements of the virtual machine through the area parameter information; Bind the unallocated storage area to the physical node number information where the virtual machine is located, so as to store the data copy generated by the virtual machine in the unallocated storage area and generate a corresponding data storage record.

7. The storage method of the data copy according to claim 1, wherein After storing and indicating the data copies generated by the M virtual machines according to the corresponding relationship, the method further includes: When it is determined that a new virtual machine needs to be created or the storage size corresponding to the M virtual machines needs to be expanded, obtain the remaining local storage size of the physical node corresponding to each of the M virtual machines; Update the configuration information of the disk space corresponding to the M virtual machines according to the storage size and the remaining local storage size.

8. The storage method of the data copy according to claim 1, wherein After storing and indicating the data copies generated by the M virtual machines according to the corresponding relationship, the method further includes: When it is determined that at least one of the M virtual machines performs IO access, determine whether the target data copy corresponding to the at least one virtual machine is located in the local storage area; When it is determined that the target data copy is located in the local storage area, prohibit the migration process of the target data copy; When it is determined that the target data copy is not located in the local storage area, trigger the migration of the target data copy, where the migration is used to add at least one copy corresponding to the target data copy to the local storage area of the distributed node corresponding to the at least one virtual machine.

9. A storage system for data replicas, characterized in that, Includes: A cluster file subsystem, used to determine N distributed nodes corresponding to M virtual machines supported by the hyper-converged architecture application; During the disk formatting process, divide the to-be-allocated storage space associated with the hyper-converged architecture application according to the N distributed nodes, so as to equally divide the to-be-allocated storage space into N storage areas; A software-defined storage subsystem, connected to the cluster file subsystem, used to determine the corresponding relationship between the N storage areas and the N distributed nodes; where each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; store and indicate the data copies generated by the M virtual machines according to the corresponding relationship.

10. The storage system for data copies according to claim 9, characterized in that, The cluster file subsystem includes: a data block allocation unit, used to determine the to-be-allocated storage space for the M virtual machines supported by the hyper-converged architecture application and evenly divide the to-be-allocated storage space into N allocable areas.

11. The storage system for data copies according to claim 9, wherein The cluster file subsystem includes: an interface addition unit, used to adjust the number of interfaces of the cluster file subsystem according to the physical node number information of the M virtual machines passed in during space allocation.

12. The storage system for data copies according to claim 9, wherein, The storage system further includes: a policy component, used to obtain the storage resources and load data corresponding to each of the N distributed nodes in real time, and adjust the split size of the to-be-allocated storage space after formatting according to the storage resources and the load data.

13. The storage system for data copies according to claim 9, wherein The cluster file subsystem further includes: a data space allocation unit, used to preferentially allocate the disk space in the storage system to the target distributed node corresponding to the target data copy on the basis of determining that the target data copy has been generated.

14. The storage system for data copies according to claim 9, characterized in that, The software-defined storage subsystem further includes: a configuration unit, configured to place at least one copy of the same logical block on the local disk of a specific node according to the area parameters corresponding to the N storage areas, where the specific node matches the node distribution of the N distributed nodes of the M virtual machines in the storage pool.

15. The storage system for data copies according to claim 9, wherein The software-defined storage subsystem further includes: a replication management unit, configured to monitor the working state of the data replicas; in the case that the working state indicates that the data replicas are lost or faulty, perform replica recovery through other distributed nodes except the current distributed node where the data replicas are located; in the case that the working state indicates that the data replicas are not lost or faulty, determine that the storage state of the data replicas is normal and no other processing operations need to be performed.

16. The storage system for data copies according to claim 9, wherein The software-defined storage subsystem further includes: a cache management unit, configured to store the first type of data blocks with an access frequency greater than or equal to a preset frequency of the M virtual machines in the cache devices corresponding to the N distributed nodes, and store the second type of data blocks with an access frequency less than the preset frequency of the M virtual machines in the non-cache devices corresponding to the N distributed nodes.

17. The storage system for data copies according to claim 9, characterized in that, The software-defined storage subsystem further includes: an interface generation unit, configured to, when determining that the cluster file subsystem has completed partitioning, receive the number of interfaces sent by the cluster file subsystem, and generate new data synchronization interfaces between the software-defined storage subsystem and the cluster file subsystem according to the number of interfaces.

18. An electronic device, characterized in that, Comprising: a memory, configured to store a computer program; a processor, configured to implement the steps of the method for storing data replicas as described in any one of claims 1 to 8 when executing the computer program.

19. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program, when executed by a processor, implements the steps of the method for storing data replicas as described in any one of claims 1 to 8.

20. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method for storing data replicas as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Device for hyper-converged infrastructure

    CN108228087A

  • Near data processing system for hyper-converged equipment

    CN111880739A

  • Deployment method and system applied to hyper-converged architecture

    CN112527325A

  • Resource management system and method for CPUs of virtual machine in cloud platform

    CN113986454A

  • Data distribution method and system of hyper-converged system

    CN114003350A

Cited By

  • Data synchronization method and device, electronic equipment, storage medium and program product

    CN120744010A