Data copy storage method, system, medium, electronic device, and program product
By dividing the storage space into storage areas equal to the number of nodes in the hyper-converged architecture and instructing data copies to be stored on the corresponding nodes, the problem of the cluster storage pool being unable to achieve IO localization in the HCI scenario is solved, and IO access efficiency and storage performance are improved.
Patent Information
- Application Number
- CN202510744585.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-06-05
AI Technical Summary
In HCI scenarios, cluster storage pools cannot achieve IO localization, resulting in low IO access efficiency.
By determining the distributed nodes corresponding to the virtual machines in the hyper-converged architecture and dividing the storage space into storage areas equal to the number of nodes when formatting the disk, a correspondence between the storage areas and the nodes is established, and the software-defined storage subsystem is instructed to store data copies on the corresponding nodes, efficient localization of data copies is achieved.
It improves IO access efficiency, optimizes storage performance, and enhances the applicability and competitiveness of hyper-converged architecture in large-scale virtualization environments without increasing additional hardware costs.
Smart Images

Figure CN120255827B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of server hyper-converged technology, in particular, to a data copy storage method and system, a medium, an electronic device and a program product. BACKGROUND
[0002] Hyper Converged Infrastructure (HCI) integrates virtualization and Software Defined Storage (SDS) to optimize IT infrastructure. SDS enhances data reliability, availability and scalability by distributing and replicating data across multiple hosts, surpassing traditional centralized storage. SDS provides block, object and file services to meet diverse application requirements. Virtualization technology improves hardware resource utilization efficiency, enhances system security and service availability, and is an important user of SDS. Both of them run on the same node in HCI, and data is directly stored on the local disk. In addition, data copies are usually saved in the form of files or data blocks on the disk. For hardware scenarios where the back-end disk is an HDD (Hard Disk Drive, HDD), there will inevitably be performance bottleneck problems.
[0003] In related technologies, performance optimization is often achieved by introducing an SSD (Solid State Drive, SSD) cache disk. However, the back-end storage of traditional server virtualization systems is usually provided by centralized storage. In HCI systems, virtualization services and software-defined storage are deployed on the same physical node, and data copies are finally saved on the local disk of the physical node. When IO (Input / Output, IO) access is performed through a virtual machine, there is a problem of whether the IO access is cross-node. Therefore, related storage pool virtualization solutions cannot achieve IO localization in HCI scenarios.
[0004] At present, there is no effective solution to the problem that the cluster storage pool cannot achieve IO localization in HCI scenarios in related technologies. SUMMARY
[0005] Embodiments of the present application provide a data copy storage method and system, a medium, an electronic device and a program product to at least solve the problem that the cluster storage pool cannot achieve IO localization in HCI scenarios in related technologies.
[0006] According to one embodiment of the present application, a data copy storage method is provided, comprising: determining N distributed nodes corresponding to M virtual machines supported by a super-converged architecture application; in a process of disk formatting, dividing a to-be-allocated storage space associated with the super-converged architecture application according to the N distributed nodes to equally divide the to-be-allocated storage space into N storage areas; determining a correspondence between the N storage areas and the N distributed nodes; wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; and storing data copies generated by the M virtual machines according to the correspondence.
[0007] According to another embodiment of the present application, a data copy storage system is provided, comprising: a cluster file subsystem configured to determine N distributed nodes corresponding to M virtual machines supported by a super-converged architecture application; in a process of disk formatting, divide a to-be-allocated storage space associated with the super-converged architecture application according to the N distributed nodes to equally divide the to-be-allocated storage space into N storage areas; and a software-defined storage subsystem connected to the cluster file subsystem and configured to determine a correspondence between the N storage areas and the N distributed nodes; wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; and store data copies generated by the M virtual machines according to the correspondence.
[0008] According to still another embodiment of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, wherein the computer program is configured to execute the steps in any of the method embodiments described above when running.
[0009] According to still another embodiment of the present application, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the method embodiments described above.
[0010] According to still another embodiment of the present application, a computer program product is provided, comprising a computer program, and the computer program is executed by a processor to implement the steps in any of the method embodiments described above.
[0011] Through the present application, firstly, N distributed nodes corresponding to M virtual machines supported by the super-converged architecture application are determined, and then in the process of disk formatting, the to-be-allocated storage space associated with the super-converged architecture application is divided according to the N distributed nodes, so as to divide the to-be-allocated storage space into N storage areas equally, so as to ensure that the number of to-be-allocated space division is consistent with the number of distributed cluster nodes, and the corresponding relationship between the N storage areas and the N distributed nodes is obtained, and on this basis, the software-defined storage subsystem is instructed to place one of the virtual machine copies to the corresponding node according to the corresponding relationship. Then, the technical effect of efficient local storage of data copies in the HCI scene is realized. Through the above method, the storage space is finely managed, the distribution of data copies is matched with the nodes where the virtual machines are located, the data transmission across nodes is reduced, the IO access efficiency is improved, the storage performance is optimized without increasing the additional hardware cost, and the applicability and competitiveness of the super-converged architecture in the large-scale virtualization environment are enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0013] Figure 1 is a hardware structure block diagram of a server device of a data copy storage method according to an embodiment of the present application;
[0014] Figure 2 is a flowchart of a data copy storage method according to an embodiment of the present application;
[0015] Figure 3 is an architecture schematic diagram of supporting cluster file system IO localization implementation according to an embodiment of the present application;
[0016] Figure 4 is a schematic diagram of a link resource organization structure according to an embodiment of the present application;
[0017] Figure 5 is a change schematic diagram of virtual disk data and SDS copy relationship according to an embodiment of the present application;
[0018] Figure 6 is a structure block diagram of a data copy storage system according to an embodiment of the present application;
[0019] Figure 7 is a computer system structure block diagram of an electronic device according to an embodiment of the present application;
[0020] In the above figure, 102 is a processor, 104 is a memory, 106 is a transmission device, 108 is an input and output device, 62 is a cluster file subsystem, 64 is a software-defined storage subsystem, 800 is a computer system, 801 is a CPU, 802 is a ROM, 803 is a RAM, 804 is a bus, 805 is an I / O interface, 806 is an input part, 807 is an output part, 808 is a storage part, 809 is a communication part, 810 is a driver, and 811 is a removable medium. DETAILED DESCRIPTION
[0021] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.
[0022] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0023] In order for those skilled in the art to better understand the technical solutions of the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0024] As an optional implementation, the method embodiments provided in the embodiments of the present application can be executed in a server device or similar computing device. Taking the case of running on a server device, Figure 1 is a hardware structure block diagram of a server device of a data copy storage method according to an embodiment of the present application. As shown in Figure 1 , the server device can include one or more (only one is shown in Figure 1 ) processor 102 (the processor 102 can include but is not limited to a microcontroller unit (MCU) or a field-programmable gate array (FPGA) processing device) and a memory 104 for storing data, wherein the above-mentioned server device can further include a transmission device 106 for communication function and an input and output device 108. Those skilled in the art can understand that Figure 1The illustrated structure is merely schematic and does not limit the structure of the server device. For example, the server device can further include more or less components than those shown, or have a different configuration of components than those shown. Figure 1 The illustrated structure is merely schematic and does not limit the structure of the server device. For example, the server device can further include more or less components than those shown, or have a different configuration of components than those shown. Figure 1 The illustrated structure is merely schematic and does not limit the structure of the server device. For example, the server device can further include more or less components than those shown, or have a different configuration of components than those shown.
[0025] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as a computer program corresponding to the data copy storage method of the embodiments of the present application. The processor 102 can execute various functional applications and data processing by running the computer programs stored in the memory 104, i.e., implement the method described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to the server device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0026] The transmission device 106 is used to receive or send data via a network. The specific examples of the network can include a wireless network provided by a communication provider of the server device. In one example, the transmission device 106 includes a network adapter (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet in a wireless manner.
[0027] In the embodiments, a data copy storage method is provided, Figure 2 The flowchart of the data copy storage method according to the embodiments of the present application is shown in FIG. 2, which includes the following steps: Figure 2
[0028] In step S202, N distributed nodes corresponding to M virtual machines supported by the hyper-converged architecture application are determined.
[0029] It can be understood that in the hyper-converged architecture (HCI), a plurality of virtual machines (M) run on a distributed cluster composed of a plurality of physical nodes. Each physical node is both a host of a virtualization environment and a node of software-defined storage (SDS), which collectively provides computing and storage resources. Through the above steps, the number of active virtual machines (M) in the current HCI system and the total number of physical nodes (i.e., distributed nodes) supporting the running of these virtual machines (N) can be identified and determined. The M virtual machines can be any number, and the N nodes are the physical basis of the hyper-converged architecture, and there is a relationship between them that M is less than or equal to N, which means that each virtual machine corresponds to at least one distributed node, but the number of nodes may exceed the number of virtual machines, providing protection for system expansion and redundancy. This determination process is a prerequisite for the implementation of the subsequent IO localization strategy, ensuring that data can be optimally distributed to specific nodes.
[0030] In step S204, during the disk formatting process, the N distributed nodes are applied to the associated to-be-allocated storage space to divide the to-be-allocated storage space into N storage areas.
[0031] It can be understood that in the traditional HCI architecture, the file system data of the virtual machine can be distributed in an unordered manner on each node of the cluster, resulting in frequent cross-node access during IO operations and limited performance. In order to overcome this challenge, a region division function is introduced in the disk formatting phase, which divides the to-be-allocated storage space into N storage areas (allocation zones), each of which corresponds to a specific distributed node. For example, when there are two nodes in the system, the available space in the storage pool will be divided into two regions, the first storage region zone1 corresponds to the first node node1, and the second storage region zone2 corresponds to the second node node2. This division ensures that when the virtual machine on each node performs data read / write operations, it can preferentially access the data copy located on the local disk, significantly reducing network transmission overhead and improving IO performance. At the same time, the region division information is passed to the SDS, further guiding the SDS when creating data copies to ensure that at least one copy is located on the corresponding node, matching the distribution strategy of the virtual machine and optimizing the overall system performance.
[0032] In step S206, the correspondence between the N storage areas and the N distributed nodes is determined; wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers.
[0033] Optionally, by defining the binding rules of each allocation zone and specific node, the subsequent data copy allocation provides clear guidance. For example, assuming N = 2, that is, there are two nodes, then the storage space will also be divided into two zones (zone1 and zone2), zone1 will correspond to node1, and zone2 will correspond to node2. This correspondence ensures that data is distributed in order between different nodes, rather than randomly or blindly scattered throughout the cluster, laying the foundation for data localization.
[0034] Step S208, according to the correspondence, the data copy generated by the M virtual machines is stored.
[0035] It can be understood that during the IO operation of the virtual machine, the storage location of the data copy is indicated according to the correspondence. This means that when the virtual machine (M) performs read and write operations, the data copy generated by the virtual machine will be guided to be stored in the storage area corresponding to the node where the virtual machine is located, so as to realize the localization of IO access. Specifically, when the operation of creating a virtual disk or having a virtual machine IO write is triggered, the cluster file system will call the optimized space allocation interface, and preferentially allocate disk space to the data copy located on the target node. For example, if the virtual machine A is located on node1, then its data copy will be stored on the physical disk corresponding to zone1 as much as possible, rather than distributed on any other node in the network. In this way, cross-node data transmission is minimized, improving the efficiency and performance of IO operations, especially when reading data within the node, which is almost not affected by network delay.
[0036] Through the above method, first, the N distributed nodes corresponding to the M virtual machines supported by the super-converged architecture application are determined, and then in the process of disk formatting, the N distributed nodes are used to divide the to-be-allocated storage space associated with the super-converged architecture application, so as to divide the to-be-allocated storage space into N storage areas, ensure that the number of to-be-allocated space division is consistent with the number of distributed cluster nodes, and obtain the correspondence between the N storage areas and the N distributed nodes. On this basis, the software-defined storage subsystem is instructed to place one of the virtual machine copies to the corresponding node according to the correspondence. Then, the efficient localization storage technology effect of the data copy in the software-defined storage is realized. Through the above method of fine management of the storage space, the distribution of the data copy is matched with the node where the virtual machine is located, the cross-node data transmission is reduced, the IO access efficiency is improved, the storage performance is optimized without increasing additional hardware cost, and the applicability and competitiveness of the super-converged architecture in a large-scale virtualization environment are enhanced.
[0037] In one example embodiment, the storage indication of the data copy generated by the M virtual machines according to the correspondence relationship comprises: determining the physical node number information corresponding to the target virtual machine generating the data copy; determining the target distributed node associated with the target virtual machine based on the physical node number information, and determining the target storage area corresponding to the target distributed node through the correspondence relationship; setting the target storage area as the storage location of the data copy of the target virtual machine to indicate the storage of the data copy.
[0038] Optionally, the above-mentioned data copy storage indication mechanism mainly focuses on determining the specific physical node number information of the "target virtual machine" generating the data copy. Simply put, since the virtual machine can span multiple nodes in the hyper-converged architecture, it is crucial to determine the physical node where it resides, which is the basis for realizing data copy localization. Specifically, when virtual machine A generates a data write operation, it first identifies the physical node where A resides, which is assumed to be node1. Subsequently, based on the physical node number information of node1, the distributed node associated with virtual machine A is determined, which is the same node node1 in the current scenario. Thus, it is ensured that the data copy can be stored on the same or nearest node as the virtual machine to reduce cross-node data transmission. Next, the correspondence relationship between the storage area and the distributed node established earlier is called to find the target storage area corresponding to node1, which is referred to as zone1 here. The storage takes zone1 as the preferred storage location of the data copy of the target virtual machine A, thereby realizing the storage indication of the data copy. When virtual machine A needs to write data, its data copy will be placed in zone1, ensuring that at least one copy is located on the node where the target virtual machine is located to realize the localization of IO access. After setting the target storage area, the storage location of the data copy is locked in the area matching the node corresponding to the target virtual machine, which not only simplifies data management, but also greatly reduces network delay and performance loss caused by cross-node access of data copies. Especially in large data centers, the number of virtual machines (M) is much smaller than the number of nodes (N), and such a mechanism can more efficiently allocate storage resources and improve overall performance.
[0039] Through the above process, the data copy of the virtual machine is precisely matched with the storage area of the physical node. This not only improves the management of data copies in HCI architecture, but also greatly improves the speed and stability of IO operations by reducing unnecessary network communication, which is a key step to realize high-performance and low-latency storage in hyper-converged environment.
[0040] In one example embodiment, after setting the target storage area as the storage location of the target virtual machine data copy, the above method further comprises: detecting the storage location setting result of the N storage areas; in the case where the setting result indicates that there is at least one storage area that is not set with a storage location, identifying the at least one storage area as a backup area of the data copy; in the case where the setting result indicates that there is no storage area that is not set with a storage location, sending a storage configuration valid target message to the management object corresponding to the storage system.
[0041] Optionally, if it is detected that there is a storage area that is not correctly configured, i.e., there is a storage area that is not associated with a specific node, these unconfigured areas will be marked as backup areas. The role of the backup area is to provide alternative storage space when the main storage area resource is full or fails, ensuring the storage safety and continuity of the virtual machine data copy. This mechanism enhances the flexibility of data management and the elasticity. When all N storage areas have been correctly set with storage locations, i.e., complete mapping between the nodes is achieved, a confirmation message will be sent to the storage management component, indicating that the storage configuration is valid. This step confirms that the storage is ready for data copy storage and management according to the IO localization strategy, marking the completion of the configuration phase and entering the normal running state.
[0042] In summary, through the above implementation, it is actively detected whether all N storage areas have been correctly set with storage locations according to the node information. According to the detection result, the storage layout is dynamically adjusted, and a backup storage area is provided to cope with the discontinuity of resource allocation or failure, while the management component is timely notified when the configuration is valid, providing a guarantee for efficient operation and management.
[0043] In one example embodiment, before determining the correspondence between the N storage areas and the N distributed nodes, the above method further comprises: setting a data transmission interface for the N storage areas, and configuring the function logic of the data transmission interface; determining a first number of a first type of storage area in the N storage areas that has been set with a data transmission interface; determining a second number of a second type of storage area in the N storage areas that has not been set with a data transmission interface based on the first number; synchronizing the second number to a cluster file subsystem to instruct the cluster file subsystem to increase the data transmission interface according to the second number.
[0044] It can be understood that by setting data transmission interfaces in each storage area (N), these interfaces are responsible for the transmission of data copies between different nodes. The function logic of configuring data transmission interfaces means defining how the interface works, including rules, priorities, encryption methods, etc. of data transmission, to ensure the efficiency and security of data during transmission. Then, based on the first number, the number of storage areas that have not yet set up data transmission interfaces will be calculated, i.e. the second number of the second type of storage areas. The purpose is to identify the configuration gap and provide a clear target for subsequent interface addition. Finally, the second number of information will be synchronized to the cluster file subsystem, instructing the subsystem to add the corresponding number of data transmission interfaces according to this number. The cluster file subsystem dynamically adds interfaces according to new requirements, which can effectively support the cross-node transmission requirements of data copies, especially when new storage areas require data copies. New data transmission interfaces will ensure that data can be quickly and accurately transmitted to the target node.
[0045] In summary, by establishing the data transmission mechanism between the cluster file system and the SDS before configuring the correspondence between the storage area and the node, and optimizing it. By pre-setting data transmission interfaces and dynamically increasing the number of interfaces according to the actual configuration situation, the system can more effectively manage and optimize the allocation of data copies between nodes and achieve the goal of IO localization.
[0046] In one example embodiment, after synchronizing the second number to the cluster file subsystem to instruct the cluster file subsystem to increase data transmission interfaces according to the second number, the above method further comprises: associating the added data transmission interfaces with the second type of storage areas that have not set up data transmission interfaces; after completing the association, according to the region parameter information of the second type of storage areas and the physical node number information of the virtual machine corresponding to the region parameter information; according to the region parameter information, the physical node number information supports the data storage function of the virtual machine.
[0047] In brief, after determining the need to increase the number of data transmission interfaces (i.e., the second number), the newly added interfaces are associated with the second type of storage areas that have not yet been configured. This means that each newly added interface will be assigned to a specific storage area to support the data transmission needs of that area. This association process ensures that each storage area has a corresponding data transmission interface, enabling cross-node transmission of data replicas. Once the data transmission interfaces are associated with the second type of storage area, further configuration is carried out based on the parameter information of the storage area and the numbering information of the virtual machine on the physical node. The area parameter information typically includes the size, location, and associated SDS node information of the storage area, while the virtual machine physical node numbering information is used to determine the specific physical node on which the virtual machine is located. Finally, these configuration information will be used to support the data storage function of the virtual machine, ensuring that the virtual machine can perform data read and write operations based on the characteristics of the node and the storage area it is located in. Specifically, when the virtual machine needs to store data, according to the physical node number of the virtual machine and the parameter information of the storage area, the data replica is preferentially stored in the storage area associated with the node where the virtual machine is located, achieving IO localization, reducing data transmission delay and network burden, and thus improving data access performance. If the available space of the current storage area is insufficient, other associated or standby storage areas will be considered to ensure data storage needs.
[0048] In summary, through the above embodiments, the close cooperation between data transmission interfaces and storage areas, as well as the reasonable matching between storage areas and virtual machine physical nodes, is ensured. Through these steps, not only the security and reliability of data storage are enhanced, but also the access efficiency of virtual machines to storage resources is maximized, achieving the improvement of data storage performance in hyper-converged architecture.
[0049] In one exemplary embodiment, supporting the data storage function of the virtual machine according to the area parameter information and the physical node numbering information includes: determining an unallocated storage area that meets the virtual machine disk space requirements through the area parameter information; storing the unallocated storage area and the virtual machine physical node numbering information to store the data replicas generated by the virtual machine in the unallocated storage area and generate corresponding data storage records.
[0050] It can be understood that when the virtual machine needs to create a new disk or expand the capacity of an existing disk, one or more currently unallocated storage areas are determined according to the requirements of the virtual machine (such as the required disk space size) and the parameter information of the storage area (such as the remaining space of the storage area and the partition layout), and these areas have enough space to meet the disk requirements of the virtual machine. When selecting a storage area, priority is given to the area that matches the physical node where the virtual machine is located to facilitate IO localization. Once a suitable unallocated storage area is found, a storage binding operation is performed, that is, the storage area is bound to the specific physical node where the virtual machine is located. This means that the data copy generated by the virtual machine will be preferentially stored in this storage area related to the node where it is located. The binding process records the physical node number of the virtual machine and the corresponding storage area information to form a data storage record for subsequent tracking and optimization of data management operations. After the storage binding is completed, the data copy generated by the virtual machine is directly stored in the unallocated storage area that has been bound, reducing the network delay caused by data transmission across nodes. At the same time, a data storage record is automatically generated, which contains detailed information of the storage operation, such as the virtual machine ID of the stored data, the location of the data copy (i.e. the bound storage area), and the timestamp of the storage operation, which helps to maintain data consistency and quickly find or migrate data copies when needed.
[0051] In summary, through the above operations, the optimization of virtual machine data storage is effectively realized, ensuring that the data copy can be stored on the physical node closest to the virtual machine under the premise of meeting performance requirements, thereby reducing cross-node data communication and improving the efficiency of IO operations. At the same time, storage binding and data storage record generation provide convenience for subsequent management and operation of the system, ensuring the transparency and controllability of data management. This strategy is particularly critical in hyper-converged architecture, as it can fully leverage the advantages of software-defined storage (SDS) while balancing the flexibility of virtualized environments and high performance of data access.
[0052] In an example embodiment, after the data copies generated by the M virtual machines are stored according to the correspondence, the method further includes: in a case where it is determined that a new virtual machine needs to be created or the storage size corresponding to the M virtual machines needs to be expanded, obtaining the remaining local storage size of the physical node corresponding to each of the M virtual machines; and updating the configuration information of the disk space corresponding to the M virtual machines according to the storage size and the remaining local storage size.
[0053] It can be understood that when a new virtual machine needs to be deployed in the hyper-converged architecture or the storage requirement of an existing virtual machine changes (such as expansion of disk capacity), the allocation of storage resources must be re-evaluated to ensure that data replicas continue to follow the principle of IO localization for storage. Collect and analyze the local storage resource status of each physical node where the virtual machine is located, specifically, obtain the size of the remaining available storage space on each node. This information is crucial for reasonable allocation of storage resources, as it directly determines which nodes can be used to create new virtual machines or expand disk space. With the remaining storage space information of each node, the disk space configuration of each virtual machine can be dynamically updated based on these data and the storage requirements of the virtual machine. For example, if a new virtual machine is created, the system will select the best physical node based on the available storage space, and then allocate disk space for it on that node, while updating the configuration of the cluster file system to ensure that the data replicas of the new virtual machine are stored preferentially on the local node. If it is to expand the storage capacity of an existing virtual machine, by checking the remaining local storage size of the node where the virtual machine is located, if sufficient, directly expand; if not, it needs to find other nodes with sufficient remaining storage space, expand the storage by migrating part of the data or adjusting the data replica distribution strategy, while updating the configuration information to reflect the new storage layout.
[0054] In summary, through the above embodiments, it is ensured that even when the number of virtual machines or storage requirements changes, the storage indication of data replicas can still follow the IO localization strategy, thereby reducing data transmission across nodes and improving the efficiency of IO operations. At the same time, the system updates the configuration based on real-time storage resource status, which can more flexibly cope with fluctuations in storage requirements, ensuring high performance and high availability of the virtualized environment under the hyper-converged architecture.
[0055] In one example embodiment, after storing the data replicas generated by the M virtual machines according to the corresponding relationship, the above method further comprises: in the case where it is determined that at least one virtual machine in the M virtual machines performs IO access, determining whether the target data replica corresponding to the at least one virtual machine is located in the local storage area; in the case where it is determined that the target data replica is located in the local storage area, prohibiting migration processing of the target data replica; in the case where it is determined that the target data replica is not located in the local storage area, triggering migration of the target data replica, wherein the migration is used to add at least one copy corresponding to the target data replica to the local storage area of the distributed node corresponding to the at least one virtual machine.
[0056] Optionally, when the system detects that any one or more of the M virtual machines is performing an IO access operation, such as reading or writing data on a disk, it initiates a series of checks and adjustment mechanisms to ensure that the data copy currently being accessed is located in the local storage area, i.e., the storage area matching the physical node on which the virtual machine is currently running. The system checks whether the target data copy, i.e., the data copy currently being accessed, is already stored in the local storage area of the node on which the virtual machine is located. This is a key step to ensure the efficiency of the IO operation, as the localization of the data copy can reduce network transmission delay and improve IO access speed. If the target data copy is already in the local storage area, the system will not perform migration processing on it, i.e., the current storage location will remain unchanged. This strategy avoids unnecessary data migration operations, saving computing and network resources, while maintaining high performance of IO access. Conversely, if the target data copy is not in the local storage area, the system will trigger the migration process of the data copy. Specifically, the system will migrate one or more copies of the target data copy to the local storage area of the node on which the virtual machine is located to achieve the localization of IO access. This operation is automatically completed through the mechanism of software-defined storage (SDS) without human intervention, thereby ensuring that the IO operation of the virtual machine can be performed on the local storage, reducing data transmission across nodes, and significantly improving the efficiency and performance of data access.
[0057] In summary, through the above implementation, the storage location of the virtual machine data copy can be continuously monitored and adjusted to ensure that the IO operation is completed on the local storage as much as possible, thereby optimizing the storage performance under the hyper-converged architecture and improving the overall response speed and user experience of the system.
[0058] The execution subject of the above steps can be a server, a terminal, etc., but is not limited thereto.
[0059] In order to facilitate understanding of the embodiments of the present application, related scenarios are explained and described, but do not limit the present application.
[0060] In related art, in the HCI scenario, when the virtual machine performs IO access, there is a problem of whether the IO is cross-node. Specifically, when the virtual machine IO data copy is located in the current node: the IO read operation does not need to be forwarded via the network, and its IO path is essentially a process of reading the local disk, with excellent performance; when the IO write operation, only the other copy needs to be forwarded to the other node, whether the performance or the network resource occupation is better than the IO data copy located in the non-current node scenario. When the virtual machine IO data copy is located in other nodes: the IO read operation needs to be read from other nodes via the network, which will be affected by network performance and data forwarding process; when the IO write operation, all copies need to be forwarded to other nodes, with large data forwarding volume and network resource occupation.
[0061] To avoid the occurrence of the above problems, as an optional implementation, the application optionally proposes an implementation method for supporting IO localization of a cluster file system. The method includes the following steps: adding a region division function in a data allocation manager of the cluster file system, supporting division of a data to-be-allocated space in a number consistent with a number of distributed cluster nodes when a disk is formatted; secondly, transmitting region division information to an SDS, and placing one of the replicas to a corresponding node according to the region division by the SDS; thirdly, adding an interface to the data space allocator of the cluster file system, supporting input of node number information when space is allocated; and finally, when a virtual disk or a virtual machine IO write triggers space allocation, the cluster file system calls the data allocation interface of the space allocation, and preferentially allocates disk space to the corresponding replica, so as to realize IO localization. The patent scheme modifies the data space allocation strategy of the cluster file system, supports linkage between cluster file system data block allocation and SDS replica data block allocation, and thus solves the problem that the cluster storage pool cannot realize IO localization in the HCI scenario.
[0062] As an optional implementation, Figure 3 is a schematic diagram of an architecture for supporting IO localization of a cluster file system according to an embodiment of the application, and includes a mapping relationship between virtual disk data block allocation and SDS underlying storage in a cluster storage pool in a typical hyper-converged scenario. Specifically, the virtual machine disk is essentially a file in the cluster storage pool, and the Hypervisor usually provides the virtual disk file to the virtual machine in the form of a block device via the virtio-blk or virtio-scsi protocol. The cluster storage pool backend is usually a large specification block device (hundreds of G or TB), which is usually provided by the SDS via the iscsi (Internet Small Computer Systems Interface, abbreviated as iscsi) protocol in the HCI scenario. The SDS node is a service component running on the physical node together with the virtualization system, and the SDS and the virtualization system interact with each other via the iscsi protocol. The SDS takes over the physical disk on the node to save data replicas, usually uses NVMe (Non-Volatile Memory Express, abbreviated as NVMe), SSD, etc. as a cache layer, and uses HDD disks to store mass replica files. Note that since the actual distribution of the virtual disk data is managed by the SDS, the real location of the data replica can be on the local disk of any node in the cluster.
[0063] Specifically, Figure 3There are two virtual machines (a first virtual machine VM1 and a second virtual machine VM2) connected to a cluster storage pool, and data transmission between the cluster storage pool and an SDS is performed through a NIC (Network Interface Card, NIC for short). The SDS mainly includes an SDS Core and an iscsi target (an iscsi protocol target storage resource). The SDC stores and processes the received data through a cache disk, an HDD disk, and the like.
[0064] Optionally, Figure 4 is a schematic diagram of a virtual disk data and SDS copy relationship according to an embodiment of the present application. In Figure 4 In the virtual disk 1 (vdisk1) and the virtual disk 2 (vdisk2), the virtual disks correspond to two virtual machines located on different nodes, respectively. The data block distribution in the virtual disk is controlled by a file system data allocator, and then the data block distribution is distributed in a CFS (Ceph File Syste, CFS for short) storage pool space. In fact, the data of the two virtual disks may be mixed and distributed on a LUN (Logical Unit Number, LUN for short). The SDS places a copy of the LUN on a local disk of any node. The figure shows a two-node scenario. For a multi-node scenario, the data distribution may not be located on the first node node1 or the second node node2.
[0065] Optionally, Figure 5 is a schematic diagram of a change of a virtual disk data and SDS copy relationship according to an embodiment of the present application. First, a region division function is added to a data allocation manager of a cluster file system, and the data allocation space division number is consistent with the number of distributed cluster nodes when formatting a disk. For example, in the figure, the available area corresponding to the CFS is evenly divided into multiple allocation zones when formatting, the number of zones is consistent with the number of cluster nodes, the first allocation zone zone1 is the allocation zone of the first node node1, and the second allocation zone zone2 is the allocation zone of the second node node2.
[0066] Second, the division region information is transmitted to the SDS, and the SDS places one of the copies to the corresponding node according to the region division. An interface is provided by the SDS. When the CFS storage pool is divided, the zone information is transmitted to the SDS system through the HCI system control plane. The SDS system will distribute a copy of an SDS LUN to the corresponding SDS node according to the zone division. As shown in the figure, the actual copy distribution of the CFS zone1 is located on the local disk of the current node, that is, the SDS node1.
[0067] Again, add an interface to the data space allocator of the cluster file system to support the transmission of node number information when allocating space.
[0068] Optionally, to ensure the efficiency of the space allocation native interface of the cluster file system, the interface generation control can be performed through the following code:
[0069] .
[0070] It should be noted that the input parameter in the above code only has the space size information to be allocated, and the node information is added and the actual space allocation is performed, and the space is allocated from the allocation zone of the specified node. Note that when the current zone has insufficient space to be allocated, the space will be allocated from other zones to ensure the availability of the function. Since the virtual machines on each node cannot guarantee the same size when applying for space, this strategy allows some nodes to apply for space from other zones when the space application limit is large.
[0071] Finally, when creating a virtual disk or a virtual machine IO write triggers space allocation, the cluster file system will call the optimized allocation interface to allocate data blocks, and preferentially allocate disk space to the replica corresponding to the current node. In this way, when the virtual machine is running, most of the IO of the current node can be localized.
[0072] To sum up, by modifying the data space allocation strategy of the cluster file system, the cluster file system data block allocation and SDS replica data block allocation are supported, thereby solving the problem that the cluster storage pool cannot achieve IO localization in the HCI scenario. Using this scheme, in the HCI scenario, the advantages of the cluster storage pool such as HA, load balancing, snapshot support, and other characteristics can be fully utilized, and the SDS IO localization feature can also be supported, achieving higher IO performance. This scheme can significantly improve the overall performance of the HCI product without increasing hardware costs, effectively enhancing the product competitiveness.
[0073] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0074] A data copy storage system is also provided in the embodiment, which is used to implement the above-mentioned embodiments and preferred embodiments, and will not be described again. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation of hardware, or a combination of software and hardware, is also possible and contemplated.
[0075] Figure 6 is a structural block diagram of a data copy storage system according to the embodiment of the present application, as shown in Figure 6 , the system comprises:
[0076] The cluster file subsystem 62 is configured to determine N distributed nodes corresponding to M virtual machines supported by the hyper-converged architecture application, and divide a to-be-allocated storage space associated with the hyper-converged architecture application into N storage areas in a disk formatting process according to the N distributed nodes.
[0077] The software-defined storage subsystem 64 is connected with the cluster file subsystem and is configured to determine a correspondence between the N storage areas and the N distributed nodes, wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; and store a data copy generated by the M virtual machines according to the correspondence.
[0078] Through the above system, firstly, N distributed nodes corresponding to M virtual machines supported by the hyper-converged architecture application are determined, and then in a disk formatting process, a to-be-allocated storage space associated with the hyper-converged architecture application is divided according to the N distributed nodes, so as to divide the to-be-allocated storage space into N storage areas, so as to ensure that the number of divided to-be-allocated spaces is consistent with the number of distributed cluster nodes, and the correspondence between the N storage areas and the N distributed nodes is obtained. On this basis, the software-defined storage subsystem is instructed to place one of the virtual machine copies to the corresponding node according to the correspondence. Then, the efficient local storage technology effect of the data copy in the software-defined storage is realized. Through the above method, the storage space is finely managed, the distribution of the data copy is matched with the node where the virtual machine is located, the data transmission across nodes is reduced, the IO access efficiency is improved, the storage performance is optimized without increasing additional hardware cost, and the applicability and competitiveness of the hyper-converged architecture in a large-scale virtualization environment are enhanced.
[0079] In an example embodiment, the software-defined storage subsystem is further configured to determine physical node number information corresponding to a target virtual machine generating the data copy; determine a target distributed node associated with the target virtual machine based on the physical node number information, and determine a target storage area corresponding to the target distributed node based on the corresponding relationship; and set the target storage area as a storage location of the data copy of the target virtual machine to indicate storage of the data copy.
[0080] In an example embodiment, the software-defined storage subsystem is further configured to, after setting the target storage area as the storage location of the data copy of the target virtual machine, detect a storage location setting result of the N storage areas; identify at least one storage area as a backup area of the data copy if the setting result indicates that there is at least one storage area without a set storage location; and send a storage configuration valid target message to a management object corresponding to the storage system if the setting result indicates that there is no storage area without a set storage location.
[0081] In an example embodiment, the cluster file subsystem is further configured to, before determining the corresponding relationship between the N storage areas and the N distributed nodes, set a data transmission interface for the N storage areas and configure function logic of the data transmission interface; determine a first quantity of a first type of storage area in the N storage areas that has a set data transmission interface; determine a second quantity of a second type of storage area in the N storage areas that does not have a set data transmission interface based on the first quantity; and synchronize the second quantity to the cluster file subsystem to instruct the cluster file subsystem to increase the data transmission interface according to the second quantity.
[0082] In an example embodiment, the cluster file subsystem is further configured to, after synchronizing the second quantity to the cluster file subsystem to instruct the cluster file subsystem to increase the data transmission interface according to the second quantity, associate the increased data transmission interface with the second type of storage area that does not have a set data transmission interface; and support data storage functions of a virtual machine based on region parameter information of the second type of storage area and physical node number information of the virtual machine corresponding to the region parameter information after the association is completed.
[0083] In an example embodiment, the cluster file subsystem is further configured to determine an unallocated storage area that meets virtual machine disk space requirements based on the region parameter information; store bind the unallocated storage area with physical node number information of a virtual machine to store a data copy generated by the virtual machine to the unallocated storage area and generate a corresponding data storage record.
[0084] In an example embodiment, the cluster file subsystem is further configured to, after the storage indication of the data replicas generated by the M virtual machines according to the correspondence relationship, in a case where it is determined that a new virtual machine needs to be created or the storage size corresponding to the M virtual machines needs to be expanded, acquire a remaining local storage size of a physical node corresponding to each of the M virtual machines; and update configuration information of disk spaces corresponding to the M virtual machines according to the storage size and the remaining local storage size.
[0085] In an example embodiment, the cluster file subsystem is further configured to, after the storage indication of the data replicas generated by the M virtual machines according to the correspondence relationship, in a case where it is determined that at least one virtual machine of the M virtual machines performs IO access, determine whether a target data replica corresponding to the at least one virtual machine is located in a local storage area; in a case where it is determined that the target data replica is located in the local storage area, prohibit migration processing of the target data replica; and in a case where it is determined that the target data replica is not located in the local storage area, trigger migration of the target data replica, wherein the migration is used to add at least one replica corresponding to the target data replica to a local storage area of a distributed node corresponding to the at least one virtual machine.
[0086] In an example embodiment, the cluster file subsystem comprises a data block allocation unit configured to determine a to-be-allocated storage space supported by a super-converged architecture application for the M virtual machines, and evenly split the to-be-allocated storage space into N allocable areas.
[0087] It should be noted that the data block allocation unit is a key component in the cluster file subsystem, and its main responsibility is to manage the allocation strategy of the storage space to meet the demand of virtual machines for storage resources. In the super-converged architecture application, the unit will:
[0088] Determine the to-be-allocated storage space: First, the data block allocation unit needs to determine how much storage space can be allocated to the M virtual machines. This usually involves evaluation of the current storage resources, including the available storage capacity on all nodes, and the existing data distribution situation.
[0089] Evenly split into N allocable areas: After determining the total to-be-allocated storage space, the data block allocation unit will evenly split these spaces into N allocable areas, where N is equal to the number of nodes in the distributed cluster. Each area will be assigned to a specific node to support subsequent data block allocation and storage operations, thereby promoting the localization of data storage.
[0090] By dividing the storage space into equal regions as the number of nodes, the data block allocation unit can ensure that each node has a fair allocation of storage resources, while also providing a foundation for the implementation of local storage of data replicas. When virtual machines generate data replicas, the system can preferentially store the replicas in the allocable region corresponding to the node where the virtual machine is located, reducing cross-node data transmission and improving the performance of IO operations. The implementation of this strategy requires close cooperation between the cluster file subsystem and software-defined storage (SDS). By passing region division information and node information, SDS can adjust the distribution of its data replicas based on this information to support IO localization.
[0091] In summary, the data block allocation unit plays a crucial role in the super-converged architecture. It not only manages the allocation of storage space, but also provides technical support for efficient and low-latency data access of virtual machines by dividing the storage space into allocable regions matching the number of nodes. This mechanism helps maintain the high-performance state of the system while ensuring the flexible use of storage resources by virtual machines.
[0092] In an exemplary embodiment, the cluster file subsystem described above includes an interface increasing unit for adjusting the number of interfaces of the cluster file subsystem according to the physical node number information of the M virtual machines transmitted during space allocation.
[0093] The cluster file subsystem (CFS) plays a core role in the hyper-converged architecture, responsible for managing data block allocation and file operations in the storage pool. In traditional implementations, the number of interfaces of the CFS is usually fixed and related to the configuration at system initialization. However, as the number of virtual machines increases or changes, and the demand for IO localization increases, the original interface configuration may no longer meet the requirements of high-performance storage access. To solve this problem, the "interface increasing unit" in this embodiment plays a crucial role: by adjusting the number of interfaces based on physical node number information, when the system is ready to allocate space for M virtual machines, the interface increasing unit receives information representing the number of physical nodes where the virtual machines are located. Based on this information, it can dynamically adjust the number of interfaces of the cluster file subsystem to ensure that each node's virtual machine can effectively utilize local storage resources. Secondly, it can also optimize data access paths: by increasing the number of interfaces matching the number of nodes, the system can more directly route data requests to the correct physical nodes, reducing unnecessary data transmission between nodes, thereby significantly improving data access efficiency and performance. That is, the interface increasing unit can ensure that when a virtual machine requests storage space, the corresponding data block allocation operation will give priority to the node's local storage resources, which is closely related to the IO localization strategy, effectively reducing cross-node data transmission delay and improving IO operation response speed. Further, the dynamic adjustment capability of the interface increasing unit also allows the system to automatically adjust the number of interfaces when nodes are added or reduced, to adapt to the dynamic changes of nodes in the hyper-converged architecture, ensuring the high availability and scalability of the system.
[0094] In summary, through the implementation of the interface increasing unit function, the cluster file subsystem can more flexibly cope with changes in virtual machines and physical nodes, optimize data access paths by increasing the number of interfaces matching the number of nodes, support IO localization strategies, and thus maintain the efficiency and performance of the storage pool while ensuring fast access to storage resources by virtual machines. This dynamic adjustment and optimization strategy is the key to intelligent storage management in the hyper-converged architecture, helping to improve overall system performance and response capabilities.
[0095] In one exemplary embodiment, the above storage system further comprises a policy component for obtaining in real time storage resources and load data corresponding to each of the N distributed nodes, and adjusting the split size after formatting and splitting the to-be-allocated storage space according to the storage resources and the load data.
[0096] It can be understood that the policy component plays the role of decision maker in the storage system, and its core task is to optimize the space splitting strategy by considering the real-time state of each distributed node during storage space allocation. The specific functions are as follows:
[0097] (1) Real-time resource and load monitoring: The policy component periodically or proactively collects storage resource usage (such as remaining disk space, cache usage, etc.) and load data (such as CPU utilization, network bandwidth usage, etc.) of N distributed nodes, which can reflect the current processing capacity and storage capacity of the nodes.
[0098] (2) Dynamic adjustment of split size: Based on the collected resource and load data, the policy component can intelligently adjust the split size of the storage space to be allocated. For example, in the case where some nodes are short of storage resources while others still have a large amount of free space, the policy component can increase the split proportion of the nodes with abundant free resources and reduce the split burden of the nodes with resource shortage.
[0099] (3) Optimization of space allocation: Dynamic adjustment of split size helps to optimize the allocation of storage space. The policy component can ensure that data replicas are stored preferentially on nodes with sufficient resources and low load, thereby improving the efficiency of IO operations and reducing performance degradation caused by node overload.
[0100] (4) Promote load balancing: Through the intervention of the policy component, the split of storage space is no longer a simple average allocation, but is adjusted according to the actual resources and load of the nodes. This dynamic split strategy helps to achieve load balancing of storage resources and avoid single node as a performance bottleneck.
[0101] (5) Support dynamic adjustment and fault recovery: The intelligent features of the policy component also lie in its ability to quickly respond to dynamic changes in the storage system, such as automatically adjusting the split strategy in the case of node failure or sudden load increase, ensuring the reliability and access performance of data are not affected.
[0102] In summary, through real-time monitoring and intelligent decision-making of the policy component, the storage system can better adapt to the changing resources and loads in a distributed environment, ensuring that the split of storage space takes into account both data safety redundancy and optimal performance. This strategic adjustment is particularly important for hyper-converged architecture, as it can maximize the use of existing resources, improve the overall storage efficiency and data access speed of the system, while reducing the complexity of management and maintenance.
[0103] In an exemplary embodiment, the above-mentioned cluster file subsystem further comprises: a data space allocation unit, configured to, on the basis of determining that a target data replica has been generated, preferentially allocate disk space in the storage system to a target distributed node corresponding to the target data replica.
[0104] That is, the data space allocation unit is a module in the cluster file subsystem responsible for data storage space allocation. Its main task is to ensure that the disk space allocation strategy in the storage system prioritizes the local storage resources of the distributed nodes corresponding to the target data replica after the target data replica is generated. Specifically: the data space allocation unit first needs to identify which data replicas are target replicas, i.e., those data sets that need to prioritize localized storage. This is usually based on the access pattern of the virtual machine, data access frequency, and hot and cold state analysis of the data. After the target data replica is identified, the data space allocation unit further analyzes the storage location of the target replica to determine which replicas should be stored on the local distributed node to reduce cross-node IO transmission.
[0105] When allocating storage space, the data space allocation unit prioritizes the local storage resources of the target distributed node. This means that when a virtual machine requests additional storage space or performs data writing, the new data block will first attempt to save on the local disk of the node, rather than being forwarded to other nodes through the network. The linkage between the data space allocation unit and the data replica layout ensures that the allocation of storage space is no longer independent of replica distribution, but can be dynamically adjusted according to the replica layout strategy, ensuring the localization of data storage and access. The priority allocation strategy of the data space allocation unit directly improves the IO performance, as the localization of data blocks reduces the distance and time of data transmission, improving the speed and efficiency of data access.
[0106] Through the above mechanism, the data space allocation unit can ensure efficient storage and access of virtual machine data, while supporting the localization of data replica storage strategy under the super-converged architecture. The introduction of this unit is of great significance for optimizing the resource utilization of the storage system, improving the IO performance, and enhancing the overall reliability and availability of the system.
[0107] In an example embodiment, the software-defined storage subsystem described above further includes a configuration unit configured to place at least one copy of the same logical block on the local disk of a specific node according to the region parameters corresponding to the N storage regions, wherein the specific node matches the node distribution of the N distributed nodes of the M virtual machines in the storage pool.
[0108] The configuration unit receives N storage area partition information from the cluster file subsystem, each area corresponding to a specific node in the virtual machine distribution. These area parameters contain important details about how to allocate data between nodes, including storage capacity, data distribution strategy, etc. Based on the received area parameters, the configuration unit will determine how to place data replicas. For any given logical block (i.e. data block in the virtual machine file system), the configuration unit will ensure that at least one copy is placed on the node corresponding to the virtual machine distribution, which is called the "specific node". When selecting the specific node for replica placement, the configuration unit will prioritize using the local disk of the node for storage rather than the disk of a remote node. This reduces network latency for data transmission and enhances the performance and efficiency of data access. The replica placement strategy of the configuration unit matches the node distribution of the N distributed nodes of the M virtual machines in the storage pool, ensuring that regardless of virtual machine migration or redistribution, data access localization is maintained, further optimizing IO performance. This node-distribution-based replica placement method also supports dynamic adjustment, allowing the configuration unit to rearrange replica placement to adapt to new node distributions when the virtual machine distribution changes, maintaining efficient and localized data access.
[0109] Through the function of the configuration unit, SDS can effectively support data replica layout optimization in hyper-converged architecture, ensuring efficient use of storage resources, while reducing data transmission distance, improving virtual machine IO access speed and overall system storage performance. This intelligent data management strategy is an important application of software-defined storage technology in hyper-converged environments, helping to achieve higher resource utilization efficiency and service levels.
[0110] In an exemplary embodiment, the above-mentioned software-defined storage subsystem further comprises a replication management unit for monitoring the working state of the data replica; in the case where the working state indicates that the data replica has lost or failed, performing replica recovery through other distributed nodes other than the current distributed node where the data replica is located; in the case where the working state indicates that the data replica has not lost or failed, determining that the storage state of the data replica is normal and no further processing operation is needed.
[0111] In an exemplary embodiment, the above-mentioned software-defined storage subsystem further comprises a cache management unit for storing a first type of data block with a frequency of access by the M virtual machines greater than or equal to a preset frequency in a cache device corresponding to the N distributed nodes, and storing a second type of data block with a frequency of access by the M virtual machines less than the preset frequency in a non-cache device corresponding to the N distributed nodes.
[0112] In an example embodiment, the software-defined storage subsystem further includes an interface generation unit configured to receive the number of interfaces sent by the clustered file subsystem upon determining that the clustered file subsystem completes the splitting, and generate a new data synchronization interface between the software-defined storage subsystem and the clustered file subsystem according to the number of interfaces.
[0113] It should be noted that the above modules can be implemented by software or hardware, and for the latter, the following implementation manners can be used, but are not limited thereto: the above modules are located in the same target processor; or the above modules are located in different target processors in any combination.
[0114] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.
[0115] In an example embodiment, the computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0116] Embodiments of the present application also provide an electronic device, which includes a target memory storing a computer program and a target processor configured to run the computer program to execute the steps in any of the above method embodiments.
[0117] Optionally, Figure 7 is a computer system structure block diagram of an electronic device according to an embodiment of the present application. As shown in Figure 7 the computer system 800 includes a CPU 801 (Central Processing Unit, CPU for short) which can perform various appropriate actions and processes according to programs stored in a ROM 802 (Read-Only Memory, ROM for short) or programs loaded from a storage part 808 into a RAM 803 (Random Access Memory, RAM for short). In the RAM 803, various programs and data required for system operation are also stored. The CPU 801, the ROM 802 and the RAM 803 are connected to each other through a bus 804. An I / O interface 805 (Input / Output interface, I / O interface for short) is also connected to the bus 804.
[0118] The following components are connected to the I / O interface 805: an input section 806 including a keyboard and a mouse, an output section 807 including a display such as a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), and the like, and a speaker, a storage section 808 including a hard disk, and the like, and a communication section 809 including a network interface card such as a Local Area Network (LAN) card, a modem, and the like. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as necessary. A removable recording medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 810 as necessary, so that a computer program read out from it is installed in the storage section 808 as necessary.
[0119] In one example embodiment, the electronic device described above can further include a transmission device connected to the central processing unit (CPU 801) and an input / output device connected to the central processing unit (CPU 801).
[0120] Embodiments of the present application also provide a computer program product including a computer program, which, when executed by a processor, implements the steps in any of the method embodiments described above.
[0121] Embodiments of the present application also provide another computer program product including a non-volatile computer readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the method embodiments described above.
[0122] Embodiments of the present application also provide a computer program including computer instructions stored in a computer readable storage medium; a processor of a computer device reads the computer instructions from the computer readable storage medium, and executes the computer instructions, so that the computer device executes the steps in any of the method embodiments described above.
[0123] The specific examples in the present embodiment can refer to the examples described in the above embodiments and example embodiments, which will not be described here again.
[0124] Those skilled in the art will further realize that the mere concepts, teachings, examples, and steps described in the foregoing description are not meant to limit or restrict the scope of the present application in any way but are merely provided to illustrate the principles and concepts of the present application. Thus, the scope of the present application should not be limited to the specific examples described herein, but should be given the broadest possible scope commensurate with the principles and concepts described herein.
[0125] Obviously, those skilled in the art should understand that each module or step of the present application described above can be realized by a general computing device, which can be centralized on a single computing device or distributed on a network composed of multiple computing devices, and can be realized by a program code executable by a computing device, so that it can be stored in a storage device and executed by a computing device, and in some cases, the steps shown or described can be executed in an order different from that shown here, or they can be made into individual integrated circuit modules, or a plurality of modules or steps can be made into a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.
[0126] The above provides a detailed introduction to the data copy storage system, method, medium, electronic device and program product provided by the present application. The principles and implementation modes of the present application are described by applying specific examples. The above example is only used to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for storing a data copy, characterized in that: include: Determine the N distributed nodes corresponding to the M virtual machines supported by the hyper-converged architecture application; During the disk formatting process, the to-be-allocated storage space associated with the super-converged architecture application is divided according to the N distributed nodes, so as to equally divide the to-be-allocated storage space into N storage areas; Determine a correspondence between the N storage areas and the N distributed nodes; wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; Performing storage instructions on the data copies generated by the M virtual machines according to the corresponding relationship; The method further includes: if it is determined that at least one virtual machine among the M virtual machines performs IO access, determining whether a target data copy corresponding to the at least one virtual machine is located in a local storage area; if it is determined that the target data copy is located in the local storage area, prohibiting migration processing of the target data copy; if it is determined that the target data copy is not located in the local storage area, triggering migration of the target data copy, wherein the migration is used to add at least one copy corresponding to the target data copy to the local storage area of the distributed node corresponding to the at least one virtual machine; After indicating the storage of the data copies generated by the M virtual machines according to the corresponding relationship, the method further includes: when it is determined that a new virtual machine needs to be created or the storage size corresponding to the M virtual machines needs to be expanded, obtaining the remaining local storage size of the physical node corresponding to each of the M virtual machines; and updating the configuration information of the disk space corresponding to the M virtual machines according to the storage size and the remaining local storage size; The method further includes: acquiring storage resources and load data corresponding to each of the N distributed nodes in real time, and adjusting the size of the partitions of the storage space to be allocated after formatting according to the storage resources and the load data.
2. The data copy storage method according to claim 1, characterized in that: Performing storage instructions on the data copies generated by the M virtual machines according to the corresponding relationship includes: Determine the physical node number information corresponding to the target virtual machine for generating the data copy; Determine a target distributed node associated with the target virtual machine based on the physical node number information, and determine a target storage area corresponding to the target distributed node through the corresponding relationship; The target storage area is set as the storage location of the target virtual machine data copy to indicate storage of the data copy.
3. The data copy storage method according to claim 2, characterized in that: After setting the target storage area as the storage location of the target virtual machine data copy, the method further includes: Detecting storage position setting results of the N storage areas; If the setting result indicates that there is at least one storage area for which no storage location is set, identifying the at least one storage area as a spare area for the data copy; In a case where the setting result indicates that there is no storage area without a storage location set, a target message indicating that the storage configuration is valid is sent to a management object corresponding to the storage system.
4. The data copy storage method according to claim 1, characterized in that: Before determining the correspondence between the N storage areas and the N distributed nodes, the method further includes: Setting a data transmission interface for the N storage areas and configuring function logic of the data transmission interface; Determine a first number of first-type storage areas having data transmission interfaces set therein among the N storage areas; Determine, based on the first number, a second number of second-type storage areas in the N storage areas where no data transmission interface is provided; The second number is synchronized to the cluster file subsystem to instruct the cluster file subsystem to increase the data transmission interface according to the second number.
5. The data copy storage method according to claim 4, characterized in that: After synchronizing the second number to the cluster file subsystem to instruct the cluster file subsystem to increase the data transmission interface according to the second number, the method further includes: Associating the added data transmission interface with the second type of storage area where no data transmission interface is set; When the association is completed, according to the area parameter information of the second type of storage area and the physical node number information of the virtual machine corresponding to the area parameter information; The data storage function of the virtual machine is supported according to the area parameter information and the physical node number information.
6. The data copy storage method according to claim 5, characterized in that: Supporting a data storage function of a virtual machine according to the area parameter information and the physical node number information includes: Determine an unallocated storage area that meets the virtual machine disk space requirement using the area parameter information; The unallocated storage area is bound to the physical node number information where the virtual machine is located, so as to store the data copy generated by the virtual machine in the unallocated storage area and generate a corresponding data storage record.
7. A data copy storage system, characterized in that: include: The cluster file subsystem is used to determine the N distributed nodes corresponding to the M virtual machines supported by the hyper-converged architecture application; During the disk formatting process, the to-be-allocated storage space associated with the super-converged architecture application is divided according to the N distributed nodes, so as to equally divide the to-be-allocated storage space into N storage areas; a software-defined storage subsystem connected to the cluster file subsystem, configured to determine a correspondence between the N storage areas and the N distributed nodes, wherein each storage area corresponds to at least one distributed node, M is less than or equal to N, and M and N are positive integers; and to indicate storage of the data copies generated by the M virtual machines according to the correspondence; The software-defined storage subsystem is further configured to, when it is determined that at least one virtual machine among the M virtual machines performs IO access, determine whether a target data copy corresponding to the at least one virtual machine is located in a local storage area; when it is determined that the target data copy is located in the local storage area, prohibit migration processing of the target data copy; when it is determined that the target data copy is not located in the local storage area, trigger migration of the target data copy, wherein the migration is configured to add at least one copy corresponding to the target data copy to a local storage area of a distributed node corresponding to the at least one virtual machine; The cluster file subsystem is further configured to, after providing storage instructions for the data copies generated by the M virtual machines according to the corresponding relationship, obtain the remaining local storage size of the physical node corresponding to each of the M virtual machines when it is determined that a new virtual machine needs to be created or the storage size corresponding to the M virtual machines needs to be expanded; and update the configuration information of the disk space corresponding to the M virtual machines according to the storage size and the remaining local storage size; The cluster file subsystem is further used to obtain storage resources and load data corresponding to each of the N distributed nodes in real time, and adjust the size of the split after formatting the allocated storage space according to the storage resources and the load data.
8. The data copy storage system according to claim 7, characterized in that: The cluster file subsystem includes: a data block allocation unit, which is used to determine the storage space to be allocated for the super-converged architecture application to support the M virtual machines, and evenly divide the storage space to be allocated into N allocatable areas.
9. The data copy storage system according to claim 7, characterized in that: The cluster file subsystem includes: an interface adding unit, which is used to adjust the number of interfaces of the cluster file subsystem according to the physical node numbering information of the M virtual machines input during space allocation.
10. The data copy storage system according to claim 7, characterized in that: The cluster file subsystem further includes: a data space allocation unit configured to preferentially allocate the disk space in the storage system to the target distributed node corresponding to the target data copy based on determining that the target data copy has been generated.
11. The data copy storage system according to claim 7, wherein: The software-defined storage subsystem also includes: a configuration unit for placing at least one copy of the same logical block on the local disk of a specific node according to the area parameters corresponding to the N storage areas, wherein the specific node matches the node distribution of the N distributed nodes of the M virtual machines in the storage pool.
12. The data copy storage system according to claim 7, characterized in that: The software-defined storage subsystem also includes: a replication management unit for monitoring the working status of the data copy; when the working status indicates that the data copy is lost or faulty, the copy is restored through other distributed nodes except the current distributed node where the data copy is located; when the working status indicates that the data copy is not lost or faulty, it is determined that the storage status of the data copy is normal and no other processing operations need to be performed.
13. The data copy storage system according to claim 7, characterized in that: The software-defined storage subsystem also includes: a cache management unit, which is used to store the first type of data blocks whose access frequency by the M virtual machines is greater than or equal to the preset frequency in the cache devices corresponding to the N distributed nodes, and store the second type of data blocks whose access frequency by the M virtual machines is less than the preset frequency in the non-cache devices corresponding to the N distributed nodes.
14. The data copy storage system according to claim 7, characterized in that: The software-defined storage subsystem also includes: an interface generation unit, which is used to receive the number of interfaces sent by the cluster file subsystem when it is determined that the cluster file subsystem has completed the division, and generate a new data synchronization interface between the software-defined storage subsystem and the cluster file subsystem according to the number of interfaces.
15. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data copy storage method as claimed in any one of claims 1 to 6 when executing the computer program.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data copy storage method according to any one of claims 1 to 6 are implemented.
17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data copy storage method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Data distribution method and system of hyper-converged system
CN114003350A
Storage space distribution method and server
CN118193479A