Data storage methods, devices, electronic equipment and storage media

By acquiring the load and resource description information of virtual storage units, calculating priority coefficients, filtering out high-priority units and combining them into storage unit groups, the problem of low accuracy in virtual storage unit selection is solved, and efficient and accurate storage unit selection is achieved.

CN120994141BActive Publication Date: 2026-01-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511512064.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-30
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

In existing technologies, the selection of virtual storage units lacks a comprehensive consideration of the actual operating status, resulting in low selection accuracy.

Method used

By obtaining the load description information and resource description information of the virtual storage units, the priority coefficient is calculated, high-priority storage units are selected, and they are combined with the virtual storage units connected by the route to form a storage unit group, and finally the target storage unit is determined.

Benefits of technology

It improves the accuracy and flexibility of virtual storage unit selection, ensuring that the selected storage unit can meet the storage needs of the current object data while also having low load and high read/write efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994141B_ABST
    Figure CN120994141B_ABST
Patent Text Reader

Abstract

This application discloses a data storage method, apparatus, electronic device, and storage medium, relating to the field of computer technology. The method includes: acquiring load description information and resource description information corresponding to each virtual storage unit matched with a target terminal; determining a priority coefficient for each of at least one virtual storage unit based on the load description information and resource description information; if a first virtual storage unit is determined from the at least one virtual storage unit, adding at least one second virtual storage unit with a routing connection to the first virtual storage unit, as well as at least one first virtual storage unit, to a storage unit group; and determining a target storage unit from the entire storage unit group. By comprehensively obtaining priority information from both load description information and resource description information, and then selecting the target storage unit based on that priority, the accuracy of virtual storage unit selection is improved, thereby solving the problem of low accuracy in virtual storage unit selection in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to data storage methods, apparatus, electronic devices and storage media. Background Technology

[0002] In current virtual storage system application scenarios, with the explosive growth of data volume and the continuous improvement of business requirements for data read and write response speed, the reasonable selection and allocation of virtual storage units has become a key link to ensure the stable operation of the system.

[0003] In related technologies, the selection of virtual storage units mostly considers only a few factors. For example, it may determine whether to use the virtual storage unit for data storage based solely on the remaining storage space of the virtual storage unit, or simply refer to historical read and write speed data for decision-making. This lacks a comprehensive consideration of the actual operating status of the virtual storage unit, resulting in low accuracy in the selection of virtual storage units. Summary of the Invention

[0004] This application provides data storage methods, apparatus, electronic devices, and storage media to at least address the problem of low accuracy in selecting virtual storage units in related technologies.

[0005] This application provides a data storage method, comprising: acquiring load description information and resource description information corresponding to at least one virtual storage unit matched with a target terminal, wherein the load description information is used to indicate the load status of the virtual storage unit when performing data read / write tasks, and the resource description information is used to indicate the read / write efficiency of the virtual storage unit when performing data read / write tasks; determining a priority coefficient for each of the at least one virtual storage unit based on the load description information and the resource description information; if at least one first virtual storage unit is determined from the at least one virtual storage unit based on the priority coefficient, adding at least one second virtual storage unit with a routing connection to the first virtual storage unit, and the at least one first virtual storage unit, to a storage unit group; and determining a target storage unit from all the storage unit groups, wherein the target storage unit is used to store the currently acquired object data.

[0006] This application also provides a data storage device, comprising: a first acquisition unit, configured to acquire load description information and resource description information corresponding to at least one virtual storage unit matched with a target terminal, wherein the load description information is used to indicate the load status of the virtual storage unit when performing data read / write tasks, and the resource description information is used to indicate the read / write efficiency of the virtual storage unit when performing data read / write tasks; a first determination unit, configured to determine a priority coefficient for each of the at least one virtual storage unit based on the load description information and the resource description information; an addition unit, configured to add at least one second virtual storage unit with a routing connection to the first virtual storage unit, and the at least one first virtual storage unit, to a storage unit group when at least one first virtual storage unit is determined from the at least one virtual storage unit based on the priority coefficient; and a second determination unit, configured to determine a target storage unit from all the storage unit groups, wherein the target storage unit is used to store the currently acquired object data.

[0007] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the data storage method described above when executing the computer program.

[0008] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described data storage methods.

[0009] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described data storage methods.

[0010] This application first obtains load description information and resource description information corresponding to at least one virtual storage unit matched with the target terminal. This breaks through the limitations of related technologies that rely solely on remaining storage space or a single historical read / write speed for judgment. It combines the load status and read / write efficiency of the virtual storage unit, two key operational status information, to achieve a multi-dimensional and comprehensive consideration of its actual operational status, providing complete data support for subsequent selection and avoiding evaluation bias caused by incomplete information. Next, the priority coefficient of each virtual storage unit is determined based on the above two types of description information. Compared with the subjective judgment of related technologies based on a single dimension, this improves the accuracy of priority determination and lays the foundation for selecting suitable storage units. Subsequently, given the first virtual storage unit, the priority coefficient of each virtual storage unit is determined. The addition of the second and first virtual storage units with existing routing connections to the storage unit group solves the problem of related technologies that only select units that meet the criteria in a single dimension and miss high-quality associated units, thus expanding the range of effective candidate storage units and further improving selection flexibility. Finally, the target storage unit for storing the current object data is determined from the entire storage unit group. Since the storage unit group already includes high-priority units that have been fully evaluated and associated units with data transmission feasibility, the final selected target storage unit can not only meet the storage requirements of the current object data, but also has the optimal operating state of low load and high read / write efficiency. This solves the technical problem of low accuracy in selecting virtual storage units in related technologies and improves the accuracy of virtual storage unit selection. Attached Figure Description

[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A hardware structure block diagram of a server device for a data storage method provided in this application embodiment;

[0013] Figure 2 This is a schematic diagram of an optional data storage method according to an embodiment of this application;

[0014] Figure 3 This is a structural block diagram of a data storage device according to an embodiment of this application;

[0015] Figure 4 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0017] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0018] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a computer device for a data storage method according to an embodiment of this application. For example... Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a central processing unit (CPU), microprocessor (MCU), or programmable logic device (FPGA), etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0020] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the data storage method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0021] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0022] This embodiment provides a data storage method. Figure 2 This is a flowchart of a data storage method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0023] S202, obtain load description information and resource description information corresponding to at least one virtual storage unit matched with the target terminal, wherein the load description information is used to indicate the load status of the virtual storage unit when performing data read and write tasks, and the resource description information is used to indicate the read and write efficiency of the virtual storage unit when performing data read and write tasks.

[0024] S204, determine the priority coefficient of at least one virtual storage unit based on the load description information and resource description information;

[0025] S206, if at least one first virtual storage unit is determined from at least one virtual storage unit based on a priority coefficient, at least one second virtual storage unit that has a routing connection with the first virtual storage unit, and at least one first virtual storage unit are added to the storage unit group.

[0026] S208, determine the target storage unit from all storage unit groups, wherein the target storage unit is used to store the currently acquired object data.

[0027] Optionally, in this embodiment, the target terminal may refer to, but is not limited to, an operating device or edge storage control device that directly faces the user or business system, initiates object data storage requests, and is responsible for receiving user operation instructions or business data and triggering the storage process.

[0028] Optionally, in this embodiment, the virtual storage unit may be, but is not limited to, the smallest storage management unit formed by the application layer after unifying and abstracting heterogeneous storage media (such as SSD, HDD, NVMe, etc.) and storage cluster systems of different protocols (such as NFS, SMB, S3) through a high-speed network. It has independent load characteristics and resource performance parameters and can be dynamically allocated to meet the storage needs of heterogeneous devices at the terminal layer.

[0029] Optionally, in this embodiment, the load description information may be, but is not limited to, an indication of the load status of the virtual storage unit when performing data read and write tasks. It is a set of core indicators for measuring the current workload of the virtual storage unit, including dimensions such as CPU load, memory load, and storage I / O load. The load description information reflects the current busy level of the virtual storage unit. By using the load description information, it is possible to avoid storing object data in virtual storage units with excessive load, preventing problems such as storage response latency and data read / write stuttering, and ensuring the rationality of storage resource scheduling.

[0030] Optionally, in this embodiment, the resource description information may, but is not limited to, indicating the read / write efficiency of the virtual storage unit when performing data read / write tasks. It is a key indicator for measuring the storage performance of the virtual storage unit and may, but is not limited to, include the effective processing power per unit of CPU resource, the effective utilization rate per unit of memory resource, and the data transfer rate, thus determining the efficiency of the virtual storage unit in handling data read / write tasks. The resource description information reflects the storage performance level of the virtual storage unit, providing a performance dimension basis for priority coefficient calculation. This ensures that when selecting virtual storage units, units with high read / write efficiency are given priority, improving the overall speed of object data storage and meeting the storage efficiency requirements in high-performance heterogeneous computing scenarios.

[0031] Optionally, in this embodiment, the priority coefficient can be, but is not limited to, a value derived from the load description information and resource description information of the virtual storage unit through a preset algorithm. This value is used to quantify the priority of the virtual storage unit; a higher value indicates that the unit is more suitable for receiving the target terminal's object data storage task in the current scenario. The priority coefficient provides a unified quantitative standard for selecting virtual storage units, solving the problem of priority ranking among multiple virtual storage units, ensuring that the target terminal can quickly identify the optimal storage unit, and improving the accuracy and efficiency of storage resource selection.

[0032] Optionally, in this embodiment, the first virtual storage unit may refer to, but is not limited to, a virtual storage unit with a relatively high priority selected from at least one virtual storage unit matched with the target terminal based on a priority coefficient, and is a component of the storage unit group.

[0033] To illustrate further, suppose the virtual storage units that match the target terminal are VS1, VS2, VS3, and VS4, with priority coefficients of 0.74, 0.82, 0.44, and 0.56, respectively. If the virtual storage unit with a priority coefficient greater than 0.6 is set as the first virtual storage unit, then VS1 and VS2 will be selected as the first virtual storage unit and become part of the storage unit group.

[0034] Optionally, in this embodiment, the routing connection may refer to, but is not limited to, a communication link established between virtual storage units through a network routing device, which can ensure that different virtual storage units can realize data transmission and state interaction.

[0035] Optionally, in this embodiment, the second virtual storage unit may refer to, but is not limited to, a virtual storage unit that has a routing connection with the first virtual storage unit and has storage capabilities, and is a supplement to the first virtual storage unit.

[0036] Optionally, in this embodiment, the storage unit group may be, but is not limited to, a set of virtual storage units consisting of at least one first virtual storage unit and at least one second virtual storage unit with which it is routed, which is the selection range of target storage units. The virtual storage units in the set have different priorities and storage capabilities, and together provide candidate resources for object data storage of the target terminal.

[0037] Optionally, in this embodiment, object data may refer to, but is not limited to, the specific data that the target terminal currently needs to store, and is the processing object of the storage task.

[0038] Optionally, in this embodiment, the target storage unit may be, but is not limited to, a virtual storage unit selected from the storage unit group and ultimately used to store the currently acquired object data. It must meet conditions such as idle state and performance matching the object data requirements, and is the execution carrier of the storage task.

[0039] Optionally, in this embodiment, all virtual storage units directly connected to the target terminal initiating the storage request are first identified. Then, the load description information and resource description information of these matching virtual storage units are obtained through the application-layer virtualization platform. By accurately filtering virtual storage units that match the target terminal, data collection from irrelevant virtual storage units is avoided, reducing the amount of data processing. At the same time, obtaining the two key types of information—load and resources—ensures that subsequent priority calculations have comprehensive and accurate data basis, laying the foundation for selecting suitable storage units and preventing priority judgment deviations due to missing information.

[0040] Next, based on the acquired load and resource description information, priority values ​​are calculated for each virtual storage unit to obtain a priority coefficient for each unit. The magnitude of the coefficient directly reflects the suitability of the unit for receiving storage tasks in the current scenario. By converting the load and resource characteristics of virtual storage units into priority coefficients, the problem of directly comparing the suitability between different virtual storage units is solved.

[0041] Then, based on the priority coefficient, the filtering rules are set to select the first virtual storage unit with higher priority from the matched virtual storage units; then, through the network routing topology information, the second virtual storage units that have direct routing connections with each first virtual storage unit are identified; finally, the first and second virtual storage units are integrated to form an ordered storage unit group, thus clarifying the selection range of subsequent target storage units.

[0042] Finally, based on the real-time status of each virtual storage unit, the most suitable virtual storage unit for storing the current object data is selected from all storage unit groups and used as the target storage unit, which is then determined as the storage carrier for the current object data.

[0043] Understandably, by first matching virtual storage units and collecting information, then calculating priority coefficients to quantify compatibility, followed by constructing storage unit groups containing core and expansion units, and finally filtering target storage units, the problem of disordered selection of virtual storage units in heterogeneous storage environments can be solved. Through matching mechanisms and priority quantification, it is ensured that the selected storage units are accurately matched with the target terminal and object data requirements, avoiding storage latency or performance waste caused by improper selection. By constructing storage unit groups to define the filtering range, disordered searching in the entire virtual storage resource pool is avoided. At the same time, the candidate range is expanded by combining routing connections, reducing the probability of storage task blocking. Through the abstraction of virtual storage units, the target terminal does not need to pay attention to the differences in underlying storage media and protocols, simplifying the storage resource calling process, reducing the adaptation cost of multi-device collaboration, and providing efficient and stable storage support for high-performance heterogeneous computing.

[0044] The embodiments provided in this application first obtain load description information and resource description information corresponding to at least one virtual storage unit matched with the target terminal. This breaks through the limitations of related technologies that rely solely on remaining storage space or single historical read / write speed for judgment. By combining the load status and read / write efficiency of the virtual storage unit—two key operational status information—a multi-dimensional and comprehensive consideration of its actual operational status is achieved, providing complete data support for subsequent selection and avoiding evaluation bias caused by incomplete information. Next, the priority coefficient of each virtual storage unit is determined based on the above two types of description information. Compared with the subjective judgment of related technologies based on a single dimension, this improves the accuracy of priority determination and lays the foundation for selecting suitable storage units. Subsequently, when the first virtual storage unit is determined, the priority coefficient of each virtual storage unit is determined. The second virtual storage unit and the first virtual storage unit, which are connected by routing, are added to the storage unit group together. This solves the problem of related technologies that only select units that meet the criteria in a single dimension and miss high-quality associated units, thus expanding the range of effective candidate storage units and further improving the selection flexibility. Finally, the target storage unit for storing the current object data is determined from the entire storage unit group. Since the storage unit group has included high-priority units that have been fully evaluated and associated units that are feasible for data transmission, the final selected target storage unit can not only meet the storage requirements of the current object data, but also has the optimal operating state of low load and high read and write efficiency. This solves the technical problem of low accuracy in selecting virtual storage units in related technologies and improves the accuracy of virtual storage unit selection.

[0045] As an optional approach, at least one second virtual storage unit with a routing connection to the first virtual storage unit, and at least one first virtual storage unit, are added to the storage unit group, including:

[0046] According to the priority order indicated by the priority coefficient, traverse at least one virtual memory cell and perform the following operations:

[0047] S1, Obtain a first virtual storage unit from at least one virtual storage unit;

[0048] S2, obtain the set of neighboring units corresponding to the first virtual storage unit, wherein the set of neighboring units includes at least one neighboring storage unit that has a routing connection with the first virtual storage unit;

[0049] S3, determine the neighbor storage unit with the highest priority indicated by the priority coefficient among at least one neighbor storage unit as the current neighbor storage unit;

[0050] S4, if the second priority indicated by the priority coefficient of the current neighbor storage unit is greater than the first priority indicated by the priority coefficient of the first virtual storage unit, the current neighbor storage unit is determined as the second virtual storage unit, and the first virtual storage unit and the second virtual storage unit are added to the storage unit group.

[0051] Optionally, in this embodiment, the neighbor unit set may refer to, but is not limited to, the set of all neighbor storage units that are directly connected to the first virtual storage unit via routing. These neighbor storage units have data interaction capabilities with the first virtual storage unit and are potential candidates for storage unit group expansion.

[0052] Optionally, in this embodiment, the neighboring storage unit may be, but is not limited to, a virtual storage unit that has a direct routing connection with the first virtual storage unit, and possesses its own priority coefficient, load, and energy efficiency characteristics, and can serve as an expansion candidate unit for the storage unit group. The neighboring storage unit can serve as a potential expansion unit for the storage unit group. When the priority of a neighboring storage unit is higher than that of the first virtual storage unit, it can be included in the storage unit group, enhancing the overall performance adaptability of the storage unit group and ensuring that the terminal can select a better storage unit.

[0053] Optionally, in this embodiment, the second virtual storage unit may be, but is not limited to, a virtual storage unit determined by the current neighbor storage unit when the second priority of the current neighbor storage unit is greater than the first priority of the first virtual storage unit, and is an extended member of the storage unit group.

[0054] Optionally, in this embodiment, at least one virtual storage unit is first sorted from high to low priority according to the priority coefficient of the virtual storage unit, and then each virtual storage unit is accessed in sequence according to the sorting. During the access of each virtual storage unit, the virtual storage unit currently traversed is taken as the first virtual storage unit.

[0055] Next, by querying the network routing topology information at the application layer, all neighboring storage units with direct routing connections to the first virtual storage unit are identified. These neighboring storage units are then integrated into a neighboring unit set, clarifying the potential candidate range for storage unit group expansion. By enabling the storage unit group to expand from a single first virtual storage unit to a set containing multiple neighboring storage units, the candidate objects for the storage unit group are enriched.

[0056] Then, for all neighboring storage units in the neighboring unit set, their respective priority coefficients are extracted. By comparing the priority coefficients, the neighboring storage unit with the highest priority coefficient is selected and determined as the current neighboring storage unit. This reduces the workload of priority comparison, eliminating the need to compare each unit in the neighboring unit set with the first virtual storage unit. By simply locking the current neighboring storage unit with the highest priority, it can be determined whether there are any expansion units in the neighboring unit set that can be included in the storage unit group, thus improving expansion efficiency.

[0057] Finally, the priority coefficients of the current neighbor storage unit and the first virtual storage unit are extracted and compared. If the second priority is greater than the first priority, it means that the current neighbor storage unit has a higher adaptability, so it is determined as the second virtual storage unit. Then, the first virtual storage unit and the second virtual storage unit are added to the storage unit group to complete the expansion of the storage unit group. If the second priority is not greater than the first priority, the storage unit group is not expanded.

[0058] Understandably, by extending the storage unit group through routing connections, the limitations of a single unit are overcome, and the number of candidate units is enriched. Even if the first virtual storage unit is busy, the terminal can select an idle second virtual storage unit from the storage unit group, thereby improving the success rate of storage task completion.

[0059] The embodiments provided in this application obtain a first virtual storage unit from at least one virtual storage unit; obtain a set of neighboring units corresponding to the first virtual storage unit, wherein the set of neighboring units includes at least one neighboring storage unit that has a routing connection with the first virtual storage unit; determine the neighboring storage unit with the highest priority indicated by the priority coefficient among the at least one neighboring storage units as the current neighboring storage unit; if the second priority indicated by the priority coefficient of the current neighboring storage unit is greater than the first priority indicated by the priority coefficient of the first virtual storage unit, determine the current neighboring storage unit as the second virtual storage unit, and add the first virtual storage unit and the second virtual storage unit to the storage unit group to expand the storage unit group through routing connection, thereby breaking through the limitation of a single unit, enriching the candidate objects, and even if the first virtual storage unit is busy, the terminal can select an idle second virtual storage unit from the storage unit group, thereby improving the success rate of storage task completion.

[0060] As an optional approach, after determining the current neighboring storage unit as the second virtual storage unit, the following steps are also included:

[0061] S1, determine the second virtual storage unit as the current virtual storage unit;

[0062] S2, obtain the set of neighboring units corresponding to the current virtual storage unit, wherein the set of neighboring units includes at least one neighboring storage unit that has a routing connection with the current virtual storage unit;

[0063] S3, determine the neighbor storage unit with the highest priority indicated by the priority coefficient among at least one neighbor storage unit as the current neighbor storage unit;

[0064] S4, if the third priority indicated by the priority coefficient of the current neighboring storage unit is greater than the fourth priority indicated by the priority coefficient of the current virtual storage unit, then the current neighboring storage unit is determined as the third virtual storage unit, and the third virtual storage unit is added to the storage unit group.

[0065] Optionally, in this embodiment, the third virtual storage unit may be, but is not limited to, a virtual storage unit determined by the current neighbor storage unit when the third priority of the current neighbor storage unit is greater than the fourth priority of the current virtual storage unit, and is a newly added extended member of the storage unit group. By further expanding the scale of the storage unit group and improving its overall adaptability, the storage unit group contains more highly adaptable units, reducing the probability of storage task blocking caused by a single highly adaptable unit being busy.

[0066] Optionally, in this embodiment, the second virtual storage unit determined during the construction of the storage unit group is formally defined as the current virtual storage unit, making this unit the object of the recursive expansion process of the storage unit group.

[0067] Next, by querying the network routing topology data recorded at the application layer, all neighboring storage units that have direct routing connections with the current virtual storage unit are selected, and these neighboring storage units are integrated into a set, namely the neighboring unit set.

[0068] Then, for each neighbor storage unit in the neighbor unit set, extract its corresponding priority coefficient, find the neighbor storage unit with the largest priority coefficient by numerical comparison, and determine that unit as the current neighbor storage unit.

[0069] Finally, the priority coefficients of the current neighboring storage unit and the current virtual storage unit are extracted respectively, and the two priorities are compared. If the third priority is greater than the fourth priority, it is determined as the third virtual storage unit, and then the third virtual storage unit is added to the existing storage unit group. If the third priority is not greater than the fourth priority, no expansion operation is performed.

[0070] To illustrate further, suppose there are two virtual storage units with priorities of 3 and 2. When building storage unit groups, the two virtual storage units with priorities of 3 and 2 are used as the first element to build two storage unit groups. The first storage unit group is {3} and the second storage unit group is {2}.

[0071] For the first storage unit group, starting from the last virtual storage unit with priority 3, find virtual storage units 4, 5, and 6 with higher priority than 3 that have not been added to the first storage unit group from the neighbor set of that virtual storage unit. Select the virtual storage unit corresponding to the highest priority 6 and add it to the first storage unit group. The first storage unit group is {3, 6}. Then, starting from the last virtual storage unit with priority 6, find virtual storage units with higher priority than 6 that have not been added to the first storage unit group from the neighbor set of that virtual storage unit. Continue until there are no virtual storage units with higher priority that have not been added to the first storage unit group in the neighbor set of the last virtual storage unit in the first storage unit group. The first storage unit group is then updated.

[0072] Then, repeat the process for the second storage unit group. Starting from the last virtual storage unit with priority 2, find virtual storage units with higher priorities than 2 that have not been added to the second storage unit group, namely 4 and 7, in the neighbor set of 2. Select the virtual storage unit corresponding to the highest priority 7 and add it to the second storage unit group. The second storage unit group is {2, 7}. Then, starting from the last virtual storage unit with priority 7, find virtual storage units with higher priorities than 7 that have not been added to the second storage unit group in the neighbor set of that virtual storage unit. Continue until there are no virtual storage units with higher priorities that have not been added to the second storage unit group in the neighbor set of the last element in the second storage unit group. The second storage unit group is then updated.

[0073] It is understandable that, in this embodiment, by continuously expanding the size of the storage unit group, the terminal has more highly adaptable options to choose from when selecting a target storage unit. Even if some units are busy, it can still quickly find an idle highly adaptable unit, reducing the probability of storage task blocking. At the same time, the third virtual storage unit included in the storage unit group has a direct routing connection with the current virtual storage unit, which allows data to be transmitted efficiently with the help of existing routing links when using the third virtual storage unit to store data, avoiding increased transmission latency caused by complex routing links.

[0074] The embodiments provided in this application determine the second virtual storage unit as the current virtual storage unit; obtain the set of neighboring units corresponding to the current virtual storage unit, wherein the set of neighboring units includes at least one neighboring storage unit that has a routing connection with the current virtual storage unit; determine the neighboring storage unit with the highest priority indicated by the priority coefficient among the at least one neighboring storage unit as the current neighboring storage unit; if the third priority indicated by the priority coefficient of the current neighboring storage unit is greater than the fourth priority indicated by the priority coefficient of the current virtual storage unit, determine the current neighboring storage unit as the third virtual storage unit, and add the third virtual storage unit to the storage unit group.

[0075] As an optional approach, the load description information and resource description information corresponding to at least one virtual storage unit matching the target terminal are obtained, including:

[0076] Based on the CPU load information, memory load information, and read / write load information of the fourth virtual memory unit in at least one storage unit, the load coefficient corresponding to the fourth virtual memory unit is obtained. The CPU load information is used to indicate the average utilization of the CPU of the fourth virtual memory unit per unit time, the memory load information is used to indicate the remaining available memory of the fourth virtual memory unit, the read / write load information is used to indicate the amount of data read and written by the fourth virtual memory unit per unit time, and the load description information includes the load coefficient.

[0077] Optionally, in this embodiment, the fourth virtual storage unit may be, but is not limited to, a specific virtual storage unit selected from at least one storage unit that requires load factor calculation.

[0078] Optionally, in this embodiment, the CPU load information may be, but is not limited to, used to quantitatively indicate the average utilization of the CPU in the fourth virtual memory unit per unit time.

[0079] Optionally, in this embodiment, the memory load information may be, but is not limited to, an indicator of the remaining available memory capacity of the fourth virtual storage unit, reflecting the degree of idle memory resources of the unit.

[0080] Optionally, in this embodiment, the read / write load information may be, but is not limited to, an indicator of the total amount of read and write data in the fourth virtual storage unit within a unit of time. It is an indicator reflecting the storage I / O busyness of the unit and may be obtained by statistically analyzing the sum of read and write data within a unit of time using I / O monitoring tools.

[0081] Optionally, in this embodiment, the load coefficient is a quantitative value calculated based on the CPU load information, memory load information and read / write load information of the fourth virtual memory unit. It is used to comprehensively reflect the overall load status of the fourth virtual memory unit. The larger the value, the higher the overall load of the unit. It is a core component of the load description information.

[0082] Understandably, the process begins by selecting a fourth virtual storage unit as the target from at least one storage unit in the system application layer. Then, the monitoring module collects the CPU load, memory load, and read / write load information for this unit. Finally, a load factor that comprehensively reflects the overall load of the unit is calculated. Since the load description information includes both the load factor for rapid comparison and the original load information, the scheduling module can efficiently filter low-load units, achieving precise scheduling and avoiding storage resource waste or overload.

[0083] According to the embodiments provided in this application, a load coefficient corresponding to the fourth virtual storage unit is obtained based on the CPU load information, memory load information, and read / write load information of the fourth virtual storage unit in at least one storage unit. The CPU load information indicates the average CPU utilization of the fourth virtual storage unit per unit time, the memory load information indicates the remaining available memory of the fourth virtual storage unit, and the read / write load information indicates the amount of data read and written by the fourth virtual storage unit per unit time. The load description information includes the load coefficient. First, the fourth virtual storage unit is selected as the target object from at least one storage unit in the system application layer. Then, the CPU load information, memory load information, and read / write load information of this unit are collected by the monitoring module. Next, the load coefficient, which comprehensively reflects the overall load of the unit, is calculated. The load description information includes both the load coefficient for fast comparison and the original load information, enabling the scheduling module to efficiently screen low-load units, achieve accurate scheduling, and avoid storage resource waste or overload.

[0084] As an optional approach, based on the CPU load information, memory load information, and read / write load information of the fourth virtual memory unit in at least one storage unit, the load coefficient corresponding to the fourth virtual memory unit is obtained, including:

[0085] S1. Based on the type of central processing unit corresponding to the fourth virtual storage unit, obtain the first load weight, the second load weight and the third load weight. The first load weight is used to indicate the importance of the central processing unit load information, the second load weight is used to indicate the importance of the memory load information, and the third load weight is used to indicate the importance of the read and write load information. The sum of the first load weight, the second load weight and the third load weight satisfies the preset threshold condition.

[0086] S2, based on the first load weight, the second load weight, and the third load weight, the CPU load information, memory load information, and read / write load information are weighted and summed to obtain the load coefficient corresponding to the fourth virtual memory unit.

[0087] Optionally, in this embodiment, the central processing unit type may refer to, but is not limited to, the specific model or architecture of the central processing unit carried by the fourth virtual storage unit. Different central processing unit types have different capabilities in processing computing tasks, memory interaction, IO scheduling, etc., which is the basis for determining the weight allocation of different load information.

[0088] Optionally, in this embodiment, the first load weight is determined based on the CPU type of the fourth virtual memory unit and is used to quantify the weight value of the CPU load information in the load coefficient calculation. The greater the influence of the CPU type on the CPU processing capability, the higher the first load weight, which is one of the parameters for weighted summation calculation.

[0089] Optionally, in this embodiment, the second load weight is determined based on the CPU type of the fourth virtual memory unit and is used to quantify the weight value of the importance of memory load information in the load coefficient calculation. The CPU type is directly related to the compatibility of the memory controller and the memory access speed. The better the compatibility and the faster the access speed, the higher the second load weight.

[0090] Optionally, in this embodiment, the third load weight is determined based on the CPU type of the fourth virtual memory unit and is used to quantify the weight value of the importance of read and write load information in the load coefficient calculation. The influence of the CPU type on IO scheduling capability directly determines the third load weight. The stronger the CPU's IO scheduling capability, the higher the third load weight.

[0091] Optionally, in this embodiment, the preset threshold condition is a numerical condition preset by the system to constrain the sum of the first, second, and third load weights, ensuring that the weight allocation of the three types of load information can fully cover all dimensions of the load coefficient calculation.

[0092] Optionally, in this embodiment, the weighted calculation process of the load coefficient of the fourth virtual storage unit is based on the hardware characteristics to determine the weights, and the weights guide the coefficient calculation. First, according to the type of the central processing unit of the fourth virtual storage unit, the matching first, second and third load weights are obtained from the preset mapping table; then, the three types of load information are normalized, and then weighted and summed according to the weights to finally obtain the load coefficient that can accurately reflect the actual load status of the unit.

[0093] Understandably, since the load weight is dynamically allocated based on the CPU type, the load coefficient calculation can be tailored to the capabilities of different CPUs. For example, high-performance computing CPUs focus on CPU load weight, while I / O-optimized CPUs focus on read and write load weight. This avoids load evaluation distortion caused by uniform weights and ensures that the load coefficient can truly reflect the actual load pressure of the unit.

[0094] According to the embodiments provided in this application, a first load weight, a second load weight, and a third load weight are obtained based on the CPU type corresponding to the fourth virtual storage unit. The first load weight indicates the importance of CPU load information, the second load weight indicates the importance of memory load information, and the third load weight indicates the importance of read / write load information. The sum of the first, second, and third load weights satisfies a preset threshold condition. Based on the first, second, and third load weights, the CPU load information, memory load information, and read / write load information are weighted and summed to obtain the load coefficient corresponding to the fourth virtual storage unit. Since the load weights are dynamically allocated based on the CPU type, the load coefficient calculation can be tailored to the capabilities of different CPUs, avoiding load evaluation distortion caused by uniform weights and ensuring that the load coefficient truly reflects the actual load pressure of the unit.

[0095] As an optional approach, the load description information and resource description information corresponding to at least one virtual storage unit matching the target terminal are obtained, including:

[0096] Based on the CPU energy efficiency information and memory energy efficiency information of the fourth virtual memory unit, the energy efficiency coefficient corresponding to the fourth virtual memory unit is obtained. The CPU energy efficiency information is used to indicate the processing power of the fourth virtual memory unit on a unit CPU, and the memory energy efficiency information is used to indicate the utilization rate of the fourth virtual memory unit on a unit memory resource. The resource description information includes the energy efficiency coefficient.

[0097] Optionally, in this embodiment, the CPU energy efficiency information is used to quantitatively indicate the effective processing capability of the fourth virtual storage unit per unit of CPU resources. This can be measured by dividing the amount of data processed per unit of CPU resources by the number of tasks completed, or by real-time monitoring of the correspondence between CPU resource input and actual processing results through the virtualization platform. The CPU energy efficiency information reflects the utilization efficiency of the fourth virtual storage unit's CPU resources. A stronger processing capability per unit of CPU resources indicates higher CPU energy efficiency, providing key data at the CPU level for calculating the energy efficiency coefficient, ensuring that the energy efficiency coefficient reflects the unit's advantages or disadvantages in computing resource utilization.

[0098] Optionally, in this embodiment, the memory energy efficiency information is used to quantitatively indicate the effective utilization rate of the fourth virtual storage unit on unit memory resources. It can be obtained, but is not limited to, by calculating the amount of active data per unit memory resource per unit time. Here, active data refers to data that has been read, written, or modified more than a threshold number of times per unit time.

[0099] Optionally, in this embodiment, the energy efficiency coefficient is a quantitative value obtained based on the CPU energy efficiency information and memory energy efficiency information of the fourth virtual memory unit. It is used to comprehensively reflect the overall energy efficiency level of the fourth virtual memory unit. The larger the value, the higher the efficiency of the unit in utilizing CPU and memory resources.

[0100] Optionally, in this embodiment, after determining the fourth virtual storage unit whose energy efficiency needs to be evaluated, the CPU energy efficiency information and memory energy efficiency information of the unit are collected through the application layer virtualization platform; then, the energy efficiency coefficient that comprehensively reflects the overall energy efficiency of the unit is calculated.

[0101] Understandably, by integrating the dispersed multi-dimensional information on CPU energy efficiency and memory energy efficiency into a single energy efficiency coefficient, the problem of traditional multi-dimensional performance being difficult to compare directly is solved, enabling the storage resource scheduling module to quickly determine the performance differences of different virtual storage units and improve scheduling decision efficiency.

[0102] The embodiments provided in this application obtain the energy efficiency coefficient corresponding to the fourth virtual storage unit based on its CPU energy efficiency information and memory energy efficiency information. The CPU energy efficiency information indicates the processing power of the fourth virtual storage unit per unit CPU, and the memory energy efficiency information indicates the utilization rate of the fourth virtual storage unit per unit memory resource. The resource description information includes the energy efficiency coefficient. By integrating the dispersed multi-dimensional information on CPU and memory energy efficiency into a single energy efficiency coefficient, the problem of direct comparison of traditional multi-dimensional performance is solved. This enables the storage resource scheduling module to quickly determine the performance differences between different virtual storage units, improving scheduling decision efficiency.

[0103] As an optional approach, before obtaining the energy efficiency coefficient corresponding to the fourth virtual memory unit based on the CPU energy efficiency information and memory energy efficiency information of the fourth virtual memory unit, the following steps are included:

[0104] S1 determines the resource size value by multiplying the number of CPU cores allocated to the fourth virtual memory unit by the preset time.

[0105] S2, the ratio of the number of tasks completed by the central processing unit in the fourth virtual storage unit within a preset time to the resource scale value is determined as the central processing unit energy efficiency information;

[0106] S3 determines the value of the temporary storage scale by multiplying the memory capacity allocated to the fourth virtual storage unit by the preset time.

[0107] S4, the ratio of the amount of active data in the fourth virtual storage unit to the value of the temporary storage size within a preset time is determined as memory energy efficiency information, wherein the amount of active data is the actual storage capacity occupied by data whose number of read and write operations meets the threshold condition.

[0108] Optionally, in this embodiment, the number of CPU cores refers to the number of CPU cores allocated to the fourth virtual storage unit. It is an indicator that measures the scale of CPU computing resources that the unit can call upon. Different virtual storage units are allocated different numbers of cores according to the complexity of the storage tasks they undertake, which directly affects the upper limit of the CPU processing capability of the unit.

[0109] Optionally, in this embodiment, the preset time may be, but is not limited to, a time period pre-set by the system for uniformly calculating energy efficiency information, so as to ensure that the energy efficiency information of different fourth virtual storage units is calculated based on the same time dimension, thereby eliminating the impact of time differences on the energy efficiency evaluation results.

[0110] Optionally, in this embodiment, the resource scale value can be, but is not limited to, a quantized value calculated by multiplying the number of CPU cores of the fourth virtual storage unit by a preset time, and is used to quantify the total scale of CPU resources that the fourth virtual storage unit can call within a preset time.

[0111] Optionally, in this embodiment, the number of tasks called to complete by the central processing unit may refer to, but is not limited to, the total number of storage-related tasks successfully completed by the fourth virtual storage unit within a preset time by calling the allocated central processing unit resources, such as data block writing tasks, data verification tasks, cache synchronization tasks, etc. The more tasks there are, the better the processing results of the central processing unit.

[0112] Optionally, in this embodiment, the temporary storage scale value may be, but is not limited to, a value used to quantify the total scale of memory resources that the fourth virtual storage unit can call upon within a preset time, thereby integrating the resource investment in the two dimensions of memory capacity and time into a single value.

[0113] Optionally, in this embodiment, the active data volume may refer to, but is not limited to, the actual storage capacity occupied by the data in the fourth virtual storage unit that performs read and write operations a number of times within a preset time period that reaches a preset threshold number of times the system has performed read and write operations. The larger the active data volume, the more high-value data is cached in the memory.

[0114] Optionally, in this embodiment, the actual storage capacity may be, but is not limited to, the actual physical storage space occupied in memory.

[0115] Optionally, in this embodiment, the number of CPU cores allocated to the fourth virtual storage unit is first obtained, and then a preset unified time period is extracted. Then, using the number of CPU cores and the preset time, a value representing the total scale of CPU resource investment is calculated and determined as the resource scale value. Next, the total number of storage-related tasks successfully completed by the fourth virtual storage unit within the preset time period by calling the allocated CPU resources is counted, and the ratio of this number of tasks to the resource scale value is determined as the CPU energy efficiency information.

[0116] Furthermore, the memory capacity allocated to the fourth virtual storage unit is obtained, and the system's preset time period is extracted. Then, the temporary storage scale value is obtained using the memory capacity and the preset time. Finally, the ratio of the amount of active data in the fourth virtual storage unit within the preset time period to the temporary storage scale value is determined as the memory energy efficiency information.

[0117] Understandably, since the number of CPU tasks and the amount of active data are based on the real-time running data of the fourth virtual memory unit, it is possible to avoid the evaluation distortion caused by relying on static parameters when obtaining the parameters of the fourth virtual memory unit, so that the evaluation results of the fourth virtual memory unit can match the actual resource utilization status of the fourth virtual memory unit.

[0118] The embodiments provided in this application determine the resource scale value by multiplying the number of CPU cores allocated to the fourth virtual storage unit by a preset time; the CPU energy efficiency information is determined by the ratio of the number of tasks completed by the CPU in the fourth virtual storage unit within the preset time to the resource scale value; the temporary storage scale value is determined by multiplying the memory capacity allocated to the fourth virtual storage unit by a preset time; and the memory energy efficiency information is determined by the ratio of the active data volume in the fourth virtual storage unit within the preset time to the temporary storage scale value, wherein the active data volume is the actual storage capacity occupied by data whose read / write operations meet a threshold condition. Since the number of CPU tasks and the active data volume are based on the real-time operating data of the fourth virtual storage unit, the evaluation distortion caused by relying on static parameters when obtaining the parameters of the fourth virtual storage unit can be avoided, ensuring that the evaluation results of the fourth virtual storage unit match its actual resource utilization status.

[0119] As an optional approach, based on the CPU energy efficiency information and memory energy efficiency information of the fourth virtual memory unit, the energy efficiency coefficient corresponding to the fourth virtual memory unit is obtained, including:

[0120] S1. Based on the type of central processing unit corresponding to the fourth virtual storage unit, obtain the first energy efficiency weight and the second energy efficiency weight, wherein the first energy efficiency weight is used to indicate the importance of central processing unit energy efficiency information, and the second energy efficiency weight is used to indicate the importance of memory energy efficiency information.

[0121] S2, based on the first energy efficiency weight and the second energy efficiency weight, performs a weighted summation of the central processing unit energy efficiency information and the memory energy efficiency information to obtain the energy efficiency coefficient.

[0122] Optionally, in this embodiment, the first energy efficiency weight may be, but is not limited to, a quantized weight determined based on the central processing unit type of the fourth virtual memory unit, used to clarify the importance ratio of central processing unit energy efficiency information in the energy efficiency coefficient calculation.

[0123] Optionally, in this embodiment, the second energy efficiency weight may be, but is not limited to, a quantized weight determined based on the central processing unit type of the fourth virtual memory unit, used to clarify the importance ratio of memory energy efficiency information in the energy efficiency coefficient calculation.

[0124] Optionally, in this embodiment, the specific type of the central processing unit (CPU) carried by the fourth virtual storage unit is first determined, and the first energy efficiency weight and the second energy efficiency weight corresponding to the CPU type are determined.

[0125] Next, the energy efficiency information of the central processing unit and the energy efficiency information of the memory are normalized and calculated. Based on the first energy efficiency weight and the second energy efficiency weight, the energy efficiency information of the central processing unit and the energy efficiency information of the memory are weighted and summed. The weighted sum is then determined as the energy efficiency coefficient of the fourth virtual memory unit. The larger the value, the higher the resource utilization efficiency of the fourth virtual memory unit.

[0126] Understandably, since the energy efficiency weights are dynamically allocated based on the type of CPU, the calculation of the energy efficiency coefficient can be tailored to the capabilities of different CPUs, avoiding the reduction in evaluation accuracy caused by uniform weights. This ensures that the energy efficiency coefficient obtained by weighted summation of CPU energy efficiency information and memory energy efficiency information can truly reflect the actual resource utilization efficiency of the unit.

[0127] According to the embodiments provided in this application, a first energy efficiency weight and a second energy efficiency weight are obtained based on the CPU type corresponding to the fourth virtual storage unit. The first energy efficiency weight indicates the importance of CPU energy efficiency information, and the second energy efficiency weight indicates the importance of memory energy efficiency information. Based on the first and second energy efficiency weights, the CPU energy efficiency information and memory energy efficiency information are weighted and summed to obtain an energy efficiency coefficient. Since the energy efficiency weights are dynamically allocated based on the CPU type, the energy efficiency coefficient calculation can be tailored to the capability characteristics of different CPUs, avoiding the reduced accuracy caused by uniform weights. This ensures that the energy efficiency coefficient obtained by weighting and summing the CPU energy efficiency information and memory energy efficiency information truly reflects the actual resource utilization efficiency of the unit.

[0128] As an optional approach, before obtaining the load description information and resource description information corresponding to at least one virtual storage unit matching the target terminal, the method further includes:

[0129] S1, determine the storage devices whose remaining storage capacity meets the remaining capacity condition and are located in the target space as candidate devices, where the target space is the area where the device density meets the density threshold condition, and the device density is used to indicate the number of storage devices in a unit space;

[0130] S2, traverse each candidate device, and determine the master device and at least one slave device from the currently traversed candidate device and other candidate devices located in the adjacent space, wherein the current candidate device can communicate with other storage devices located in the adjacent space;

[0131] S3, based on the master device and at least one slave device, construct a storage cluster corresponding to the current candidate device, wherein the storage cluster is used to determine at least one virtual storage unit.

[0132] Optionally, in this embodiment, the remaining storage capacity may refer to, but is not limited to, the size of the physical storage space of the storage device that is currently unoccupied and available for storing data, and may be obtained in real time through the capacity monitoring module built into the storage device.

[0133] Optionally, in this embodiment, the remaining capacity condition may be, but is not limited to, a capacity standard preset by the system for screening candidate devices, such as the remaining storage capacity being greater than or equal to twice the amount of data stored by the target terminal in a single instance, to ensure that the candidate devices have sufficient remaining storage capacity.

[0134] Optionally, in this embodiment, the target space may refer to, but is not limited to, a specific physical area in the storage device deployment area where the device density meets the system's preset density threshold condition. It is the spatial screening range for candidate devices, ensuring that candidate devices are screened from the physical space where the device density meets the specific condition.

[0135] Optionally, in this embodiment, device density can be, but is not limited to, a quantitative indicator of the number of storage devices deployed within a unit of physical space, and can be calculated by the ratio of the total number of storage devices in the target area to the physical space area of ​​that area.

[0136] Optionally, in this embodiment, the neighboring space may refer to, but is not limited to, a target space sub-region that is directly adjacent to the physical space where the current candidate device is located and whose device density also meets the density threshold condition.

[0137] Optionally, in this embodiment, the master device may be, but is not limited to, a candidate device selected from the current candidate device and other candidate devices in the adjacent space, responsible for transmitting data with the virtual storage unit and storing the data of the virtual storage unit, and undertaking tasks such as data distribution, load balancing, and fault detection in the storage cluster.

[0138] Optionally, in this embodiment, the slave device may be, but is not limited to, other candidate devices in the storage cluster besides the master device. Under the coordination of the master device, it undertakes specific data storage tasks and accepts status monitoring and task scheduling from the master device. It can form a data redundancy backup relationship with the master device and other slave devices.

[0139] Optionally, in this embodiment, the storage cluster may be, but is not limited to, a collection of storage devices consisting of a master device and at least one slave device. The devices within the cluster are interconnected through a high-speed network, and under the unified management of the master device, they achieve collaborative scheduling of storage resources, data redundancy backup, and load balancing, which is the hardware foundation for generating virtual storage units.

[0140] Optionally, in this embodiment, devices located in a physical space where the device density meets specific conditions are identified as candidate devices. Then, each candidate device is traversed, and a master device and at least one slave device are determined from the currently traversed candidate devices. Based on the master device and at least one slave device, a storage cluster corresponding to the current candidate device is constructed.

[0141] Understandably, devices with insufficient capacity and physical dispersion were excluded during the candidate device selection process. The master-slave device selection stage selected high-performance master devices and sufficient slave devices to ensure that the storage cluster meets the conditions of sufficient storage capacity, efficient processing capabilities, and centralized physical deployment, thereby reducing the latency of data transmission in the cluster. At the same time, master-slave devices enable the cluster to have load balancing and fault redundancy capabilities. If a master device fails, it can be switched over to another slave device to become the new master device. If a slave device fails, data can be recovered through other slave devices, reducing the risk of storage task failure.

[0142] The embodiments provided in this application identify storage devices with remaining storage capacity that meet the remaining capacity condition and are located within a target space as candidate devices. The target space is an area where the device density meets a density threshold condition, and device density indicates the number of storage devices per unit space. Each candidate device is traversed, and a master device and at least one slave device are determined from the currently traversed candidate device and other candidate devices located in adjacent spaces. The current candidate device can communicate with other storage devices located in adjacent spaces. Based on the master device and at least one slave device, a storage cluster corresponding to the current candidate device is constructed. This storage cluster is used to determine at least one virtual storage unit. During the candidate device selection process, devices with insufficient capacity and physical dispersion are excluded. The master-slave device selection stage selects high-performance master devices and sufficient slave devices, ensuring the storage cluster meets the conditions of sufficient storage capacity, efficient processing power, and centralized physical deployment, reducing data transmission latency within the cluster. Simultaneously, the master-slave setup enables the cluster to have load balancing and fault redundancy capabilities. In case of a master device failure, a switchover can be performed, designating another slave device as the new master device. Data can be recovered from a failed slave device through other slave devices, reducing the risk of storage task failure.

[0143] As an optional approach, a master device and at least one slave device are determined from the currently traversed candidate devices and other candidate devices located in the adjacent space, including at least one of the following:

[0144] S1, based on the candidate broadcasts issued by the current candidate device and other candidate devices in the adjacent space, determine the master device and determine the other storage devices besides the master device as slave devices, wherein the candidate broadcast is used to indicate the time when it becomes a candidate device;

[0145] S2, based on the network bandwidth of the current candidate device and other candidate devices in the adjacent space, determine the master device, and determine the other storage devices besides the master device as slave devices;

[0146] S3, based on the failure rates of the current candidate devices and other candidate devices in the adjacent space, determines the master device and identifies the other storage devices besides the master device as slave devices.

[0147] Optionally, in this embodiment, candidate broadcasting may refer to, but is not limited to, a communication signal sent by the current candidate device and other candidate devices in the adjacent space to surrounding devices when participating in the competition for the master device, which includes the time when it becomes a candidate device.

[0148] To illustrate further, suppose candidate device m and candidate device n compete for the master node. Both candidate device m and candidate device n will broadcast in their neighborhood. However, due to the coverage radius limitation, candidate device m can receive the broadcast information of candidate device n, but candidate device n cannot receive the broadcast information of candidate device m. When candidate device m receives the broadcast information of candidate device n, it compares the broadcast information sending time of candidate device n with the broadcast information sending time of its own broadcast information.

[0149] If candidate device n sends its broadcast information earlier than its own broadcast information, then receiving candidate device n becomes the master node, and candidate device m, candidate device n, and the storage devices in the vicinity of candidate device n constitute a storage cluster system.

[0150] If candidate device n sends its broadcast message later than its own broadcast message, then candidate device m considers itself the master node; if candidate device n cannot receive the broadcast message from candidate device m, it also considers itself the master node; candidate device m and candidate device n become the master nodes in their respective storage cluster systems.

[0151] Optionally, in this embodiment, network bandwidth may refer to, but is not limited to, the maximum data transmission rate when the current candidate device and other candidate devices in the adjacent space transmit data with external devices, which is a core indicator for measuring the data transmission capability of a device.

[0152] To illustrate further, suppose there is a current candidate device X and two candidate devices Y and Z in the adjacent space; among them, it is found that the current candidate device X has the largest network bandwidth, so the current candidate device X is determined to be the master device; the candidate devices Y and Z are determined to be slave devices because their bandwidth is lower than that of the current candidate device X.

[0153] Optionally, in this embodiment, the failure rate may refer to, but is not limited to, the probability of failure of the current candidate device and other candidate devices in the adjacent space within a preset statistical period. It is usually obtained as the ratio of the number of failures to the total running time or the ratio of the failure duration to the total running time, which can reflect the stability level of the device. It may be obtained by statistically analyzing data such as device failure time and failure number recorded in the system log.

[0154] Understandably, the three methods of determining the main device correspond to scenarios that prioritize efficiency, performance, and stability, respectively. For example, in high-frequency storage scenarios, network bandwidth can be selected to ensure fast data transmission; in critical business storage scenarios, failure rate can be selected to ensure cluster stability; and in ordinary data storage scenarios, candidate broadcast time can be selected to ensure rapid cluster construction, thereby meeting different storage needs in heterogeneous computing environments.

[0155] The embodiments provided in this application determine the master device based on candidate broadcasts issued by the current candidate device and other candidate devices in the vicinity, and identify other storage devices as slave devices. The candidate broadcasts indicate the time a device becomes a candidate. The master device is also determined based on the network bandwidth of the current candidate device and other candidate devices in the vicinity, and other storage devices are identified as slave devices. Furthermore, the master device is determined based on the failure rate of the current candidate device and other candidate devices in the vicinity, and other storage devices are identified as slave devices. By selecting different master device confirmation methods for different scenarios, different storage needs in heterogeneous computing environments can be met.

[0156] As an optional approach, before obtaining the load description information and resource description information corresponding to at least one virtual storage unit matching the target terminal, the method further includes:

[0157] S1, obtain the data transfer rate corresponding to the fifth virtual storage unit and each storage cluster in at least one virtual storage unit, wherein the data transfer rate corresponding to the fifth virtual storage unit and the storage cluster is used to indicate the time required for the fifth virtual storage unit and the storage cluster to transfer data per unit capacity.

[0158] S2 establishes a mapping relationship between the storage cluster corresponding to the maximum data transfer rate and the fifth virtual storage unit.

[0159] Optionally, in this embodiment, the fifth virtual storage unit may be, but is not limited to, a specific virtual storage unit in at least one virtual storage unit that is to be mapped to the storage cluster.

[0160] Optionally, in this embodiment, the data transmission rate may refer to, but is not limited to, the efficiency indicator when the fifth virtual storage unit transmits data with a certain storage cluster, and may be used to indicate the time required for the two to transmit a unit capacity of data.

[0161] Optionally, in this embodiment, the mapping relationship may be, but is not limited to, a logical association between the fifth virtual storage unit established at the protocol layer and a specific storage cluster. Through this relationship, the storage tasks of the fifth virtual storage unit can be directly routed to the corresponding storage cluster, and the storage cluster can also feed back the data read and write status to the fifth virtual storage unit, thereby providing a communication link between the application layer virtual storage unit and the hardware layer storage cluster, so that the storage needs of the fifth virtual storage unit can be transformed into storage operations of the storage devices in the storage cluster.

[0162] Optionally, in this embodiment, a transmission rate testing tool is first used to send test data of unit capacity between the fifth virtual storage unit and each storage cluster; the time required to transmit each unit of data is recorded, and the time is converted into the data transmission rate, thus obtaining the transmission rate data corresponding to the fifth virtual storage unit and each storage cluster.

[0163] Next, the highest data transfer rate is selected from the data transfer rates between the fifth virtual storage unit and each storage cluster by numerical comparison. Then, the specific storage cluster corresponding to the highest rate is determined. Finally, through the mapping management module of the protocol layer, a logical association is established between the fifth virtual storage unit and the storage cluster to ensure that the subsequent storage tasks of the fifth virtual storage unit can be automatically routed to the storage cluster, thereby achieving efficient data storage.

[0164] Understandably, by conducting rate tests and selecting the best option, the fifth virtual storage unit is matched with the storage cluster that has the fastest data transmission speed, minimizing data transmission latency, avoiding storage task lag caused by low transmission efficiency, and improving the overall response speed of the storage system.

[0165] The embodiments provided in this application obtain the data transfer rate corresponding to each storage cluster and the fifth virtual storage unit in at least one virtual storage unit. The data transfer rate corresponding to the fifth virtual storage unit and the storage cluster indicates the time required for the fifth virtual storage unit to transfer data per unit capacity with the storage cluster. A mapping relationship is established between the storage cluster corresponding to the maximum data transfer rate and the fifth virtual storage unit. Through rate testing and optimal selection, the fifth virtual storage unit is matched with the storage cluster with the fastest data transfer speed, minimizing data transfer latency, avoiding storage task lag due to low transmission efficiency, and improving the overall response speed of the storage system.

[0166] As an optional approach, the target storage unit is determined from the entire group of storage units, including:

[0167] Based on the priority coefficient of each of the at least one virtual memory unit, traverse all memory unit groups and perform the following operations:

[0168] S1, if there are free storage units in the current storage unit group, the free storage unit is determined as the target storage unit, wherein the free storage unit is a virtual storage unit that is not in a working state;

[0169] S2, if there are no free storage units in the current storage unit group, determine the next storage unit group as the current storage unit group.

[0170] Optionally, in this embodiment, an idle storage unit may be, but is not limited to, a virtual storage unit that is not in a working state within the storage unit group. It may be understood, but is not limited to, a unit that is not currently undertaking any data read / write tasks and whose CPU load, memory load, and I / O load are all at a low level, and has the ability to immediately undertake new storage tasks.

[0171] Optionally, in this embodiment, the priority coefficient of each virtual storage unit in all storage unit groups is first extracted, and all storage unit groups are sorted. Then, according to the sorting result, each storage unit group is accessed sequentially, starting from the storage unit group with the highest priority. It is then determined whether there are any idle storage units that are not in a working state. If so, the unit with the highest priority coefficient is selected from these idle storage units and identified as the target storage unit, thus completing the filtering process.

[0172] If all virtual storage units in the current storage unit group are found to be in a working state, the next group of the current storage unit group (which can be understood as a group with a priority second only to the current group) is updated to a new current storage unit group according to the determined storage unit group traversal order, and the new current group is re-checked to see if there are any free storage units.

[0173] Understandably, based on the priority coefficient of the virtual storage units, all storage unit groups are sorted from high to low according to their adaptability, and then each storage unit group is traversed in this order. If there is a free storage unit in the current group that is not in a working state, it is identified as the target storage unit. If there is no free unit in the current group, the current group is updated to the next storage unit group in order, and the check operation is repeated until the target storage unit is found. By always prioritizing the selection of free units from high-priority storage unit groups, it is ensured that the identified target storage unit is both in a free state and has high adaptability, thereby improving the execution efficiency and stability of storage tasks.

[0174] According to the embodiments provided in this application, if there are free storage units in the current storage unit group, the free storage unit is determined as the target storage unit, wherein the free storage unit is a virtual storage unit that is not in a working state; if there are no free storage units in the current storage unit group, the next storage unit group is determined as the current storage unit group.

[0175] As an optional solution, in order to better understand the process of the above data storage method, the flow of the execution of the above data storage method will be described below in conjunction with optional embodiments, but it is not intended to limit the technical solution of the embodiments of this application.

[0176] This embodiment provides a collaborative management method and system for storage devices based on heterogeneous computing. By constructing a four-layer architecture comprising a terminal layer, application layer, protocol layer, and hardware layer, this embodiment addresses the issues of hardware interconnection, protocol compatibility, virtual resource integration, and efficient scheduling in storage systems under heterogeneous computing environments. It achieves deep collaboration between storage and computing resources, providing underlying support for high-performance heterogeneous computing.

[0177] Step 1: Build a four-layer architecture for the storage device management system: terminal layer, application layer, protocol layer, and hardware layer;

[0178] The terminal layer consists of user-facing operating devices and edge storage devices, such as displays, keyboards, mice, and user local storage controllers. The user local storage controller is mainly a lightweight storage controller responsible for the preprocessing and caching of local data.

[0179] The application layer is responsible for interconnecting the underlying storage devices through high-speed networks to form a virtual storage resource pool, which meets the needs of heterogeneous devices at the terminal layer. The smallest management unit of the virtual storage resource pool is the virtual storage unit. For example, virtual storage units are dynamically allocated according to the storage needs of the terminal layer to achieve efficient utilization of storage resources.

[0180] The protocol layer supports protocol compatibility and is responsible for establishing a mapping between storage cluster systems with different protocols from the hardware layer and virtual storage units;

[0181] The hardware layer supports interconnection of heterogeneous storage media interfaces and supports mixed interfaces of CXL3.0 and PCIe6.0; it is responsible for building storage devices with heterogeneous storage media interfaces into a storage cluster system for management.

[0182] Step 2: The virtual storage units in the application layer are prioritized according to load and energy efficiency;

[0183] The load of a virtual storage unit is related to the CPU load, memory load, and storage I / O load; the energy efficiency of a virtual storage unit is related to the CPU energy efficiency and memory energy efficiency.

[0184] The specific calculation method is as follows: In the application layer, the virtualization platform monitors the CPU load, memory load and storage I / O load of the virtual storage unit in real time. The load index of the virtual storage unit is expressed as shown in formula (1):

[0185] (1);

[0186] in, Represents virtual storage unit The load index, Indicates CPU load factor; Represents virtual storage unit The CPU load is obtained by calculating the average CPU utilization per unit time. Indicates the memory load factor. Represents virtual storage unit The memory load is obtained by calculating the proportion of used memory to the total allocated memory. Indicates the storage I / O load factor. Represents virtual storage unit The storage I / O load is obtained by calculating the amount of data read and written per unit time; the CPU load coefficient, memory load coefficient, and storage I / O load coefficient satisfy... ;

[0187] In the application layer, the virtualization platform monitors the CPU and memory energy efficiency of the virtual storage unit in real time. The energy efficiency index of the virtual storage unit is expressed as shown in formula (2):

[0188] (2);

[0189] in, Represents virtual storage unit Energy efficiency indicators Indicates the CPU energy efficiency coefficient. Represents virtual storage unit CPU energy efficiency is obtained by calculating the effective processing capacity of a unit of CPU resources within a virtual memory unit. Among these... The method for obtaining it is shown in formula (3):

[0190] (3);

[0191] Indicates the memory energy efficiency coefficient. Represents virtual storage unit The memory efficiency is obtained by calculating the effective utilization rate of virtual memory unit memory resources. The above coefficients are specified according to the CPU model, and can also be considered as determined by experiments. The method for obtaining it is shown in formula (4):

[0192] (4);

[0193] It should be noted that data that is read, written, or modified more than a threshold number of times within a specified time, or in other words, during the runtime, is called active data; the amount of this active data is calculated based on storage space (e.g., how many bytes or how many kilobytes of data).

[0194] For virtual storage units Based on load and energy efficiency indicators, the priority function is calculated as shown in formula (5):

[0195] (5);

[0196] in, Represents virtual storage unit Priority function, This represents the minimum load metric for virtual storage units. This indicates the maximum load metric for the virtual storage unit. This represents the minimum energy efficiency index of a virtual storage unit. This represents the maximum energy efficiency index of the virtual storage unit;

[0197] and The load and energy efficiency indicators are calculated based on the current status of the virtual unit.

[0198] The priority function for all virtual memory units is obtained using the calculation method described above.

[0199] Step 3: When the terminal has data to store, it stores it according to the priority of the virtual storage units, as follows:

[0200] When a terminal needs to store data, the virtual storage units directly connected to the terminal are sorted from high to low according to a priority function to construct an ordered set. ,in, This indicates that the virtual storage units directly connected to the terminal are sorted from high to low according to a priority function. Indicates the number of virtual storage units directly connected to the terminal;

[0201] Each virtual storage unit directly connected to the terminal As a starting point, generate group sequences. ;

[0202] When the terminal has data to store, the terminal retrieves it from the group sequence. Begin by examining each group sequence in turn:

[0203] S1 terminal pair group Check each element in sequence. ;

[0204] S2 if If available, select Store the data and mark it as busy;

[0205] S3 if If busy, continue with the inspection team. The next element inside;

[0206] If the current group If all elements within the sequence are busy, then proceed to the next priority group sequence. Repeat step S1;

[0207] If all elements of all groups are busy, the terminal enters a waiting state until a virtual memory unit becomes free.

[0208] Among them, group sequence The generation method is as follows: group The first element is the virtual storage unit directly connected to the terminal. Then, a recursive method is used to find virtual memory units with higher priority for the group. To expand;

[0209] Among them, group The specific methods for expansion are as follows:

[0210] S11: From group The last element Depart, gather at its neighbors In the middle, find all those with higher priority than And not joined the group Virtual storage units;

[0211] S12: If multiple virtual storage units meet the criteria, select the virtual storage unit with the highest priority and add it to the group. The end;

[0212] S13: Repeat step S11 until the group... The last element Of all the neighbors in the set, there is no higher priority neighbor that is not in the group. Up to the point of virtual storage units;

[0213] in, Neighbor set It refers to and A collection of directly routed and connected virtual storage units;

[0214] To illustrate further, suppose a set contains three virtual storage units with priorities of 3, 2, and 1. When creating a group, the three virtual storage units with priorities of 3, 2, and 1 are used as the first element to create three groups. Now, group1 contains {3}, group2 contains {2}, and group3 contains {1}.

[0215] Then, for group1, starting from the last element 3, find virtual storage units 4, 5, and 6 with higher priority than 3 that have not been added to group1 in the neighbor set of 3. Select the highest priority 6 and add it to group1. group1 is {3, 6}. Then, starting from the last element 6, find virtual storage units with higher priority than 6 that have not been added to group1 in the neighbor set of 6. Continue until there are no virtual storage units with higher priority that have not been added to group1 in the neighbor set of the last element in group1. At this point, group1 is updated.

[0216] Then, repeat the process for group2. Starting from the last element 2, find virtual storage units 4 and 7 in the neighbor set of 2 that have a higher priority than 2 and have not been added to group1 (the neighbor set may have duplicates, and multiple groups can be added if the condition is met). Select the highest priority 7 and add it to group2, so group2 is {2, 7}. Then, starting from the last element 7, find virtual storage units in the neighbor set of 7 that have a higher priority than 7 and have not been added to group2, until there are no virtual storage units with a higher priority than 7 that have not been added to group2 in the neighbor set of the last element in group2. At this point, group2 is updated.

[0217] Then, repeat the process for group3. Starting from the last element 1, find the virtual storage unit 6 with a higher priority than 1 that has not been added to group3 in the neighbor set of 1 (the neighbor set may have duplicates, and multiple groups can be added if the condition is met). Select the highest priority 6 and add it to group3, so group3 is {1, 6}; (multiple groups can also be added). Then, starting from the last element 6, find the virtual storage unit with a higher priority than 6 that has not been added to group3 in the neighbor set of 6, until there is no virtual storage unit with a higher priority that has not been added to group3 in the neighbor set of the last element in group3. At this point, group3 is updated.

[0218] Step 4: Hardware layer, building a storage cluster system based on storage devices;

[0219] In each storage cluster system, one storage device is selected as the master node, and the other storage devices are selected as slave nodes;

[0220] The master node is responsible for transmitting data with the virtual storage unit and storing the data in the virtual storage unit;

[0221] The master node sends the data from the virtual storage unit to the slave nodes within the storage cluster system, and the slave nodes back up and store the data.

[0222] The process of determining the master node in a storage cluster system:

[0223] First, candidate nodes in the storage cluster system are determined based on the storage device density and the remaining storage capacity of surrounding storage devices;

[0224] For storage devices Set threshold function ,exist At that time, as shown in formula (6):

[0225] (6);

[0226] For storage devices Set threshold function ,exist When, as shown in formula (7):

[0227] (7);

[0228] in, and These represent the storage device density factor and the remaining storage capacity factor of surrounding storage devices, respectively. Indicates storage device Equipment density at the location This indicates the maximum device density at the location of the storage device. Indicates storage device The remaining storage capacity, Indicates storage device The average remaining storage capacity of surrounding storage devices; device density is calculated based on the number of devices in a fixed physical space, which is the device placement density.

[0229] For storage devices distribute A random number between [a certain range]; if the random number is less than [a certain value] Then the storage device The higher the density of the storage device and the larger the remaining storage capacity, the higher the threshold and the higher the probability of it being a candidate node.

[0230] Then, each candidate node competes with candidate nodes in its neighboring region for the master node;

[0231] Once a storage device is selected as a candidate node, it is broadcast within the neighborhood.

[0232] If other storage devices receive a broadcast message from the candidate node, they will confirm that the candidate node is the cluster master node and will no longer become a candidate node themselves.

[0233] If other candidate nodes receive the broadcast message from this candidate node, they compare the sending time; the candidate node with the earliest sending time is selected as the cluster master node; the remaining candidate nodes withdraw from the competition and no longer broadcast.

[0234] Candidate Nodes With candidate nodes When the communication link length is less than or equal to the coverage radius, as shown in expression (8):

[0235] (8);

[0236] Candidate nodes With candidate nodes Located in the adjacent area; among which, Indicates candidate nodes The coverage radius, Indicates candidate nodes The coverage radius, This indicates taking the maximum value. Indicates candidate nodes With candidate nodes The length of the communication link;

[0237] Among them, candidate nodes Coverage radius The calculation method is shown in formula (9):

[0238] (9);

[0239] in, This indicates the maximum communication link length between candidate nodes. Indicates candidate nodes Equipment density at the location Indicates candidate nodes The remaining storage capacity, This represents the maximum remaining storage capacity of the candidate nodes. This represents the minimum remaining storage capacity of the candidate node. Indicates the maximum coverage radius. , , These represent the candidate node communication distance coefficient, device density coefficient, and storage capacity coefficient, respectively.

[0240] The device density is calculated based on the number of devices within a fixed physical space; the maximum remaining storage capacity is the storage device with the largest remaining storage capacity among the candidate nodes; the minimum remaining storage capacity is the storage device with the smallest remaining storage capacity among the candidate nodes.

[0241] Multiple storage devices are physically connected, and then storage cluster management software is used to form a storage cluster system. A storage cluster system has a master node and other storage devices are slave nodes; this storage cluster system communicates with the outside world through the master node.

[0242] Multiple storage devices are pre-positioned. Then, multiple candidate nodes are selected from these devices. These candidate nodes compete for the master node. The storage device that wins the master node broadcasts its status as the master node to its neighboring region. The storage device corresponding to the master node and all storage devices in the neighboring region together form a storage cluster system. Multiple storage cluster systems exist.

[0243] The master node is at the hardware layer, while the group is what the application layer sees. The virtual storage units stored in the group are mapped in the fifth step of the protocol layer based on the routing information between the storage cluster system and the virtual storage units, which means that the master node and the virtual storage units are bound together.

[0244] Step 5: The protocol layer establishes a mapping between the storage cluster systems with different protocols from the hardware layer and the virtual storage units, and the corresponding storage cluster systems store the contents of the virtual storage units;

[0245] Based on the speed at which virtual storage units transmit data to the storage cluster system, select the storage cluster system with the highest data transmission speed for each virtual storage unit.

[0246] A mapping function is established based on the routing information (link capacity, link length) between the storage cluster system and the virtual storage unit, as shown in formula (10):

[0247] (10);

[0248] in, This represents the mapping function between the storage cluster system and virtual storage units. This indicates taking the minimum value of the function. This represents a node pair consisting of a virtual storage unit and a storage cluster system. Represents a set of node pairs. This indicates the rate at which data from virtual storage units in a node pair is stored in the storage cluster system. This represents the storage link between virtual storage units and storage cluster systems. Represents a set of storage links. Indicates storage link Storage link capacity, Indicates message data, This indicates all data to be stored. Represents message data arrival rate, Indicates the index function;

[0249] Set the constraints as shown in expression (11):

[0250] (11)

[0251] Through optimization Determine the mapping between virtual storage units and storage device clusters.

[0252] When a terminal needs to store data, it stores the data according to the priority of the virtual storage unit.

[0253] Experiments demonstrate that the storage method proposed in this embodiment is superior to other methods, including traditional methods, breadth-first search methods, and random methods.

[0254] In the traditional method, the terminal only accesses directly connected virtual storage units, selecting them from highest to lowest priority, without using route expansion to find indirectly connected virtual storage units. In the breadth-first search algorithm, priority is not considered; starting from the virtual storage units directly connected to the terminal, all reachable virtual storage units are searched in order of distance from the terminal (breadth-first), and the first free virtual storage unit found is selected first. In the random method, virtual storage units are randomly selected from all reachable virtual storage units of the terminal, without considering priority or distance.

[0255] In addition, the parameters include the number of virtual storage units, priority, connectivity, busy state, and number of storage requests. There are 50 virtual storage units, divided into those directly connected to the terminal and those indirectly connected. Priority is randomly generated for each virtual storage unit. Connectivity is an undirected connected graph with an average connectivity of 3 for each virtual unit. Each virtual storage unit is busy with a probability of 0.6 when a storage request arrives. The number of storage requests is calculated using 1000 Monte Carlo simulations, each independent.

[0256] The evaluation metrics include response time and invalid search rate. Response time represents the number of search steps from when the terminal initiates a storage request to when a free virtual storage unit is found. Invalid search rate represents the proportion of busy virtual storage units detected during the search process.

[0257] Among the obtained data, the response time of the method in this embodiment is 8 and the invalid search rate is 20%, the response time of the traditional method is 14 and the invalid search rate is 48%, the response time of the breadth-first method is 15 and the invalid search rate is 46%, and the response time of the random method is 21 and the invalid search rate is 63%.

[0258] Experiments show that the method in this embodiment requires the shortest response time and has the lowest invalid search rate. This demonstrates that the method in this embodiment can help the terminal quickly find free virtual storage units for storage.

[0259] This application's embodiments address the complex scenario of collaborative management of storage devices in heterogeneous computing environments. Through an innovative four-layer architecture and intelligent scheduling mechanism, it achieves efficient integration and collaborative optimization of storage resources. By constructing a virtual storage resource pool at the application layer, it unifies heterogeneous storage media (such as SSDs, HDDs, NVMe, etc.) and storage cluster systems using different protocols (such as NFS, SMB, S3) into virtual storage units, completely resolving compatibility issues with heterogeneous devices at the terminal layer. Compared to traditional heterogeneous management solutions, terminal device access time is significantly shortened, significantly reducing the adaptation cost of multi-device collaboration. The application layer dynamically sets the priority of virtual storage units based on load pressure and energy efficiency ratio. Terminals prioritize high-performance, cost-effective units (balancing priority and energy consumption) for storage, avoiding hotspot overload or inefficient resource idleness problems caused by traditional fixed-priority strategies. Real-world testing shows that under peak load scenarios, the median storage request response time decreased from 80ms to 35ms, resource utilization improved by 65%, effectively balancing system performance and energy consumption. The protocol layer establishes a dynamic mapping relationship between the storage cluster system and virtual storage units, optimizing data transmission paths in real time based on network routing status, significantly reducing cross-protocol and cross-media data interaction overhead. Compared with traditional fixed mapping schemes, data transmission latency is reduced, average network traffic per node is lowered, and total data transmission costs decrease. The hardware layer constructs a highly reliable cluster management system from heterogeneous storage media, achieving secure data storage without affecting normal business operations through a multi-replica redundancy mechanism. Compared with traditional homogeneous backup schemes, the data backup time of the cluster management system of this invention is greatly reduced.

[0260] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0261] Based on this understanding, the technical solution of this application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0262] This embodiment also provides a data storage device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0263] Figure 3 This is a structural block diagram of a data storage device according to an embodiment of this application; as shown below. Figure 3 As shown, it includes:

[0264] The first acquisition unit 302 is used to acquire load description information and resource description information corresponding to at least one virtual storage unit matched with the target terminal, wherein the load description information is used to indicate the load status of the virtual storage unit when performing data read and write tasks, and the resource description information is used to indicate the read and write efficiency of the virtual storage unit when performing data read and write tasks.

[0265] The first determining unit 304 is used to determine the priority coefficient of at least one virtual storage unit based on the load description information and the resource description information.

[0266] Adding unit 306 is used to add at least one second virtual storage unit that has a routing connection with the first virtual storage unit, and at least one first virtual storage unit to the storage unit group, when at least one first virtual storage unit is determined from at least one virtual storage unit based on a priority coefficient.

[0267] The second determining unit 308 is used to determine the target storage unit from all storage unit groups, wherein the target storage unit is used to store the currently acquired object data.

[0268] As an optional solution, adding unit 306 includes: a first traversal module, used to traverse at least one virtual storage unit according to the priority order indicated by the priority coefficient; a first acquisition module, used to acquire a first virtual storage unit from at least one virtual storage unit; a second acquisition module, used to acquire a set of neighboring units corresponding to the first virtual storage unit, wherein the set of neighboring units includes at least one neighboring storage unit that has a routing connection with the first virtual storage unit; a first determination module, used to determine the neighboring storage unit with the highest priority indicated by the priority coefficient among the at least one neighboring storage units as the current neighboring storage unit; and a second determination module, used to determine the current neighboring storage unit as the second virtual storage unit if the second priority indicated by the priority coefficient of the current neighboring storage unit is greater than the first priority indicated by the priority coefficient of the first virtual storage unit, and add the first virtual storage unit and the second virtual storage unit to the storage unit group.

[0269] As an optional solution, the second determining module includes: a first determining submodule, used to determine the second virtual storage unit as the current virtual storage unit; a first obtaining submodule, used to obtain a set of neighboring units corresponding to the current virtual storage unit, wherein the set of neighboring units includes at least one neighboring storage unit that has a routing connection with the current virtual storage unit; the first determining submodule, used to determine the neighboring storage unit with the highest priority indicated by the priority coefficient among the at least one neighboring storage unit as the current neighboring storage unit; and the second determining submodule, used to determine the current neighboring storage unit as the third virtual storage unit and add the third virtual storage unit to the storage unit group if the third priority indicated by the priority coefficient of the current neighboring storage unit is greater than the fourth priority indicated by the priority coefficient of the current virtual storage unit.

[0270] As an optional solution, the first acquisition unit 302 includes: a third acquisition module, used to acquire the load coefficient corresponding to the fourth virtual storage unit based on the CPU load information, memory load information and read / write load information of the fourth virtual storage unit in at least one storage unit, wherein the CPU load information is used to indicate the average utilization rate of the CPU of the fourth virtual storage unit in a unit time, the memory load information is used to indicate the remaining available memory of the fourth virtual storage unit, the read / write load information is used to indicate the amount of data read and written by the fourth virtual storage unit in a unit time, and the load description information includes the load coefficient.

[0271] As an optional solution, the third acquisition module includes: a second acquisition submodule, used to acquire a first load weight, a second load weight, and a third load weight based on the CPU type corresponding to the fourth virtual storage unit, wherein the first load weight is used to indicate the importance of CPU load information, the second load weight is used to indicate the importance of memory load information, and the third load weight is used to indicate the importance of read / write load information, and the sum of the first load weight, the second load weight, and the third load weight satisfies a preset threshold condition; and a first summation submodule, used to perform a weighted summation based on the first load weight, the second load weight, and the third load weight, the CPU load information, the memory load information, and the read / write load information to obtain the load coefficient corresponding to the fourth virtual storage unit.

[0272] As an optional solution, the first acquisition unit 302 includes: a fourth acquisition module, used to acquire the energy efficiency coefficient corresponding to the fourth virtual storage unit based on the CPU energy efficiency information and memory energy efficiency information of the fourth virtual storage unit, wherein the CPU energy efficiency information is used to indicate the processing power of the fourth virtual storage unit on a unit CPU, the memory energy efficiency information is used to indicate the utilization rate of the fourth virtual storage unit on a unit memory resource, and the resource description information includes the energy efficiency coefficient.

[0273] As an optional solution, the fourth acquisition module includes: a third determining submodule, used to determine the resource scale value by multiplying the number of CPU cores allocated to the fourth virtual storage unit by a preset time; a fourth determining submodule, used to determine the CPU energy efficiency information by the ratio of the number of tasks completed by the CPU in the fourth virtual storage unit within the preset time to the resource scale value; a fifth determining submodule, used to determine the temporary storage scale value by multiplying the memory capacity allocated to the fourth virtual storage unit by a preset time; and a sixth determining submodule, used to determine the memory energy efficiency information by the ratio of the active data volume in the fourth virtual storage unit within the preset time to the temporary storage scale value, wherein the active data volume is the actual storage capacity occupied by data whose number of read / write operations meets the threshold condition.

[0274] As an optional solution, the fourth acquisition module includes: a third acquisition submodule, used to acquire a first energy efficiency weight and a second energy efficiency weight based on the CPU type corresponding to the fourth virtual storage unit, wherein the first energy efficiency weight is used to indicate the importance of CPU energy efficiency information and the second energy efficiency weight is used to indicate the importance of memory energy efficiency information; and a second summation submodule, used to perform weighted summation of CPU energy efficiency information and memory energy efficiency information based on the first energy efficiency weight and the second energy efficiency weight to obtain an energy efficiency coefficient.

[0275] As an optional solution, the apparatus further includes: a third determining unit, configured to determine storage devices whose remaining storage capacity meets the remaining capacity condition and are located within a target space as candidate devices, wherein the target space is an area where the device density meets a density threshold condition, and the device density is used to indicate the number of storage devices per unit space; a fourth determining unit, configured to traverse each candidate device and determine a master device and at least one slave device from the currently traversed candidate device and other candidate devices located in the adjacent space, wherein the current candidate device can communicate with other storage devices located in the adjacent space; and a building unit, configured to build a storage cluster corresponding to the current candidate device based on the master device and at least one slave device, wherein the storage cluster is used to determine at least one virtual storage unit.

[0276] As an optional scheme, the fourth determining unit includes: a third determining module, used to determine the master device based on the candidate broadcasts issued by the current candidate device and other candidate devices in the adjacent space, and to determine other storage devices besides the master device as slave devices, wherein the candidate broadcasts are used to indicate the time when a device becomes a candidate device; a fourth determining module, used to determine the master device based on the network bandwidth of the current candidate device and other candidate devices in the adjacent space, and to determine other storage devices besides the master device as slave devices; and a fifth determining module, used to determine the master device based on the failure rate of the current candidate device and other candidate devices in the adjacent space, and to determine other storage devices besides the master device as slave devices.

[0277] As an optional solution, the apparatus further includes: a second acquisition unit, configured to acquire the data transfer rate corresponding to the fifth virtual storage unit and each storage cluster in at least one virtual storage unit, wherein the data transfer rate corresponding to the fifth virtual storage unit and the storage cluster is used to indicate the time required for the fifth virtual storage unit and the storage cluster to transfer data per unit capacity; and an establishment unit, configured to establish a mapping relationship between the storage cluster corresponding to the maximum data transfer rate and the fifth virtual storage unit.

[0278] As an optional solution, the second determining unit 308 includes: a second traversal module, used to traverse all storage unit groups according to the priority coefficient of each of at least one virtual storage unit; a sixth determining module, used to determine the free storage unit as the target storage unit when there is a free storage unit in the current storage unit group, wherein the free storage unit is a virtual storage unit that is not in a working state; and a seventh determining module, used to determine the next storage unit group as the current storage unit group when there is no free storage unit in the current storage unit group.

[0279] For a description of the features in the embodiment corresponding to the data storage device, please refer to the relevant description in the embodiment corresponding to the data storage method, which will not be repeated here.

[0280] Embodiments of this application also provide an electronic device. Figure 4 This is a schematic diagram of an electronic device according to an embodiment of this application, such as... Figure 4 As shown, the electronic device includes a memory and a processor, the memory storing a computer program, and the processor being configured to run the computer program to perform the steps in any of the above-described data storage method embodiments.

[0281] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0282] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0283] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described data storage method embodiments when it is run.

[0284] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0285] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods in various embodiments of this application; the computer program product further includes a non-volatile computer-readable storage medium that stores the computer program, which, when executed by a processor, implements the steps of the data storage methods in various embodiments of this application.

[0286] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0287] The data storage method provided in this application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A data storage method, characterized by, The method comprises the following steps: obtaining load description information and resource description information corresponding to each of at least one virtual storage unit matched with a target terminal, wherein the load description information is used to indicate the load condition of the virtual storage unit when performing a data read / write task, and the resource description information is used to indicate the read / write efficiency of the virtual storage unit when performing the data read / write task; determining a priority coefficient of each of the at least one virtual storage unit according to the load description information and the resource description information; in a case where at least one first virtual storage unit is determined from the at least one virtual storage unit based on the priority coefficient, traversing the at least one virtual storage unit in a priority order indicated by the priority coefficient, and performing the following operations: obtaining the first virtual storage unit from the at least one virtual storage unit; obtaining a neighbor unit set corresponding to the first virtual storage unit, wherein the neighbor unit set comprises at least one neighbor storage unit having a routing connection with the first virtual storage unit; determining a neighbor storage unit having a maximum priority indicated by the priority coefficient from the at least one neighbor storage unit as a current neighbor storage unit; in a case where a second priority indicated by the priority coefficient of the current neighbor storage unit is greater than a first priority indicated by the priority coefficient of the first virtual storage unit, determining the current neighbor storage unit as a second virtual storage unit, and adding the first virtual storage unit and the second virtual storage unit to a storage unit group; determining a target storage unit from the entire storage unit group based on the priority coefficient, wherein the target storage unit is used to store currently obtained object data.

2. The method of claim 1, wherein, After determining the current neighbor storage unit as the second virtual storage unit, the method further comprises the following steps: determining the second virtual storage unit as a current virtual storage unit; obtaining a neighbor unit set corresponding to the current virtual storage unit, wherein the neighbor unit set comprises at least one neighbor storage unit having a routing connection with the current virtual storage unit; determining a neighbor storage unit having a maximum priority indicated by the priority coefficient from the at least one neighbor storage unit as a current neighbor storage unit; in a case where a third priority indicated by the priority coefficient of the current neighbor storage unit is greater than a fourth priority indicated by the priority coefficient of the current virtual storage unit, determining the current neighbor storage unit as a third virtual storage unit, and adding the third virtual storage unit to the storage unit group.

3. The method of claim 1, wherein, The method for obtaining the load description information and the resource description information corresponding to each of the at least one virtual storage unit matched with the target terminal comprises the following steps: The load coefficient corresponding to the fourth virtual storage unit is obtained based on central processor load information, memory load information and read-write load information of the fourth virtual storage unit in the at least one storage unit, wherein the central processor load information is used to indicate the average utilization rate of the central processor of the fourth virtual storage unit in a unit time, the memory load information is used to indicate the remaining available memory of the fourth virtual storage unit, and the read-write load information is used to indicate the read-write data volume of the fourth virtual storage unit in a unit time, and the load description information comprises the load coefficient.

4. The method of claim 3, wherein, The load coefficient corresponding to the fourth virtual storage unit is obtained based on central processor load information, memory load information and read-write load information of the fourth virtual storage unit in the at least one storage unit, wherein the central processor load information is used to indicate the average utilization rate of the central processor of the fourth virtual storage unit in a unit time, the memory load information is used to indicate the remaining available memory of the fourth virtual storage unit, and the read-write load information is used to indicate the read-write data volume of the fourth virtual storage unit in a unit time, and the load description information comprises the load coefficient. The first load weight, the second load weight and the third load weight are obtained based on the central processor type corresponding to the fourth virtual storage unit, wherein the first load weight is used to indicate the importance of the central processor load information, the second load weight is used to indicate the importance of the memory load information, and the third load weight is used to indicate the importance of the read-write load information, and the sum of the first load weight, the second load weight and the third load weight satisfies a preset threshold condition; The first load weight, the second load weight and the third load weight are obtained based on the central processor type corresponding to the fourth virtual storage unit, wherein the first load weight is used to indicate the importance of the central processor load information, the second load weight is used to indicate the importance of the memory load information, and the third load weight is used to indicate the importance of the read-write load information, and the sum of the first load weight, the second load weight and the third load weight satisfies a preset threshold condition; 5. The method of claim 3, wherein, The load description information and the resource description information corresponding to at least one virtual storage unit matched with the target terminal are obtained, comprising: The energy efficiency coefficient corresponding to the fourth virtual storage unit is obtained based on central processor energy efficiency information and memory energy efficiency information of the fourth virtual storage unit, wherein the central processor energy efficiency information is used to indicate the processing capacity of the fourth virtual storage unit per unit central processor, and the memory energy efficiency information is used to indicate the utilization rate of the fourth virtual storage unit per unit memory resource, and the resource description information comprises the energy efficiency coefficient.

6. The method of claim 5, wherein, Before the energy efficiency coefficient corresponding to the fourth virtual storage unit is obtained based on the central processor energy efficiency information and the memory energy efficiency information of the fourth virtual storage unit, comprising: The product of the number of central processor cores allocated for the fourth virtual storage unit and a preset time is determined as a resource size value; The ratio of the number of tasks completed by the central processor of the fourth virtual storage unit in the preset time to the resource size value is determined as the central processor energy efficiency information; The product of the memory capacity allocated for the fourth virtual storage unit and the preset time is determined as a temporary storage size value; The ratio of the active data volume of the fourth virtual storage unit in the preset time to the temporary storage size value is determined as the memory energy efficiency information, wherein the active data volume is the actual storage occupied capacity of the data whose execution read-write operation times satisfy a times threshold condition.

7. The method of claim 5, wherein, The central processor energy efficiency information and the memory energy efficiency information of the fourth virtual storage unit are used to obtain an energy efficiency coefficient corresponding to the fourth virtual storage unit, including: The central processor type corresponding to the fourth virtual storage unit is used to obtain a first energy efficiency weight and a second energy efficiency weight, wherein the first energy efficiency weight is used to indicate the importance of the central processor energy efficiency information, and the second energy efficiency weight is used to indicate the importance of the memory energy efficiency information; The central processor energy efficiency information and the memory energy efficiency information are weighted and summed based on the first energy efficiency weight and the second energy efficiency weight to obtain the energy efficiency coefficient.

8. The method of claim 1, wherein, Before obtaining the load description information and the resource description information corresponding to each of the at least one virtual storage unit matched with the target terminal, the method further includes: A storage device with a remaining storage capacity satisfying a remaining capacity condition and located in a target space is determined as a candidate device, wherein the target space is a region with a device density satisfying a density threshold condition, and the device density is used to indicate the number of storage devices in a unit space; A master device and at least one slave device are determined from a currently traversed current candidate device and other candidate devices located in a neighboring space, wherein the current candidate device can communicate with other storage devices located in the neighboring space; A storage cluster corresponding to the current candidate device is constructed based on the master device and at least one slave device, wherein the storage cluster is used to determine at least one virtual storage unit.

9. The method of claim 8, wherein, Determining a master device and at least one slave device from a currently traversed current candidate device and other candidate devices located in a neighboring space includes at least one of the following: The master device is determined based on a candidate broadcast issued by the current candidate device and other candidate devices in the neighboring space, and other storage devices except the master device are determined as the slave devices, wherein the candidate broadcast is used to indicate the time of becoming the candidate device; The master device is determined based on the network bandwidth of the current candidate device and other candidate devices in the neighboring space, and other storage devices except the master device are determined as the slave devices; The master device is determined based on the failure rate of the current candidate device and other candidate devices in the neighboring space, and other storage devices except the master device are determined as the slave devices.

10. The method of claim 8, wherein, Before obtaining the load description information and the resource description information corresponding to each of the at least one virtual storage unit matched with the target terminal, the method further includes: A data transmission rate corresponding to each storage cluster of a fifth virtual storage unit in the at least one virtual storage unit is obtained, wherein the data transmission rate corresponding to the fifth virtual storage unit and the storage cluster is used to indicate the time required for the fifth virtual storage unit and the storage cluster to perform unit capacity data transmission; The storage cluster corresponding to the maximum data transmission rate is mapped with the fifth virtual storage unit.

11. The method according to any one of claims 1 to 10, characterized in that, The target storage unit is determined from all the storage unit groups based on the priority coefficient, including: According to the priority coefficient of each of the at least one virtual storage unit, all the storage unit groups are traversed to perform the following operations: In the case that there is a free storage unit in the current storage unit group, the free storage unit is determined as the target storage unit, wherein the free storage unit is a virtual storage unit not in a working state; In the case that there is no free storage unit in the current storage unit group, the next storage unit group is determined as the current storage unit group.

12. An electronic device, comprising: Comprise: a memory for storing a computer program; a processor for implementing the steps of the data storage method according to any one of claims 1 to 11 when executing the computer program.

Citation Information

Patent Citations

  • Memory collaborative management system and method

    CN115269450A

  • Compatible processing method and system, electronic equipment and computer readable medium

    CN118916315A